Disease inspection and detection data processing method, system, device, equipment and medium

By integrating the real-time data warehouse architecture with the Flink stream-batch integrated processing, the real-time performance and storage cost issues of the disease detection and testing data processing system are solved, high-concurrency data query and high-throughput analysis are achieved, and the emergency response capabilities of public health events are improved.

CN120600196APending Publication Date: 2025-09-05CENT FOR DISEASE CONTROL & PREVENTION OF THE NORTHERN THEATER COMMAND OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510667297.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing disease detection and testing data processing system has problems such as insufficient real-time performance, high storage cost, low data query efficiency, complex system and high maintenance cost, especially serious delays in public health event monitoring and emergency decision-making.

Method used

It adopts a lake-warehouse integrated real-time data warehouse architecture and combines it with Flink to achieve integrated stream and batch processing. Through a tiered storage strategy of hot, warm, and cold storage layers, it manages data based on data type and timeliness. It also uses Flink CDC to achieve real-time data capture and migration, supporting high-concurrency data queries and high-throughput analysis.

Benefits of technology

It enables efficient real-time query and analysis of data, reduces storage costs, simplifies system development and maintenance, and improves emergency response capabilities for public health incidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600196A_ABST
    Figure CN120600196A_ABST
Patent Text Reader

Abstract

The invention discloses a disease inspection and detection data processing method, system, device and equipment and a medium. The disease inspection and detection data processing method comprises the following steps: collecting original data from a plurality of disease data sources; processing the original data according to the data type of the original data to obtain target data; according to the data type of the target data and the timeliness of the target data, storing the target data to different storage layers; and under the condition that a data query request is received, processing the data of the corresponding storage layer based on the request type, and outputting a processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data, and in particular to a method, system, device, equipment and medium for processing disease detection data. Background Art

[0002] With the rapid development of medical information technology, the efficient management and application of disease detection and testing data has become increasingly critical. In the medical field, disease detection and testing data is an important basis for clinical diagnosis, disease research, and medical decision-making. How to better process this data to facilitate data analysis and decision-making by data users has become a research focus.

[0003] The data processing system proposed by the related technology lacks the real-time performance in collecting disease data, the cost of storing disease data is high, the unreasonable storage method leads to low data query efficiency, low data real-time query and analysis capabilities, and the system development is complex and the maintenance cost is high. Summary of the Invention

[0004] The embodiments of the present invention are intended to at least partially address one of the technical problems in the related art. To this end, one object of the present invention is to provide a disease detection data processing method, system, apparatus, device, and medium that enable real-time querying of high-concurrency data and analysis of high-throughput data.

[0005] An embodiment of the present invention provides a disease inspection and detection data processing method, which includes: collecting original data from multiple disease data sources; processing the original data according to the data type of the original data to obtain target data; storing the target data in different storage layers according to the data type of the target data and the timeliness of the target data; when receiving a data query request, processing the data of the corresponding storage layer based on the request type, and outputting the processing result.

[0006] Exemplarily, different storage layers include a hot storage layer, a warm storage layer, and a cold storage layer; according to the data type of the target data and the timeliness of the target data, the target data is stored in different storage layers, including: according to the data type of the target data, the target data is stored in the warm storage layer or the cold storage layer; according to the timeliness of the target data stored in the warm storage layer, the target data stored in the warm storage layer is synchronously stored in the hot storage layer.

[0007] Exemplarily, the data type of the target data includes a structured type and an unstructured type; according to the data type of the target data, the target data is stored in a warm storage layer or a cold storage layer, including: when the data type of the target data is a structured type, the target data is stored in the warm storage layer; when the data type of the target data is an unstructured type, the target data is stored in the cold storage layer.

[0008] Exemplarily, based on the timeliness of the target data stored in the warm storage layer, the target data stored in the warm storage layer is synchronously stored in the hot storage layer, including: when the timeliness of the target data stored in the warm storage layer meets the timeliness condition, the target data stored in the warm storage layer is standardized and synchronously stored in the hot storage layer.

[0009] Exemplarily, the request types include real-time request types and batch request types; when a data query request is received, the data of the corresponding storage layer is processed based on the request type, and the processing result is output, including: when the request type of the data query request is a real-time request type, the data of the hot storage layer is processed based on the stream processing mode, and the processing result is output; when the request type of the data query request is a batch request type, the data of the hot storage layer or the warm storage layer is processed based on the batch processing mode, and the processing result is output; when the request type of the data query request includes a real-time request type and a batch request type, the data of the hot storage layer or the warm storage layer is processed based on the stream processing mode and the batch processing mode, and the processing result is output.

[0010] Exemplarily, the data type of the original data includes at least one of a structured type and an unstructured type, and the processing method includes at least one of a cache processing, an extraction processing, and an encoding processing; according to the data type of the original data, the original data is processed to obtain the target data, including: when the data type of the original data is a structured type, the original data is cached to obtain structured data of the temporary message middleware as the target data; when the data type of the original data is an unstructured type, the original data is extracted and / or encoded to obtain the extracted structured data and / or the encoded unstructured data as the target data.

[0011] Exemplarily, the method further includes: migrating the target data of the storage layer based on at least one of timeliness data, access frequency, query request, and storage space of the target data.

[0012] Exemplarily, based on at least one of the timeliness, access frequency, query request, and storage space of the target data, the target data of the storage layer is migrated, including: for the target data of the hot storage layer, when the timeliness data of the target data is at least one of a first preset time, the access frequency is a first preset frequency, and the storage space is a first preset threshold, the target data of the hot storage layer is migrated to the warm storage layer; for the target data of the warm storage layer, when the timeliness data of the target data is at least one of a second preset time, the access frequency is a second preset frequency, and the storage space is a second preset threshold, the target data of the warm storage layer is migrated to the cold storage layer; for the target data of the cold storage layer, when the user submits a historical data query request, the target data of the cold storage layer is migrated to the hot storage layer or the warm storage layer based on the time range of the historical data query request.

[0013] Exemplarily, the original data collection method includes at least one of full acquisition, incremental acquisition, and real-time acquisition.

[0014] Another embodiment of the present invention provides a disease detection data processing system, which includes a data acquisition module, a batch-flow integrated module and a storage module. The disease detection data processing system is used to execute the steps of the method of any of the above embodiments.

[0015] Another embodiment of the present invention provides a disease inspection and detection data processing device, which includes: an acquisition module for collecting raw data from multiple disease data sources; a first processing module for processing the raw data according to the data type of the raw data to obtain target data; a storage module for storing the target data in different storage layers according to the data type of the target data and the timeliness of the target data; a second processing module for processing the data of the corresponding storage layer based on the request type when a data query request is received, and outputting the processing result.

[0016] An embodiment of the present invention provides an electronic device including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the method of any one of the above embodiments when executing the computer program.

[0017] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method of any one of the above embodiments are implemented.

[0018] In the above embodiment, the disease detection data processing method includes: collecting raw data from multiple disease data sources; processing the raw data according to the data type of the raw data to obtain target data; storing the target data in different storage layers according to the data type of the target data and the timeliness of the target data; and processing the data of the corresponding storage layer based on the request type when a data query request is received, and outputting the processing result. This method stores the data in different storage layers according to the target data type and timeliness, realizing the reasonable storage and management of data, which can not only meet the storage requirements of different data, but also improve the utilization efficiency of storage resources and reduce storage costs; by collecting raw data from multiple disease data sources, data silos are broken and data utilization is improved. When a data query request is received, the corresponding storage layer data is processed based on the request type, quickly responding to diverse data requirements, improving data query efficiency, and realizing real-time query of high-concurrency data and analysis of high-throughput data.

[0019] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A flow chart of a method for processing disease detection data provided by an embodiment of the present invention;

[0021] Figure 2 This is an architecture diagram of a disease detection data processing system provided by another embodiment of the present invention;

[0022] Figure 3 A schematic diagram of synchronization between the lake warehouse and the real-time data warehouse provided in an embodiment of the present invention;

[0023] Figure 4 A flow chart showing the integration of batch and stream processing provided by an embodiment of the present invention;

[0024] Figure 5 A logical diagram of storage layer target data migration provided by an embodiment of the present invention;

[0025] Figure 6 A block diagram of a disease detection data processing device provided in another embodiment of the present invention;

[0026] Figure 7 A block diagram of an electronic device provided in accordance with another embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0028] With the rapid development of medical information technology, the efficient management and application of disease detection and testing data has become increasingly critical. In the medical field, disease detection and testing data is an important basis for clinical diagnosis, disease research, and medical decision-making. How to better process this data to facilitate data analysis and decision-making by data users has become a research focus.

[0029] The disease detection and testing data processing system includes functions such as data integration, data processing, centralized storage, and data services. The processing system connects with the medical institution or testing and testing institution system to complete data collection. The collected raw data is processed in the data processing system and then stored centrally according to unified standards. It is then provided to the data demander through the service interface. The existing disease detection and testing data processing has the following defects: (1) Insufficient real-time performance: Traditional ETL (Extract-Transform-Load) cannot capture incremental data changes of testing equipment or LIS system (Laboratory Information System) in real time, resulting in delays in public health event monitoring. (2) High storage cost: Unstructured data (such as CT images, gene sequences) takes up a lot of storage space and lacks hierarchical management of hot and cold data. (3) Lake-warehouse separation: The data lake (raw data storage) is separated from the data warehouse (analysis optimization), resulting in low cross-modal query efficiency. (4) Weak real-time analysis: Traditional data warehouses (such as Hive) cannot support low-latency OLAP (Online Analytical Processing) queries, affecting the efficiency of emergency decision-making. (5) Batch-stream separation: Stream processing (real-time alarms) and batch processing (statistical models) need to be developed independently, resulting in a complex system and high maintenance costs.

[0030] Related technologies use Flink for stream processing, but lack deep integration of CDC (Change Data Capture) collection and lake warehouse integration. Real-time data warehouses often rely on proprietary commercial solutions (such as Snowflake), which are costly and have limited scalability. Lake warehouse solutions based on Hudi and Iceberg face performance bottlenecks in real-time updates and transaction support.

[0031] In view of this, the implementation method of this application proposes a disease inspection and detection data processing method, which adopts an advanced lake-warehouse integrated real-time data warehouse architecture, and has high-concurrency data real-time query and high-throughput data analysis capabilities, and uses an open technology stack to achieve good ecological support and compatibility.

[0032] Figure 1 This is a flow chart of the disease detection data processing method provided in an embodiment of the present invention.

[0033] like Figure 1 As shown, the disease detection data processing method 100 includes steps S110 to S140.

[0034] Step S110: collecting raw data from multiple disease data sources.

[0035] Exemplarily, the raw data includes structured data from business databases (such as MySQL / Oracle) (such as nucleic acid testing data from various regions), third-party APIs (Application Programming Interface), inspection equipment log files, and other data sources (Simple Storage Service, object storage service), HDFS (Hadoop Distributed File System), and the collection method can be real-time collection, incremental collection, full collection and other collection methods.

[0036] Step S120 , processing the original data according to the data type of the original data to obtain target data.

[0037] Exemplarily, the data types of raw data include structured data (data tables) and unstructured data (medical images, test equipment log files, CSV (Comma-Separated Values) test reports), and the processing of raw data includes caching processing, extraction processing, and encoding processing.

[0038] Step S130 : storing the target data in different storage layers according to the data type and timeliness of the target data.

[0039] Exemplarily, the data types of target data include structured types and unstructured types, where structured data includes target data obtained by caching the original data through the message middleware and target data obtained by extracting and processing the unstructured data (such as file name, number, date, storage address, and file ownership). Unstructured data includes target data obtained by encoding the original data (such as data obtained by encoding medical images). The storage layer includes a hot storage layer, a warm storage layer, and a cold storage layer. The timeliness of the target data is, for example, three months, one year, etc.

[0040] Step S140: When a data query request is received, the data of the corresponding storage layer is processed based on the request type, and the processing result is output.

[0041] For example, data query requests include real-time requests and batch requests. Real-time requests, for example, include displaying a large screen showing the status of public health events (such as the positivity rate) in real time. Batch requests can aggregate and generate hourly load reports for testing agencies or analyze disease data from the past year. Data processing at the corresponding storage layer involves processing stream and batch data based on the Flink streaming and batch integration platform, and outputting the calculation results to the requesting end.

[0042] According to the embodiments of the present application, data is stored in different storage layers based on the target data type and timeliness, so as to realize reasonable storage and management of data, which can not only meet the storage requirements of different data, but also improve the utilization efficiency of storage resources and reduce storage costs; by collecting the original data from multiple disease data sources, breaking the data silos, improving data utilization, and when receiving a data query request, processing the corresponding storage layer data based on the request type, quickly responding to diversified data requirements, improving data query efficiency, and realizing real-time query of high-concurrency data and analysis of high-throughput data.

[0043] Figure 2 This is an architecture diagram of a disease detection data processing system provided by another embodiment of the present invention.

[0044] The disease inspection and detection data processing system includes a data acquisition module, a batch-flow integration module and a storage module.

[0045] like Figure 2 As shown, the disease detection data processing system includes a data acquisition module ( Figure 2 Data source layer and collection layer in the batch-stream integration module ( Figure 2 The computing layer in the ) and the storage module ( Figure 2The disease processing system also provides auxiliary modules for metadata management and monitoring alarms.

[0046] Specifically, the data collection module uses Flink CDC (Change Data Capture) and Java programming to support the collection of multiple data sources, including database systems (MySQL / Oracle, etc.), inspection equipment logs, and third-party APIs (Application Programming Interfaces). The collection module supports real-time, incremental, and full collection methods, and can be used according to data source type and business needs.

[0047] Hive Metastore is used for unified metadata management, including Paimon, Doris, MinIO, and Flink. Prometheus is used for system monitoring and alerting, with monitoring indicators including Flink Metrics (task throughput), Doris Query (response time / QPS), and MinIO storage health (number of objects / capacity).

[0048] The batch-stream integration module uses Apache Flink to implement batch-stream fusion processing. Stream processing is used to calculate positivity rates and device anomaly alerts in real time, with the results written to Kafka or pushed to the dashboard. Batch processing is used to trigger FlinkBatch jobs daily to generate regional public health event trend reports. Flink SQL is used to unify the writing of stream and batch tasks, reducing code redundancy.

[0049] The storage modules are layered and organized into a tiered storage system. The hot storage layer (Doris, a real-time data warehouse) stores near-real-time data (e.g., the past three months) and stores it in the real-time data warehouse for real-time queries and analysis. A real-time analytics engine built on Apache Doris synchronizes hot data from Paimon to Doris in real time, providing sub-second OLAP (Online Analytical Processing) response times. The warm storage layer (Paimon, a lake-warehouse-integrated system) stores mild data (e.g., mild data from the past year) and stores it in Paimon tables. This supports frequent updates and OLAP queries. Paimon's Schema Evolution feature automatically generates structured tables, supporting ACID (Atomicity-Consistency-Isolation-Durability) transactions and efficient queries. The cold storage layer (MinIO, an object storage system) stores historical data (e.g., data older than one year). This data is compressed using Zstandard and archived in MinIO object storage, providing cost-effective, highly available storage services with S3 protocol compatibility.

[0050] Different storage layers include hot storage layer, warm storage layer and cold storage layer; target data is stored in different storage layers according to the data type and timeliness of the target data, including: storing the target data in the warm storage layer or the cold storage layer according to the data type of the target data; synchronously storing the target data stored in the warm storage layer to the hot storage layer according to the timeliness of the target data stored in the warm storage layer.

[0051] For example, Figure 2 As shown, the storage layer includes a hot storage layer (real-time data warehouse Doris), a warm storage layer (lake-warehouse integrated Paimon), and a cold storage layer (object storage MinIO). The target data types include structured and unstructured data, and the timeliness of the target data can be three months or one year. Synchronization uses Flink CDC (Change Data Capture) provided by the Flink engine, which can capture change data in the business database or API (Application Programming Interface) in real time. The captured data is written to the Paimon table (warm storage layer). The data stored in Paimon is standardized and stored in Doris (hot storage layer) to complete data synchronization.

[0052] The data type of the target data includes a structured type and an unstructured type; according to the data type of the target data, the target data is stored in a warm storage layer or a cold storage layer, including: when the data type of the target data is a structured type, the target data is stored in the warm storage layer; when the data type of the target data is an unstructured type, the target data is stored in the cold storage layer.

[0053] For example, structured data includes data tables obtained by processing raw data provided by a relational database, and unstructured data includes inspection equipment log files, medical imaging images, etc. obtained by processing data imported from other data sources.

[0054] In the above embodiment, by adopting the object storage MinIO and optimizing the data storage model by performing hot and cold tiering of data, the long-term storage cost of massive data (especially images and logs) is reduced, and a storage service with high cost performance and high availability is provided.

[0055] Figure 3 A schematic diagram of the synchronization between the lake warehouse and the real-time data warehouse provided in an embodiment of the present invention.

[0056] According to the timeliness of the target data stored in the warm storage layer, the target data stored in the warm storage layer is synchronously stored in the hot storage layer, including: when the timeliness of the target data stored in the warm storage layer meets the timeliness condition, the target data stored in the warm storage layer is standardized and synchronously stored in the hot storage layer.

[0057] For example, Figure 3 As shown in the figure, when new data enters the lake and warehouse, the newly collected raw data will first be saved in the lake-warehouse integrated Paimon, and then the standard detailed data after ETL (Extract-Transform-Load) through Flink CDC will be synchronized to the real-time data warehouse Doris.

[0058] Specifically, data synchronization from Paimon to Doris is implemented using Flink CDC Connector + Doris StreamLoad. Only aggregation results that require high-frequency access are synchronized, and the synchronization delay is less than 1 minute.

[0059] In another example, the real-time data warehouse Doris and the integrated lake warehouse (Paimon) also support reverse writeback: reverse writeback of data from Doris to Paimon is used in scenarios such as data correction, dimension completion, or re-running historical tasks. The writeback task is manually triggered on demand, and Doris exports data files and loads them in batches into the Paimon table.

[0060] Regarding the query division of labor strategy: the Paimon lake warehouse is responsible for complex analysis scenarios such as full scans, time travel, schema changes, and machine learning; the Doris data warehouse is responsible for scenarios such as sub-second interactive queries, real-time large screens, ad hoc analysis, and low-latency API responses.

[0061] In the above embodiment, based on the linkage between the lake warehouse and the real-time data warehouse, high-frequency analysis (such as real-time large screen) is achieved by directly querying Doris, and deep analysis (such as associating historical images with detection results) is achieved by performing joint queries across Paimon and MinIO through the Trino / Doris external table.

[0062] Figure 4 This is a batch-stream integration flowchart provided by an embodiment of the present invention.

[0063] Request types include real-time request type and batch request type. When a data query request is received, the data in the corresponding storage layer is processed based on the request type and the processing results are output, including:

[0064] When the data query request type is a real-time request type, the data of the hot storage layer is processed based on the stream processing mode, and the processing result is output;

[0065] When the request type of the data query request is a batch request type, processing the data of the hot storage layer or the warm storage layer based on the batch processing mode and outputting the processing result;

[0066] When the request type of the data query request includes a real-time request type and a batch request type, data of the hot storage layer or the warm storage layer is processed based on the stream processing mode and the batch processing mode, and a processing result is output.

[0067] Specifically, batch-stream fusion processing is implemented based on Apache Flink, which uses a unified architecture to process both stream and batch data. Stream processing: Using Paimon stream tables as the data source (synchronized from Paimon to Doris), Flink performs real-time computations on data from the real-time data warehouse (hot storage layer), counting the number of tests in each region in real time and triggering threshold alerts (e.g., a hospital's positive rate exceeding 5%). For example, real-time calculations of positive rates and device anomaly alerts are performed, with the results written to Kafka or pushed to a dashboard.

[0068] For example, batch processing involves triggering a Flink Batch job daily to generate a regional public health event trend report. Flink then batch processes data based on the hot storage layer. If a user requests analysis of six months of data, Flink batch processes data based on both the hot and warm storage layers. Flink analyzes the first three months of data based on the hot storage layer, and the last three months of data based on the warm storage layer. Finally, the combined processing results are returned to the user. Table 1 shows the key data flow technology implementation of the disease detection data processing system proposed in this invention, as shown in Table 1:

[0069] Table 1

[0070]

[0071] This system implements batch and stream integration through Flink, uses the same SQL to process streams / batches, checkpoint fault tolerance (stream), and dynamic scaling (batch). It implements Schema Evolution, Time Travel Query, Merge-On-Read, and a unified stream and batch storage interface based on the Paimon lake warehouse capability. It achieves sub-second response based on Doris real-time analysis, a vectorized execution engine, and materialized view acceleration. It also implements object version control, automatic dumping to low-cost storage media, and cross-region replication based on the MinIO cold storage strategy.

[0072] In the above embodiment, Flink is used to unify the APIs for stream processing and batch processing, simplify the development process, enable high-concurrency real-time data query and high-throughput data analysis, and flexibly allocate resources between batch processing tasks and stream processing tasks, thus avoiding resource waste and improving overall resource utilization.

[0073] The data type of the original data includes at least one of a structured type and an unstructured type, and the processing method includes at least one of a cache processing, an extraction processing and an encoding processing; according to the data type of the original data, the original data is processed to obtain the target data, including: when the data type of the original data is a structured type, the original data is cached to obtain structured data of the temporary message middleware as the target data; when the data type of the original data is an unstructured type, the original data is extracted and / or encoded to obtain the extracted structured data and / or the encoded unstructured data as the target data.

[0074] Specifically, the data collection module uses Flink CDC and Java programming to support the collection of multiple data sources (raw data), including database systems (MySQL / Oracle, etc.), inspection equipment logs, and third-party APIs. Structured data is sent to the message middleware (Kafka / Pulsar) transmission engine for subsequent storage processing. Unstructured data (such as text, log files, images, and videos) is processed through attribute extraction to obtain structured data information such as file name, number, date, storage address, and file ownership. Unstructured data itself (such as text, log files, images, and videos) is encoded and stored in MinIO object storage.

[0075] Figure 5 This is a logical diagram of storage layer target data migration provided by an embodiment of the present invention.

[0076] The method further includes: migrating the target data of the storage layer based on at least one of the timeliness data, access frequency, query request, and storage space of the target data.

[0077] For example, storage layers are marked according to data timestamps and access frequency, and data migration policies are set. Data is automatically migrated between different layers of storage according to the migration policies. For example, hot data is cooled and migrated from Doris to Paimon, and warm data is cooled and migrated from Paimon to object storage MinIO.

[0078] Exemplarily, migrating target data in a storage layer based on at least one of timeliness, access frequency, query request, and storage space of the target data includes:

[0079] For target data in the hot storage layer, when the timeliness data of the target data is at least one of a first preset time, an access frequency is at least one of a first preset frequency, and a storage space is at least one of a first preset threshold, the target data in the hot storage layer is migrated to the warm storage layer;

[0080] For target data in the warm storage layer, when the timeliness of the target data is at least one of a second preset time, an access frequency is at least one of a second preset frequency, and a storage space is at least one of a second preset threshold, the target data in the warm storage layer is migrated to the cold storage layer;

[0081] For the target data in the cold storage layer, when the user submits a historical data query request, the target data in the cold storage layer is migrated to the hot storage layer or the warm storage layer based on the time range of the historical data query request.

[0082] Core attributes and division of labor of each layer: Hot layer, Doris+SSD, stores frequently accessed hot data, and is used in real-time OLAP query and interactive analysis scenarios; warm layer, Paimon+HDD / cloud storage, stores medium and low-frequency accessed warm data, and is used in batch ETL and streaming data merging scenarios; cold layer, uses HDD disk to privately deploy MinIO cluster, archives and stores infrequently accessed cold data, and is used in compliance storage, disaster backup and other scenarios.

[0083] The target data migration for the storage layer involves migrating from the hot layer to the warm layer (automatically cooling down). Trigger conditions include the Doris table partition access frequency being less than a threshold (e.g., no queries for one day) (the first preset frequency), Doris storage space pressure being greater than a threshold (the first preset threshold), and data TTL expiration (e.g., 30-day retention) (the first preset time). Technical implementation: Doris data is exported and written to the Paimon master table via a Flink job.

[0084] Warm tier to cold tier (scheduled archiving). Trigger conditions include the Paimon partition being created > 360 days ago (the second preset time), the storage cost optimization policy being triggered (the second preset threshold), and data access volume approaching zero (the second preset frequency). Technical Implementation: Data is copied to the Paimon bucket via S3 DistCP.

[0085] Cold to warm / hot tier migration (on-demand): Trigger conditions include user requests for historical data queries, data recalculation requirements (such as model training), and data audit compliance requirements. Technical implementation: Flink batch jobs are used to extract data from MinIO and load it into Paimon or Doris.

[0086] In the above embodiment, migrating data based on storage space can avoid wasting storage space, make more rational use of storage resources, and improve the overall efficiency of storage devices. By analyzing access frequency and query requests and migrating frequently accessed and frequently queried data to high-speed storage devices or more optimal storage locations, data access latency can be reduced and data read speeds can be accelerated, thereby improving the overall data processing and response performance of the system and providing more efficient data support for business applications.

[0087] In another example, the original data is collected in at least one of full data collection, incremental data collection, and real-time data collection.

[0088] Full data acquisition involves collecting all data from a data source into a target storage or processing system at once, regardless of how the data has changed. For example, copying all records in a database table to a data warehouse, regardless of whether the data is newly created or has existed for a long time, is a common approach.

[0089] Incremental acquisition only retrieves data that has changed since the last acquisition, including new data, modified data, and deleted data records (usually requiring an additional mechanism to mark deletion operations). For example, the database records the timestamp of each data modification. By comparing timestamps, you can retrieve data that has changed since the last acquisition time.

[0090] Real-time acquisition means collecting and processing data immediately, with virtually no data delay. For example, real-time data from sources like sensors and log files can be received in real time via a message queue (such as Kafka). As new data is generated, it is immediately transferred to the processing system for analysis.

[0091] In the above embodiments, raw data is collected through a variety of data collection methods, making data collection more flexible, the usage scenarios more comprehensive, and able to adapt to diverse business needs, thereby achieving the unity of data integrity and timeliness.

[0092] Compared with related technologies, the above-mentioned disease detection data processing method has the following technical effects: (1) Real-time breakthrough: CDC data synchronization delay is <500ms, and Doris real-time data warehouse query response time is <1 second. The alarm delay of public health event detection results (such as positive results) is <800ms, which meets the emergency response needs of public health events. (2) Storage cost reduction: Hot and cold tiering and MinIO object storage reduce long-term storage costs by more than 55%. (3) Analysis efficiency improvement: Under the Paimon lake warehouse + Doris real-time data warehouse architecture, the mixed load query performance is improved by 65%. (4) Development efficiency optimization: Flink's batch and stream integrated API reduces the amount of code by 40%, and the operation and maintenance complexity is reduced. (5) Scalability enhancement: Supports data writing of tens of thousands per second, and dynamic expansion to meet the demand for tens of millions of nucleic acid tests per day.

[0093] The present invention proposes alternative solutions for the disease detection data processing system: (1) CDC collection alternative solution: Debezium+Kafka is used to replace Flink CDC, but additional message queue maintenance is required, which increases the system complexity. A timed polling database is used, sacrificing real-time performance in exchange for compatibility. (2) Lake warehouse integrated alternative solution: Apache Hudi or Iceberg is used to replace Paimon, but Hudi has weak support for stream processing and Iceberg has low real-time update efficiency. The traditional HDFS+Hive solution is used, which lacks transaction support and Schema Evolution capabilities. (3) Real-time data warehouse alternative solution: ClickHouse or StarRocks is used to replace Doris, but ClickHouse lacks strong consistency transaction support and StarRocks has weak ecological compatibility. Commercial solutions such as Snowflake are used, which are costly and difficult to deploy privately. (4) Object storage alternative solution: AWSS3 or Alibaba Cloud OSS is used to replace MinIO, but it relies on public cloud services and has limited data sovereignty. Using Ceph to build private object storage significantly increases the complexity of operation and maintenance. (5) Alternative solution for batch and stream processing: Use Spark Structured Streaming + Spark Batch, but its real-time performance is weaker than Flink, and the two APIs differ significantly. Using independent stream processing (such as KafkaStreams) and batch processing (such as Hive) systems requires maintaining two computing clusters.

[0094] The disease inspection and detection data processing method and system proposed in the present invention are designed with an advanced lake-warehouse integrated real-time data warehouse architecture. The system adopts an open technology stack to achieve good ecological support and compatibility, and solves the real-time incremental capture of multi-source data such as LIS systems and testing equipment, as well as data standard conversion problems; by adopting object storage and optimizing the data storage model by hot and cold stratification of data, the long-term storage cost of massive data (especially images and logs) is reduced; based on Paimon, the lake-warehouse integration is realized, and combined with the Doris real-time data warehouse to support mixed load analysis, it has high-concurrency real-time query and high-throughput data analysis capabilities; through Flink, the API unification of stream processing and batch processing is realized to simplify the development process, and the system supports container platform deployment. With the help of the elastic scaling capabilities of the container platform, the system can be dynamically expanded and reduced to respond to the data peaks of public health emergencies (such as large-scale nucleic acid testing).

[0095] Figure 6 This is a block diagram of a disease detection data processing device provided in another embodiment of the present invention.

[0096] The embodiment of the present invention provides a disease detection data processing device 600, see Figure 6The disease detection data processing device 600 includes: an acquisition module 610 , a first processing module 620 , a storage module 630 , and a second processing module 640 .

[0097] Exemplarily, the acquisition module 610 is used to acquire raw data from multiple disease data sources.

[0098] Exemplarily, the first processing module 620 is configured to process the original data according to the data type of the original data to obtain target data.

[0099] Exemplarily, the storage module 630 is configured to store the target data in different storage layers according to the data type and timeliness of the target data.

[0100] Exemplarily, the second processing module 640 is configured to process data in the corresponding storage layer based on the request type when receiving a data query request, and output a processing result.

[0101] Exemplarily, different storage layers include a hot storage layer, a warm storage layer, and a cold storage layer; the storage module 630 is also used to store the target data in the warm storage layer or the cold storage layer according to the data type of the target data; and synchronize the target data stored in the warm storage layer to the hot storage layer according to the timeliness of the target data stored in the warm storage layer.

[0102] Exemplarily, the data type of the target data includes a structured type and an unstructured type; the storage module 630 is also used to store the target data in a warm storage layer when the data type of the target data is a structured type; and to store the target data in a cold storage layer when the data type of the target data is an unstructured type.

[0103] Exemplarily, the storage module 630 is further configured to normalize the target data stored in the warm storage layer and synchronously store the target data in the hot storage layer when the timeliness of the target data stored in the warm storage layer meets the timeliness condition.

[0104] Exemplarily, the request types include real-time request types and batch request types; the second processing module 640 is also used to, when the request type of the data query request is a real-time request type, process the data of the hot storage layer based on the stream processing mode and output the processing result; when the request type of the data query request is a batch request type, process the data of the hot storage layer or the warm storage layer based on the batch processing mode and output the processing result; when the request type of the data query request includes a real-time request type and a batch request type, process the data of the hot storage layer or the warm storage layer based on the stream processing mode and the batch processing mode, and output the processing result.

[0105] Exemplarily, the data type of the original data includes at least one of a structured type and an unstructured type, and the processing method includes at least one of a cache processing, an extraction processing, and an encoding processing; the first processing module 620 is also used to perform cache processing on the original data when the data type of the original data is a structured type, and obtain structured data of the temporary message middleware as the target data; when the data type of the original data is an unstructured type, perform extraction processing and / or encoding processing on the original data to obtain extracted structured data and / or encoded unstructured data as the target data.

[0106] Exemplarily, the disease detection data processing device 600 further includes: migrating the target data in the storage layer based on at least one of the timeliness data, access frequency, query request, and storage space of the target data.

[0107] Exemplarily, based on at least one of the timeliness, access frequency, query request, and storage space of the target data, the target data of the storage layer is migrated, including: for the target data of the hot storage layer, when the timeliness data of the target data is at least one of a first preset time, the access frequency is a first preset frequency, and the storage space is a first preset threshold, the target data of the hot storage layer is migrated to the warm storage layer; for the target data of the warm storage layer, when the timeliness data of the target data is at least one of a second preset time, the access frequency is a second preset frequency, and the storage space is a second preset threshold, the target data of the warm storage layer is migrated to the cold storage layer; for the target data of the cold storage layer, when the user submits a historical data query request, the target data of the cold storage layer is migrated to the hot storage layer or the warm storage layer based on the time range of the historical data query request.

[0108] Exemplarily, the original data collection method includes at least one of full acquisition, incremental acquisition, and real-time acquisition.

[0109] It can be understood that for the specific description of the disease detection data processing device 600, reference can be made to the above description of the disease detection data processing method, which will not be repeated here.

[0110] Figure 7 A block diagram of an electronic device provided in accordance with another embodiment of the present invention.

[0111] An embodiment of the present application provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above method when executing the computer program.

[0112] like Figure 7 As shown, for ease of understanding, the embodiment of the present application shows a specific electronic device 700.

[0113] The electronic device 700 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0114] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0115] Multiple components in the electronic device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0116] The computing unit 701 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods described above. For example, in some embodiments, any one or more of the above-described methods can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of any one or more of the various methods described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform any one or more of the above-described methods by any other appropriate means (e.g., by means of firmware).

[0117] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method of any one of the above embodiments are implemented.

[0118] It should be noted that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such instruction execution system, apparatus, or device. For purposes of the present invention, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0119] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0120] In the description of the present invention, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In the present invention, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0121] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0122] In addition, the terms "first" and "second" used in the embodiments of the present invention are only used for descriptive purposes and should not be understood as indicating or implying relative importance, or implicitly indicating the number of technical features indicated in this embodiment. Therefore, the features defined by the terms "first" and "second" in the embodiments of the present invention can explicitly or implicitly indicate that the embodiment includes at least one of such features. In the description of the present invention, the word "plurality" means at least two or two or more, such as two, three, four, etc., unless otherwise clearly and specifically defined in the embodiments.

[0123] In the present invention, unless otherwise clearly specified or limited in the embodiments, the terms "installed," "connected," "connect," and "fixed" appearing in the embodiments should be understood in a broad sense. For example, the connection may be a fixed connection, a detachable connection, or an integral connection. It can also be a mechanical connection, an electrical connection, etc.; of course, it can also be a direct connection, or an indirect connection through an intermediate medium, or it can be internal communication between two elements, or an interaction between two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood based on the specific implementation.

[0124] In the present invention, unless otherwise expressly specified or limited, when a first feature is "above" or "below" a second feature, it may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediary. Furthermore, when a first feature is "above," "above," or "above" a second feature, it may mean that the first feature is directly above or diagonally above the second feature, or simply means that the first feature is at a higher level than the second feature. When a first feature is "below," "below," or "below" a second feature, it may mean that the first feature is directly below or diagonally below the second feature, or simply means that the first feature is at a lower level than the second feature.

[0125] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A disease detection data processing method, characterized in that: The method comprises: Collect raw data from multiple disease data sources; Processing the original data according to the data type of the original data to obtain target data; Storing the target data in different storage layers according to the data type of the target data and the timeliness of the target data; When a data query request is received, the data in the corresponding storage layer is processed based on the request type and the processing result is output.

2. The method according to claim 1, characterized in that The different storage layers include a hot storage layer, a warm storage layer, and a cold storage layer; storing the target data in different storage layers according to the data type of the target data and the timeliness of the target data includes: storing the target data in the warm storage layer or the cold storage layer according to a data type of the target data; The target data stored in the warm storage layer is synchronously stored in the hot storage layer according to the timeliness of the target data stored in the warm storage layer.

3. The method according to claim 2, characterized in that The data type of the target data includes a structured type and an unstructured type; and storing the target data in the warm storage layer or the cold storage layer according to the data type of the target data includes: When the data type of the target data is a structured type, storing the target data in a warm storage layer; In a case where the data type of the target data is an unstructured type, the target data is stored in a cold storage layer.

4. The method according to claim 2, characterized in that The step of synchronously storing the target data stored in the warm storage layer to the hot storage layer according to the timeliness of the target data stored in the warm storage layer includes: When the timeliness of the target data stored in the warm storage layer meets the timeliness condition, the target data stored in the warm storage layer is standardized and synchronously stored in the hot storage layer.

5. The method according to claim 1, wherein The request type includes a real-time request type and a batch request type; When a data query request is received, the data of the corresponding storage layer is processed based on the request type and the processing result is output, including: When the request type of the data query request is a real-time request type, processing the data of the hot storage layer based on a stream processing mode and outputting the processing result; When the request type of the data query request is a batch request type, processing the data of the hot storage layer or the warm storage layer based on a batch processing mode, and outputting a processing result; In a case where the request type of the data query request includes a real-time request type and a batch request type, the data of the hot storage layer or the warm storage layer is processed based on the stream processing mode and the batch processing mode, and the processing result is output.

6. The method according to claim 1, wherein The data type of the original data includes at least one of a structured type and an unstructured type, and the processing method includes at least one of a cache process, an extraction process, and an encoding process; and the processing of the original data according to the data type of the original data to obtain the target data includes: In the case where the data type of the original data is a structured type, caching the original data to obtain structured data of a temporary message middleware as the target data; In the case where the data type of the original data is an unstructured type, the original data is subjected to extraction processing and / or encoding processing to obtain extracted structured data and / or encoded unstructured data as the target data.

7. The method according to any one of claims 2 to 4, characterized in that The method further comprises: The target data of the storage layer is migrated based on at least one of timeliness data, access frequency, query request, and storage space of the target data.

8. The method according to claim 7, characterized in that The migrating target data of the storage layer based on at least one of timeliness, access frequency, query request, and storage space of the target data includes: For target data in the hot storage layer, when the timeliness data of the target data is at least one of a first preset time, an access frequency is a first preset frequency, and a storage space is a first preset threshold, the target data in the hot storage layer is migrated to the warm storage layer; For target data in the warm storage layer, when the timeliness of the target data is at least one of a second preset time, a second preset access frequency, and a second preset storage space, the target data in the warm storage layer is migrated to the cold storage layer; For the target data in the cold storage layer, when a user submits a historical data query request, the target data in the cold storage layer is migrated to the hot storage layer or the warm storage layer based on the time range of the historical data query request.

9. The method according to claim 1, characterized in that The original data is collected in a manner including at least one of full data collection, incremental data collection, and real-time data collection.

10. A disease detection data processing system, characterized in that: The disease detection data processing system includes a data acquisition module, a batch-flow integration module and a storage module, and the disease detection data processing system is used to execute the method according to any one of claims 1-9.

11. A disease detection data processing device, characterized in that: The device comprises: An acquisition module, used to collect raw data from multiple disease data sources; A first processing module, configured to process the original data according to the data type of the original data to obtain target data; A storage module, configured to store the target data in different storage layers according to the data type of the target data and the timeliness of the target data; The second processing module is used to process the data of the corresponding storage layer based on the request type when receiving a data query request, and output the processing result.

12. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Gene data storage method and device based on mixed medium and storage medium

    CN121191603A

  • Hybrid medium-based gene data storage method, device and storage medium

    CN121191603B