Multi-source heterogeneous medical data aggregation and federal query method

By building Kafka cluster and Apache Doris database, combining ETL servers to clean and standardize multi-source medical data, the complexity and cost problems of integration and query of multi-source heterogeneous medical data in the existing technology are solved, and efficient and flexible data processing and analysis are achieved.

CN120030046APending Publication Date: 2025-05-23THE FIRST AFFILIATED HOSPITAL OF SOOCHOW UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510055701.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently integrate and query multi-source heterogeneous medical data, resulting in complex data processing, high cost, and difficult to achieve fast and accurate high-quality data views.

Method used

By building a Kafka cluster and Apache Doris database managed by Zookeeper, combining ETL servers for data cleaning, format standardization and structure processing, the aggregation of multi-source medical data and federal query are realized.

Benefits of technology

It greatly reduces the custom development workload and maintenance costs required for data integration, improves the flexibility and scalability of data processing, and realizes fast and accurate high-quality data query and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030046A_ABST
    Figure CN120030046A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source heterogeneous medical data aggregation and federal query method, which comprises the following steps of: establishing a Kafka cluster, and configuring a Broker; the method comprises the following steps: deploying an Apache Dores cluster by using Docker, configuring front-end and rear-end nodes, and appointing an IP address; configuring a Kafka producer, and sending the Kafka producer to the theme; deploying an ETL server as a Kafka consumer, adding the Kafka consumer into the same consumption group, and configuring a multi-thread processing message; data processing is realized in the ETL server; using an NLP model to carry out tagging processing and error identification on patient records; storing the processed data in a database, generating a unique primary key and adopting a partitioning strategy; and deploying a Web server and providing a patient data query interface. Multi-source heterogeneous medical data are accessed and processed in a unified mode, data originally dispersed in different systems are subjected to message transfer and buffering through Kafka, and flowing and seamless integration of the data are effectively achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical information technology, and in particular to a multi-source heterogeneous medical data aggregation and federation query method. Background Art

[0002] In the current advancement of medical informatization, different medical institutions and departments often establish multiple independent information systems, including HIS (Hospital Information System), PACS (Picture Archiving and Communication System), EMR (Electronic Medical Record System), LIS (Laboratory Information System), and others. Due to the lack of unified standards and specifications at the outset of each system's design, their underlying architectures, data models, storage formats, and access interfaces vary significantly, resulting in a high degree of heterogeneity in medical data at the source. Existing technologies typically integrate multi-source data through separate docking, individual conversions, and manual scripting. These approaches are not only labor-intensive and costly, but also require ongoing maintenance and updates. In particular, when medical institutions seek to conduct joint analysis and decision support on cross-system data, they must extract data from multiple systems, then rely on manual or additional middleware to clean and standardize the data format before completing secondary analysis in a unified data warehouse or data lake.

[0003] This process, due to the multiple data sources, complex data structures, and varying data quality, results in complex and lengthy steps such as data cleaning, deduplication, missing value filling, and data standardization, making it difficult to quickly and accurately obtain a high-quality, unified data view. Existing solutions lack a unified architecture and efficient integration methods for data acquisition and processing. Faced with incompatible data formats and interfaces across different systems, developers must continuously write separate docking logic and data conversion rules for each source, severely impacting the flexibility and scalability of data processing.

[0004] During the data query and analysis phase, users often need to jump between multiple systems to individually retrieve, compare, and integrate data from different sources. This not only significantly reduces data utilization efficiency, but also makes it difficult to obtain comprehensive, global analysis results in a timely manner, thereby affecting the decision-making efficiency and quality of medical services. At the same time, with the continuous expansion of medical data scale and the continuous increase in data types, traditional single-point query and centralized processing methods also face huge challenges in performance and scalability, and are unable to meet the requirements of fast acquisition, flexible query, and high concurrent access. Because of this, existing technologies urgently need a method that can more efficiently aggregate, standardize, and uniformly query multi-source heterogeneous medical data to reduce the difficulty and cost of data processing and improve the overall efficiency of data analysis and decision-making. Summary of the Invention

[0005] To overcome the shortcomings of the existing technology, the present invention proposes a multi-source heterogeneous medical data aggregation and federated query method. Through the ETL server, the data is cleaned, formatted and structured, and the field names, data formats and time standards between different systems are unified, thereby significantly reducing the custom development workload and maintenance costs required for data integration.

[0006] To achieve the above objectives, the present invention provides a multi-source heterogeneous medical data aggregation and federated query method, comprising the following steps:

[0007] Step 1: Build a Kafka cluster managed by Zookeeper and configure multiple Brokers for high availability.

[0008] Step 2: Use Docker to deploy the Apache Doris cluster, configure the front-end (FE) and back-end (BE) nodes and specify the IP addresses;

[0009] Step 3: Configure the Kafka producer to ensure that each message contains a unique messageId field and is sent to the specified topic;

[0010] Step 4: Deploy the ETL server as a Kafka consumer, join the same consumer group, and configure multi-threaded message processing;

[0011] Step 5: Implement data processing in the ETL server, including missing value identification, duplicate detection, format standardization, and text cleaning;

[0012] Step 6: Use NLP models to label patient records and identify errors.

[0013] Step 7: Store the processed data into the Apache Doris database, generate a unique primary key and adopt a partitioning strategy;

[0014] Step 8: Deploy a web server to provide a patient data query interface and ensure the security and performance of the interface.

[0015] Furthermore, step 1 is as follows:

[0016] Step 1.1: Download and decompress the stable version of Zookeeper, configure the zoo.cfg file, set the data directory, client port (default 2181), and cluster node information (if it is a cluster deployment), then start the Zookeeper service and ensure it is running normally and listening on the specified port;

[0017] Step 1.2: Download the appropriate version of the Kafka binary package from the Apache Kafka official website and unzip it to the target installation directory;

[0018] Step 1.3: Create a separate configuration file for each broker (such as server-1.properties, server-2.properties, and server-3.properties). Set a unique broker.id in each configuration file, configure zookeeper.connect to point to the address of the Zookeeper cluster, specify listeners as the IP address and port that the broker listens on (such as PLAINTEXT: / / host:ip:9092), define log.dirs as the Kafka data storage directory, set the number of partitions to num.partitions = 3 and the replication factor to default.replication.factor = 3, and configure min.insync.replicas to 2 to ensure reliable message writing.

[0019] Step 1.4: Execute the startup script on each broker node, for example, bin / kafka-server-start.shconfig / server-1.properties, and verify that each broker successfully connects to Zookeeper and registers with the Kafka cluster.

[0020] Step 1.5: Use the Kafka command-line tool to check the status of the brokers and confirm that all brokers are running and can communicate with each other.

[0021] Furthermore, step 2 is as follows:

[0022] Step 2.1: Make sure Docker and Docker Compose (if necessary) are installed on the target server, and pull the Apache Doris official image or build a custom image as needed;

[0023] Step 2.2: Write a Dockerfile for the Doris FE node. Based on the apache / doris:build-env-for-2.0 image, create a Doris installation directory, copy and unzip the Doris binary package, set the environment variable DORIS_HOME, and expose the necessary ports (9030, 9050, 9060, 9010). Write a similar Dockerfile for the Doris BE node, ensuring that the startup command is start_be.sh.

[0024] Step 2.3: Use the Docker Builder command to build the FE and BE images respectively;

[0025] Step 2.4: Run the command sysctl -w vm.max_map_count=2000000 on the host machine to set the kernel parameters to ensure the resources required for Doris to run;

[0026] Step 2.5: Use the Docker command to start the FE container, set the environment variables FE_SERVERS and FE_ID, mount the data directory to the host machine, and specify the network mode as host;

[0027] Step 2.6: Use the Docker command to start the BE container, set the environment variables FE_SERVERS and BE_ADDR, mount the data directory to the host, and specify the network mode as host;

[0028] Step 2.7: Use the MySQL client to connect to the FE node (for example, mysql-client -h fe_host -Pquery_port -uroot) to check whether the FE and BE nodes are registered and communicating with each other, and confirm that the cluster status is healthy.

[0029] Furthermore, step 3 is as follows:

[0030] Step 3.1: Choose a suitable Kafka producer library based on the programming language you are using, such as kafka-clients for Java or confluent-kafka for Python.

[0031] Step 3.2: Set the bootstrap server address of the Kafka cluster (bootstrap.servers), enable idempotence configuration (enable.idempotence=true) to ensure message uniqueness, configure the number of retries (retries) and the acks level for message sending (acks=all) to improve reliability, and optimize the buffer size and batch sending parameters (such as batch.size and linger.ms) according to business needs;

[0032] Step 3.3: Generate a globally unique messageId using UUID, snowflake algorithm, or other mechanism;

[0033] Step 3.4: In the producer code, ensure that each message contains the unique messageId field. Use a method with a callback mechanism when sending messages to ensure confirmation and error handling of message sending, and send the patient data in a unified format to the specified Kafka topic.

[0034] Furthermore, step 4 is as follows:

[0035] Step 4.1: Choose a suitable Kafka consumer library based on your programming language, such as kafka-clients for Java.

[0036] Step 4.2: Set the consumer group ID (spring.kafka.consumer.group-id=etl-server-group), the automatic offset reset policy (spring.kafka.consumer.auto-offset-reset=latest), disable auto-commit offsets (spring.kafka.consumer.enable-auto-commit=false), configure the deserializers (spring.kafka.consumer.key-deserializer and spring.kafka.consumer.value-deserializer), and set the number of concurrent threads (spring.kafka.listener.concurrency=5) to improve processing efficiency;

[0037] Step 4.3: Subscribe to the specified Kafka topic, listen for and pull new messages, save the messages to the local database first, ensure that the save is successful, and then manually submit the offset to the Kafka broker;

[0038] Step 4.4: Use multithreading or thread pooling to process messages in parallel, improving data processing throughput and response speed, and ensuring that the ETL server can efficiently process large amounts of data.

[0039] Step 4.5: Deploy the ETL service on the target server, ensure that it can run stably and continuously consume messages from Kafka, and regularly monitor its operating status and performance indicators.

[0040] Furthermore, step 5 is as follows:

[0041] Step 5.1: Write logic to identify and handle missing values, ensure that the patient's unique identifier (such as ID number, medical number) exists in each visit or hospitalization record, and fill or mark missing data to ensure data integrity.

[0042] Step 5.2: Implement a duplicate data detection mechanism to check whether there are identical records in the database using unique identifiers (such as messageId or patient ID) and perform deduplication to ensure data uniqueness.

[0043] Step 5.3: Standardize the date and time formats. Convert all dates to yyyy-MM-dd format and all times to yyyy-MM-dd HH:mm:ss format to ensure data consistency and facilitate subsequent processing.

[0044] Step 5.4: Clean the text to remove useless spaces and special characters, and process the beginning and end of the string to ensure the neatness and readability of the data.

[0045] Step 5.5: During the data processing process, if a conversion exception is encountered, an error email is sent to notify the system administrator, and a log is recorded. The exception message status is updated to exception for subsequent processing;

[0046] Step 5.6: Convert the cleaned and standardized data into a unified structured format, ensuring that field names (e.g., name, gender) are consistent across different systems, and prepare the data for subsequent NLP processing and storage.

[0047] Step 5.7: After data cleaning and standardization, use the NLP model to label and identify errors in the patient's hospitalization records, diagnosis information, and bone density data to further improve the data availability and accuracy.

[0048] Furthermore, step 6 is as follows:

[0049] Step 6.1: Choose an NLP model suitable for the medical field, such as BERT or a custom-built model. Ensure that the model has been pre-trained or fine-tuned for medical text (e.g., electronic medical records, diagnostic reports).

[0050] Step 6.2: Integrate the NLP model into the ETL server and ensure that the model can receive pre-processed text data and output labeled results.

[0051] Step 6.3: Input the cleaned and standardized text data, such as patient hospitalization records, outpatient records, and diagnosis information, into the NLP model, ensuring that the data format meets the model's input requirements.

[0052] Step 6.4: Use the NLP model to analyze the input text data, extract and generate patient labels (such as hypertension, diabetes, etc.) and identify incorrect information in fracture type and medical history;

[0053] Step 6.5: Merge the labels and recognition results generated by the NLP model with the original data to form a structured data record containing labeled information;

[0054] Step 6.6: For errors or abnormal data identified by the NLP model, mark these records and record relevant information in the error log for subsequent manual review and processing;

[0055] Step 6.7: Based on feedback and error recognition results from actual applications, continue to optimize the NLP model to improve its accuracy and applicability by increasing training data and adjusting model parameters.

[0056] Furthermore, step 7 is as follows:

[0057] Step 7.1: Configure the connection to the Apache Doris database in the ETL server to ensure that data can be written to the Doris cluster smoothly;

[0058] Step 7.2: Use the snowflake algorithm or other distributed unique ID generation mechanism to generate a globally unique primary key ID for each record to ensure data uniqueness and traceability;

[0059] Step 7.3: Map the structured data after ETL processing to the table structure in the Doris database, ensuring the consistency of field names and data types;

[0060] Step 7.4: Use the hash partitioning strategy to evenly distribute data to different partitions based on the ID field to avoid excessive data in a single partition and improve query performance.

[0061] Step 7.5: Efficiently write the converted data into the Doris database through batch insert or streaming write to ensure data integrity and consistency;

[0062] Step 7.6: Configure the number of partition copies in the Doris cluster to 3 to ensure that data has redundant backups on multiple nodes, ensuring that data is not lost in the event of a single point of failure and that the system continues to operate normally;

[0063] Step 7.7: Use Doris's query tools (such as MySQL client) to verify whether the data is stored correctly, check the uniqueness of the primary key and the uniformity of the partition distribution, and ensure that the data storage process is correct.

[0064] Furthermore, step 8 is as follows:

[0065] Step 8.1: Select an appropriate web framework (such as Spring Boot, Django, Express) based on project requirements to quickly develop and deploy query interfaces.

[0066] Step 8.2: Design a multi-dimensional patient data query interface based on business requirements, and determine the query parameters that need to be supported and the returned data format.

[0067] Step 8.3: Write query logic in the web server to interact with the Doris database to ensure that the required patient data can be retrieved from Doris efficiently.

[0068] Step 8.4: Implement authentication (e.g., OAuth2, JWT) and permission control mechanisms to ensure that only authorized users can access and query patient data, protecting data privacy and security.

[0069] Step 8.5: Use a caching mechanism (such as Redis) to cache frequently used query results, reduce direct access to the Doris database, and improve query response speed. At the same time, optimize database indexes and query statements to improve query efficiency.

[0070] Step 8.6: Deploy the web server to the production environment and use a load balancer (such as Nginx or HAProxy) to distribute traffic and ensure that the system can handle high concurrent query requests.

[0071] Step 8.7: Establish a comprehensive monitoring system (such as Prometheus and Grafana) to monitor the performance indicators and interface calls of the web server in real time, record access logs and error logs, and promptly identify and resolve potential problems.

[0072] Step 8.8: Perform comprehensive functional and performance testing to ensure that all query interfaces work as expected and can maintain stable and efficient responses under high load.

[0073] Step 8.9: Based on user feedback and business needs, continuously optimize and expand the query interface, regularly update security policies and performance optimization measures, and ensure the long-term stable operation of the Web service.

[0074] Compared with the prior art, the present invention has the following beneficial effects:

[0075] 1. The present invention provides a multi-source heterogeneous medical data aggregation and federated query method, which standardizes and integrates medical data from different sources and in non-uniform formats, reducing the complexity of cross-system data aggregation.

[0076] 2. The present invention provides a method for aggregating and federating medical data from multiple sources and heterogeneous structures. It uses the high-performance distributed database Doris for data storage and query, and combines NLP technology to achieve automatic labeling and anomaly identification of unstructured data such as medical records, diagnostic reports, and imaging information. This improves data quality while making it easier for subsequent analysis and decision support.

[0077] 3. The present invention provides a multi-source heterogeneous medical data aggregation and federated query method, which uses a distributed database for high-performance storage and query to meet the needs of rapid retrieval and analysis of multi-source data.

[0078] 4. The present invention provides a method for aggregating and federating medical data from multiple sources and heterogeneous structures, provides a unified access interface, implements federated queries on previously independent data sources, reduces multi-system switching and complex operations, improves data utilization efficiency and scalability, and provides more timely and comprehensive data support for medical decision-making and scientific research analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0080] Figure 1 This is a schematic diagram of the steps of the present invention

[0081] Figure 2 This is a system architecture diagram

[0082] Figure 3 Schematic diagram of node health status

[0083] Figure 4 System topology diagram

[0084] Figure 5 It is the core process of etl message processing DETAILED DESCRIPTION

[0085] The technical solutions of the present invention will be more clearly and completely explained below through description of preferred embodiments of the present invention in conjunction with the accompanying drawings.

[0086] Explanation of terms:

[0087] Kafka: Apache Kafka is a distributed stream processing platform for building real-time data pipelines and stream processing applications.

[0088] Zookeeper: Apache Zookeeper is a coordination service framework that provides unified naming, configuration management, distributed synchronization, and group services for distributed applications. It is used to store and manage cluster metadata in Kafka.

[0089] Docker: Docker is an open source containerization platform that packages applications and their dependent environments into lightweight, portable containers for rapid deployment and operation in different environments.

[0090] Apache Doris (FE, BE): Doris is a modern distributed columnar analytical database. The FE (FrontEnd) front-end node is responsible for metadata management and SQL parsing optimization, and the BE (BackEnd) back-end node is responsible for data storage and query calculation.

[0091] ETL: refers to the process of extracting, transforming and loading data, which is used to uniformly extract, cleanse and transform data from different sources and formats, and load it into the target database or data warehouse.

[0092] NLP: Natural language processing technology is used to understand and analyze text data. In this method, it is used to extract disease labels, diagnostic information, and abnormal data from medical text.

[0093] UUID: Universally unique identifier, which generates a globally unique ID through an algorithm to ensure the uniqueness of different data records in a distributed system.

[0094] Snowflake algorithm: A distributed ID generation algorithm that generates globally unique, trend-ordered IDs by combining timestamps, machine IDs, and serial numbers. It is often used to ensure the uniqueness of data primary keys.

[0095] CIDR: Classless Inter-Domain Routing notation, a method used to describe IP addresses and network prefixes so that programs (such as Doris) can select a matching local IP based on the specified network segment.

[0096] MIMIC-III, MedQuAD, PubMed: MIMIC-III is a public database of intensive care unit clinical data, MedQuAD is a medical question-answering dataset, and PubMed is a biomedical literature database. These data sources can be used for training or fine-tuning NLP models.

[0097] Callback (message callback): The callback function configured when sending a message on the message production side. When the message is sent successfully or unsuccessfully, the callback is triggered for corresponding processing to ensure the reliability of message sending and error handling capabilities.

[0098] MySQL client: MySQL database client tool used to connect to and execute SQL queries on MySQL or MySQL-compatible databases (such as Doris FE).

[0099] Eureka: Eureka is Netflix's open source service registration and discovery component, which is used to enable different services to automatically register and discover each other in a microservice architecture, thereby achieving dynamic service management.

[0100] Nginx: Nginx is a high-performance HTTP server and reverse proxy server.

[0101] HAProxy: HAProxy is an open source software that provides a proxy.

[0102] Redis: Redis is an open source in-memory data structure storage system.

[0103] OAuth2: OAuth 2.0 is an authorization protocol.

[0104] JWT: A compact and self-contained JSON-based access token.

[0105] SQL: Structured Query Language,

[0106] Prometheus: An open source monitoring and alerting system.

[0107] Grafana: An open source data visualization tool.

[0108] BERT: Pre-trained language model.

[0109] DNS: Domain Name System.

[0110] JDBC: Java database connectivity technology.

[0111] FE_SERVERS, FE_ID, BE_ADDR: Doris cluster environment variable configuration items. FE_SERVERS specifies FE node information, FE_ID is the current FE node ID, and BE_ADDR is the current BE node address.

[0112] priority_networks: Doris configuration item, used to give priority to the IP that matches the specified network segment (CIDR format) as the local IP when there are multiple IPs.

[0113] bootstrap.servers: One of the configuration properties of Kafka producers and consumers, used to specify the bootstrap server address list of the Kafka cluster.

[0114] enable.idempotence: Kafka producer configuration item, enabling message idempotence to ensure that messages are not persisted repeatedly when retrying to send, thereby ensuring message uniqueness.

[0115] acks: Kafka producer configuration item, which specifies the confirmation policy after the message is sent.

[0116] batch.size: Kafka producer configuration item.

[0117] linger.ms: Kafka producer configuration item.

[0118] enable-auto-commit: Kafka consumer configuration item, whether to automatically commit message offsets.

[0119] auto-offset-reset: Kafka consumer configuration item, where to start consuming when no offset is available (such as the latest message or the earliest message).

[0120] deserializer: Kafka consumer or producer configuration item used to specify the deserialization / serialization class of message keys and values.

[0121] Concurrency: In this method, it refers to the number of concurrent Kafka consumers or ETL processing threads, which is used to improve processing throughput.

[0122] Quartz: An open source distributed task scheduling framework for managing and executing scheduled tasks.

[0123] like Figure 1 As shown, the present invention is:

[0124] Step 1: Build a Kafka cluster and configure Broker to achieve high availability;

[0125] Step 2: Use Docker to deploy the Apache Doris cluster, configure the front-end and back-end nodes, and specify the IP addresses;

[0126] Step 3: Configure the Kafka producer to ensure that each message contains a unique messageId field and send it to the topic;

[0127] Step 4: Deploy the ETL server as a Kafka consumer, join the same consumer group, and configure multi-threaded message processing;

[0128] Step 5: Implement data processing in the ETL server, including missing value identification, duplicate detection, format standardization, and text cleaning;

[0129] Step 6: Use NLP models to label patient records and identify errors.

[0130] Step 7: Store the processed data into the Apache Doris database, generate a unique primary key and adopt a partitioning strategy;

[0131] Step 8: Deploy a web server to provide a patient data query interface.

[0132] A multi-source heterogeneous medical data aggregation and federated query system, suitable for a multi-source heterogeneous medical data aggregation and federated query method, the system is built as follows Figure 2 As shown in the figure, on the left are various medical data sources (HIS, PACS, EMR, LIS), which send corresponding data to respective topics in the Kafka cluster (such as HIS topic, PACS topic, EMR topic, LIStopic). The Kafka cluster acts as a message transfer and buffer layer, providing subscription capabilities for unified and orderly data.

[0133] In the central federated query system, the ETL server, acting as a data processing unit, subscribes to the corresponding Kafka topic, cleans, converts, and standardizes the pulled messages, and then stores the organized data in the Doris database cluster. Doris, a high-performance distributed database, is used to centrally store organized medical data, enabling efficient query and analysis.

[0134] On the right side of the federated query system is the Web service layer, which performs query and analysis operations by accessing the Doris database and provides a unified data query interface. Ultimately, front-end clients (such as client 1, client 2, and client 3) can obtain the required analysis results and query data from the underlying multi-source heterogeneous data by calling the Web service interface.

[0135] The details are as follows:

[0136] 1.1 Kakfa cluster construction

[0137] This patent uses Kafka's own Zookeeper to manage the cluster.

[0138] The broker configuration is as follows (taking 3 nodes as an example):

[0139] #Kafka Broker ID, each node must be unique

[0140] broker.id=1

[0141] #Zookeeper connection address

[0142] zookeeper.connect=host:ip

[0143] #Kafka listening address

[0144] listeners=PLAINTEXT: / / host:ip

[0145] #Kafka data storage directory

[0146] log.dirs= / var / lib / kafka-logs

[0147] #Set the number of partitions and replicas

[0148] num.partitions=3

[0149] default.replication.factor=3

[0150] #The minimum number of synchronized Fubens to ensure successful message writing

[0151] min.insync.replicas=2

[0152] replica.lag.time.max.ms

[0153] 1.2 Kafka producer

[0154] Kafka producer configuration is set by different producers themselves. Producers can adjust related parameters according to their business characteristics and hardware resource configuration. Just ensure that there is only one message sent.

[0155] 1.3 Kafka consumer

[0156] ETL is the consumer of the message, and the relevant configuration is as follows:

[0157] #etle clusters belong to the same groupid

[0158] spring.kafka.consumer.group-id=etl-server-group

[0159] #Start consuming from the latest news

[0160] spring.kafka.consumer.auto-offset-reset=latest

[0161] # Turn off automatic message submission

[0162] spring.kafka.consumer.enable-auto-commit=false

[0163] #Serialization, deserialization configuration

[0164] spring.kafka.consumer.key-deserializer=org.apache.kafka.common.serialization.StringDeserializer

[0165] spring.kafka.consumer.value-deserializer=org.apache.kafka.common.serialization.StringDeserializer

[0166] #Concurrency thread count

[0167] spring.kafka.listener.concurrency=5

[0168] 1.4 Create kafka topic

[0169] kafka-topics.sh --create --topic <your-topic-name>--bootstrap-server <kafka-broker>:9092--partitions <num-partitions>--replication-factor <num-replicas>

[0170] Parameter Description:

[0171] --create: Create a topic.

[0172] --topic <your-topic-name>:Specify the name of the topic.

[0173] --bootstrap-server <kafka-broker>:9092: The boot server address of the Kafka cluster (any broker address can be used).

[0174] --partitions <num-partitions>: Set the number of partitions for the topic.

[0175] --replication-factor <num-replicas>: Set the replication factor to specify the number of replicas for each partition.

[0176] According to this command, create topics for the upstream data source: pacs-source-topic, emr-source-topic, lis-source-topic, and his-source-topic.

[0177] 2. Dorisdb cluster construction plan

[0178] This patent uses Docker containerization to deploy a DorisDB cluster, with one Doris FE front-end node and five Doris BE nodes. The Docker files for Doris FE and BE are as follows:

[0179] fe:

[0180] #doris official mirror basic environment

[0181] FROM apache / doris:build-env-for-2.0

[0182] ENTRYPOINT["sh","-c"]

[0183] RUN mkdir-p / opt / doris

[0184] #Copy the downloaded Doris binary file into the Docker container

[0185] COPY apache-doris-2.0.0-bin-x64.tar.gz / opt / doris

[0186] #Unzip the Doris binary file

[0187] RUN tar-zxvf / opt / doris / apache-doris-2.0.0-bin-x64.tar.gz-C / opt / doris

[0188] #Set the environment variable to point to the decompressed Doris directory

[0189] ENV DORIS_HOME= / opt / doris / apache-doris-2.0.0-bin-x64

[0190] #Expose the default ports of Doris FE and BE

[0191] EXPOSE 9030 9050 9060 9010

[0192] #Start Doris FE and BE services

[0193] CMD[" / opt / doris / apache-doris-2.0.0-bin-x64 / fe / bin / start_fe.sh"]

[0194] be:

[0195] #doris official mirror basic environment

[0196] FROM apache / doris:build-env-for-2.0

[0197] ENTRYPOINT["sh","-c"]

[0198] RUN mkdir-p / opt / doris

[0199] #Copy the downloaded Doris binary file into the Docker container

[0200] COPY apache-doris-2.0.0-bin-x64.tar.gz / opt / doris

[0201] #Unzip the Doris binary file

[0202] RUN tar-zxvf / opt / doris / apache-doris-2.0.0-bin-x64.tar.gz-C / opt / doris

[0203] #Set the environment variable to point to the decompressed Doris directory

[0204] ENV DORIS_HOME= / opt / doris / apache-doris-2.0.0-bin-x64

[0205] #Expose the default ports of Doris FE and BE

[0206] EXPOSE 9030 9050 9060 9010

[0207] #Start Doris FE and BE services

[0208] CMD[" / opt / doris / apache-doris-2.0.0-bin-x64 / be / bin / start_be.sh"]

[0209] Before starting, execute the following command on the host: sysctl -w vm.max_map_count = 2000000

[0210] Regarding IP binding, due to the presence of multiple network cards or the presence of virtual network cards caused by the installation of environments such as docker, the same host may have multiple different IPs. Currently, Doris cannot automatically identify available IPs. Therefore, when there are multiple IPs on the deployment host, the priority_networks configuration item must be used to force the specification of the correct IP. For example, priority_networks = 10.1.3.0 / 24, which is a CIDR representation method. FE or BE will look for a matching IP based on this configuration item as its own local IP

[0211] Use the docker buidler command to build the fe and be images respectively. The image names are apache / doris:2.0.0-fe and apache / doris:2.0.0-be

[0212] Start fe:

[0213] docker run-itd\

[0214] --name=fe-01\

[0215] --env FE_SERVERS="fe1:{fehost}:9010"\

[0216] --env FE_ID=1\

[0217] -v / root / data / doris / fe / doris-meta: / opt / doris / apache-doris-2.0.0-bin-x64 / fe / doris-meta\

[0218] -v / root / data / doris / fe / log: / opt / doris / apache-doris-2.0.0-bin-x64 / fe / log\

[0219] -v / root / data / doris / fe / conf: / opt / doris / apache-doris-2.0.0-bin-x64 / fe / conf\

[0220] --network=host\

[0221] apache / doris:2.0.0-fe

[0222] Start be:

[0223] docker run-itd\

[0224] --name=be-01\

[0225] --env FE_SERVERS="fe1:{fehost}:9010"\

[0226] --env BE_ADDR="{behost}:9050"\

[0227] -v / root / data / doris / be / storage: / opt / doris / apache-doris-2.0.0-bin-x64 / be / storage\

[0228] -v / root / data / doris / be / log: / opt / doris / apache-doris-2.0.0-bin-x64 / be / log\

[0229] -v / root / data / doris / be / conf: / opt / doris / apache-doris-2.0.0-bin-x64 / be / conf\

[0230] --network=host\

[0231] apache / doris:2.0.0-be

[0232] Replace {fehost} in the above two startup commands with the real IP address of the machine where fe is located, and replace {behost} with the real IP address of the machine where be is located.

[0233] After fe and be are started successfully, connect to fe through the mysql client tool: . / mysql-client-h fe_host-Pquery_port-uroot.

[0234] After logging in, execute the following command to add each BE:

[0235] ALTER SYSTEM ADD BACKEND"be_host:heartbeat-service_port";

[0236] When the cluster is successfully deployed, all our data is saved in the node where be is located, and all data is mapped to the local directory. When the docker process has a problem or is killed, a new process is started and the data is not lost.

[0237] After the cluster is successfully started, check the health status of each node, such as Figure 3 As shown;

[0238] Dorsi creates a table:

[0239] Each table created by this patent contains a unique logical primary key ID (the ID is guaranteed to be globally unique according to the snowflake algorithm) and uses a hash partitioning strategy on the ID to evenly distribute data and avoid excessive data volume in some partitions. The number of partition replicas is set to 3 to ensure that Doris can continue to operate normally without data loss in the event of a single point of failure.

[0240] System topology such as Figure 4 The figure shows the complete processing flow from data access to result presentation. On the left are various medical data sources (PACS, EMR, HIS, LIS). This data is sent to the Kafka cluster via a data exchange protocol (Push mode). The Kafka cluster serves as a message middleware, providing external subscription capabilities. The ETL server consumes these messages from Kafka via Pull mode, extracting, cleaning, and transforming them before storing them in the underlying Doris database.

[0241] Downstream of the ETL server, there is an AI model service (AI model service) cluster. These services provide multi-instance model reasoning capabilities through DNS load balancing. After the ETL server completes the structuring and standardization of the data, it can pass the data to the AI ​​model service through the DNS load balancing mechanism for intelligent analysis and labeling, and finally store the processing results in the Doris database cluster through a JDBC connection. The Doris database cluster includes front-end (Doris FE) and back-end (Doris BE) components, providing the system with high-performance query and analysis capabilities through a distributed architecture.

[0242] Dynamic service registration and discovery are performed between the ETL server and the web server through the Eureka service registry. As a service governance component, the Eureka Server ensures that microservice modules, such as the ETL server and the web server, can recognize and communicate with each other. Upstream data is ultimately exposed through a web server query interface. The web server cluster is then connected to Nginx as a unified reverse proxy and load balancing gateway, centrally dispatching external requests (such as those from PC clients, applets, and third-party systems). Users and applications simply connect to the web layer through Nginx to access integrated medical data and federated query results.

[0243] The entire topology implements a complete closed loop, from data source acquisition (Push to Kafka) to ETL processing (Pull from Kafka), AI analysis (accessing AI model services via DNS load balancing), data storage (Doris DB cluster), service discovery (Eureka), and unified access portal (Nginx). This architecture allows for flexible access to various data sources, loosely coupled interoperability between different components, and ultimately presents a highly available and scalable federated query service.

[0244] 1. The core process of system data processing is as follows Figure 5 As shown, the details are as follows:

[0245] Each message producer (data source) sends its data to a Kafka server through a Kafka topic, which is then subscribed to by an ETL. Producers must enable idempotence for messages to ensure message uniqueness and use a send method with a callback mechanism to ensure message uniqueness and reliability. By default, each message contains a messageId field, which producers must ensure is globally unique.

[0246] ETL obtains subscription messages from Kafka. ETL monitors the topic of the upstream data source. When a new message is pulled, it will first save the message to the local database. After successful saving, it commits the message to the Kafka broker.

[0247] Use Quartz middleware to start a distributed scheduled task, batch retrieve messages from the local message table, and evenly distribute them to each ETL server node for parallel computing based on the node data of the ETL server cluster. Each ETL node starts a thread pool to process messages concurrently.

[0248] Each parameter is set according to the online hardware configuration.

[0249] To update existing data, the ETL first queries the database based on the message's unique identifier, messageId (a globally unique messageId is sent by default when each message is sent from the producer to the Kafka broker). If an existing record is found, the data is logically deleted from the database to ensure data uniqueness. This step primarily handles updates to old data, meaning it deletes the data first and then saves it.

[0250] Check and process the default values ​​of specific messages, such as the patient's unique identifier, outpatient number, hospitalization number, etc. These fields are not saved in some upstream data sources and need to be filled in to facilitate various data query operations of subsequent businesses. Since different message types require different fields to be filled in and the filling methods are also different, this patent adopts a strategy pattern to process the fields that need to be filled in different types of messages. The specific processing method is to dynamically obtain a specific processing class based on the key value in the message and execute the corresponding strategy pattern method in the class.

[0251] Date fields are formatted uniformly. The date format of each system is different. This system uniformly formats the date and time into yyyy-MM-dd and yyyy-MM-dd HH:mm:ss. Special characters or spaces at the beginning and end of the string are processed. If the conversion is abnormal, an error email is sent to notify the system administrator in real time. The log is recorded, the local message table is updated, the status is abnormal, and the processing is ended. After the administrator receives the alarm email, he needs to log in to the error message management page, which will display a list of error messages and the cause of the error. The administrator can delete the message and let the upstream data source process the erroneous data and resend it to Kafka, or manually correct the erroneous data. The system will save the data update to the local message table, update the status of the message to unprocessed, and process the message after the next scheduled task starts.

[0252] Convert the data into a unified data structure and unify the field descriptions. For example, the patient's name and gender use different field descriptions in different systems. ETL needs to convert these fields into a unified name.

[0253] AI NLP (natural language parsing) technology is used to process patient hospitalization records, outpatient records, and diagnostic records. This primarily examines abnormalities in patients' past medical records, patient label definitions, and fracture type diagnoses. This patent utilizes a self-built model because it is adaptable to specific medical domain corpora, particularly text data containing specialized terms, disease names, and other specific concepts. Fine-tuning the model allows for optimization for specific medical needs (such as patient medical records and diagnostic reports).

[0254] Fine-tuning steps:

[0255] Data annotation: Collect and annotate electronic medical record data in the medical field, especially datasets containing medical errors (such as diagnostic errors, medication errors, treatment plan errors, etc.). The dataset should cover a variety of common error types to ensure that the model can identify and correct these errors.

[0256] Task Design: When fine-tuning, design different inputs and outputs based on the task. For example, for an error detection task, the input is medical record text, and the output is the error category, location, and cause of the error.

[0257] Corpus source: This patent uses public databases and real patient data from hospitals as corpus sources.

[0258] Public medical corpora: Use databases such as MIMIC-III, MedQuAD, and PubMed as corpora for pre-training or fine-tuning of models.

[0259] Internal hospital data: With appropriate patient privacy authorization, electronic medical records and medical records within the hospital provide more targeted real-world data for model training. The data size ranges from 5,000 to 10,000 electronic medical records and medical records.

[0260] Abnormal data check prompt:

[0261] You are a chief physician with decades of clinical experience. You are currently reviewing a patient's medical records. Your task is to use your expertise and experience to identify any errors in these records, such as inconsistencies in chronic disease descriptions or inadequate descriptions of necessary information (chief complaint, present medical history, past medical history, allergies, family history, personal history, etc.). The results will be returned in JSON format, where the key contains the case number, the error data, and the cause of the error. List of past cases: {Case List}.

[0262] Check the results returned by the AI. If there are any data anomalies, log them and send an email to notify the system administrator in real time. After receiving the alert email, the administrator needs to log in to the error message management page, which displays a list of error messages and the cause of the error. The administrator can delete the message and have the upstream data source process the erroneous data and resend it to Kafka. Alternatively, after manually correcting the erroneous data, the system will save the data updates to the local message table, update the message status to unprocessed, and process the message at the next scheduled task.

[0263] Tag processing prompt:

[0264] You are a chief physician with decades of clinical experience. You are reviewing a patient's medical records. Your task is to use your expertise and experience to identify any chronic conditions they may have suffered from. These include hypertension, diabetes, kidney disease, rheumatoid arthritis, gastrointestinal diseases (inflammatory bowel disease, malabsorption, celiac disease, peptic ulcers, etc.), thyroid or parathyroid disorders (hyperthyroidism or hypoparathyroidism), lung diseases (chronic obstructive pulmonary disease, etc.), prolonged immobilization, sexually transmitted infections (such as HIV and syphilis), sexual dysfunction (gonadal dysplasia), and immune system disorders. The results are returned in JSON format, with the key being the disease list. Past case list: {case list}.

[0265] Fracture type diagnosis prompt:

[0266] You are a chief physician with decades of clinical experience. You are reviewing a patient's diagnostic information. Your task is to use your expertise and experience to determine the patient's fracture type (if any). The main fracture types include: femur, tibia, fibula, humerus, forearm, wrist, ankle, spine, sternocostal, and pelvic fractures. The results are returned in JSON format, where the key is a list of fracture types. Diagnostic information: {diagnostic information}.

[0267] Save the NLP result information to the database.

[0268] Convert message processing into ETL standard data structure and save it to the database.

[0269] Delete successfully processed messages from the local message table.

[0270] As a specific implementation, different medical data source systems (such as PACS systems, EMR systems, LIS systems, HIS systems, etc.) send patient-related data to the Kafka cluster in the form of messages. These data sources publish data to their corresponding Kafka topics (such as pacs-source-topic, emr-source-topic, lis-source-topic, and his-source-topic) through configured Kafka producers. To ensure message reliability, each message contains a globally unique messageId field, and idempotence configuration and callback mechanism are enabled to ensure message uniqueness and reliable transmission.

[0271] A Kafka cluster consists of multiple Broker nodes, managed by Zookeeper. Each Broker node has a unique broker.id and is configured with multiple partitions and replication factors to achieve high data availability and load balancing. After a Kafka producer sends a message to a specified topic, the Kafka cluster is responsible for distributing the message to the corresponding Broker node and ensuring redundant data backup based on the configured partitioning and replication strategies. To do this, you need to download and unzip the stable version of Zookeeper, configure the zoo.cfg file, set the data directory, client port (default 2181), and cluster node information (if a cluster deployment is used), then start the Zookeeper service and ensure that it is running properly and listening on the specified port. Download the appropriate version of the Kafka binary package from the Apache Kafka official website and unzip it to the target installation directory. Create a separate configuration file for each broker, set a unique broker.id, configure zookeeper.connect to point to the address of the Zookeeper cluster, specify listeners as the IP address and port that the broker listens on (for example, PLAINTEXT: / / host:ip:9092), define log.dirs as the Kafka data storage directory, and set the number of partitions num.partitions = 3 and the replication factor default.replication.factor = 3. Configure min.insync.replicas = 2 to ensure reliable message writing. Execute the startup script on each broker node, for example, bin / kafka-server-start.sh config / server-1.properties, and use the Kafka command-line tool to check the status of the brokers to confirm that all brokers are running and can communicate with each other.

[0272] As a Kafka consumer, the ETL server joins the same consumer group (etl-server-group) and is configured to manually commit offsets to ensure reliable and consistent data processing. The ETL server uses a multi-threaded concurrent mechanism to listen for and pull new messages from Kafka in real time. Each message pulled is first saved to a local database. After successful saving, the server manually commits the offset to prevent data loss due to processing failures. When deploying the ETL server, select an appropriate Kafka consumer client library and configure it based on your programming language, such as kafka-clients for Java. Configure consumer parameters, including the consumer group ID, automatically reset offsets to the latest message, disable automatic offset commit, configure the deserializer, and set the number of concurrent threads to improve processing efficiency. Write the ETL server code to subscribe to a specified Kafka topic, listen for and pull new messages, first save the messages to a local database, and then manually commit the offsets to the Kafka broker after successful saving. Use multi-threading or a thread pool mechanism to parallelize message processing, improving data processing throughput and responsiveness, ensuring the ETL server can efficiently handle large amounts of data.

[0273] The ETL server performs a series of cleansing and standardization processes on the received medical data. This includes identifying and handling missing values, ensuring that unique patient identifiers (such as ID number, medical consultation number, and hospitalization number) are present in every medical or hospitalization record, and filling or marking missing data to ensure data integrity. Unique identifiers such as messageId or patient ID are used to detect whether identical records exist in the database. If so, duplicates are removed to avoid data redundancy. All date fields are converted to the yyyy-MM-dd format, and time fields are converted to the yyyy-MM-dd HH:mm:ss format to ensure data consistency and facilitate subsequent processing. Text cleaning is performed to remove unnecessary spaces, special characters, and other noise from fields to ensure the neatness and readability of the text data. If any conversion exceptions are encountered during data processing, an error email is sent to the system administrator, a log is recorded, and the exception message status is updated to exception for subsequent processing. The cleaned and standardized data is converted into a unified structured format, ensuring that field names (such as name and gender) are consistent across different systems, and preparing the data for subsequent NLP processing and storage.

[0274] After completing data cleaning and standardization, the ETL server calls the pre-trained NLP model to analyze and process the patient's hospitalization records, diagnostic information, and bone density data. First, select an NLP model suitable for the medical field, such as BERT or a self-built dedicated model, and integrate it into the ETL process to support the processing of medical text data. Through the NLP model, various labels for patients (such as hypertension, diabetes, osteoporosis, etc.) are extracted and generated, and potential erroneous information in fracture types and medical history is identified. For abnormal data identified by the NLP model, these records are marked and recorded in the error log, and then a notification email is sent to the system administrator through the alarm system to facilitate subsequent manual review and processing. Based on feedback and error recognition effects in actual applications, the NLP model is continuously optimized, and the accuracy and applicability of the model are improved by increasing training data and adjusting model parameters.

[0275] After cleansing, standardization, and labeling, the ETL server stores the data in an Apache Doris database. First, the ETL server configures a connection to the Doris database to ensure smooth data writing to the Doris cluster. A globally unique primary key ID is generated using the snowflake algorithm to ensure uniqueness and traceability of each record in the database. The structured data after ETL processing is mapped to the table structure in the Doris database, ensuring consistency in field names and data types. A hash partitioning strategy is used based on the ID field to evenly distribute data across partitions to avoid excessive data in a single partition. Furthermore, the number of partition replicas is set to 3 to ensure data redundancy across multiple nodes, ensuring high availability and data security. Structured data is efficiently written to the Doris database through batch inserts or streaming writes, ensuring data integrity and consistency. A MySQL client is used to connect to the FE node to verify that the FE and BE nodes are properly registered and communicating with each other, confirming the cluster's health, and verifying that the data is correctly stored. The uniqueness of the primary key and the uniformity of the partition distribution are also checked to ensure that the data storage process is correct.

[0276] The web server deployed in the system serves as a unified service entry point, providing a multi-dimensional patient data query interface that allows users to access and analyze aggregated data using various query criteria. Select an appropriate web framework, such as Spring Boot, Django, or Express, and design and implement the multi-dimensional patient data query interface based on project requirements, determining supported query parameters and the return data format. Develop query logic within the web server that interacts with the Apache Doris database to ensure efficient retrieval of required patient data from Doris. Integrate authentication mechanisms (such as OAuth2 or JWT) and permission control to ensure only authorized users can access and query sensitive medical data, protecting data privacy and security. To optimize query performance, introduce a caching mechanism (such as Redis) to cache frequently used query results, reducing direct access to the Doris database. Optimize database indexes and query statements to improve query efficiency. Deploy the web server in a production environment and use a load balancer (such as Nginx or HAProxy) to distribute query requests, ensuring the system can handle high-concurrency queries and dynamically scale server instances as needed to accommodate traffic growth. Establish a complete monitoring system to monitor the performance indicators and interface calls of the Web server in real time, record access logs and error logs, promptly discover and resolve potential problems, and ensure the stable operation and efficient response of the Web service.

[0277] To ensure stable system operation and efficient performance, comprehensive monitoring and maintenance measures are implemented. Monitoring tools such as Prometheus and Grafana are used to monitor the operational status, resource utilization, and performance metrics of the Kafka cluster, Apache Doris database cluster, and ETL server in real time to ensure the healthy operation of each component. Thresholds and alert rules for key indicators are configured to ensure timely notification and resolution of system anomalies, preventing escalation. A regular backup mechanism is established to back up important data in the Doris database, ensuring rapid recovery and data security in the event of data loss or system failure. To enhance system stability and performance, system components such as the software versions and configuration parameters of Kafka, Doris, and the ETL server are regularly updated and optimized based on actual operational conditions and business needs, improving overall system performance and reliability. Furthermore, based on user feedback and data analysis results, NLP models and data processing processes are continuously optimized to enhance data processing accuracy and efficiency. Regular security audits and vulnerability scans are conducted to ensure system security, prevent data leaks and unauthorized access, and maintain long-term stable system operation.

[0278] The above-described specific embodiments merely describe preferred embodiments of the present invention and do not limit the scope of protection of the present invention. Any modifications, substitutions, and improvements made to the technical solution of the present invention by a person skilled in the art based on the textual description and drawings provided herein, without departing from the design concept and spirit of the present invention, shall fall within the scope of protection of the present invention. The scope of protection of the present invention is determined by the claims.

Claims

1. A multi-source heterogeneous medical data aggregation and federated query method, characterized in that: The following steps are involved: Step 1: Build a Kafka cluster and configure Broker to achieve high availability; Step 2: Use Docker to deploy the Apache Doris cluster, configure the front-end and back-end nodes, and specify the IP addresses; Step 3: Configure the Kafka producer to ensure that each message contains a unique messageId field and is sent to the topic; Step 4: Deploy the ETL server as a Kafka consumer, join the same consumer group and configure multi-threaded message processing; Step 5: Implement data processing in the ETL server, including missing value identification, duplicate detection, format standardization, and text cleaning; Step 6: Use NLP models to label patient records and identify errors. Step 7: Store the processed data in the Apache Doris database, generate a unique primary key and adopt a partitioning strategy; Step 8: Deploy a Web server to provide a patient data query interface.

2. A multi-source heterogeneous medical data aggregation and federated query method according to claim 1, characterized in that: Step 1 is as follows: Step 1.1: Download and decompress the stable version of Zookeeper, configure the zoo.cfg file, set the data directory, client port, and cluster node information, and then start the Zookeeper service; Step 1.2: Download the Kafka binary package and unzip it to the target installation directory; Step 1.3: Create a separate configuration file for each Broker, set a unique broker.id in each configuration file, configure zookeeper.connect to point to the address of the Zookeeper cluster, specify listeners as the IP address and port that the Broker listens on, define log.dirs as the Kafka data storage directory, and set the number of partitions num.partitions = 3 and the replication factor default.replication.factor = 3, and configure min.insync.replicas = 2 to ensure reliable writing of messages; Step 1.4: Execute the startup script on each Broker node and verify that each Broker successfully connects to Zookeeper and registers in the Kafka cluster; Step 1.5: Use the Kafka command-line tool to check the status of the brokers and confirm that all brokers are running and can communicate with each other.

3. A multi-source heterogeneous medical data aggregation and federated query method according to claim 1, characterized in that: Step 2 is as follows: Step 2.1: Make sure Docker and Docker Compose are installed on the target server, and pull the Apache Doris official image or build a custom image as required; Step 2.2: Write a Dockerfile for the Doris FE node. Based on the apache / doris:build-env-for-2.0 image, create a Doris installation directory, copy and unzip the Doris binary package, set the environment variable DORIS_HOME, and expose the necessary ports. Write a similar Dockerfile for the Doris BE node, and make sure the startup command is start_be.sh. Step 2.3: Use the Docker Builder command to build the front-end and back-end images respectively; Step 2.4: Execute the command sysctl-w vm.max_map_count=2000000 on the host machine to set the kernel parameters to ensure the resources required for Doris to run; Step 2.5: Use the Docker command to start the front-end container, set the environment variables FE_SERVERS and FE_ID, mount the data directory to the host, and specify the network mode as host; Step 2.6: Use the Docker command to start the backend container, set the environment variables FE_SERVERS and BE_ADDR, mount the data directory to the host, and specify the network mode as host; Step 2.7: Use the MySQL client to connect to the front-end node, check whether the front-end and back-end nodes are registered normally and communicate with each other, and confirm that the cluster status is healthy.

4. The multi-source heterogeneous medical data aggregation and federated query method according to claim 1, characterized in that: Step 3 is as follows: Step 3.1: Select the Kafka producer library based on the programming language used; Step 3.2: Set the bootstrap server address of the Kafka cluster, enable idempotence configuration, configure the number of retries and acks level for message sending to improve reliability, and optimize the buffer size and batch sending parameters according to business needs; Step 3.3: Generate a globally unique messageId using UUID, snowflake algorithm or other mechanism; Step 3.4: In the producer code, ensure that each message contains the unique messageId field, and use a method with a callback mechanism when sending messages to ensure confirmation and error handling of message sending, and send the patient data to the specified Kafka topic in a unified format.

5. The multi-source heterogeneous medical data aggregation and federated query method according to claim 1, characterized in that: Step 4 is as follows: Step 4.1: Select Kafka consumer library based on programming language; Step 4.2: Set the consumer group ID, automatic offset reset strategy, turn off automatic offset commit, configure the deserializer, and set the number of concurrent threads to improve processing efficiency; Step 4.3: Subscribe to the Kafka topic, listen for and pull new messages, save the messages to the local database first, and manually submit the offset to the Kafka broker after ensuring that the message is saved successfully; Step 4.4: Use multithreading or thread pool mechanism to process messages in parallel; Step 4.5: Deploy the ETL service on the target server and regularly monitor its running status and performance indicators.

6. A multi-source heterogeneous medical data aggregation and federated query method according to claim 1, characterized in that: Step 5 is as follows: Step 5.1: Write logic to identify and handle missing values, ensure that a unique identifier for the patient is present in each visit or hospitalization record, and fill in or mark missing data; Step 5.2: Implement a duplicate data detection mechanism, check whether there are exactly the same records in the database through unique markers, and perform deduplication processing; Step 5.3: Unify the date and time formats, convert all dates to yyyy-MM-dd format, and convert time to yyyy-MM-dd HH:mm:ss format; Step 5.4: Perform text cleaning, remove useless spaces and special characters, and process the beginning and end of the string; Step 5.5: During the data processing process, if a conversion exception is encountered, an error email is sent to notify the system administrator, and a log is recorded, and the exception message status is updated to exception; Step 5.6: Convert the cleaned and standardized data into a unified structured format and prepare the data for subsequent NLP processing and storage; Step 5.7: After data cleaning and standardization, call the NLP model to label and identify errors in the patient’s hospitalization records, diagnosis information, and bone density data.

7. The multi-source heterogeneous medical data aggregation and federated query method according to claim 1, characterized in that: Step 6 is as follows: Step 6.1: Choose an NLP model for the medical domain and make sure the model has been pre-trained or fine-tuned for medical texts. Step 6.2: Integrate the NLP model in the ETL server and ensure that the model can receive the preprocessed text data and output the labeled results; Step 6.3: Input the cleaned and standardized text data into the NLP model, making sure the data format meets the model’s input requirements; Step 6.4: Use the NLP model to analyze the input text data, extract and generate patient labels, and identify incorrect information in fracture type and medical history; Step 6.5: Merge the labels and recognition results generated by the NLP model with the original data to form a structured data record containing labeled information; Step 6.6: For errors or abnormal data identified by the NLP model, mark these records and record relevant information in the error log; Step 6.7: Based on the feedback and error recognition results in actual applications, continue to optimize the NLP model, and improve the accuracy and applicability of the model by increasing training data and adjusting model parameters.

8. The multi-source heterogeneous medical data aggregation and federated query method according to claim 1, characterized in that: Step 7 is as follows: Step 7.1: Configure the connection to the Apache Doris database in the ETL server to ensure that data can be written to the Doris cluster smoothly; Step 7.2: Use the snowflake algorithm or other distributed unique ID generation mechanism to generate a globally unique primary key id for each record; Step 7.3: Map the structured data after ETL processing to the table structure in the Doris database; Step 7.4: Use the hash partitioning strategy to evenly distribute data to different partitions based on the id field to avoid excessive data in a single partition; Step 7.5: Efficiently write the converted data into the Doris database through batch insertion or streaming writing; Step 7.6: Configure the number of partition copies in the Doris cluster to 3 to ensure that the data has redundant backup on multiple nodes; Step 7.7: Use Doris' query tool to verify that the data is stored correctly, check the uniqueness of the primary key and the uniformity of the partition distribution.

9. The multi-source heterogeneous medical data aggregation and federated query method according to claim 1, characterized in that: Step 8 is as follows: Step 8.1: Select a web framework based on project requirements; Step 8.2: Design a multi-dimensional patient data query interface based on business requirements, and determine the query parameters that need to be supported and the returned data format; Step 8.3: Write query logic in the web server to interact with the Doris database to ensure that the required patient data can be retrieved from Doris efficiently; Step 8.4: Implement authentication and permission control mechanisms to ensure that only authorized users can access and query patient data; Step 8.5: Use a cache mechanism to cache common query results, reduce direct access to the Doris database, and optimize database indexes and query statements; Step 8.6: Deploy the web server to the production environment and use a load balancer to distribute traffic to ensure that the system can handle high concurrent query requests; Step 8.7: Monitor the performance indicators and interface calls of the Web server in real time, and record access logs and error logs; Step 8.8: Conduct comprehensive functional and performance testing to ensure that all query interfaces work as expected and can maintain stable and efficient responses under high load; Step 8.9: Based on user feedback and business needs, continue to optimize and expand the query interface, and regularly update security policies and performance optimization measures.

Citation Information

Cited By

  • Mass process data presentation method based on Websocket

    CN122064754A