Full-scale Medical Data Synchronization System and Method

By introducing specific processors and API interfaces into Apache NiFi, the interaction between the data governance platform and NiFi is achieved, the data flow state feedback problem is solved, the complexity of medical data synchronization is reduced, and synchronization efficiency and maintainability are improved.

CN114168678BActive Publication Date: 2025-07-18TIANJIN HEALTH CARE BIG DATA CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111317707.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-09
Publication Date
2025-07-18
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

In the prior art, Apache NiFi lacks an interactive mechanism with external systems, resulting in the data flow status that cannot be fed back to the data governance platform and cannot meet the requirements of the data governance ETL processing logic, increasing the complexity of health and medical data synchronization and reducing synchronization efficiency.

Method used

By introducing process start processors, text segmentation processors, data query processors, data library processors and process end processors in Apache NiFi, combining exception monitoring processors and API interfaces, the interaction between the data governance platform and NiFi is realized, meeting the visual process orchestration and automated scheduling of data flows, and monitoring the data flow status.

Benefits of technology

It realizes the synchronization of a full-scale medical data at a single time, reduces the complexity of data synchronization, improves synchronization efficiency and maintainability, and has the characteristics of being simple and easy to use and strong maintainability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114168678B_ABST
    Figure CN114168678B_ABST
Patent Text Reader

Abstract

The present invention discloses a full-volume medical data synchronization system and method, belonging to the technical field of data synchronization. The technical problem to be solved by the present invention is how to reduce the complexity of health care data synchronization, improve the efficiency and maintainability of data synchronization. The adopted technical solution is as follows: The system realizes the visual process orchestration and automated scheduling of data streams based on NIFI, and realizes the interaction between the data governance platform and NIFI through the combination of NIFI processors, exception monitoring processors and corresponding API interfaces, meeting the requirements of controlling the logic of data streams and monitoring the status of data streams, so as to realize the synchronization of single-time full-volume medical data; among them, the NIFI processor combination includes a process start processor, a text segmentation processor, a data query processor, a data storage processor and a process end processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data synchronization, and specifically to a full-volume medical data synchronization system and method. Background Art

[0002] To build a health care big data platform, it is necessary to integrate the data in the HIS, EMR, LIS and other systems of each medical institution. The ETL of these multi-source heterogeneous data from different manufacturers and different systems is a key technical problem to be solved.

[0003] Apache NiFi is an easy-to-use, powerful and reliable data pulling, data processing and distribution system. Apache NiFi is designed for data streams. It supports highly configurable data routing, transformation and system mediation logic of flowcharts, and supports dynamically pulling data from multiple data sources. NiFi was originally a project of the NSA and is currently open-source. It is one of the top-level projects of the Apache Foundation. NiFi is Java-based and uses Maven to support package build management. NiFi works based on the web mode and is scheduled on the server in the background. Users can define a process for data processing and then process it. The background has components such as a data processing engine and task scheduling.

[0004] However, NIFI lacks an interaction mechanism with external systems, and the data stream status cannot be fed back to the data governance platform, which does not meet the requirements of the data governance ETL processing logic.

[0005] Therefore, how to reduce the complexity of health care data synchronization, improve the efficiency and maintainability of data synchronization is a technical problem to be solved urgently at present. Summary of the Invention

[0006] The technical task of the present invention is to provide a full-volume medical data synchronization system and method to solve the problems of how to reduce the complexity of health care data synchronization, improve the efficiency and maintainability of data synchronization.

[0007] The technical task of the present invention is realized in the following way. A full-volume medical data synchronization system is based on NIFI to realize visual process orchestration and automated scheduling of data streams, and realizes the interaction between the data governance platform and NIFI through a combination of NIFI processors, an exception monitoring processor (Monitor processor) and corresponding API interfaces, meeting the requirements of controlling the data stream logic and monitoring the data stream status, so as to realize the synchronization of a single full-volume medical data;

[0008] Among them, the combination of NIFI processors includes

[0009] A Begin processor, which is used to notify the data governance platform that data synchronization has started;

[0010] A GenerateAndSplitText processor, which is used to batch process multiple SQLs according to a custom delimiter;

[0011] An ExecuteSQL processor, which is used to obtain data, execute an SQL query, and transform the data;

[0012] An IPutJdbc processor, which is used to store data into a target database;

[0013] An End processor, which is used to notify the data governance platform that the data synchronization process has ended.

[0014] Among them, NIFI is an easy-to-use, powerful, and reliable data ingestion, data processing, and distribution system for automating the management of data flows between systems. By integrating with NIFI, the data governance platform can achieve visual process orchestration and automated scheduling functions for data flows, thereby reducing the complexity of health care data synchronization and improving the efficiency and maintainability of data synchronization.

[0015] Preferably, the Begin processor includes:

[0016] A creation module, which is used to create a blank flow file and record the information required for the entire process into the metadata attributes of the flow file;

[0017] A notification module, which is used to call the notification interface URL of the data governance platform to notify downstream processors such as the GenerateAndSplitText processor to start data processing:

[0018] When the call to the notification interface URL of the data governance platform is successful, it feeds back a successful output;

[0019] When the call to the notification interface URL of the data governance platform fails, it feeds back a call failure message;

[0020] A transfer module, which is used to transfer metadata attribute information in a RestAPI manner after calling the notification interface URL of the data governance platform; among them, the information of the metadata attributes includes the process start time, process number, and process status information;

[0021] A stop module, which is used to stop its own processor after the transfer of the metadata attribute information is completed, and control the entire data flow to execute only once.

[0022] Preferably, the text segmentation processor (GenerateAndSplitText processor) includes:

[0023] A splitting module for receiving the stream file passed by the process start processor (Begin processor) through the event notification mechanism, starting data processing, and splitting multiple SQL texts configured by itself according to specified rules;

[0024] A writing module 1 for generating corresponding numbers of stream files according to the number of split SQLs. The newly generated stream files will inherit the attributes of the stream file passed by the process start processor (Begin processor), write the SQLs into the metadata attributes, and pass them to downstream processors such as the data query processor (ExecuteSQL processor), realizing that one data stream supports multiple data synchronization tasks.

[0025] Preferably, the data query processor (ExecuteSQL processor) includes:

[0026] A receiving module 1 for receiving the stream file passed by the text segmentation processor (GenerateAndSplitText processor) and obtaining the SQL information in the metadata attributes of the stream file;

[0027] A connection module for connecting to the original database using JDBC connection technology to perform data query operations;

[0028] A writing module 2 for writing the obtained data into the content of the stream file and passing the data information to downstream processors such as the data storage processor (IPutJdbc processor).

[0029] Preferably, the data storage processor (IPutJdbc processor) includes:

[0030] A receiving module 2 for receiving the stream file passed by the data query processor (ExecuteSQL processor);

[0031] An analysis module for parsing the data information contained in the stream file into SQL statements for data insertion;

[0032] An insertion module for connecting to the target database using JDBC technology, executing the SQL statement for data insertion, and inserting the data in the stream file into the target database.

[0033] Preferably, the process end processor (End processor) includes:

[0034] A receiving module 3 for receiving the stream file processed by the data storage processor (IPutJdbc processor) and performing summary statistics according to the metadata information in the stream file.

[0035] A calling module, which is used to call the configured notification interface in the way of RestAPI after determining that all data has been processed, and transfer the statistical information of the entire process processing to the governance platform, and inform the data governance platform that the current data synchronization process has ended; wherein, the statistical information of the entire process processing includes the specific information of the process end time, process number and the number of synchronized records.

[0036] A deletion module, which is used to automatically delete the data stream by the data governance platform calling an interface after the data synchronization is completed, saving system resources and ensuring the normal operation of the business system.

[0037] Preferably, the exception monitoring processor (Monitor processor) includes

[0038] A monitoring module, which is used to monitor the execution of the entire data synchronization process by circularly calling the Rest interface to obtain the information in the bulletin-board.

[0039] An exception detection module, which is used to call the notification interface to send information to the data governance platform after obtaining relevant information, and the data governance platform parses the log and detects exceptions, so as to perform classification processing for different situations.

[0040] A full-volume medical data synchronization method, and the method is as follows:

[0041] S1. The process start processor (Begin processor) is used as the process start node to notify the data governance platform that the data synchronization has started.

[0042] S2. The text segmentation processor (GenerateAndSplitText processor) receives the data stream passed by the process start processor (Begin processor) and configures multiple data mapping SQLs.

[0043] S3. The data query processor (ExecuteSQL processor) receives the stream file passed by the text segmentation processor (GenerateAndSplitText processor), obtains the stream file data and queries and converts the data.

[0044] S4. The data storage processor (IPutJdbc processor) receives the data of the data query processor (ExecuteSQL processor) and stores it in the target database.

[0045] S5. The process end processor (End processor) receives the stream file processed by the data storage processor (IPutJdbc processor), and judges whether the data synchronization is completed. After the data synchronization is completed, it calls the notification interface and notifies the end of the synchronization process.

[0046] S6. The anomaly monitoring processor (Monitor processor) monitors the execution of the entire process and passes the anomaly information to the data governance platform through an interface for subsequent processing.

[0047] Preferably, the working process of the process start processor (Begin processor) is as follows:

[0048] S101. Create a blank stream file and record the information required for the entire process in the metadata attributes of the stream file;

[0049] S102. Call the notification interface URL of the data governance platform to notify downstream processors such as the text segmentation processor (GenerateAndSplitText processor) to start data processing:

[0050] ①. When the call to the notification interface URL of the data governance platform is successful, feedback and output success;

[0051] ②. When the call to the notification interface URL of the data governance platform fails, feedback the call failure information;

[0052] S103. After calling the notification interface URL of the data governance platform, use the RestAPI method to transfer the metadata attribute information; among them, the information of the metadata attributes includes the process start time, process number, and process status information;

[0053] S104. After the transfer of the metadata attribute information is completed, stop its own processor to control the entire data stream to execute only once;

[0054] The working process of the text segmentation processor (GenerateAndSplitText processor) is as follows:

[0055] S201. Receive the stream file passed by the process start processor (Begin processor) through the event notification mechanism, start data processing, and split multiple SQL texts configured by itself according to the specified rules;

[0056] S202. Generate the corresponding number of stream files according to the number of split SQLs. The newly generated stream files will inherit the stream file attributes passed by the process start processor (Begin processor), write the SQLs into the metadata attributes, and transfer them to downstream processors such as the data query processor (ExecuteSQL processor) to implement a data stream supporting multiple data synchronization tasks;

[0057] The working process of the data query processor (ExecuteSQL processor) is as follows:

[0058] S301. Receive the stream file passed by the text splitting processor (GenerateAndSplitText processor) and obtain the SQL information in the stream file metadata attributes;

[0059] S302. Use JDBC connection technology to connect to the original database and perform data query operations;

[0060] S303. Write the obtained data into the stream file content and pass the data information to downstream processors such as the data storage processor (IPutJdbc processor).

[0061] More preferably, the working process of the data storage processor (IPutJdbc processor) is specifically as follows:

[0062] S401. Receive the stream file passed by the data query processor (ExecuteSQL processor);

[0063] S402. Parse the data information contained in the stream file into SQL statements for data insertion;

[0064] S403. Use JDBC technology to connect to the target database, execute the SQL statement for data insertion, and insert the data in the stream file into the target database;

[0065] The working process of the process end processor (End processor) is specifically as follows:

[0066] S501. Receive the stream file processed by the data storage processor (IPutJdbc processor) and perform summary statistics according to the metadata information in the stream file;

[0067] S502. After determining that all data has been processed, use the RestAPI method to call the configured notification interface, pass the entire process processing statistics information to the governance platform, and inform the data governance platform that the data synchronization process has ended; among them, the entire process processing statistics information includes specific information such as the process end time, process number, and number of synchronized items;

[0068] S503. After the data synchronization is completed, the data governance platform automatically deletes the data stream by calling the interface, saves system resources, and ensures the normal operation of the business system;

[0069] The working process of the exception monitoring processor (Monitor processor) is specifically as follows:

[0070] S601. Monitor the execution situation of the entire data synchronization process by circularly calling the Rest interface to obtain the information in the bulletin-board;

[0071] S602. After obtaining relevant information, call the notification interface to send the information to the data governance platform. The data governance platform parses the log and detects anomalies, and then classifies and processes different situations.

[0072] The full-volume medical data synchronization system and method of the present invention have the following advantages:

[0073] (1) Based on NIFI, the present invention realizes the visual process orchestration and automated scheduling functions of the data flow. By combining NIFI processors, developing status notification, exception monitoring processors and corresponding API interfaces, the interaction between the data governance platform and NIFI is realized, meeting the requirements of controlling the data flow logic and monitoring the data flow status, so as to realize the synchronization of single-time full-volume medical data.

[0074] (2) The present invention provides the visual process orchestration and automated scheduling functions of the data flow, meeting the requirements of monitoring and alarming the full life cycle operation of the data flow. It can realize the synchronization of single-time full-volume data, reduce the complexity of health care data synchronization, and improve the efficiency of data synchronization. It has the characteristics of being simple to use and easy to maintain.

[0075] (3) The present invention monitors the execution of the entire process through the Monitor processor, and transfers the exception information to the data governance platform through the interface for subsequent processing.

[0076] (4) After the Begin processor receives the information, the Begin processor stops itself to ensure that the entire data flow is executed only once, realizing the single-time operation of the NIFI data flow.

[0077] (5) The present invention can configure multiple data mapping SQLs through the GenerateAndSplitText processor to realize the function of completing multiple data synchronization tasks with one data flow.

[0078] (6) The present invention has the functions of start, end, and exception event notification, and can realize the joint control of the data synchronization logic with the data governance platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] The present invention will be further described below with reference to the accompanying drawings.

[0080] Attached Figure 1 is the structural block diagram of the full-volume medical data synchronization system;

[0081] Attached Figure 2 is the flow block diagram of the full-volume medical data synchronization method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0082] The following provides a detailed description of the full - volume medical data synchronization system and method of the present invention with reference to the accompanying drawings of the specification and specific embodiments.

[0083] Embodiment 1:

[0084] As shown Figure 1 in the figure, the full - volume medical data synchronization system of the present invention is based on NIFI to realize the visual process orchestration and automated scheduling of data streams, and realizes the interaction between the data governance platform and NIFI through the combination of NIFI processors, the exception monitoring processor (Monitor processor) and the corresponding API interfaces, meeting the requirements of controlling the data stream logic and monitoring the data stream status, so as to realize the synchronization of a single full - volume medical data;

[0085] Among them, the combination of NIFI processors includes,

[0086] The process start processor (Begin processor) is used to notify the data governance platform that the data synchronization has started;

[0087] The text segmentation processor (GenerateAndSplitText processor) is used to batch - process multiple SQLs according to a custom delimiter;

[0088] The data query processor (ExecuteSQL processor) is used to obtain data, execute SQL queries and convert data;

[0089] The data storage processor (IPutJdbc processor) is used to store data into the target database;

[0090] The process end processor (End processor) is used to notify the data governance platform that the data synchronization process has ended.

[0091] The process start processor (Begin processor) in this embodiment includes,

[0092] A creation module, which is used to create a blank flow file and record the information required for the entire process into the metadata attributes of the flow file;

[0093] A notification module, which is used to call the data governance platform notification interface URL to notify downstream processors such as the text segmentation processor (GenerateAndSplitText processor) to start data processing:

[0094] When the call to the data governance platform notification interface URL is successful, it feeds back and outputs success;

[0095] When the call to the data governance platform notification interface URL fails, it feeds back the call failure information;

[0096] A transfer module, which is used to transfer metadata attribute information in the RestAPI manner after calling the notification interface URL of the data governance platform; among them, the information of the metadata attributes includes the process start time, process number, and process status information.

[0097] A stop module, which is used to stop its own processor after the transfer of the metadata attribute information is completed, and control the entire data stream to execute only once.

[0098] The text splitting processor (GenerateAndSplitText processor) in this embodiment includes

[0099] A splitting module, which is used to receive the flow file passed by the process start processor (Begin processor) through the event notification mechanism, start data processing, and split multiple SQL texts configured by itself according to the specified rules.

[0100] A writing module 1, which is used to generate corresponding numbers of flow files according to the number of split SQLs. The newly generated flow files will inherit the attributes of the flow file passed by the process start processor (Begin processor), write the SQLs into the metadata attributes, and transfer them to downstream processors such as the data query processor (ExecuteSQL processor), so as to realize that one data stream supports multiple data synchronization tasks.

[0101] The data query processor (ExecuteSQL processor) in this embodiment includes

[0102] A receiving module 1, which is used to receive the flow file passed by the text splitting processor (GenerateAndSplitText processor) and obtain the SQL information in the metadata attributes of the flow file.

[0103] A connection module, which is used to connect to the original database using JDBC connection technology to perform data query operations.

[0104] A writing module 2, which is used to write the obtained data into the content of the flow file and transfer the data information to downstream processors such as the data warehousing processor (IPutJdbc processor).

[0105] The data warehousing processor (IPutJdbc processor) in this embodiment includes

[0106] A receiving module 2, which is used to receive the flow file passed by the data query processor (ExecuteSQL processor).

[0107] An analysis module, which is used to parse the data information contained in the flow file into SQL statements for data insertion.

[0108] An insertion module, which is used to connect to a target database using JDBC technology, execute an SQL statement for data insertion, and insert the data in the stream file into the target database.

[0109] The process end processor (End processor) in this embodiment includes

[0110] A third receiving module, which is used to receive the stream file processed by the data storage processor (IPutJdbc processor), and perform summary statistics according to the metadata information in the stream file;

[0111] A calling module, which is used to call the configured notification interface in the RestAPI manner after determining that all data has been processed, and transfer the statistical information of the entire process processing to the governance platform, and inform the data governance platform that the data synchronization process has ended; wherein, the statistical information of the entire process processing includes specific information such as the process end time, process number, and number of synchronized records;

[0112] A deletion module, which is used to automatically delete the data stream by the data governance platform through an interface call after data synchronization is completed, save system resources, and ensure the normal operation of the business system.

[0113] The exception monitoring processor (Monitor processor) in this embodiment includes

[0114] A monitoring module, which is used to monitor the execution of the entire data synchronization process by circularly calling the Rest interface to obtain information in the bulletin-board;

[0115] An exception detection module, which is used to call the notification interface to send information to the data governance platform after obtaining relevant information, and the data governance platform parses the log and detects exceptions, so as to perform classification processing for different situations.

[0116] Embodiment 2:

[0117] As shown in the appendix Figure 2 The full-volume medical data synchronization method of the present invention is as follows:

[0118] S1. The process start processor (Begin processor) serves as the process start node and notifies the data governance platform that data synchronization has started;

[0119] S2. The text segmentation processor (GenerateAndSplitText processor) receives the data stream passed by the process start processor (Begin processor) and configures multiple data mapping SQLs;

[0120] S3. The data query processor (ExecuteSQL processor) receives the stream file passed by the text segmentation processor (GenerateAndSplitText processor), obtains the stream file data, and queries and converts the data;

[0121] S4. The data storage processor (IPutJdbc processor) receives the data from the data query processor (ExecuteSQL processor) and stores it in the target database;

[0122] S5. The process end processor (End processor) receives the stream file processed by the data storage processor (IPutJdbc processor), and determines whether the data synchronization is completed. After the data synchronization is completed, it calls the notification interface and notifies the end of the synchronization process;

[0123] S6. The exception monitoring processor (Monitor processor) monitors the execution of the entire process, and passes the exception information to the data governance platform through the interface for subsequent processing.

[0124] The working process of the process start processor (Begin processor) in this embodiment is specifically as follows:

[0125] S101. Create a blank stream file, and record the information required for the entire process in the metadata attributes of the stream file;

[0126] S102. Call the notification interface URL of the data governance platform to notify downstream processors such as the text segmentation processor (GenerateAndSplitText processor) to start data processing:

[0127] ①. When the call to the notification interface URL of the data governance platform is successful, feedback and output success;

[0128] ②. When the call to the notification interface URL of the data governance platform fails, feedback the call failure information;

[0129] S103. After calling the notification interface URL of the data governance platform, use the RestAPI method to transfer the metadata attribute information; among them, the information of the metadata attributes includes the process start time, process number, and process status information;

[0130] S104. After the transfer of the metadata attribute information is completed, stop its own processor to control the entire data stream to execute only once;

[0131] The working process of the text segmentation processor (GenerateAndSplitText processor) in this embodiment is specifically as follows:

[0132] S201. Receive the flow file passed by the process start processor (Begin processor) through the event notification mechanism, start data processing, and split multiple SQL texts configured by itself according to specified rules;

[0133] S202. Generate flow files with the corresponding quantity according to the number of split SQLs. The newly generated flow files will inherit the attributes of the flow file passed by the process start processor (Begin processor), write the SQLs into the metadata attributes, and pass them to downstream processors such as the data query processor (ExecuteSQL processor), realizing that one data stream supports multiple data synchronization tasks;

[0134] The working process of the data query processor (ExecuteSQL processor) in this embodiment is specifically as follows:

[0135] S301. Receive the flow file passed by the text splitting processor (GenerateAndSplitText processor), and obtain the SQL information in the metadata attributes of the flow file;

[0136] S302. Use the JDBC connection technology to connect to the original database and execute data query operations;

[0137] S303. Write the obtained data into the content of the flow file, and pass the data information to downstream processors such as the data storage processor (IPutJdbc processor).

[0138] The working process of the data storage processor (IPutJdbc processor) in this embodiment is specifically as follows:

[0139] S401. Receive the flow file passed by the data query processor (ExecuteSQL processor);

[0140] S402. Parse the data information contained in the flow file into SQL statements for data insertion;

[0141] S403. Use the JDBC technology to connect to the target database, execute the SQL statement for data insertion, and insert the data in the flow file into the target database;

[0142] The working process of the process end processor (End processor) in this embodiment is specifically as follows:

[0143] S501. Receive the flow file processed by the data storage processor (IPutJdbc processor), and perform summary statistics according to the metadata information in the flow file;

[0144] S502. After determining that all data has been processed, use the RestAPI method to call the configured notification interface, and transfer the statistical information of the entire process to the governance platform, and inform the data governance platform that the data synchronization process has ended; among them, the statistical information of the entire process includes the specific information of the process end time, process number, and the number of synchronized records.

[0145] S503. After the data synchronization is completed, the data governance platform automatically deletes the data stream by calling the interface, saving system resources and ensuring the normal operation of the business system.

[0146] The working process of the exception monitoring processor (Monitor processor) in this embodiment is specifically as follows:

[0147] S601. Monitor the execution of the entire data synchronization process by repeatedly calling the Rest interface to obtain information in the bulletin-board.

[0148] S602. After obtaining the relevant information, call the notification interface to send the information to the data governance platform, and the data governance platform analyzes the log and detects exceptions, and then classifies and processes different situations.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A full - volume medical data synchronization system, characterized in that, The system realizes the visual process orchestration and automated scheduling of data streams based on NIFI, and realizes the interaction between the data governance platform and NIFI through the combination of NIFI processors, exception monitoring processors and corresponding API interfaces, meeting the requirements of controlling the logic of data streams and monitoring the status of data streams, so as to realize the synchronization of single-time full-volume medical data; Among them, the NIFI processor combination includes, A process start processor, which is used to notify the data governance platform that the data synchronization has started; A text segmentation processor, which is used to batch process multiple SQLs according to a custom delimiter; A data query processor, which is used to obtain data, execute SQL queries and convert data; A data storage processor, which is used to store data in the target database; A process end processor, which is used to notify the data governance platform that the data synchronization process has ended; Among them, the process start processor includes, A creation module, which is used to create a blank flow file and record the information required for the entire process into the metadata attributes of the flow file; A notification module, which is used to call the data governance platform notification interface URL to notify the text segmentation processor to start data processing: When the call to the data governance platform notification interface URL is successful, it will feedback and output success; When the call to the data governance platform notification interface URL fails, it will feedback the call failure information; A transfer module, which is used to transfer metadata attribute information in the RestAPI manner after calling the data governance platform notification interface URL; among them, the information of the metadata attributes includes the process start time, process number and process status information; A stop module, which is used to stop its own processor after the transfer of metadata attribute information is completed, controlling the entire data stream to execute only once; The text segmentation processor includes, A splitting module, which is used to receive the flow file passed by the process start processor through the event notification mechanism, start data processing, and split the configured multiple SQL texts according to the specified rules; A writing module 1, which is used to generate corresponding numbers of flow files according to the number of split SQLs. The newly generated flow files will inherit the flow file attributes passed by the process start processor, write the SQLs into the metadata attributes, and transfer them to the data query processor, realizing that one data stream supports multiple data synchronization tasks; The data query processor includes, A receiving module 1, which is used to receive the flow file passed by the text segmentation processor and obtain the SQL information in the flow file metadata attributes; A connection module, which is used to connect to the original database using JDBC connection technology to execute data query operations; A writing module 2, which is used to write the obtained data into the flow file content and transfer the data information to the data storage processor; The data storage processor includes, A receiving module 2, which is used to receive the flow file passed by the data query processor; An analysis module, which is used to parse the data information contained in the flow file into SQL statements for data insertion; An insertion module, which is used to connect to the target database using JDBC technology, execute the SQL statement for data insertion, and insert the data in the flow file into the target database; The process end processor includes, Receiving module three, which is used to receive the stream file processed by the data storage processor and perform summary statistics according to the metadata information in the stream file; Invocation module, which is used to, after determining that all data has been processed, call the configured notification interface in the RestAPI manner, transfer the entire process processing statistics information to the governance platform, and inform the data governance platform that the current data synchronization process has ended; among them, the entire process processing statistics information includes specific information such as the process end time, process number, and number of synchronized records; Deletion module, which is used to, after the data synchronization is completed, the data governance platform automatically deletes the data stream by calling the interface; The exception monitoring processor includes Monitoring module, which is used to monitor the execution of the entire data synchronization process by circularly calling the Rest interface to obtain information in the bulletin-board; Exception detection module, which is used to, after obtaining relevant information, call the notification interface to send information to the data governance platform, and the data governance platform parses the log and detects exceptions, so as to perform classification processing for different situations.

2. A full - volume medical data synchronization method, characterized in that, The method is specifically as follows: S1. The process start processor, as the process start node, notifies the data governance platform that the data synchronization has started; S2. The text segmentation processor receives the data stream passed by the process start processor and configures multiple data mapping SQLs; S3. The data query processor receives the stream file passed by the text segmentation processor, obtains the stream file data and queries and converts the data; S4. The data storage processor receives the data from the data query processor and stores it in the target database; S5. The process end processor receives the stream file processed by the data storage processor, determines whether the data synchronization is completed, and calls the notification interface after the data synchronization is completed, and notifies the synchronization process to end; S6. The exception monitoring processor monitors the execution of the entire process, and transfers the exception information to the data governance platform through the interface for subsequent processing; Among them, the working process of the process start processor is specifically as follows: S101. Create a blank stream file and record the information required for the entire process in the metadata attributes of the stream file; S102. Call the data governance platform notification interface URL to notify the text segmentation processor to start data processing: ①. When the call to the data governance platform notification interface URL is successful, then feedback and output success; ②. When the call to the data governance platform notification interface URL fails, then feedback the call failure information; S103. After calling the data governance platform notification interface URL, use the RestAPI method to transfer the metadata attribute information; among them, the information of the metadata attributes includes the process start time, process number, and process status information; S104. After the transfer of the metadata attribute information is completed, stop its own processor and control the entire data stream to execute only once; The working process of the text segmentation processor is specifically as follows: S201. Receive the stream file passed by the process start processor through the event notification mechanism, start data processing, and split the multiple SQL texts configured by itself according to the specified rules; S202. Generate stream files with the corresponding quantity according to the number of split SQLs. The newly generated stream files will inherit the attributes of the stream files passed by the process start processor, write the SQLs into the metadata attributes, and pass them to the data query processor, so as to support multiple data synchronization tasks with one data stream; The working process of the data query processor is specifically as follows: S301. Receive the stream files passed by the text segmentation processor and obtain the SQL information in the metadata attributes of the stream files; S302. Connect to the original database using the JDBC connection technology to perform data query operations; S303. Write the obtained data into the content of the stream file and pass the data information to the data storage processor; The working process of the data storage processor is specifically as follows: S401. Receive the stream files passed by the data query processor; S402. Parse the data information contained in the stream file into SQL statements for data insertion; S403. Connect to the target database using the JDBC technology, execute the SQL statements for data insertion, and insert the data in the stream file into the target database; The working process of the process end processor is specifically as follows: S501. Receive the stream files processed by the data storage processor and perform summary statistics according to the metadata information in the stream files; S502. After determining that all data has been processed, use the RestAPI method to call the configured notification interface, pass the processing statistics information of the entire process to the governance platform, and inform the data governance platform that the data synchronization process has ended; among them, the processing statistics information of the entire process includes specific information such as the process end time, process number, and number of synchronized records; S503. After the data synchronization is completed, the data governance platform automatically deletes the data stream by calling the interface; The working process of the exception monitoring processor is specifically as follows: S601. Monitor the execution status of the entire data synchronization process by repeatedly calling the Rest interface to obtain the information in the bulletin-board; S602. After obtaining the relevant information, call the notification interface to send the information to the data governance platform. The data governance platform parses the log and detects exceptions, and then classifies and processes different situations.

Citation Information

Patent Citations

  • A system and method for writing data into a graph database based on NiFi

    CN109376153A

  • Streaming data batch conversion method and system based on NiFi and state value thereof

    CN110647548A

  • Heterogeneous data real-time synchronization system and device

    CN112256796A