Data verification early warning implementation method for multiple heterogeneous data sources
Through the modularly designed data verification method, the data consistency of multiple heterogeneous data sources is automatically verified, which solves the data consistency and accuracy problems in software engineering, and realizes an efficient and flexible data check and early warning mechanism, which improves the scalability and reliability of the system.
Patent Information
- Application Number
- CN202510358898.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
In software engineering, there are difficulties in how to verify and ensure data consistency and accuracy between multiple heterogeneous data sources, especially when data transmission and storage between different systems, the existing technology has poor scalability and high maintenance costs.
Adopting a modular design, through the UI configuration detection task, the system automatically verifies data consistency, including data source custom module, verification rule definition module, early warning channel definition module, data loading module, verification task generation module, verification task scheduling module, verification task execution module, execution result analysis module and verification exception notification module, to realize an automated and intelligent data verification process.
Improve data consistency and accuracy, enhance system scalability, flexibility and reliability, reduce the risk of management and change, improve task execution efficiency and resource utilization, and ensure that problems can be quickly identified and resolved.
Smart Images

Figure CN120296800A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology software development, and particularly to a method for implementing data verification and warning for multiple heterogeneous data sources. Background Art
[0002] Currently, in the field of software engineering, the DDD model is usually adopted for software design and development to form various domain services. The communication between services is carried out through the RPC method, and the data storage methods and storage formats between different systems are different. For example, some use relational databases such as Mysql and Oracle for storage, while others use NoSql storage such as Redis, MongoDB, and ES. It is very difficult to verify and ensure the consistency and accuracy of business data between different systems. For example, for an order data and order details data, it is difficult to verify the consistency and accuracy of data in trading systems, financial systems, risk control systems, warehousing systems, logistics systems, membership systems, and statistical systems. Whether it is based on code verification mode or distributed transaction method, there are problems such as difficult implementation, poor scalability, and high maintenance costs. Summary of the Invention
[0003] The purpose of this application is to provide a method for implementing data verification and warning for multiple heterogeneous data sources. By configuring a series of detection tasks through the UI, the system automatically verifies the data according to business rules, discovers data inconsistencies in advance, and warns relevant personnel, thereby ensuring the consistency and accuracy of data to solve the problems in the background art.
[0004] A method for implementing data verification and warning for multiple heterogeneous data sources provided by this application adopts the following technical solutions, including: Data source custom module: used to manage and configure data sources, and ensure sufficient extensibility of the system by adding, deleting, and modifying data sources; Verification rule definition module: used to manage verification rules, and adapt to different verification scenarios by adding, deleting, and modifying verification rules; Warning channel definition module: used to manage warning channels, and adapt to different warning methods by configuring warning channels; Data loading module: used to load multiple data sources; Verification task generation module: used for the system to generate verification tasks according to the information submitted by users; Verification task scheduling module: used for the system to automatically build a cluster based on the master election algorithm for multiple nodes, select the master node, and schedule and dispatch tasks to appropriate nodes for execution in real time; Verification task execution module: used for the system to automatically build a cluster based on the master election algorithm for multiple nodes, and non-master nodes automatically become working nodes. The working nodes receive tasks dispatched by the master node and call the corresponding rule engine; Execution result analysis module: After the worker node finishes executing the tasks dispatched by the master node, it determines whether it passes according to the execution result. If it does not pass, it automatically analyzes the position of the data anomaly point; Verification task scheduling module: Used for the system to support multi-task combination and aggregation, construct DAG graph detection tasks, and the system automatically schedules these tasks; Verification exception notification module: After the worker node determines that the detection rule passes, it sends a detection exception notification and warning through the warning channel to assist developers in discovering problems and troubleshooting the system in a timely manner.
[0005] By adopting the above technical solutions, a series of detection tasks are configured by the UI, and the system automatically verifies the data according to business rules, discovers data inconsistencies in advance, and warns relevant personnel, thus ensuring the consistency and accuracy of the data.
[0006] Preferably, the data source in the data source customization module includes name, address, account, and type.
[0007] By adopting the above technical solutions, by supporting flexible management of data sources, the system can adapt to a variety of different data platforms and storage requirements, providing high scalability and flexibility. The centralized data source configuration management simplifies the system integration, maintenance, and operation and maintenance processes, improves the reliability, performance, and resource utilization rate of the system, and at the same time reduces the risks of management and change.
[0008] Preferably, the verification rules in the verification rule definition module include a rule execution engine and rule content.
[0009] By adopting the above technical solutions, by supporting flexible management of verification rules and adapting to multiple rule engines, the system can select the most suitable verification execution method according to different business scenarios, enhancing the flexibility, execution efficiency, and scalability of the verification rules. The flexible rule configuration and management also improve the adaptability and maintenance convenience of the system, can quickly respond to changing requirements, and enhance the performance and operability of the overall verification process.
[0010] Preferably, the warning channels in the warning channel definition module include DingTalk robots, enterprise WeChat, email notifications, and custom URL callbacks.
[0011] By adopting the above technical solutions, managing warning channels and supporting multiple warning methods and custom configurations not only enhances the flexibility and adaptability of the system, but also improves the transmission efficiency and accuracy of warning information. This enables the system to better serve different teams and scenarios, optimize the team's response ability, and ensure that key issues can be quickly identified and resolved.
[0012] Preferably, the data loading module supports multiple data loading methods, including HTTP data loader, WEBSOCKET data loader, MYSQL data loader, MOGODB loader, and ES data loader.
[0013] By adopting the above technical solution, the support for multiple data loading methods enables the system to adapt to different data sources and scenarios, improving the flexibility, real-time performance, and efficiency of data loading. Whether it is batch data processing, real-time data streams, or cross-system data integration, the system can select the most suitable data loading method according to actual needs. This not only enhances the compatibility, scalability, and fault tolerance of the system but also improves the overall performance and resource utilization rate, ensuring that the system can efficiently and stably process various types of data loading tasks.
[0014] Preferably, the user-submitted information in the verification task generation module includes task configuration, data source, verification rules, and scheduling rules.
[0015] By adopting the above technical solution, by automatically generating verification tasks based on information such as task configuration, data source, verification rules, and scheduling rules submitted by users, the system can improve the efficiency of task creation, ensure the standardization and consistency of task execution, and reduce manual intervention and configuration errors. This process enhances the flexibility, traceability, and adaptability of task management, ensuring that tasks can be executed at the right time, place, and in the correct manner, improving the overall task execution efficiency and the optimal use of system resources.
[0016] Preferably, in the verification task scheduling module, the master node is responsible for constructing a task trigger according to the task scheduling rules and then real-time scheduling and dispatching tasks to appropriate nodes for execution.
[0017] By adopting the above technical solution, through automated and flexible scheduling rules and real-time task dispatching, it is possible to improve the efficiency, stability, and manageability of task execution, optimize resource allocation and load balancing, enhance the system's response speed and scalability, and make the entire system operate more efficiently and reliably, quickly adapting to changes in business requirements.
[0018] Preferably, it improves the flexibility, efficiency, resource utilization rate, and scalability of task execution, simplifies the task configuration and management process, enabling the entire system to execute complex verification tasks more efficiently and adapt to changing requirements.
[0019] By adopting the above technical solution, non-master nodes can automatically become working nodes and call the corresponding rule engine.
[0020] Preferably, the task configuration execution method includes the following steps: S1. The user configures tasks through the UI interface, submits task names, data sources, data query parameters, execution frequencies, time offsets, overdue warning times, warning channels, verification rules, remarks and other parameters to generate detection tasks; S2. After the detection task is generated, the cluster master node automatically schedules according to the task configuration, and distributes the task to the appropriate cluster working node for execution according to the round-robin load balancing algorithm at an appropriate time; S3. After receiving the execution command from the master node, the cluster working node performs verification operations according to the task configuration. If the verification result fails, it sends a notification message according to the task warning channel, and the system personnel handle it after receiving the notification.
[0021] By adopting the above technical solutions, it helps to build an efficient, automated and reliable data proofreading and verification mechanism, optimize the data processing process, and be able to respond quickly when problems occur.
[0022] Preferably, the data verification method includes the following steps: S10. The system loads and obtains data from multiple data sources to generate proofreading parameters; S20. The system verification module loads the proofreading parameters and executes the verification operation by loading the rule engine according to the verification rules configured in the task; S30. The system sends a notification to the system personnel through the warning channel configured in the task according to the verification operation result if it fails.
[0023] By adopting the above technical solutions, an efficient, reliable and flexible task management and scheduling system is realized, which can automatically and intelligently complete task execution and monitoring, improve system performance and quickly respond to abnormal situations.
[0024] In summary, the present application includes at least one of the following beneficial technical effects: The method of the present application adopts a modular design, including a data source custom module, a verification rule definition module, a warning channel definition module, a data loading module, a verification task generation module, a verification task scheduling module, a verification task execution module, an execution result analysis module, a verification task orchestration module, a verification exception notification module, etc. A series of detection tasks are configured by the UI, and the system automatically verifies the data according to business rules, discovers data inconsistencies in advance, and warns relevant personnel, thereby ensuring data consistency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a schematic diagram of the system modules in the embodiment of the present application; Figure 2 is a flow chart of the configuration task in the embodiment of the present application; Figure 3It is the data task verification flowchart in the embodiments of this application. Specific embodiments
[0026] The following will Figure 1 - Appendix Figure 3 be used to further elaborate on this application in detail.
[0027] A method for realizing data verification and warning of multiple heterogeneous data sources, referring to Figure 1 , includes: Data source customization module: used to manage and configure data sources. By adding, deleting, and modifying data sources, the system can be fully extended. The data source includes information such as name, address, account, type, etc., such as MYSQL ORACLE databases, MONGODBES, etc.; Verification rule definition module: used to manage verification rules. By adding, deleting, and modifying verification rules, different verification scenarios can be adapted. The verification rules include rule execution engines, rule content, etc. The system supports AVIATOR rule engines, V8 rule engines, GROOVY rule engines, JAVA class rule engines, etc.; Warning channel definition module: used to manage warning channels. By configuring warning channels, different warning methods can be adapted. The warning channels include DingTalk robots, enterprise WeChat, email notifications, custom URL callbacks, etc.; Data loading module: used to load multiple data sources. The data loading module supports multiple data loading methods, including HTTP data loaders, WEBSOCKET data loaders, MYSQL data loaders, MOGODB loaders, ES data loaders; Verification task generation module: used for the system to generate verification tasks according to the information submitted by users. The information submitted by users includes task configuration, data sources, verification rules, and scheduling rules; Verification task scheduling module: used for the system to automatically build a cluster based on the master election algorithm on multiple nodes and elect a master node. The master node is responsible for building a task trigger according to the task scheduling rules and then scheduling and dispatching tasks to appropriate nodes for execution in real time; Verification task execution module: used for the system to automatically build a cluster based on the master election algorithm on multiple nodes. Non-master nodes automatically become worker nodes. Worker nodes receive tasks dispatched by the master node and call corresponding rule engines according to the type of execution engine and rule content; Execution result analysis module: used for the worker nodes to determine whether the task dispatched by the master node is passed according to the execution result after the task is completed. If it fails, the position of the data anomaly point is automatically analyzed; Verification task orchestration module: used for the system to support multi-task combination and aggregation, build a DAG graph detection task, and the system automatically orchestrates this task; Verification exception notification module: After determining that the detection rule passes, the working node sends a detection exception notification warning through the warning channel to assist developers in promptly discovering problems and troubleshooting the system.
[0028] This application adopts a modular design, enabling the system to achieve a high degree of automation, intelligence, and flexibility, optimizing various links such as task generation, scheduling, execution, result analysis, and warning notification. The system has strong scalability, can adapt to diverse requirements, and can quickly respond, locate, and handle exceptions during task execution, ensuring the stable operation of the system and the continuity of business.
[0029] A method for realizing data verification and warning of multiple heterogeneous data sources, referring to Figure 2 , and its task configuration and execution method includes the following steps: S1. The user configures tasks through the UI interface, submits task names, data sources, data query parameters, execution frequencies, time offsets, overdue warning times, warning channels, verification rules, remarks, and other parameters to generate detection tasks; S2. After the detection task is generated, the cluster master node automatically schedules according to the task configuration and distributes the task to the appropriate cluster working node for execution at an appropriate time according to the round-robin load balancing algorithm; S3. After receiving the execution command from the master node, the cluster working node performs verification operations according to the task configuration. If the verification result fails, it sends a notification message according to the task warning channel, and the system personnel handle it after receiving the notification.
[0030] With the above task configuration execution method, users can flexibly configure task parameters through the UI interface, such as task name, data source, query parameters, execution frequency, etc. This configuration method makes the definition of tasks more simple and intuitive, and can quickly adjust and generate adaptive detection tasks according to different requirements; the cluster master node can automatically schedule tasks according to the task configuration, and dispatch tasks to appropriate cluster worker nodes through the round-robin load balancing algorithm; this can effectively improve the efficiency of task execution, avoid overloading or idling of some nodes in the cluster, and at the same time ensure the balanced distribution of tasks and enhance the overall processing capacity of the system; the cluster worker nodes can process multiple tasks in parallel, reducing the time for task execution; through the distributed computing architecture, the cooperation of multiple nodes can significantly improve the processing efficiency, especially when facing a large amount of data and frequent tasks, it can accelerate the completion of the verification operation; when the verification operation fails, the system will automatically send a notice according to the configured warning channel; this automated warning mechanism can ensure that problems are discovered and reported to relevant personnel at the first time of occurrence, reduce the response time, and improve the efficiency of problem solving; Task traceability and manageability: The tasks submitted through the UI interface include detailed configuration parameters (such as task name, execution frequency, etc.), which are convenient for later management, tracking, and modification; at the same time, the execution process and results of tasks can be logged according to preset rules, enhancing the traceability and transparency of task execution; Enhance the stability and reliability of the system: The automatic scheduling and real-time warning feedback mechanism can detect and handle anomalies in a timely manner, reducing manual intervention and delays; coupled with the distributed processing of the cluster, it can effectively improve the stability and fault tolerance of the system, and avoid the failure of the entire task caused by a single node failure; Improve resource utilization: Through the dynamic scheduling and load balancing algorithm of the cluster, the utilization efficiency of resources is maximized; the cluster worker nodes can automatically allocate tasks according to the current load, avoiding resource waste or imbalance, and improving the overall efficiency of the system.
[0031] A method for realizing data verification and warning of multiple heterogeneous data sources, referring to Figure 3 , and its data verification method includes the following steps: S10. The system loads and obtains data from multiple data sources to generate calibration parameters; S20. The system verification module loads the calibration parameters and executes the verification operation by loading the rule engine according to the verification rules configured for the task; S30. The system, according to the result of the verification operation, if it fails, sends a notice to the system personnel through the warning channel configured for the task.
[0032] With the above data verification method, the system can obtain and generate verification parameters from multiple data sources to ensure the accuracy and consistency of data. Through the verification and validation operations of the verification parameters, potential errors or inconsistencies can be detected before the data enters the subsequent processes, thereby improving the data quality. Moreover, by loading the verification parameters and executing the verification rules, the system can automatically perform verification, reducing the workload of manual inspection, saving time, and avoiding human errors. This automation can greatly improve the processing efficiency, especially when dealing with a large amount of data. The system has set up a warning channel, and once the verification fails, it can notify the relevant personnel in a timely manner. This instant feedback mechanism can help the team quickly locate and correct problems, preventing data errors from affecting the normal operation or decision-making of the system. By managing the verification rules and warning channels through task configuration, the verification parameters and notification methods can be flexibly adjusted, and customized configurations can be made according to the requirements of different tasks, enhancing the adaptability and operability of the system. Through the verification operations of the rule engine, the entry of incorrect data into the system can be effectively reduced, the system failures or delays caused by inaccurate data can be decreased, and the stability and reliability of the entire system can be improved.
[0033] The data verification algorithm of a method for realizing data verification and warning of multiple heterogeneous data sources includes the following: 1) In the custom data source 1, load and obtain data according to the task configuration of the data source and data query parameters (supporting dynamic conditions) to form dataset A. For example, calculate the start time as [system current time - time offset - execution frequency - 10S] and the end time as [system current time - time offset] according to the configuration, and create order data within this time period for the order table in the transaction database. 2) Load and obtain data from the custom data source 2 according to the task configuration of the data source and data query parameters (supporting dynamic conditions) or the data in dataset A to form dataset B. For example, obtain the payment records in the financial system (dataset B) according to the order numbers in the transaction system order records (dataset A). 3) Verify whether various information in dataset A and dataset B is consistent and accurate. For example, the order number must conform to the rules defined by the business, there must be corresponding financial payment records for the order in the paid status, and the order amount must be equal to the financial payment amount, etc. 4) The system supports verifying the business data within a specified time period. After submitting the time period detection record, the platform automatically and asynchronously executes the data verification for this time period. When the verification fails, the abnormal data is automatically sent to the warning channel to notify the system personnel.
[0034] The present invention provides a method for realizing data verification and warning of multiple heterogeneous data sources, which can achieve the detection and warning of data consistency in multiple systems, and also realizes the query of multiple heterogeneous dynamic data sources. It can also implement multiple rule engines to execute custom verification rules and support the detection of specified time periods and real-time data.
[0035] The embodiments of the specific implementation manners are all preferred embodiments of the present application, and do not limit the protection scope of the present application thereby. The same components are denoted by the same reference numerals. Therefore, all equivalent changes made according to the structure, shape and principle of the present application shall be covered within the protection scope of the present application.
Claims
1. A method for realizing data verification and early warning of multiple heterogeneous data sources, characterized in that Including: Data source customization module: Used to manage and configure data sources. By adding, deleting, and modifying data sources, it ensures sufficient extensibility of the system; Verification rule definition module: Used to manage verification rules. By adding, deleting, and modifying verification rules, it adapts to different verification scenarios; Early warning channel definition module: Used to manage early warning channels. By configuring early warning channels, it adapts to different early warning methods; Data loading module: Used to load multiple data sources; Verification task generation module: Used for the system to generate verification tasks based on the information submitted by users; Verification task scheduling module: Used for the system to automatically build a cluster based on the master election algorithm on multiple nodes, elect a master node, and schedule and dispatch tasks to appropriate nodes for execution in real time; Verification task execution module: Used for the system to automatically build a cluster based on the master election algorithm on multiple nodes. Non-master nodes automatically become worker nodes, and worker nodes receive tasks dispatched by the master node and call the corresponding rule engine; Execution result analysis module: Used for the worker node to determine whether it passes according to the execution result after completing the tasks dispatched by the master node. If it does not pass, it automatically analyzes the location of data anomalies; Verification task orchestration module: Used for the system to support multi-task combination and aggregation, build DAG graph detection tasks, and the system automatically orchestrates these tasks; Verification exception notification module: After the worker node determines that the detection rule passes, it sends a detection exception notification and warning through the early warning channel to assist developers in discovering problems and troubleshooting the system in a timely manner.
2. The method for realizing data verification and warning of multiple heterogeneous data sources according to claim 1, characterized in that, In the data source customization module, the data source includes name, address, account, and type.
3. A method for implementing data verification and warning of multiple heterogeneous data sources according to claim 1, characterized in that In the verification rule definition module, the verification rule includes a rule execution engine and rule content.
4. A method for implementing data verification and warning of multiple heterogeneous data sources according to claim 1, characterized in that, In the early warning channel definition module, the early warning channel includes DingTalk robot, WeCom, email notification, and custom URL callback.
5. A method for implementing data verification and warning of multiple heterogeneous data sources according to claim 1, characterized in that, The data loading module supports multiple data loading methods, including HTTP data loader, WEBSOCKET data loader, MYSQL data loader, MOGODB loader, and ES data loader.
6. A method for implementing data verification and warning of multiple heterogeneous data sources according to claim 1, characterized in that In the verification task generation module, the information submitted by users includes task configuration, data source, verification rule, and scheduling rule.
7. A method for implementing data verification and warning of multiple heterogeneous data sources according to claim 1, characterized in that In the verification task scheduling module, the master node is responsible for building a task trigger according to the task scheduling rule and then scheduling and dispatching tasks to appropriate nodes for execution in real time.
8. A method for implementing data verification and warning of multiple heterogeneous data sources according to claim 1, characterized in that, In the verification task execution module, after the worker node receives the task dispatched by the master node, it calls the corresponding rule engine according to the execution engine type and rule content.
9. A method for realizing data verification and early warning of multiple heterogeneous data sources according to claim 1, characterized in that The method for executing the task configuration includes the following steps: S1. The user configures the task through the UI interface, submits parameters such as task name, data source, data query parameters, execution frequency, time offset, overdue warning time, warning channel, verification rule, and remarks to generate a detection task; S2. After the detection task is generated, the cluster master node automatically schedules according to the task configuration and dispatches the task to an appropriate cluster worker node for execution at an appropriate time according to the round-robin load balancing algorithm; S3. After receiving the execution command from the master node, the cluster worker node performs verification operations according to the task configuration. If the verification result fails, it sends a notification message according to the task warning channel, and the system personnel handle it after receiving the notification.
10. A method for implementing data verification and warning of multiple heterogeneous data sources according to claim 1, characterized in that, The data verification method includes the following steps: S10. The system loads and obtains data from multiple data sources to generate verification parameters; S20. The system verification module loads the verification parameters and executes the verification operation by the rule engine according to the verification rules configured for the task; S30. The system, according to the result of the verification operation, if it fails, sends a notification to the system personnel through the warning channel configured for the task.