Data cleaning method and device based on Drools rule engine and electronic equipment

By adopting the data cleaning method based on the Drools rule engine in the data integration process, problems such as insufficient flexibility and low automation in the existing technology are solved, and efficient, flexible and scalable data cleaning functions are realized, improving data quality and automation.

CN120196623APending Publication Date: 2025-06-24CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510322892.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing data integration and data cleaning technologies have problems such as insufficient flexibility, low degree of automation, poor integration, and challenges in processing efficiency and real-time.

Method used

Using the data cleaning method based on the Drools rule engine, the dynamic configuration and management of cleaning rules is realized through the Drools rule engine, the data synchronization framework is integrated, and an efficient, flexible and scalable data cleaning system is built.

Benefits of technology

It improves the flexibility, efficiency and maintainability of data processing, realizes efficient, flexible and scalable data cleaning functions, improves data quality and automation, and meets the needs of real-time and high throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196623A_ABST
    Figure CN120196623A_ABST
Patent Text Reader

Abstract

The invention relates to a data cleaning method and device based on a Drools rule engine. The method comprises the following steps: acquiring original data from various data sources through a data acquisition module; converting the collected original data into a data format which can be processed by a Drools rule engine, and inserting the data format into a working memory; the data cleaning module carries out cleaning processing on the data according to a predefined cleaning rule, wherein the cleaning processing comprises operations such as duplicate removal, missing value filling and format correction; outputting the cleaned data to a target system through a data output module; and recording a cleaning log through the monitoring module, counting a cleaning result, and providing visual display. According to the scheme, a Drools rule engine is adopted, dynamic loading and updating rules are supported, and the dynamic loading and updating rules can take effect without restarting a system; the flexibility is high, and updating of various data sources and dynamic rules is supported; the cleaning efficiency is high, and large-scale data can be rapidly processed through the rule engine; the monitoring function is complete, real-time log recording and visual display are provided, and the cleaning process is convenient to manage and optimize.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data cleaning, and particularly relates to a data cleaning method, device, computer-readable storage medium, and electronic device in the data integration process based on the Drools rule engine. Background Art

[0002] Data integration is a key link in data processing, and its core goal is to integrate data from different sources to support applications such as data analysis and decision-making. With the continuous expansion of the data scale and the increasing diversification of data sources, data integration and data cleaning technologies play a crucial role in modern data processing. However, the existing data integration and data cleaning technologies still have many defects and deficiencies, mainly reflected in the following aspects: (1) Poor flexibility and scalability When dealing with multi-source heterogeneous data, the format, standard, and semantic differences between different data sources are significant, which increases the difficulty of data cleaning. The existing cleaning methods are difficult to dynamically adjust rules and cannot adapt to complex and changeable business requirements.

[0003] (2) Insufficient automation Although automated data cleaning technologies have made certain developments, the existing tools still highly rely on manual intervention when dealing with complex data problems. In addition, the existing tools are insufficient in dealing with unstructured data (such as text), and it is difficult to achieve fully automated data cleaning.

[0004] (3) Low integration degree of tools and platforms The integration degree of data cleaning tools and data integration platforms is not high, resulting in a disconnection between the data cleaning process and the data integration process. For example, although many ETL tools (such as Kettle) provide data cleaning functions, their functions are still limited when dealing with complex data cleaning tasks. In addition, the integration of data cleaning tools with other big data processing platforms (such as data warehouses, data lakes) is not tight enough, affecting the overall efficiency of data processing.

[0005] (4) Challenges in processing efficiency and real-time performance In the big data environment, the data volume is huge, and traditional data cleaning methods are difficult to meet the requirements in terms of processing speed and real-time performance. Although distributed computing frameworks (such as Hadoop, Spark) can improve processing efficiency, when facing real-time data streams, the latency problem of data cleaning still exists. Especially in the online process control scenario, the real-time requirement of data cleaning is relatively high, and the existing technologies often cannot meet it.

[0006] (5) High time and cost investment Data cleaning is a time-consuming and costly process, especially when dealing with large volumes of data of poor quality, which requires a significant amount of time and resources for cleaning. Additionally, the effectiveness of data cleaning is often difficult to quantify, resulting in a disproportionate investment-return ratio for enterprises in data cleaning. Summary of the Invention

[0007] To address the above issues, this application proposes a new data cleaning method based on the Drools rule engine. This solution realizes the dynamic configuration and management of cleaning rules through the Drools rule engine, enabling data cleaning during the data integration process and improving the flexibility, efficiency, and maintainability of data processing.

[0008] This solution aims to achieve efficient, flexible, and scalable data cleaning functions through the powerful rule definition and processing capabilities of the rule engine. By integrating the Drools rule engine with the data synchronization framework, this solution constructs a complete data cleaning system suitable for the cleaning requirements of multi-source heterogeneous data, which can effectively improve data quality and provide support for subsequent data analysis and decision-making. The overall concept and main implementation process of this solution are as follows: 1. System Architecture Design This solution designs a data cleaning method and device applicable to data integration. The overall architecture includes a data collection module, a Drools rule engine, a data output module, and a monitoring module. These modules cooperate closely to form an efficient, automated, and monitorable data cleaning process.

[0009] 2. Data Cleaning Rule Management This solution uses the Drools rule engine as the core component, responsible for the definition, storage, loading, update, and execution of data cleaning rules. The Drools rule engine supports rule definition based on the declarative rule language (DRL), allowing users to write cleaning rules in an intuitive and flexible manner. The rules are stored in a rule library and can be dynamically loaded and updated, enabling rapid adjustment and optimization of cleaning rules to adapt to complex and changing business requirements.

[0010] 3. Data Collection and Input This solution collects raw data from various data sources (such as relational databases, file systems, message queues, etc.) through the data collection module and transfers it to the Drools rule engine. The data collection module supports multiple data formats and protocols and can efficiently obtain data from heterogeneous data sources. The collected raw data is preliminarily parsed and formatted and then input into the Drools rule engine as the basic data for subsequent cleaning operations.

[0011] 4. Data Cleaning Execution After receiving the input data, the Drools rule engine cleans the data according to predefined cleaning rules. The cleaning rules cover various operations such as data verification, duplicate removal, format conversion, outlier handling, and missing value filling. Through an efficient rule matching and execution mechanism, the rule engine quickly identifies and corrects quality problems in the data, ensuring the accuracy and consistency of the output data. The cleaning process supports parallel processing and distributed computing, which can effectively improve the efficiency of data cleaning and meet the requirements of large-scale data processing.

[0012] 5. Data Output and Storage The cleaned data is transmitted to the target end through the data output module, such as a data warehouse, a data lake, or other downstream application systems. The data output module supports multiple output formats and storage methods, and can flexibly adapt to different business scenarios. At the same time, the output module also provides data verification and integrity check functions to ensure that the cleaned data is accurately stored in the target location.

[0013] 6. Monitoring and Visualization The system is equipped with a monitoring module to implement log recording, metric statistics, alarm, and visualization display of the data cleaning process. The monitoring module records key information in the data cleaning process in real time, including the amount of data processed, the execution of cleaning rules, and the statistics of abnormal data. Through real-time monitoring and analysis of these metrics, problems in the data cleaning process can be discovered in a timely manner, and the alarm mechanism can be triggered to remind the operation and maintenance personnel to handle them. In addition, the monitoring module also provides a visualization interface to intuitively display the effect and performance metrics of data cleaning in the form of charts and reports, facilitating users to manage and optimize the cleaning process.

[0014] In the implementation process of this solution, the following design key points are included: (1) Dynamic Rule Loading and Execution During the data synchronization process, this solution can dynamically load Drools rules to implement data cleaning. The Drools rule engine automatically matches and executes corresponding cleaning rules according to data characteristics, without restarting the system or interrupting the data processing flow, thus ensuring the flexibility and real-time nature of data cleaning.

[0015] (2) Seamless Integration Architecture This solution designs a lightweight integration architecture, integrating the Drools rule engine as a plugin or extension module of DataX (or other data synchronization frameworks). This architecture does not require additional complex configurations, and realizes efficient communication and data exchange between Drools and the data synchronization framework through APIs, ensuring the close coordination between the data cleaning function and the data synchronization process.

[0016] (3) Multi-dimensional Data Verification In the process of data cleaning, this solution combines the rule definition ability of the Drools rule engine and the data transmission ability of DataX to achieve multi-dimensional data cleaning functions, including but not limited to deduplication, integrity verification, format correction, etc. This solution supports users to customize cleaning rules and automatically records cleaning results and cleaning logs for subsequent monitoring and optimization.

[0017] (4)Scalability and Flexibility This solution supports users to customize Drools rules to meet the data cleaning requirements in different business scenarios. Through modular design, users can easily expand and integrate other data processing tools to further enhance the functions and applicability of the system.

[0018] Specifically, this application provides the following technical solutions: In the first aspect of this application, a data cleaning method based on the Drools rule engine is provided. The method includes: S1. Collect raw data from the data source through the data collection module; S2. Convert the collected raw data into a data format that can be processed by the Drools rule engine and insert it into the working memory of the Drools rule engine; S3. The data cleaning module cleans the data according to the predefined cleaning rules; S4. Output the cleaned data to the target system through the data output module; S5. Record the cleaning log, count the cleaning results, and provide visual display through the monitoring module.

[0019] Further, the data collection module in step S1 of the method of this application supports collecting data from relational databases, Kafka, file systems, message queues, and API interfaces; The target system in step S4 includes but is not limited to: data warehouses, data lakes, or downstream application systems.

[0020] Further, the cleaning process in step S3 of the method of this application includes but is not limited to the following operations: Data verification: Identify data that does not meet the requirements through format verification, range verification, and type verification; Deduplication: Delete duplicate records by comparing specific fields (such as id, name, etc.) of the records; Missing value filling: Fill in missing fields with the mean or default value according to the business logic; Outlier handling: Identify extreme values that do not conform to the normal distribution or business logic, and perform deletion, correction, or marking on the outliers; Format conversion: Convert the data format according to the business requirements; Format correction: Correct data that does not conform to the preset format.

[0021] Furthermore, step S3 of the method of the present application further includes: (1) Define a data model According to the business data to be cleaned, a dynamic data model is constructed using the Map collection in Java; (2) Load data The original data obtained by the data acquisition module is encapsulated into a Map object dynamicObject corresponding to the dynamic data model; then, by calling the kieSession.insert(dynamicObject) method, the dynamicObject object is inserted into the working memory of Drools; (3) Execute rules When it is necessary to execute the data cleaning rules, the kieSession.fireAllRules() method is called to trigger all predefined rules; after the rules are executed, the kieSession.dispose() method is called to release resources.

[0022] Furthermore, the predefined rules described in step S3 of the method of the present application are stored in a database, supporting user-defined and real-time updates to adapt to the data cleaning requirements in different business scenarios.

[0023] Furthermore, the monitoring module described in step S5 of the method of the present application has the following functions: Log recording function: Record the timestamp, rule name, and cleaning status information of data cleaning; Index statistics function: Statistically calculate the cleaning success rate, failure rate, number of times the cleaning rule is triggered, execution time, and usage of system resources; Alarm function: Send an alarm notification when the data cleaning failure rate exceeds the preset threshold, a serious error occurs during the cleaning process, the system resource usage rate exceeds the safety threshold, or the data cleaning task times out and is not completed; Visualization function: Display the data cleaning progress, rule execution status, and key indicators through a visualization interface.

[0024] The second aspect of the present application provides a data cleaning device based on the Drools rule engine, and the device includes: Data acquisition module: Used to collect original data from a data source and convert the collected original data into a data format that can be processed by the Drools rule engine; Drools rule engine module: Used to define, store, load, update, and execute data cleaning rules; Data cleaning module: used to receive raw data and perform data cleaning operations through the Drools rule engine; Data output module: used to output the cleaned data to the target system; Monitoring module: used to monitor the data cleaning process, record cleaning logs, count cleaning results, and provide visual display.

[0025] When the device runs, it implements the steps of the aforementioned data cleaning method based on the Drools rule engine.

[0026] Furthermore, in the device of the present application, the Drools rule engine module is integrated as a plugin or extension module of DataX (or other data synchronization frameworks), and efficient communication and data exchange between the Drools rule engine module and DataX are achieved through the API.

[0027] Furthermore, the working process of the data cleaning module in the device of the present application includes: (1) Define the data model According to the business data to be cleaned, a dynamic data model is constructed using the Map collection in Java; (2) Load data The raw data obtained by the data acquisition module is encapsulated into a Map object dynamicObject corresponding to the dynamic data model; then, by calling the kieSession.insert(dynamicObject) method, the dynamicObject object is inserted into the working memory of Drools; (3) Execute rules When it is necessary to execute the data cleaning rules, the kieSession.fireAllRules() method is called to trigger all predefined rules; after the rules are executed, the kieSession.dispose() method is called to release resources.

[0028] The third aspect of the present application provides an electronic device, including: a memory and a processor; Memory: used to store computer programs; Processor: used to execute the computer program to implement the steps of the aforementioned data cleaning method based on the Drools rule engine.

[0029] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the aforementioned data cleaning method based on the Drools rule engine.

[0030] In summary, by integrating the Drools rule engine with the data synchronization framework, this solution has achieved efficient, flexible, and scalable data cleaning capabilities, effectively addressing the problems of insufficient flexibility, low automation, and poor integration in traditional data cleaning technologies, providing reliable technical support for modern data processing. This solution has the following significant advantages: (1) Enhanced flexibility This solution uses the Drools rule engine to define data cleaning rules, supporting dynamic loading and updating of rules, which take effect without restarting the system. This feature enables the system to quickly adapt to complex and changing business requirements, significantly enhancing the flexibility of data cleaning.

[0031] (2) Improved cleaning efficiency The Drools rule engine has efficient execution capabilities, which can significantly improve the speed of data cleaning, especially suitable for large-scale data scenarios. In addition, through parallel processing and distributed computing technologies, combined with the efficient rule matching mechanism of the Drools rule engine, this solution can quickly process large-scale data, effectively enhancing the overall efficiency of data integration and meeting the requirements of real-time and high throughput.

[0032] (3) Enhanced data quality With the precise rule definition of the Drools rule engine, this solution can effectively identify and eliminate errors, inconsistencies, and duplicates in the data. At the same time, it supports cleaning real-time data streams to ensure that the data meets high-quality standards before entering the target system, thereby enhancing the overall quality of the data.

[0033] (4) Increased automation This solution uses the Drools rule engine to automatically identify and handle problems such as missing values, duplicate data, and outliers. From data collection to the execution of cleaning rules, to data output and monitoring, the entire process is highly automated, reducing manual intervention. Through the automation mechanism, the efficiency and accuracy of data cleaning are significantly improved, while the labor cost is reduced.

[0034] (5) Strong scalability This solution seamlessly integrates the Drools rule engine with the data synchronization framework, achieving a deep integration of rule engine technology with the existing data integration process. This solution supports seamless integration with other data processing tools and platforms, can be widely applied to various data synchronization frameworks or tools, and has good scalability and usability.

[0035] (6) Wide applicability This solution can be applied to both batch data integration scenarios and real-time data stream integration scenarios. It has strong compatibility and can be integrated with most data integration tools or frameworks, having wide applicability.

[0036] Other features and advantages of the present application will be elaborated in detail in the subsequent specification, or can be understood by implementing the relevant technical solutions of the present application. The objectives and other advantages of the present application can be achieved by the technical features and means clearly pointed out in the specification, claims, and drawings, and can be obtained through the implementation process of these technical contents. Brief Description of the Drawings

[0037] In order to more clearly elaborate the technical solutions of the embodiments of the present application, the drawings involved in the description of the embodiments will be briefly introduced below. It should be noted that the drawings only show some embodiments of the present application. For those skilled in the art, other relevant drawings can be derived based on these drawings without creative labor.

[0038] Figure 1 It is a schematic diagram of the overall design architecture of the solution of the present application.

[0039] Figure 2 It is the overall implementation flowchart of the data cleaning method of the present application based on the Drools rule engine.

[0040] Figure 3 It is the composition structure diagram of the data cleaning device of the present application based on the Drools rule engine.

[0041] Figure 4 It is the schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed Embodiments

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. It should be clear that the described embodiments are only some embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative labor belong to the protection scope of the present application.

[0043] In this document, the term "including" and any form of its deformation (such as "including", "including") are open-ended expressions and should be understood as "including but not limited to", that is, the listed content is not an exhaustive list and may also include other content not explicitly mentioned. The term "based on" should be understood as "at least partially based on", that is, the referred basis or condition may not be the only factor and may also involve other relevant factors. The term "an embodiment" should be understood as "at least one embodiment", that is, the described embodiment is not the only possible implementation and there may be other similar embodiments.

[0044] In this application, when the terms "one" and "multiple" are used to modify relevant elements or features, their expressions are illustrative rather than restrictive. Unless otherwise clearly stated in the context, "one" should be understood as "at least one", and "multiple" should be understood as "at least two". Those skilled in the art should make reasonable interpretations of these terms based on the semantic and logical relationships in the context to ensure that the possibility of "one or more" is covered.

[0045] Term Explanation: Working Memory: One of the core components of the Drools rule engine, used to store and manage data participating in rule matching. The rule engine retrieves data from the working memory and performs pattern matching with predefined rules. When developing an application, users only need to insert data objects into the working memory to enable these data to participate in rule matching and execution. For example, in this embodiment, the kieSession.insert(dynamicObject) method is called to insert the dynamicObject object into the working memory, making it the data source processed by the rule engine.

[0046] Fact: Refers to the data object formed after inserting an ordinary Java object (usually a JavaBean) into the working memory in the Drools rule application. For example, the dynamicObject object in this embodiment becomes a Fact object after being inserted into the working memory. The Fact object is the bridge for data interaction between the application and the rule engine, enabling the data of the application to be recognized, matched, and processed by the rule engine. Through the Fact object, the application can pass business data to the rule engine and operate on the data according to the rule definition.

[0047] Figure 2 The overall implementation process of the data cleaning method based on the Drools rule engine provided by this application is shown as follows, including the following steps: S1. Collect raw data from the data source through the data collection module; S2. Convert the collected raw data into a data format that can be processed by the Drools rule engine and insert it into the working memory of the Drools rule engine; S3. The data cleaning module performs data cleaning processing according to predefined cleaning rules; S4. Output the cleaned data to the target system through the data output module; S5. Record the cleaning log, count the cleaning results, and provide visual display through the monitoring module.

[0048] To more clearly elaborate on the technical solution of this application, such asFigure 1 As shown below, it will be further described through embodiments of specific scenarios.

[0049] (I) Rule Engine The rule engine, namely the Business Rule Management System (BRMS for short), is a software system used to manage and execute business rules. Its core idea is to separate the business decision-making part in the application program, write business decisions (i.e., business rules) through predefined semantic modules, and be configured and managed by users or developers as needed.

[0050] After introducing the rule engine into the system, business rules are no longer directly embedded in the system in the form of program code, but stored in an independent rule library. The rule engine is responsible for loading and executing these rules, thus realizing the decoupling of business rules and system code. Business personnel can manage business rules like managing data, including operations such as querying, adding, updating, statistics, and submission. These rules are loaded into the rule engine for the application system to call.

[0051] (II) Data Acquisition Module The data acquisition module is one of the core components of this solution, and its main function is to acquire raw data from multiple data sources. This module supports multiple data interfaces, including but not limited to relational databases, Kafka, API interfaces, and file systems, etc., and can meet the data acquisition requirements in different scenarios.

[0052] Taking the open-source data synchronization framework DataX as an example, DataX is an offline synchronization tool for heterogeneous data sources open-sourced by Alibaba. It adopts the "Framework + Plugin" architecture, abstracting the data reading and writing functions into Reader and Writer plugins respectively. Users can support new data sources by extending the plugins, thus achieving seamless synchronization with existing data sources.

[0053] In the real-time data integration scenario, the implementation method of the data acquisition module is similar to offline synchronization, reading raw data from various data sources through read plugins to ensure the real-time and accuracy of data.

[0054] (III) Rule Management Module (i.e., Drools Rule Engine Module) The rule management module is the core component in this solution for defining, updating, storing, and loading data cleaning rules. This module supports visual rule editing, reducing the technical threshold of rule definition, enabling business personnel and developers to manage rules more conveniently.

[0055] In this solution, the rule management module uses Drools as the management and execution engine for data cleaning rules. Drools is an open-source rule engine developed based on the Java language provided by the JBoss organization. It can separate complex and variable business rules from hard-coded programs and store them in files or specific storage media (such as databases) in the form of rule scripts. In this way, changes to business rules can take effect immediately in the production environment without modifying the project code or restarting the server, greatly improving the flexibility and response speed of the system, and at the same time reducing the maintenance cost.

[0056] 1. Definition of Rules The Drools rule engine defines rules through rule files with the.drl suffix. A rule consists of a condition (Left-HandSide, LHS) and an action (Right-Hand Side, RHS). In this solution, rules are mainly defined for common scenarios in data cleaning during the data integration process. The following are specific examples: 1.1 Definition of Duplicate Removal Rules The purpose of duplicate removal rules is to delete duplicate records by comparing specific fields (such as id, name, etc.) of the records:

[0057] 1.2 Definition of Missing Value Filling Rules Missing value filling rules can fill in missing data according to business logic. For example, filling in fields with missing values using the mean value, or filling in fields with missing values using default values:

[0058] 1.3 Definition of Format Correction Rules Format correction rules are used to correct data formats that do not meet the requirements. For example, correcting the email address format:

[0059] The above only shows examples of rule definitions for common scenarios of duplicate removal, missing value filling, and format correction in data cleaning. Users can define different rules according to specific business requirements. In addition, when dealing with multi-source heterogeneous data, different rules can be defined to adapt to the characteristics of different data sources, with high scalability.

[0060] 2. Storage of Rules Drools supports two rule storage methods: one is to store rules through rule files with the.drl suffix, and the other is to store rules through a database. In this solution, to facilitate the management, maintenance, and extension of rules, a database is used to store data cleaning rules.

[0061] Taking the MySQL database as an example, the table creation statement is as follows:

[0062] Among them, the rule field is specified as the TEXT type and is used to store the rule string.

[0063] 3. Loading and Updating of Rules Drools supports dynamic loading and updating of rules without restarting the service, which is crucial for the stability of online systems. When the rules are updated, just update the rules stored in the database and they will take effect immediately. For example, the SQL statement for updating the rules is: UPDATE Rules SET rule = 'new rule string' WHERE id = 1; In this solution, by listening for the insert and update operations of the rules, dynamic loading is triggered. When new data cleaning rules are inserted or updated, a listening event will be triggered, enabling Drools to load the new rules.

[0064] (4) Data Cleaning Module The data cleaning module is one of the core components of this solution. Its main function is to receive the raw data, trigger the predefined data cleaning rules, and perform the corresponding data cleaning operations. This module is integrated into the data synchronization framework in the form of a plugin and is responsible for passing the raw data obtained by the data collection module to the Drools rule engine for cleaning.

[0065] The working process of the data cleaning module mainly includes the following steps: (1) Defining the data model In Drools, the data model is defined as Fact. The Fact object is the bridge for data interaction between the application system and the rule engine. The data model is usually defined according to the business data to be cleaned. For example, a typical data model Java class is defined as follows:

[0066] In this solution, due to the integration of data from multiple tables, the data models corresponding to each table may be different. Therefore, it is necessary to frequently create and modify the data models. To handle the dynamic model more flexibly, this solution uses the Map collection in Java to construct the dynamic data model, and its usage is as follows:

[0067] (2) Loading data After the data acquisition module obtains the original data (such as from a relational database, Kafka, API interface, or file, etc.), it is necessary to convert this original data into Fact objects that Drools can process. In the example of this solution, the original data is encapsulated as a Map object dynamicObject corresponding to the dynamic data model; subsequently, by calling the kieSession.insert(dynamicObject) method, the dynamicObject object is inserted into the Drools Working Memory for subsequent rule matching and execution.

[0068] (3)Execute rules When it is necessary to execute the data cleaning rules, call the kieSession.fireAllRules() method to trigger all predefined rules. After the rule execution is completed, call the kieSession.dispose() method to release resources.

[0069] The data cleaning module realizes efficient and flexible data cleaning functions by defining a dynamic data model, loading the original data, and executing predefined rules. This module closely cooperates with the data acquisition module and outputs the cleaned data to the target system, providing reliable support for data integration.

[0070] (V)Data output module The main task of the data output module is to process and output the data after data cleaning. After the data cleaning is completed, this solution obtains the set of cleaned data by calling the kieSession.getObjects() method, and these data are then passed to the target system or storage medium.

[0071] Taking the open-source data synchronization framework DataX as an example (similar to the implementation method of the real-time data integration scenario), the data output module writes the cleaned data to the target end by configuring the corresponding write plugin (Writer Plugin), such as a relational database, file system, message queue, or other data storage systems. The write plugin is responsible for converting and storing the data according to the format and requirements of the target system to ensure the integrity and consistency of the data.

[0072] (VI)Monitoring module The main function of the monitoring module is to monitor the entire process of data cleaning, record the cleaning log, count the cleaning results, and display the status of data cleaning through a visual interface. The monitoring module includes the following core parts: (1)Log recording The monitoring module will record the results of each data cleaning in detail, including but not limited to the following: The timestamp of the cleaning; Name of the cleaning rule used; Cleaning status (success, failure, exception, etc.); Number of data records involved; Changes in data before and after cleaning.

[0073] These log messages provide important basis for subsequent auditing, problem troubleshooting, and optimization.

[0074] (2)Indicator Statistics The monitoring module conducts statistics and analysis on key indicators during the data cleaning process, including but not limited to: Success rate and failure rate of data cleaning; Trigger times and execution times of each cleaning rule; Comparison of data before and after cleaning (such as deduplication rate, missing value filling rate, etc.); Usage of system resources (such as CPU and memory occupancy).

[0075] These indicators can help users evaluate the effectiveness and efficiency of data cleaning and provide data support for optimizing rules.

[0076] (3)Alarm Mechanism The monitoring module has a real-time alarm function. When the following situations occur during the data cleaning process, alarm notifications will be automatically triggered: The failure rate of data cleaning exceeds the preset threshold; Serious errors or exceptions occur during the cleaning process; The usage rate of system resources exceeds the safety threshold; The data cleaning task times out and is not completed.

[0077] Alarm notifications can be sent through various methods, such as emails, text messages, instant messaging tools, etc., so that operation and maintenance personnel can discover and handle problems in a timely manner.

[0078] (4)Visual Interface The monitoring module provides a visual console, which displays information such as the progress of data cleaning, rule execution status, and key indicators through intuitive charts, reports, and dashboards. Users can monitor the status of data cleaning tasks in real time through this interface, quickly locate problems, and evaluate the cleaning effect. The visual interface supports custom configuration, and users can display different data dimensions and indicators according to actual needs.

[0079] The data output module and the monitoring module are important components of the data cleaning system. The data output module is responsible for securely and efficiently writing the cleaned data into the target system, while the monitoring module provides comprehensive monitoring and management support for the data cleaning process through log recording, metric statistics, alert mechanisms, and visual displays, ensuring the efficient execution and quality assurance of the data cleaning tasks.

[0080] Figure 3 The following shows a data cleaning device based on the Drools rule engine proposed in this application. The device includes: Data acquisition module: Used to collect raw data from the data source and convert the collected raw data into a data format that the Drools rule engine can process; Drools rule engine module: Used to define, store, load, update, and execute data cleaning rules; Data cleaning module: Used to receive raw data and perform data cleaning operations through the Drools rule engine; Data output module: Used to output the cleaned data to the target system; Monitoring module: Used to monitor the data cleaning process, record cleaning logs, count cleaning results, and provide visual displays.

[0081] When the above device runs, it implements the steps of the data cleaning method based on the Drools rule engine disclosed in this application.

[0082] The flowcharts and block diagrams in the accompanying drawings show the possible implementation manners of the devices, methods, and computer program products according to various embodiments of this application, including the architecture, functions, and operations. In these figures, each block may represent a module, a program segment, or a part of the code, which contains one or more executable instructions for implementing the specified logical function. It should be noted that each block in the block diagram and / or flowchart, as well as the combinations of these blocks, can be implemented by a dedicated hardware-based system to implement the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0083] As Figure 4 shown, an embodiment of this application also discloses an electronic device, including: a processor 310, a communication interface 320, a memory 330 for storing computer programs executable by the processor, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete mutual communication through the communication bus 340. The processor 310 runs the executable computer program to implement the steps of the above data cleaning method based on the Drools rule engine.

[0084] It is understandable that, in addition to including a memory and a processor, the electronic device may further include an input device (such as a keyboard), an output device (such as a display), and other communication modules. These input devices, output devices, and other communication modules communicate with the processor through an I / O interface (i.e., an input / output interface).

[0085] The operations of this application can be implemented by writing computer program code using one or more programming languages or combinations thereof. The programming languages include, but are not limited to, the following types: Object-oriented programming languages, such as Java, Smalltalk, C++, etc.; Conventional procedural programming languages, such as the "C" language or similar programming languages.

[0086] The execution modes of the program code include, but are not limited to: Executing entirely on the user's computer; Partially executing on the user's computer and partially on a remote computer; Executing as an independent software package; Executing entirely on a remote computer or server.

[0087] In scenarios involving a remote computer, the remote computer can be connected to the user's computer through any type of network. The network includes, but is not limited to, a local area network (LAN) or a wide area network (WAN). In addition, the remote computer can also be connected to an external computer through an Internet service provider, such as by using the Internet for connection.

[0088] Furthermore, this application also discloses a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute each step of the data cleaning method based on the Drools rule engine disclosed in this application.

[0089] In the context of this application, a computer-readable storage medium refers to a tangible medium that can store computer program code and related data. Specific examples include, but are not limited to, the following: (1) Portable computer disks: Removable magnetic storage media such as floppy disks.

[0090] (2) Hard disks: Fixed storage devices including mechanical hard disks and solid-state drives.

[0091] (3) Random access memory (RAM): A volatile storage medium for temporarily storing data and program code.

[0092] (4) Read-only memory (ROM): A non-volatile storage medium for storing fixed programs and data.

[0093] (5) Erasable Programmable Read-Only Memory (EPROM) or Flash Memory: A non-volatile storage medium that supports multiple erasures and programming.

[0094] (6) Fiber optic storage device: A storage medium based on fiber optic technology.

[0095] (7) Portable Compact Disc Read-Only Memory (CD-ROM): A read-only medium that stores data in the form of an optical disc.

[0096] (8) Optical storage device: Storage media based on optical principles such as DVDs, Blu-ray discs, etc.

[0097] (9) Magnetic storage device: Storage media based on magnetic principles such as magnetic tapes, magnetic disks, etc.

[0098] (10) Any suitable combination of the above: For example, combining multiple storage media to meet different storage requirements.

[0099] These computer-readable storage media can be used to store the program code and related data described in this application to support the operation of the program and the persistent storage of data.

[0100] In particular, according to the embodiments of this application, the processes described in the flowcharts can be implemented as computer software programs. For example, the embodiments of this application relate to a computer program product that includes a computer program carried on a non-transitory computer-readable medium. This computer program contains program code for executing the data cleaning method based on the Drools rule engine disclosed in this application. When this computer program is executed by a processing device, it can implement the above functions defined in the embodiments of this application.

[0101] Although the above discussion contains several specific implementation details, these details should not be construed as limiting the scope of this application. The above description is only for the preferred embodiments of this application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the disclosure involved in this application is not limited to the technical solutions formed by the specific combinations of the above technical features. At the same time, this application should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept.

[0102] Those skilled in the art should also understand that they can modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements on some of the technical features without departing from the spirit and scope of the technical solutions of the embodiments of this application. These modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the core spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A data cleaning method based on Drools rule engine, characterized in that: The method comprises: S1. Collecting raw data from the data source through the data acquisition module; S2. Convert the collected raw data into a data format that can be processed by the Drools rule engine and insert it into the working memory of the Drools rule engine; S3. The data cleaning module cleans the data according to predefined cleaning rules; S4. Output the cleaned data to the target system through the data output module; S5. Record cleaning logs, count cleaning results, and provide visual display through the monitoring module.

2. The method according to claim 1, characterized in that The data collection module described in step S1 supports collecting data from relational databases, Kafka, file systems, message queues, and API interfaces; The target system described in step S4 includes: a data warehouse, a data lake, or a downstream application system.

3. The method according to claim 1, characterized in that The cleaning process described in step S3 includes the following operations: Data verification: Identify data that does not meet the requirements through format verification, range verification, and type verification; Deduplication: Remove duplicate records by comparing specific fields of the records; Missing value filling: fill missing fields with mean or default value according to business logic; Outlier processing: Identify extreme values ​​that do not conform to normal distribution or business logic, and delete, correct or mark the outliers; Format conversion: convert data formats according to business needs; Format correction: Correct data that does not conform to the preset format.

4. The method according to claim 1, characterized in that Step S3 also includes: (1) Define the data model Based on the business data that needs to be cleaned, a dynamic data model is constructed using the Map collection in Java; (2) Loading data Encapsulate the original data obtained by the data acquisition module into the Map object dynamicObject corresponding to the dynamic data model; then insert the dynamicObject object into the working memory of Drools by calling the kieSession.insert(dynamicObject) method; (3) Implementation rules When data cleansing rules need to be executed, call the kieSession.fireAllRules() method to trigger all predefined rules; after the rules are executed, call the kieSession.dispose() method to release resources.

5. The method according to claim 1, characterized in that The predefined rules described in step S3 are stored in the database, supporting user customization and real-time updating to meet the data cleaning requirements in different business scenarios.

6. The method according to claim 1, characterized in that The monitoring module described in step S5 has the following functions: Logging function: record the timestamp, rule name, and cleaning status information of data cleaning; Indicator statistics function: statistics on cleaning success rate, failure rate, cleaning rule triggering times, execution time, and system resource usage; Alarm function: send an alarm notification when the data cleaning failure rate exceeds the preset threshold, a serious error occurs during the cleaning process, the system resource usage exceeds the safety threshold, or the data cleaning task times out and is not completed; Visualization function: Display data cleaning progress, rule execution status and key indicators through a visual interface.

7. A data cleaning device based on Drools rule engine, characterized in that: The device comprises: Data collection module: used to collect raw data from the data source and convert the collected raw data into a data format that can be processed by the Drools rule engine; Drools rule engine module: used to define, store, load, update and execute data cleaning rules; Data cleaning module: used to receive raw data and perform data cleaning operations through the Drools rule engine; Data output module: used to output the cleaned data to the target system; Monitoring module: used to monitor the data cleaning process, record cleaning logs, count cleaning results, and provide visual display.

8. The device according to claim 7, characterized in that The device integrates the Drools rule engine module as a plug-in or extension module of DataX, and implements efficient communication and data exchange between the Drools rule engine module and DataX through an API.

9. The device according to claim 7, characterized in that The workflow of the data cleaning module includes: (1) Define the data model Based on the business data that needs to be cleaned, a dynamic data model is constructed using the Map collection in Java; (2) Loading data Encapsulate the original data obtained by the data acquisition module into the Map object dynamicObject corresponding to the dynamic data model; then insert the dynamicObject object into the working memory of Drools by calling the kieSession.insert(dynamicObject) method; (3) Implementation rules When data cleansing rules need to be executed, call the kieSession.fireAllRules() method to trigger all predefined rules; after the rules are executed, call the kieSession.dispose() method to release resources.

10. An electronic device, characterized in that: include: Memory and processor; Memory: used to store computer programs; Processor: used to execute the computer program to implement the steps of the data cleaning method based on the Drools rule engine as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Big data cleaning and process arrangement system, method and equipment and medium

    CN120469992A

  • Heterogeneous data automatic cleaning method and system based on configurable rule engine

    CN121350019A

  • A heterogeneous data automatic cleaning method and system based on a configurable rule engine

    CN121350019B

  • Industrial big data-oriented batch flow integrated data cleaning system

    CN121455941A

  • Rule engine-based coal mine multi-source data dynamic processing and cleaning method

    CN122111991A