Data cleaning method and data cleaning engine based on rule data driving

Through a rule-driven data cleaning method, multi-source data from different business systems is processed, and the problems of high complexity, high cost and high risk in the prior art are solved, and efficient and accurate data cleaning effect is achieved.

CN119938662AActive Publication Date: 2025-05-06BEIJING LIUJINSUIYUE TECH CO LTD

Patent Information

Application Number
CN202510433442.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately process multi-source data from different business systems, especially when facing complex data sources, diverse business needs, and large-scale distributed data processing.

Method used

A data cleaning method based on rule data is adopted to construct a data feature fingerprint by receiving the original data streams of multiple heterogeneous service systems. Then, based on the conditional features in the preset rule library, the data feature fingerprints are similarly matched, the rules whose matching degree is higher than the preset threshold are filtered, and the rules whose input and output dependencies are analyzed, the rules are output, the conflicting rule pairs are marked, and the rule execution sequence is generated through topological sorting, and the original data stream is cleaned.

Benefits of technology

An efficient and reliable data cleaning system is realized, which can handle the relationship between different systems flexibly and accurately, effectively detect and handle rule conflicts, reduce interference from conflicting rules, ensure data quality, and improve data cleaning throughput and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938662A_ABST
    Figure CN119938662A_ABST
Patent Text Reader

Abstract

The invention relates to a data cleaning method based on rule data driving and a data cleaning engine, and belongs to the technical field of data processing. The data cleaning method comprises the following steps: receiving original data streams of a plurality of heterogeneous service systems; extracting a multi-source data feature and a cross-system field association feature, and constructing to obtain a data feature fingerprint; similarity matching is carried out on the data feature fingerprints, rules with the matching degree higher than a preset threshold value are screened, and a candidate rule set is obtained; converting the candidate rule set into a directed acyclic graph; marking conflict rule pairs to obtain a conflict early warning list; a virtual isolation layer is inserted into the directed acyclic graph, and a path identification table is generated; performing topological sorting according to the directed acyclic graph, generating a rule execution sequence, and performing cleaning processing on the original data stream; distributing a data execution channel according to the path identification table; and outputting the cleaned clean data stream and the cleaning log containing the rule execution path. According to the invention, data from different business systems can be efficiently and accurately processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data cleaning method and a data cleaning engine based on rule-based data driving. Background Art

[0002] With the rapid development of information technology, the amount of data has increased exponentially, especially in various enterprises. Various heterogeneous business systems have generated a large amount of structured, semi-structured and unstructured data. These data come from various sources, including customer transaction records, product information, user behavior data, sensor data, social media texts, etc. At the same time, more and more enterprises are gradually adopting emerging technologies such as big data technology, cloud computing technology and the Internet of Things (IoT), which further promote the generation, transmission and storage of enterprise data. Enterprise business systems are also becoming increasingly complex, and often use different technology stacks (such as relational databases, NoSQL databases, file systems, log systems, etc.) to store and manage different types of data.

[0003] However, the heterogeneity and dispersion of data and the ever-expanding diversified needs make data management and cleaning extremely complex. For various business systems, how to accurately extract useful information from multiple different data sources, integrate them reasonably, and ensure the accuracy, consistency, integrity and efficiency of data is an important challenge in the current enterprise informatization process.

[0004] At present, common data cleaning methods often focus on the cleaning and processing of a single data source, and the processing capabilities and intelligence level of multi-source data are still insufficient. Especially when processing data from different business systems, there are often large differences in data formats, transmission protocols, and storage structures between these systems, resulting in high complexity, high costs, and high risks in the data processing process.

[0005] Therefore, faced with complex data sources, diverse business needs and large-scale distributed data processing, how to efficiently and accurately process data from different business systems has become a key issue that needs to be solved urgently. Summary of the invention

[0006] In order to efficiently and accurately process data from different business systems, the present application provides a data cleaning method and a data cleaning engine based on rule-based data drive.

[0007] In the first aspect, the present application provides a data cleaning method based on rule-based data drive, which adopts the following technical solution: A data cleaning method based on rule data drive, the data cleaning method comprising: Receive raw data streams from multiple heterogeneous business systems; Extracting multi-source data features and cross-system field association features in the original data stream to construct a data feature fingerprint; Based on the conditional features in the preset rule base, similarity matching is performed on the data feature fingerprint, and rules with matching degrees higher than a preset threshold are screened to obtain a candidate rule set; Converting the candidate rule set into a directed acyclic graph, wherein the nodes of the directed acyclic graph represent rules and the edges represent data dependencies; Analyze the input-output dependency of each candidate rule in the candidate rule set, mark conflicting rule pairs, and obtain a conflict warning list; Insert a virtual isolation layer into the directed acyclic graph according to the conflict warning list, bind the conflict rule pair to the isolation layer identifier, and generate a path identifier table; Performing topological sorting on the directed acyclic graph after inserting the virtual isolation layer to generate a rule execution sequence; Clean the original data stream according to the rule execution sequence, and add a track label to each data; the track label includes a rule path identifier, a field modification history, and an isolation layer identifier; Allocating a data execution channel according to the path identification table; wherein, when it is detected that the isolation layer identification in the track tag matches the conflict warning list, triggering data rerouting; Output the cleaned data stream and the cleaning log containing the rule execution path.

[0008] By adopting the above technical solutions, an efficient and reliable data cleaning system is formed through multi-source data fusion, rule dependency modeling, conflict isolation and topological execution. This technical solution can flexibly and accurately handle the relationship between different systems when facing multi-source data, and can effectively detect and handle rule conflicts, reducing the interference of conflicting rules, ensuring data quality, and improving data cleaning throughput and efficiency, thereby providing accurate and stable data support for various business systems.

[0009] Optionally, the original data stream includes structured numerical data, text data or a binary file; and the step of extracting multi-source data features from the original data stream includes: Performing dynamic distribution statistics on the structured numerical data to generate data distribution features; Extracting semantic vector features from the text data; Parsing the binary file and extracting metadata features; Multi-source data features are obtained according to the data distribution features, semantic vector features and metadata features.

[0010] By adopting the above technical solutions, a multi-source data feature set is formed, which provides rich information for subsequent data cleaning, anomaly detection and pattern recognition. It not only improves the accuracy and efficiency of data processing, but also can intelligently process diverse heterogeneous data, ensuring the comprehensiveness and high quality of data cleaning.

[0011] Optionally, the step of analyzing the input-output dependency of each candidate rule in the candidate rule set, marking conflicting rule pairs, and obtaining a conflict warning list includes: Receive a candidate rule set, parse the input field set and output field set of each candidate rule, and generate a rule input and output field table; Based on the input and output field table, construct a rule dependency graph; Traversing the rule dependency graph, detecting rule pairs with intersections in output fields and no dependency paths, and marking them as conflicting rule pairs; Determining the confidence of the conflicting rule pair according to the cross-system field association feature in the data feature fingerprint; The corresponding preset routing strategy is matched according to the confidence of each conflicting rule pair, and a conflict warning list is generated.

[0012] By adopting the above technical solutions, rule dependency graphs, cross-system field association features and confidence adjustment mechanisms are used in combination to effectively identify and handle rule conflicts in the data cleaning process. By clarifying the dependencies between rules, potential conflicts are discovered in a timely manner, the confidence of conflicts is adjusted, and the processing flow is optimized according to the preset routing strategy, an efficient and intelligent conflict warning list and processing path are finally generated, thereby improving the accuracy, processing efficiency and flexibility of data cleaning.

[0013] Optionally, the steps of inserting a virtual isolation layer in the directed acyclic graph according to the conflict warning list, binding the conflict rule pair with an isolation layer identifier, and generating a path identifier table include: According to the directed acyclic graph and the conflict warning list, locate the nearest common ancestor node of the conflicting rule pair; Inserting a virtual isolation layer node after the outgoing edge of the nearest common ancestor node pointing to the conflicting rule pair to generate an updated directed acyclic graph; Injecting an isolation layer identifier, a conflict rule pair, and a bypass copy generation mark into the virtual isolation layer node; A path identification table is generated according to the preset routing strategy in the conflict warning list and the cross-system association characteristics of the data feature fingerprint.

[0014] By adopting the above technical solution, the root cause of the conflicting rules is accurately found based on LCA positioning and a virtual isolation layer is inserted at the appropriate location to successfully isolate the conflicting path. Through intelligent routing strategies and analysis of cross-system characteristics, it is ensured that data can flow along the appropriate path, effectively solving problems such as incomplete isolation of data conflicts and poor consistency across systems.

[0015] Optionally, the step of generating a rule execution sequence by topologically sorting the directed acyclic graph after inserting the virtual isolation layer includes: Injecting node labels into the directed acyclic graph after the virtual isolation layer is inserted according to the path identification table; Initialize a high priority queue and a low priority queue according to the node label type, and sort the nodes in the queue based on the business priority strategy; the node label types include main chain nodes, branch nodes, and virtual isolation layer nodes; hierarchically traverse the high priority queue and the low priority queue to generate a main execution sequence and a branch execution sequence; A rule execution sequence is obtained according to the main execution sequence and the branch execution sequence.

[0016] By adopting the above technical solution, based on the injection of node labels, the system can clearly identify the main chain nodes and branch nodes, and reasonably schedule resources according to the priority strategy; at the same time, hierarchical traversal ensures that the rules are executed according to priority, making the execution process of complex rules more stable and efficient.

[0017] Optionally, the step of hierarchically traversing the high priority queue and the low priority queue to generate a main execution sequence and a branch execution sequence includes: Circularly extract the main chain nodes in the high priority queue and add them to the main execution sequence; When the high priority queue is empty, extract the current branch node in the low priority queue and add it to the branch execution sequence; Traverse all outgoing edges of the current branch node and filter out successor nodes whose node label type is a branch; The in-degree of the successor node is updated, and the successor node whose in-degree is reset to zero and whose node label type is branch is added to a low priority queue.

[0018] By adopting the above technical solutions, the system can efficiently manage the execution order and priority of nodes, ensure the priority execution of main chain nodes and timely processing of branch nodes after the main chain nodes are executed, flexibly schedule tasks according to the node dependencies and resource status, optimize resource utilization, avoid disordered execution and conflicts, and improve the parallelism and efficiency of task execution.

[0019] Optionally, after the step of outputting the cleaned data stream and the associated cleaning log, the following step is further included: Based on the rule execution success rate and business feedback data in the cleaning log, eliminate the rules whose execution success rate is lower than the preset success rate threshold and modify the feature extraction parameters; The historical rule execution paths are encoded into a rule gene library that can be genetically optimized, and the rule library version is updated iteratively.

[0020] By adopting the above technical solution, during the rule execution process, the system can monitor the success rate of the rules through cleaning logs, eliminate inefficient rules, and continuously correct feature extraction parameters, thereby ensuring the efficiency and accuracy of the rule base. At the same time, by converting the historical rule execution path into a genetically optimized gene library and using genetic algorithms to iteratively update the rules, the rule base has the ability to self-optimize. This technical solution improves the system's adaptive ability in complex data processing and business execution, ensuring that the rule base always remains efficient, accurate, and sustainably optimized in the face of ever-changing business needs.

[0021] In the second aspect, the present application provides a data cleaning engine based on rule-based data drive, which adopts the following technical solutions: A data cleaning engine based on rule data drive, the data cleaning engine comprising: A receiving module, used to receive original data streams from multiple heterogeneous business systems; A feature extraction module is used to extract multi-source data features and cross-system field association features in the original data stream to construct a data feature fingerprint; A rule screening module is used to perform similarity matching on the data feature fingerprint based on the conditional features in the preset rule base, screen the rules with matching degrees higher than a preset threshold, and obtain a candidate rule set; A conversion module, used to convert the candidate rule set into a directed acyclic graph; wherein the nodes of the directed acyclic graph represent rules and the edges represent data dependencies; A conflict marking module, used to analyze the input-output dependency of each candidate rule in the candidate rule set, mark conflicting rule pairs, and obtain a conflict warning list; A path identification table generating module, used for inserting a virtual isolation layer in the directed acyclic graph according to the conflict warning list, binding the conflict rule pair with the isolation layer identification, and generating a path identification table; A topological sorting module, used to perform topological sorting according to the directed acyclic graph after the virtual isolation layer is inserted, and generate a rule execution sequence; A cleaning processing module, used for cleaning the original data stream according to the rule execution sequence, and adding a track label to each data; the track label includes a rule path identifier, a field modification history and an isolation layer identifier; An execution channel allocation module is used to allocate data execution channels according to the path identification table; wherein, when it is detected that the isolation layer identification in the track label matches the conflict warning list, data rerouting is triggered; The output module is used to output the clean data stream after cleaning and the cleaning log containing the rule execution path.

[0022] In a third aspect, the present application provides a computer device, which adopts the following technical solution: A computer device comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method as described in the first aspect.

[0023] In a fourth aspect, the present application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium stores a computer program that can be loaded by a processor and execute any one of the methods in the first aspect.

[0024] In summary, the present application includes at least one of the following beneficial technical effects: the present application can effectively extract and match data features from multiple heterogeneous business systems, optimize the order of rule execution, and process data dependencies across systems. By constructing a directed acyclic graph of rules, marking conflicting rule pairs, and introducing a virtual isolation layer, the reliability of rule execution and the accuracy of data are ensured. At the same time, the generation of trajectory labels and the data rerouting mechanism provide detailed execution records and dynamic control for the data cleaning process, so that the data processing path can be flexibly adjusted according to business needs. This technical solution not only improves the efficiency and accuracy of data cleaning, but also provides comprehensive traceability and transparency for data quality control, greatly enhances the intelligence and automation level of data processing, can cope with complex heterogeneous data environments, improve data quality, and support more accurate data decisions and business operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a first flow chart of a data cleaning method based on rule-based data drive according to one of the embodiments of the present application.

[0026] Figure 2 This is a second flow chart of a data cleaning method based on rule-based data drive according to one of the embodiments of the present application.

[0027] Figure 3 This is a third flow chart of a data cleaning method based on rule-based data drive according to one of the embodiments of the present application.

[0028] Figure 4 This is a fourth flow chart of a data cleaning method based on rule-based data drive according to one of the embodiments of the present application.

[0029] Figure 5 This is a fifth flow chart of a data cleaning method based on rule-based data drive according to one of the embodiments of the present application.

[0030] Figure 6 This is a sixth flow chart of a data cleaning method based on rule-based data drive according to one of the embodiments of the present application.

[0031] Figure 7 This is the seventh flow chart of a data cleaning method based on rule-based data drive according to one of the embodiments of the present application. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solutions and advantages of this application more clear, the following Figure 1-7 It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0033] The embodiment of the present application discloses a data cleaning method based on rule-based data driving.

[0034] Reference Figure 1 , a data cleaning method based on rule data driven, the data cleaning method includes: Step S101, receiving original data streams of multiple heterogeneous business systems; Among them, the original data stream includes structured numerical data tables, text data and binary files; Specifically, data is uniformly accessed from multiple heterogeneous business systems. Since business systems often use different technology stacks (such as relational databases, NoSQL databases, log systems, or binary files, etc.), this step requires the use of adapter patterns to encapsulate different data source access interfaces to achieve unified access to different data formats.

[0035] For example, the e-commerce system pushes order data (structured data) through the Kafka message queue, the logistics system uploads receipt photos (binary files) through SFTP, and the payment system returns transaction flow (Protobuf format) through the gRPC interface. After these data are connected through the adapter, they are uniformly converted into Parquet format for storage.

[0036] Step S102, extracting multi-source data features and cross-system field association features in the original data stream, and constructing a data feature fingerprint; Among them, the data feature fingerprint (DFP) is a global feature set, including data distribution features, semantic vector features and cross-system field association features of multiple business systems.

[0037] Specifically, numerical distribution features: for example, calculating statistics such as mean, variance, skewness, and kurtosis for data fields such as transaction amount and user age; semantic vector features: using pre-trained language models (such as BERT) to convert unstructured text into semantic vectors to capture the deep semantics of the text; cross-system field association features: building a field association network between different business systems, such as recording the association between the order system and the payment system through a graph database.

[0038] It is understandable that through multi-dimensional feature fusion, DFP can provide rich contextual information for subsequent rule matching and enhance the accuracy of data cleaning. Numerical features can identify abnormal patterns, text features can accurately match specific content, and cross-system correlation features are used to ensure the consistency of rules in different business systems.

[0039] Step S103, based on the conditional features in the preset rule base, similarity matching is performed on the data feature fingerprints, and rules with matching degrees higher than a preset threshold are screened to obtain a candidate rule set; Among them, the data feature fingerprint (DFP) is matched with similarity based on the conditional features in the preset rule base, and the rules with matching degree higher than the preset threshold are screened out. Each rule can be composed of a condition part (such as the amount is greater than 1000) and an action part (such as marking as a suspicious transaction).

[0040] Specifically, the matching process includes structured condition matching, semantic condition matching, and cross-system field matching. Structured condition matching, such as the numerical condition in the amount field matching rule; semantic condition matching, such as the semantic matching of text fields, calculates the similarity between the semantic vector of the data feature fingerprint and the rule condition; cross-system field matching, such as when the rule depends on fields in two systems, checks whether the field association strength in the data feature fingerprint meets the requirements.

[0041] Step S104, converting the candidate rule set into a directed acyclic graph; wherein the nodes of the directed acyclic graph represent the rules, and the edges represent the data dependency; Specifically, the dependency relationship between each rule in the candidate rule set is converted into a directed acyclic graph (DAG) to clarify the order of rule execution. Nodes represent rules and edges represent the dependency relationship between rules. Dependencies include explicit dependencies (such as the output of rule R1 is the input of rule R2) and implicit dependencies (such as business logic requires rule R1 to be executed first).

[0042] For example, in a credit risk control scenario, rule R201 (calculating credit score) and rule R202 (approving credit limit based on credit score) constitute explicit dependencies, while rules R203 (verifying proof of work) and R204 (verifying income flow) have no data dependencies but need to be executed in the order of business logic.

[0043] It is understandable that DAG modeling clearly displays the dependencies between rules, ensuring that the rules are executed in the correct order and avoiding execution errors. By detecting circular dependencies, the risk of rule design is reduced and the failure rate in the production environment is reduced.

[0044] Step S105, analyzing the input-output dependency of each candidate rule in the candidate rule set, marking conflicting rule pairs, and obtaining a conflict warning list; Among them, conflicting rules refer to multiple rules that modify the same field without dependencies, which may lead to inconsistent data results. This step traverses the dependencies between rules, analyzes the input and output of the rules, identifies all possible conflicting rule pairs, and obtains a conflict warning list.

[0045] For example, in a logistics system, rule R301 (timeout marked as "delayed") and rule R302 (manually marked as "normal" by customer service) modify the same "logistics status" field and have no dependency, so they are marked as a conflicting rule pair.

[0046] Step S106, inserting a virtual isolation layer into the directed acyclic graph according to the conflict warning list, binding the conflict rule to the isolation layer identifier, and generating a path identifier table; The virtual isolation layer is a logical marking node inserted into the directed acyclic graph to split the execution path of the conflicting rule pair. The insertion position is after the nearest common ancestor (LCA) node of the conflicting rule pair to ensure that the conflicting rules are executed in different branches. The path identification table records the isolation layer ID, the conflicting rule pair, and the rerouting target (such as the manual review interface or the cross-system verification service).

[0047] For example, rules R7 and R8 conflict, and the LCA node is rule R6. A virtual isolation layer L3 is inserted after R6, and a path identification table is generated, indicating that conflicting rules R7 and R8 need to be manually reviewed and processed.

[0048] It is understandable that the virtual isolation layer ensures that conflicting rules do not affect the main process and prevents erroneous data from entering the downstream. The insertion management of the path identification table makes the system more maintainable and reduces the response time of manual intervention in production.

[0049] Step S107, performing topological sorting according to the directed acyclic graph after the virtual isolation layer is inserted, and generating a rule execution sequence; In one of the embodiments of the present application, topological sorting is implemented based on the Kahn algorithm: initialize the nodes with in-degree 0 (rules without predecessor dependencies), remove the nodes in turn and update the in-degree of the successor nodes until a complete sequence is generated. For nodes with multiple in-degrees (such as multiple rules that depend on the same predecessor rule), sort them according to business priority. The conflict-free path (main execution sequence) and the conflicting path (branch execution sequence) are sorted separately, the main chain is executed in the order of dependency, and the branches are processed asynchronously.

[0050] For example, the final sorting result is R1→R2→R4→R5 (main chain), the conflicting branches R3→L3→R6 are executed asynchronously, the main chain rules are assigned to the high-priority computing nodes, and the branch rules are assigned to the low-priority queues.

[0051] It can be understood that topological sorting can ensure the logical correctness of the main execution chain, and conflicting branches are handled asynchronously to avoid blocking the main process.

[0052] Step S108, cleaning the original data stream according to the rule execution sequence, and adding a track label to each data; the track label includes a rule path identifier, a field modification history, and an isolation layer identifier; Specifically, in the data cleaning process, the track label is the key metadata used to track the flow of data. Among them, the rule path identifier is a unique identifier generated by a hash algorithm for the data flow through the rule sequence. The path identifier is dynamically generated based on the topologically sorted rule execution sequence to ensure the uniqueness and tamper-proofness of the execution chain. The field modification history stores the modification records of the field in the form of key-value pairs, including the original value, modified value, execution rule ID and timestamp. These modifications are incrementally updated through the version control mechanism, supporting data rollback and historical version comparison.

[0053] In addition, when data enters the virtual isolation layer, the isolation layer identifier and the current processing status (e.g., L3→Pending) are injected and bound to the conflicting rule pairs in the path identification table to trigger subsequent rerouting decisions.

[0054] Step S109, allocating a data execution channel according to the path identification table; wherein, when it is detected that the isolation layer identification in the track tag matches the conflict warning list, data rerouting is triggered; Specifically, the path identification table is used to record the mapping relationship between the virtual isolation layer and the conflicting rule pair. When the matching conditions are met, the data rerouting is triggered according to the preset routing strategy in the path identification table: a data copy is created and sent to the specified channel (such as manual review), and the main thread caches the intermediate state and waits for the result.

[0055] For example, after a certain order data executes R1 (format verification) → R2 (risk control scoring), the trajectory label is [R1, R2]; if it enters the isolation layer L3, the copy is sent to manual review, and the label is updated to [R1, R2, L3→pending], and R6 is continued after manual confirmation.

[0056] It is understandable that trajectory tags support data lineage tracking and help locate the root cause of cleaning errors. The rerouting mechanism ensures that conflicting data does not pollute the clean stream, and the quality of the final output data is guaranteed.

[0057] Step S110, outputting the cleaned data stream and the cleaning log including the rule execution path.

[0058] Specifically, the data stream that does not trigger any conflicting rule pairs during the cleaning process is cleaned rule by rule according to the topological sequence to ensure that the output data meets the preset quality standards. The cleaning log will record the final rule execution path and field modification history for downstream system analysis.

[0059] Specifically, the cleaning log contains execution metadata (data ID, cleaning start and end time, resource consumption), rule execution path details (successful / failed rule list, conflict handling records), and performance indicators (single rule time consumption, cross-system call delay).

[0060] In the above implementation, an efficient and reliable data cleaning system is formed through multi-source data fusion, rule dependency modeling, conflict isolation and topology execution. This technical solution can flexibly and accurately handle the relationship between different systems when facing multi-source data, and can effectively detect and handle rule conflicts, reducing the interference of conflicting rules, ensuring data quality, and improving data cleaning throughput and efficiency, thereby providing accurate and stable data support for various business systems.

[0061] Reference Figure 2 As an implementation of step S102, the step of extracting multi-source data features in the original data stream includes: Step S201, performing dynamic distribution statistics on structured numerical data to generate data distribution features; Among them, structured numeric data is usually in the form of a table, containing multiple numeric fields. In this process, the system performs dynamic distribution statistical analysis on each numeric field. This means that the system calculates the statistical characteristics of these data fields, such as mean, variance, skewness, kurtosis, maximum value, minimum value, etc.

[0062] It is understandable that the purpose of dynamic distribution statistics is to analyze the distribution characteristics of data, understand the central tendency of the data, the degree of dispersion, and whether there are outliers. Through these statistical characteristics, the system can identify potential problems in the data (such as skewed distribution or extreme outliers), providing a basis for subsequent data cleaning and processing.

[0063] Step S202, extracting semantic vector features from text data; Among them, text data is usually unstructured, such as log files, comments, articles, etc. By using natural language processing (NLP) technology, pre-trained language models (such as BERT, Word2Vec, etc.) can be used to convert text data into semantic vectors. Semantic vector features can capture deep semantic information in the text, rather than just simple features based on word frequency. Through this step, text data can be converted from abstract, unstructured content to numerical features that can be processed by computers.

[0064] Step S203, parsing the binary file and extracting metadata features; Among them, binary files (such as images, audio, video files, compressed files, etc.) contain structured binary data and cannot be directly used for traditional data analysis. Metadata features in these files can be extracted through specialized parsing tools (such as parsing EXIF ​​data of images, parsing metadata of audio, etc.). Metadata usually includes information such as the creation date, size, format, resolution, file author, encoding method, etc. of the file. These metadata help to understand the basic characteristics of binary data and provide key information for subsequent processing and analysis.

[0065] Step S204, obtaining multi-source data features according to data distribution features, semantic vector features and metadata features.

[0066] Among them, a comprehensive multi-source data feature set is generated by integrating the distribution features of structured data, the semantic features of text data, and the metadata features of binary data. The integration of these multi-source data features can strengthen the connection between data and improve the intelligence and accuracy of the data cleaning process.

[0067] In the above implementation, a multi-source data feature set is formed, which provides rich information for subsequent data cleaning, anomaly detection and pattern recognition, which not only improves the accuracy and efficiency of data processing, but also can intelligently process diverse heterogeneous data to ensure the comprehensiveness and high quality of data cleaning.

[0068] Reference Figure 3 As an implementation method of step S105, the steps of analyzing the input-output dependency of each candidate rule in the candidate rule set, marking conflicting rule pairs, and obtaining a conflict warning list include: Step S301, receiving a candidate rule set, parsing the input field set and output field set of each candidate rule, and generating a rule input and output field table; Among them, each rule usually consists of conditions (input fields) and actions (output fields), so these fields need to be extracted and classified. The input field set refers to the conditional fields required for the rule to be triggered, while the output field set refers to the data fields that are changed or generated after the rule is applied. The hybrid parsing engine is used in the parsing process, combining the syntax tree parsing module and the semantic entity recognition module to ensure the structural and semantic understanding of the rule conditions.

[0069] It is understandable that by generating the rule input and output field table, the system can clearly identify the dependent fields and target fields of each rule, laying the foundation for the subsequent construction of the rule dependency graph and conflict detection. The technical effect of this step is to provide accurate field mapping for rule management in the data cleaning process, so that subsequent rule dependency and conflict analysis have efficient support.

[0070] Step S302, constructing a rule dependency graph based on the input and output field tables; The rule dependency graph is a directed graph, where nodes represent candidate rules and edges represent explicit or implicit dependencies between rules. Explicit dependency means that when the output field of rule A is used as the input field of rule B, rule A explicitly depends on rule B; implicit dependency means that the execution of certain rules must be carried out in sequence based on business logic constraints. The establishment of this dependency helps the system understand the relationship between rules and provides support for subsequent conflict detection and priority scheduling.

[0071] For example, assume that the output field of rule A is "suspicious transaction flag", and the input field of rule B also includes "suspicious transaction flag". If rule A is executed first, rule B will depend on the output field of rule A, then in the rule dependency graph, we will add an explicit dependency edge from rule A to rule B.

[0072] Step S303, traversing the rule dependency graph, detecting rule pairs with intersections in output fields and no dependency paths, and marking them as conflicting rule pairs; The system traverses the rule dependency graph, analyzes the dependencies between the rules, and checks whether there are rule pairs with intersecting output fields. If the output fields of two rules intersect, and there is no dependency path between the two rules (that is, they are independent of each other), then the two rules may conflict and need to be marked as conflicting rule pairs. For example, when two rules try to modify the same field without a dependency in the execution order, data inconsistency or uncertainty will result.

[0073] Step S304, determining the confidence of the conflicting rule pair according to the cross-system field association feature in the data feature fingerprint; The system determines the confidence of conflicting rule pairs based on the cross-system field association features in the data feature fingerprint (DFP). The data feature fingerprint evaluates the actual impact between conflicting rule pairs by modeling the strength of association between cross-system fields. For example, two rules may modify related fields in different systems. If their cross-system association strength is high, the risk of conflict is greater; on the contrary, if their cross-system association is weak, the possibility of conflict is low.

[0074] For example, suppose that rule A and rule B modify fields in two different business systems respectively, but there is a strong cross-system correlation between these two fields (such as the order status and payment status of the same user). The system will increase the confidence of the conflicting rule pair based on this correlation.

[0075] Step S305 , matching the corresponding preset routing strategy according to the confidence of each conflict rule pair, and generating a conflict warning list.

[0076] Among them, the conflict warning list lists all conflicting rule pairs and their confidence levels. After that, the system can select the appropriate processing method based on the preset routing policy template and assign the conflicting rule pairs to a specific adjudication interface. The routing policy template decides whether to handle the conflicting rules through manual review, rerouting, etc. based on the confidence level of the conflicting rule pairs and business needs.

[0077] For example, assuming that the conflicting rules have a higher confidence level for A and B, the system will route them to the manual review interface, while the rules with lower conflicting confidence levels will directly enter the automated processing flow.

[0078] In the above implementation, rule dependency graphs, cross-system field association features, and confidence adjustment mechanisms are used in combination to effectively identify and handle rule conflicts in the data cleaning process. By clarifying the dependencies between rules, timely discovering potential conflicts, adjusting the confidence of conflicts, and optimizing the processing flow according to the preset routing strategy, an efficient and intelligent conflict warning list and processing path are finally generated, thereby improving the accuracy, processing efficiency, and flexibility of data cleaning.

[0079] Reference Figure 4 As an implementation method of step S106, inserting a virtual isolation layer into the directed acyclic graph according to the conflict warning list, binding the conflict rule to the isolation layer identifier, and generating a path identifier table include: Step S401, locating the nearest common ancestor node of the conflicting rule pair according to the directed acyclic graph and the conflict warning list; Among them, by analyzing the rule dependencies in the directed acyclic graph (DAG), the nearest common ancestor (LCA) node of the conflicting rule pair is located. The LCA node is the last common node between two rules in the DAG, which means that the two rules have a common execution path starting from this node.

[0080] In some embodiments, a breadth-first search (BFS) algorithm may be used to traverse, identify common paths between conflicting rule pairs, and find their LCA nodes. LCA positioning is a key step in resolving rule conflicts, because all conflicting rules depend on paths starting from LCA nodes, and a virtual isolation layer will be inserted after the LCA node to ensure that rule conflicts are isolated.

[0081] For example, suppose there is a conflict between rule A and rule B, and they share a common node C in the DAG as the LCA node. Then after LCA positioning, the dependency path of the conflicting rule pairs A and B will be marked as the path starting from node C.

[0082] Step S402, inserting a virtual isolation layer node after the outgoing edge of the nearest common ancestor node pointing to the conflicting rule pair, to generate an updated directed acyclic graph; Once the LCA node is determined, the system will modify the dependency path between the LCA node and the conflicting rule pair. Specifically, the system will cut off the original edge from the LCA node to the conflicting rule pair and insert a virtual isolation layer node between the LCA node and the conflicting rule pair. The virtual isolation layer is a logical node that separates conflicting rules to ensure that they do not share the same execution path, thereby avoiding conflicts.

[0083] For example, suppose the LCA node is C, and conflicting rules A and B depend on the output path of node C. In this case, the system cuts off the original dependency edge from C to A and B, and inserts a virtual isolation layer node L between C node and A and B, generating new branch paths C→L→A and C→L→B.

[0084] Step S403, injecting an isolation layer identifier, a conflict rule pair, and a bypass copy generation mark into the virtual isolation layer node; After inserting the virtual isolation layer node, the system needs to inject the isolation layer ID, conflict rule pair, and bypass replica generation tag into the node. These identifiers ensure the uniqueness and traceability of the isolation layer. The isolation layer ID is the unique identifier of the isolation layer, which is used to track the execution path of the layer; the conflict rule pair records the rules isolated in the isolation layer; the bypass replica generation tag indicates whether a replica needs to be generated and processed through bypass. These tags are crucial for subsequent routing decisions, data processing, and isolation management.

[0085] For example, assume that a virtual isolation layer L is inserted into the DAG, the system assigns an isolation layer ID (such as "L1") to the L node, and records the execution of conflicting rules A and B in the isolation layer. If the L node triggers replica generation, the bypass replica mark indicates that the data needs to take another path for processing.

[0086] Step S404: Generate a path identification table according to the preset routing strategy in the conflict warning list and the cross-system correlation characteristics of the data feature fingerprint.

[0087] The system generates a path identification table based on the preset routing strategy in the conflict warning list and the cross-system correlation features in the data feature fingerprint (DFP). The path identification table contains the isolation layer ID, rerouting target, and cross-system verification rules of the conflict rule pair.

[0088] Specifically, the routing strategy will decide whether the data needs manual review or verification through cross-system services based on the confidence of the conflicting rules. The cross-system correlation feature matrix provides an assessment of the strength of the correlation between fields in the systems. If the strength exceeds the preset threshold, the cross-system verification service is triggered.

[0089] For example, if the confidence of conflicting rules A and B is high and the cross-system association strength between them is also strong, the system will route it to the manual review interface and mark "manual review" as the target in the path identification table. If the association strength is low, the system may directly pass the data to the automatic processing path.

[0090] In the above implementation, the root cause of the conflicting rules is accurately found based on LCA positioning and a virtual isolation layer is inserted at the appropriate location to successfully isolate the conflicting path. Through the analysis of intelligent routing strategies and cross-system characteristics, it is ensured that data can flow along the appropriate path, effectively solving problems such as incomplete isolation of data conflicts and poor consistency across systems.

[0091] Reference Figure 5 As an implementation method of step S107, the step of generating a rule execution sequence by topologically sorting the directed acyclic graph after inserting the virtual isolation layer includes: Step S501, injecting node labels into the directed acyclic graph after inserting the virtual isolation layer according to the path identification table; the node label types include main chain nodes, branch nodes and virtual isolation layer nodes; Among them, a label is injected into each node in the directed acyclic graph (DAG) after the virtual isolation layer is inserted to distinguish different types of nodes. The labels include main chain nodes, branch nodes and virtual isolation layer nodes. The main chain node represents the rule node that does not rely on the virtual isolation layer, the branch node represents the rule node after the isolation layer is branched, and the virtual isolation layer node itself represents the isolated logical layer. The path identification table helps the system understand the execution status of each node and the conflicting path of the rules, thereby providing input for the subsequent rule execution sequence generation.

[0092] Step S502, initializing a high priority queue and a low priority queue according to the node label type, and sorting the nodes in the queue based on the service priority policy; Among them, the system initializes two priority queues according to the type of node label (main chain node, branch node, virtual isolation layer node): high priority queue and low priority queue. Main chain nodes usually have higher business priority because they do not depend on other nodes, while branch nodes and virtual isolation layer nodes are given lower priority. The nodes in the queue are sorted according to the business priority strategy (such as time requirements, computational complexity, etc.) to ensure that the main chain nodes are executed first, and the branch nodes are processed later according to business needs.

[0093] Specifically, the initialization step includes: screening the main chain nodes with an in-degree of 0 to add to the high priority queue, and screening the branch nodes with an in-degree of 0 to add to the low priority queue.

[0094] Step S503, hierarchically traverse the high priority queue and the low priority queue to generate a main execution sequence and a branch execution sequence; The system extracts the main chain nodes from the high priority queue and adds them to the main execution sequence through hierarchical traversal. When the high priority queue is empty, the system extracts the branch nodes from the low priority queue and adds them to the branch execution sequence. During the traversal, the system updates the in-degree of the successor nodes and determines whether to add these nodes to the low priority queue based on the node dependencies (especially for branch nodes, if the nodes they depend on have been executed, their in-degree will be reset to zero, and the system will add them to the low priority queue to wait for execution).

[0095] Step S504: Obtain a rule execution sequence according to the main execution sequence and the branch execution sequence.

[0096] In the above implementation, based on the injection of node labels, the system can clearly identify the main chain nodes and branch nodes, and reasonably schedule resources according to the priority strategy; at the same time, the hierarchical traversal ensures that the rules are executed according to the priority, making the execution process of complex rules more stable and efficient.

[0097] Reference Figure 6As an implementation of step S503, the step of hierarchically traversing the high priority queue and the low priority queue to generate the main execution sequence and the branch execution sequence includes: Step S601, cyclically extracting the main chain nodes in the high priority queue and adding them to the main execution sequence; The system extracts nodes from the high-priority queue in a loop and adds them to the main execution sequence. Main chain nodes are executed first because they usually do not rely on the execution of other nodes and represent the backbone of the process. Therefore, the main chain nodes in the high-priority queue should be processed first, which ensures that the most core part of the business process is completed first.

[0098] For example, assume that the nodes in the high priority queue are A, B, and C, and these nodes are all main chain nodes. The system extracts these main chain nodes from the queue in sequence and adds them to the main execution sequence in order. For example, rule A is extracted and added to the main execution sequence, and then rules B and C are added in sequence.

[0099] Step S602, when the high priority queue is empty, extract the current branch node in the low priority queue and add it to the branch execution sequence; Among them, when the main chain nodes in the high priority queue have all been executed, the system will extract the branch nodes in the low priority queue and add them to the branch execution sequence. The low priority queue usually contains branch nodes with more complex dependencies or requiring fewer resources, so they are executed after the main chain nodes are processed.

[0100] Step S603, traverse all outgoing edges of the current branch node, and filter out the successor nodes whose node label type is a branch; After the current branch node is extracted and added to the branch execution sequence, the system will traverse all outgoing edges of the node and filter out the successor nodes. These successor nodes need to be checked for their type, especially to determine whether the node label type is a branch node. If the successor node is a branch type, then they will become the next branch node to be processed and continue to be added to the low priority queue waiting for execution.

[0101] Step S604, updating the in-degree of the successor node, and adding the successor node whose in-degree is reset to zero and whose node label type is branch to a low priority queue.

[0102] The system will update the in-degree of the selected successor nodes. Whenever a node is executed, the in-degree of all its successor nodes will decrease by 1. When the in-degree of a successor node reaches zero and the node is a branch node, the system will add the node to the low priority queue and wait for subsequent execution. The in-degree reaching zero means that all the predecessor nodes of the node have been completed and the node can be executed.

[0103] For example, suppose rule C is executed, and the out-edges of rule C point to rule D and rule E. Suppose the in-degree of rule E is 1, and the in-degree of rule D is 2. When rule C is executed, the in-degrees of rules D and E are reduced by 1 respectively. Suppose the in-degree of rule D is reduced to 0, and rule D is a branch node, then rule D will be added to the low priority queue to wait for processing.

[0104] It is understandable that by updating the in-degree, the system can ensure that the nodes are executed in order according to the dependencies. The judgment of in-degree zeroing can avoid deadlock or execution order errors, ensure that all dependencies are handled correctly, and effectively organize and optimize the execution process.

[0105] In the above implementation, the system can efficiently manage the execution order and priority of the nodes, ensure the priority execution of the main chain nodes and the timely processing of the branch nodes after the main chain nodes are executed, flexibly schedule tasks according to the node dependencies and resource status, optimize resource utilization, avoid disordered execution and conflicts, and improve the parallelism and efficiency of task execution.

[0106] Reference Figure 7 As a further implementation of the data cleaning method, after the step of outputting the cleaned data stream and the associated cleaning log, the method further includes: Step S701, based on the rule execution success rate and business feedback data in the cleaning log, eliminate the rules whose execution success rate is lower than the preset success rate threshold and modify the feature extraction parameters; The effectiveness of each rule is evaluated by analyzing the rule execution success rate and business feedback data recorded in the cleaning log. If the execution success rate of a rule is lower than the preset success rate threshold, the rule will be considered inefficient and may need to be eliminated or adjusted. Eliminating inefficient rules can avoid waste of resources and ensure that the rules in the rule base can effectively promote business processes. At the same time, through the combination of business feedback data and cleaning logs, the system can also correct feature extraction parameters to improve the accuracy and efficiency of rule execution.

[0107] Specifically, the rule execution success rate is calculated by the number of successes and failures in the execution log, while business feedback data usually comes from actual business results, reflecting the actual effect of rule execution. If a rule fails frequently or fails to achieve the expected effect, the system will optimize the rule based on this data. Correcting feature extraction parameters refers to adjusting the conditions or algorithms used for data extraction in the rule to improve the performance of the rule.

[0108] Step S702: Encode the historical rule execution path into a rule gene library that can be genetically optimized, and iteratively update the rule library version.

[0109] Among them, the historical rule execution path is converted into a rule gene library that can be genetically optimized, and the rule library version is iteratively updated. By analyzing the historical rule execution path, the system can identify the optimal execution path and rule combination. These historical paths are encoded in the form of "genes". This encoding method enables the rule library to have the ability of "genetic optimization", that is, to continuously optimize the rule execution process by simulating natural selection and genetic algorithms.

[0110] Specifically, in the process of encoding the execution path of the rules into genes, the system will extract the key parameters, execution order, and dependency information of each rule execution, and then store them in the form of gene encoding. These rule genes can be continuously evolved through the optimization process of genetic algorithms in subsequent execution to find the best rule combination and execution order.

[0111] For example, assuming that the rule execution path is rule A-> rule B-> rule C, in the genetic algorithm, each rule and its execution order will be encoded as a gene. As the execution path is optimized, the system may find a better path, for example, the order of rule B and rule C can be swapped to improve the overall execution efficiency. Through the genetic algorithm, the system gradually eliminates inappropriate rule combinations and retains the optimal path.

[0112] In the above implementation, during the rule execution process, the system can monitor the success rate of the rules by cleaning the logs, eliminate inefficient rules, and continuously correct the feature extraction parameters, thereby ensuring the efficiency and accuracy of the rule base. At the same time, by converting the historical rule execution path into a genetically optimized gene library and using genetic algorithms to iteratively update the rules, the rule base has the ability to self-optimize. This technical solution improves the system's adaptive ability in complex data processing and business execution, ensuring that the rule base always remains efficient, accurate, and sustainably optimized in the face of ever-changing business needs.

[0113] The embodiment of the present application also discloses a data cleaning engine based on rule-based data drive.

[0114] A data cleaning engine based on rule data drive, the data cleaning engine includes: A receiving module, used to receive original data streams from multiple heterogeneous business systems; The feature extraction module is used to extract multi-source data features and cross-system field correlation features in the original data stream and construct data feature fingerprints; A rule screening module is used to perform similarity matching on data feature fingerprints based on conditional features in a preset rule base, screen rules with matching degrees higher than a preset threshold, and obtain a candidate rule set; A conversion module, used to convert the candidate rule set into a directed acyclic graph; wherein the nodes of the directed acyclic graph represent the rules and the edges represent the data dependency; The conflict marking module is used to analyze the input-output dependency of each candidate rule in the candidate rule set, mark the conflicting rule pairs, and obtain a conflict warning list; A path identification table generation module is used to insert a virtual isolation layer into the directed acyclic graph according to the conflict warning list, bind the conflict rule pair to the isolation layer identification, and generate a path identification table; A topological sorting module is used to perform topological sorting according to the directed acyclic graph after the virtual isolation layer is inserted, and generate a rule execution sequence; The cleaning processing module is used to clean the original data stream according to the rule execution sequence and add a track label to each data; the track label includes the rule path identifier, field modification history and isolation layer identifier; An execution channel allocation module is used to allocate data execution channels according to the path identification table; wherein, when it is detected that the isolation layer identification in the trajectory label matches the conflict warning list, data rerouting is triggered; The output module is used to output the clean data stream after cleaning and the cleaning log containing the rule execution path.

[0115] A rule-based data-driven data cleaning engine in an embodiment of the present application can implement any of the above-mentioned data cleaning methods, and the specific working process of each module in the data cleaning engine can refer to the corresponding process in the above-mentioned method embodiment.

[0116] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are only illustrative; for example, the division of a certain module is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0117] The embodiment of the present application also discloses a computer device.

[0118] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a rule-based data-driven data cleaning method as described above is implemented.

[0119] The embodiment of the present application also discloses a computer-readable storage medium.

[0120] A computer-readable storage medium stores a computer program that can be loaded by a processor and execute any one of the above-mentioned rule-based data-driven data cleaning methods.

[0121] Among them, computer-readable storage media can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus or device; the program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0122] It should be noted that in the above embodiments, the description of each embodiment has different emphases, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0123] The above are all preferred embodiments of the present application, and are not intended to limit the protection scope of the present application. Any feature disclosed in this specification (including the abstract and drawings), unless otherwise stated, can be replaced by other equivalent or alternative features with similar purposes. That is, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.

Claims

1. A data cleaning method based on rule data driven, characterized in that: The data cleaning method comprises: Receive raw data streams from multiple heterogeneous business systems; Extracting multi-source data features and cross-system field association features in the original data stream to construct a data feature fingerprint; Based on the conditional features in the preset rule base, similarity matching is performed on the data feature fingerprint, and rules with matching degrees higher than a preset threshold are screened to obtain a candidate rule set; Converting the candidate rule set into a directed acyclic graph, wherein the nodes of the directed acyclic graph represent rules and the edges represent data dependencies; Analyze the input-output dependency of each candidate rule in the candidate rule set, mark conflicting rule pairs, and obtain a conflict warning list; Insert a virtual isolation layer into the directed acyclic graph according to the conflict warning list, bind the conflict rule pair to the isolation layer identifier, and generate a path identifier table; Performing topological sorting on the directed acyclic graph after inserting the virtual isolation layer to generate a rule execution sequence; Clean the original data stream according to the rule execution sequence, and add a track label to each data; the track label includes a rule path identifier, a field modification history, and an isolation layer identifier; Allocating a data execution channel according to the path identification table; wherein, when it is detected that the isolation layer identification in the track tag matches the conflict warning list, triggering data rerouting; Output the cleaned data stream and the cleaning log containing the rule execution path.

2. The data cleaning method based on rule data drive according to claim 1, characterized in that: The original data stream includes structured numerical data, text data or binary files; The step of extracting multi-source data features from the original data stream includes: Performing dynamic distribution statistics on the structured numerical data to generate data distribution features; Extracting semantic vector features from the text data; Parsing the binary file and extracting metadata features; Multi-source data features are obtained according to the data distribution features, semantic vector features and metadata features.

3. The data cleaning method based on rule data drive according to claim 1 is characterized in that: The steps of analyzing the input-output dependency of each candidate rule in the candidate rule set, marking conflicting rule pairs, and obtaining a conflict warning list include: Receive a candidate rule set, parse the input field set and output field set of each candidate rule, and generate a rule input and output field table; Based on the input and output field table, construct a rule dependency graph; Traversing the rule dependency graph, detecting rule pairs with intersections in output fields and no dependency paths, and marking them as conflicting rule pairs; Determining the confidence of the conflicting rule pair according to the cross-system field association feature in the data feature fingerprint; The corresponding preset routing strategy is matched according to the confidence of each conflicting rule pair, and a conflict warning list is generated.

4. The data cleaning method based on rule data drive according to claim 3 is characterized in that: The steps of inserting a virtual isolation layer into the directed acyclic graph according to the conflict warning list, binding the conflict rule pair with the isolation layer identifier, and generating a path identifier table include: According to the directed acyclic graph and the conflict warning list, locate the nearest common ancestor node of the conflicting rule pair; Inserting a virtual isolation layer node after the outgoing edge of the nearest common ancestor node pointing to the conflicting rule pair to generate an updated directed acyclic graph; Injecting an isolation layer identifier, a conflict rule pair, and a bypass copy generation mark into the virtual isolation layer node; A path identification table is generated according to the preset routing strategy in the conflict warning list and the cross-system association characteristics of the data feature fingerprint.

5. The data cleaning method based on rule data drive according to claim 4 is characterized in that: The step of topologically sorting the directed acyclic graph after inserting the virtual isolation layer and generating a rule execution sequence comprises: Injecting node labels into the directed acyclic graph after the virtual isolation layer is inserted according to the path identification table; Initialize a high priority queue and a low priority queue according to the node label type, and sort the nodes in the queue based on the business priority strategy; the node label types include main chain nodes, branch nodes, and virtual isolation layer nodes; hierarchically traverse the high priority queue and the low priority queue to generate a main execution sequence and a branch execution sequence; A rule execution sequence is obtained according to the main execution sequence and the branch execution sequence.

6. The data cleaning method based on rule data drive according to claim 5 is characterized in that: The step of hierarchically traversing the high priority queue and the low priority queue to generate a main execution sequence and a branch execution sequence includes: Circularly extract the main chain nodes in the high priority queue and add them to the main execution sequence; When the high priority queue is empty, extract the current branch node in the low priority queue and add it to the branch execution sequence; Traverse all outgoing edges of the current branch node and filter out successor nodes whose node label type is a branch; The in-degree of the successor node is updated, and the successor node whose in-degree is reset to zero and whose node label type is branch is added to a low priority queue.

7. A data cleaning method based on rule-based data drive according to any one of claims 1 to 6, characterized in that: After the step of outputting the cleaned data stream and the associated cleaning log, the method further includes: Based on the rule execution success rate and business feedback data in the cleaning log, eliminate the rules whose execution success rate is lower than the preset success rate threshold and modify the feature extraction parameters; The historical rule execution paths are encoded into a rule gene library that can be genetically optimized, and the rule library version is updated iteratively.

8. A data cleaning engine based on rule data drive, characterized in that: The data cleaning engine comprises: A receiving module, used to receive original data streams from multiple heterogeneous business systems; A feature extraction module is used to extract multi-source data features and cross-system field association features in the original data stream to construct a data feature fingerprint; A rule screening module is used to perform similarity matching on the data feature fingerprint based on the conditional features in the preset rule base, screen the rules with matching degrees higher than a preset threshold, and obtain a candidate rule set; A conversion module, used to convert the candidate rule set into a directed acyclic graph; wherein the nodes of the directed acyclic graph represent rules and the edges represent data dependencies; A conflict marking module, used to analyze the input-output dependency of each candidate rule in the candidate rule set, mark conflicting rule pairs, and obtain a conflict warning list; A path identification table generating module, used for inserting a virtual isolation layer in the directed acyclic graph according to the conflict warning list, binding the conflict rule pair with the isolation layer identification, and generating a path identification table; A topological sorting module, used to perform topological sorting according to the directed acyclic graph after the virtual isolation layer is inserted, and generate a rule execution sequence; A cleaning processing module, used for cleaning the original data stream according to the rule execution sequence, and adding a track label to each data; the track label includes a rule path identifier, a field modification history and an isolation layer identifier; An execution channel allocation module is used to allocate data execution channels according to the path identification table; wherein, when it is detected that the isolation layer identification in the track label matches the conflict warning list, data rerouting is triggered; The output module is used to output the clean data stream after cleaning and the cleaning log containing the rule execution path.

9. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.

10. A computer-readable storage medium, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Bus real-time geographic information data cleaning method and system

    CN103699680A

  • An efficient data cleaning conversion method based on CIM

    CN109308290A

  • Base data cleaning method and device based on rule engine, equipment and medium

    CN118467522A

  • Rule matching method and device based on graph theory, computer equipment and storage medium

    CN119557656A

  • Method and apparatus for computing priorities between conflicting rules for network services

    US20040177139A1

Cited By

  • Heterogeneous multi-modal data format synchronization method, system and terminal

    CN121009143A

  • Alarm linkage endless loop detection method and system in rule trigger chain

    CN121257669A

  • Heterogeneous data automatic cleaning method and system based on configurable rule engine

    CN121350019A

  • A heterogeneous data automatic cleaning method and system based on a configurable rule engine

    CN121350019B