Data auditing method, device and equipment
By generating task sequences and performing causal relationship analysis using semantic knowledge graphs, the problem of scattered master data across multiple systems was solved, enabling efficient and accurate data auditing, adapting to business changes, reducing errors, and improving efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA RESOURCES POWER INVESTMENT CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-12
AI Technical Summary
After the centralized deployment of existing technologies in large group ERP systems, master data is scattered across multiple systems, resulting in high procurement costs, inventory backlogs, and compliance risks. Furthermore, traditional auditing methods cannot perceive changes in business semantics in real time, leading to low auditing accuracy.
The system uses a semantic knowledge graph to generate task sequences, identifies abnormal data through anomaly detection tools, performs causal correlation analysis, and dynamically updates the knowledge graph to adapt to business changes, ensuring that audit tasks are consistent with the latest business semantics.
It significantly improves the accuracy of data auditing, reduces auditing errors caused by semantic disconnect, and enhances auditing efficiency and adaptability.
Smart Images

Figure CN122019233A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a data auditing method, apparatus and equipment. Background Technology
[0002] After large groups complete the centralized deployment of Enterprise Resource Planning (ERP), the master data accumulated over the years in multiple systems is still scattered, causing business pain points such as high procurement costs, thus creating an urgent need for master data quality auditing.
[0003] Data auditing methods in related technologies mainly rely on predefined static rule bases (such as format validation and dictionary matching) and fixed metadata models for logical judgment. However, these methods cannot perceive, understand, and adapt to the constantly changing business semantic context in real time. When data changes due to business semantic shifts, the static auditing logic in these technologies struggles to make accurate identifications and judgments, resulting in low accuracy in data auditing. Summary of the Invention
[0004] This application provides a data auditing method, apparatus, and equipment that can improve the accuracy of data auditing.
[0005] In a first aspect, embodiments of this application provide a data auditing method, the method comprising: In response to receiving an audit task, a task sequence corresponding to the audit task is generated based on the audit task and the semantic knowledge graph corresponding to the current time. The semantic knowledge graph includes multiple knowledge entities and business relationship edges between the multiple knowledge entities. The multiple knowledge entities are master data stored in different business systems. Based on the anomaly checking tool corresponding to the task sequence, obtain the abnormal data that exists in the target master data indicated by the task sequence; Based on business relationship edges, perform causal correlation analysis on abnormal data to determine the cause of the abnormality. Based on the confidence level of the abnormal data and the cause of the abnormality, the semantic knowledge graph is updated, and the updated semantic knowledge graph is used to execute the next received audit task.
[0006] Secondly, this application provides a data auditing device, which includes: The generation module is used to respond to receiving an audit task and generate a task sequence corresponding to the audit task based on the audit task and the semantic knowledge graph corresponding to the current time. The semantic knowledge graph includes multiple knowledge entities and business relationship edges between the multiple knowledge entities. The multiple knowledge entities are master data stored in different business systems. The acquisition module is used to acquire abnormal data that exists in the target master data indicated by the task sequence based on the anomaly inspection tool corresponding to the task sequence; The determination module is used to perform causal correlation analysis on the abnormal data based on the business relationship edge to determine the abnormal cause of the abnormal data; The update module is used to update the semantic knowledge graph based on the confidence level of the abnormal data and the cause of the abnormality, and to use the updated semantic knowledge graph to execute the next received audit task.
[0007] Thirdly, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions; When the processor executes computer program instructions, it implements the data auditing method as described in any of the embodiments of the first aspect.
[0008] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the data auditing method as described in any of the embodiments of the first aspect.
[0009] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform a data auditing method as described in any of the embodiments of the first aspect above.
[0010] In the data auditing method, apparatus, and device provided in this application embodiment, when generating a task sequence, this application does not rely on predefined static rules, but dynamically constructs the task sequence by combining the semantic knowledge graph at the current moment with the audit task, enabling the task sequence to accurately match the business association logic of the master data. In the abnormal data identification stage, anomaly checking tools corresponding to the task sequence are used to verify the business relationship edges in the semantic knowledge graph. In the anomaly cause analysis stage, causal association analysis is performed based on the business relationship edges in the semantic knowledge graph to obtain the anomaly cause of the abnormal data. Simultaneously, the semantic knowledge graph is continuously updated through the confidence levels of abnormal data and anomaly causes, enabling the graph to adapt to the dynamic changes in business semantics in real time, ensuring that the audit task always remains consistent with the latest business semantics, thereby significantly reducing audit errors caused by semantic disconnect and significantly improving the accuracy of data auditing. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is one of the flowcharts illustrating the data auditing method provided in the embodiments of this application; Figure 2 This is a second schematic flowchart of the data auditing method provided in the embodiments of this application; Figure 3 This is the third flowchart illustrating the data auditing method provided in the embodiments of this application; Figure 4 This is the fourth flowchart illustrating the data auditing method provided in the embodiments of this application; Figure 5 This is a flowchart illustrating the zero-sample causal inference engine provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a data auditing device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0013] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0014] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0015] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0016] After large groups complete the centralized deployment of ERP systems, incremental data has been standardized and implemented through unified templates and interfaces. However, the existing master data accumulated over the years in multiple systems such as ERP-Customer, ERP-Accounting, Supplier Relationship Management (SRM), and Warehouse Management System (WMS) remains in a scattered, "raw" state. This issue directly leads to the "three highs" business pain points: high procurement costs, high inventory backlog, and high compliance risks. The quality dilemma of existing master data manifests as follows: (1) The same supplier has multiple records in the ERP-Customer database with different names, taxpayer identification numbers and statuses; (2) The same material has multiple codes in SRM and WMS, which makes it impossible to collect procurement contracts and dilutes the scale effect of centralized procurement; (3) The provinces and cities of the customers were filled in arbitrarily, which caused the audit trajectory of the "three flows" of invoices, funds and logistics to be broken.
[0017] Existing master data quality tools generally adopt a single-point architecture of "rule engine + metadata + isolated AI detection". The inherent flaws of this traditional auditing approach are as follows: (1) Rigid rules: Pre-set regular expressions, dictionaries or decision trees cannot adapt to the semantic shifts of business such as "merger of customers, adjustment of product categories, and change of tax rate", resulting in an expansion of the rule base and an exponential increase in maintenance costs; (2) Semantic fragmentation: The metadata only describes the static table structure and lacks a dynamic characterization of the entire business semantics of "customer-contract-order-invoice-payment". The audit results are disconnected from the real business context; (3) Zero-sample blind zone: When a new anomaly that has not been labeled in history appears, the model needs to be re-labeled and trained, and the response cycle is measured in weeks, missing the risk blocking window.
[0018] To address the problems existing in related technologies, embodiments of this application provide a data auditing method, apparatus, and device.
[0019] The data auditing method provided in the embodiments of this application will be introduced first below. For example... Figure 1 As shown, the method specifically includes the following steps: S100, in response to receiving an audit task, generate a task sequence corresponding to the audit task based on the audit task and the semantic knowledge graph corresponding to the current time; the semantic knowledge graph includes multiple knowledge entities and business relationship edges between the multiple knowledge entities; the multiple knowledge entities are master data stored in different business systems.
[0020] Optionally, in this embodiment, the audit task is a specific verification request for master data quality initiated by a user or business system, serving as the trigger point for the entire audit process. Its core objective is to verify quality indicators such as the completeness, consistency, accuracy, and compliance of the master data. Examples include "verifying whether the taxpayer identification numbers of the same supplier are consistent in the ERP and SRM systems," "checking the uniqueness of material codes in the WMS system," and "verifying the standardization of the province and city information entered for customer locations." Tasks are typically submitted in natural language or standardized instructions, clearly defining the audit object, scope, and quality requirements.
[0021] Semantic knowledge graphs are dynamic knowledge carriers supporting the execution of audit tasks. They are structured knowledge models that integrate the master data association logic and business semantics of multiple business systems. They are not static dictionaries, but rather online graphs that continuously evolve with business changes. Their core components include knowledge entities, business relationship edges between entities, temporal features of entities and relationships, and semantic weights. Their core function is to map the entire business logic chain, such as "customer-contract-order-invoice-payment," in real time, capturing semantic shifts in business logic such as tax rate adjustments and customer consolidation. This provides dynamic semantic support for the parsing of audit tasks, sub-task decomposition, and causal relationship analysis, ensuring that the audit logic remains consistent with the real business context.
[0022] Knowledge entities are the core nodes in a semantic knowledge graph, corresponding to the core business objects of master data stored in different business systems. They are the concrete representation of master data in the graph. These entities cover key business entities in the enterprise operation process, including core objects such as customers, products, suppliers, orders, materials, warehouses, and shipping documents. Each entity is bound to a unique identifier and core attributes (such as the supplier's taxpayer identification number, the specifications and models of materials, and the customer's registered address). Knowledge entities unify and abstract master data scattered in systems such as ERP, SRM, and WMS into graph nodes, providing a foundation for cross-system data association and verification.
[0023] Business systems serve as the storage carriers of master data, referring to the various core business support systems deployed by enterprises during their digital operations. These systems encompass multiple business areas, including resource management, supply chain management, and warehousing and logistics. Examples include ERP systems, SRM systems, and WMS systems. Each of these systems stores master data specific to its business scenario. For instance, the ERP system stores basic customer and supplier information, the SRM system stores procurement collaboration-related data, and the WMS system stores material warehousing information, serving as the direct source of master data required for auditing tasks.
[0024] Master data is the core data of critical business entities that is shared, used, and maintained across multiple systems, applications, and business processes within an organization. Unlike business data that records only a single transaction, master data is characterized by high stability, high reusability, and cross-domain sharing. Core categories include core information about business objects such as customers, products, suppliers, employees, materials, and warehouses, such as supplier names, taxpayer identification numbers, material codes, customer registered addresses, and contact information. The quality of master data directly determines operational efficiency and the level of decision-making intelligence, making it a core object of data auditing.
[0025] Business relationship edges are directed association links connecting different knowledge entities in a semantic knowledge graph, used to represent the business logic associations and dependencies between master data entities. These relationships are built based on the full-link business logic such as "customer-contract-order-invoice-payment," for example, "supplier-providing-materials," "order-related-invoice," and "customer-signing-contract," while also including temporal features (such as "order creation time is earlier than payment time") and semantic weights (representing the credibility and business importance of the relationship). Relationship edges evolve dynamically with business changes; for example, after customer mergers, the "customer-order" association link is automatically updated to ensure the accuracy of cross-system entity associations.
[0026] A task sequence is a set of executable subtasks formed after semantic parsing and decomposition of an audit task. It is an intermediate product that transforms an abstract audit task into specific operational steps. Specifically, by planning an intelligent agent based on the business semantics and relational logic of a semantic knowledge graph, the original audit task can be decomposed into multiple parallel or dependent subtasks. Each subtask clarifies the executing subject, the operation object, the tool selection, and the execution priority. Examples include "extracting basic supplier information from the ERP system", "extracting purchase contract data of the same supplier from the SRM system", "comparing the consistency of the supplier's taxpayer identification number between the two systems", and "detecting the causal relationship of inconsistent data".
[0027] Optionally, in one feasible implementation of this application, when the system receives an audit task submitted by a user or business system (such as "check the consistency of the same material code across systems" in natural language form), the planning agent first performs semantic parsing on the audit task, transforms the task into a high-dimensional vector form through an encoding model, and then performs vector matching with knowledge entities and business relationship edges in the semantic knowledge graph to recall the master data entities and related relationships associated with the audit task.
[0028] Based on the business relationship edges between knowledge entities in the semantic knowledge graph, the planning agent filters out valid associations whose edge weights meet preset thresholds and binds them with corresponding audit rules (such as integrity and consistency verification rules). Subsequently, combining the temporal features and causal constraints in the semantic knowledge graph, a dependency graph between subtasks is automatically generated, clarifying the execution order and parallel relationship of each subtask. Each subtask includes the target business system, the master data field to be audited, the execution priority, and the corresponding anomaly checking tool identifier.
[0029] Finally, priority weights are assigned to each subtask through weight normalization to form a complete task sequence. This ensures that the task sequence can accurately match the current business semantic context, providing clear and actionable execution guidelines for subsequent abnormal data acquisition.
[0030] S200, based on the anomaly checking tool corresponding to the task sequence, obtain the abnormal data that exists in the target master data indicated by the task sequence.
[0031] Optionally, in this embodiment, the anomaly detection tool is a set of specialized technical tools for performing audit tasks and detecting master data quality issues, providing core support for the implementation of the task sequence. Its types include at least one of graph computing engines, rule engines, and anomaly detection models: the graph computing engine relies on the business relationship edges of the semantic knowledge graph to trace the source path of master data and determine its completeness; the rule engine has a built-in master data quality rule library covering verification rules such as completeness, consistency, and accuracy, which can accurately match the quality requirements of the target master data; the anomaly detection model learns the distribution characteristics of normal master data through algorithms and automatically identifies abnormal situations that deviate from these characteristics.
[0032] The target master data is the core object of the audit task, referring to the core business data that needs to be quality verified, which is clearly indicated by the task sequence, comes from different business systems, and is subject to quality verification.
[0033] Anomaly data refers to master data records that, after verification by anomaly detection tools, are determined to fail to meet master data quality requirements. These non-compliance issues primarily manifest in dimensions such as completeness, consistency, and accuracy. Examples include records of inconsistent taxpayer identification numbers for the same supplier in different systems, coding conflicts due to "one item, multiple codes" for materials, and records with incorrectly entered province and city information for customers. These data deviate from the expected standard state of master data and may lead to problems such as business process blockages, decision-making biases, and compliance risks.
[0034] Optionally, in one feasible implementation of this application, the executing agent first parses the target requirements of each subtask in the task sequence, including the target master data to be audited, the quality verification dimensions (completeness, consistency, accuracy, etc.) and the corresponding anomaly checking tool type.
[0035] If a graph computing engine is specified for the subtask, the executing agent will trace the source path of the target master data based on the business relationship edges in the semantic knowledge graph, and determine whether there are any anomalies such as missing or broken data in the target master data through path integrity analysis, node association verification, etc.
[0036] If a rule engine is specified, the system will retrieve quality rules (such as encoding format specifications, field non-null constraints, cross-system data consistency standards, etc.) that match the target master data from the master data quality rule library, perform record-by-record verification on the target master data, and filter out records that do not conform to the rules.
[0037] If an anomaly detection model is specified, the standardized target master data is input into the model. The model learns the distribution characteristics of normal master data and identifies anomalous data that deviates from these characteristics.
[0038] During tool execution, the executing agent synchronously invokes the Extract-Transform-Load (ETL) tool to extract, format, and perform lightweight cleaning of cross-system data. It also invokes the Structured Query Language (SQL) tool to align cross-source data and generate a candidate anomaly list, ensuring the standardization and validity of the input data. Finally, the executing agent summarizes the results from each tool, removes duplicate records, and labels the anomaly types (such as data duplication, encoding conflicts, missing fields, etc.), forming a complete set of anomaly data.
[0039] S300, based on the business relationship edge, perform causal correlation analysis on the abnormal data to determine the abnormal cause of the abnormal data.
[0040] Optionally, in this embodiment, the cause of an anomaly refers to the core factor that leads to the anomaly in the target master data, determined after conducting causal association analysis through the business relationship edges of the semantic knowledge graph. For example, the anomaly of inconsistent taxpayer identification numbers of the same supplier in multiple systems may be due to the failure of multi-source entities to align; the anomaly of "one item, multiple codes" for materials may be rooted in the lack of unified coding rules across systems.
[0041] Optionally, in one feasible implementation of this application, the executing agent first traces all knowledge entities associated with the abnormal data (such as master data such as orders, invoices, and materials associated with the abnormal supplier data) based on the business relationship edges in the semantic knowledge graph, and filters out candidate reasons that may cause the abnormality (such as misalignment of cross-system entities, semantic conflict of fields, failure to execute business rules, etc.).
[0042] Subsequently, the abnormal data and candidate causes are mapped to the causal latent space. Causal features are extracted through a pre-trained model, and combined with domain common sense and constraints in a pre-set causal knowledge base, the potential causal relationship between the two is initially constructed.
[0043] Next, using outlier data as the outcome variable and candidate causes as the variables to be analyzed, structural equation modeling was used to identify confounding factors that satisfy the backdoor criterion, eliminate interference from irrelevant variables, and construct a directed acyclic graph (DAG) relationship diagram. Based on this relationship diagram, intervention operations were performed, and counterfactual conditional probability distributions were generated to obtain causal effect values such as average treatment effect or individual treatment effect.
[0044] Next, the significance of the causal effect values was verified: a null hypothesis of "no causal relationship between candidate causes and outliers" was generated based on the relationship graph structure; a control sample set was generated through time series permutation to construct a null distribution; the significance p-values corresponding to the causal effect values were calculated; and significant causal relationships with p < 0.05 were selected. Finally, the core factors with the highest impact on the outliers were determined as the causes of the outliers based on the ranking of the causal effect values.
[0045] S400, based on the confidence level of the abnormal data and the cause of the abnormality, update the semantic knowledge graph, and use the updated semantic knowledge graph to execute the next received audit task.
[0046] Optionally, in the embodiments of this application, confidence level refers to the degree of credibility of the causal relationship between the abnormal cause and the abnormal data determined through causal association analysis. It is the core indicator for quantifying the reliability of the causal relationship, and its value range is usually [0,1]. The higher the confidence level, the greater the probability that the abnormal cause leads to the corresponding abnormal data, and the more reliable the inference result; conversely, it indicates that the causal relationship is more uncertain, and there may be unexcluded confounding factors or inference bias.
[0047] Optionally, in one feasible implementation of this application, the verification agent first summarizes the abnormal data, causes of abnormalities and corresponding confidence levels output by S300, and calculates the confidence level by combining historical repair cases of similar abnormalities in the memory module through cosine similarity.
[0048] For highly reliable anomaly-root cause pairs with a confidence level higher than a preset threshold (e.g., 0.8), the system transforms them into standardized triple vectors (anomaly data node, causal relationship edge, anomaly cause node), and uses a unified encoding function to generate a graph embedding increment Δg. This triple is then added as a new knowledge unit to the semantic knowledge graph, enabling the expansion of entities and relationships. Simultaneously, based on the elastic weight formula σ... ij (t+1)=βσ ij (t)+(1-β)s (where β is the forgetting coefficient, s is the confidence level, σ) ij(t) represents the edge weight of the business relationship edge between the i-th knowledge entity and the j-th knowledge entity in the semantic knowledge graph at time t). The weights of the original related business relationship edges in the graph are updated. High confidence results will strengthen the weight of the corresponding relationship edge, while low confidence results will weaken the weight, thus realizing the accumulation of experience.
[0049] For low-confidence results with a confidence level below the threshold, the system does not add them directly to the graph, but records them to the queue to be verified. At the same time, it sends a replanning signal to the planning agent to adjust the weights of subsequent subtasks for further verification.
[0050] Therefore, when the system receives a new data audit task, the planning agent will directly parse the task using the updated semantic knowledge graph. The new graph includes high-confidence anomalies—root cause triples—accumulated from the current audit, strengthened or weakened business relationship edge weights, and more accurate cross-system master data semantic associations. Thus, when generating task sequences, the planning agent will prioritize business links with increased weights, automatically avoiding paths already verified as low-risk, significantly reducing invalid probing and redundant verification. When acquiring anomalous data, the execution agent can directly utilize the newly added causal relationships in the graph to quickly locate potential anomaly sources without needing to re-derive all associations from the original data. When evaluating results, the verification agent will also combine the latest causal constraints and confidence information in the graph to make a more accurate judgment on the reliability of anomalous data and root causes. In this way, each audit task execution evolves based on previous experience, enabling the system to have stronger adaptability and higher audit efficiency when facing new business changes or unknown anomalies.
[0051] In the data auditing method provided in this application embodiment, when generating the task sequence, this application does not rely on predefined static rules, but dynamically constructs the task sequence by combining the semantic knowledge graph at the current moment with the audit task, so that the task sequence can accurately match the business association logic of the master data. In the abnormal data identification stage, the anomaly checking tool corresponding to the task sequence is used to check based on the business relationship edges in the semantic knowledge graph. In the anomaly cause analysis stage, causal association analysis is performed based on the business relationship edges in the semantic knowledge graph to obtain the anomaly cause of the abnormal data. At the same time, the semantic knowledge graph is continuously updated by the confidence of abnormal data and anomaly cause, so that the graph can adapt to the dynamic changes of business semantics in real time, ensuring that the task sequence generation, anomaly detection and causal analysis of subsequent audit tasks are always consistent with the latest business semantics, thereby greatly reducing the audit error caused by semantic disconnect and significantly improving the accuracy of data auditing.
[0052] In one embodiment, the anomaly detection tool includes at least one of a graph computing engine, a rule engine, and an anomaly detection model; the rule engine includes a master data quality rule base. The anomaly detection tool based on the task sequence obtains abnormal data in the master data indicated by the task sequence that contains anomalies, including: Based on the anomaly detection tool corresponding to the task sequence, output the first data in the master data indicated by the task sequence that contains an anomaly; Based on the first data, the abnormal data is determined; Specifically, when the anomaly detection tool is the graph computing engine, the source path of the target master data is determined based on the semantic knowledge graph; the completeness of the target master data is determined based on the source path; and the first data is determined from the target master data based on the completeness. When the anomaly detection tool is the rule engine, a target quality rule matching the target master data is obtained from the master data quality rule library; based on the target quality rule, the first data is determined from the target master data; the target quality rule is used to determine at least one of the completeness, consistency, and accuracy of the target master data; When the anomaly detection tool is the anomaly detection model, the target master data is input into the anomaly detection model to obtain the first data.
[0053] Optionally, in this embodiment, the source path refers to the complete flow link of the target master data in the semantic knowledge graph from the original data collection source to the current storage location, formed by connecting knowledge entities and business relationship edges. For example, the flow link of "ERP system customer table → ETL synchronization → SRM system procurement table" provides a traceable basis for determining whether the master data is complete.
[0054] Completeness refers to the degree of completeness of master data in three aspects: key attributes, relationships, and flow paths. Key attribute completeness requires that core fields (such as the supplier's taxpayer identification number or material specifications) are not missing; relationship completeness requires that the connection between master data and upstream / downstream related data (such as customers and orders) is unbroken; flow path completeness requires that the source path is not missing or interrupted. Failure to meet completeness standards indicates that the master data has problems such as missing key information or broken relationships.
[0055] The master data quality rule base is a core component of the rule engine, serving as a structured database storing various master data quality verification rules. The rules in the rule base are formulated based on business logic, compliance requirements, and data standards, covering multi-dimensional verification logic such as completeness, consistency, and accuracy. Examples include rules for "supplier name not empty," "cross-system material code consistency," and "customer address conforming to administrative division standards." The rule base supports dynamic updates, iterating synchronously with changes in business semantics (such as tax rate adjustments or changes in customer classification standards), ensuring that verification rules always match the business status.
[0056] Target quality rules are specific verification rules selected from the master data quality rule library that completely match the target master data indicated by the current task sequence. They define specific verification standards and judgment logic based on the type of target master data (e.g., customer, supplier, material), the business scenario (e.g., procurement, warehousing), and quality requirements. For example, there is a "consistency verification rule" for supplier data in ERP and SRM systems (taxpayer identification numbers must be completely consistent).
[0057] Consistency refers to the degree to which information related to the same master data entity remains consistent and conflict-free across different business systems, storage locations, or time points. It includes dimensions such as cross-system consistency (e.g., the same supplier's name and code are consistent in ERP and SRM systems), field format consistency (e.g., date format is uniformly "YYYY-MM-DD"), and logical consistency (e.g., the correspondence between order status and payment status is consistent). Insufficient consistency can lead to data conflicts and business process blockages (e.g., the inability to aggregate purchase contracts).
[0058] Accuracy refers to the degree to which the attribute values and semantic descriptions of the target master data are consistent with actual business conditions and standard specifications. It requires that the field values of the master data be true and valid (e.g., the supplier's registered business address matches the address entered), that the semantic descriptions conform to unified standards (e.g., order status codes match business definitions), and that the calculation results be accurate (e.g., material unit prices match pricing rules). Failure to meet accuracy standards can lead to decision-making biases and compliance risks (e.g., discrepancies between invoice flow and actual business transactions), directly impacting the reliability and compliance of enterprise operations.
[0059] Anomaly detection models are intelligent verification tools built on machine learning algorithms. By learning the characteristic distribution patterns of normal master data, they automatically identify anomalous data that deviates from this distribution. The model learns the attribute value ranges, relationship patterns, and semantic features of normal master data through training data, such as the tax number format of normal suppliers and the coding rules of material codes. When target master data is input into the model, it calculates the degree of deviation from the normal feature distribution. If the deviation exceeds a preset threshold, it is identified as the first piece of data with an anomaly. This can efficiently identify hidden anomalies (such as implicit duplicate data and semantically conflicting data) that are difficult for rule engines to cover.
[0060] Optionally, in one specific implementation of this application, the executing agent selectively invokes a graph computing engine, a rule engine, or anomaly detection model to perform operations based on the tool type specified in the task sequence. After invoking the corresponding tool, the tool first outputs the first data indicating an anomaly in the master data indicated by the task sequence; subsequently, the system further determines the final anomaly data based on the first data.
[0061] If the task sequence clearly requires the joint invocation of multiple anomaly detection tools (such as the calculation engine + rule engine, rule engine + anomaly detection model, or the three working together), the system will combine the first data output by each tool to obtain the final anomaly data: specifically, it can use methods such as intersection and merging, full set integration, or hierarchical filtering.
[0062] The system employs a three-tiered approach: intersection and merging, retaining only the first data points identified as anomalous by multiple tools to ensure high reliability; full set integration, aggregating the first data points from all tools to avoid overlooking potential anomalies; and hierarchical filtering, filtering the first data points sequentially based on tool priority (e.g., rule engines prioritize basic compliance checks, while anomaly detection models supplement the identification of latent anomalies), ultimately forming a logically coherent set of anomalous data covering different quality dimensions. Simultaneously, the system labels each piece of anomalous data with its source tool and judgment criteria, ensuring the traceability of the anomalous data.
[0063] If the graph computing engine is invoked, it will be based on the semantic knowledge graph and trace the flow link (i.e. the source path) of the target master data across systems through the path analysis function. It will check the integrity of the association and the continuity of data flow of each node in the link one by one to determine the integrity of the target master data. Master data whose integrity does not reach the preset threshold (such as missing key nodes or broken links) will be marked as the first data.
[0064] If the rule engine is invoked, it will first filter the target quality rules from the master data quality rule library according to the type of the target master data (such as customer, material) and the business scenario to which it belongs. Then, it will verify the target master data one by one according to the rules and filter out the abnormal records that do not conform to the rules as the first data.
[0065] If an anomaly detection model is invoked, the target master data will first undergo standardized preprocessing (such as format unification and feature extraction) and then be input into the model that has been trained with historical normal master data. The model will identify and output data that deviates from the normal distribution as the first data by comparing the feature distribution differences between the target data and the normal data.
[0066] In these alternative embodiments, tools can be flexibly selected according to task requirements to ensure that the identification of abnormal data not only fits the business semantic context, but also covers quality issues in different dimensions, thus ensuring the comprehensiveness and accuracy of audit results.
[0067] In one embodiment, the step of performing causal correlation analysis on the abnormal data based on the business relationship edge to determine the cause of the abnormality of the abnormal data includes: Based on the business relationship edge, at least one candidate cause associated with the abnormal data is determined; The anomalous data and the candidate causes are mapped to a causal latent space to obtain causal features; the causal latent space represents the potential causal relationship between the anomalous data and the candidate causes. Based on the causal characteristics and the preset causal knowledge base, the causal relationship between the abnormal data and the candidate causes is determined; Based on the causal relationship, the causal effect value of each candidate cause on the abnormal data is determined; the causal effect value is used to characterize the degree of influence of the candidate cause on the abnormal data. Based on the causal effect value, the abnormal cause is determined from each of the candidate causes.
[0068] Optionally, in this embodiment of the application, the candidate reason is a set of potential factors that may cause the anomaly, selected from the knowledge entities and business links associated with the abnormal data, based on the business relationship edges in the semantic knowledge graph.
[0069] The causal latent space is a high-dimensional feature space used to characterize the potential causal relationships between anomalous data and candidate causes. It is constructed using pre-trained models (such as the causal Transformer). This space discards the interference of surface correlations in the data, retaining only the core information that reflects the "cause-effect" logic. In this space, anomalous data with genuine causal relationships will exhibit stronger feature correlations with candidate causes.
[0070] Causal features are high-dimensional feature vectors obtained by mapping outlier data and candidate causes to a causal latent space, and they are a quantitative representation of the potential causal relationship between the two. For example, when the candidate cause is "coding rule conflict" and the outlier data is "inconsistent material coding," the causal features will highlight core information such as "rule difference degree" and "code generation time sequence difference." These features can accurately characterize the impact logic of candidate causes on outlier data.
[0071] The pre-defined causal knowledge base is a structured database storing known causal relationship rules and domain common sense within the master data business domain. It is the core foundation supporting zero-shot causal reasoning. The database contains business-verified causal rules such as "customer-merchant merger → duplicate supplier records" and "missing logistics tracking number → inventory data deviation," while also covering temporal causal constraints such as "order creation time must be earlier than payment time." Its function is to provide prior evidence for determining causal relationships and reduce blind searches. For example, when analyzing the "missing invoice header" anomaly, the rule "valid tax number → header should exist" in the knowledge base can quickly help determine the rationality of the candidate cause "tax number verification not executed."
[0072] Causality refers to the "cause-effect" logical relationship between candidate causes and abnormal data, meaning that the existence of a candidate cause directly leads to the generation of abnormal data. For example, the causal relationship between "multi-source entity misalignment" and "inconsistent records from multiple systems of the same supplier" must satisfy the logical constraint that "after excluding other rule conflicts, misalignment will inevitably lead to record inconsistency."
[0073] The causal effect value is a numerical indicator that quantifies the degree of influence of a candidate cause on outlier data. Its value range is usually [0,1]. The closer the value is to 1, the more significant the impact of the candidate cause on the outlier data. For example, when the candidate cause is "address not standardized", the causal effect value will quantify the probability that this factor will lead to the "three-stream audit trajectory break" anomaly.
[0074] In these alternative embodiments, the closed loop of association filtering, feature mapping, knowledge base verification, and effect quantification accurately locates the cause of anomalies. This not only breaks through the limitations of traditional methods that only detect anomalies, but also eliminates interference by using causal logic, thereby improving the accuracy and reliability of root cause determination.
[0075] In one embodiment, determining the causal relationship between the abnormal data and the candidate causes based on the causal features and a preset causal knowledge base includes: Using the abnormal data as the outcome variable and the candidate causes as the variables to be analyzed, a structural equation model is used to determine the potential causal relationship between the outcome variable and the variables to be analyzed. The potential causal relationships are filtered using the causal constraints stored in the causal knowledge base to obtain the causal relationship between the abnormal data and the candidate causes.
[0076] Optionally, in the embodiments of this application, the outcome variable is the object being explained in the causal association analysis, that is, the clearly existing abnormal data, which is the "effect" that the candidate cause may produce.
[0077] The variable to be analyzed is the potential factor used to explain the outcome variable in causal association analysis. It is the candidate cause selected based on business relationship edges, which is the "cause" that may cause abnormal data.
[0078] Structural equation modeling (SEM) is a statistical modeling tool used to analyze complex causal relationships between variables. It can handle both direct and indirect causal relationships and is suitable for scenarios involving multiple interacting factors in master data operations. By constructing mathematical equations that include the variables to be analyzed, the outcome variables, and potential confounding factors, it quantifies the strength and direction of the association between variables. For example, it can quantify the impact of "inconsistent coding rules" on "material coding conflicts" through equations, while simultaneously eliminating the interference of confounding factors such as "data entry delays," providing a quantitative basis for identifying potential causal relationships.
[0079] Potential causal relationships are preliminary associations between the variables to be analyzed and the outcome variables obtained through structural equation modeling analysis, but have not yet been verified by business logic. Such associations may include genuine causal relationships and spurious associations (such as correlations that only occur due to time synchronization). For example, model analysis may reveal a potential association between "cross-system data synchronization delay" and "inconsistent supplier information," but whether it conforms to the actual logic of master data business still needs further verification in conjunction with causal constraints.
[0080] Causal constraints are business-verified causal logic rules and domain common sense stored in a pre-defined causal knowledge base. They are the core basis for filtering genuine causal relationships. These constraints include explicit causal association rules (such as "merger of customers and merchants → duplicate supplier records"), time constraints (such as "order creation time must be earlier than payment time"), and business logic constraints (such as "if the tax number is valid, the invoice header should exist"). They can eliminate false associations in potential causal relationships and ensure that the final causal relationship conforms to the actual operating logic of the master data business.
[0081] In these alternative embodiments, potential causal relationships are discovered through structural equation modeling and then filtered using causal constraints. This process not only eliminates confounding factors but also aligns with actual business logic, accurately removing spurious associations. This improves the accuracy and reliability of causal relationship determination and lays a solid foundation for subsequent causal effect quantification and anomaly cause localization.
[0082] In one embodiment, determining the causal effect value of each candidate cause on the abnormal data based on the causal relationship includes: Based on the causal relationship, a relationship graph is constructed, where the vertices of the relationship graph represent the abnormal data and the candidate causes, and the edges of the relationship graph represent the causal relationship between the abnormal data and the candidate causes. Based on the relationship diagram, an intervention operation is performed on the candidate causes to determine the counterfactual conditional probability distribution of the abnormal data under the intervention operation; the intervention operation is used to change the value of the variable corresponding to the candidate cause to simulate the impact of the value change on the abnormal data; the counterfactual conditional probability distribution is used to characterize the probability distribution of the abnormal data as the outcome variable under the intervention operation. Based on the counterfactual conditional probability distribution, determine the causal effect value of each candidate cause on the abnormal data.
[0083] Optionally, in this embodiment, the relationship graph is a directed graph structure that visually presents the causal relationship between abnormal data and candidate causes. Its core consists of vertices and edges. Vertices correspond to abnormal data (outcome variables) and each candidate cause (variables to be analyzed), clearly marking the core identifiers of the data and potential causes. Edges are directional connecting lines, pointing from candidate causes to abnormal data, intuitively representing the logical relationship between cause and effect. They can also include association strength indicators, providing a structured visual analysis foundation for subsequent intervention operations and causal effect calculations.
[0084] Intervention operations are simulations that adjust the values of variables corresponding to candidate causes based on the relationship diagram. The purpose is to isolate and analyze the independent impact of the candidate cause on outlier data. The operation method is to keep other variables unchanged and only change the value of the target candidate cause (e.g., changing "coding rules are not unified" from "yes" to "no"), thereby eliminating confounding factors and accurately simulating the possible changes in the state of outlier data after the change in the value of the cause. This is a key prerequisite for quantifying causal effects.
[0085] The counterfactual conditional probability distribution characterizes the probability distribution of anomalous data as an outcome variable under intervention. In other words, it's the set of probabilities that anomalous data would exhibit different states if the value of a candidate cause were changed. For example, it could be the distribution of the probability and severity of the "material code conflict" anomaly after the intervention "coding rules are not unified" is changed to "unified". This distribution is derived through multiple sampling simulations and can quantify the magnitude and likelihood of the impact of candidate causes on anomalous data.
[0086] Alternatively, in one specific implementation of this application, it is implemented in the following manner: 1. Candidate cause screening: Based on the business relationship edges in the semantic knowledge graph, trace all knowledge entities and business links associated with abnormal data, screen out potential causes at the data level and business level, and form a set of candidate causes; 2. Causal Feature Extraction: A pre-trained model is used to map abnormal data and candidate causes to the causal latent space. Causal features are extracted through regularization or contrastive learning to ensure that the features retain only the core information of "cause → effect" and discard confounding biases. 3. Potential causal association mining: Using outlier data as the outcome variable and candidate causes as the variables to be analyzed, a mathematical equation containing confounding factors is constructed through structural equation modeling to quantify the strength and direction of the association between variables and determine the potential causal association between them. 4. Causal Relationship Filtering: Invoke the causal constraints in the preset causal knowledge base to eliminate false correlations in potential associations and obtain the true causal relationships that conform to the business logic of the master data; 5. Relationship Graph Construction: Using abnormal data and candidate causes as vertices and actual causal relationships as directed edges, a visual directed relationship graph is constructed to clearly present the "cause → effect" logic; 6. Intervention and Counterfactual Reasoning: Based on the relationship diagram, intervention operations are performed on candidate causes (other variables are fixed, and the values of the target candidate cause are changed), and the counterfactual conditional probability distribution is calculated using the formula:
[0087] in, To express the probability that variable Y will take a certain value after an intervention operation is performed on variable X (forcing it to take the value x), do( X is an intervention quantifier for causal inference, representing the active setting of the value of X, rather than passively observing X=x. X is the "causal variable" in the causal relationship, corresponding to the candidate anomaly causes in master data quality audits; Z is the "confounding factor set" in the causal relationship, referring to variables that simultaneously affect X and Y and may lead to spurious correlations. These variables must satisfy the "backdoor criterion" (i.e., blocking all association paths between X and Y except the causal path); "Probability of Y taking a certain value given that X is interfered with as x and the confounding factor Z takes a specific value" reflects the direct correlation strength between X=x and Y after excluding the interference of Z; The marginal probability of the confounding factor Z refers to the probability of the confounding factor Z taking a specific value. It is used to perform a weighted average of the conditional probabilities under different Z values.
[0088] 7. Causal effect value calculation: Based on the counterfactual conditional probability distribution, the causal effect value (range [0,1]) is calculated using the formula for average treatment effect or individual treatment effect. The closer the value is to 1, the more significant the impact of the candidate cause on the outlier data. 8. Identification of Anomalies: Sort the causal effect values from high to low, select the most significant core factors as the causes of anomalies, and form a complete analysis result of "abnormal data - causal link - cause of anomalies".
[0089] In these alternative embodiments, by constructing a relationship graph to visualize causal logic, the impact of intervention operations on candidate causes on abnormal data is simulated, and the causal effect value is quantified by combining counterfactual conditional probability distribution. This can accurately quantify the degree of influence of each candidate cause, eliminate false association interference, provide objective quantitative basis for screening core abnormal causes, and improve the accuracy and reliability of abnormal cause location.
[0090] In one embodiment, determining the anomalous cause from among the candidate causes based on the causal effect value includes: The abnormal data is divided into multiple data blocks according to a preset time window. For the first variable of the candidate cause corresponding to each data block, the first variable is randomly rearranged and the timestamp of the first variable is reassigned, while keeping the timestamps of other data in the data block other than the first variable unchanged, to obtain the control sample set; A null distribution is constructed based on the control sample set, the position of the cumulative distribution function of the causal effect value in the null distribution is determined, and the significance corresponding to the causal effect value is obtained; the null distribution is used to characterize that the first variable has no causal relationship with the outlier data. If the significance is greater than a preset significance threshold, the candidate cause corresponding to the causal effect value is determined as the abnormal cause.
[0091] Optionally, in this embodiment, a data block is the smallest data unit obtained by dividing abnormal data according to a preset time window. Each data block contains complete master data records related to the abnormal data within that time window and corresponding candidate cause information. The time window can be flexibly set according to the business scenario (such as hour, day, week). For example, "monthly customer data anomalies" can be divided into data blocks by "day" to facilitate verification of the impact of candidate causes in each time period.
[0092] The first variable is the quantitative variable corresponding to a certain candidate cause in each data block, and it is the core intervention object in causal effect analysis.
[0093] The control sample set is a comparative dataset generated by randomly rearranging the first variable and reassigning timestamps. It is used to simulate the scenario where "candidate causes and anomalous data have no causal relationship". During construction, the timestamps of other data in the data block besides the first variable (such as the anomalous data itself and other candidate cause variables) are kept unchanged. Only the temporal distribution of the first variable is shuffled. For example, the 1 / 0 values of the "missing logistics tracking number" variable are randomly redistributed to the original timestamps to eliminate false associations caused by temporal coincidences.
[0094] The null distribution is a probability distribution constructed based on a control sample set. It is used to characterize the possible distribution range of causal effect values when "the first variable (candidate cause) has no causal relationship with the outlier data". It is formed by repeatedly generating control sample sets, calculating the corresponding causal effect values, and statistically analyzing the distribution characteristics (such as mean, variance, and quantiles) of these effect values. For example, after repeatedly randomly rearranging the variable "missing logistics tracking number", multiple sets of causal effect values without a real causal relationship are obtained, and their distribution is the null distribution.
[0095] The cumulative distribution function (CDF) position refers to the specific point on the CDF corresponding to the actual calculated causal effect value. It quantifies the rarity of the effect value in the zero distribution. The CDF describes the probability that the causal effect value is less than or equal to a certain value in the zero distribution. For example, if the actual causal effect value corresponds to the 95th percentile on the CDF, it indicates that only 5% of the effect values in the zero distribution are greater than this actual value, suggesting a high degree of deviation from the "no causal association" hypothesis.
[0096] Significance is a quantitative indicator based on the position of the cumulative distribution function. It is used to determine the reliability of the causal relationship between candidate causes and outlier data, and its value range is usually [0,1]. It reflects the probability that the actual causal effect value comes from the null distribution (i.e., no causal relationship). The closer the significance is to 1, the more reliable the causal relationship is; the closer it is to 0, the more likely it is a spurious relationship.
[0097] Optionally, in one specific implementation of this application, firstly, the abnormal data is split according to a preset time window (such as hours or days) to obtain multiple time-series continuous data blocks. Each data block contains abnormal data within the corresponding time period and related candidate cause information, ensuring the integrity of the data's temporal characteristics. Next, for the first variable corresponding to the candidate cause in each data block, the timestamps are randomly rearranged and reassigned, while keeping the timestamps of other data within the data block unchanged, thus constructing a control sample set where "the candidate cause has no real correlation with the abnormal data".
[0098] Subsequently, causal effect values were calculated multiple times based on the control sample set. The distribution characteristics of these values were statistically analyzed to construct a null distribution. Then, the specific position of the actual causal effect value in this null distribution was determined by the cumulative distribution function, and the significance (range [0,1]) was calculated. Finally, the significance was compared with a preset threshold (e.g., 0.95). If the significance was greater than the threshold, the causal relationship between the candidate cause and the abnormal data was deemed reliable, and it was identified as an abnormal cause; otherwise, the candidate cause was excluded, ensuring the scientific and accurate location of abnormal causes.
[0099] In these alternative embodiments, by dividing data into time-series blocks, constructing a control sample set to generate a zero distribution, and combining the cumulative distribution function to determine significance, the causes of anomalies can be screened using quantitative indicators. This can eliminate the interference of time-series coincidences and spurious associations. Significance verification ensures the credibility of the causal relationship between the causes of anomalies and the abnormal data, improves the accuracy and scientific nature of root cause localization, and provides a reliable basis for data governance.
[0100] In one embodiment, updating the semantic knowledge graph based on the confidence levels of the abnormal data and the causes of the abnormality includes: Based on the stated cause of the anomaly and the stated data, an anomaly group is constructed; The similarity between the abnormal group and the historical group is determined as the confidence level; the historical master data corresponding to the historical group matches the target master data; If the confidence level is greater than a preset confidence threshold, the edge weights of the business relationship edges corresponding to the abnormal data are updated based on a weight update formula; the weight update formula is σ. ij (t+1)=βσ ij (t)+(1-β)s; where σ is the forgetting rate coefficient, β∈(0,1), and σ is the forgetting rate coefficient. ij (t) represents the original edge weight of the business relationship edge at the current time t, and s represents the confidence level.
[0101] Optionally, in this embodiment, the anomaly group is a related data unit constructed based on the abnormal data identified in the current audit task and the corresponding abnormal cause. It is the core carrier representing the "problem-root cause" correspondence. For example, a combination of "abnormal data: conflict between ERP and SRM system supplier tax number + abnormal cause: coding rules not synchronized".
[0102] Historical groups are historical data units stored in the system memory module that match the current target master data type. They correspond to verified "abnormal data - abnormal cause" combinations in historical audit tasks. Their core characteristic is that the historical master data and the current target master data belong to the same business type (e.g., both are supplier master data or material master data), and contain complete historical anomaly details, corresponding root causes, and verification results. For example, archived combinations such as "inconsistent supplier names between ERP and OMS systems + anomaly cause: name standardization not implemented" provide comparable historical evidence for current confidence level calculations.
[0103] Edge weight is a quantitative indicator of the importance of business relationship edges connecting two knowledge entities (master data) in a dynamic semantic knowledge graph, denoted by σ. ijThis indicates that the value range is usually [0,1], directly reflecting the credibility and priority of the business relationship in audit reasoning. For example, the weight of the relationship edge between "supplier entity - order entity" will be adjusted according to the confidence level of the abnormal association between the two. The higher the confidence level, the greater the weight, and the more likely the relationship will be called first in subsequent audits, ensuring that the graph can continuously accumulate effective business association experience.
[0104] Optionally, in one specific implementation of this application, firstly, the abnormal data obtained from the current audit is integrated with the corresponding abnormal causes to construct a structured abnormal group, and then converted into a standardized vector through an encoding function. Next, the verification agent calls the historical group stored in the memory module and calculates the similarity between the two using the cosine similarity formula s=cos(abnormal group vector, historical group vector). This similarity is the confidence score s (within the range [0,1]), which is used to characterize the degree of matching between the current abnormal association and the historical valid cases.
[0105] Subsequently, the confidence level s is compared with the preset confidence threshold τ. If s ≥ τ, it indicates that the current abnormal association is credible, based on the elastic weight formula σ. ij (t+1)=βσ ij (t)+(1-β)s updates the weight of the corresponding business relationship edge in the graph. High confidence will strengthen the weight of the corresponding relationship edge and realize experience accumulation. If s<τ, the weight is not updated, and a replanning signal is sent to the planning agent to adjust the priority of subsequent subtasks for further verification.
[0106] In these alternative embodiments, by constructing anomaly groups and comparing them with historical groups to determine the confidence level, and using an elastic weight formula to update the graph edge weights when the confidence level is high, the semantic knowledge graph can accumulate effective audit experience, dynamically adapt to changes in business semantics, and filter low-confidence associations to ensure the accuracy of the graph, providing more accurate knowledge support for subsequent audits and improving audit efficiency and reliability.
[0107] In one embodiment, before generating a task sequence corresponding to the audit task based on the audit task and the semantic knowledge graph corresponding to the current time in response to receiving the audit task, the method further includes: Entity relationships are extracted from the different business systems; the entity relationships include the master data, the data associations between the master data, and the timestamps corresponding to the master data; wherein, semantic alignment is performed on the master data that refers to the same data in the different business systems; Based on the timestamps, the entity relationships are arranged in a time sequence to obtain a time-series data stream; Based on the time-series data stream, determine the causal links between the master data; Based on the causal chain, determine the relation weight corresponding to each of the data associations; The semantic knowledge graph is obtained by taking the master data as the knowledge entity, the association relationship as the business relationship edge, and the relationship weight as the edge weight of the business relationship edge.
[0108] Optionally, in one specific implementation of this application, firstly, the intelligent agent calls the extraction tool to extract entity relationships from different business systems such as ERP, WMS, and SRM, including master data (such as customers and materials), data association relationships (such as the association between customers and orders), and corresponding timestamps.
[0109] Meanwhile, through multi-source entity alignment and referential resolution techniques, semantic alignment is performed on master data that refer to the same thing in different systems (such as "customer ID" in ERP and "buyer ID" in OMS) to eliminate ambiguity of synonyms with different names or the same name with different meanings.
[0110] Next, the integrated entity relationships are arranged chronologically according to the extracted timestamps, forming an ordered time-series data stream that clearly presents the time evolution trajectory of the master data association.
[0111] Then, the planning agent calls the temporal causal module to learn the causal chain of the temporal data stream. By analyzing the temporal sequence relationship and dependency logic between master data, it can discover real causal links such as "customer-order-invoice" and eliminate false associations caused only by time synchronization.
[0112] Then, based on the credibility of the causal link and the importance of the business, the relationship weight corresponding to each data relationship is quantitatively calculated, and the weight reflects the business priority of the relationship.
[0113] Finally, the master data is used as the knowledge entity, the data relationship is used as the business relationship edge, and the calculated relationship weight is used as the initial edge weight of the business relationship edge, which are then integrated to form a semantic knowledge graph.
[0114] In these alternative embodiments, by extracting entity relationships from multiple systems and semantically aligning them, and then arranging them temporally and mining causal links, weights are assigned to the relationships to construct a semantic knowledge graph. This can eliminate ambiguity in the master data of multiple systems, clarify the causal relationships and importance between data, provide accurate and dynamic business semantic support for the generation of subsequent audit task sequences, and improve the targeting and efficiency of audits.
[0115] It should be noted that the various optional implementation methods described in the embodiments of this application can be combined with each other or implemented individually without conflict, and the embodiments of this application do not limit this.
[0116] The data auditing method of this application is described below with a complete embodiment.
[0117] like Figure 2As shown, this application achieves parallel perception and collaborative decision-making through a multi-agent system with clearly defined roles, utilizes a dynamically evolving semantic knowledge graph to understand the business context in real time, and leverages zero-shot causal reasoning to locate the causes of anomalies in the absence of historical samples, thereby constructing an adaptive, traceable, and highly efficient intelligent master data quality auditing system. Specifically, it can be divided into three logical stages: The first stage is initialization and task reception: After the user submits the master data audit task to the system, the system creates role instances of planning, execution, and verification agents, builds these agents and configures toolkits and system messages, completes the collaboration preparation, and generates the agent set, initialization toolkit and empty result cache accordingly.
[0118] The second stage is a multi-agent collaborative loop: the planning agent understands the task and formulates an execution plan, the execution agent selects tools to execute the task, and the verification agent evaluates the results and provides feedback, forming a closed-loop iteration of "planning-execution-verification-feedback" until the result meets the quality threshold, corresponding to the loop logic of sub-task plan generation, intermediate result output and multi-dimensional scoring feedback.
[0119] The third stage is result aggregation, evaluation, and archiving: the outputs of subtasks are merged in the order of dependencies to form a complete result, the intelligent agent is verified to score the result in multiple dimensions, the collaborative trajectory is then stored and archived, and finally the audit results with scores and an interpretable report are returned to the user, corresponding to the process of result merging, quality scoring, and trajectory storage.
[0120] Specifically, (I) Initialization and Task Reception: After a user submits an audit task, the system immediately generates three types of intelligent agent instances: "planning," "execution," and "verification," and binds them with toolkits, role permissions, and message buses to complete the initialization of the collaborative environment. Given a master data audit task T, the system generates an intelligent agent set A={P,E,V}, where P is the planning intelligent agent, E is the execution intelligent agent, and V is the verification intelligent agent; the initialization toolkit M and the empty result cache R0= An empty result cache is a temporary table (or memory object) that is initially empty, used to accumulate intermediate results generated in each round of agent collaboration.
[0121] (II) Multi-Agent Collaborative Loop: The planning agent decomposes the overall task into parallel or dependent subtasks, dynamically generating an execution plan (i.e., a task sequence). The execution agent: Based on the plan, it invokes tools such as data exploration, consistency verification, and anomaly detection to generate intermediate results. The verification agent: It scores the intermediate results in multiple dimensions (accuracy, completeness, consistency) and provides real-time feedback to the planning and execution agents, triggering adaptive adjustments to the plan or parameters, forming a closed loop of "plan-execution-verification-feedback" until the results converge. Let the subtask plan π output by the planning agent in the k-th iteration be... kThe intermediate result r returned by the executing agent k and verify the multidimensional quality score s given by the intelligent agent. k The following are examples: π k = P(T, R k-1 (1) r k = E(π k ; M) (2) s k = V(r k ),s k ∈ [0,1] (3) Where T refers to the original audit request submitted by the user (i.e., the audit task); R k-1 The term "r" refers to the set of all intermediate results generated by the (k-1)th round of agent collaboration; "M" refers to the set of tools that the executing agent can call, including ETL tools (cross-system data extraction and standardization), SQL tools (cross-database join queries and difference filtering), code executors (zero-shot causal inference computation), etc. The executing agent processes the data by calling the corresponding tools to generate intermediate results r. k ; If s k If the value is greater than or equal to τ (the quality convergence threshold), then stop the loop and proceed to stage (3); otherwise, update R. k = R k-1 ∪ {r k The loop continues. This closed loop guarantees that the quality of the result increases monotonically with k, and the expected error exponent decreases as follows: E[1-s k ] ≤ ρ k ,ρ∈(0,1) (4) In formula (4), ρ is the error attenuation coefficient; E[·] is the expectation operator, which represents the average error.
[0122] (III) Result Aggregation, Evaluation, and Archiving: The outputs of each subtask are merged according to the dependency graph (i.e., the relationships between subtasks in the task sequence), and the agent is verified to perform a final multi-dimensional score on the merged result. The system stores the entire collaborative link, decision trajectory, and interpretable report in the "Storage Record Module" and returns the final answer R* with a score to the user, achieving full traceability. According to the dependency graph G, {r k Perform topological merging: = {k∈topo(G)} r k (5) Final score = V( ), along with the trajectory {(π k , rk , s k Store in the storage module and return: Output = ( , , Trace). (6) In formula (5), topo(G) represents the node sequence obtained by performing topological sorting on the subtask dependency graph G; Operators that represent merging results in dependency order; This represents the final, aggregated, complete audit result. In formula (6), The overall quality score of the final result is represented by Trace; Trace represents the complete collaborative trajectory {(π k , r k , s k )}, used for interpretable archiving.
[0123] like Figure 3 As shown, automated quality control of master data across business systems is achieved through dynamic semantic knowledge graphs and zero-shot causal reasoning technology. Essentially, it treats the "audit task" as a multi-agent collaborative game, completing the entire automated process from "anomaly detection" to "root cause localization" and then to "strategy self-optimization" under the dual drive of dynamic semantic knowledge graphs and zero-shot causal reasoning. The three-agent collaborative process is as follows: (I) Planning Intelligent Agent: Task Semantic Parsing and Policy Generation For a natural language prompt Q (e.g., "Find materials with the same description but different codes"), the planning agent uses a dynamic semantic knowledge graph G(t) to parse Q into cross-system field mappings and business rules, outputting an executable subtask sequence π = {t1, t2, ..., t}. n}, each t i Includes target table, fields, and weight w i The semantic parsing calculation formula is as follows: q = Enc(Q)∈ d (7) In formula (7), q refers to the semantic vector obtained after encoding the natural language auditing task Q; Enc(·) represents the semantic encoding function; d : Represents the d-dimensional real space where q is located, where d is a pre-defined vector dimension (such as 128, 256, 512), which must be consistent with the embedding dimension of entities and relations in the dynamic semantic knowledge graph (G(t)).
[0124] The formula for extracting association rules is as follows: R ij = σij · (ent i ,ent j ;G(t))(8) In formula (8), ( Let G(t) be the graph embedding inner product, and G(t) be the dynamic semantic knowledge graph at time t, where nodes are the main data entities ent and edge weights σ. ij σ represents the degree of semantic equivalence across systems. ij When the value is ≥τ, establish field mapping and generate subtask t. i and weight w i = softmax(σ ij ) is used for subsequent parallel scheduling. Its advantages are as follows: (a) Zero-sample start-up: No historical annotation is required, and the graph provides semantic alignment such as "ERP.material description ≈ OMS.product name"; (b) Dynamic strategy update: When a new system is connected, the graph is automatically expanded, and the planning body instantly recognizes the mapping of new fields without the need for manual rewriting of rules.
[0125] (ii) Executing intelligent agent: tool invocation + zero-shot causal inference Given subtask t i Given the current graph G(t), the main actions of the agent are as follows: (1) Call the tool ψ ∈ Ψ (ETL, SQL, code executor) to complete cross-system data extraction and standardization; Ψ is the tool package.
[0126] (2) For candidate abnormal records (x) i , x j Perform zero-sample causal inference: (9) In formula (9), X represents candidate reasons (such as "missing logistics tracking number"); Y represents business indicators (such as "inventory quantity deviation"), i.e., abnormal data, do( ) represents differential calculation; ACE(t) i The value of ) represents the average causal influence of the candidate cause (X) on the abnormal data (Y) in the i-th subtask; do(X=1) is the "interference factor" in causal inference, which means forcing the candidate cause X to be in the "existence / occurrence" state (i.e., actively setting the value of X to 1) to exclude interference from other confounding factors (such as "user access rights" and "system synchronization frequency"); do(X=0) means "forcing the candidate cause X not to occur (X=0)". When |ACE|>ε and the significance p<0.05, the causal chain is determined to be valid; (3) Output anomaly-root cause pair set r(t) i ); Its advantages are as follows: (a) Unlabeled reasoning: No need for historical abnormal samples, differential calculation can be completed directly by using the semantic structure and domain common sense in the graph; (b) Cross-system semantic consistency: The graph provides real-time mapping of "TMS.WaySlip Number → WMS.WarehouseInbound Number" to avoid misjudgment caused by "same name, different meaning".
[0127] (III) Validation of the agent: result verification and closed-loop feedback For the result r(t) of the agent execution i By comparing the historical repair cases in the memory module with the data, the reliability score of the agent's computational results is verified. s = cos( r(t i ), Mhist ) ∈[0,1](10) In formula (10), s is the confidence score; Mhist is the historical repair vector of the same type of anomaly in the memory module. If s ≥ τ → pass; otherwise, trigger "replanning" and update the task priority. Then, memory update is performed, and the anomaly-root cause-repair triplet is written into the graph G(t+1) = G(t) ⊕ Δg(r, s) to realize experience accumulation. Among them, Δg is the embedding increment of the current anomaly node and the repair edge, which is consolidated with elastic weight: σ ij (t+1)=βσ ij (t)+(1-β)s, β∈(0,1) controls the forgetting rate, achieving "experience accumulation". If s<τ, a replanning signal Δπ = f(s) is sent to the planning agent, adjusting the weights of the subtasks in the next round w′=w·exp(-γs), where Δπ is the adjustment instruction sent by the verification agent to the planning agent; f(·) represents the replanning mapping function, w′ is the execution weight of the subtask in the next iteration; w is the initial weight of the subtask in the current round; exp( γs) is the weight decay factor Its technical advantages are as follows: (a) Self-optimizing closed loop: Each round of audit feeds back into the graph, making subsequent reasoning more and more accurate; (b) Explainable output: The causal chain is persisted in the form of graph edges, which can directly show the auditor the path of "missing logistics order → inventory deviation".
[0128] like Figure 3As shown, this application designs a "dynamic, temporally aware knowledge graph" generation framework for AI agents. The core objective is to enable machines to "remember the past, understand the present, and reason about the future" like humans in a continuously changing data environment, providing a complete support path for master data quality auditing in multi-agent collaboration. Unlike traditional static knowledge bases or one-time RAG solutions, it incorporates "time" as a first-class citizen into the data model, allowing the graph to automatically grow, correct, and retain a complete history as business evolves. The following is a description of the layered design of the Dynamic Semantic Knowledge Graph (DSKG): "Data Access Layer (Collaboration Entry) → Graph Construction Layer (Knowledge and Rule Collaboration) → Computational Reasoning Layer (Intelligent Decision-Making) → Application Service Layer (Output)". (i) Data Access Layer: The “Data Collaboration Entry Point” for Multi-Agent Collaboration This layer includes Representational State Transfer (REST) interfaces, WebSocket, Management and Control Platform (MCP) services, and Remote Procedure Call (RPC) interfaces, serving as the infrastructure for data interaction and task collaboration among multiple agents. In master data quality auditing, different agents (such as the agent responsible for data collection and the agent executing rule verification) need to utilize these interfaces. (a) Data synchronization: Real-time acquisition of raw data from the main data source (such as business systems, ETL tools); (b) Task collaboration: Transmit audit task instructions (such as “perform integrity audit on the master data of system X”) through protocols such as WebSocket, or implement cross-agent rule calls through RPC (such as agent A calling agent B’s “semantic consistency verification” service).
[0129] (ii) Graph Construction Layer: The "Knowledge and Rule Collaboration Layer" for Multi-Agent Cooperation The three main modules of this layer (extraction and fusion, temporal causality, and dynamic evolution) are the core knowledge carriers for multi-agent collaboration, providing support for auditing in terms of rules, causal relationships, and dynamic updates. (a) Extraction and Fusion Module: Through "entity-relationship-practice extraction," "reference resolution," and "multi-source entity alignment," the module integrates structured (e.g., database tables) and unstructured (e.g., documents, logs) information from master data to form a unified knowledge graph. Multiple agents can share "relationship rules between master data entities" (e.g., the association constraints of "customer-order-product") based on this graph, avoiding redundant modeling.
[0130] (b) Temporal Causality Module: Through "event temporalization," "causal chain learning," and "causal edge weight update," the temporal dimension rules of the master data are sorted out (such as "order creation time must be earlier than payment time"). When multiple agents collaborate, data anomalies can be judged based on causal chains (such as agent A detecting "order creation time is later than payment time," agent B calling the causal chain rules to confirm whether it is an anomaly).
[0131] (c) Dynamic Evolution Module: Through "semantic drift detection," "conflict resolution," and "version snapshots," it addresses dynamic changes in master data (such as updates to business rules and additions to data sources). When multiple agents collaborate, it can synchronize "semantic drift" (such as "changes in the meaning of the customer address field") and "conflict rules" (such as "conflicts in the definition of 'order status' in different systems") in real time, ensuring that audit rules always match the business status.
[0132] (III) Computational Reasoning Layer: The "Intelligent Decision Engine" for Multi-Agent Collaboration This layer's graph computing engine, rule engine, machine learning model, and zero-shot causal inference engine constitute an intelligent toolset for multi-agent auditing tasks. (a) Graph computing engine: Based on the knowledge graph structure, multiple agents can collaboratively perform "path analysis" (such as tracing the source and flow of master data and judging data integrity) or "subgraph matching" (such as detecting whether master data conforms to predefined "business subgraph" rules).
[0133] (b) Rule Engine: Multiple agents share the "Master Data Quality Rule Base" (such as integrity, consistency, and accuracy rules), and the rule engine automatically executes rule verification (such as agent A detecting "customer name cannot be empty", and agent B detecting "order amount must be consistent with payment amount").
[0134] (c) Machine learning model: Multiple agents can collaboratively train an “anomaly detection model” (such as a “master data anomaly classifier” trained based on historical data) or share a “semantic understanding model” (such as parsing the natural language description of master data and judging semantic consistency).
[0135] (d) Zero-sample causal reasoning engine: When encountering a new scenario without prior rules (such as abnormal master data caused by sudden business changes), multiple agents can collaboratively call the causal reasoning engine to deduce the cause of the abnormality based on the "causal relationship" (such as whether the "abnormal order status" is caused by the "payment system failure"), and realize intelligent auditing under "zero sample".
[0136] (iv) Application Service Layer: "Output and Human-Machine Collaboration Layer" of Multi-Agent Collaboration The graph query, semantic interface, and visualization functions of this layer serve as the external output and human-computer interaction interface for the results of multi-agent collaboration. (a) Graph query: The "master data quality audit results" (such as abnormal data list and rule matching results) generated by multi-agent collaboration can be called by business personnel or downstream systems (such as data governance platform) through the graph query interface to realize the "traceability and reusability of audit results".
[0137] (b) Semantic interface: Multiple agents can interact with business personnel based on the "semantic interface" (e.g., agent A parses the business personnel's instruction "query the data quality of order X", agent B executes the query and returns the result), reducing the threshold for human-machine collaboration.
[0138] (c) Visualization: The results of master data quality analysis (such as abnormal data distribution and rule execution efficiency) of multi-agent collaboration can be presented through visualization tools to help business personnel quickly locate problems (such as "the master data integrity audit pass rate of a certain system is less than 90%).
[0139] Figure 4 The complete technical architecture of the zero-shot causal inference engine is presented, consisting of two main modules: the core causal inference process and zero-shot capability support. Through multi-stage collaboration, it achieves the derivation from observed data to causal effect estimation. This architecture, through a causal inference chain of "observed data → representation → discovery → mapping → inference → estimation," combined with zero-shot capability support of "pre-trained knowledge base + prior constraints," achieves: (a) generalization ability: causal inference can quickly adapt to new domains / new data without requiring a large amount of labeled data, thanks to prior knowledge and constraints; (b) efficiency improvement: the knowledge base and constraints reduce blind searches, accelerating causal structure mining and effect estimation; (c) result credibility: prior knowledge and constraints ensure that the inference process conforms to domain logic, improving the reliability of causal effect estimation. Its system architecture is described below: (I) Core causal reasoning process (logical chain from input to output): 1. Input the raw observation data, denoted as: (11) in As a feature, The outcome variable has no intervention information; n is the sample size. 2. Causal Representation Learning (Extracting Data Features): The goal is to map observed data to a causal latent space, distinguishing between correlation and causality. Pre-trained models (such as Large Language Models (LLMs) or Causal Transformers) are used to extract representations. (12) Z i It is the master data feature X i The low-dimensional dense vector obtained after causal representation learning; X i This refers to the original attributes of master data extracted from multiple systems. Through regularization or contrastive learning, the representation is improved. Preserve causal information rather than confounding bias; θ represents the causal representation learning function, which is derived from a pre-trained model with regularization constraints, and θ is the learnable parameter of the model.
[0140] 3. Causal Discovery Engine (Uncovering causal relationships between variables): Under zero-sample conditions, causal structures between variables are identified using causal representations and prior knowledge. Common methods include conditional independence tests or structural equation modeling. By combining semantic priors provided by a pre-trained causal knowledge base, the search space is reduced; X j →X k Represents the master data feature X j It is X k The direct causal cause; X j It is a candidate causal feature, X k These are target features that may be affected by them; Excluding X j and X k In addition, all other master data feature sets that may simultaneously affect both (i.e., hybrid features); "∣" is a condition symbol, read as "under the condition of...".
[0141] 4. Construct a causal graph (visualize the causal structure): Organize the discovered causal relationships into a directed acyclic graph (DAG) or a structural causal model (SCM): , , (13) in, Let represent a set of vertices, with d nodes in total, where each node represents a random variable or feature; Represents a set of edges, where each edge is an ordered pair ( ), indicating from point to The directed edges in the causal graph yes The direct cause; This is a cause-and-effect diagram; 5. Intervention and Counterfactual Reasoning (Simulating Intervention and Deriving Results): Based on the causal diagram, perform differentiation or counterfactual reasoning to estimate the intervention effect: (14) in, As a confounding factor, it satisfies the backdoor criterion. Within the same set of background factors... If below, Forced to set , The value is determined by the structural equation. Completely determined; through multiple sampling The counterfactual distribution can then be obtained. ,in, The structural equation representing the abnormal results of the master data is the core function that describes the causal relationship between "candidate cause and abnormal result"; The set of background confounding factors affecting the abnormal outcome Y encompasses all factors that may indirectly influence Y but are not explicitly included as candidate causes; x: refers to the intervention value for the candidate cause variable X; U: is A simplified expression; It means "follows a certain probability distribution"; The joint probability distribution representing the background confounding factor U is automatically learned based on historical operation logs and business scenario features from a dynamic semantic knowledge graph. It is a counterfactual conditional probability distribution. The counterfactual reasoning used to estimate causal effects is generated by the following formula through the model: (15) 6. Output Causal Effect Estimation (Quantifying the Impact of Intervention): The final output is the causal effect size, such as the average treatment effect (ATE). (16) Or Individual Treatment Effect (ITE): (17) Among them, ITE i Represents the candidate cause X in the i-th master data sample. i The abnormality index Y of this sample i The magnitude of the individual causal effect; Y i : The anomaly quantification index for the i-th sample, corresponding to the anomaly state of the specific master data record; do(X i =1) / do(X i =0): For the interference estimator of the i-th sample, force the candidate cause X of that sample respectively. i It is in the state of "occurring" or "not occurring".
[0142] (II) Zero-Shot Capability Support (Key Module for Enhancing Reasoning Generalization and Efficiency): A pre-trained causal knowledge base provides prior causal knowledge to assist in data representation and causal relationship mining. Causal prior constraints standardize the search scope and reasoning boundaries of causal discovery, improving efficiency and reliability. The core value of the Zero-Shot Causal Inference Engine (ZSCR) is to achieve efficient causal reasoning in scenarios with little or no data through knowledge guidance and process collaboration, outputting accurate quantification results of causal effects.
[0143] This application transcends the traditional scope of "data quality tools" and becomes a "semantic immune system" for master data operations. It is applicable to scenarios with strong master data compliance, such as power, finance, manufacturing, healthcare, industrial internet, military, and retail e-commerce. Its main technical effects are as follows: (1) Dynamic semantic knowledge graph-driven self-evolution of audit rules: Construct a three-dimensional knowledge graph of master data entities, business behaviors, and temporal states, and absorb operation logs from systems such as ERP and CRM in real time (such as customer merging and supplier discontinuation). Use a temporal graph neural network to capture semantic drift patterns. When business rules change (such as the definition of "distributor" being changed from "payment cycle < 30 days" to "credit rating > A"), DSKG automatically triggers the graph neural rule engine (GNRE) to generate new audit rules without manual rewriting. Achieve 100% automation of the rule lifecycle; (2) Zero-sample causal reasoning to deal with unknown anomalies: The master data anomaly detection is transformed into a counterfactual causal reasoning problem. A pre-trained large model is used to learn the causal structure of "normal master data" (such as the dependency chain of "customer tax number → invoice header → contract number"). Counterfactual explanations are generated for unseen anomalies (such as "tax number exists but invoice header is empty") ("if the tax number is valid, the probability of missing header is <0.1%"). Through causal representation decoupling technology, the anomaly pattern is decomposed into business factors (such as "distributor cross-selling") and data factors (such as "field truncation") to achieve accurate positioning in zero-sample scenarios.
[0144] (3) Human-machine collaborative "semantic sandbox" repair mechanism: For high confidence anomalies (such as "the supplier's address conflicts with the business registration location"), the system automatically calls the toolchain to complete the correction; for low confidence anomalies (such as "the customer's industry classification is ambiguous"), the semantic sandbox is activated to generate a natural language audit report (including DSKG visualization path and counterfactual explanation). Business personnel can use conversational BI (such as ChatBI) to ask follow-up questions in natural language, and the system returns the graph reasoning path in real time, which greatly reduces the workload of manual review; (4) Multi-agent collaboration: Through the division of labor and negotiation among role-based intelligent agents (such as data collection agent, consistency agent, and causal reasoning agent), parallel perception and decision fusion of distributed master data quality are achieved. This transforms master data auditing from "high expert dependence" to "automated intelligent service," reducing enterprise-level data governance costs.
[0145] Figure 6 A schematic diagram of the structure of a data auditing device provided in another embodiment of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0146] Reference Figure 6 The data auditing device may include: The generation module 601 is used to generate a task sequence corresponding to the audit task in response to receiving an audit task, based on the audit task and the semantic knowledge graph corresponding to the current time; the semantic knowledge graph includes multiple knowledge entities and business relationship edges between the multiple knowledge entities; the multiple knowledge entities are master data stored in different business systems; The acquisition module 602 is used to acquire abnormal data in the target master data indicated by the task sequence based on the anomaly checking tool corresponding to the task sequence; The determination module 603 is used to perform causal correlation analysis on the abnormal data based on the business relationship edge to determine the abnormal cause of the abnormal data; The update module 604 is used to update the semantic knowledge graph based on the confidence level of the abnormal data and the cause of the abnormality, and to use the updated semantic knowledge graph to execute the next received audit task.
[0147] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application, and are devices corresponding to the above-mentioned methods. All implementation methods in the above-mentioned method embodiments are applicable to the embodiments of this device. For details on its specific functions and the technical effects it brings, please refer to the method embodiment section, which will not be repeated here.
[0148] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0149] Figure 7 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0150] The device may include a processor 701 and a memory 702 storing program instructions.
[0151] When processor 701 executes the program, it implements the steps in any of the above method embodiments.
[0152] For example, the program can be divided into one or more modules / units, one or more of which are stored in memory 702 and executed by processor 701 to complete this application. The one or more modules / units can be a series of program instruction segments capable of performing a specific function, which describe the execution process of the program in the device.
[0153] Specifically, the processor 701 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0154] Memory 702 may include mass storage for data or instructions. For example, and not limitingly, memory 702 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 702 may include removable or non-removable (or fixed) media. Where appropriate, memory 702 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 702 is non-volatile solid-state memory.
[0155] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.
[0156] The processor 701 implements any of the methods described in the above embodiments by reading and executing program instructions stored in the memory 702.
[0157] In one example, the electronic device may also include a communication interface 703 and a bus 710. The processor 701, memory 702, and communication interface 703 are connected via the bus 710 and communicate with each other.
[0158] The communication interface 703 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0159] Bus 710 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 710 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0160] Furthermore, in conjunction with the methods in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores program instructions; when these program instructions are executed by a processor, they implement any of the methods in the above embodiments.
[0161] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0162] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0163] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.
[0164] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0165] The functional modules shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on machine-readable media or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable media" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer grids such as the Internet, intranets, etc.
[0166] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0167] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0168] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A data auditing method, characterized in that, The method includes: In response to receiving an audit task, a task sequence corresponding to the audit task is generated based on the audit task and the semantic knowledge graph corresponding to the current time; the semantic knowledge graph includes multiple knowledge entities and business relationship edges between the multiple knowledge entities; the multiple knowledge entities are master data stored in different business systems; Based on the anomaly detection tool corresponding to the task sequence, obtain the abnormal data that exists in the target master data indicated by the task sequence; Based on the business relationship edges, perform causal correlation analysis on the abnormal data to determine the cause of the abnormality of the abnormal data; Based on the confidence level of the abnormal data and the cause of the abnormality, the semantic knowledge graph is updated, and the updated semantic knowledge graph is used to execute the next received audit task.
2. The method according to claim 1, characterized in that, The anomaly detection tool includes at least one of a graph computing engine, a rule engine, and an anomaly detection model; the rule engine includes a master data quality rule library. The anomaly detection tool based on the task sequence obtains abnormal data in the master data indicated by the task sequence that contains anomalies, including: Based on the anomaly detection tool corresponding to the task sequence, output the first data in the master data indicated by the task sequence that contains an anomaly; Based on the first data, the abnormal data is determined; Specifically, when the anomaly detection tool is the graph computing engine, the source path of the target master data is determined based on the semantic knowledge graph; the completeness of the target master data is determined based on the source path; and the first data is determined from the target master data based on the completeness. When the anomaly detection tool is the rule engine, a target quality rule matching the target master data is obtained from the master data quality rule library; based on the target quality rule, the first data is determined from the target master data; the target quality rule is used to determine at least one of the completeness, consistency, and accuracy of the target master data; When the anomaly detection tool is the anomaly detection model, the target master data is input into the anomaly detection model to obtain the first data.
3. The method according to claim 1, characterized in that, The step of performing causal correlation analysis on the abnormal data based on the business relationship edge to determine the cause of the abnormality includes: Based on the business relationship edge, at least one candidate cause associated with the abnormal data is determined; The anomalous data and the candidate causes are mapped to a causal latent space to obtain causal features; the causal latent space represents the potential causal relationship between the anomalous data and the candidate causes. Based on the causal characteristics and the preset causal knowledge base, the causal relationship between the abnormal data and the candidate causes is determined; Based on the causal relationship, the causal effect value of each candidate cause on the abnormal data is determined; the causal effect value is used to characterize the degree of influence of the candidate cause on the abnormal data. Based on the causal effect value, the abnormal cause is determined from each of the candidate causes.
4. The method according to claim 3, characterized in that, The step of determining the causal relationship between the abnormal data and the candidate causes based on the causal features and a preset causal knowledge base includes: Using the abnormal data as the outcome variable and the candidate causes as the variables to be analyzed, a structural equation model is used to determine the potential causal relationship between the outcome variable and the variables to be analyzed. The potential causal relationships are filtered using the causal constraints stored in the causal knowledge base to obtain the causal relationship between the abnormal data and the candidate causes.
5. The method according to claim 3, characterized in that, The step of determining the causal effect value of each candidate cause on the abnormal data based on the causal relationship includes: Based on the causal relationship, a relationship graph is constructed, where the vertices of the relationship graph represent the abnormal data and the candidate causes, and the edges of the relationship graph represent the causal relationship between the abnormal data and the candidate causes. Based on the relationship diagram, an intervention operation is performed on the candidate causes to determine the counterfactual conditional probability distribution of the abnormal data under the intervention operation; the intervention operation is used to change the value of the variable corresponding to the candidate cause to simulate the impact of the value change on the abnormal data; the counterfactual conditional probability distribution is used to characterize the probability distribution of the abnormal data as the outcome variable under the intervention operation. Based on the counterfactual conditional probability distribution, determine the causal effect value of each candidate cause on the abnormal data.
6. The method according to claim 5, characterized in that, The step of determining the abnormal cause from each of the candidate causes based on the causal effect value includes: The abnormal data is divided into multiple data blocks according to a preset time window. For the first variable of the candidate cause corresponding to each data block, the first variable is randomly rearranged and the timestamp of the first variable is reassigned, while keeping the timestamps of other data in the data block other than the first variable unchanged, to obtain the control sample set; A null distribution is constructed based on the control sample set, the position of the cumulative distribution function of the causal effect value in the null distribution is determined, and the significance corresponding to the causal effect value is obtained; the null distribution is used to characterize that the first variable has no causal relationship with the outlier data. If the significance is greater than a preset significance threshold, the candidate cause corresponding to the causal effect value is determined as the abnormal cause.
7. The method according to claim 1, characterized in that, The step of updating the semantic knowledge graph based on the confidence levels of the abnormal data and the causes of the abnormality includes: Based on the stated cause of the anomaly and the stated data, an anomaly group is constructed; The similarity between the abnormal group and the historical group is determined as the confidence level; the historical master data corresponding to the historical group matches the target master data; If the confidence level is greater than a preset confidence threshold, the edge weights of the business relationship edges corresponding to the abnormal data are updated based on a weight update formula; the weight update formula is σ. ij (t+1)=βσ ij (t)+(1 β)s; where β∈(0,1), σ ij (t) represents the original edge weight of the business relationship edge at the current time t, and s represents the confidence level.
8. The method according to claim 1, characterized in that, Before generating a task sequence corresponding to the audit task based on the audit task and the semantic knowledge graph corresponding to the current time in response to receiving the audit task, the method further includes: Entity relationships are extracted from the different business systems; the entity relationships include the master data, the data associations between the master data, and the timestamps corresponding to the master data; wherein, semantic alignment is performed on the master data that refers to the same data in the different business systems; Based on the timestamps, the entity relationships are arranged in a time sequence to obtain a time-series data stream; Based on the time-series data stream, determine the causal links between the master data; Based on the causal chain, determine the relation weight corresponding to each of the data associations; The semantic knowledge graph is obtained by taking the master data as the knowledge entity, the association relationship as the business relationship edge, and the relationship weight as the edge weight of the business relationship edge.
9. A data auditing device, characterized in that, The device includes: The generation module is used to respond to receiving an audit task and generate a task sequence corresponding to the audit task based on the audit task and the semantic knowledge graph corresponding to the current time. The semantic knowledge graph includes multiple knowledge entities and business relationship edges between the multiple knowledge entities. The multiple knowledge entities are master data stored in different business systems. The acquisition module is used to acquire abnormal data that exists in the target master data indicated by the task sequence based on the anomaly inspection tool corresponding to the task sequence; The determination module is used to perform causal correlation analysis on the abnormal data based on the business relationship edge to determine the abnormal cause of the abnormal data; The update module is used to update the semantic knowledge graph based on the confidence level of the abnormal data and the cause of the abnormality, and to use the updated semantic knowledge graph to execute the next received audit task.
10. An electronic device, characterized in that, include: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the data auditing method as described in any one of claims 1-8.