Data management-oriented data integration and quality assurance collaborative verification method and system and medium

By generating integration task trigger identifiers, pushing data verification requests, and invoking data verification rules, the problem of disconnection between the data integration group and the assurance group is solved, enabling real-time data quality assurance and intelligent decision-making, and improving the efficiency and transparency of government and enterprise data governance.

CN121808276APending Publication Date: 2026-04-07SUZHOU LONGSHI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In government and enterprise data governance, the disconnect between the data integration team and the data assurance team often leads to quality problems being discovered only after integration is completed. This results in high costs for tracing and repairing the problem, a lack of efficient collaboration mechanisms, and an inability to monitor the health status and progress of data during the integration phase in real time, which severely restricts the overall efficiency of data governance.

Method used

By generating an integrated task trigger identifier, pushing data verification requests to the assurance task queue, calling data verification rules for parsing and verification, generating quality verification tags, and updating the node status of the business execution process, the system utilizes workflow engines, state machine algorithms, and rule engines to achieve fully automated monitoring and intelligent decision-making throughout the entire process.

Benefits of technology

It enables real-time data quality assurance and intelligent decision-making, significantly improves data governance efficiency, enhances transparency and control capabilities throughout the entire process, and improves the reliability, efficiency, and collaboration efficiency of business execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808276A_ABST
    Figure CN121808276A_ABST
Patent Text Reader

Abstract

The invention relates to a data management-oriented data integration and quality assurance collaborative verification method and system and a medium, and relates to the technical field of data management. The data integration and quality assurance collaborative verification method comprises the following steps: associating and monitoring integration nodes and assurance nodes according to a service execution process, and generating an integration task trigger identifier; according to the integrated task trigger identifier, the integrated node pushes a data verification request to a guarantee task queue to generate a queue voucher; according to the queue voucher, the guarantee node calls a data verification rule to analyze and verify the data verification request, triggers an exception handling decision, and generates a quality verification label; updating the node state of the business execution process according to the quality verification label, and matching and executing a process operation decision; through deep embedding of a data guarantee link and an active access data integration process, preposition and left shift of quality verification are realized, problems are quickly found and positioned from the source, data management efficiency is improved, and transparency and management and control capability of the whole process are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data governance technology, and in particular to a data integration and quality assurance collaborative verification method, system and medium for data governance. Background Technology

[0002] In government and enterprise data governance practices, the work of the data integration team and the data assurance team (or data quality team) is often disconnected, resulting in fragmented processes. The current typical model is that the data integration team is responsible for collecting data from source systems to the data platform, and then the data assurance team conducts quality checks and verifications on the implemented data.

[0003] This sequential and passive "integrate first, govern later" model has significant drawbacks: First, it is impossible to control the key stages of data entering the lake and warehouse in advance, and quality problems are often only discovered after integration is completed, resulting in high costs for tracing and repair. Secondly, the two teams lack an efficient collaboration mechanism and often communicate through unstructured methods such as documents and emails, resulting in slow response to problems and easy shirking of responsibility.

[0004] Existing patents disclose a method for optimizing and monitoring the data governance process, relating to the field of data governance technology; including: Step 1: Accessing dynamic data: configuring data sources, supporting access to multiple data sources, including interface data and database data; Step 2: Configuring rules: configuring execution rules and monitoring rules for data governance based on a dynamic rule engine, supporting real-time updates and adjustments; Step 3: Real-time monitoring: monitoring the data governance process in real time, identifying abnormal data and violations; Step 4: Intelligent analysis: using machine learning algorithms to analyze monitoring data, predict potential problems and provide optimization suggestions; Step 5: Visualization: displaying monitoring results and analysis reports through visualization tools, facilitating users to quickly understand the progress and status of data governance; Step 6: Automatic optimization: automatically prompting data governance strategies or triggering alarm mechanisms based on analysis results.

[0005] The existing technical solutions described above have the following drawbacks: 1. Traditional approaches lack a unified, holistic perspective, making it impossible to monitor the health status and governance progress of data during the integration phase in real time. This hinders centralized cross-functional control and efficient collaboration, severely restricting the overall efficiency of business-theme-oriented data governance. Summary of the Invention

[0006] To address the shortcomings of existing technologies, the purpose of this application is to provide a collaborative verification method, system, and medium for data integration and quality assurance in data governance. By deeply embedding the data assurance process and proactively reaching the data integration process, it achieves the forward and leftward shift of quality verification, enabling rapid discovery and location of problems from the source, significantly improving data governance efficiency. Through a visualized global control view, it enhances the transparency and control capabilities of the entire process.

[0007] This was achieved using the following technical solutions: Firstly, this application provides a collaborative verification method for data integration and quality assurance in data governance, including: Based on the business execution process, the integration nodes and assurance nodes are associated and monitored to generate integration task trigger identifiers; Based on the integration task trigger identifier, the integration node pushes a data verification request to the security task queue and generates a queue credential. Based on the queue credentials, the node ensures that the data verification rules are parsed and the data verification request is verified, triggers the exception handling decision, and generates the quality verification label. Update the node status of the business execution process based on the quality verification label, and match and execute the corresponding process operation decisions.

[0008] By adopting the above technical solution, the workflow engine algorithm is used to associate and monitor integration nodes and assurance nodes in real time, automatically generate integration task trigger identifiers, and drive the message queue to push data verification requests and generate queue credentials. Subsequently, the assurance node calls the rule engine algorithm to parse and verify the request data and generate quality verification tags. Then, the state of the business process nodes is updated based on the state machine algorithm, and the decision tree algorithm is used to match and execute the corresponding process operation strategy. This achieves automated monitoring of the entire process, real-time data quality assurance, and intelligent decision-making, significantly improving the reliability, efficiency, and accuracy of business execution.

[0009] This application further specifies: based on the business execution process, it associates and monitors integration nodes and assurance nodes to generate integration task trigger identifiers, including: Based on the data processing flow, the business execution flow is deconstructed, and the data integration group and data assurance group are extracted; Based on the data flow direction, the data integration group and data assurance group are node-based, generating integration nodes and assurance nodes; Based on the data operation dependencies, the integration nodes and the guarantee nodes are associated and mapped to determine the interaction chain between nodes; The integration nodes are monitored based on the interaction chain between nodes to determine the real-time status of the nodes; Based on the real-time status of the nodes and the preset pre-association rules, an integrated task trigger identifier is generated.

[0010] By adopting the above technical solution, the business process is deconstructed and nodeified based on graph algorithms and topological sorting, and a node interaction chain is constructed through dependency mapping. Then, the real-time status of the nodes is monitored by a state machine, and finally, the integrated task triggering identifier is automatically generated based on the pre-association rules driven by the rule engine. This achieves automated orchestration, real-time status tracking and intelligent triggering of the entire process, which significantly improves the collaborative efficiency, reliability and response speed of business execution.

[0011] This application further specifies that: based on the integration task trigger identifier, the integration node pushes a data verification request to the security task queue, and generates a queue credential, including: The integrated task trigger identifier is parsed to obtain the interface dependency hash, trigger timestamp, and node verification code; The integrated nodes are verified and sorted according to the node checksum, and an integrated processing sequence is generated. The data model in the integration node is called according to the integration processing sequence, and the data source is extracted and loaded into the model by combining the interface dependency hash for verification. Based on the data source loading model, the business data source is transformed and deconstructed to extract the business data stream; Based on the metadata in the integration node, the business data flow is sorted to obtain physical data entities; Physical data entities are grouped according to data type to obtain single-type data entities; Based on the data collection time, duplicate values ​​are removed, invalid values ​​are removed, and missing values ​​are filled for each type of data entity to obtain standard data for that type. Based on the business scope, all single-category standard data are merged to obtain business standard data; The trigger timestamp, process stage code, and business standard data or single data entity or single standard data are structured and encapsulated to generate a data verification request and assign a priority mark. Based on the service load and response latency of the protection nodes, the protection routing path is selected, and a protection task queue is generated; Based on priority markers and guaranteed queue status, traffic control is applied to data verification requests, and the queuing timestamp and queuing sequence number are recorded to generate queue credentials. If the priority is marked as high, the current data verification request is inserted at the head of the guarantee task queue. If the queue status indicates that the queue capacity has reached its limit, the backpressure mechanism is triggered, and the integration node is notified to delay the push.

[0012] By adopting the above technical solution, the integration nodes are verified and sorted through dependency resolution and hash verification algorithms. The data model is called to extract business data streams. Standard data is obtained through data transformation and quality control algorithms (including deduplication, invalid value processing, and missing value filling). Then, based on the rule engine, it is encapsulated into data verification requests with priority tags. The guaranteed route is dynamically selected according to load balancing and response latency. Finally, traffic control is performed through priority queue management and back pressure mechanism to generate queue credentials. This realizes the full-process automation and intelligent scheduling of data from extraction, cleaning to delivery, which significantly improves data processing quality, system throughput, and resource utilization efficiency.

[0013] This application is further configured to: based on the queue credentials, ensure that the node calls the data verification rules to parse and verify the data verification request, trigger an exception handling decision, and generate a quality verification label, including: Based on the queue credentials, the node ensures that the enqueue timestamp and enqueue sequence number are verified. If all verifications pass, load and parse the data verification request bound to the queue credentials to obtain the process stage code and the data to be verified. The data verification rules are matched against the pre-defined data verification rules based on the process stage codes to obtain the stage verification rules; The data to be verified is parsed in a specific format and then structured extracted using a preset data pattern to obtain the data fields to be verified. Static validation is performed on the data fields to be validated according to the phase validation rules and the rule priority. If the static validation of the data field to be verified fails, an invalid data conclusion code will be generated. If not, perform business logic validation on the current data field to be validated; if the business logic validation fails, correct the dependencies and state transitions between the metadata of the current data field to be validated, extract the logical skeleton, record the rule score matrix, and generate a data part warning conclusion code or data invalid conclusion code. If not, perform data source validation on the current data field to be validated; if the data source validation fails, record the unmatched data field, generate a data validation exception conclusion code, and trigger exception handling decision; otherwise, generate a data validation complete success conclusion code. Based on the data verification conclusion code, rule score matrix, and evidence chain summary, and combined with the current data fields to be verified, a quality verification label is generated. Exception handling decisions include: Analyze the data verification anomaly conclusion code based on the interaction status between nodes to determine the cause of the anomaly; If the cause of the anomaly is an external dependency error, the logical skeleton will be re-matched a limited number of times according to the phase verification rules to determine the final data verification conclusion code. If the cause of the anomaly is a certain error, replace the data verification anomaly conclusion code with a data invalid conclusion code or a data partial warning conclusion code, and record the violation data segment; If the cause of the anomaly is a rule conflict, the stage verification rules are sorted according to the rule priority to determine the final data verification conclusion code.

[0014] By adopting the above technical solution, based on queue credential verification and rule engine, and through multi-layer verification algorithms such as priority-driven static verification, dependency analysis and data source comparison, the data to be verified is comprehensively reviewed. The system also uses intelligent decision-making mechanisms such as rule conflict resolution and logical skeleton re-matching to handle anomalies, and finally generates a quality verification label containing a conclusion code and evidence chain. This achieves a high degree of automation, intelligence and high reliability in the data verification process, and significantly improves the accuracy of data quality control.

[0015] This application is further configured to: update the node status of the business execution process based on the quality verification label, and match and execute the corresponding process operation decisions, including: The quality verification labels are parsed and mapped according to the preset state mapping rule table to determine the target node of the business execution process and generate a state update instruction. The state variables of the target node are locked, and a new state code is registered in conjunction with the state update instruction to obtain the node to be executed; Based on the business process diagram, perform topology analysis on the nodes to be executed, generate state change events, and update the node state view; The decision rule set is matched with the state change event to obtain the phase operation decision; Based on the phased operational decisions, rule reasoning is performed on the nodes to be executed to determine the node execution strategy and allocate the corresponding operational resources. If the node execution strategy is to continue execution, then release the node lock of the currently pending node, send a ready signal to the subsequent process nodes, and allocate computing resources. If the node execution strategy is to retry, then the amount of data on the node to be executed is adjusted, and a limited number of retries are initiated in the resource isolation environment; If the node execution strategy is process rerouting, then the execution path of the process instance is dynamically reconstructed, a backup backup node is loaded, and the data fields to be verified are injected. If the node execution strategy is external intervention, then create and push the node execution work order, and freeze the process clock.

[0016] By adopting the above technical solution, based on state machine mapping and rule engine, the target node is locked and its state is updated by parsing quality verification tags. State change events are generated using business graph topology, which in turn drive decision rule set matching and reasoning. Differentiated strategies such as "continue execution", "retry execution", "process rerouting" or "external intervention" are dynamically generated and executed. Resource scheduling algorithms are linked to allocate resources or reconstruct paths, realizing intelligent, automated driving and elastic management of business process status, which significantly improves the resilience, adaptability and operation and maintenance efficiency of process execution.

[0017] This application further specifies that the data validation rule invocation algorithm includes: Data quality verification rules are sharded and stored according to the rule table name, and a phase rule index is constructed. ; Among them, Index table For table name indexes, ⊕ represents distributed aggregation operations, and N... node Let k be the total number of storage nodes, k be the node index, and r be the number of nodes. i Let be the i-th data quality validation rule, R be the rule base, and table be... i Let be the i-th rule table, ↔ be the rule mapping, and Index be... field For field name indexing, DistHash is the hash function, and field... i This refers to the name of the i-th field. Based on the rule complexity and memory usage, the execution schedule of the data quality verification rules is determined to establish the rule scheduling priority. ; Where PQ() is the rule-based scheduling priority, Sort() is the elastic scheduling function, and s i For the question level, w i Here, PCost() represents the rule weight, PCost() represents the rule complexity, and MCost() represents the memory usage. Based on the rule-based heat index function, streaming rule matching is performed on the business standard data within a preset time window, and abnormal matching data is aggregated. Where d represents standard business data, Stream(T) represents the data stream within the time window T, Match() represents batch streaming matching, and Result... ri For abnormal matching data, f k For the k-th field, Ψ typei For typed validation functions, params i For business data, l hot () represents the regularity heat characteristic function, errori Warning level; The abnormal matching data is incrementally updated according to the data processing batch t to obtain the increment.

[0018] By adopting the above technical solution, storing the rule base and establishing a dynamic index through a distributed sharding algorithm, calculating rule priorities by combining elastic scheduling functions, and processing abnormal data in real time using streaming matching and incremental update mechanisms, efficient scheduling and adaptive optimization of the data verification process are achieved, significantly improving the real-time performance of large-scale data quality monitoring and system resource utilization.

[0019] Secondly, this application also provides a collaborative verification system for data integration and quality assurance in data governance, employing the following technical solution: A collaborative verification system for data integration and quality assurance oriented towards data governance, comprising a method for collaborative verification of data integration and quality assurance, including: The conversion trigger module is used to generate integration task trigger identifiers based on the association and monitoring of integration nodes and assurance nodes in the business execution process; The integration trigger module is used to push data verification requests from the integration node to the security task queue based on the integration task trigger identifier, and generate queue credentials. The quality verification module is used to ensure that nodes call data verification rules to verify data verification requests based on queue credentials and generate quality verification tags. The decision matching module is used to update the node status of the business execution process based on the quality verification label, and to match and execute the corresponding process operation decisions.

[0020] By adopting the above technical solution, the transformation trigger module monitors nodes based on state machines and rule engines and generates trigger identifiers. The driving integrated trigger module pushes verification requests using message queues and load balancing algorithms. The quality verification module outputs quality labels with the help of rule engines and multi-layer verification algorithms. The decision matching module dynamically executes process decisions based on the labels through state machine mapping and resource scheduling algorithms. This achieves fully automated and intelligent data quality control and adaptive process optimization, significantly improving the reliability, efficiency, and resilience of business execution.

[0021] Thirdly, this application also provides an electronic device, comprising: One or more processors; Memory, used to store one or more programs; When one or more programs are executed by one or more processors, the one or more processors implement any of the methods in the above scheme.

[0022] Fourthly, this application also provides a storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the data integration and quality assurance collaborative verification method for data governance as described above.

[0023] In summary, the beneficial technical effects of this application are as follows: Based on state machine mapping and rule engine, the system locks the target node and updates its state by parsing the quality verification label, drives the matching and reasoning of decision rule sets, dynamically generates and executes differentiated strategies, and performs resource allocation or path reconstruction. This achieves intelligent, automated driving and elastic management of business process state, significantly improving the resilience and adaptability of process execution. By deconstructing and node-based business processes, constructing node interaction chains through dependency mapping, monitoring real-time node status using a state machine, and finally automatically generating integrated task trigger identifiers based on pre-association rules driven by a rule engine, the system achieves automated orchestration, real-time status tracking, and intelligent triggering of the entire process, significantly improving the collaborative efficiency, reliability, and response speed of business execution. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the overall process of the collaborative verification method for data integration and quality assurance in this application; Figure 2 This is a schematic diagram of the data flow between the data integration group and the data assurance group in this application; Figure 3 This is a flowchart illustrating step S2 in the collaborative verification method for data integration and quality assurance in this application; Figure 4 This is a flowchart illustrating step S3 in the collaborative verification method for data integration and quality assurance in this application; Figure 5 This is a data processing flowchart of Embodiment 1 in this application; Figure 6 This is a schematic diagram of the data integration and quality assurance collaborative verification system in this application. Detailed Implementation

[0025] The present application will be further described in detail below with reference to the accompanying drawings.

[0026] Reference Figure 1 This application discloses a collaborative verification method for data integration and quality assurance oriented towards data governance, comprising: S1: Generate an integration task trigger identifier based on the business execution process association and monitoring integration nodes and assurance nodes; S2: Based on the integration task trigger identifier, the integration node pushes a data verification request to the security task queue and generates a queue credential; S3: Based on the queue credentials, ensure that the node calls the data verification rules to parse and verify the data verification request, trigger the exception handling decision, and generate the quality verification label; S4: Update the node status of the business execution process based on the quality verification label, and match and execute the corresponding process operation decisions.

[0027] The implementation principle of this embodiment is as follows: the integration node and the assurance node in the business execution process are monitored in real time by a state machine and a rule engine, and an integration task trigger identifier is automatically generated; then, based on a message queue and a load balancing algorithm, the data verification request is pushed to the assurance task queue and a queue credential is generated; the assurance node calls the rule engine and a multi-layer verification algorithm to parse the verification request and triggers an exception handling decision mechanism to generate a quality verification label; finally, based on the label, the node state is updated by a state mapping and decision tree matching algorithm, and the corresponding process operation decision is dynamically executed.

[0028] Preferably, step S1 includes: Based on the data processing flow, the business execution flow is deconstructed, and the data integration group and data assurance group are extracted. Figure 2 ); Based on the data flow direction, the data integration group and data assurance group are node-based, generating integration nodes and assurance nodes; Based on the data operation dependencies, the integration nodes and the guarantee nodes are associated and mapped to determine the interaction chain between nodes; The integration nodes are monitored based on the interaction chain between nodes to determine the real-time status of the nodes; Based on the real-time status of the nodes and the preset pre-association rules, an integrated task trigger identifier is generated.

[0029] In this embodiment, the business execution process is structured and modeled, decomposing it into a series of ordered integration nodes (functional units that execute core business logic) and guarantee nodes (auxiliary units that provide status monitoring, fault tolerance, and resource guarantees). The inputs, outputs, preconditions, and postconditions of each node are defined.

[0030] Establish a dependency and trigger relationship mapping between integration nodes and assurance nodes. This is typically achieved through a directed graph or rule table, including: Strong dependency: The execution of integration node A must wait for the return of a ready signal from guarantee node B; Event triggering relationship: When the protection node C detects a specific event, it triggers the integration node D to start execution; Status monitoring relationship: When the integration node E is running, it is necessary to ensure that node F continuously provides status feedback.

[0031] Real-time collection of status data from all nodes (such as "Running", "Idle", "Fault", "Data Ready") and fusion processing. Continuous evaluation by monitoring logic: Does the execution status of the integration node meet the expected process? Are the guarantee conditions of the guarantee node continuously met (such as sufficient resources and environmental security)? Is the timeliness and consistency of data flow between nodes guaranteed?

[0032] Trigger judgment based on monitoring data and predefined association rules: When all the prerequisites for a certain integration node (including the outputs from other integration nodes and the status of associated guarantee nodes) are met, the system determines that the node can be triggered.

[0033] Then, a globally unique integrated task trigger identifier is generated. This identifier typically includes: task sequence number, target node ID, trigger timestamp, dependency hash (used to verify the integrity of the trigger conditions), and priority identifier.

[0034] The generated trigger identifier is distributed to the target integration node and its associated support nodes. Upon receiving the identifier, the target integration node verifies its validity (e.g., by checking the hash) and then initiates execution. The associated support nodes simultaneously enter the enhanced monitoring mode for this task.

[0035] Preferably, refer to Figure 3 Step S2 includes: The integrated task trigger identifier is parsed to obtain the interface dependency hash, trigger timestamp, and node verification code; The integrated nodes are verified and sorted according to the node checksum, and an integrated processing sequence is generated. The data model in the integration node is called according to the integration processing sequence, and the data source is extracted and loaded into the model by combining the interface dependency hash for verification. Based on the data source loading model, the business data source is transformed and deconstructed to extract the business data stream; Based on the metadata in the integration node, the business data flow is sorted to obtain physical data entities; Physical data entities are grouped according to data type to obtain single-type data entities; Based on the data collection time, duplicate values ​​are removed, invalid values ​​are removed, and missing values ​​are filled for each type of data entity to obtain standard data for that type. Based on the business scope, all single-category standard data are merged to obtain business standard data; The trigger timestamp, process stage code, and business standard data or single data entity or single standard data are structured and encapsulated to generate a data verification request and assign a priority mark. Based on the service load and response latency of the protection nodes, the protection routing path is selected, and a protection task queue is generated; Based on priority markers and guaranteed queue status, traffic control is applied to data verification requests, and the queuing timestamp and queuing sequence number are recorded to generate queue credentials. If the priority is marked as high, the current data verification request is inserted at the head of the guarantee task queue. If the queue status indicates that the queue capacity has reached its limit, the backpressure mechanism is triggered, and the integration node is notified to delay the push.

[0036] In this embodiment, after receiving the trigger identifier, the integration node first parses the task metadata (such as the target node ID and dependency hash) carried in the identifier, then extracts the output data and process status information required for local execution, and encapsulates them into a structured verification request package according to a preset protocol format (such as JSONSchema). This request package includes at least: a copy of the trigger identifier, a snapshot of the data to be verified, the current process stage code, a data validity timestamp, and a request priority flag.

[0037] Based on the node relationships implicit in the trigger identifier, the routing strategy engine matches the types of protection nodes that need to be coordinated (such as data consistency verification nodes and resource availability monitoring nodes). The system dynamically selects the best routing path based on the service load and response latency of the protection nodes and attaches the corresponding protection service tag to the request packet.

[0038] The packaged request is submitted to the admission interface of the task queue. The queue manager performs flow control based on the request priority flag and the current queue depth: high-priority requests can be inserted at the head of the queue; if the queue is full, a backpressure mechanism is triggered, notifying the integration node to postpone the push. At the same time, the system records the enqueue timestamp and sequence number, and generates a queue credential.

[0039] Preferably, refer to Figure 4 Step S3 includes: Based on the queue credentials, the node ensures that the enqueue timestamp and enqueue sequence number are verified. If all verifications pass, load and parse the data verification request bound to the queue credentials to obtain the process stage code and the data to be verified. The data verification rules are matched against the pre-defined data verification rules based on the process stage codes to obtain the stage verification rules; The data to be verified is parsed in a specific format and then structured extracted using a preset data pattern to obtain the data fields to be verified. Static validation is performed on the data fields to be validated according to the phase validation rules and the rule priority. If the static validation of the data field to be verified fails, an invalid data conclusion code will be generated. If not, perform business logic validation on the current data field to be validated; if the business logic validation fails, correct the dependencies and state transitions between the metadata of the current data field to be validated, extract the logical skeleton, record the rule score matrix, and generate a data part warning conclusion code or data invalid conclusion code. If not, perform data source validation on the current data field to be validated; if the data source validation fails, record the unmatched data field, generate a data validation exception conclusion code, and trigger exception handling decision; otherwise, generate a data validation complete success conclusion code. Based on the data verification conclusion code, rule score matrix, and evidence chain summary, and combined with the current data fields to be verified, a quality verification label is generated. Exception handling decisions include: Analyze the data verification anomaly conclusion code based on the interaction status between nodes to determine the cause of the anomaly; If the cause of the anomaly is an external dependency error, the logical skeleton will be re-matched a limited number of times according to the phase verification rules to determine the final data verification conclusion code. If the cause of the anomaly is a certain error, replace the data verification anomaly conclusion code with a data invalid conclusion code or a data partial warning conclusion code, and record the violation data segment; If the cause of the anomaly is a rule conflict, the stage verification rules are sorted according to the rule priority to determine the final data verification conclusion code.

[0040] In this embodiment, when any verification rule fails or returns an uncertain result, an exception handling sub-process is triggered: If the error is retryable (such as network jitter), a limited number of retries will be performed according to the retry policy configured in the rule. If it is a deterministic failure, record the detailed error code, the violation data fragment, and the triggered rule ID; If multiple rule results conflict, the rule arbitrator (based on a preset priority or voting mechanism) is invoked to determine the final validity conclusion.

[0041] In this embodiment, the safeguard node verifies the validity of the sequence number, timestamp, and digital signature of the received queue credential. After confirming it as a legitimate call, it loads the complete data verification request package bound to the credential from the distributed cache. Simultaneously, based on the process stage code and trigger identifier within the request package, it loads the corresponding data verification rule set (including syntax rules, business logic rules, and consistency rules) and the required external parameter context from the rule base.

[0042] The system parses the format of the snapshot data to be verified in the request packet (e.g., JSON / XML / binary deserialization) and extracts it in a structured manner according to the data pattern defined in the rule set. The system executes the rule matching engine to dynamically associate the extracted data fields with predefined verification rules, establishing a mapping matrix of "data fields - verification rules". If the data format is abnormal or key fields are missing, an "invalid format" label is generated directly and the process is terminated early.

[0043] Perform multi-level validation according to rule priority: Basic validation layer: Performs static rule validations such as data type, range, and enumeration value; Logic validation layer: Performs cross-field business logic consistency verification (such as dependency relationship verification and state transition validity); External consistency verification layer: Calls associated external system interfaces (such as databases and directory services) to compare the authenticity of data; Each verification unit is an atomic operation. During execution, resource usage (CPU / memory) and timeout are monitored in real time. Once a timeout or an exception occurs, the rule verification is marked as "unknown" and the exception context is recorded.

[0044] By combining the validation results of all rules, a structured quality validation label is generated. This label includes: core conclusion codes (such as "VALID" (all passed), "INVALID" (failed), "WARNING" (partial warning), "UNCERTAIN" (validation anomaly)), a detailed score matrix (recording the pass status, time taken, and confidence level of each rule), an evidence chain summary (linking the key data fingerprints of success or failure to the rule ID), and remediation suggestions (for failed items, attaching links to predefined remediation guidelines or data templates in the rule base).

[0045] At the same time, additional context for this verification is constructed, including the verification time window, the rule version number used, and the identifier of the protection node host.

[0046] Preferably, step S4 includes: The quality verification labels are parsed and mapped according to the preset state mapping rule table to determine the target node of the business execution process and generate a state update instruction. The state variables of the target node are locked, and a new state code is registered in conjunction with the state update instruction to obtain the node to be executed; Based on the business process diagram, perform topology analysis on the nodes to be executed, generate state change events, and update the node state view; The decision rule set is matched with the state change event to obtain the phase operation decision; Based on the phased operational decisions, rule reasoning is performed on the nodes to be executed to determine the node execution strategy and allocate the corresponding operational resources. If the node execution strategy is to continue execution, then release the node lock of the currently pending node, send a ready signal to the subsequent process nodes, and allocate computing resources. If the node execution strategy is to retry, then the amount of data on the node to be executed is adjusted, and a limited number of retries are initiated in the resource isolation environment; If the node execution strategy is process rerouting, then the execution path of the process instance is dynamically reconstructed, a backup backup node is loaded, and the data fields to be verified are injected. If the node execution strategy is external intervention, then create and push the node execution work order, and freeze the process clock.

[0047] In this embodiment, after receiving the quality verification tag, its core conclusion code and detailed score matrix are parsed, and the verification conclusion is mapped to a specific node status update instruction according to a preset state mapping rule table. For example: VALID → "Ready", INVALID → "Blocked - Data Abnormality", WARNING → "Ready - Requires Manual Review", UNCERTAIN → "Retry Pending". At the same time, the evidence chain summary and repair suggestions in the tag are extracted as additional status information.

[0048] The node state update is performed using a distributed transaction mechanism, including: locking the target node state variable to prevent concurrent conflicts; writing the new state (including sub-state codes), timestamp, and verification tag digest to the node state register; updating the topological state of the node in the flowchart (e.g., propagating the "blocked" state back along the dependency chain to the preceding node); generating a state change event and publishing it to the process event bus to ensure that the state views of all listening components are eventually consistent.

[0049] Based on the updated node status combination and the rule ID in the quality verification label, the process execution decision engine is triggered: load the decision rule set (such as retry rules, branch jump rules, and compensation transaction rules) that matches the current process stage, node type, and exception mode; input the status data, verification evidence, and historical process performance indicators as facts into the rule engine (such as Drools); execute rule reasoning and generate decision output: continue execution, retry the current node, jump to the backup node, start manual intervention in the process, or trigger process rollback.

[0050] Based on the decision output type, perform the following operations: Continue execution: Release the node lock, send a ready signal to subsequent nodes, and allocate computing resources; Retry execution: Adjust the input data or parameters according to the repair suggestions in the tag, and initiate a limited number of retries in a resource isolation environment (following the exponential backoff strategy); Process rerouting: When the decision is a branch jump, dynamically reconstruct the execution path of the process instance, load the backup node component, and inject context data; Manual intervention: When the decision is manual intervention, automatically create a work order and push it to the designated personnel's terminal, while freezing the process clock; and all execution operations are associated with the original verification tag ID to ensure full-link traceability.

[0051] Establish a real-time monitoring loop for decision execution effectiveness: collect key process indicators after decision execution (such as node recovery time, resource utilization, and anomaly resolution rate); evaluate the long-term utility of different decision rules in similar scenarios through online learning algorithms (such as multi-armed slot machine models); when the decision effect continues to deviate from expectations, automatically generate rule optimization suggestions (such as adjusting thresholds, adding or deleting rule conditions), and hot update them to the rule base after security approval.

[0052] The implementation principle of this embodiment is as follows: First, based on directed graph modeling and rule engine monitoring, the business process is deconstructed into nodes and the dependency relationship is mapped. The node status is tracked in real time and the integrated task triggering identifier is automatically generated according to the pre-associated rules. Then, the data model is invoked and the data is cleaned (duplicates, invalid values ​​are handled, missing values ​​are filled) and fusion algorithms are used to obtain standard data. After being encapsulated into priority verification requests, the requests are pushed to the guarantee task queue based on load balancing and routing algorithms. Priority queue management and back pressure mechanisms are used for traffic control and queue credentials are generated. Subsequently, the rules are verified through the loading stage of the rule engine. The multi-layer verification algorithm, from static verification and business logic verification to data source comparison, is executed on the data to be verified. Combined with the exception handling decision mechanism (such as rule conflict arbitration and logical skeleton re-matching), a quality verification label containing a conclusion code and evidence chain is generated. Finally, the node state is updated through state machine mapping based on the label, and differentiated strategies such as "continue execution", "retry execution", "process change" or "external intervention" are dynamically generated and executed based on decision rule set matching and reasoning (such as decision tree). This is linked with resource scheduling algorithms to realize intelligent driving and elastic management of the process, thereby building a full-link automated closed loop from process monitoring, data extraction, intelligent verification to adaptive decision-making, which significantly improves the reliability, resilience, data quality and overall efficiency of business execution.

[0053] Preferably, the data validation rule invocation algorithm includes: Data quality verification rules are sharded and stored according to the rule table name, and a phase rule index is constructed. ; Among them, Index table For table name indexes, ⊕ represents distributed aggregation operations, and N... node Let k be the total number of storage nodes, k be the node index, and r be the number of nodes. i Let be the i-th data quality validation rule, R be the rule base, and table be... i Let be the i-th rule table, ↔ be the rule mapping, and Index be... field For field name indexing, DistHash is the hash function, and field... i This refers to the name of the i-th field. Based on the rule complexity and memory usage, the execution schedule of the data quality verification rules is determined to establish the rule scheduling priority. ; Where PQ() is the rule-based scheduling priority, Sort() is the elastic scheduling function, and s i For the question level, w i Here, PCost() represents the rule weight, PCost() represents the rule complexity, and MCost() represents the memory usage. Based on the rule-based heat index function, streaming rule matching is performed on the business standard data within a preset time window, and abnormal matching data is aggregated. Where d represents standard business data, Stream(T) represents the data stream within the time window T, Match() represents batch streaming matching, and Result... ri For abnormal matching data, f k For the k-th field, Ψ typei For typed validation functions, params i For business data, l hot () represents the regularity heat characteristic function, error i Warning level; The abnormal matching data is incrementally updated according to the data processing batch t to obtain the increment.

[0054] In this embodiment, during the birth registration verification of the provincial government data quality platform, the "Birth Medical Certificate" table and the "Permanent Resident Population Information" table are first matched using a distributed rule index. The "Date of Birth" field triggers a standardization check (Ψregex) and a value range check (Ψrange), while the "Mother's ID Number" field triggers a cross-table association check (Ψforeignasync) and a logical consistency check (Ψlogic).

[0055] During the rule execution phase, the "Date of Birth" field is first quickly checked for duplicate values ​​using a Bloom filter (e.g., the ID number already exists in the permanent resident table). If the field cardinality exceeds one million (|Hf|>10^6), an approximate deduplication strategy is adopted. For the "Mother's ID Number" field, the relationship is quickly verified through a local cache (Cache(Tref)). If the cache misses, it is submitted to Spark SQL for asynchronous batch query (AsyncQuery) to ensure the real-time performance and accuracy of cross-table verification. Meanwhile, the rule execution priority is calculated based on problem level (si), weight (wi), and resource consumption (PCost, MCost) (PQ). For example, ID card format verification is prioritized due to its high frequency and low resource consumption. Anomalies are aggregated through micro-batch processing (Batch(tw)). If an invalid "date of birth" format is found (e.g., "20251332" does not conform to the YYYYMMDD specification) or a "mother's ID number" is missing from the permanent resident population table (Ψforeignasync), the anomaly type (e.g., "value range out of limit" or "foreign key missing") is output, and the anomaly report (FinalReport) is updated using incremental calculation (ΔResultri). Ultimately, through versioned hot updates and dynamic adjustments using Bloom filters, efficient verification of 1.5 billion pieces of government data per day is achieved. Example

[0056] Reference Figure 5 Within a graphical interface, users can drag and drop to build an end-to-end workflow for a specific business theme (such as "customer information governance"). Logical dependencies between nodes are defined: it is determined and confirmed that the data integration node is a prerequisite for the data assurance node; that is, subsequent quality verification nodes are only meaningful and activated after the integration task is successfully executed. Furthermore, the required quality verification rule sets are pre-bound or configured for the data assurance node, embedding two previously isolated nodes into a unified, visual collaboration framework.

[0057] Continuously monitor the running status of all data integration nodes, and trigger subsequent actions only when the status of an integration node changes from "Executing" to "Successful". Once this condition is met, automatically mark the status of the current business topic process instance as "Pending Verification" and generate a structured verification request; The logic for generating verification requests includes: automatically capturing and encapsulating the key context of this integration task, such as data samples (for actual value checks), metadata (such as table structure and field types), and integration logs (for analyzing incremental or error situations), and pushing them to the corresponding task queue of the data assurance group.

[0058] Once the data assurance node is triggered, it first matches the metadata information (such as data source type, table name, and business theme) carried in the request with the pre-set rule base, automatically filters and loads all applicable quality verification rules (for example, performing format compliance verification on the "ID number" field and numerical range and non-empty verification on the "transaction amount" field).

[0059] Next, the verification scripts or tools corresponding to these rules are automatically called or executed, and the integrated data samples are used as input for batch evaluation (the whole process does not require the assurance team to manually initiate the inspection, but automatically and accurately completes the first round of quality scanning and generates a verification report containing detailed pass, alarm or failure entries).

[0060] After the verification component completes its execution, the results of the potentially scattered verification rules need to be aggregated into a unified, machine-readable conclusion. Based on comprehensive judgment logic, for example: if any "failure" item exists, the overall result is "failure"; if there are no "failures" but there is an "alarm," it is "alarm"; only if all items "pass" is it considered "pass." The integrated results are then sent back to the process engine in real time. Upon receiving the results, the engine immediately updates the status of the data assurance node in the process instance (e.g., from "Verification in progress" to "Verification failed") and synchronously updates the status flags in the global process view.

[0061] After receiving the verification result, the process engine will immediately match it with the "strategy set" pre-configured for the process and execute the corresponding automated decision.

[0062] If the result is "passed", the engine will trigger subsequent nodes (such as data development nodes and data application nodes); if the result is "alarm", the engine may execute the strategy of "sending a notification to the person in charge but continuing the process"; if the result is "failed", the engine will strictly execute the strategy of "blocking the process" - that is, suspending the progress of the entire workflow and automatically triggering alarm actions (such as directly connecting with the relevant team in the collaboration platform or sending an email).

[0063] It collects and aggregates status data (such as "integrating", "pending verification", "verification failed"), detailed logs, and quality verification reports from different process instances and nodes in real time from data sources (such as process engine database and log system).

[0064] Then, through a unified monitoring dashboard or view, using intuitive formats such as colors (red for failure, green for success), progress bars, and topology diagrams, managers are provided with a panoramic view of data governance work.

[0065] Reference Figure 6A collaborative verification system for data integration and quality assurance oriented towards data governance, applied to collaborative verification methods for data integration and quality assurance, includes: The conversion trigger module is used to generate integration task trigger identifiers based on the association and monitoring of integration nodes and assurance nodes in the business execution process; The integration trigger module is used to push data verification requests from the integration node to the security task queue based on the integration task trigger identifier, and generate queue credentials. The quality verification module is used to ensure that nodes call data verification rules to verify data verification requests based on queue credentials and generate quality verification tags. The decision matching module is used to update the node status of the business execution process based on the quality verification label, and to match and execute the corresponding process operation decisions.

[0066] The implementation principle of this embodiment is as follows: the transformation trigger module monitors the node status based on the directed graph and rule engine and generates trigger identifiers, drives the integrated trigger module to push verification requests and generate queue credentials using message queues and load balancing algorithms, the quality verification module outputs quality labels with the help of rule engines and multi-layer verification algorithms, and the decision matching module dynamically updates the status and executes process decisions based on the labels through state machine mapping and decision tree reasoning, thereby realizing full-link automated and intelligent data quality control and flexible process scheduling.

[0067] An electronic device, comprising: One or more processors; Memory, used to store one or more programs; When one or more programs are executed by one or more processors, the one or more processors implement any of the methods in the above scheme.

[0068] A storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the data integration and quality assurance collaborative verification method as described above.

[0069] The embodiments described in this specific implementation are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A collaborative verification method for data integration and quality assurance oriented towards data governance, characterized in that, include: Based on the business execution process, the integration nodes and assurance nodes are associated and monitored to generate integration task trigger identifiers; Based on the integrated task trigger identifier, the integrated node pushes a data verification request to the security task queue and generates a queue credential; Based on the queue credentials and data verification rules, the algorithm is invoked, the guarantee node parses and verifies the data verification request, triggers anomaly handling decisions, and generates a quality verification label. The node status of the business execution process is updated based on the quality verification label, and the corresponding process operation decision is matched and executed.

2. The data integration and quality assurance collaborative verification method for data governance according to claim 1, characterized in that, The step of generating an integration task trigger identifier based on the association and monitoring of integration nodes and assurance nodes according to the business execution process includes: Based on the data processing flow, the business execution flow is deconstructed, and the data integration group and data assurance group are extracted; Based on the data flow direction, the data integration group and the data protection group are node-based to generate integration nodes and protection nodes; Based on the data operation dependencies, the integration node and the protection node are associated and mapped to determine the interaction chain between the nodes; The integrated node is monitored based on the inter-node interaction chain to determine the real-time status of the node; Based on the real-time status of the node and the preset pre-association rules, an integrated task trigger identifier is generated.

3. The collaborative verification method for data integration and quality assurance oriented towards data governance according to claim 1, characterized in that, The step of the integration node pushing a data verification request to the guarantee task queue and generating a queue credential based on the integration task trigger identifier includes: The integrated task trigger identifier is parsed to obtain the interface dependency hash, trigger timestamp, and node checksum. The integrated nodes are verified and sorted according to the node checksums to generate an integrated processing sequence; The data model in the integration node is called according to the integration processing sequence, and the data is verified by combining the interface dependency hash to extract the data source and load the model. Based on the data source loading model, the business data source is transformed and deconstructed to extract the business data stream; The business data stream is sorted based on the metadata in the integration node to obtain physical data entities; The physical data entities are grouped according to their data types to obtain single-type data entities; Based on the data collection time, duplicate values ​​are removed, invalid values ​​are removed, and missing values ​​are filled for the single-type data entities to obtain single-type standard data. Based on the business scope, all the single-category standard data are fused to obtain business standard data.

4. The data integration and quality assurance collaborative verification method for data governance according to claim 3, characterized in that, The step of pushing a data verification request to the security task queue and generating a queue credential based on the integration task trigger identifier also includes: The trigger timestamp, process stage code, and business standard data or single data entity or single standard data are structured and encapsulated to generate a data verification request and assign a priority mark. Based on the service load and response latency of the protection nodes, the protection routing path is selected, and a protection task queue is generated; Based on the priority flag and the guaranteed queue status, traffic control is performed on the data verification request, and the queuing timestamp and queuing sequence number are recorded to generate a queue credential. If the priority is marked as high, the current data verification request is inserted at the head of the guarantee task queue. If the status of the guarantee queue is that the queue capacity has reached the upper limit, the back pressure mechanism is triggered, and the integration node is notified to delay the push.

5. The collaborative verification method for data integration and quality assurance oriented towards data governance according to claim 1, characterized in that, The process of invoking the algorithm based on the queue credentials and data verification rules, the safeguard node parsing and verifying the data verification request, triggering anomaly handling decisions, and generating quality verification tags includes: Based on the queue credentials, the node ensures that the enqueue timestamp and enqueue sequence number are verified. If all verifications pass, the data verification request bound to the queue credential is loaded and parsed to obtain the process stage code and the data to be verified. The algorithm is invoked according to the data verification rules to verify and match the process stage codes, thereby obtaining the stage verification rules; The data to be verified is parsed in a specific format and then extracted in a structured manner using a preset data pattern to obtain the data fields to be verified. The data fields to be verified are statically validated according to the rule priority based on the stage verification rules. If the static validation of the data field to be verified fails, an invalid data conclusion code is generated. If not, perform business logic validation on the current data field to be validated; if the business logic validation fails, correct the dependencies and state transitions between the metadata of the current data field to be validated, extract the logical skeleton, record the rule score matrix, and generate a data part warning conclusion code or data invalid conclusion code. If not, perform data source validation on the current data field to be validated; if the data source validation fails, record the unmatched data field, generate a data validation exception conclusion code, and trigger exception handling decision; otherwise, generate a data validation complete success conclusion code. Based on the data verification conclusion code, the rule score matrix, and the evidence chain summary, and combined with the current data field to be verified, a quality verification label is generated.

6. The collaborative verification method for data integration and quality assurance oriented towards data governance according to claim 1 or 5, characterized in that, The anomaly handling decision includes: Analyze the data verification anomaly conclusion code based on the interaction status between nodes to determine the cause of the anomaly; If the cause of the anomaly is an external dependency error, the logical skeleton will be re-matched a limited number of times according to the phase verification rules to determine the final data verification conclusion code. If the cause of the anomaly is a certain error, then replace the data verification anomaly conclusion code with a data invalid conclusion code or a data partial warning conclusion code, and record the violation data fragment; If the cause of the anomaly is a rule conflict, the stage verification rules are sorted according to the rule priority to determine the final data verification conclusion code.

7. The data integration and quality assurance collaborative verification method for data governance according to claim 1, characterized in that, The data verification rule invocation algorithm includes: Data quality verification rules are sharded and stored according to the rule table name, and a phase rule index is constructed. ; Among them, Index table For table name indexes, ⊕ represents distributed aggregation operations, and N... node Let k be the total number of storage nodes, k be the node index, and r be the number of nodes. i Let be the i-th data quality validation rule, R be the rule base, and table be... i Let be the i-th rule table, ↔ be the rule mapping, and Index be... field For field name indexing, DistHash is the hash function, and field... i This refers to the name of the i-th field. Based on the rule complexity and memory usage, the execution schedule of the data quality verification rules is determined to establish the rule scheduling priority. ; Where PQ() is the rule-based scheduling priority, Sort() is the elastic scheduling function, and s i For the question level, w i Here, PCost() represents the rule weight, PCost() represents the rule complexity, and MCost() represents the memory usage. Based on the rule-based heat index function, streaming rule matching is performed on the business standard data within a preset time window, and abnormal matching data is aggregated. Where d represents standard business data, Stream(T) represents the data stream within the time window T, Match() represents batch streaming matching, and Result... ri For abnormal matching data, f k For the k-th field, Ψ typei For typed validation functions, params i For business data, l hot () represents the regularity heat characteristic function, error i Warning level; The abnormal matching data is incrementally updated according to the data processing batch t to obtain the increment.

8. The data integration and quality assurance collaborative verification method for data governance according to claim 1, characterized in that, The step of updating the node status of the business execution process based on the quality verification label, and matching and executing the corresponding process operation decision, includes: The quality verification labels are parsed and mapped according to the preset state mapping rule table to determine the target node of the business execution process and generate a state update instruction. The state variables of the target node are locked, and a new state code is registered in conjunction with the state update instruction to obtain the node to be executed; Based on the business process diagram, perform topology analysis on the nodes to be executed, generate state change events, and update the node state view; The decision rule set is matched based on the state change event to obtain the stage operation decision; Based on the stage operation decision, rule reasoning is performed on the node to be executed to determine the node execution strategy and allocate the corresponding operation resources. If the node execution strategy is to continue execution, then release the node lock of the currently pending node, send a ready signal to the subsequent process nodes, and allocate computing resources. If the node execution strategy is to retry execution, then the amount of data of the current node to be executed is adjusted, and a limited number of retries are initiated in the resource isolation environment; If the node execution strategy is process track change, then the execution path of the process instance is dynamically reconstructed, a backup support node is loaded, and the data field to be verified is injected. If the node execution strategy is external intervention, then create and push the node execution work order, and freeze the process clock.

9. A collaborative verification system for data integration and quality assurance oriented towards data governance, used to implement the collaborative verification method for data integration and quality assurance as described in any one of claims 1-8, characterized in that, include: The conversion trigger module is used to generate integration task trigger identifiers based on the association and monitoring of integration nodes and assurance nodes in the business execution process; An integrated triggering module is used to push a data verification request to the guarantee task queue based on the integrated task triggering identifier, and generate a queue credential. The quality verification module is used to call the algorithm based on the queue credentials and data verification rules, and the guarantee node verifies the data verification request and generates a quality verification tag. The decision matching module is used to update the node status of the business execution process based on the quality verification label, and to match and execute the corresponding process operation decision.

10. A storage medium storing at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the data integration and quality assurance collaborative verification method as described in any one of claims 1 to 8.