Data auditing method based on big data

By building a cross-form dependency graph and quantized state operations, combined with a self-calibration loop mechanism and distributed fault-tolerant control, the problems of cross-form data dependency breaks and cascading rollback storms in big data platforms are solved, achieving high-stability and high-precision data auditing.

CN120803812AInactive Publication Date: 2025-10-17SHENZHEN JINMAILI MEDIA TECH CO LTD

Patent Information

Application Number
CN202510852994.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies in big data platforms suffer from cross-form data dependency fractures, cascading rollback storms, lack of fault tolerance in distributed environments, and long-term state drift, resulting in reduced system stability and audit accuracy.

Method used

By building a cross-form dependency graph, adopting incremental state rollback, anti-rollback storm mechanism and distributed fault-tolerant control, combined with quantized state operation and self-calibration loop mechanism, accurate cascade review and high-stability operation across business forms can be achieved.

Benefits of technology

It achieves precise mapping and cascading audit of cross-form data dependencies, suppresses cascading rollback storms, ensures high availability and long-term audit accuracy in distributed environments, and ensures the stability and accuracy of the system in high-frequency data change scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803812A_ABST
    Figure CN120803812A_ABST
Patent Text Reader

Abstract

The invention discloses a data auditing method based on big data, and relates to the technical field of big data, and the method comprises the steps: obtaining multi-modal data in a big data platform form, extracting features of the multi-modal data to generate a feature set, creating a rule base based on the feature set, and constructing a cross-form dependency graph according to the feature set and the rule base; cascade auditing is executed through a cross-form dependency map, when auditing fails, an affected field is reversely positioned according to the connection direction of the map, and incremental state rollback and re-auditing are executed on the affected field; when it is detected that the high-frequency data is changed, an anti-rollback storm mechanism is started, a cascade load index is obtained, and a rollback instruction set is generated through cascade blocking based on the index; and executing distributed fault-tolerant control, constructing a node collaborative protection system through cascade loads, and outputting a fault-tolerant operation instruction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data, and particularly relates to a data auditing method based on big data. BACKGROUND

[0002] In the current data auditing practice of the big data platform, the business scenarios of complex form association are facing persistent technical challenges. The traditional auditing scheme has significant limitations in processing cross-form data dependencies. The logical association between different business forms is difficult to establish effective mapping. The existing auditing failure processing mechanism generally adopts a global rollback strategy. This approach is prone to cause cascading operation diffusion in a high-frequency data change environment, leading to increased system stability risks. In a distributed deployment architecture, there is inherent delay in the state synchronization process between nodes. Data version divergence is likely to occur when there is network partitioning or node failure, and the system recovery process is uncertain. The long-term running data state maintenance mechanism lacks continuous calibration capability, and the auditing accuracy shows a gradual deviation trend over time, directly affecting the accuracy of key business decisions.

[0003] The Chinese invention patent with the publication number CN111966673B provides a data auditing method, device and storage medium based on big data. The patent improves the efficiency of single-form batch auditing, but only optimizes the auditing efficiency of isomorphic forms, and cannot solve the problem of cross-form business links. SUMMARY

[0004] The present application provides a data auditing method based on big data, which solves the problems of cross-form data dependency breakage, cascade rollback storm, distributed environment fault tolerance deficiency, and long-term state drift in the prior art, and achieves the technical effects of precise cascade auditing of cross-business forms, high stability operation against storm, distributed fault self-healing, and long-term precision maintenance.

[0005] The present application provides a data auditing method based on big data, which includes:

[0006] S1: Obtain multi-modal data in the big data platform form, extract features of the multi-modal data to generate a feature set, create a rule base based on the feature set, and construct a cross-form dependency graph based on the feature set and the rule base;

[0007] S2: Perform cascade auditing through the cross-form dependency graph. When the auditing fails, locate the affected fields in the reverse direction according to the connection direction of the graph, and perform incremental state rollback and re-auditing on the affected fields;

[0008] S3: When high-frequency data changes are detected, start the anti-rollback storm mechanism, obtain cascade load indicators, and generate a rollback instruction set based on the indicators through cascade blocking;

[0009] S4: Perform distributed fault-tolerant control, build node cooperative protection system through cascading load, and output fault-tolerant operation instructions.

[0010] Further, the multi-modal data includes field metadata, historical audit logs, and business rule documents.

[0011] The feature set includes: performing structured parsing on field metadata to generate structural features, performing statistical analysis on historical audit logs to generate behavior features, and performing semantic segmentation on business rule documents to generate logical semantic features.

[0012] The rule base includes: strong dependency rules, weak association rules, dynamic evolution rules, and conflict resolution rules.

[0013] Further, the cross-form dependency graph includes: generating a node set based on the feature set, each node associated with a form field; cross-form node connection is realized through a double-linked list, and the connection direction represents the audit influence relationship between fields; the connection corresponding to the strong dependency rule is marked as a strong dependency type, and the connection corresponding to the weak association rule is marked as a weak association type.

[0014] Further, the cascading audit further includes: performing complexity sorting audit on the current form field; activating cross-form dependency graph scanning to identify associated form field combinations; batch loading associated field data and performing audit according to the rule types in the rule base.

[0015] Further, the incremental state rollback includes: first, dynamic radius calculation is performed to determine the accurate rollback range through graph association distance; then, quantum state operation is implemented, and a quantum state bitmap is created for strong dependency fields; transaction version rollback is performed on strong dependency fields within the radius, and weak association fields are only reset for verification marks, and fields outside the radius remain unchanged; finally, a hierarchical review mechanism is used to synchronize real-time review for strong dependency fields, asynchronous batch processing for weak association fields, and automatic skipping for fields outside the radius, to cooperatively eliminate redundant reviews.

[0016] Further, the quantum state operation further includes deploying a state health monitoring module: periodically generating a hash check code for the quantum state bitmap of the strong dependency field; locating the verification range of the associated field based on the cross-form dependency graph; when the distortion rate of the bitmap is greater than 5%, activating the hierarchical review mechanism to perform full review.

[0017] Further, the health monitoring module further includes a self-calibration cycle mechanism, which includes: triggering calibration through timing, real-time filtering of key field sets, correcting data distortion through hierarchical repair strategies, and finally updating system quantum state parameters to achieve cross-period stability control.

[0018] Further, the self-calibration cycle mechanism further comprises: according to the periodically generated hash check code, constructing an error model through a drift compensation algorithm, and pre-compensating offset correction before the rollback operation is executed; using the error model to iteratively optimize the parameters, and realizing prediction and correction of cumulative errors.

[0019] Further, the anti-rollback storm mechanism comprises: monitoring the rollback trigger frequency and the cascading load index in real time; dynamically executing a hierarchical fuse strategy based on the index; optimizing the storage structure of the change record while executing the fuse strategy; filtering key rollback paths based on a cross-form dependency graph to generate a lightweight rollback instruction set; the cascading load index comprises a cascading depth and a resource load index.

[0020] Further, the distributed fault-tolerant control comprises: eliminating quantum state delay through versioned state synchronization, avoiding network splitting risk by using a partition-aware arbitration mechanism, preventing historical version error recovery based on space-time consistency verification, and deploying an error propagation blocker to suppress chain diffusion.

[0021] One or more technical solutions provided in the present application have at least the following technical effects or advantages:

[0022] By constructing a cross-form dependency graph to dynamically map the logical relationship between heterogeneous business fields, combining incremental rollback and anti-storm mechanism to realize precise rollback range control and system chain collapse suppression, the stable operation of high-frequency scenarios is ensured; according to the distributed fault-tolerant architecture, node fault self-healing and error propagation blocking are completed to ensure high availability in a distributed environment; with the help of quantum state self-calibration technology, long-term running cumulative errors are continuously corrected to maintain the consistency of audit accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 A data audit method flowchart based on big data in the embodiments of the present application. DETAILED DESCRIPTION

[0024] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the related drawings; the preferred embodiments of the present application are shown in the drawings, but the present application can be realized in many different forms and is not limited to the embodiments described herein; on the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs; the terms used herein in the specification of the present application are only for the purpose of describing the specific embodiments and are not intended to limit the present application; the term "and / or" used herein includes any and all combinations of one or more related listed items.

[0026] Embodiment one: as shown, a big data-based data auditing method. Figure 1

[0027] S1: Obtain multi-modal data in a big data platform form, extract features of the multi-modal data to generate a feature set, create a rule library based on the feature set, and construct a cross-form dependency graph based on the feature set and the rule library;

[0028] The multi-modal data includes field metadata, historical audit logs, and business rule documents.

[0029] Specifically, the field identifier, data constraints, and cross-form foreign keys are obtained from the metadata storage layer of the big data platform form; the historical audit logs are obtained from the distributed log system; and the business rule documents are obtained from the structured document repository and the contract file.

[0030] The feature set includes: performing structured analysis on field metadata to generate structural features, statistically analyzing historical audit logs to generate behavior features, and performing semantic segmentation on business rule documents to generate logical semantic features.

[0031] Specifically, the input field metadata is subjected to structured analysis: identify the field type and bind the basic verification rules, bind the range verification and format verification for numerical fields, bind the effective period verification and time sequence verification for date type fields, and bind the regular expression for string fields; establish cross-field verification relationships based on business logic, automatically associate upstream and downstream time effectiveness verification for date types, automatically bind the primary key verification of the associated table for foreign key fields, and finally generate structural features; input historical audit logs for statistical analysis: identify high-frequency error patterns and their potential associations, parse the field co-occurrence error patterns in the historical audit logs based on the initial topology structure of the cross-form dependency graph, construct a Bayesian network to calculate the conditional probability, and simultaneously weaken the weight of obsolete data through an exponential decay mechanism and realize real-time updating, finally generating behavior features; generate logical semantic features by analyzing business rule documents: perform semantic segmentation on the text to extract key elements, construct a rule relationship graph with quantified intensity, and finally obtain an interpretable rule library with traceable sources.

[0032] The rule library includes: strong dependency rules, weak association rules, dynamic evolution rules, and conflict resolution rules.

[0033] ​Specifically, input the structural features and logical semantic features, extract the foreign key primary key association in the structural features and fuse the rigid constraints in the semantic features, output the strong dependency rules that violate and block the process; input the behavior features and logical semantic features, filter the behavior association with a confidence greater than 75% and bind the advisory constraints in the semantic features, generate the weak association rules that are not mandatory; input the behavior features, load the initial impact probability and start the exponential decay engine, output the self-adaptive dynamic evolution rules; input the structural features, behavior features and logical semantic features, establish a priority decision tree to obtain the conflict resolution rules; and create the rule library of cross-form dependencies according to the strong dependency rules, weak association rules, dynamic evolution rules and conflict resolution rules.

[0034] The cross-form dependency graph includes: generating a node set based on the feature set, each node being associated with a form field; connecting the cross-form nodes through a double-linked list, the connection direction representing the review influence relationship between the fields; and the connection marked as a strong dependency type corresponding to the strong dependency rules, and the connection marked as a weak association type corresponding to the weak association rules.

[0035] Specifically, the cross-form dependency graph is constructed according to the feature set and the rule library; the nodes of each form field are generated according to the field metadata in the structural features; the cross-form nodes are connected through a double-linked list, wherein the forward link represents the review influence of the source field change on the target field, and the reverse link supports the abnormal source tracing of the target field; and the connection type is marked according to the rule library, the strong dependency rules corresponding to the red solid line connection, and the weak association rules corresponding to the blue dashed line connection. The graph inherits the dynamic evolution capability of the behavior features, and each connection is real-time marked with the current connection probability, and the calculation formula is:

[0036] P c =α·S r +β·P b +γ·F f ,

[0037] Wherein, P c is the connection probability, S r is the rule coefficient, the strong dependency rules are 1.0, the weak association rules are 0.6, and the dynamic evolution rules are 0.8; P b is the behavior probability, which is obtained by calculating the Bayesian network, and the value range is [0, 1]; F f is the data freshness, which is obtained by calculating the exponential decay function, and the value range is [0, 1]; α, β, γ are corresponding weight coefficients, and the sum is 1, and here α=0.3, β=0.6, and γ=0.1;

[0038] The updating is continuous through an exponential decay mechanism, and the strong dependence red solid line is automatically switched to the weak association blue dashed line when the connection probability falls below 0.5, realizing real-time visual mapping and risk transmission tracking of cross-form review influence.

[0039] S2: performing cascade review through the cross-form dependency graph;

[0040] The cascade review further includes: performing complexity sorting review on the current form field; activating cross-form dependency graph scanning to identify associated form field combinations; batch loading associated field data and performing review according to the rule types in the rule library.

[0041] Specifically, all field metadata of the current form are extracted, and the weight value of each field is calculated, with the formula being:

[0042] W=S r ×EF×log2(DD+1),

[0043] wherein W is the field weight value, S r is a rule coefficient, the strong dependence rule is 1.0, the weak association rule is 0.6, and the dynamic evolution rule is 0.8; EF is an error rate factor, with the formula being: EF=(0.7×ER)+(0.3×EH), ER is the review failure rate of the field in the last 7 days, EH is the cumulative historical review failure rate of the field, 0.7 and 0.3 are corresponding weights, and the recent 7 days data is focused on; log2 is the logarithm with base 2; DD is the dependency depth, which is the total number of associated fields in the cross-form dependency graph; and 1 is a depth offset to avoid invalid calculation of the logarithm when DD=0;

[0044] A priority queue is generated in descending order of weight, and the top 50 key fields are reviewed preferentially;

[0045] When the form field is changed or the review fails, the cross-form dependency graph scanning is automatically activated, the strong dependence connection is forced to scan, the weak association connection is included when the connection probability is greater than or equal to 0.5, and the real-time connection probability is less than 0.3 to skip and cache; a bidirectional depth-first search algorithm is used, the core field is taken as the forward scanning along the review influence path, and the end field is taken as the backward scanning back to the core field; during the scanning process, the weak connection with a connection probability less than 0.3 is directly skipped and cached with a high frequency path; according to the scanning result, the strong dependence fields are generated into a strong dependence group, and the weak association fields are generated into a weak association group; all field data of the strong dependence group and the weak association group are batch extracted to avoid multiple I / O operations;

[0046] Based on different rule types in the rule library, a differentiated audit strategy is adopted: real-time synchronous audit is adopted for strong dependence rules, strict field-by-field verification is adopted, any failure will block the process and trigger incremental rollback; asynchronous batch audit is adopted for weakly associated rules, parallel verification is adopted, failure only records an alarm and does not block the process, and subsequent asynchronous repair; dynamic evolution rules are executed as an independent rule type, when the connection probability is less than 0.3, the verification is skipped, when the connection probability is greater than or equal to 0.7, the rule is upgraded to a strong rule audit, and when the connection probability is between 0.3 and 0.7, the original rule type audit is maintained; when there are multiple rule constraints in the same field, a conflict resolution rule is triggered, the final effective rule is output according to the decision tree, and the audit strategy is updated in real time.

[0047] The technical solutions in the embodiments of the application have at least the following technical effects or advantages:

[0048] The application adopts multi-modal feature fusion and dynamic graph modeling to achieve the technical effects of precise influence control and real-time adaptive optimization of cross-form cascade audit, constructs a feature set that can dynamically map cross-form field associations by fusing field metadata, historical audit logs and business rule documents, solves the problem of logical discontinuity across business forms in traditional solutions; a cross-form dependency graph is generated based on the feature set and the rule library, a bidirectional linked list is used to realize node connection and influence direction identification, when the connection probability is less than 0.5, the weak association type is automatically downgraded, realizing real-time visual management of business association relationships; a weight ordering mechanism is used to prioritize key fields, and a field group is dynamically identified by combining graph scanning, and differentiated synchronous strong rule audit and asynchronous weak rule verification is performed, avoiding invalid computational resource consumption; a priority decision tree is constructed by a conflict resolution rule, automatically resolving multiple rule constraint conflicts and ensuring continuous execution of the audit process.

[0049] Embodiment two: in embodiment one, by constructing a cross-form dependency graph that fuses multi-modal features, combining a dynamic connection probability model and a rule library grading mechanism, real-time visual mapping and cascade audit influence control of logical associations between heterogeneous business forms are realized, solving the core problem of cross-form data dependency discontinuity in traditional solutions, but the dependency graph only provides the influence direction when the audit fails, and does not define the precise rollback range, which can easily cause redundant operations, and lacks a state calibration mechanism, to solve these problems, embodiment two further improves embodiment one.

[0050] Step S2 further comprises: when the audit fails, the affected fields are located in reverse according to the connection direction of the graph, and incremental state rollback and re-audit are performed on the affected fields;

[0051] The incremental state rollback comprises: firstly, calculating a dynamic radius, determining an accurate rollback range by a graph correlation distance; then, implementing a quantumized state operation, and creating a quantum state bitmap for a strongly dependent field; performing a transaction version rollback on the strongly dependent field within the radius, resetting a check mark for a weakly correlated field, and keeping the state of a field outside the radius unchanged; finally, using a hierarchical review mechanism, synchronously reviewing the strongly dependent field in real time, asynchronously processing the weakly correlated field in batches, and automatically skipping the field outside the radius, and cooperatively eliminating redundant reviews.

[0052] Specifically, when the review fails, the affected field is accurately located by a cross-form dependency graph, the rollback range is minimized, and resource waste of traditional global rollback is avoided. The topological structure of the cross-form dependency graph and the coordinates of the failed review field are input, the rollback radius is determined based on the graph correlation distance, and the calculation formula is:

[0053] R = max(R d ) x p,

[0054] wherein R is the rollback radius, R d is the distance of the strongly dependent path, and p is a dynamic attenuation factor for inhibiting the excessive expansion of the rollback radius, and the value range is (0, 1], and the rollback radius centered on the failed field is obtained by the calculation formula;

[0055] The quantumized state operation is implemented, a binary state bitmap is created for the strongly dependent field, 1 bit represents 1 field, the value 1 needs to be rolled back to the historical version, 0 is normal state, each field is bound with a transaction version number, and the last effective version is extracted from the log to cover when rolling back; the weakly correlated field only resets the check mark and does not trigger data writeback; the bitmap operation is bound with the transaction log to ensure the atomicity of the state rollback;

[0056] The hierarchical review mechanism is used to synchronously review the strongly dependent field in real time, perform multi-thread parallel check, strictly review each field based on the strongly dependent rules in the rule library, and bind a distributed transaction lock to ensure that the review process is not interrupted; the weakly correlated field is processed asynchronously in batches, the field is marked as a state to be asynchronously checked, and the batch task is triggered in a fixed time window; the fields outside the radius are dynamically marked by topological analysis of the cross-form dependency graph; the fields skipped in the quantum state bitmap keep the original version number and do not trigger any operation.

[0057] The quantumized state operation further comprises deploying a state health monitoring module: periodically generating a hash check code of the quantum state bitmap of the strongly dependent field; locating the check range of the associated field based on the cross-form dependency graph; when the bitmap distortion rate is greater than 5%, the hierarchical review mechanism is activated to perform full review.

[0058] Specifically, the SHA-256 hash value of the quantum state bitmap of the strongly dependent field is calculated every ten minutes, the associated field bitmap within the radius R is selected according to the cross-form dependency map, and the difference rate between the current hash value and the historical reference value is calculated:

[0059]

[0060] wherein S z is the distortion rate, H n is the number of hash difference bits of the current bitmap, H s is the total number of bitmap bits;

[0061] When the quantum state bitmap distortion rate exceeds 5%, the associated field range is locked through the cross-form dependency map, and the hierarchical review mechanism is started to perform full-data repair, ensuring that the system state returns to consistency.

[0062] The health degree monitoring module further includes a self-calibration cycle mechanism, which includes: real-time screening of a key field set through a timing trigger calibration, correction of data distortion through a hierarchical repair strategy, and finally updating of system quantum state parameters to achieve cross-cycle stability control.

[0063] Specifically, full calibration is performed every 24 hours, covering all strongly dependent fields, and when distortion rate > 3% is detected for three consecutive times, the period is automatically shortened to 8 hours, and an emergency weight is calculated through an exponential decay model, with the formula being:

[0064]

[0065] wherein W e is the emergency weight, 0.5 is the time decay coefficient for controlling the decay rate, t c is the current time, t0 is the last calibration time, and W e > 0.7 triggers emergency calibration;

[0066] Through a composite weight calculation model, core field screening is performed, with the formula being:

[0067] W c = 0.5·ER+0.3·log2(DD+1)+0.2·BP,

[0068] wherein W c is the composite weight, BP is the business criticality, 0.5, 0.3, and 0.2 are the corresponding weight coefficients, and the sum is 1, the field with W c > 0.7 is selected as the core field, and the top 50 fields are sorted in descending order of W c ;

[0069] When S zWhen ≤3%, take incremental calibration, only recalculate the core field bitmap, and use local verification; when 3% < S z When ≤5%, take local review, expand to the associated field within radius R, and use transaction log version comparison; when 5% < S z When >5%, take full review, activate the hierarchical review mechanism, and force reconstruction of the full quantum state;

[0070] In the system state update process, first, a new quantum state bitmap version is generated and its SHA-256 hash value H new is calculated; then, pre-commit verification is performed to ensure data integrity; subsequently, H new is synchronized to all replica nodes through a distributed consensus protocol, and a consensus is reached when more than half of the nodes return an acknowledgement signal; after the consensus is completed, the transaction is submitted, and the strongly dependent fields take effect immediately, while the weakly associated fields are written into an asynchronous queue and take effect in batches within a 5-minute dynamic window; at the same time, the system retains the last three historical version bitmaps to form a version snapshot chain, supporting second-level rollback in abnormal scenarios. If the hash verification between nodes fails, the system automatically reverts to the previous stable version and links to the hierarchical review mechanism to recalibrate.

[0071] The self-calibration cycle mechanism further includes: according to the periodically generated hash check code, an error model is constructed through a drift compensation algorithm to perform pre-compensation offset correction before the rollback operation is executed; the error model is used for parameter iterative optimization to predict and correct cumulative errors.

[0072] Specifically, a hash check code sequence H = {H1, H2,..., H n} of the quantum state bitmap is automatically generated every 10 minutes, and the system calculates the error value of the adjacent period check code:

[0073]

[0074] where ΔH t is the error value, H t is the hash value generated at the current period (time point t), H t-1 is the hash value generated at the previous period (time point t-1), and L is the length of the hash code, which is fixed at 256 bits;

[0075] By quantifying the degree of state change, a space-time feature vector is extracted:

[0076] F t = [ΔH t , τ t , ρ t ],

[0077] where F t is the space-time feature vector, τ t is the time decay factor, and ρ t is the space decay factor.e is an index, t is a current time point, t0 is a reference time point; p t is a space correlation factor, N is the total number of node pairs participating in the calculation, w ij is the connection weight of node i and node j, d ij is the graph distance from node i to node j.

[0078] An error model is constructed by a drift compensation algorithm:

[0079] e t+1 = f1AH t + f2AH t-1 + b1t t + b2p t ,

[0080] where e t+1 is the prediction error value, t+1 is the next time, and the value range is [0, 1]; f1, f2 are time coefficients, and b1, b2 are space coefficients.

[0081] If e t+1 > 0.02, the system automatically generates a pre-compensation instruction, flips a specific bit in the quantum state bitmap, adjusts the generation logic of the transaction version number, calculates the new hash value of the compensated bitmap, and ensures that the corrected state approaches the expected state.

[0082] After each actual execution of the quantization operation, the real error AH a is recorded, and the deviation between the prediction error and the real error is calculated:

[0083] d = |e t+1 - AH a |,

[0084] where d is the cumulative error, and if d > 0.01, the parameter iterative optimization is triggered, and the coefficients f1, f2, b1, b2 are updated using gradient descent.

[0085] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:

[0086] The present application calculates the accurate rollback range by dynamic radius, combines the quantum state operation and the hierarchical review mechanism, realizes the resource optimization of incremental rollback, introduces the self-calibration cycle system, periodically generates the hash check code and constructs the error prediction model, actively suppresses the quantum state drift through pre-compensation offset correction, and guarantees the state consistency of the system in the distributed environment by means of parameter iterative optimization and distributed consensus protocol, and finally solves the problems of uncontrolled rollback range, long-term precision decay and distributed collaboration failure.

[0087] Embodiment three: when the system faces high-frequency data changes during the rollback operation in embodiment two, a chain rollback effect occurs, single-point changes trigger multi-level dependent rollback, and intermediate state loss causes data inconsistency. This phenomenon is called "rollback storm", which seriously affects system stability and business continuity. To solve these problems, this embodiment further improves embodiment two.

[0088] S3: when high-frequency data changes are detected, start the anti-rollback storm mechanism, obtain the cascading load index, and generate a rollback instruction set based on the index through cascading blocking;

[0089] The anti-rollback storm mechanism includes: real-time monitoring of rollback trigger frequency and cascading load index; dynamically executing a hierarchical fuse strategy based on the index; optimizing the storage structure of change records while executing the fuse strategy; filtering key rollback paths based on a cross-form dependency graph to generate a lightweight rollback instruction set; the cascading load index includes cascading depth and resource load index.

[0090] Specifically, a distributed probe cluster is deployed to collect host and container layer indicators every 100 ms and push them to the Kafka message queue; a stream computing engine consumes and aggregates multi-dimensional indicator data streams in real time to synchronously monitor three key dimensions: the number of rollback operations per second (time dimension), the cascading influence depth tracked in real time through the cross-form dependency graph (space dimension), and the resource load of CPU utilization, memory occupancy, and network loan usage (system dimension);

[0091] The system forms a dynamic fuse decision by real-time fusion of indicator data of the three key dimensions. First, continuously monitor the rollback operation frequency, and immediately determine it as a high-risk scenario when the number of triggers per second exceeds the 150 threshold; at the same time, track the cascading propagation depth, and once the influence path exceeds 8 hops of deep cascading, it indicates that the risk is rapidly spreading; in addition, analyze the system resource load comprehensively, especially when CPU utilization is continuously higher than 85% and memory occupancy is more than 90%, it reflects that the underlying resources are severely insufficient; therefore, the fuse decision baseline is established by quantifying the threshold: rollback operation frequency ≤ 50 times / second is the safe threshold, rollback operation frequency 51-150 times / second is the warning threshold, and rollback operation frequency > 150 times / second is the dangerous threshold; cascading influence depth ≤ 3 hops is the safe threshold, cascading influence depth 4-8 hops is the warning threshold, and cascading influence depth > 8 hops is the dangerous threshold; CPU utilization ≤ 75%, memory occupancy ≤ 80%, and network bandwidth ≤ 85% are the safe thresholds, CPU utilization 76%-85%, memory occupancy 81%-90%, and network bandwidth 86%-95% are the warning thresholds, and CPU utilization > 85% for 5 seconds, memory occupancy > 90%, and network bandwidth > 95% for 3 seconds are the dangerous thresholds;

[0092] According to the quantitative index, a joint decision mechanism is established: when any dimension reaches the danger threshold (such as frequency > 150 times / sec or depth > 8 jumps), the highest level fuse is triggered; if two dimensions deteriorate synchronously (such as high frequency operation accompanied by CPU overload), the medium level fuse is started; and for single dimension slight abnormality (such as frequency temporarily exceeding 50 times / sec), preventive current limiting is implemented. All judgments introduce time decay factor, and the index change in the last 5 minutes is given higher weight, ensuring that the system maintains acute response to sudden conditions, while avoiding false triggering caused by temporary fluctuations.

[0093] The quantum compression technology is used to optimize the storage structure of change records, and the continuous change records are incrementally encoded, and only the difference bits between adjacent versions are stored; based on the change frequency, the dynamic hierarchical storage is performed: the high frequency change field (> 5 times / sec) retains the complete track, the low frequency field (< 1 times / hour) only stores the first and last states, and the invalid field is directly filtered. At the same time, according to the topological relationship of the cross-form dependency graph, the strongly associated fields are merged into data blocks for storage. Run-length encoding is used to compress continuous unchanged fields (such as 20 unchanged fields compressed into a single byte mark), and 2-bit Huffman short encoding is allocated for high frequency change mode, combined with shared code table technology to reduce 30% metadata overhead.

[0094] According to the field composite weight and real-time cascade depth, a pruning algorithm is used to obtain a path key value:

[0095]

[0096] Among them, PCV is the path key value, D c is the cascade depth, PL is the path length, the logarithmic term suppresses the weight expansion of long path, and avoids excessive attention to super-long links; during normal operation, the PCV threshold is 0.7; during medium level fusing, the PCV threshold is 0.8; and during high level fusing, the PCV threshold is 0.9;

[0097] All paths in the dependency graph are traversed, the PCV value of each path is calculated, when the PCV is greater than the current threshold, the key path is marked, the complete rollback link is retained, and high priority computing resources are allocated; otherwise, the decay path is marked, moved into cold storage, and 45% of the memory occupation is released, and finally the lightweight rollback instruction set is generated.

[0098] The technical solutions in the embodiments of the application have at least the following technical effects or advantages:

[0099] The application constructs a dynamic fuse decision mechanism by monitoring rollback trigger frequency, cascade depth and system resource load indicators in real time, and triggers a hierarchical fuse strategy when any dimension breaks through the danger threshold; quantum compression technology is used to perform incremental encoding and hierarchical storage on change records, and a graph pruning algorithm is used to calculate path key values, and by dynamically adjusting the PCV threshold (screening key rollback paths, generating a lightweight instruction set to eliminate redundant operations, and finally suppressing chain rollback diffusion when resources are overloaded to ensure stable operation of core business links.

[0100] Embodiment four: Embodiment three dynamically triggers a fuse strategy by monitoring rollback frequency, cascade depth and resource load indicators in real time, and generates a lightweight rollback instruction set to suppress chain rollback storm in high-frequency scenarios by combining quantum compression storage and graph pruning algorithm, but lacks a state synchronization mechanism when a node fails and a cross-node error propagation blocker is not configured, resulting in a gap in the fault tolerance capability of the distributed environment. Embodiment four further improves embodiment three.

[0101] S4: Perform distributed fault-tolerant control, build a node cooperative protection system through cascading load, and output fault-tolerant operation instructions.

[0102] The distributed fault-tolerant control includes: eliminating quantum state delay through versioned state synchronization;

[0103] Specifically, a dynamic vector clock synchronization mechanism is used to maintain a vector clock identifier (format: {node ID: sequence number}) for each distributed node. When the quantum state bitmap is updated, the additional vector clock mark is synchronized. When resolving conflicts, the version with the higher clock sequence number is preferred, and when the sequence numbers are the same, the arbitration is performed according to the preset node priority. Through incremental synchronization optimization, only the changed bitmap segment is transmitted, and a reverse entropy protocol is used to achieve efficient synchronization: node A sends a hash value to node B, B requests the missing segment after comparing the difference, and A pushes the difference bit data.

[0104] The distributed fault-tolerant control also includes: using a partition-aware arbitration mechanism to avoid network splitting risk;

[0105] Specifically, a node connectivity matrix is constructed, and if the network delay between nodes is ≤50ms, it is marked as connected (value 1), otherwise disconnected (value 0). When the matrix is broken into independent subgraphs, arbitration is automatically triggered, and each node submits the current weight W i , the calculation formula is:

[0106] W i = 0.6·CPU u + 0.4·SZ,

[0107] where CPU uFor CPU utilization, SZ is the total number of strongly dependent fields of the node, and is normalized to [0, 1]; the total weight is calculated i If the total weight i >0.7, the write operation is allowed, otherwise the cluster enters the read-only mode and obtains the arbitration decision instruction.

[0108] The distributed fault-tolerant control further comprises: preventing historical version error recovery based on space-time consistency verification;

[0109] Specifically, to prevent the distributed node from being recovered to an invalid or outdated historical version, a space-time consistency verification module is deployed to ensure that the target version is not expired and the current business topology is compatible. The historical version V t to be recovered is received, which contains a version number, a timestamp T v and a space topology S v . Time consistency verification is performed to obtain the current system time T c , and the time difference is calculated:

[0110] ΔT=|T c -T v |,

[0111] Wherein, ΔT is the time difference, if ΔT>300 seconds, it is determined that the version is expired, and the recovery is refused, and a version recovery instruction is obtained.

[0112] Space consistency verification is performed, the field node set N v associated with T v is obtained through the cross-form dependency graph, the S c node set N C of the current graph topology is extracted, and the topology difference degree is calculated:

[0113]

[0114] If ΔS>0, it is determined that the version space is invalid;

[0115] According to the time consistency verification and the space consistency verification, the comprehensive verification result is output, if both pass, the recovery is allowed, otherwise the recovery is refused, and the system will record the warning log and notify the operation and maintenance.

[0116] The distributed fault-tolerant control further comprises: deploying an error propagation blocker to suppress chain diffusion;

[0117] Specifically, a fault propagation prediction model is constructed based on the cross-form dependency graph to quantify the fault transmission risk between any two nodes, and the calculation formula is:

[0118]

[0119] Wherein, P prP pr [i][j] is the probability of node i failure affecting j, L ij is the number of failure paths from node i to j, L s is the total number of paths;

[0120] Calculate the propagation risk value R j for each node j, the calculation formula is:

[0121] R j =∑P pr [i][j]×G i ,

[0122] wherein G i is the failure state: normal = 0, warning = 0.5, failure = 1.0;

[0123] According to R j , the three-level protection is dynamically executed: R j <0.3, no intervention is triggered, and the business flow is smooth; 0.3≤R j <0.6, call asynchronous batch processing, temporarily store and batch process the request, and absorb the burst traffic; R j ≥0.6, disconnect the high-risk connection of the node and repair the affected field, and obtain the blocking execution instruction.

[0124] Through the cascading load indexes of the total number of strong dependency fields, failure propagation probability, propagation risk value, and topology difference, according to the collaborative operation of partition-aware arbitration, error propagation blocker, and spatiotemporal consistency check, a node collaborative protection system is constructed and arbitration decision instructions, version recovery instructions, and blocking execution instructions are outputted, which constitute fault-tolerant operation instructions.

[0125] The technical solutions in the embodiments of the application have at least the following technical effects or advantages:

[0126] The application adopts a distributed fault-tolerant control system, and realizes node failure self-healing, error chain diffusion inhibition, and high availability in a distributed environment through the collaborative operation of four core technical modules of versioned state synchronization, partition-aware arbitration, spatiotemporal consistency check, and error propagation blocker. The measured single-node downtime recovery time is compressed from 8.2 seconds to 0.3 seconds, the error propagation blocking rate is 98.7%, and the application supports deep integration with cascading audit and storm fuse mechanism to form a closed-loop protection.

[0127] The above only describes the preferred embodiments of the application and is not used to limit the application. For those skilled in the art, the application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall be included in the protection scope of the application.

Claims

1. A data audit method based on big data, characterized in that: include: S1: Obtain multimodal data from forms on the big data platform, extract features from the multimodal data to generate a feature set, create a rule base based on the feature set, and construct a cross-form dependency graph based on the feature set and the rule base; S2: Cascading audits are performed through a cross-form dependency graph. If an audit fails, the affected fields are located in reverse according to the connection direction of the graph, and incremental state rollback and re-audit are performed on the affected fields. S3: When high-frequency data changes are detected, the rollback storm resistance mechanism is activated, cascade load indicators are obtained, and a rollback instruction set is generated based on the indicators through cascade blocking. S4: Execute distributed fault-tolerant control, build a node collaborative protection system through cascading loads, and output fault-tolerant operation instructions.

2. A data audit method based on big data according to claim 1, characterized in that: The multimodal data includes field metadata, historical audit logs, and business rule documents; The feature set includes: performing structured parsing on field metadata to generate structural features, performing statistical analysis on historical audit logs to generate behavioral features, and performing semantic segmentation on business rule documents to generate logical semantic features; The rule base includes: strong dependency rules, weak association rules, dynamic evolution rules and conflict resolution rules.

3. A data audit method based on big data according to claim 2, characterized in that: The cross-form dependency graph includes: generating a node set based on a feature set, each node is associated with a form field; realizing cross-form node connection through a bidirectional linked list, and the connection direction represents the audit impact relationship between fields; the connection corresponding to the strong dependency rule is marked as a strong dependency type, and the connection corresponding to the weak association rule is marked as a weak association type.

4. The data audit method based on big data according to claim 1, characterized in that: The cascade audit also includes: performing complexity sorting audit on the current form fields; activating cross-form dependency graph scanning to identify related form field combinations; batch loading related field data, and performing audit according to the rule type in the rule library.

5. The data audit method based on big data according to claim 1, characterized in that: The incremental state rollback includes: first performing dynamic radius calculation and determining the precise rollback range through graph association distance; then implementing quantized state operations and creating quantum state bitmaps for strongly dependent fields; performing transaction version rollback on strongly dependent fields within the radius, resetting only the check mark for weakly associated fields, and keeping the state of fields outside the radius unchanged; finally, using a hierarchical review mechanism, conducting real-time synchronous review of strongly dependent fields, asynchronous batch processing of weakly associated fields, and automatically skipping fields outside the radius, collaboratively eliminating redundant reviews.

6. A data audit method based on big data according to claim 5, characterized in that: The quantized state operation also includes deploying a state health monitoring module: periodically generating hash check codes for the quantum state bitmap of strongly dependent fields; locating the check range of associated fields based on the cross-form dependency graph; and activating a hierarchical review mechanism to perform a full review when the bitmap distortion rate is greater than five percent.

7. A data audit method based on big data according to claim 6, characterized in that: The health monitoring module also includes a self-calibration cycle mechanism, which includes: triggering calibration through time, screening key field sets in real time, correcting data distortion through a hierarchical repair strategy, and finally updating the system quantum state parameters to achieve cross-cycle stability control.

8. The data audit method based on big data according to claim 7, characterized in that: The self-calibration loop mechanism also includes: constructing an error model through a drift compensation algorithm based on a periodically generated hash check code, and performing pre-compensation offset correction before the rollback operation is executed; using the error model to iteratively optimize parameters to predict and correct accumulated errors.

9. The data audit method based on big data according to claim 1, characterized in that: The anti-rollback storm mechanism includes: real-time monitoring of rollback trigger frequency and cascade load indicators; dynamic execution of hierarchical circuit breaking strategies based on indicators; optimizing the storage structure of change records while executing the circuit breaking strategies; screening key rollback paths based on cross-form dependency graphs to generate a lightweight rollback instruction set; the cascade load indicators include cascade depth and resource load indicators.

10. The data audit method based on big data according to claim 1, characterized in that: The distributed fault-tolerant control includes: eliminating quantum state delays through versioned state synchronization, adopting a partition-aware arbitration mechanism to avoid network split risks, preventing historical version error recovery based on spatiotemporal consistency verification, and deploying an error propagation blocker to suppress chain diffusion.

Citation Information

Patent Citations

  • Data auditing methods, devices, and storage media based on big data

    CN111966673B

Cited By

  • Document auditing system and method based on image recognition

    CN121214470A

  • Enterprise behavior deviation detection method and system based on cross report comparison

    CN121524189A

  • A method and system for detecting enterprise behavior deviation based on cross-report comparison

    CN121524189B

  • Clinical test data table collaborative auditing method and system based on file state driving

    CN122067684A

  • A File-State-Driven Collaborative Review Method and System for Clinical Trial Data Tables

    CN122067684B