A method, system, and storage medium for business data testing

Through alliance blockchain and differential privacy technology, the distributed data processing traceability platform is built, which solves the data consistency verification problem of existing business data testing methods in multi-source heterogeneous scenarios, and realizes efficient, secure and automated data testing and repair, improving the credibility and consistency of data processing.

CN120105494BActive Publication Date: 2025-07-11SHENZHEN FARBEN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510594413.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-11
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing business data testing methods lack support for the consistency verification of encrypted data sharding and hash chains in multi-source heterogeneous and dynamically changing data processing scenarios. The consensus mechanism is high latency and poor real-time, and the lack of compact proof structure design, resulting in a rough granularity of data integrity verification and an increase in storage overhead.

Method used

Adopting alliance blockchain and differential privacy technology, a distributed data processing traceability platform is built, a synthetic test data set is generated through a data isolation verification environment, combined with smart contracts for automated verification, and a directed acyclic graph of risk calculations is generated, and an automated correction workflow is realized through incremental compensation processing equipment.

Benefits of technology

A highly trusted, automated and traceable data testing system has been realized, which improves test security, accuracy and efficiency, accurately locates different sources and realizes system consistency verification, reduces manual intervention, and improves the intelligence level of data quality management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105494B_ABST
    Figure CN120105494B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data transfer and verification, and particularly to a business data testing method, system, and storage medium. The method includes the following steps: obtaining an original business data stream; deploying a regulatory reporting shadow processing sandbox based on the original business data stream to obtain a data isolation and verification environment; configuring a consortium blockchain server based on the original business data stream to construct a distributed data processing traceability platform; using the original business data stream on the distributed data processing traceability platform to generate a data processing proof, obtaining a data processing hash chain; configuring differential privacy protection processor parameters based on the data isolation and verification environment to obtain a privacy protection test data generator; training the privacy protection test data generator to generate a synthetic test data set. By combining the consortium blockchain and differential privacy technologies, the present invention constructs a highly reliable, automated, and traceable data testing system, significantly improving data security, verification efficiency, and system consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data transfer and verification, and particularly to a business data testing method, system and storage medium. Background Art

[0002] Traditional business data testing methods mostly rely on centralized database systems for data verification and tracking, making it difficult to meet the data processing scenarios of multi-source heterogeneous and dynamically changing data. In recent years, the rise of consortium blockchain technology has provided a new solution for the secure processing and trusted verification of data. By virtue of the immutability, timestamp mechanism and consensus protocol of the blockchain, the verifiability and traceability of the whole process of business data processing can be realized. However, the existing testing methods still have the following deficiencies: most solutions lack effective support for the consistency verification of encrypted data sharding and hash chains, resulting in rough data integrity verification granularity; the existing consensus mechanisms have high latency and poor real-time performance, and are not suitable for high-frequency business scenarios; there is a lack of a compact proof structure design, causing redundancy in indexes and backup data and increasing storage overhead. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a business data testing method, system and storage medium to solve at least one of the above technical problems.

[0004] To achieve the above object, a business data testing method includes the following steps:

[0005] Step S1: Obtain the original business data stream; deploy a regulatory reporting shadow processing sandbox based on the original business data stream to obtain a data isolation verification environment;

[0006] Step S2: Configure a consortium blockchain server based on the original business data stream to build a distributed data processing traceability platform;

[0007] Step S3: Generate a data processing proof based on the original business data stream by using the distributed data processing traceability platform to obtain a data processing hash chain;

[0008] Step S4: Configure differential privacy protection processor parameters based on the data isolation verification environment to obtain a privacy protection test data generator; train the privacy protection test data generator to generate a synthetic test data set;

[0009] Step S5: Deploy a smart contract based on the synthetic test data set to the distributed data processing traceability platform to obtain an automated verification engine; conduct a data comparison analysis between the source system and the target system based on the data processing hash chain and the automated verification engine to generate a risk calculation difference report; configure a risk calculation graph visualization server based on the synthetic test data set and the data isolation verification environment and generate a risk calculation directed acyclic graph;

[0010] Step S6: preset a difference positioning processor based on the risk calculation difference report and the risk calculation directed acyclic graph operation, configure an incremental compensation processing device, generate an automated correction workflow; execute the automated correction workflow, and generate a system consistency verification result.

[0011] The present invention constructs a highly reliable, automated, and traceable data testing system by introducing alliance blockchain and differential privacy technology, which is significantly better than the traditional data verification method that relies on centralized databases. First, the isolation of business data and test environment is realized to ensure that the real business system is not affected during the test process, and the test security is improved; secondly, with the help of the blockchain's immutability and timestamp mechanism, the authenticity and verifiability of the entire data processing process are ensured, and the shortcomings of the existing methods in data sharding encryption and hash chain consistency verification are solved; thirdly, the differential privacy protection mechanism is introduced to effectively shield sensitive information when generating test data, and improve test compliance and data security; in addition, automated verification and difference analysis are realized by combining synthetic test data with smart contracts, reducing manual intervention and improving test efficiency and accuracy; the risk calculation directed acyclic graph constructed provides a clear visualization of risk causal paths, which helps to accurately locate the source of differences; finally, through the automated correction workflow linked by difference positioning and incremental compensation mechanism, an integrated closed-loop processing from difference discovery to system repair is realized, ensuring that the system finally meets the consistency verification requirements, and significantly improving the intelligence and automation level of data quality management.

[0012] Preferably, the present invention further provides a business data testing system for executing the above-mentioned business data testing method, the business data testing system comprising:

[0013] The data isolation sandbox deployment module is used to obtain the original business data flow; deploy the regulatory reporting shadow processing sandbox based on the original business data flow to obtain the data isolation verification environment;

[0014] The blockchain traceability platform construction module is used to configure the alliance blockchain server based on the original business data flow and build a distributed data processing traceability platform;

[0015] The data processing proof generation module is used to generate data processing proof based on the distributed data processing traceability platform using the original business data stream to obtain the data processing hash chain;

[0016] A privacy synthetic test data generation module is used to configure the differential privacy protection processor parameters based on the data isolation verification environment to obtain a privacy protection test data generator; train the privacy protection test data generator to generate a synthetic test data set;

[0017] Automated Difference Analysis and Risk Map Construction Module, which is used to deploy smart contracts for the distributed data processing traceability platform based on the synthetic test data set to obtain an automated verification engine; perform data comparison and analysis between the source system and the target system based on the data processing hash chain and the automated verification engine to generate a risk calculation difference report; configure a risk calculation graph visualization server based on the synthetic test data set and the data isolation verification environment, and generate a risk calculation directed acyclic graph;

[0018] Difference Repair and Consistency Verification Module, which is used to operate a preset difference location processor based on the risk calculation difference report and the risk calculation directed acyclic graph, and configure an incremental compensation processing device to generate an automated correction workflow; execute the automated correction workflow to generate a system consistency verification result.

[0019] By organically integrating the data isolation sandbox, blockchain traceability platform, privacy synthetic test data generation, automated difference analysis and risk map construction, and difference repair and consistency verification modules, the present invention forms an efficient, secure, and automated data verification and repair system. First, the data isolation sandbox deployment module provides isolation and security protection for the original business data stream, ensuring privacy and security during the data processing process; the blockchain traceability platform construction module provides traceability and immutability for data processing through blockchain technology, ensuring transparency and credibility throughout the data processing process; the privacy synthetic test data generation module generates a synthetic test data set while ensuring privacy protection, providing compliant data that does not disclose sensitive information for subsequent testing; the automated difference analysis and risk map construction module accurately identifies the differences between systems through the automated deployment of smart contracts and data comparison analysis, and uses visualization technology to help decision-makers quickly grasp the risk sources and impacts; finally, the difference repair and consistency verification module ensures the automation, accuracy, and efficiency of the system repair process through the application of incremental compensation and automated correction workflows, significantly improving the speed and quality of system consistency verification. Overall, this method greatly improves the accuracy, operability, and security of data processing verification, and is applicable to data consistency verification and risk management across systems and platforms.

[0020] Preferably, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the business data testing method described above is implemented. Description of the Drawings

[0021] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes, and advantages of the present invention will become more obvious:

[0022] Figure 1 It is a schematic diagram of the step flow of the business data testing method of the present invention;

[0023] Figure 2 For Figure 1 the detailed step flow schematic diagram of step S1 in Specific implementation mode

[0024] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.

[0025] To achieve the above object, please refer to Figures 1 to 2 , the present invention provides a business data testing method, and the method includes the following steps:

[0026] Step S1: Obtain the original business data stream; deploy a regulatory reporting shadow processing sandbox based on the original business data stream to obtain a data isolation verification environment;

[0027] Step S2: Configure a consortium blockchain server based on the original business data stream to build a distributed data processing traceability platform;

[0028] Step S3: Use the original business data stream on the distributed data processing traceability platform to generate data processing proof, and obtain a data processing hash chain;

[0029] Step S4: Configure differential privacy protection processor parameters based on the data isolation verification environment to obtain a privacy protection test data generator; train the privacy protection test data generator to generate a synthetic test data set;

[0030] Step S5: Deploy a smart contract on the distributed data processing traceability platform based on the synthetic test data set to obtain an automated verification engine; perform data comparison and analysis between the source system and the target system based on the data processing hash chain and the automated verification engine to generate a risk calculation difference report; configure a risk calculation graph visualization server based on the synthetic test data set and the data isolation verification environment, and generate a risk calculation directed acyclic graph;

[0031] Step S6: Operate a preset difference positioning processor based on the risk calculation difference report and the risk calculation directed acyclic graph, and configure an incremental compensation processing device to generate an automated correction workflow; execute the automated correction workflow to generate a system consistency verification result.

[0032] In the embodiment of the present invention, referring to Figure 1 shown, it is the step flow schematic diagram of a business data testing method of the present invention. In this example, the business data testing method includes the following steps:

[0033] Step S1: Obtain the original business data stream; deploy a regulatory reporting shadow processing sandbox based on the original business data stream to obtain a data isolation verification environment;

[0034] In the embodiment of the present invention, first, the original business data stream is obtained through the transaction database centralized management system. The data stream includes four fields: transaction date, transaction amount, transaction account, and transaction type, and the data volume is 5 million structured records per day. Then, a server rack array is deployed based on the original business data stream, specifically configured with 10 HP ProLiant DL380 servers, each equipped with an Intel Xeon E5-2699 v4 processor, 256GB ECC memory, and a 10TB RAID 10 storage array to build the physical infrastructure of the shadow processing sandbox. Then, based on the physical infrastructure of the shadow processing sandbox, core component service orchestration is performed, and 50 business microservice nodes are deployed using Docker container technology. Each node is configured with a 4GB memory limit and a 2-core CPU resource limit, and automatic scaling is achieved through the Kubernetes cluster management system version 1.22.5 to obtain a microservice cluster. After that, parameter tuning of the data normalization engine is performed based on the original business data stream, including setting a field mapping matrix, configuring a data type conversion rule set (unifying all date formats to the ISO-8601 standard and unifying the amount to two decimal places), performing UTF-8 encoding standardization, setting the data cleaning threshold to 99.95% accuracy, and finally obtaining a standardized data conversion module. Subsequently, a security isolation mechanism is deployed based on the standardized data conversion module to implement a three-layer network partition (DMZ area, security area, core area), configure a deep packet inspection firewall blocking rate of 99.99%, implement a strict access control list (ACL), limit the IP address segment access range to 192.168.1.0 / 24, and establish a data transmission encryption channel using the AES-256-GCM algorithm to obtain a data isolation control unit. Finally, the regulatory reporting shadow processing sandbox is integrally deployed based on the standardized data conversion module and the data isolation control unit, configured with a real-time data synchronization mechanism (synchronization delay not exceeding 5 seconds), implemented database table space isolation (each regulatory reporting module has an independent table space without interference), deployed a memory isolation protection mechanism (address space layout randomization ASLR technology), and implemented an image data backup strategy (incremental backup interval of 15 minutes and full backup once every 24 hours), and finally obtain a data isolation verification environment.

[0035] Step S2: Configure a consortium blockchain server based on the original business data stream to build a distributed data processing traceability platform;

[0036] In the embodiments of the present invention, first, the hardware resources of the high-performance computing server cluster are allocated based on the original business data stream. Seven Dell PowerEdge R740 servers are configured, each equipped with a dual-way Intel Xeon Gold 6258R processor (28 cores and 56 threads), 768GB of DDR4-3200MHz ECC memory, an 8TB NVMe SSD storage array, and a 100Gbps network interface card to form an array of physical nodes of the consortium blockchain. Then, based on the array of physical nodes of the consortium blockchain, the parameters of the consensus mechanism engine are configured. The Practical Byzantine Fault Tolerance (PBFT) consensus algorithm is adopted, the number of verification nodes is configured to be 4, and the number of fault-tolerant nodes is f = 1 (i.e., (3f + 1) = 4 verification nodes). The block generation period is set to 2 seconds, the upper limit of the number of transactions per single block is 10,000, and the transaction verification timeout is set to 500 milliseconds to obtain a high-throughput consensus network. Subsequently, based on the high-throughput consensus network, a blockchain server cluster is deployed. The core components of the Hyperledger Fabric v2.4 version are installed, 4 organizational (Organization) nodes are configured, each organization contains 2 peer nodes (Peer Node) and 1 ordering node (OrderingNode), the number of channels (Channel) is set to 2, the execution environment of the smart contract (Chaincode) is a Docker container, and the maximum number of parallel transaction processing is 5,000 transactions per second to obtain the core services of the consortium blockchain. After that, a hardware accelerator for cryptographic security components is installed, and the Intel QuickAssist Technology acceleration card QAT8970-AEPW is deployed. The RSA-4096 signature verification throughput is configured to reach 20,000 times per second, the Elliptic Curve Digital Signature Algorithm (ECDSA) P-256 signature verification throughput reaches 40,000 times per second, the SHA-256 hash calculation throughput reaches 40Gbps, and it has a dedicated hardware random number generator (TRNG) to obtain a high-performance signature verification unit. Subsequently, a secure channel for the inter-node communication protocol is built based on the high-performance signature verification unit, and TLS1.3 Encryption protocol, configure the parameters of the Elliptic Curve Diffie-Hellman Ephemeral (ECDHE) algorithm, use the TLS_AES_256_GCM_SHA384 cipher suite, set the certificate rotation period to 7 days, update the communication key every hour, set the data transmission encryption strength to 256 bits, establish a two-way authentication mechanism between nodes, and set the maximum bandwidth limit of the transport layer to 10 Gbps to obtain an encrypted data transmission network; finally, deploy a distributed data processing and traceability platform based on the consortium blockchain core service and the encrypted data transmission network to achieve blockchain storage of transaction data (transaction hash index links to the original transaction), configure a data processing event recording mechanism (processing time accurate to milliseconds), implement a traceability query interface (support multi-dimensional combined queries, query response time not exceeding 200 milliseconds), deploy a real-time monitoring system (system status refresh frequency of 1 second), set up a full-link trace log (log storage period of 180 days), and complete the construction and configuration of the distributed data processing and traceability platform.

[0037] Step S3: Generate data processing proofs using the original business data stream based on the distributed data processing and traceability platform to obtain a data processing hash chain;

[0038] In the embodiments of the present invention, the original business data stream is preprocessed by format conversion. The XML-to-JSON conversion algorithm is applied to convert unstructured XML data (with a multi-layer nesting depth not exceeding 5 layers) into a standardized JSON format. The data cleaning process includes null value filling (using -1 to replace numerical null values and using "N / A" to replace character null values), outlier handling (removing values exceeding 3 standard deviations), and field standardization (unifying the date format to the ISO-8601 standard and the monetary precision to 2 decimal places) to obtain a standardized data packet. Then, an encryption component is loaded into the standardized data packet, and the AES-256-GCM encryption algorithm is applied to encrypt the key fields (the encryption key length is 256 bits, the initialization vector IV length is 12 bytes, and the authentication tag length is 16 bytes). The HMAC-SHA512 message authentication code of the overall data packet is calculated (the key length is 512 bits) to obtain an encrypted and protected data set. Then, the encrypted and protected data set is sharded based on the consortium blockchain server (the single-piece size is 1MB, and the overlap rate is 5%). Based on the sharding result, the hash processor parameters are configured. The hash algorithm combination uses SHA3-512 and Blake2b for double hashing. The size of the hash calculation thread pool is set to 128 threads, the hash buffer size is set to 256MB, and the batch processing threshold is set to 10,000 records to obtain a hash generation engine. Then, the standardized data packet is synchronized with the timestamp. The Network Time Protocol (NTP) server is used to synchronize with the atomic clock (the synchronization accuracy is at the microsecond level). Based on the hash generation engine, multiple hash calculations are performed on the timestamp synchronization result. The calculation formula is HashFinal = SHA3-512(Blake2b(data) + timestamp + random entropy value). The random entropy value is taken from the 32-byte random number output by the hardware random number generator to obtain the original proof of data processing. Subsequently, the original proof of data processing is processed by a proof aggregation processor. The Merkle Tree structure is adopted, the tree height limit is 16 layers, the leaf node hash value uses the original proof, and the calculation formula for the parent node hash value is H(left child node hash + right child node hash). The root node calculation uses the weighted hash algorithm with the weight coefficient being the reciprocal of the child node depth to obtain a proof aggregation tree. Then, the proof aggregation tree is consensus-verified based on the distributed data processing traceability platform. The multi-round PBFT consensus process is adopted (at least 3f + 1 nodes reach an agreement, where f is the number of fault-tolerant nodes). The consensus threshold is set to 75% node approval, the consensus timeout is set to 5 seconds, and the consensus verification result is written into the high-speed chain memory (the write speed is not less than 200,000 IOPS) to obtain the persistent proof data. After that, the proof compression processor is operated based on the persistent proof data. The zlib compression algorithm is adopted (the compression level is set to 9), the compression ratio is not less than 60%, and the compression block size is 64KB to obtain a compact proof structure;Redundancy check device monitoring is performed on the compact proof structure. The Reed-Solomon forward error correction code (RS(10,4), that is, 10 data blocks and 4 check blocks) is adopted, and the error correction ability is such that the data can still be recovered even if any 4 block data are lost, obtaining a reliability-enhanced proof. Based on the reliability-enhanced proof, a proof index generator is called, and a B+ tree index structure is adopted (the leaf node capacity is 128 records, and the internal node branching factor is 64). The index key is a combined key of timestamp + transaction hash, and the total number of index entry points does not exceed 16, obtaining a proof index table. Then, chain encapsulation processing is performed on the proof index table, and a chained hash pointer structure is adopted (each block contains the hash value of the previous block). The block size is fixed at 4MB, and the block header contains the current block hash value, the previous block hash value, the Merkle root, the timestamp, and the nonce, and it is written into the hash chain backup device (using a RAID 10 storage array with a total capacity of 50TB), and a timed snapshot backup is performed (once per hour), obtaining a distributed backup hash chain. Finally, hash chain consistency verification is performed on the distributed backup hash chain, and a chained backtracking verification algorithm is adopted (backtracking from the latest block to the genesis block), and the verification frequency is once every 10 minutes. The consistency determination criterion is that the hash values of all nodes match 100%, obtaining the final data processing hash chain.;

[0039] Step S4: Configure the differential privacy protection processor parameters based on the data isolation verification environment to obtain a privacy protection test data generator; train the privacy protection test data generator to generate a synthetic test data set;

[0040] In the embodiments of the present invention, first, a data distribution feature set is extracted based on a data isolation verification environment. By calculating the statistical features of each field in the original business data stream, including mean, variance, median, quartile, kurtosis, skewness, correlation coefficient between fields, integrity features (missing rate not exceeding 0.1%), dimensionality features (high-dimensional sparsity not exceeding 20%), time-series features (autocorrelation coefficient calculation window set to 30 days), and field importance features (calculated by information gain ratio, threshold set to 0.05); then, based on the data distribution feature set, a differential privacy sensitivity analyzer is calibrated, and the field sensitivities are classified into layers (high sensitivity: ID number, account number, mobile phone number; medium sensitivity: name, address, amount; low sensitivity: transaction type, transaction time). The global sensitivity δ value is calculated as the maximum change in the L1 norm of each field, set to 10.5, and a sensitivity threshold matrix is generated (high sensitivity threshold 0.2, medium sensitivity threshold 0.5, low sensitivity threshold 1.0); then, the noise injection ratio of the sensitivity threshold matrix is calculated, and the Laplace noise mechanism is adopted. The noise ratio formula is b = Δf / ε, where Δf is the sensitivity and ε is the privacy budget. The overall privacy budget ε = 3.0 is set and allocated according to field sensitivity (high sensitivity ε1 = 0.5, medium sensitivity ε2 = 1.0, low sensitivity ε3 = 1.5) to obtain a privacy budget configuration table; then, based on the privacy budget configuration table, the parameters of the differential privacy protection processor are loaded. The seed value of the noise generator is set to a 128-bit random number generated by a hardware random number generator, the noise distribution parameter b = Δf / ε (different for each field) is configured, the query sensitivity threshold is set (the impact of a single query result does not exceed 0.1%), and the data access frequency limit is set (the number of accesses to the same data set per hour does not exceed 10 times) to obtain an initialized processor instance; then, a privacy-utility balance test is performed on the initialized processor instance, the data availability index is calculated (statistical significance retention rate not less than 95%), the privacy leakage risk index is analyzed (differential privacy attack success rate does not exceed 0.1%), and the utility-privacy balance coefficient α = 0.7 (weight biased towards data availability) is calculated. The privacy budget allocation is adjusted by the bisection method, the number of iterations is set to 50 times, and the convergence threshold is set to 0.001 to obtain an optimized processor configuration; subsequently, based on the optimized processor configuration, a data generation engine is assembled, and a generator architecture is constructed, including a data distribution modeling module (separate modeling of the distribution of each field), a correlation preservation module (based on the Gaussian Copula method, maintaining the correlation error between fields not exceeding 5%), and a privacy budget allocation module (dynamically adjusting the privacy budget of each field, with the total budget remaining 3.0) to obtain a prototype of a privacy protection test data generator; then, a privacy protection strength verification device test is performed on the prototype of the privacy protection test data generator, and a membership inference attack test is executed (the attack success rate needs to be lower than 0.5%)), the attribute inference attack test (the accuracy of sensitive attribute inference needs to be close to the random guessing level of 50%), the differential attack test (the detection probability of the output difference of adjacent datasets does not exceed e^ε = 20.1%), and obtain the generator compliance report; finally, based on the generator compliance report, fine-tune the parameters of the privacy protection test data generator prototype, adjust the privacy budget allocation ratio (high sensitivity ε1 = 0.4, medium sensitivity ε2 = 1.1, low sensitivity ε3 = 1.5), optimize the noise distribution parameters (increase the noise intensity by 10% for highly sensitive fields), adjust the data generation batch size to 10,000 records per batch, set the generation timeout limit to 30 seconds per batch, and obtain the privacy protection test data generator; then, based on the data isolation verification environment, extract training samples from the original business data stream, adopt a stratified sampling strategy (stratify by business type, and the sample size of each layer is the same as the original ratio), the sampling ratio is 20% of the total data volume, the minimum sample quantity limit is 50,000 records, and the KS test p-value for sample representativeness verification is required to be greater than 0.05 to obtain the training dataset; then, input the training dataset into the privacy protection test data generator, execute the parameter learning process, set the number of iterations to 1,000 times, the learning rate to 0.001, the batch size to 128 records, and the convergence condition to that the loss function changes less than 0.0001 for 5 consecutive iterations to obtain the generator model during training; then, based on the generator model during training, perform data distribution consistency detection, calculate the JS divergence between the original distribution and the generated distribution (it needs to be less than 0.05), execute the two-sample KS test (the p-value needs to be greater than 0.05), calculate the correlation retention degree between fields (the difference in Pearson correlation coefficients does not exceed 0.1), and evaluate the feature retention rate (the key business feature retention rate is not less than 95%) to obtain the model performance evaluation report; subsequently, perform parameter tuning device analysis on the model performance evaluation report, identify the poorly performing feature dimensions (fields with JS divergence greater than 0.05), adjust the corresponding dimension generation parameters accordingly (increase the number of training rounds by 50%, improve the sampling accuracy by 30%), retrain the problem dimensions 5 times, and verify the adjustment effect to obtain the optimized generator model; finally, based on the optimized generator model, call the high-performance data synthesis engine, set the generated data volume to 10 times the original data volume, the generation speed requirement to be not less than 10,000 records per second, set the number of parallel generation threads to 16, perform real-time quality monitoring during the generation process (check the distribution consistency every 100,000 records generated), and execute batch-level data merging (the memory buffer size is set to 1GB) to obtain the final synthetic test dataset.

[0041] Step S5: Deploy smart contracts to the distributed data processing traceability platform based on the synthetic test dataset to obtain an automated verification engine; perform data comparison and analysis between the source system and the target system based on the data processing hash chain and the automated verification engine to generate a risk calculation difference report; configure a risk calculation graph visualization server based on the synthetic test dataset and the data isolation verification environment, and generate a risk calculation directed acyclic graph;

[0042] In the embodiments of the present invention, first, the identity authentication of the U shield physical encryption module is initialized. By accessing the Feitian Integrity ePass3000 model U shield that conforms to the PKCS#11 standard interface, a two-factor authentication process is executed, including PIN code verification (8-digit PIN code, with a limit of 3 error attempts) and fingerprint recognition (FAR false acceptance rate < 0.0001%, FRR false rejection rate < 1%). The X.509 digital certificate stored inside the U shield is read (RSA-4096-bit key, with a certificate validity period of 1 year), and a certificate chain verification is performed (a three-level verification chain of root certificate, intermediate certificate, and user certificate). The online certificate status check is completed (the OCSP response timeout is set to 2 seconds), and a hardware security authenticator is obtained. Then, based on the hardware security authenticator, a digital signature of the smart contract code is performed. First, the verification logic is encapsulated into a Solidity language smart contract (the number of code lines is controlled within 500 lines, and the number of functions does not exceed 20). The contract includes a data hash verification function, a comparison and analysis logic, and a difference report generation function. The smart contract is compiled to generate bytecode (the compilation optimization level is set to the highest). The U shield is used to execute the ECDSA-secp256k1 signature algorithm (the hash function is SHA3-256) to sign the bytecode. The length of the signature data is 72 bytes, and the correctness of the signature is verified to obtain a secure and trustworthy deployment package. Then, based on the synthetic test dataset and the secure and trustworthy deployment package, the smart contract is deployed on the distributed data processing traceability platform. A contract deployment transaction is executed (the gas limit is set to 10,000,000 units, and the gas price is set to 20 Gwei), and the transaction confirmation is awaited (the number of confirmations is set to 12 blocks). The validity of the deployed contract address is verified, and the contract initialization function is executed (the deployment initial parameters are set, such as the data source address, the target system address, and the permission control list) to obtain an automated verification engine. Subsequently, based on the data processing hash chain and the automated verification engine, the verification instructions of the U shield security chip are loaded. A verification command sequence is written through the U shield APDU instruction set (ISO / IEC 7816 standard). The instructions include SELECT (select the application, with an AID length of 16 bytes), VERIFY (verify the user, with a verification data length of 8 bytes), COMPUTE (perform hash calculation, supporting SHA-256 / 384 / 512 algorithms, with a data block size of 4KB), COMPARE (perform comparison, with a comparison speed not lower than 1GB / minute). The U shield read / write speed is 48MB / second, and the parallel processing thread limit is 4, to obtain a hardware-accelerated verification unit;Next, based on the hardware-accelerated verification unit, data comparison and analysis between the source system and the target system are carried out. The comparison and analysis adopt a three-stage process: First, perform metadata structure comparison (field type consistency comparison, field length comparison, field constraint comparison). Then, perform data content hash comparison (generate hash values in batches of 1000 records and compare the consistency of hash values). Finally, perform statistical feature comparison (calculate the statistical distribution differences of each field, and set the difference threshold to 1%). Record all difference items (the difference record format is JSON, including difference fields, difference types, difference values, and location information), and generate a risk calculation difference report. Then, based on the synthetic test dataset and the data isolation verification environment, configure the risk calculation graph visualization server, deploy the NVIDIA Tesla T4 GPU acceleration card (16GB video memory, 2560 CUDA cores), configure the graphics rendering parameters (resolution set to 4K, refresh rate 60Hz, color depth 32 bits), deploy the Neo4j graph database engine (version 4.4.0, index type is B-tree, cache size set to 32GB), set the graph layout algorithm to force-directed layout (spring coefficient k = 0.1, gravity coefficient g = 0.05, number of iterations 1000 times), and configure the risk calculation graph visualization server. Next, based on the synthetic test dataset, input data to the risk calculation graph visualization server, control the data import rate within 10GB / hour, perform data regularization processing (normalize the attribute values to the interval [0,1]), construct the graph data structure (nodes are business entities, edges are business relationships, attributes are business characteristics), calculate the node centrality metrics (degree centrality, betweenness centrality, closeness centrality), perform the community discovery algorithm (Louvain algorithm, modularity threshold set to 0.6), and obtain the risk data preprocessing results. Subsequently, based on the risk data preprocessing results and the data isolation verification environment, perform U shield instruction interaction, transmit the rendering instructions through the U shield secure channel (using the SCP03 secure channel protocol). The instructions include graph layout update instructions (update frequency 1 time / second), node attribute calculation instructions (calculation time not exceeding 100 milliseconds / node), and risk level marking instructions (the risk level is divided into 5 levels and marked with different colors), perform GPU-accelerated graphics rendering (using the OpenGL 4.6 API), and obtain the hardware-accelerated rendering unit. Finally, based on the hardware-accelerated rendering unit, construct the risk data relationship topology, control the number of nodes within 10000, control the number of edges within 100000, perform the force-directed layout algorithm to layout the nodes (avoid the node overlap rate being controlled below 5%), calculate the critical path (using the Dijkstra algorithm, edge weight is the risk value), mark the risk nodes (high-risk nodes are red, medium-risk nodes are yellow, low-risk nodes are green), draw the risk propagation path (the arrow points to the risk propagation direction), annotate the risk value (accurate to two decimal places), and complete the generation and visualization of the risk calculation directed acyclic graph.

[0043] Step S6: Based on the risk calculation difference report and the risk calculation directed acyclic graph, operate a preset difference positioning processor, configure an incremental compensation processing device, and generate an automated correction workflow; execute the automated correction workflow to generate a system consistency verification result.

[0044] In the embodiment of the present invention, when the U shield identity authentication is authorized based on the risk calculation difference report, the system starts the RSA-2048-bit asymmetric encryption algorithm to perform three-factor authentication. After the user inserts the U shield physical device, enters an 8-digit digital PIN code and completes fingerprint recognition, the system compares the recognition result with the template stored in the secure area. When the comparison similarity reaches more than 98%, a 256-bit high-level management permission authentication token is generated. Then, the system loads the node relationship matrix (dimension is n×n, n is the number of business data processing nodes) in the risk calculation directed acyclic graph into the preset difference positioning processor. This processor uses the graph theory depth-first search algorithm to traverse all node paths and calculate the weight value of each path (weight value calculation formula = ∑(r i ×d i ), r i is the risk value of node i, d i is the depth coefficient of node i), and generates a 200×200 dimension difference correlation analysis matrix; Subsequently, the ARM security coprocessor built in the U shield executes the SHA-256 hash algorithm at a clock frequency of 1024MHz to perform root cause analysis on the difference correlation analysis matrix. By calculating the normalized centrality of each node (value range 0-1, calculation formula = k i / (n-1), k iis the connection degree of node i, and n is the total number of nodes) to determine the root cause of the difference, and form a root cause analysis report of the difference including the severity score of the difference (1-10 points, 10 points being the most serious); the system configures incremental compensation processing equipment according to the difference correlation analysis matrix, sets the data correction threshold to 0.05 (that is, when the data deviation exceeds 5%, the correction is triggered), and the compensation strategy adopts the two-phase commit protocol to ensure transaction consistency. The correction operation is decomposed into a preparation phase and a commit phase, and all correction rules are stored as a JSON-format compensation strategy configuration file; subsequently, the correction rules in the root cause analysis report of the difference are written into the EEPROM security memory (with a capacity of 256KB) built into the USB key. The memory uses AES-256-bit encryption and sets HMAC signature verification to form a tamper-proof correction rule library; the system uses the directed acyclic graph topological sorting algorithm (with a complexity of O(V+E), where V is the number of nodes and E is the number of edges) based on the compensation strategy configuration file and the tamper-proof correction rule library to arrange the execution order of the workflow, ensuring that dependent tasks are executed first. Each task node is set with a 15-second timeout limit and a 3-time retry mechanism to generate an automated correction workflow in the BPMN2.0 standard format; finally, the system calls the advanced management permission authentication token to execute the tasks in the automated correction workflow with 5 parallel threads. Data snapshots are taken before and after each correction operation (the snapshot interval is 200 milliseconds). By calculating the difference rate of the data hash values before and after the correction (the allowable error range is 0.001%) and recording the execution time (in milliseconds), resource consumption (CPU usage rate, memory occupancy), success rate (percentage), etc. of each operation, a system consistency verification result is formed. This result includes the complete execution log of all correction operations and the final data consistency metric value (the value range is 0-100%, and it is required to reach more than 99.99% to be judged as verification successful).

[0045] The present invention constructs a highly reliable, automated, and traceable data testing system by introducing alliance blockchain and differential privacy technology, which is significantly better than the traditional data verification method that relies on centralized databases. First, the isolation of business data and test environment is realized to ensure that the real business system is not affected during the test process, and the test security is improved; secondly, with the help of the blockchain's immutability and timestamp mechanism, the authenticity and verifiability of the entire data processing process are ensured, and the shortcomings of the existing methods in data sharding encryption and hash chain consistency verification are solved; thirdly, the differential privacy protection mechanism is introduced to effectively shield sensitive information when generating test data, and improve test compliance and data security; in addition, automated verification and difference analysis are realized by combining synthetic test data with smart contracts, reducing manual intervention and improving test efficiency and accuracy; the risk calculation directed acyclic graph constructed provides a clear visualization of risk causal paths, which helps to accurately locate the source of differences; finally, through the automated correction workflow linked by difference positioning and incremental compensation mechanism, an integrated closed-loop processing from difference discovery to system repair is realized, ensuring that the system finally meets the consistency verification requirements, and significantly improving the intelligence and automation level of data quality management.

[0046] Preferably, step S1 comprises the following steps:

[0047] Step S11: Obtaining the original service data stream;

[0048] Step S12: deploying a server rack array based on the original business data flow to obtain a shadow processing sandbox physical infrastructure;

[0049] Step S13: Perform core component service orchestration based on the shadow processing sandbox physical infrastructure to obtain a microservice cluster;

[0050] Step S14: Optimize the parameters of the data normalization engine based on the original business data stream to obtain a standardized data conversion module;

[0051] Step S15: deploying a security isolation mechanism based on the standardized data conversion module to obtain a data isolation control unit;

[0052] Step S16: Based on the standardized data conversion module and the data isolation control unit, the regulatory reporting shadow processing sandbox is integrated and deployed to obtain a data isolation verification environment.

[0053] In the embodiments of the present invention, first, a data collection gateway is used to connect to the core business system database at a bandwidth of 10 Gbps (including the Oracle 19c financial core business database, the DB2 v11.5 customer information system, and the MySQL 8.0 transaction log database). The ETL task is executed every 15 minutes in an incremental extraction manner to obtain the original business data stream. The extracted data includes structured data such as transaction log tables (with 53 fields), customer profile tables (with 128 fields), and account information tables (with 76 fields), and the data volume is controlled to not exceed 500 MB per batch. Then, based on the obtained original business data stream, 4 Dell PowerEdge R940xa servers are deployed (each configured with 4 Intel Xeon Gold 6248R processors, 768 GB of DDR4 memory, and a 30 TB NVMe storage array) to form a server rack array. The servers are interconnected through a 40 Gbps InfiniBand network, and the VMware ESXi 7.0 virtualization platform is configured to achieve resource isolation. 32 virtual CPU cores and 256 GB of memory are allocated specifically for the shadow processing sandbox to run, and a physical network isolation area is set (VLAN ID is set to 4096) to build the physical infrastructure of the shadow processing sandbox. Subsequently, based on the physical infrastructure of the shadow processing sandbox, a Kubernetes v1.24 cluster is deployed for core component service orchestration. 5 master nodes and 12 worker nodes are configured, and Pod resource limits are set (the maximum CPU usage rate is 75%, and the memory limit is 192 GB). 30 microservice containers are deployed (including data access services, transaction parsing services, rule verification services, report generation services, etc.), and each service is configured with 3 replica instances to ensure 99.999% high availability. Service mesh architecture is used for communication between containers, and the request timeout threshold is limited to 500 milliseconds to obtain a complete microservice cluster. Then, based on the original business data stream, data normalization processing is performed. The Apache Spark 3.2 data processing engine is deployed (the cluster is configured with 16 executors, and each executor is allocated 8 GB of memory). Data cleaning rules are written (including field filling rate checks, data type conversions, duplicate record removal, outlier detection, etc.), field mapping conversion rules are set (the mapping table from source field names to target field names contains a total of 257 field mapping relationships), a normalization conversion task is executed (the conversion rate reaches 5000 records / second), and the data normalization engine parameters are modified (including the batch size set to 10000 records and the memory occupancy ratio set to 0.6. Perform performance tuning by setting the garbage collection policy to G1GC to obtain a standardized data conversion module; implement a three-layer security isolation mechanism based on the standardized data conversion module. First, deploy a physical isolation network gateway device (unidirectional data flow control) to ensure unidirectional data flow. Subsequently, configure firewall policies (restrict access to 34 necessary IP ports). Finally, implement data-level access control (adopt a role-based access control matrix, define 4 types of roles and 18 operation permissions), and at the same time implement operation audit tracking (record all data access operations, and the audit log retention period is 180 days) to obtain a data isolation control unit with complete access control capabilities. Finally, integrate the standardized data conversion module with the data isolation control unit, deploy a dedicated processing environment for regulatory reporting, implement a data processing pipeline (the single-pipeline throughput reaches 3,000 records per second), configure a real-time monitoring dashboard (monitor 15 key indicators, and the refresh frequency is 5 seconds per time), set the exception alarm threshold (trigger an alarm when the CPU usage rate exceeds 85%, the memory usage rate exceeds 80%, and the response time exceeds 2 seconds), complete the integration and deployment of the regulatory reporting shadow processing sandbox, and finally obtain a complete, independent and fully functional data isolation verification environment.

[0054] The present invention realizes a data isolation verification environment with high controllability, high security and high modularity by constructing a shadow processing sandbox and performing componentized service orchestration. First, through the deployment of the physical infrastructure driven by the original business data, it ensures the physical isolation of the verification environment from the real business system, and guarantees business continuity and system stability during the testing process; second, through the core component orchestration of the microservice cluster, it constructs a test service system with elastic scaling and rapid deployment capabilities, improving the system response speed and resource utilization efficiency; third, through the parameter tuning of the data normalization engine, it realizes the format unification and preprocessing standardization of multi-source heterogeneous data, enhancing the compatibility and adaptation capabilities of the test system for complex data streams; in addition, combined with the deployed security isolation mechanism, it further strengthens the security protection of data access control and processing processes, preventing the leakage or misuse of sensitive data; finally, integrate and deploy the standardized module and the isolation control mechanism to construct a complete data isolation verification environment, providing a solid underlying support for subsequent synthetic data testing, risk assessment and consistency verification, and effectively improving the maintainability, security and scalability of the entire test platform.

[0055] Preferably, step S2 includes the following steps:

[0056] Step S21: Allocate the hardware resources of the high-performance computing server cluster based on the original business data stream to obtain an array of physical nodes of the consortium blockchain;

[0057] Step S22: Configure the consensus mechanism engine parameters based on the array of physical nodes of the consortium blockchain to obtain a high-throughput consensus network;

[0058] Step S23: Deploy a blockchain server cluster based on a high-throughput consensus network to obtain the core services of the consortium blockchain;

[0059] Step S24: Install a hardware accelerator for the cryptographic security component to obtain a high-performance signature verification unit;

[0060] Step S25: Build a secure channel for the inter-node communication protocol based on the high-performance signature verification unit to obtain an encrypted data transmission network;

[0061] Step S26: Deploy a distributed data processing and traceability platform based on the core services of the consortium blockchain and the encrypted data transmission network.

[0062] In the embodiments of the present invention, first, the required hardware resources are quantitatively calculated based on the processing requirements of the original business data stream. 8 high-performance IBM Power System S924 servers are configured according to the peak processing volume of 7,500 transaction records per second (each configured with 12 cores of POWER9 processors, a frequency of 3.8 GHz, 512 GB of DDR4 memory, and 12 TB of NVMe storage). A high-speed network interconnection architecture is built through 10 Cisco Nexus 9336C-FX2 switches (each configured with 36 100 GbE ports), the network latency is controlled within 10 microseconds, the RedHat Enterprise Linux 8.6 operating system is deployed and the real-time kernel patch (RT-PREEMPT) is enabled, dedicated computing resources are allocated to each node (the number of CPU cores is 10, the memory capacity is 384 GB, and the storage space is 8 TB) and NUMA affinity binding is set to form a consortium blockchain physical node array with physical isolation characteristics; then, a Practical Byzantine Fault Tolerance (PBFT) consensus mechanism engine is configured based on the consortium blockchain physical node array, and the consensus parameters are set including the block generation time interval of 2 seconds, the upper limit of the size of a single block of 16 MB, and the maximum number of transactions processed per second of 3,500. The number of fault-tolerant nodes is configured as f=(n-1) / 3 (where n is the total number of nodes 8, so the number of fault-tolerant nodes is 2), a time synchronization service is deployed (using the PTP protocol, the time error is controlled within 50 nanoseconds), the size of the network communication buffer is adjusted to 64 MB and the zero-copy technology is enabled, and the size of the transaction verification thread pool is set to the number of cores × 2, finally achieving a high-throughput consensus network with a TPS (transactions per second) of 3,200; subsequently, Hyperledger Fabric v2. is deployed based on the high-throughput consensus network.4 blockchain server clusters, configure 7 organizations (corresponding to different business departments), 28 endorsement nodes (4 for each organization), and 8 ordering nodes, set channel configuration parameters (including a block cut-off timeout of 200 milliseconds, a maximum message count of 500, and a preferred maximum byte count of 2MB), deploy a chaincode (smart contract) execution environment and set the chaincode execution timeout limit to 30 seconds, configure a distributed ledger storage strategy (using CouchDB as the state database and setting the batch write size to 1000 records), implement an event subscription mechanism (with a maximum event buffer size of 8192 events), to obtain the core services of the consortium blockchain; then install a cryptographic accelerator based on dedicated hardware (Intel QuickAssist Technology acceleration card, each card providing a throughput of 100Gbps), integrate an HSM (Hardware Security Module) hardware security module to protect private keys (FIPS 140-2 Level 4 certification), configure ECC (elliptic curve cryptography) parameters for the P-384 curve to achieve digital signatures, deploy an NIST-certified random number generator (DRBG SP 800-90A), integrate a TEE (trusted execution environment) to achieve secure computing isolation, and achieve a signature verification performance of 8000 operations per second through 16 parallel processing units, to obtain a high-performance signature verification unit; build a TLS 1.3 secure channel based on the high-performance signature verification unit, configure the cipher suite as TLS_AES_256_GCM_SHA384, implement a two-way authentication mechanism (based on X.509 v3 certificates, with a key length of 4096 bits), deploy a certificate transparency log monitoring service, set the secure channel session timeout to 10 minutes, configure OCSP (Online Certificate Status Protocol) response binding to ensure certificate validity checking, implement link layer encryption (using the 256-bit AES-GCM encryption algorithm and rotating the session key every 24 hours), deploy a security audit log system that records all communication events, to obtain an encrypted data transmission network with high throughput and low latency characteristics (average latency less than 5 milliseconds); finally, deploy a distributed data processing traceability platform based on the consortium blockchain core services and the encrypted data transmission network, integrate a data ingestion module (with a processing rate of 5000 records per second), configure a data processing pipeline (including processing steps such as data verification, business rule checking, data transformation, and transaction packaging), implement a distributed transaction management mechanism (using the two-phase commit protocol with a timeout set to 5 seconds), deploy a distributed ledger query API (limiting the size of the single query result set to 1000 records and the query timeout to 3 seconds), configure a traceability index service (implemented based on a B+ tree, with an index node size of 4KB), set the system monitoring and alarm thresholds (trigger an alarm when the blockchain network latency exceeds 500 milliseconds, the transaction rejection rate exceeds 0.1%, or the node CPU usage exceeds 80%), to complete all the deployment work of the distributed data processing traceability platform.

[0063] The present invention constructs a high-performance consortium blockchain node architecture and a secure communication mechanism to create a distributed data processing platform with high throughput, strong security, and traceability capabilities. First, by means of a high-performance computing server cluster, the resource scheduling of the consortium blockchain physical nodes is realized, ensuring that the system has excellent computing power and concurrent processing capabilities when facing large-scale business data processing; secondly, through the configuration optimization of the high-throughput consensus mechanism, the efficiency of on-chain transaction confirmation is effectively improved, solving the performance bottleneck existing in traditional blockchains in high-frequency trading or high-concurrency scenarios; the formed blockchain server cluster deploys a stable and reliable core service system, enhancing the overall availability and scalability of the system; further, by integrating a cryptographic hardware accelerator, the signature verification efficiency is significantly improved, and the consensus verification delay is effectively reduced; the encrypted data transmission network constructs a high-security-level transmission channel for node-to-node communication, preventing data from being tampered with or stolen during transmission; finally, through the integrated deployment of the above-mentioned hardware and protocol layers, the completed distributed data processing and traceability platform can realize verifiable tracking of business data throughout the process and all nodes, comprehensively improving the system's capabilities in data compliance processing, chain auditing, and security supervision.

[0064] Preferably, step S3 includes the following steps:

[0065] Step S31: Perform format conversion preprocessing on the original business data stream to obtain a standardized data packet; load an encryption component for the standardized data packet to obtain an encrypted protection data set;

[0066] Step S32: Perform sharding processing on the encrypted protection data set based on the consortium blockchain server, and configure the hash processor parameters based on the sharding processing results to obtain a hash generation engine;

[0067] Step S33: Synchronize the timestamps of the standardized data packets, and perform multiple hash calculations on the timestamp synchronization results based on the hash generation engine to obtain the original proof of data processing;

[0068] Step S34: Aggregate, consensus verify, and compress and optimize the original proof of data processing to generate a compact proof structure, and construct an anti-tampering storage link through chain encapsulation and distributed backup to generate a data processing hash chain.

[0069] In the embodiments of the present invention, the format conversion processor performs data format unification operations on the original service data stream, converts data in different formats such as JSON, XML, CSV, etc. into the standard JSON format encoded in UTF-8 to form a standardized data packet. Subsequently, the system calls the AES-256 encryption algorithm module to perform full-field encryption processing on the standardized data packet. Each field uses an independent encryption key, and the key length is fixed at 256 bits to generate an encrypted protection data set. The consortium blockchain server calls the data sharding processor to shard the encrypted protection data set according to a size of 256KB for each shard, and each shard is appended with a 32-byte check code. Then, based on the sharding processing results, a hash processor configured with the SHA-256 hash algorithm is set. The buffer size of the processor is set to 4MB, and the single processing capacity is 10 million hashes per second to form a hash generation engine. A timestamp issued by a timestamp server conforming to the RFC3161 standard with a millisecond-level precision is added to each data record in the standardized data packet, and the timestamp synchronization result is input into the hash generation engine. The hash generation engine uses a parallel computing method of the double SHA-256 algorithm and the SHA-3 algorithm to generate a multi-hash value with a length of 512 bits, which constitutes the original proof of data processing. The proof aggregation processor aggregates the original proof of data processing according to the Merkle tree structure, with the tree depth set to 20 layers, and each node contains at most 64 child nodes to generate a proof aggregation tree. The distributed data processing traceability platform uses the PBFT consensus algorithm to perform consensus verification on the proof aggregation tree, requiring at least 75% of the nodes to reach an agreement. The verified result is written into the SSD high-speed chain memory at a write speed of not less than 3GB / second to obtain persistent proof data. The proof compression processor operating the ZK-SNARK algorithm with a compression ratio of 10:1 processes the persistent proof data to obtain a compact proof structure. Subsequently, a 4+2 check mode monitoring is performed through the Reed-Solomon redundancy check device with a redundancy ratio of 30% to obtain a reliability-enhanced proof. Based on this, a proof index generator with a B+ tree structure is called, with an index depth of 5 layers, to establish a proof index table. The proof index table is processed by the RSA-4096 algorithm for chained encapsulation, with an encapsulation block size of 1MB, and the encapsulation result is written into the distributed hash chain backup device. The backup device adopts a 5-node off-site redundant architecture with a write latency of no more than 200 milliseconds to form a distributed backup hash chain. Through the hash chain consistency verification by comparing the SHA-256 fingerprints of the distributed backup hash chain, with a verification period of 5 minutes and a consistency requirement of 100% matching, the final data processing hash chain is obtained.

[0070] The present invention constructs an efficient, secure, and verifiable data processing proof system by introducing multi-layer processing mechanisms such as standardized preprocessing, encryption protection, hash calculation, and chained encapsulation. First, through the standardization and encryption operations on the original business data stream, the unification of data structures and the security protection of content are achieved, laying a foundation for subsequent processing. Subsequently, combined with the blockchain server, the encrypted data is fragmented and the hash parameters are optimized, making the hash generation engine more targeted and high-performance, significantly improving the adaptation ability of multi-source data processing and the hash calculation efficiency. Based on the timestamp synchronization and multi-hash technology, the original data processing proof is generated, ensuring the integrity and verifiability of the data processing process, effectively preventing the risks of tampering and backtracking. Further, through aggregation, consensus verification, and compression optimization, a large number of original proof structures are transformed into a compact data structure, greatly reducing the storage burden. Finally, with the help of chained encapsulation and distributed backup mechanisms, a tamper-proof and highly available data processing hash chain is constructed, not only improving the data auditing ability of the system, but also enhancing the traceability and evidence effectiveness in regulatory compliance scenarios.

[0071] Particularly importantly, step S34 includes the following steps:

[0072] Step S341: Process the original data processing proof through a proof aggregation processor to obtain a proof aggregation tree;

[0073] Step S342: Based on the distributed data processing traceability platform, conduct consensus verification on the proof aggregation tree and write the consensus verification result into a high-speed chained memory to obtain persistent proof data;

[0074] Step S343: Operate a proof compression processor based on the persistent proof data to obtain a compact proof structure; monitor the compact proof structure through a redundancy check device to obtain a reliability-enhanced proof; call a proof index generator based on the reliability-enhanced proof to obtain a proof index table;

[0075] Step S344: Conduct chained encapsulation processing on the proof index table and write it into a hash chain backup device to obtain a distributed backup hash chain;

[0076] Step S345: Conduct hash chain consistency verification on the distributed backup hash chain to obtain the final data processing hash chain.

[0077] In the embodiments of the present invention, it is proved that the aggregation processor organizes the data processing original proof according to the Merkle tree algorithm, sets the tree height to 20 layers, and each parent node is associated with at most 64 child nodes. The aggregation processor uses GPU acceleration to build the proof aggregation tree at a speed of processing 1 million hash values per second. Each node contains a 256-bit hash value and a 32-bit timestamp, and the hash value of the root node of the proof aggregation tree is used as the unique identifier of the entire batch of data. The distributed data processing traceability platform calls the PBFT consensus algorithm to perform consensus verification on the proof aggregation tree, sets the consensus threshold to 75% of the total number of nodes, and the consensus timeout to 5 seconds. After successful verification, the consensus result is written into the high-speed chained memory. The memory is configured with an NVMe SSD array, the writing speed is 3GB / second, and the data redundancy is 2.0 to form persistent proof data. Activate the proof compression processor of the ZK-SNARK algorithm, set the compression parameters to make the compression ratio reach 10:1, process the persistent proof data to obtain a compact proof structure. The size of the compact proof structure does not exceed 10% of the original data and retains 100% verification ability. Subsequently, through the redundancy check device of Reed-Solomon coding, a 4+2 check mode is monitored, the check block size is set to 4KB, and the redundancy ratio is 30%. The check result meets the reliability requirement of 99.9999% to obtain a reliability-enhanced proof. Then, based on this proof, call the proof index generator using the B+ tree structure. The depth of the index tree is 5 layers, each node contains at most 256 index entries, and the index key length is 32 bytes to build a proof index table. The proof index table is processed by chain encapsulation using the RSA-4096 asymmetric encryption algorithm. The size of each encapsulation block is 1MB, and the inter-block is chained by using the hash value of the previous block as the encryption seed of the next block. The encapsulated data is written into the hash chain backup device with a 5-node remote deployment architecture. The geographical distance between nodes is not less than 300 kilometers, and the data synchronization delay is controlled within 200 milliseconds. All nodes adopt the full-data replication strategy to form a distributed backup hash chain. The system calculates the data fingerprint of each node in the distributed backup hash chain through the SHA-256 hash algorithm. The fingerprint length is 256 bits, and the data fingerprints of all nodes are compared within a fixed 5-minute verification cycle. The verification standard is set to 100% exact match, and any mismatch will trigger the automatic data repair mechanism. The repair uses the majority voting method to determine the correct data version, and the verified data constitutes the final data processing hash chain.

[0078] Through multi-level proof structure optimization and chained encapsulation processing, the present invention constructs an efficient, reliable and traceable data processing proof mechanism. First, through the aggregation of original data processing proofs, a structured proof aggregation tree is formed, realizing the efficient merging and logical organization of original scattered proofs, and improving the verification efficiency; the consensus verification completed in the distributed traceability platform ensures the global credibility of the aggregation result, and is persisted through a high-speed chained memory, enhancing the data writing ability and anti-loss ability of the system under high concurrency; then, through compression processing and redundancy check, a reliable proof with a compact structure and enhanced verification ability is constructed, greatly reducing the storage and transmission overhead, and at the same time improving the data integrity detection ability; the further generated proof index table facilitates subsequent efficient query and rapid positioning of relevant proof data, supporting large-scale concurrent retrieval requirements; finally, through chained encapsulation and writing to the hash chain backup device, a tamper-proof distributed backup mechanism is realized, combined with the hash chain consistency verification, ensuring that the finally generated data processing hash chain meets high standards in terms of integrity, consistency and audit availability, providing a solid technical support for business data testing and compliance auditing.

[0079] Preferably, configuring the differential privacy protection processor parameters based on the data isolation verification environment in step S4 includes:

[0080] Extracting the data distribution feature set from the data isolation verification environment;

[0081] Calibrating the differential privacy sensitivity analyzer based on the data distribution feature set to generate a sensitivity threshold matrix;

[0082] Calculating the noise injection ratio for the sensitivity threshold matrix to obtain a privacy budget configuration table;

[0083] Loading the differential privacy protection processor parameters based on the privacy budget configuration table to obtain an initialized processor instance;

[0084] Performing a privacy utility balance test on the initialized processor instance to obtain an optimized processor configuration;

[0085] Assembling a data generation engine based on the optimized processor configuration to obtain a prototype of a privacy protection test data generator;

[0086] Performing a privacy protection strength verification device test on the prototype of the privacy protection test data generator to obtain a generator compliance report;

[0087] Performing parameter fine-tuning on the prototype of the privacy protection test data generator based on the generator compliance report to obtain a privacy protection test data generator.

[0088] In the embodiments of the present invention, first, perform a feature extraction operation on the original business data in the data isolation verification environment. Through the data analysis engine, perform statistical analysis on 257 business fields, calculate the numerical distribution of each field (including minimum value, maximum value, average value, median, standard deviation, quantile), frequency distribution (frequency of occurrence of Top-K values, K is set to 10), and correlation matrix (calculated using the Pearson correlation coefficient, threshold set to ±0.7). Apply corresponding feature extraction algorithms for different data types (Z-score normalization for numerical types, One-hot encoding for categorical types, Fourier transform to extract frequency features for time series), and finally form a data distribution feature set containing 1024 feature dimensions; Subsequently, calibrate the differential privacy sensitivity analyzer based on the data distribution feature set, calculate the global sensitivity value △f for each feature dimension (defined as the maximum difference between the query functions f of any two adjacent data sets D and D', i.e., △f = max|f(D) - f(D')|). Use the interval clamping method to limit the sensitivity of continuous features within the range of [μ - 3σ, μ + 3σ] (μ is the mean, σ is the standard deviation), calculate the influence range of feature changes for discrete features through frequency analysis, and use an exponential decay function (decay coefficient α = 0.85) to adjust the sensitivity as the influence range increases. Organize the sensitivity values of all features into a 257×4 sensitivity threshold matrix (the 4 dimensions are maximum sensitivity, average sensitivity, weighted sensitivity, and combined sensitivity); Then calculate the noise injection ratio for the sensitivity threshold matrix, use the Laplace mechanism to determine the noise scale b = △f / ε (where ε is the privacy budget parameter), set different privacy budgets for different business scenarios (ε = 0.1 for core transaction data, ε = 0.5 for customer information data, ε = 1.0 for statistical summary data), and calculate the final noise injection ratio in combination with the business importance weight of each feature (determined by business analysis, range from 0 to 1), generating a privacy budget configuration table containing the noise parameters required for each query operation; Load the differential privacy protection processor parameters based on the privacy budget configuration table. The processor adopts a GPU acceleration architecture (4 NVIDIA A100 GPUs, each with 40GB of video memory), the parallel processing ability is 10000 queries per second, set the memory cache size to 128GB, limit the setting range of the privacy parameter ε within [0.01, 2.0], and use a high-precision random number generator for noise generation (the entropy source is a quantum random number generator, entropy value > 7.5 bits / byte). Pre-compile and optimize the processing function for each type of query operation (including count statistics, sum calculation, mean calculation, median calculation, quantile calculation) to obtain an initialized processor instance; Subsequently, perform a privacy-utility balance test on the initialized processor instance. The test set contains 1000 typical query scenarios, and design a test matrix to cover different privacy budgets ε (from 0.01 to 2.0, step size of 0.01) Calculate the utility loss rate (defined as |original result - privacy processed result| / |original result|) and privacy protection strength (quantified using KL divergence) for each configuration with different query complexities (from single-field queries to 25-field combined queries). Optimize the parameter configuration through the binary search algorithm (find the optimal parameter combination with a utility loss rate < 15% and meeting privacy requirements), and apply the parameter adaptive adjustment mechanism (dynamically adjust the privacy budget allocation according to query frequency and data sensitivity) to obtain the optimized processor configuration. Assemble a data generation engine based on the optimized processor configuration. The engine architecture includes a data reading module (throughput 500MB / second), a feature extraction module (supporting 257 field types), a privacy parameter control module (implementing budget management and consumption recording), a noise generation module (supporting Laplace distribution and Gaussian distribution), and a data synthesis module (synthesis rate 1000 records / second). Integrate the system with an HDFS storage backend (capacity set to 50TB) and a Spark distributed computing framework (configured with 64 executors, each allocated 4 cores and 16GB of memory) to obtain a prototype of a privacy protection test data generator. Perform a privacy protection strength verification test on the prototype of the privacy protection test data generator. The test methods include differential attack tests (observing output changes after inserting / deleting a single record), link attack tests (attempting to associate synthetic data with external data sources), and reconstruction attack tests (attempting to reconstruct the original data from synthetic data). The test results require that the privacy leakage risk is lower than 0.001 (i.e., the attack success rate < 0.1%), the generator utilization rate (defined as the proportion of synthetic data retaining useful information of the original data) is greater than 85%, and the comprehensive score (privacy × 0.6 + utilization rate × 0.4) is greater than 0.8. Organize all test metrics into a generator compliance report. Finally, perform parameter fine-tuning on the prototype of the privacy protection test data generator based on the generator compliance report. Adjust the parameters of the noise distribution function (the standard deviation is adjusted to 0.95 times the original value), update the budget allocation strategy (increase the budget by 10% for high-frequency queries), optimize the cache management algorithm (LRU replacement strategy, cache hit rate increased to 92%), and adjust the memory usage configuration (limit the working set size within 75% of the total memory). After fine-tuning, the system improves the data utilization rate by 3.5 percentage points and reduces the query response time by 25% while maintaining the same privacy guarantee, and completes the final configuration of the privacy protection test data generator.

[0089] Through the dynamic configuration and iterative optimization of differential privacy parameters, the present invention constructs a privacy protection test data generation system with high adaptability, high security, and strong compliance. First, by extracting the data distribution characteristics in the data isolation verification environment, the configuration of differential privacy parameters is made more suitable for the characteristics of real business data, improving the pertinence and effectiveness of privacy protection strategies; combining sensitivity analysis and privacy budget calculation ensures the mathematical rationality of privacy protection and the controllability of the protection strength; through privacy utility balance testing, the contradiction between data availability and privacy protection is effectively reconciled, ensuring that the generated data is representative and has application value in testing; the constructed data generation engine has strong configurability and flexible response, facilitating integration into diverse testing scenarios; in the link of generator compliance verification and parameter fine-tuning, it effectively ensures that the generator output complies with relevant data protection regulations and security specifications, reducing the risk of data leakage caused by improper synthetic data; the finally obtained privacy protection test data generator can not only generate synthetic data with statistical authenticity of high quality, but also take into account regulatory compliance and system security, providing a robust and secure input basis for business data testing and algorithm verification.

[0090] Particularly importantly, training the privacy protection test data generator in step S4 to generate a synthetic test data set includes:

[0091] Extracting training samples from the original business data stream based on the data isolation verification environment to obtain a training data set;

[0092] Inputting the training data set into the privacy protection test data generator to obtain a generator model during training;

[0093] Performing data distribution consistency detection based on the generator model during training to obtain a model performance evaluation report;

[0094] Analyzing the model performance evaluation report with a parameter tuning device to obtain an optimized generator model;

[0095] Invoking a high-performance data synthesis engine based on the optimized generator model to obtain a synthetic test data set.

[0096] In the embodiment of the present invention, first, the system performs a training sample extraction operation on the original business data stream based on the data isolation verification environment, extracts 20% of the records from the original data through the random stratified sampling algorithm while keeping the data distribution characteristics unchanged, and obtains a training data set with a training data capacity of 500MB. Then, the system inputs the training data set into the privacy protection test data generator through the TLS1.3 secure channel. The generator adopts the Gaussian distribution noise injection mechanism and sets the noise scale parameters ε = 0.1 and δ = 10^-5, and performs 10,000 iterations of training in the 512GB memory space to obtain the generator model during training. Next, the system uses the Kolmogorov-Smirnov test method to detect the data distribution consistency of the generator model during training, sets the confidence level α = 0.05, calculates the KS statistic of the numerical fields and the chi-square value of the categorical fields, and generates a model performance evaluation report containing 95 feature distribution comparison results. Subsequently, based on the 12 distribution deviation points identified in the model performance evaluation report, the system calls the parameter tuning device for analysis. The device adopts the Bayesian optimization algorithm, sets the learning rate to 0.001, the batch size to 256, and performs 50 rounds of exploration and search. The weight parameters are adjusted for each deviation point to obtain an optimized generator model with the deviation mean reduced to less than 0.03. Finally, the system calls the high-performance data synthesis engine configured with 32-core CPUs and 128GB of RAM based on the optimized generator model. The engine sets the batch size to 1,000 records, performs 500 batch generation operations, applies the forward propagation algorithm and the probability sampling method, and generates a synthetic test data set containing 500,000 records, retaining the original data structure and with a total size of 2.5GB within 30 minutes, while ensuring that the statistical characteristic difference between the generated data and the original data is controlled within 5% and does not contain any explicit identifiers that can be directly mapped to the original records.

[0097] The present invention realizes the generation of a synthetic test data set with high quality and strong privacy protection capabilities by constructing a closed-loop generator training and evaluation optimization process. First, by extracting training samples in the data isolation verification environment, it ensures that the training data is representative and isolated from the real production system, effectively preventing the leakage of original data. Second, by training the privacy protection test data generator and combining data distribution consistency detection means, it realizes a high degree of fitting of the model with the original data in terms of data structure, statistical features, etc., ensuring that the generated data has real business characteristics. The model performance evaluation and parameter tuning process improve the generalization ability and generation quality of the generator, making the synthetic data maintain stability and availability in multi-scenario tests. The finally generated synthetic test data set neither leaks the original sensitive information nor retains the real business rules of the data, providing a high-quality and low-risk synthetic data basis for subsequent applications such as testing, verification, and risk analysis, and significantly improving the security, effectiveness, and compliance of data testing.

[0098] Preferably, step S5 includes the following steps:

[0099] Step S51: Initialize the identity authentication of the U shield physical encryption module to obtain a hardware security authenticator;

[0100] Step S52: Perform digital signature on the smart contract code based on the hardware security authenticator to obtain a secure and trusted deployment package;

[0101] Step S53: Deploy the smart contract to the distributed data processing traceability platform based on the synthetic test data set and the secure and trusted deployment package to obtain an automated verification engine;

[0102] Step S54: Load the verification instruction of the U shield security chip based on the data processing hash chain and the automated verification engine to obtain a hardware-accelerated verification unit;

[0103] Step S55: Perform data comparison and analysis between the source system and the target system based on the hardware-accelerated verification unit to generate a risk calculation difference report;

[0104] Step S56: Configure the risk calculation graph visualization server based on the synthetic test data set and the data isolation verification environment;

[0105] Step S57: Input data to the risk calculation graph visualization server based on the synthetic test data set to obtain a risk data preprocessing result;

[0106] Step S58: Perform U shield instruction interaction based on the risk data preprocessing result and the data isolation verification environment to obtain a hardware-accelerated rendering unit;

[0107] Step S59: Construct a risk data relationship topology based on the hardware-accelerated rendering unit to obtain a risk calculation directed acyclic graph.

[0108] In an embodiment of the present invention, the system is connected to a USB key physical encryption module, and the identity authentication process is initialized through a three-way handshake protocol. First, an 8-digit PIN code needs to be entered during the authentication process. If the PIN code is entered incorrectly more than 3 times, the USB key will be locked. After successful authentication, the system calls the built-in RSA-4096 key pair generator of the USB key to create a temporary session key. The validity period of the session key is 30 minutes. The system uses this key to establish a 256-bit encrypted communication channel with the USB key. After the channel is established, the system obtains a hardware security authenticator, which contains a 512-bit digital certificate and a 128-bit session token; the smart contract code to be deployed is submitted to the hardware security authenticator. The contract code is written in Solidity language, and the code length is controlled within 10KB. The system calls the built-in ECDSA-P256 algorithm of the USB key through the hardware security authenticator to digitally sign the smart contract code. The signing process is completed in the physically isolated environment of the USB key. The length of the signature result is 64 bytes. The system packs the signature result and the smart contract code together into a secure and trusted deployment package, and the deployment package is encapsulated in PKCS#7 format; based on the synthetic test dataset, deployment preparation is carried out for the distributed data processing and tracing platform. The fuel limit for contract deployment is set to 5 million units. The secure and trusted deployment package is submitted to the contract deployment service of the distributed data processing and tracing platform through the HTTPS protocol. After the service verifies the digital signature in the deployment package, it executes the bytecode compilation of the smart contract. The size of the compiled code does not exceed 15KB. Subsequently, the compiled bytecode is deployed to all nodes of the blockchain network. The deployment uses a two-phase commit protocol to ensure consistency. After the deployment is completed, the system obtains an automated verification engine instance, which contains the contract address and ABI interface definition; the verification path and hash digest are extracted from the data processing hash chain. The path depth does not exceed 20 layers, and the digest length is fixed at 256 bits. These information are formatted into APDU instructions that conform to the USB key security chip interface specification. The instruction length is controlled within 256 bytes. For ultra-long instructions, fragment transmission is used. The verification instructions are loaded into the USB key security chip at a rate of 480Mbps through the USB 3.0 interface. After the chip receives the instructions, it starts the built-in SHA-256 hash verification circuit. The working frequency of this circuit is 1.2GHz, and the verification throughput reaches 5 million hash operations per second. After the verification is completed, the system obtains a hardware-accelerated verification unit; based on the hardware-accelerated verification unit, the data comparison and analysis process between the source system and the target system is started. The standardized business data records are extracted from the source system, and the total number of records is controlled within 100,000 per batch. At the same time, the corresponding processed result data is extracted from the target system. The checksum of both sides of the data is calculated using the same 256-bit SHA-3 algorithm. The system performs a field-by-field exact comparison through the hardware-accelerated verification unit. The comparison calculates the difference using the bitwise exclusive OR method, and the comparison results of each field are summarized and statistically analyzed. The difference exceeds the threshold 0.01% of the fields are marked as risk points. The system sorts all risk points by severity and generates a structured risk calculation difference report, which includes the total difference, difference distribution, and typical difference samples. Based on the synthetic test dataset and the data isolation verification environment, configure the risk calculation graph visualization server. The server is configured with a 32-core CPU, 256GB of memory, and 4TB of NVMe storage. A distributed graph database is deployed on the server, and the database is configured as a 5-node cluster, with each node having a processing capacity of 6000 TPS, and the overall throughput of the cluster reaching 30000 TPS. Import the synthetic test dataset into the risk calculation graph visualization server. The import process uses the batch parallel loading method, with the batch size set to 1000 records and the parallelism set to 16. Clean and standardize the imported data. The processing includes type conversion, null value filling, and outlier filtering. The processing parameters are set to uniformly use UTF-8 encoding, the date and time format is unified to the ISO-8601 standard, and the numerical precision is unified to 4 decimal places. After the processing is completed, the system obtains the risk data preprocessing result. Connect to the U shield physical device through an encrypted channel based on the risk data preprocessing result and the data isolation verification environment. The connection uses the TLS 1.3 protocol to ensure security. Send a graphics rendering acceleration instruction to the U shield. The instruction includes a data structure description and rendering parameters. After receiving the instruction, the U shield activates the built-in graphics processing unit, which has 1024 CUDA cores, a processing frequency of 1.5 GHz, and a memory bandwidth of 336 GB / s. The system uses this hardware resource to accelerate the graphics rendering calculation. The rendering uses the force-directed algorithm for graph layout optimization. The number of iterations is set to 500 times, and the convergence threshold is set to 0.001. After the calculation is completed, the system obtains the hardware acceleration rendering unit. Build the relationship topology between risk data based on the hardware acceleration rendering unit. The topology construction uses a hierarchical structure, with the maximum hierarchical depth set to 5 layers, and the number of nodes per layer limited to 200. The connection weights between nodes are calculated based on data correlation. The correlation is calculated using the Pearson correlation coefficient method, and the correlation coefficient threshold is set to 0.7. Nodes pairs exceeding the threshold establish connections. Organize the constructed topological structure into a directed acyclic graph form to ensure no circular dependencies. The total number of nodes in the graph is controlled within 1000, and the number of edges is controlled within 5000. Finally, generate a risk calculation directed acyclic graph.

[0109] The present invention constructs a secure, trustworthy, highly automated, and clearly visually presented business data verification system by integrating a USB key physical encryption module, secure deployment of smart contracts, high-performance verification, and visual rendering mechanisms. First, by leveraging the identity authentication mechanism of the USB key hardware, the identity trustworthiness of key system links and the physical security of instruction execution are ensured, effectively preventing illegal tampering and signature forgery. Through digital signature and trustworthy deployment of smart contracts, the source of the verification logic can be traced and the execution process can be audited. Combined with a hardware-accelerated verification unit, the processing speed and accuracy of data comparison and risk analysis are significantly improved, making it suitable for high-frequency and high-concurrency data verification scenarios. Synthetic test data and an isolated environment are introduced to achieve dual guarantees of the security and effectiveness of business data testing. In terms of risk calculation, a graph visualization server and a USB key instruction interaction mechanism are integrated, and a directed acyclic graph of risk calculation with clear structure and rigorous logic is constructed through a hardware-accelerated rendering unit, making the risk propagation path of complex data relationships more intuitively visible. Overall, not only is the automation and accuracy of data consistency verification and difference analysis improved, but also highly secure and fully controllable technical support is provided for scenarios such as regulatory auditing and compliance testing.

[0110] Preferably, step S54 includes the following steps:

[0111] Step S541: Extract the key verification path and data digest index from the data processing hash chain to generate a verification task instruction set;

[0112] Step S542: Format the verification task instruction set into an instruction stream that conforms to the USB key security chip interface protocol;

[0113] Step S543: Start the USB key communication driver in the automated verification engine, and load the verification instruction stream into the USB key security chip through the USB / Type-C interface;

[0114] Step S544: Start the hardware verification engine in the USB key security chip, complete key matching, digital signature verification, and hash chain recalculation, and generate a verification result;

[0115] Step S545: Compare the verification result with the data processing hash chain to generate a verification consistency flag;

[0116] Step S546: Report the verification consistency flag to the automated verification engine, trigger the risk difference analysis / consistency confirmation process, and form a hardware-accelerated verification unit.

[0117] In the embodiment of the present invention, a hash chain analyzer is invoked to extract the key verification path from the data processing hash chain. The analyzer traverses the hash chain structure using the depth-first search algorithm to identify the path nodes necessary for verification. At the same time, the data digest index associated with each path node is extracted. The index format is a 64-bit hexadecimal string, and each index corresponds to a specific business data processing operation. After extraction, the system organizes these paths and indexes into a structured XML format verification task instruction set. The instruction set includes the number of verification nodes, the verification path topology, and the digest index value of each node. The format conversion processor is started to convert the verification task instruction set into an APDU instruction stream that conforms to the U shield security chip interface protocol. The APDU instruction header length is fixed at 5 bytes, and the maximum length of the instruction body is 255 bytes. For ultra-long instructions, the fragmentation transmission technology is used, and the fragmentation size is 240 bytes. Each fragment is appended with 16 bytes of check information. The U shield communication driver in the automated verification engine is started. The driver uses the USB 3.1 or Type-C interface to establish a communication connection with the U shield security chip. The communication rate is set at 480 Mbps, and the transmission uses the asynchronous mode. The timeout threshold is set to 2 seconds. After successfully establishing the connection, the system loads the formatted verification instruction stream into the U shield security chip in sequence through the established communication channel. The loading process uses the block transmission and confirmation mechanism to ensure that each instruction block is correctly received. After receiving the verification instruction stream, the U shield security chip activates the built-in hardware verification engine. The operating frequency of the verification engine is 1.2 GHz. First, the RSA-2048 algorithm is executed for key matching operations to compare whether the key fingerprint in the instruction is consistent with the master key stored in the security chip. Then, the verification engine uses the ECDSA-P256 algorithm to verify the validity of the digital signature. After verification, the engine invokes the SHA-256 algorithm to recalculate the hash values of the key nodes in the hash chain. The hash calculation speed is 5 million hash operations per second. After completion of the calculation, a verification result data packet containing all the recalculated hash values is generated. The recalculated hash values in the verification result data packet are compared one by one with the original hash values stored in the data processing hash chain. The comparison algorithm uses the method of summing after bitwise exclusive OR operation. A comparison result of zero indicates a complete match, and any non-zero result indicates an inconsistency. After the comparison is completed, the system generates a binary format verification consistency flag, where 1 indicates complete consistency and 0 indicates a difference. The verification consistency flag and the detailed verification comparison data are reported to the automated verification engine through the U shield communication driver. The reporting uses an encrypted channel and the AES-256-GCM algorithm to ensure the security of data transmission. When the received consistency flag value is 1, the automated verification engine triggers the consistency confirmation process to generate a confirmation report. When the consistency flag value is 0, the risk difference analysis process is triggered to generate a difference details report. The above verification results and processing logic together constitute the hardware acceleration verification unit.

[0118] By integrating the instruction interaction and hardware acceleration verification capabilities of the USB key security chip, the present invention realizes efficient, trustworthy, and consistent verification of the data processing hash chain. By extracting the key verification path and data digest index from the hash chain, it can accurately focus on the core verification content, improving the pertinence and execution efficiency of the verification task; the standardized processing of the instruction set ensures its compatibility with the USB key hardware interface, enhancing the universality and stability of the system; combined with the USB key communication driver and interface loading mechanism, it realizes the automated and high-speed injection of verification tasks, reducing the manual participation cost and error risk; during the hardware verification process, key verification, signature verification, and hash recalculation are completed, with extremely high computing performance and security level, effectively preventing software-level attacks or tampering behaviors; through the automatic comparison mechanism between the verification result and the original hash chain, a consistency flag can be quickly generated and the result reported, further driving the difference analysis and consistency confirmation process, and constructing a complete hardware acceleration verification unit; the overall solution can significantly improve the automation, security, and response speed of the data processing verification process, and is particularly suitable for business scenarios with extremely high requirements for verification intensity, real-time performance, and non-repudiation.

[0119] Preferably, step S6 includes the following steps:

[0120] Step S61: Based on the risk calculation difference report, authorize the verification of the USB key identity authentication to obtain a high-level management authority authentication token;

[0121] Step S62: Based on the risk calculation, load the topological data of the preset difference positioning processor for the directed acyclic graph to obtain a difference correlation analysis matrix;

[0122] Step S63: Perform USB key security calculation co-processing on the preset difference positioning processor to obtain a difference root cause analysis report;

[0123] Step S64: Configure the operation parameters of the incremental compensation processing device based on the difference correlation analysis matrix to obtain a compensation strategy configuration file;

[0124] Step S65: Write the correction rules into the built-in security memory of the USB key based on the difference root cause analysis report to obtain an anti-tampering correction rule library;

[0125] Step S66: Orchestrate the correction workflow engine strategy based on the compensation strategy configuration file and the anti-tampering correction rule library to obtain an automated correction workflow;

[0126] Step S67: Based on the high-level management authority authentication token, perform task scheduling and execution on the automated correction workflow to obtain a system consistency verification result.

[0127] In an embodiment of the present invention, after receiving a risk calculation difference report, a U shield identity authentication process is triggered. The process adopts a multi-factor authentication mechanism, including physical contact verification, 8-digit PIN code verification, and fingerprint biometric verification. After all three verifications are passed, the system calls the RSA-4096 key generator built into the U shield to create a temporary management key with a validity period of 3 hours. The key length is 4096 bits. The system uses this key to establish an AES-256-GCM encryption channel with the management backend. After the channel is successfully established, the system obtains an advanced management permission authentication token, which contains a 512-bit digital signature and a 256-bit timestamp; topological data is derived from the risk calculation directed acyclic graph. The topological data contains a structured description with no more than 1000 nodes and no more than 5000 edges. The system loads the topological data into a preset difference localization processor, which is configured with a 32-core CPU, 128GB of memory, and 2TB of NVMe storage. The system performs a sparse matrix transformation on the loaded data. The matrix dimension does not exceed 1000×1000, and the matrix elements use 32-bit floating-point numbers to represent the difference weights. The system performs a singular value decomposition algorithm on the transformed matrix, sets the singular value truncation threshold to 0.01, and obtains a difference correlation analysis matrix in compressed representation; the difference correlation analysis matrix is transmitted to the U shield security calculation co-processing module. The transmission is in a block-by-block manner, with each block size of 4KB and a transmission rate of 400Mbps. After receiving the data, the U shield starts the built-in graph algorithm accelerator, which has 256 dedicated computing cores with a frequency of 1.8GHz. The system executes the PageRank algorithm and the community detection algorithm in the U shield security environment. The number of algorithm iterations is set to 200 times, and the convergence threshold is set to 0.0001. The result of the algorithm execution is signed by the RSA-2048 algorithm and then returned to the main system. The system constructs a difference root cause analysis report based on the returned result. The report contains up to 10 key root cause nodes and 50 secondary impact nodes; an incremental compensation processing device is configured based on the difference correlation analysis matrix. The device adopts a hardware acceleration architecture based on FPGA. The FPGA chip is of the Xilinx Ultrascale+ series, and the number of logic units is 1.3M, with 6840 DSP units. The system compiles the differential data and compensation strategy into FPGA firmware, and the firmware size does not exceed 20MB. The system loads the firmware into the FPGA chip and initializes the operating environment. The environment configuration includes 32-bit floating-point operation precision, 128KB cache size, and 16-channel parallel processing ability. The system generates a compensation strategy configuration file, which is in XML format and has a size of no more than 5MB, containing a complete description of the compensation rules and execution order; converts the differential root cause analysis report into a binary correction rule format, and the rule set size is controlled within 2MB. Each rule contains a 128-bit condition description and a 256-bit execution operation. The system writes the rule set into the secure memory built into the USB token through a secure channel. The memory is designed to prevent tampering, and the writing speed is 50MB / s. After writing, the system performs SHA-256 hash verification to ensure data integrity. After passing the verification, the system obtains a tamper-proof correction rule library; calls the workflow orchestration engine based on the compensation strategy configuration file and the tamper-proof correction rule library. The engine defines the workflow using the BPMN 2.0 standard. The system first constructs a task dependency graph in the form of a DAG, which contains no more than 100 task nodes, and each node is associated with a specific correction operation. The system executes a topological sorting algorithm on the graph to determine the task execution order. After sorting, the system allocates task execution resources, including the number of CPU cores, memory capacity, and storage bandwidth. The system generates an automated correction workflow containing a complete execution plan. The workflow file is in JSON format and has a size of no more than 10MB; schedules and executes the automated correction workflow based on the advanced management permission authentication token. The system first verifies the digital signature and validity period of the token. After passing the verification, the system starts a distributed task scheduler. The scheduler uses a master-slave architecture. The master node is responsible for global scheduling, and the slave nodes execute specific tasks. The communication between nodes uses a message queue based on ZeroMQ. The system decomposes the workflow into atomic tasks, and the execution time of each atomic task does not exceed 5 minutes. The system uses a transaction processing mechanism to ensure the atomicity of task execution. Data verification is performed before and after each task execution, and the CRC32 algorithm is used for verification. When the verification value does not match, it is automatically rolled back. After all tasks are executed, the system summarizes the execution results and generates a system consistency verification result containing consistency metrics and correction details. The result is presented in an HTML format report, and the report size does not exceed 20MB, including two display forms: charts and data tables.

[0128] By integrating USB key identity authentication, secure computing, differential root cause analysis, and automated correction processes, the present invention effectively constructs a highly secure, high-precision, and fully automated system consistency repair mechanism. The USB key is used to achieve secure authorization of high-level management permissions, ensuring that the initiation and scheduling of the correction process have strong identity guarantees and preventing unauthorized operations; the loading of the differential positioning and correlation analysis matrix realizes a structured understanding of complex data dependency relationships and precise modeling of differential propagation paths, improving the depth and breadth of problem identification; root cause analysis is completed through the co-processing ability of the USB key, and combined with the hardware trust environment, highly credible differential identification results are provided; the combination of compensation strategy configuration and anti-tampering correction rules not only realizes the generation of flexible and dynamic correction schemes but also guarantees the integrity and non-tamperability of the correction rules during execution; finally, the differential closed-loop correction and system consistency verification are completed through the automated workflow engine, effectively reducing manual intervention and improving the response speed and execution accuracy; this mechanism is widely applicable to application scenarios such as business data synchronization, regulatory compliance auditing, and cross-system consistency verification, providing solid technical support for constructing a trustworthy business data processing and correction system.

[0129] Preferably, the present invention further provides a business data testing system for performing the above-mentioned business data testing method. The business data testing system includes:

[0130] A data isolation sandbox deployment module for obtaining the original business data stream; deploying a regulatory reporting shadow processing sandbox based on the original business data stream to obtain a data isolation verification environment;

[0131] A blockchain traceability platform construction module for configuring a consortium blockchain server based on the original business data stream to construct a distributed data processing traceability platform;

[0132] A data processing proof generation module for generating a data processing hash chain by generating a data processing proof based on the distributed data processing traceability platform using the original business data stream;

[0133] A privacy synthetic test data generation module for configuring differential privacy protection processor parameters based on the data isolation verification environment to obtain a privacy protection test data generator; training the privacy protection test data generator to generate a synthetic test data set;

[0134] An automated difference analysis and risk map construction module for deploying an intelligent contract on the distributed data processing traceability platform based on the synthetic test data set to obtain an automated verification engine; performing data comparison analysis between the source system and the target system based on the data processing hash chain and the automated verification engine to generate a risk calculation difference report; configuring a risk calculation graph visualization server based on the synthetic test data set and the data isolation verification environment and generating a risk calculation directed acyclic graph;

[0135] A difference repair and consistency verification module is used to operate a preset difference locator based on a risk calculation difference report and a risk calculation directed acyclic graph, and configure an incremental compensation processing device to generate an automated correction workflow; execute the automated correction workflow to generate a system consistency verification result.

[0136] Preferably, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the business data testing method described above is implemented.

Claims

1. A business data testing method, characterized in that, It includes the following steps: Step S1: Obtain the original business data stream; deploy a regulatory reporting shadow processing sandbox based on the original business data stream to obtain a data isolation verification environment; Step S2: Configure a consortium blockchain server based on the original business data stream to build a distributed data processing traceability platform; Step S3: Use the original business data stream on the distributed data processing traceability platform to generate data processing proof, obtaining a data processing hash chain; Step S4: Configure differential privacy protection processor parameters based on the data isolation verification environment to obtain a privacy protection test data generator; Train the privacy protection test data generator to generate a synthetic test data set; Step S5: Deploy smart contracts on the distributed data processing traceability platform based on the synthetic test data set to obtain an automated verification engine; Perform data comparison and analysis between the source system and the target system based on the data processing hash chain and the automated verification engine to generate a risk calculation difference report; Configure a risk calculation graph visualization server based on the synthetic test data set and the data isolation verification environment, and generate a risk calculation directed acyclic graph; Step S6: Operate a preset difference positioning processor based on the risk calculation difference report and the risk calculation directed acyclic graph, and configure an incremental compensation processing device to generate an automated correction workflow; Execute the automated correction workflow to generate a system consistency verification result.

2. The service data testing method according to claim 1, wherein Step S1 includes the following steps: Step S11: Obtain the original business data stream; Step S12: Deploy a server rack array based on the original business data stream to obtain the physical infrastructure of the shadow processing sandbox; Step S13: Perform core component service orchestration based on the physical infrastructure of the shadow processing sandbox to obtain a microservice cluster; Step S14: Optimize the parameters of the data normalization engine based on the original business data stream to obtain a standardized data conversion module; Step S15: Deploy a security isolation mechanism based on the standardized data conversion module to obtain a data isolation control unit; Step S16: Integrate and deploy a regulatory reporting shadow processing sandbox based on the standardized data conversion module and the data isolation control unit to obtain a data isolation verification environment.

3. The service data testing method according to claim 1, characterized in that Step S2 includes the following steps: Step S21: Allocate the hardware resources of a high-performance computing server cluster based on the original business data stream to obtain a consortium blockchain physical node array; Step S22: Configure the consensus mechanism engine parameters based on the consortium blockchain physical node array to obtain a high-throughput consensus network; Step S23: Deploy a blockchain server cluster based on the high-throughput consensus network to obtain the core services of the consortium blockchain; Step S24: Install a hardware accelerator for cryptographic security components to obtain a high-performance signature verification unit; Step S25: Build a secure communication channel for the inter-node communication protocol based on the high-performance signature verification unit to obtain an encrypted data transmission network; Step S26: Deploy a distributed data processing traceability platform based on the core services of the consortium blockchain and the encrypted data transmission network.

4. The service data testing method according to claim 1, characterized in that Step S3 includes the following steps: Step S31: Perform format conversion preprocessing on the original business data stream to obtain a standardized data packet; load an encryption component into the standardized data packet to obtain an encrypted protection data set; Step S32: Shard the encrypted and protected data set based on the consortium blockchain server, and configure the hash processor parameters based on the sharding result to obtain a hash generation engine; Step S33: Synchronize the timestamps of the standardized data packets, and perform multiple hash calculations on the timestamp synchronization result based on the hash generation engine to obtain the original proof of data processing; Step S34: Aggregate, consensus verify, and compress and optimize the original proof of data processing to generate a compact proof structure, and build a tamper-proof storage link through chained encapsulation and distributed backup to generate a data processing hash chain.

5. The service data testing method according to claim 1, wherein The differential privacy protection processor parameters configured based on the data isolation verification environment in Step S4 include: Extract the data distribution feature set from the data isolation verification environment; Calibrate the differential privacy sensitivity analyzer based on the data distribution feature set to generate a sensitivity threshold matrix; Calculate the noise injection ratio for the sensitivity threshold matrix to obtain a privacy budget configuration table; Load the differential privacy protection processor parameters based on the privacy budget configuration table to obtain an initialized processor instance; Perform a privacy utility balance test on the initialized processor instance to obtain an optimized processor configuration; Assemble a data generation engine based on the optimized processor configuration to obtain a prototype of a privacy protection test data generator; Perform a privacy protection strength verification device test on the prototype of the privacy protection test data generator to obtain a generator compliance report; Fine-tune the parameters of the prototype of the privacy protection test data generator based on the generator compliance report to obtain a privacy protection test data generator.

6. The service data testing method according to claim 1, characterized in that Step S5 includes the following steps: Step S51: Initialize the identity authentication of the U shield physical encryption module to obtain a hardware security authenticator; Step S52: Perform a digital signature on the smart contract code based on the hardware security authenticator to obtain a secure and trusted deployment package; Step S53: Deploy the smart contract to the distributed data processing traceability platform based on the synthetic test data set and the secure and trusted deployment package to obtain an automated verification engine; Step S54: Load the verification instructions of the U shield security chip based on the data processing hash chain and the automated verification engine to obtain a hardware-accelerated verification unit; Step S55: Perform a data comparison and analysis between the source system and the target system based on the hardware-accelerated verification unit to generate a risk calculation difference report; Step S56: Configure a risk calculation graph visualization server based on the synthetic test data set and the data isolation verification environment; Step S57: Input data to the risk calculation graph visualization server based on the synthetic test data set to obtain a risk data preprocessing result; Step S58: Perform U shield instruction interaction based on the risk data preprocessing result and the data isolation verification environment to obtain a hardware-accelerated rendering unit; Step S59: Build a risk data relationship topology based on the hardware-accelerated rendering unit to obtain a risk calculation directed acyclic graph.

7. The service data testing method according to claim 6, wherein Step S54 includes the following steps: Step S541: Extract the key verification path and data digest index from the data processing hash chain to generate a verification task instruction set; Step S542: Format the verification task instruction set into an instruction stream that conforms to the U shield security chip interface protocol; Step S543: Start the U shield communication driver in the automated verification engine, and load the verification instruction stream to the U shield security chip through the USB / Type-C interface; Step S544: Start the hardware verification engine in the U shield security chip, complete key matching, digital signature verification and hash chain recalculation, and generate a verification result; Step S545: Compare the verification result with the data processing hash chain to generate a verification consistency flag; Step S546: Report the verification consistency flag to the automated verification engine, trigger the risk difference analysis / consistency confirmation process, and form a hardware acceleration verification unit.

8. The service data testing method according to claim 1, wherein Step S6 includes the following steps: Step S61: Authorize the verification of the U shield identity authentication based on the risk calculation difference report to obtain a high-level management authority authentication token; Step S62: Load the topological data of the preset difference localization processor based on the risk calculation directed acyclic graph to obtain a difference correlation analysis matrix; Step S63: Perform U shield security calculation co-processing on the preset difference localization processor to obtain a difference root cause analysis report; Step S64: Configure the operation parameters of the incremental compensation processing device based on the difference correlation analysis matrix to obtain a compensation strategy configuration file; Step S65: Write the correction rules into the built-in security memory of the U shield based on the difference root cause analysis report to obtain a tamper-proof correction rule library; Step S66: Orchestrate the correction workflow engine policy based on the compensation strategy configuration file and the tamper-proof correction rule library to obtain an automated correction workflow; Step S67: Perform task scheduling and execution on the automated correction workflow based on the high-level management authority authentication token to obtain a system consistency verification result.

9. A business data testing system, characterized in that, For executing the business data test method as described in claim 1, the business data test system includes: A data isolation sandbox deployment module, configured to obtain the original business data stream; deploy a regulatory reporting shadow processing sandbox based on the original business data stream to obtain a data isolation verification environment; A blockchain traceability platform construction module, configured to configure a consortium blockchain server based on the original business data stream and construct a distributed data processing traceability platform; A data processing proof generation module, configured to generate a data processing proof based on the distributed data processing traceability platform using the original business data stream to obtain a data processing hash chain; A privacy synthetic test data generation module, configured to configure the differential privacy protection processor parameters based on the data isolation verification environment to obtain a privacy protection test data generator; train the privacy protection test data generator to generate a synthetic test data set; An automated difference analysis and risk map construction module, configured to deploy a smart contract on the distributed data processing traceability platform based on the synthetic test data set to obtain an automated verification engine; perform data comparison and analysis between the source system and the target system based on the data processing hash chain and the automated verification engine to generate a risk calculation difference report; configure a risk calculation graph visualization server based on the synthetic test data set and the data isolation verification environment, and generate a risk calculation directed acyclic graph; A difference repair and consistency verification module is used to operate a preset difference location processor based on a risk calculation difference report and a risk calculation directed acyclic graph, and configure an incremental compensation processing device to generate an automated correction workflow; execute the automated correction workflow to generate a system consistency verification result.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the business data testing method according to any one of claims 1-8.

Citation Information

Patent Citations

  • High-throughput distributed account book system based on DAG

    CN116846674A

  • Partially-ordered blockchain

    US20210194672A1