System and method for auditing and supervising real-time traceability of multi-source data
By deploying blockchain nodes at the source of data, collecting and recording data feature information in real time, and calculating the credibility coefficient, the single point failure risk and lack of real-time performance of the multi-source data traceability method in the existing technology are solved, and real-time traceability auditing and efficient supervision of multi-source data are realized.
Patent Information
- Application Number
- CN202511115296.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-11
AI Technical Summary
The existing multi-source data traceability method relies on a centralized system, which has the risk of single point failure, making it difficult to achieve full-link trusted records. It also lacks real-time and dynamic adjustment capabilities, and is unable to accurately locate the root cause of data anomalies, resulting in inefficient data supervision.
Deploy blockchain nodes at the source of data to collect the original feature information of the data in real time and generate a unique data lineage identifier, record data flow information and calculate the credibility coefficient. Through chain association and periodic verification, locate the maximum attenuation point of data anomalies and dynamically update the attenuation coefficient comparison table.
It realizes real-time traceability audit of multi-source data, can promptly detect data quality anomalies, accurately locate the specific abnormal operational behaviors, and improve the audit efficiency and the accuracy of long-term supervision.
Smart Images

Figure CN120632744A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data governance technology, and specifically to a system and method for real-time traceability, auditing and supervision of multi-source data. Background Art
[0002] With the rapid development of information technology, data has become multi-source, heterogeneous, and frequently circulated. Data is prone to distortion, tampering, or quality anomalies during generation, transmission, and processing, posing challenges to data credibility and security. Existing data traceability methods mostly rely on centralized systems, which pose a risk of single-point failure and make it difficult to achieve full-link trusted records. Traditional auditing methods are mostly post-event static verification, lacking real-time and dynamic adjustment capabilities, and unable to accurately locate the root cause of data anomalies. At the same time, existing technologies do not adequately quantify the credibility of data during the data flow process, making it difficult to assess the degree of data quality degradation through unified standards. This results in inefficient data supervision, delayed anomaly tracing, and an inability to meet the needs of real-time supervision of multi-source data. Summary of the Invention
[0003] The purpose of the present invention is to provide a system and method for real-time traceability, auditing and supervision of multi-source data to solve the problems raised in the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a method for real-time traceability, auditing and supervision of multi-source data, the method comprising: Deploy blockchain nodes at the source of data. When data is generated, automatically collect the original feature information of the data, write it into the blockchain, generate a unique data lineage identifier, and establish a data lineage index table. When data flows, the node identification, timestamp and operation type are recorded in real time. The attenuation value is obtained from the attenuation coefficient comparison table based on the operation type and complexity. The current credibility coefficient is calculated and the flow node information, attenuation value and current credibility coefficient are added to the chain as a new node to form a chain association. Periodically verify data integrity. When data quality anomalies are detected, extract the current feature information of the abnormal data, retrieve the complete lineage chain corresponding to the abnormal data, analyze the field mapping deviation and value range deviation, and associate the attenuation value of each link with the accumulated offset; Filter links where the credibility drop is greater than the threshold and mark them as high-risk attenuation points. Combine the feature offset jump points to determine the maximum attenuation point. Retrieve the chain record of the maximum attenuation point to locate the specific operation behavior. Visually output the root cause link, data lineage chain and feature offset change curve, upload the traceability conclusion as an audit record to the chain, and dynamically update the attenuation coefficient comparison table.
[0005] According to the above scheme, the original feature information includes data field structure, field value range, generation timestamp, collection device identification and initial credibility coefficient; The original feature information is serialized according to a predefined structured format. A hash operation is performed on the serialized original feature information to generate a feature summary. This summary is then combined with the current blockchain state information to generate a data lineage identifier. The complete feature information is then written into a new blockchain block. The data lineage index table includes a data lineage identifier field, an original data storage address field, and a blockchain transaction hash field; a mapping relationship is established between the data lineage identifier and the original data physical storage address, as well as an index relationship between the data lineage identifier and the blockchain transaction record.
[0006] According to the above scheme, the initial credibility coefficient includes: The initial credibility coefficient is obtained through a multi-dimensional feature evaluation model, based on a comprehensive evaluation of data type factors, data source authentication level, data integrity check value, timestamp credibility, and data generation frequency. The formula is as follows: C0=w1×DT+w2×DA+w3×DC+w4×TR+w5×DF; Among them, C0 represents the initial credibility coefficient; w1, w2, w3, w4 and w5 represent weighting coefficients; DT represents the data type factor; DA represents the data source authentication level; DC represents the data integrity check value; TR represents the timestamp credibility; DF represents the data generation frequency; The timestamp credibility TR is obtained based on the time difference between the data generation time and the blockchain writing time. The formula is TR=e (-λ×Δt) , where λ represents the attenuation coefficient; The data type factor is divided according to the standardization of the data structure; the data source authentication level is determined by the validity of the digital certificate and the access control level of the data source device; the data integrity check value is calculated based on the SHA-3, BLAKE3 or SHA-256 algorithm to calculate the checksum matching degree; the data generation frequency is the number of times data is generated per unit time, which is used to evaluate the stability of data generation.
[0007] According to the above scheme, data flow includes the data collection, transmission, storage or sharing stages; Complexity is assessed based on the number of fields involved, the depth of the calculation logic, and the number of associated data records, and is divided into three levels: simple, medium, and complex. Simple operations include data format conversion, field renaming, and data filtering; medium operations include data aggregation, associated queries, and basic statistical analysis; and complex operations include machine learning model training, advanced data mining, multidimensional data analysis, and the application of complex algorithms. According to the operation type and complexity, the attenuation value is obtained from the attenuation coefficient comparison table to calculate the current credibility coefficient. The formula is as follows: C t =C t-1 ×(1-α); Among them, C t Expressed as the credibility coefficient of the current node; C t-1 It is represented as the credibility coefficient of the previous node, and α is represented as the attenuation value corresponding to the current operation; The attenuation coefficient comparison table analyzes historical data governance cases, calculates the average data distortion rate under different operation types, and sets the benchmark attenuation coefficient with reference to international standards; Serialize the transfer node information, attenuation value, and current credibility coefficient according to a predefined structured format, perform a hash operation on the serialized information to generate a feature summary, and combine it with the current blockchain state information to generate a new data transfer node identifier. Write the complete transfer node information, attenuation value, current credibility coefficient, feature summary, and data transfer node identifier into a new block of the blockchain, forming a chain association with the previous node. The flow node information includes the physical address, logical address, device ID, system ID, user ID, operation IP address, operation port number, operation protocol type, operation command, operation parameters, operation result status code and operation result description of the operation node.
[0008] According to the above scheme, the chain association adopts the Merkle tree structure, and each data flow node contains a pointer to its parent node. The complete data flow path can be obtained by tracing back step by step; at the same time, a data flow index table is established in the blockchain. The data flow index table includes a data lineage identifier field, a data flow node identification field, a flow timestamp field, an operation type field, an attenuation value field, and a credibility coefficient field, and establishes a mapping relationship from the data lineage identifier to the data flow node, as well as a temporal association relationship between the data flow nodes.
[0009] According to the above scheme, data integrity includes data field structure, field value range, data relationship constraints and data integrity constraints; When data quality anomalies are detected, the exception handling process is automatically triggered to extract the current feature information of the abnormal data. The current feature information includes data field structure, field value range, current timestamp, current storage node identifier and current credibility coefficient; Compare the current feature information of the abnormal data with the original feature information to analyze the field mapping deviation and value range deviation. The field mapping deviation is obtained by calculating the deviation between the current field structure and the original field structure. The deviation includes the difference in the number of fields, the difference in field names, the difference in field types, the difference in field order, and the difference in field constraints. The value range deviation is obtained by calculating the deviation between the current field value range and the original field value range. The deviation includes the upper limit offset of the value range, the lower limit offset of the value range, the value range distribution form offset, and the value range dispersion offset. The field mapping deviation and value range deviation are correlated with the attenuation value of each circulation link, and the cumulative offset of each link is calculated. The cumulative offset includes the cumulative field mapping deviation, the cumulative value of value range deviation and the cumulative value of credibility attenuation.
[0010] According to the above scheme, a time series analysis algorithm is used to perform time series correlation analysis on the attenuation value, characteristic offset, and cumulative offset of each link, identifying the changing trend, mutation point, and periodic fluctuation pattern of the characteristic offset. By comparing the credibility drop of each link with the preset threshold, links with a credibility drop greater than the threshold are screened out and marked as high-risk attenuation points. The credibility drop threshold is established by collecting records of credibility coefficient changes at each stage of the operation process, distinguishing the natural credibility drop range of normal data flow and the credibility drop range of abnormal data. Statistical analysis methods are used to determine the critical value between normal and abnormal attenuation, and the critical value is used as the initial credibility drop threshold. Combined with the position and amplitude of the feature offset jump point, the attenuation point with the greatest impact on data quality is determined. The feature offset jump point is obtained by calculating the rate of change of the feature offset of adjacent links. The point where the rate of change exceeds the preset jump threshold is marked as a feature offset jump point. The jump threshold is determined by extracting the rate of change of characteristic offsets in each link in historical data, calculating the distribution range of the rate of change under normal operation and the peak value of the rate of change corresponding to abnormal operation; and determining the jump threshold through ROC curve analysis. Retrieve the on-chain record of the maximum attenuation point. The on-chain record includes the physical address, logical address, device identification, system identification, user identification, operation IP address, operation port number, operation protocol type, operation command, operation parameters, operation result status code and operation result description of the operation node; by analyzing the on-chain record, locate the specific operation behavior that causes data quality abnormality.
[0011] According to the above scheme, the field mapping deviation degree is as follows: D m =(n1+n2+n3+n4+n5) / N; Among them, D mIt is expressed as the field mapping deviation; n1 is the field quantity difference value, that is, the absolute value of the difference between the current field quantity and the original field quantity; n2 is the field name difference value, that is, the number of fields with mismatched names; n3 is the field type difference value, that is, the number of fields with mismatched types; n4 is the field order difference value, that is, the number of fields with inconsistent order with the original; n5 is the field constraint difference value, that is, the number of fields with mismatched constraints; N is the total number of original fields; The value range deviation is as follows: D r =(|U c -U o |+|L c -L o |+|F c -F o |+|S c -S o |) / (U o +L o +F o +S o ); Among them, D r Expressed as the deviation of the value range; U c Indicates the upper limit of the current field value range; U o Indicates the upper limit of the original field value range; L c Indicates the lower limit of the current field value range; L o Indicates the lower limit of the original field value range; F c Represents the distribution morphological parameters of the current field value range; F o Expressed as the original field value range distribution morphological parameter; S c It is expressed as the discrete degree parameter of the current field value range; S o Expressed as the discrete degree parameter of the original field value range; The characteristic offset change rate is as follows: R=(F t -F t-1 ) / F t-1 ; Among them, R represents the rate of change of characteristic offset; F t Expressed as the characteristic offset of the current link; F t-1 Expressed as the feature offset of the previous link; The cumulative offset is as follows: C m =ΣD mi ; C r =ΣD ri ; C a =Σα i ; Among them, Cm Expressed as the cumulative amount of field mapping deviation; D mi It is expressed as the field mapping deviation of the i-th link; C r Expressed as the cumulative deviation from the value range; D ri Expressed as the value range deviation of the i-th link; C a Expressed as the cumulative amount of credibility decay; α i Expressed as the attenuation value of the i-th link.
[0012] According to the above scheme, the traceability conclusion includes abnormal data identification, node information of the root cause link, feature offset analysis results, credibility change trends, operation behavior positioning results and rectification suggestions; the traceability conclusion is serialized according to a predefined structured format, and a hash operation is performed on the serialized traceability conclusion to generate a conclusion summary. This is combined with the current blockchain state information to generate an audit record identifier. The complete traceability conclusion, conclusion summary and audit record identifier are then written to a new blockchain block; Dynamically updating the attenuation coefficient comparison table includes calculating the actual attenuation impact coefficient of each type of operation based on the operation type, feature offset and credibility change data in the traceability conclusion, comparing the corresponding values in the original attenuation coefficient comparison table, and updating the attenuation value of the corresponding operation type and complexity combination when the deviation exceeds the preset adjustment threshold; adjusting the threshold by calculating the deviation distribution between the actual attenuation impact coefficient in the historical traceability conclusion and the value in the original attenuation coefficient comparison table to determine the acceptable error range; combining the weight of the attenuation coefficient for credibility assessment, setting the graded adjustment standard, and obtaining the adjustment threshold; serializing and hashing the updated attenuation coefficient comparison table to generate a comparison table summary, which is written into the blockchain together with the update timestamp and operation node identifier.
[0013] A real-time traceability audit and supervision system for multi-source data, comprising: a data acquisition module, a data quality verification module, a credibility assessment module, a traceability audit module and a dynamic optimization module; The data collection module includes the original data collection module and the flow information collection module; the original data collection module automatically collects original feature information, performs serialization and hash operations to generate feature summaries, and generates unique data lineage identifiers; the flow information collection module collects flow information in real time, generates new node identifiers, and uploads them to the chain; The data quality verification module includes an integrity verification module and an anomaly analysis module. The integrity verification module verifies the integrity of the data according to a preset period and triggers the anomaly handling process when an anomaly is detected. The anomaly analysis module extracts the current feature information of the abnormal data, retrieves the corresponding complete lineage chain, and locates the specific operation behavior that caused the anomaly. Credibility assessment module, used for initial assessment and dynamic calculation of data credibility coefficient, and quantifying changes in credibility during data flow; The traceability audit module includes a visualization module and a record-on-chain module. The visualization module outputs the root cause link, data lineage chain, and feature offset change curve through a graphical interface, intuitively presenting the data anomaly traceability results. The record-on-chain module serializes the traceability conclusion and generates a hash summary, which is combined with the blockchain status information to generate an audit record identifier and then uploaded to the chain. The dynamic optimization module dynamically optimizes the attenuation coefficient comparison table based on the traceability results to improve the accuracy of credibility assessment.
[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention records data flow information in real time and calculates the credibility coefficient, combined with periodic integrity verification, to achieve timely discovery and rapid response to data quality anomalies; 2. This invention quantifies the field mapping deviation, value range deviation, and feature offset change rate, combined with credibility decay analysis, to accurately locate the maximum decay point and specific operation behavior of data anomalies, thereby improving audit efficiency; 3. The present invention dynamically updates the attenuation coefficient comparison table based on the traceability conclusion, so that the credibility assessment model is continuously optimized with the actual application scenario, thereby improving the accuracy and adaptability of long-term data supervision. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a structural diagram of a multi-source data real-time traceability audit and supervision system according to the present invention; Figure 2 A flowchart of the steps for establishing a data lineage index table for a method for real-time traceability, auditing, and supervising multi-source data according to the present invention; Figure 3 This is a flowchart of the steps of data flow analysis for a method for real-time traceability, auditing and supervision of multi-source data according to the present invention; Figure 4 The present invention is a flowchart of the steps of locating abnormal data in a method for real-time traceability, auditing and supervision of multi-source data. DETAILED DESCRIPTION
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0017] Example: Figure 1-Figure 4 As shown, the present invention provides a technical solution, a method for real-time traceability auditing and supervision of multi-source data, the method comprising the steps of: S1. Deploy blockchain nodes at the source of data. When data is generated, automatically collect the original feature information of the data, write it into the blockchain, generate a unique data lineage identifier, and establish a data lineage index table. Specifically, the original feature information includes the data field structure, field value range, generation timestamp, collection device identification and initial credibility coefficient; the original feature information is serialized according to a predefined structured format, and the serialized original feature information is hashed to generate a feature summary, which is combined with the current blockchain status information to generate a data lineage identifier, and the complete feature information is written to the new block of the blockchain; the data lineage index table includes the data lineage identifier field, the original data storage address field and the blockchain transaction hash field; a mapping relationship between the data lineage identifier and the physical storage address of the original data is established, as well as an index relationship between the data lineage identifier and the blockchain transaction record.
[0018] For example, deploying blockchain nodes at the source of data generation to collect and store raw data features. When data is generated, the following information is automatically collected: Data field structure: includes field A, field B, field C and field D; Field value range: field A∈[1,100], field C∈[0,1000]; Generation timestamp: 2025-01-01 00:00:00; Collection device identification: source server device number; Furthermore, the initial credibility coefficient includes: The initial credibility coefficient is obtained through a multi-dimensional feature evaluation model, based on a comprehensive evaluation of data type factors, data source authentication level, data integrity check value, timestamp credibility, and data generation frequency. The formula is as follows: C0=w1×DT+w2×DA+w3×DC+w4×TR+w5×DF; Among them, C0 represents the initial credibility coefficient; w1, w2, w3, w4 and w5 represent weighting coefficients; DT represents the data type factor; DA represents the data source authentication level; DC represents the data integrity check value; TR represents the timestamp credibility; DF represents the data generation frequency; The timestamp credibility TR is obtained based on the time difference between the data generation time and the blockchain writing time. The formula is TR=e (-λ×Δt) , where λ represents the attenuation coefficient; The data type factor is divided according to the standardization of the data structure; the data source authentication level is determined by the validity of the digital certificate and the access control level of the data source device; the data integrity check value is calculated based on the SHA-3, BLAKE3 or SHA-256 algorithm to calculate the checksum matching degree; the data generation frequency is the number of times data is generated per unit time, which is used to evaluate the stability of data generation.
[0019] For example: Data type factor (DT): structured data, DT=0.9; Data source authentication level (DA): the device holds a secondary digital certificate, access control is internal level, DA=0.8; Data integrity check value (DC): hash value calculated using SHA-256, matching degree 100%, DC=1.0; Timestamp reliability (TR): the time difference between data generation and blockchain writing Δt=1 second, λ=0.1, TR=TR=e (-λ×Δt) =0.9048; data generation frequency (DF): 200 data items per hour, medium stability, DF=0.7; weighting coefficients w1, w2, w3, w4 and w5 are 0.2, 0.2, 0.3, 0.1 and 0.2 respectively, and it is calculated that: C0=0.2×0.9+0.2×0.8+0.3×1.0+0.1×0.9048+0.2×0.7=0.8505, and the initial credibility coefficient is 0.8505.
[0020] S2. When data flows, the node identifier, timestamp, and operation type are recorded in real time. Based on the operation type and complexity, the attenuation value is obtained from the attenuation coefficient comparison table, and the current credibility coefficient is calculated. The flow node information, attenuation value, and current credibility coefficient are added to the chain as a new node to form a chain association. Specifically, data flow includes data collection, transmission, storage or sharing; Complexity is assessed based on the number of fields involved in the operation, the depth of the calculation logic, and the number of associated data records, and is divided into three levels: simple, medium, and complex. Simple operations include data format conversion, field renaming, and data filtering; medium operations include data aggregation, associated queries, and basic statistical analysis; and complex operations include machine learning model training, advanced data mining, multidimensional data analysis, and complex algorithm applications. For example, operation type: filtering, simple complexity, involving 2 fields, and associated with 50 records; According to the operation type and complexity, the attenuation value is obtained from the attenuation coefficient comparison table to calculate the current credibility coefficient. The formula is as follows: C t =C t-1 ×(1-α); Among them, C t Expressed as the credibility coefficient of the current node; C t-1 It is represented as the credibility coefficient of the previous node, and α is represented as the attenuation value corresponding to the current operation; The attenuation coefficient comparison table analyzes historical data governance cases, calculates the average data distortion rate under different operation types, and sets the benchmark attenuation coefficient with reference to international standards; Serialize the transfer node information, attenuation value, and current credibility coefficient according to a predefined structured format, perform a hash operation on the serialized information to generate a feature summary, and combine it with the current blockchain state information to generate a new data transfer node identifier. Write the complete transfer node information, attenuation value, current credibility coefficient, feature summary, and data transfer node identifier into a new block of the blockchain, forming a chain association with the previous node. The transfer node information includes the physical address, logical address, device ID, system ID, user ID, operation IP address, operation port number, operation protocol type, operation command, operation parameters, operation result status code and operation result description of the operation node; For example: Attenuation coefficient comparison table: the attenuation value of simple complexity filtering operation is α=0.03; the credibility coefficient is C t =C0×(1-α)=0.8505×0.97=0.8250; on-chain information: Node ID: processing server number, timestamp: 2025-01-01 00:05:00, operation command: filter records with field A < 10, decay value 0.03, credibility 0.8250, generate a new node and form a chain with the source node; In another possible embodiment, when data flows, the node identifier, timestamp and operation type are recorded in real time. The operation type is aggregation, medium complexity, merging the average value of field A and field C; the decay value α=0.06, the corresponding value of the medium complexity aggregation operation; the credibility coefficient is C t =C0×(1-α)=0.8250×(1-0.06)=0.7755; on-chain information: node identification: processing server number, timestamp: 2025-01-01 00:10:00, operation parameters: aggregation period: 5 minutes, forming a chain association. This is only an example and is not a limitation.
[0021] Furthermore, the chain association adopts a Merkle tree structure, and each data flow node contains a pointer to its parent node. The complete data flow path can be obtained by tracing back step by step; at the same time, a data flow index table is established in the blockchain. The data flow index table includes a data lineage identifier field, a data flow node identification field, a flow timestamp field, an operation type field, an attenuation value field, and a credibility coefficient field, and establishes a mapping relationship from the data lineage identifier to the data flow node, as well as a temporal association relationship between the data flow nodes.
[0022] S3. Periodically verify data integrity. When data quality anomalies are detected, extract the current feature information of the abnormal data, retrieve the complete lineage chain corresponding to the abnormal data, analyze the field mapping deviation and value range deviation, and associate the attenuation value of each link with the accumulated offset; Specifically, data integrity includes data field structure, field value range, data relationship constraints and data integrity constraints; when data quality abnormalities are detected, the abnormality handling process is automatically triggered to extract the current feature information of the abnormal data, which includes data field structure, field value range, current timestamp, current storage node identifier and current credibility coefficient; the current feature information of the abnormal data is compared with the original feature information, and the field mapping deviation and value range deviation are analyzed; the field mapping deviation is obtained by calculating the deviation between the current field structure and the original field structure, and the deviation includes the difference in field quantity, field name, field type, field order and field constraint condition; the value range deviation is obtained by calculating the deviation between the current field value range and the original field value range, and the deviation includes the upper limit offset of the value range, the lower limit offset of the value range, the value range distribution form offset and the value range dispersion offset; the field mapping deviation and value range deviation are correlated and analyzed with the attenuation value of each flow link, and the offset accumulation of each link is calculated, and the offset accumulation includes the field mapping deviation accumulation, the value range deviation accumulation and the credibility attenuation accumulation.
[0023] Furthermore, the field mapping deviation is calculated as follows: D m =(n1+n2+n3+n4+n5) / N; Among them, D m It is expressed as the field mapping deviation; n1 is the field quantity difference value, that is, the absolute value of the difference between the current field quantity and the original field quantity; n2 is the field name difference value, that is, the number of fields with mismatched names; n3 is the field type difference value, that is, the number of fields with mismatched types; n4 is the field order difference value, that is, the number of fields with inconsistent order with the original; n5 is the field constraint difference value, that is, the number of fields with mismatched constraints; N is the total number of original fields; The value range deviation is as follows: D r =(|U c -U o |+|L c -L o |+|F c -F o |+|S c -S o |) / (U o +L o +F o+S o ); Among them, D r Expressed as the deviation of the value range; U c Indicates the upper limit of the current field value range; U o Indicates the upper limit of the original field value range; L c Indicates the lower limit of the current field value range; L o Indicates the lower limit of the original field value range; F c Represents the distribution morphological parameters of the current field value range; F o Expressed as the original field value range distribution morphological parameter; S c It is expressed as the discrete degree parameter of the current field value range; S o Expressed as the discrete degree parameter of the original field value range; For example: Periodic verification: Data integrity is verified every hour, using the BLAKE3 algorithm to compare the hash values of each node's data; Anomaly detection: During the verification at 2025-01-01 10:00:00, it was found that the value range of field C of the secondary processing node was abnormal. The current field C value ∈ [0, 1500] deviated from the original range [0, 1000], triggering the exception handling process; Extract current feature information: field value range: field C∈[0,1500]; current timestamp: 2025-01-01 10:00:00; current credibility coefficient: 0.7755; storage node identifier: processing server number; For example: retrieve the complete lineage chain: trace the full chain record of the source through the data lineage identifier; calculate the field mapping deviation D m :The original number of fields N=4, the current number of fields is the same (n1=0); the field names and types are the same (n2=n3=0); the field order is different: Field C and Field D are swapped (n4=1); the field constraints are different: the constraint of Field C is changed to Field C≤1500, the original is Field C≤1000, n5=1; D m =(0+0+0+1+1) / 4=0.5; Calculated value range deviation D r :Upper limit of value range U o =1000, U c =1500; lower limit of value range L o =0 and L c =0 (no deviation); distribution shape parameter F o =2 (normal distribution), F c =2 (no bias); dispersion parameter S o =200, S c =300; Get the value range deviation D r=(500+0+0+100) / (1000+0+2+200)=600 / 1202≈0.4992.
[0024] Offset accumulation: Secondary processing link: C m =0.5, C r =0.4992, C a =0.03+0.06=0.09.
[0025] S4. Filter links where the credibility drop is greater than the threshold and mark them as high-risk attenuation points. Combined with the feature offset jump point, determine the maximum attenuation point. Retrieve the chain record of the maximum attenuation point to locate the specific operation behavior. Specifically, a time series analysis algorithm is used to perform time series correlation analysis on the attenuation value, characteristic offset, and cumulative offset of each link to identify the changing trend, mutation point, and periodic fluctuation pattern of the characteristic offset. By comparing the credibility drop of each link with the preset threshold, links with a credibility drop greater than the threshold are screened out and marked as high-risk attenuation points. The credibility drop threshold is established by collecting records of credibility coefficient changes at each stage of the operation process, distinguishing the natural credibility drop range of normal data flow and the credibility drop range of abnormal data. Statistical analysis methods are used to determine the critical value between normal and abnormal attenuation, and the critical value is used as the initial credibility drop threshold. Combined with the position and amplitude of the feature offset jump point, the attenuation point with the greatest impact on data quality is determined. The feature offset jump point is obtained by calculating the rate of change of the feature offset of adjacent links. The point where the rate of change exceeds the preset jump threshold is marked as a feature offset jump point. The jump threshold is determined by extracting the rate of change of characteristic offsets in each link in historical data, calculating the distribution range of the rate of change under normal operation and the peak value of the rate of change corresponding to abnormal operation; and determining the jump threshold through ROC curve analysis. Retrieve the on-chain record of the maximum attenuation point. The on-chain record includes the physical address, logical address, device identification, system identification, user identification, operation IP address, operation port number, operation protocol type, operation command, operation parameters, operation result status code and operation result description of the operation node; by analyzing the on-chain record, locate the specific operation behavior that causes data quality abnormality.
[0026] Furthermore, the characteristic offset change rate is as follows: R=(F t -F t-1 ) / F t-1 ; Among them, R represents the rate of change of characteristic offset; F t Expressed as the characteristic offset of the current link; F t-1Expressed as the feature offset of the previous link; The cumulative offset is as follows: C m =ΣD mi ; C r =ΣD ri ; C a =Σα i ; Among them, C m Expressed as the cumulative amount of field mapping deviation; D mi It is expressed as the field mapping deviation of the i-th link; C r Expressed as the cumulative deviation from the value range; D ri Expressed as the value range deviation of the i-th link; C a Expressed as the cumulative amount of credibility decay; α i Expressed as the attenuation value of the i-th link.
[0027] For example: the decrease in the first processing stage is: 0.8505-0.8250=0.0255; the decrease in the second processing stage is: 0.8255-0.7755=0.0495; the preset credibility decrease threshold is 0.04, and the decrease in the second processing stage is 0.0495>0.04, which is marked as a high-risk attenuation point; The feature offset F of the first processing step t-1 =0.05 (slight deviation); secondary processing link characteristic offset F t =0.5+0.4992=0.9992; rate of change R=(0.9992-0.05) / 0.05=18.98, preset jump threshold = 4 (determined by ROC curve analysis), R>4 is marked as the feature offset jump point; Combining the high-risk attenuation points and jump points, the secondary processing link is determined to be the maximum attenuation point, and its chain record is retrieved: the operation command is to adjust the value range of field C, and the operation parameter is to manually expand the upper limit to 1500.
[0028] S5. Visually output the root cause link, data lineage chain, and feature offset change curve, upload the traceability conclusion as an audit record, and dynamically update the attenuation coefficient comparison table; Specifically, the traceability conclusion includes abnormal data identification, node information of the root cause link, feature offset analysis results, credibility change trend, operation behavior positioning results and rectification suggestions; the traceability conclusion is serialized according to a predefined structured format, and the serialized traceability conclusion is hashed to generate a conclusion summary, which is combined with the current blockchain state information to generate an audit record identification, and the complete traceability conclusion, conclusion summary and audit record identification are written into a new block of the blockchain; the dynamic update of the attenuation coefficient comparison table includes calculating the actual attenuation impact coefficient of each type of operation based on the operation type, feature offset and credibility change data in the traceability conclusion, comparing the corresponding values in the original attenuation coefficient comparison table, and updating the attenuation value of the corresponding operation type and complexity combination when the deviation exceeds the preset adjustment threshold; adjusting the threshold by calculating the deviation distribution between the actual attenuation impact coefficient in the historical traceability conclusion and the value in the original attenuation coefficient comparison table to determine the acceptable error range; combining the weight of the attenuation coefficient on the credibility assessment, setting the graded adjustment standard and obtaining the adjustment threshold; serializing and hashing the updated attenuation coefficient comparison table to generate a comparison table summary, which is written into the blockchain together with the update timestamp and operation node identification.
[0029] The present invention provides another technical solution, a real-time traceability audit supervision system for multi-source data, which includes: a data acquisition module, a data quality verification module, a credibility assessment module, a traceability audit module and a dynamic optimization module; The data collection module includes the original data collection module and the flow information collection module; the original data collection module automatically collects original feature information, performs serialization and hash operations to generate feature summaries, and generates unique data lineage identifiers; the flow information collection module collects flow information in real time, generates new node identifiers, and uploads them to the chain; The data quality verification module includes an integrity verification module and an anomaly analysis module. The integrity verification module verifies the integrity of the data according to a preset period and triggers the anomaly handling process when an anomaly is detected. The anomaly analysis module extracts the current feature information of the abnormal data, retrieves the corresponding complete lineage chain, and locates the specific operation behavior that caused the anomaly. Credibility assessment module, used for initial assessment and dynamic calculation of data credibility coefficient, and quantifying changes in credibility during data flow; The traceability audit module includes a visualization module and a record-on-chain module. The visualization module outputs the root cause link, data lineage chain, and feature offset change curve through a graphical interface, intuitively presenting the data anomaly traceability results. The record-on-chain module serializes the traceability conclusion and generates a hash summary, which is combined with the blockchain status information to generate an audit record identifier and then uploaded to the chain. The dynamic optimization module dynamically optimizes the attenuation coefficient comparison table based on the traceability results to improve the accuracy of credibility assessment.
[0030] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
Claims
1. A method for real-time traceability, auditing and supervision of multi-source data, characterized by: The method includes: Deploy blockchain nodes at the source of data. When data is generated, automatically collect the original feature information of the data, write it into the blockchain, generate a unique data lineage identifier, and establish a data lineage index table. When data flows, the node identification, timestamp and operation type are recorded in real time. The attenuation value is obtained from the attenuation coefficient comparison table based on the operation type and complexity. The current credibility coefficient is calculated and the flow node information, attenuation value and current credibility coefficient are added to the chain as a new node to form a chain association. Periodically verify data integrity. When data quality anomalies are detected, extract the current feature information of the abnormal data, retrieve the complete lineage chain corresponding to the abnormal data, analyze the field mapping deviation and value range deviation, and associate the attenuation value of each link with the accumulated offset; Filter links where the credibility drop is greater than the threshold and mark them as high-risk attenuation points. Combine the feature offset jump points to determine the maximum attenuation point. Retrieve the chain record of the maximum attenuation point to locate the specific operation behavior. Visually output the root cause link, data lineage chain and feature offset change curve, upload the traceability conclusion as an audit record to the chain, and dynamically update the attenuation coefficient comparison table.
2. A method for real-time traceability, auditing and supervision of multi-source data according to claim 1, characterized in that: The original feature information includes data field structure, field value range, generation timestamp, collection device identification and initial credibility coefficient; The original feature information is serialized according to a predefined structured format, a hash operation is performed on the serialized original feature information to generate a feature digest, and the digest is combined with the current blockchain state information to generate a data lineage identifier, and the complete feature information is written into a new blockchain block; The data lineage index table includes a data lineage identifier field, an original data storage address field, and a blockchain transaction hash field; a mapping relationship between the data lineage identifier and the original data physical storage address, and an index relationship between the data lineage identifier and the blockchain transaction record are established.
3. A method for real-time traceability, auditing and supervision of multi-source data according to claim 2, characterized in that: The initial credibility coefficient includes: The initial credibility coefficient is obtained through a multi-dimensional feature evaluation model, based on a comprehensive evaluation of data type factors, data source authentication level, data integrity check value, timestamp credibility, and data generation frequency. The formula is as follows: C0=w1×DT+w2×DA+w3×DC+w4×TR+w5×DF; Among them, C0 represents the initial credibility coefficient; w1, w2, w3, w4 and w5 represent weighting coefficients; DT represents the data type factor; DA represents the data source authentication level; DC represents the data integrity check value; TR represents the timestamp credibility; DF represents the data generation frequency; The timestamp credibility is obtained based on the time difference between the data generation time and the blockchain writing time, and the formula is TR=e (-λ×Δt) , where λ represents the attenuation coefficient and Δt represents the difference between the data generation timestamp and the blockchain writing timestamp.
4. The method for real-time traceability, auditing and supervision of multi-source data according to claim 1, characterized in that: The data flow includes data collection, transmission, storage or sharing; The complexity is assessed based on the number of fields involved in the operation, the depth of the calculation logic, and the number of associated data records, and is divided into three levels: simple, medium, and complex. According to the operation type and complexity, the attenuation value is obtained from the attenuation coefficient comparison table to calculate the current credibility coefficient. The formula is as follows: C t =C t-1 ×(1-α); Among them, C t Expressed as the credibility coefficient of the current node; C t-1 It is represented as the credibility coefficient of the previous node, and α is represented as the attenuation value corresponding to the current operation; The transfer node information, attenuation value and current credibility coefficient are serialized in a predefined structured format, and a hash operation is performed on the serialized information to generate a feature summary. This is combined with the current blockchain state information to generate a new data transfer node identifier. The complete transfer node information, attenuation value, current credibility coefficient, feature summary and data transfer node identifier are written into a new block of the blockchain to form a chain association with the previous node.
5. A method for real-time traceability, auditing and supervision of multi-source data according to claim 4, characterized in that: The chain association adopts a Merkle tree structure, and each data flow node contains a pointer to its parent node. The complete data flow path can be obtained by tracing back step by step; at the same time, a data flow index table is established in the blockchain. The data flow index table includes a data lineage identifier field, a data flow node identification field, a flow timestamp field, an operation type field, an attenuation value field and a credibility coefficient field, and establishes a mapping relationship from the data lineage identifier to the data flow node, as well as a temporal association relationship between the data flow nodes.
6. The method for real-time traceability, auditing and supervision of multi-source data according to claim 1, characterized in that: The data integrity includes data field structure, field value range, data relationship constraints and data integrity constraints; When data quality anomalies are detected, the exception handling process is automatically triggered to extract the current feature information of the abnormal data, which includes the data field structure, field value range, current timestamp, current storage node identifier and current credibility coefficient; Compare the current feature information of the abnormal data with the original feature information, and analyze the field mapping deviation and value range deviation; the field mapping deviation is obtained by calculating the deviation between the current field structure and the original field structure, and the deviation includes the difference in the number of fields, the difference in field names, the difference in field types, the difference in field order, and the difference in field constraints; the value range deviation is obtained by calculating the deviation between the current field value range and the original field value range, and the deviation includes the upper limit offset of the value range, the lower limit offset of the value range, the value range distribution form offset, and the value range dispersion offset; The field mapping deviation and value range deviation are correlated with the attenuation value of each circulation link, and the offset accumulation of each link is calculated. The offset accumulation includes the field mapping deviation accumulation, the value range deviation accumulation and the credibility attenuation accumulation.
7. A method for real-time traceability, auditing and supervision of multi-source data according to claim 6, characterized in that: Using the time series analysis algorithm, we conduct time series correlation analysis on the attenuation value, characteristic offset and cumulative offset of each link to identify the changing trend, mutation point and periodic fluctuation pattern of characteristic offset; By comparing the credibility drop of each link with the preset threshold, the links with a credibility drop greater than the threshold are screened out and marked as high-risk attenuation points; Determine the attenuation point with the greatest impact on data quality based on the location and amplitude of the feature offset jump point. The feature offset jump point is obtained by calculating the rate of change of the feature offset of adjacent links. Points where the rate of change exceeds a preset jump threshold are marked as feature offset jump points. Retrieve the on-chain record of the maximum attenuation point, where the on-chain record includes the physical address, logical address, device identifier, system identifier, user identifier, operation IP address, operation port number, operation protocol type, operation command, operation parameters, operation result status code, and operation result description of the operation node; and locate the specific operation behavior that causes data quality abnormality by analyzing the on-chain record.
8. A method for real-time traceability, auditing and supervision of multi-source data according to claim 7, characterized in that: The field mapping deviation is as follows: <h2 style=";text-align:left;direction:ltr">D<h2 style=";text-align:left;direction:ltr"> m <h2 style=";text-align:left;direction:ltr"> =(n1+n2+n3+n4+n5) / N; Among them, D m It represents the field mapping deviation; n1 represents the difference in the number of fields; n2 represents the difference in the field name; n3 represents the difference in the field type; n4 represents the difference in the field order; n5 represents the difference in the field constraint; N represents the total number of original fields; The value range deviation is as follows: D r =(|U c -U o |+|L c -L o |+|F c -F o |+|S c -S o |) / (U o +L o +F o +S o ); Among them, D r Expressed as the deviation of the value range; U c Indicates the upper limit of the current field value range; U o Indicates the upper limit of the original field value range; L c Indicates the lower limit of the current field value range; L o Indicates the lower limit of the original field value range; F c Represents the distribution morphological parameters of the current field value range; F o Expressed as the original field value range distribution morphological parameter; S c It is expressed as the discrete degree parameter of the current field value range; S o Expressed as the discrete degree parameter of the original field value range; The characteristic offset change rate is as follows: R=(F t -F t-1 ) / F t-1 ; Among them, R represents the rate of change of characteristic offset; F t Expressed as the characteristic offset of the current link; F t-1 Expressed as the feature offset of the previous link; The cumulative offset is expressed as follows: C m =ΣD mi ;C r =ΣD ri ;C a =Sa i ; Among them, C m Expressed as the cumulative amount of field mapping deviation; D mi It is expressed as the field mapping deviation of the i-th link; C r Expressed as the cumulative deviation from the value range; D ri Expressed as the value range deviation of the i-th link; C a Expressed as the cumulative amount of credibility decay; α i Expressed as the attenuation value of the i-th link.
9. The method for real-time traceability, auditing and supervision of multi-source data according to claim 1, characterized in that: The traceability conclusion includes abnormal data identification, node information of the root cause link, feature offset analysis results, credibility change trends, operation behavior positioning results and rectification suggestions; the traceability conclusion is serialized according to a predefined structured format, a hash operation is performed on the serialized traceability conclusion to generate a conclusion summary, which is combined with the current blockchain state information to generate an audit record identification, and the complete traceability conclusion, conclusion summary and audit record identification are written into a new blockchain block; The dynamically updating attenuation coefficient comparison table includes calculating the actual attenuation impact coefficient of each type of operation based on the operation type, feature offset, and credibility change data in the traceability conclusion, comparing the corresponding values in the original attenuation coefficient comparison table, and updating the attenuation value of the corresponding operation type and complexity combination when the deviation exceeds a preset adjustment threshold; The updated attenuation coefficient comparison table is serialized and hashed to generate a comparison table summary, which is written to the blockchain together with the update timestamp and operation node identifier.
10. A real-time traceability audit and supervision system for multi-source data, characterized by: The system includes: data acquisition module, data quality verification module, credibility assessment module, traceability audit module and dynamic optimization module; The data acquisition module includes an original data acquisition module and a flow information acquisition module; the original data acquisition module automatically collects original feature information, performs serialization and hash operations to generate feature summaries, and generates unique data lineage identifiers; the flow information acquisition module collects flow information in real time, generates new node identifiers, and uploads them to the chain; The data quality verification module includes an integrity verification module and an anomaly analysis module; the integrity verification module performs integrity verification on the data according to a preset period and triggers the anomaly handling process when an anomaly is detected; the anomaly analysis module extracts the current feature information of the abnormal data, retrieves the corresponding complete lineage chain, and locates the specific operation behavior that caused the anomaly; The credibility assessment module is used for the initial assessment and dynamic calculation of the data credibility coefficient and quantifies the credibility changes during the data flow process; The traceability audit module includes a visualization module and a record-uploading module. The visualization module outputs the root cause link, data lineage chain, and feature offset change curve through a graphical interface, intuitively presenting the data anomaly traceability results. The record-uploading module serializes the traceability conclusion and generates a hash summary, which is combined with the blockchain status information to generate an audit record identifier and then uploaded to the chain. The dynamic optimization module dynamically optimizes the attenuation coefficient comparison table based on the traceability results to improve the accuracy of the credibility assessment.
Citation Information
Patent Citations
Block chain-based audit data evidence traceability method and terminal
CN113836233A
Positioning method and device for data exception reason tracing based on data consanguinity
CN117708103A
Data verification method and device based on block chain and electronic equipment
CN118796809A
Method for high-performance distributed storage of block data and timestamp, cross-chain communication and data collaboration
WO2023050555A1
Cited By
Unified social credit code data quality control method based on traceability technology
CN120975806A
Drug enterprise financial data full-link traceability query system and method based on block chain
CN121213268A
Multi-source index data intelligent storage management method and system
CN121412211A
Data table cleaning method and device, storage medium and electronic equipment
CN121745972A