A power engineering audit information management method and system based on data governance
By combining data dictionaries and anchor keys for engineering objects, the problems of scattered data sources and difficulty in version management in power engineering audits have been solved, achieving unified data management and traceability of the audit process, and improving the information management level of power engineering audits.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU JUNNAN TECH CO LTD
- Filing Date
- 2026-06-23
- Publication Date
- 2026-07-21
AI Technical Summary
In power engineering audits, data sources are scattered, data versions are difficult to manage uniformly, and the audit process lacks complete and traceable data support, resulting in inconsistent data definitions, poor traceability and verifiability of audit results, and a lack of systematic version management mechanisms.
Establish a data dictionary, generate unique entity identifiers and summary fingerprints, achieve unified data management through data monotonic chains and engineering object anchor keys, construct version maps and process topology diagrams, generate audit clues and perform evidence package summary processing, forming a closed-loop mechanism for auditing and data governance.
It has enabled the unified organization and management of power engineering audit data, improved the traceability and standardization of audit analysis, ensured that audit conclusions have clear data basis and verifiable paths, and improved the overall level of information management.
Smart Images

Figure CN122434473A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of audit information management technology, specifically relating to a method and system for managing power engineering audit information based on data governance. Background Technology
[0002] With the continuous expansion of power engineering construction, a large amount of data related to project cost, material usage, contract execution, and on-site management is generated during the project initiation, design, procurement, construction, and settlement stages. This data typically originates from multiple business systems and management processes, exhibits diverse data types and significant structural differences, and undergoes continuous changes during project implementation, thus introducing high data management complexity to power engineering auditing.
[0003] Current power engineering audits primarily rely on manual data processing or decentralized system analysis. Audit data is often stored separately according to the source system or business module, lacking unified data semantic constraints and structural standards. This leads to inconsistencies in data definitions between different systems and difficulty in accurately associating the same business object. Especially during the verification of project costs, material costs, and contract amounts, frequent data version changes and the easy overwriting or replacement of historical data make it difficult for auditors to fully reconstruct the cost formation process and data evolution path, affecting the traceability and reproducibility of audit results.
[0004] Furthermore, existing technologies for identifying audit data version differences largely rely on manual comparison or ad-hoc rules, lacking a systematic version management mechanism. The correspondence between audit status and data changes is unclear, and the process of generating audit clues lacks a unified basis. After problems are discovered and rectification is completed, the relationship between the rectified data and the original data is often not effectively preserved, governance rules are difficult to dynamically adjust with audit results, and it is difficult to form a closed-loop mechanism that links auditing and data governance. Summary of the Invention
[0005] This invention provides a data governance-based method and system for managing power engineering audit information, which solves the technical problems in related technologies such as the dispersed sources of power engineering audit data, the difficulty in unifying the management of data versions, and the lack of complete and traceable data support for the audit process.
[0006] This invention provides a data governance-based method for managing audit information in power engineering projects, comprising the following steps: Step 1: Establish a data dictionary and generate unique entity identifiers. Standardize the original records to obtain summary fingerprints. Write the unique entity identifier, source system, collection time, and summary fingerprint into the registration block and append them to form a data monotonic chain. Step 2: Generate anchor keys for engineering objects based on the project unique identifier, object domain, and object local identifier, and bind the entity unique identifier and the engineering object anchor keys to the engineering object associated data; Step 3: Based on the project object association data and data dictionary, obtain the total project cost, total material cost, and contract amount, generate consistent calculation results, and generate process topology convergence results based on the process topology diagram; Step 4: Construct a version graph based on the data monotonic chain to obtain version difference information, and transfer and reference the registration block and version difference information according to the audit state set and state transition rules; Step 5: Generate audit trails based on the consistency calculation results and version difference information; Step 6: Select the registration block referenced by the audit clue from the data monotonic chain, assemble the evidence package and generate the evidence package summary fingerprint. Generate the document summary fingerprint based on the structured audit text and the evidence package summary fingerprint and write it into the registration block. Step 7: Based on the audit clues, identify the rectification targets and collect the rectified data, write it into the data monotonic chain, repeat steps 3 to 6, and update the data dictionary to the new version.
[0007] This invention provides a power engineering audit information management system based on data governance, comprising: The data governance module establishes a data dictionary and generates unique entity identifiers. It standardizes the original records to obtain summary fingerprints, writes the unique entity identifier, source system, collection time, and summary fingerprint into the registration block, and appends them to form a data monotonic chain. The project object anchoring module generates project object anchor keys based on the project unique identifier, object domain, and object local identifier, and binds the entity unique identifier with the project object anchor key as project object associated data; The cost consistency aggregation module obtains the total project cost, total material cost, and contract amount based on the project object association data and data dictionary, generates consistent calculation results, and generates process topology aggregation results based on the process topology diagram; The version evolution control module constructs a version graph based on the data monotonic chain to obtain version difference information, and transfers and references the registration block and version difference information according to the audit status set and status transition rules. The audit trail generation module generates audit trails based on consistency calculation results and version difference information. The audit evidence storage module selects the registration block referenced by the audit clues from the data monotonic chain to assemble the evidence package and generate the evidence package summary fingerprint. Based on the structured audit text and the evidence package summary fingerprint, it generates the document summary fingerprint and writes it into the registration block. The rectification closed-loop evolution module identifies rectification targets based on audit clues and collects data after rectification, writes it into the data monotonic chain, aggregates duplicate expenses to the audit evidence storage module, and updates the data dictionary to the new version.
[0008] The beneficial effects of this invention are as follows: Based on the concept of data governance, this invention provides unified organization and management of multi-source data generated during power engineering audits. It can continuously record the data version evolution process without changing the original data content, providing a complete and stable data foundation for audit activities. By using a data dictionary to uniformly constrain field semantics, value rules, and structural relationships, it ensures consistency of data from different source systems under the same governance framework, reducing the impact of inconsistent audit data definitions.
[0009] This invention introduces an engineering object anchoring mechanism to establish a stable association between various business data in engineering projects and specific engineering objects. This gives clear object-oriented targeting to cost calculations, consistency analysis, and audit trail generation, improving the traceability of audit analysis. Through joint analysis of data version differences and cost discrepancies, audit trails can simultaneously reflect data changes and business calculation results, helping auditors accurately identify key areas of concern.
[0010] Furthermore, this invention incorporates the selection, assembly, and summarization of audit evidence into a unified data governance process, and updates the data and governance rules in a versioned manner after audit rectification, forming a closed-loop mechanism that links audit and data governance. This makes the information management process of power engineering audit more standardized and transparent, and ensures that audit conclusions have clear data basis and verifiable paths, thereby improving the overall information management level of power engineering audit. Attached Figure Description
[0011] Figure 1 This is a flowchart of a power engineering audit information management method based on data governance according to the present invention. Detailed Implementation
[0012] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.
[0013] like Figure 1 As shown, a data governance-based method for managing power engineering audit information includes the following steps: Step 1: Establish a data dictionary and generate unique entity identifiers. Standardize the original records to obtain summary fingerprints. Write the unique entity identifier, source system, collection time, and summary fingerprint into the registration block and append them to form a data monotonic chain. Step 2: Generate anchor keys for engineering objects based on the project unique identifier, object domain, and object local identifier, and bind the entity unique identifier and the engineering object anchor keys to the engineering object associated data; Step 3: Based on the project object association data and data dictionary, obtain the total project cost, total material cost, and contract amount, generate consistent calculation results, and generate process topology convergence results based on the process topology diagram; Step 4: Construct a version graph based on the data monotonic chain to obtain version difference information, and transfer and reference the registration block and version difference information according to the audit state set and state transition rules; Step 5: Generate audit trails based on the consistency calculation results and version difference information; Step 6: Select the registration block referenced by the audit clue from the data monotonic chain, assemble the evidence package and generate the evidence package summary fingerprint. Generate the document summary fingerprint based on the structured audit text and the evidence package summary fingerprint and write it into the registration block. Step 7: Based on the audit clues, identify the rectification targets and collect the rectified data, write it into the data monotonic chain, repeat steps 3 to 6, and update the data dictionary to the new version.
[0014] In one embodiment of the present invention, a data dictionary is established and a unique entity identifier is generated. The original records are standardized to obtain a summary fingerprint. The unique entity identifier, source system, collection time, and summary fingerprint are written into a registration block and appended to form a data monotonic chain, including: Step 11: Establish a unified data dictionary to govern and constrain multi-source heterogeneous data involved in power engineering audits. The data dictionary defines field names, data types, units, precision, value ranges, mandatory fields, and field order rules for different object domains such as projects, contracts, lists, materials, processes, equipment, personnel, invoices, and on-site evidence. It also registers the correspondence between fields from the source system and fields in the data dictionary. Through this method, the data dictionary logically forms a unified semantic definition and structural specification for various types of audit data, providing a consistent data standard for subsequent data processing. Specifically, object domains distinguish the audit scope of different business objects, and field order rules constrain the arrangement order of fields during data processing and calculation, thereby avoiding data ambiguity caused by differences in field arrangement.
[0015] Step 12: Collect original records from the source systems related to power engineering audits. These original records include at least field names, field values, and information about the source system to which each field belongs. For each collected original record, determine its object domain based on the data dictionary. Extract the values of the entity local identifier field set marked in the data dictionary from the original records. The entity local identifier field set refers to several fields that can uniquely represent a business entity, and their composition and order are predefined by the data dictionary. Based on this, concatenate the object domain with the values of the entity local identifier field set according to the field order rules specified in the data dictionary, and perform irreversible fixed-length identifier calculation on the concatenated result to generate a unique entity identifier. The irreversible fixed-length identifier calculation refers to performing summary processing on the input content according to predetermined rules to generate a fixed-length identifier value, and this identifier value cannot be used to reverse the input content. The unique entity identifier generated in this way maintains consistency across data collected from different source systems and at different times, enabling stable identification and association of the same business entity during subsequent audits.
[0016] Step 13: Perform standardization processing on the original records. This standardization process, based on a data dictionary, unifies the structure and numerical values of the fields in the original records. Specifically, it includes: rearranging fields according to field order rules; converting units based on field units and rounding values according to field precision; validating field values based on their range; and writing a pre-defined missing value representation into the data dictionary for fields that are required but missing in the original records. Through this standardization process, differences in field order, units of measurement, numerical precision, and missing value representation between different systems are eliminated, ensuring that the standardized original records maintain consistency in structure and numerical scope.
[0017] Step 14: After standardization, perform irreversible fixed-length identifier calculation on the standardized original records to obtain a summary fingerprint. The summary fingerprint characterizes the overall content features of the standardized original records; its length is fixed and the calculation result is irreversible. Subsequently, the entity unique identifier, source system, acquisition time, and summary fingerprint are written into the registration block, and the registration block is written into the data monotonic chain in an append-only manner. The data monotonic chain is a chain structure formed by appending registration blocks in chronological order. The append-only method means that only adding registration blocks to the end of the existing registration block sequence is allowed, without deleting, overwriting, replacing, or modifying the content of existing registration blocks, thereby ensuring the continuity and traceability of audit data in the time dimension.
[0018] Through the above technical solutions, this invention, in the process of power engineering audit information management, uses a data dictionary as the core governance carrier to achieve unified semantic constraints, structural specifications, and consistent processing of multi-source audit data; through the generation of unique entity identifiers and summary fingerprints, it achieves stable identification of business entities and data content; and through the construction of registration blocks and data monotonic chains, it achieves versioned recording and non-overwriteable storage of audit data, effectively solving the problems of diverse data sources, inconsistent structures, and difficulty in tracing versions in the process of power engineering auditing.
[0019] In one embodiment of the present invention, generating an anchor key for an engineering object and binding the entity's unique identifier to the anchor key for the engineering object as associated data includes: Step 21: Extract the project unique identifier, object domain, and object partial identifier from the original records based on the data dictionary, and perform standardization processing on the extraction results. The project unique identifier is used to identify the specific project scope to which the power engineering belongs; the object domain is used to distinguish the business object category to which the original record belongs, and its value is limited to one of engineering, contract, list, material, process, equipment, personnel, invoice, and on-site evidence; the object partial identifier refers to the field or combination of fields used to distinguish different business instances within the same object domain, and its field composition is predefined by the data dictionary. By performing standardization processing consistent with the previous steps, the project unique identifier, object domain, and object partial identifier are kept consistent in field order, value expression, and data scope, providing a unified input for subsequent calculations.
[0020] Step 22: After standardization, the standardized project unique identifier, standardized object domain, and standardized object local identifier are concatenated according to the field order rules specified in the data dictionary. An irreversible fixed-length identifier calculation is then performed on the concatenated result to generate the engineering object anchor key. The engineering object anchor key is an identifier value used to uniquely identify a specific business object within a given engineering project. Its calculation process is irreversible and has a fixed length, ensuring stability and consistency across data from different source systems and at different collection times. By introducing the engineering object anchor key, different business objects within the same engineering project can obtain a unified and reusable anchor identifier within the data governance system.
[0021] Step 23: Locate the registration block containing the entity's unique identifier in the data monotonic chain, and establish a binding relationship between the entity's unique identifier and the project object's anchor key to form project object association data. This project object association data describes the correspondence between entity-level data and project object-level data, ensuring that entity data written into the data monotonic chain can be accurately aggregated under the corresponding project object. In this invention, the project object association data does not modify the original registration block but is maintained as an independent association, thereby preserving the integrity and immutability of historical data in the data monotonic chain.
[0022] Through the above technical solution, this invention achieves unified object-level anchoring and cross-system association of power engineering audit data within a data governance framework. The engineering object anchor key uniformly encodes the project dimension, object domain dimension, and object local identifier dimension, solving the problem of difficult unified aggregation of data from multiple systems and multiple objects. By binding the entity's unique identifier with the engineering object anchor key, a stable mapping relationship is established from entity-level data to engineering object-level data, effectively improving the standardization, relevance, and traceability of data governance, and supporting consistent data management and analysis throughout the audit process.
[0023] In one embodiment of the present invention, the total project cost, total material cost, and contract amount are obtained based on the project object association data and data dictionary, a consistent calculation result is generated, and a process topology convergence result is generated based on the process topology diagram, including: Step 31: Using the anchor key of the project object as the aggregation key, extract the quantity field, unit price field, material quantity field, material unit price field, and contract amount field specified by the data dictionary from the associated data of the project object. The quantity field, unit price field, material quantity field, material unit price field, and contract amount field are all predefined by the data dictionary, and their meanings, units of measurement, and precision rules are clearly constrained within the data dictionary. After extracting the above fields, the field values are converted to units according to the field units defined by the data dictionary, and the values are rounded according to the field precision, thereby eliminating differences in measurement standards and numerical expressions between different source systems.
[0024] Step 32: After completing field alignment, calculate the costs for records with the same anchor key in the project object association data. Specifically, multiply the corresponding values of the quantity field and the unit price field to obtain the project cost details, and sum the project cost details under the same project object anchor key to obtain the total project cost; multiply the corresponding values of the quantity field and the unit price field to obtain the material cost details, and sum the material cost details under the same project object anchor key to obtain the total material cost. Simultaneously, determine the contract amount from the extracted contract amount field according to the contract amount value rules defined in the data dictionary. These contract amount value rules include object domain limitation rules, field order priority rules, and data collection time order rules, used to constrain the value selection method when multiple contract amount records exist, thereby ensuring the uniqueness and determinism of the contract amount in the object dimension.
[0025] Step 33: After obtaining the total project cost, total material cost, and contract amount, perform a consistency calculation on the anchor key for each project object. Specifically, calculate the difference between the total project cost and the contract amount to obtain the project difference; calculate the difference between the total material cost and the total project cost to obtain the material difference. Then, write the project object anchor key, project difference, and material difference into the consistency calculation results. These consistency calculation results are used to characterize the degree of consistency of cost data under different sources and calculation methods, and are an important data basis for generating subsequent audit clues.
[0026] Step 34: Based on the completed consistency calculation, perform aggregation analysis on cost-related data according to the process topology diagram. The process topology diagram is a directed graph structure used to describe the business process structure, including node identifiers and directed connections between nodes. Node identifiers are used to identify different business nodes in the process, and directed connections are used to describe the sequential or dependent relationships between nodes. According to the correspondence between object domains, field names, and node identifiers registered in the data dictionary, the associated data of engineering objects, total engineering costs, total material costs, contract amounts, and consistency calculation results are mapped to the corresponding nodes in the process topology diagram. After completing the mapping, according to the directed connections defined in the process topology diagram, perform bottom-up aggregation calculations on each node to obtain the process topology aggregation result. Specifically, based on the field values mapped to the node identifiers, perform summation to obtain the node input values, perform bottom-up aggregation calculations on the node input values of each node to obtain the node aggregation value, and generate a process topology aggregation result containing engineering object anchor keys, node identifiers, and node aggregation values corresponding to the node identifiers. The process topology aggregation results reflect the hierarchical aggregation of cost data across business processes, supporting the analysis of cost formation paths and structural relationships during the audit process.
[0027] Through the above steps, this invention achieves object-level aggregation, unified calculation, and process-oriented convergence analysis of power engineering audit cost data. By adhering to the unified constraints of field meanings, units, and precision in the data dictionary, the standardization and comparability of the calculation processes for total project costs, total material costs, and contract amounts are ensured. The introduction of project object anchor keys and process topology diagrams enables consistent organization and aggregation of cost data across both project object and business process dimensions. This effectively solves the problems of scattered cost data sources, inconsistent calculation methods, and difficulty in tracing process relationships in power engineering audits.
[0028] In one embodiment of the present invention, version difference information is obtained by constructing a version graph based on a data monotonic chain, and the registration block and version difference information are transferred and referenced according to the audit state set and state transition rules, including: Step 41: Using the entity's unique identifier as the search key, retrieve the registration block containing the entity's unique identifier in the data monotonic chain. Since the registration blocks in the data monotonic chain are written in an append-only manner according to the acquisition time, multiple registration blocks corresponding to the same entity's unique identifier naturally constitute a set of data records evolving over time. Sort the retrieved registration blocks according to the acquisition time to form a version sequence, and use each registration block in the version sequence as a node in the version graph. The version graph is a directed graph structure used to describe the evolution relationship of data versions of the same entity, where nodes correspond to data versions at different points in time.
[0029] Step 42: After forming the nodes of the version graph, directed edges are established for registration blocks with adjacent collection times in the version sequence. Each directed edge represents the data version evolution relationship of the same entity at adjacent time points. The directed edges in the version graph are independent of the directed connections in the process topology graph, serving version difference analysis and process convergence calculation respectively. Subsequently, for the starting and ending registration blocks corresponding to each directed edge, the corresponding standardized original records are read, and a field-by-field comparison is performed according to the field names and field order rules defined in the data dictionary. Through the above field-by-field comparison, the fields that have changed between adjacent versions are identified, generating version difference information. The version difference information includes at least an entity unique identifier, a summary fingerprint pair for identifying the two versions, and a set of field names where records have changed, wherein the set of difference fields is used to clarify the specific range of differences between data versions.
[0030] Step 43: After obtaining version difference information, an audit status set and status transition rules are introduced to standardize and control the audit process. The audit status set describes different audit stages or processing states that may occur during the audit process, and the status transition rules define the transition relationships between audit states when specific conditions are met. By predefining the audit status set and status transition rules, during audit execution, audit states are transitioned according to the status transition rules. At each audit status transition, the corresponding registration block and version difference information are referenced to generate a status transition record. The status transition record is used to associate specific data versions and their differences, providing clear data evidence for changes in audit states.
[0031] Through the above technical solution, this invention achieves structured management of the version evolution process of power engineering audit data within a data governance framework. Based on data monotonic chains and version graphs, it can completely preserve the data versions of the same entity at different points in time and their evolutionary relationships, achieving traceable governance of audit data. By combining version difference information with the audit status transition process, changes in audit status are based on clear data differences, avoiding reliance on human experience or subjective judgment in the audit process. This effectively improves the standardization, consistency, and verifiability of data version management and audit process control in power engineering audit information management.
[0032] In one embodiment of the present invention, generating audit clues based on consistency calculation results and version difference information includes: Step 51: Using the project object anchor key as the index key, retrieve the entity unique identifier corresponding to the project object anchor key from the project object association data, and form a set of entity unique identifiers from the retrieved multiple entity unique identifiers. Simultaneously, establish the correspondence between the project object anchor key and the entity unique identifier set. Through this method, a clear mapping is established between the data at the project object dimension and the data at the entity dimension, providing a foundation for subsequently associating version difference information at the entity level.
[0033] Step 52: After establishing the above correspondence, based on the consistency calculation results, read the engineering difference and material difference corresponding to the anchor key of the engineering object. According to the difference determination rules defined in the data dictionary, perform zero-value determination or non-zero-value determination on the engineering difference and material difference respectively. The difference determination rules are pre-registered in the data dictionary and are used to determine the status of the engineering difference and material difference. Specifically, based on the field precision defined in the data dictionary, perform zero-value determination or non-zero-value determination on the engineering difference and material difference respectively; when the difference is equal to zero under the field precision constraint, it is determined to be in a zero-value state; when the difference is not equal to zero under the field precision constraint, it is determined to be in a non-zero-value state. The data dictionary further defines whether the triggering condition is met when the difference is in a non-zero-value state to determine whether to generate a difference-type audit clue. The difference-type audit clue at least includes the anchor key of the engineering object and the corresponding engineering difference or material difference.
[0034] Step 53: Based on the generated audit trails for discrepancies, according to the correspondence between the anchor keys of engineering objects and the sets of unique entity identifiers, the anchor keys of engineering objects in the audit trails for discrepancies are mapped to the corresponding sets of unique entity identifiers. The sets of unique entity identifiers are then associated with version difference information to generate the final audit trail. The audit trail comprehensively includes engineering discrepancies, material discrepancies, summary fingerprint pairs, and a set of difference fields. The summary fingerprint pairs are used to identify adjacent data versions where discrepancies occur, and the set of difference fields is used to clarify the specific range of fields that have changed. Through the above association method, a direct link is established between cost inconsistency anomalies and specific data version changes.
[0035] Through the above technical solution, this invention achieves multi-dimensional integrated identification of anomalies in power engineering audits. Based on the data dictionary and engineering object anchor keys, it unifies the cost consistency calculation results with the data at the engineering object level. By using unique entity identifiers and version difference information, it further drills down to entity-level data version changes, ensuring that the generation process of audit clues has clear data sources, rule bases, and traceable paths. This effectively improves the standardization, interpretability, and data governance consistency of anomaly identification in the power engineering audit information management process.
[0036] In one embodiment of the present invention, selecting the registration block referenced by the audit clue from the data monotonic chain to assemble the evidence package and generate the evidence package summary fingerprint, and generating the document summary fingerprint based on the structured audit text and the evidence package summary fingerprint and writing it into the registration block, includes: Step 61: Based on the entity unique identifier and summary fingerprint pair contained in the audit lead, using the entity unique identifier as the search key, retrieve registration blocks containing the entity unique identifier in the data monotonic chain to obtain a registration block set. The registration block set refers to multiple data version records associated with the same entity unique identifier in the data monotonic chain. Subsequently, select registration blocks from the registration block set whose summary fingerprints match the summary fingerprint pairs in the audit lead, thereby limiting the range of data versions directly related to the audit lead. This method ensures that the data entering the evidence packaging and assembly process has a clear source and a controlled scope.
[0037] Step 62: After completing the registration block screening, select the starting and ending registration blocks from the matched registration blocks that correspond to the summary fingerprint pairs and are arranged in chronological order of collection time. These will be used as the registration blocks referenced by the audit trail. The starting and ending registration blocks represent the two data versions that differed. Subsequently, the registration blocks referenced by the audit trail are summarized and assembled into an evidence package. The evidence package is a aggregated representation of the data versions involved in the audit trail, used to centrally carry the data evidence required for audit analysis.
[0038] Step 63: After the evidence package is assembled, the registration blocks within the evidence package are sorted according to the collection time, and a fixed order is determined accordingly. Following this fixed order, the summary fingerprints recorded in each registration block are read sequentially and concatenated. An irreversible fixed-length identifier calculation is performed on the concatenated result to generate the evidence package summary fingerprint. The evidence package summary fingerprint uniquely represents the overall content characteristics of all registration blocks within the evidence package, and its calculation result remains stable and consistent when the content of the evidence package remains unchanged.
[0039] Step 64: Based on the generated evidence package summary fingerprint, the structured audit text and the evidence package summary fingerprint are concatenated in a predetermined fixed order. An irreversible fixed-length identifier calculation is then performed on the concatenated result to generate a document summary fingerprint. The structured audit text refers to the audit description content organized according to a predetermined structure, used to record audit analysis conclusions and related explanations. Subsequently, the document summary fingerprint is written into a new register block, and this register block is written into the data monotonic chain in an append-only manner, thereby forming a record of the audit document in the data monotonic chain.
[0040] Through the above technical solution, this invention achieves standardized assembly and summary management of audit evidence for power engineering projects. Relying on data monotonic chains and unique entity identifiers, it accurately locates and fully references the data versions involved in audit leads. By generating evidence package summary fingerprints and document summary fingerprints, audit evidence and documents possess stable content identification and traceability, effectively supporting the management needs of verifiable evidence sources and traceable audit conclusions during power engineering audits.
[0041] In one embodiment of the present invention, selecting the registration block referenced by the audit clue from the data monotonic chain to assemble the evidence package and generate the evidence package digest fingerprint includes: Step 71: After filtering the registration block set in Step 61, first determine the set of field names of the standardized original records corresponding to each registration block based on the data dictionary. By extracting the set of field names, each registration block becomes structurally comparable, providing a unified basis for subsequent filtering based on field differences.
[0042] Step 72: After obtaining the set of field names for each registration block, select registration blocks sequentially from the set of registration blocks to add to the evidence package, using the set of difference fields contained in the audit trail as the coverage target. The set of difference fields refers to the set of field names identified in the version difference information that have changed between adjacent data versions. Specifically, during the selection process, priority is given to adding the registration block with the largest number of field names belonging to the difference field set to the evidence package, ensuring that each selection covers as many difference fields as possible. This method allows for the priority inclusion of the data version most valuable for interpreting the audit trail in the evidence package.
[0043] Step 73: When multiple registration blocks have the same number of differential fields covered, to ensure the determinism and repeatability of the selection process, the selection order is further determined according to the lexicographical order of the registration blocks' summary fingerprints, and the registration blocks are selected accordingly to be added to the evidence package. The above selection and judgment process is repeated until all field names contained in the differential field set in the audit trail are covered by the registration blocks in the evidence package. This method ensures that the evidence package covers all differential fields while containing as few registration blocks as possible, and that the selection rules are clear and the result is unique.
[0044] Through the above technical solution, this invention achieves refined control over the audit evidence selection process. By relying on the unified constraints of field names and structures through a data dictionary, comparability between registration blocks is ensured. Through a screening mechanism that uses the set of differing fields as the coverage target, the evidence package focuses on data changes directly related to audit leads, avoiding redundancy and interference from irrelevant data. This makes the evidence package preparation process characterized by clear rules and verifiable results, enhancing the standardization and relevance of evidence governance in the power engineering audit information management process.
[0045] In one embodiment of the present invention, the rectification targets are identified based on audit clues, and the rectified data is collected and written into a data monotonic chain. Steps 3 to 6 are repeated, and the data dictionary is version-updated to a new version. Step 81: Based on the entity unique identifier and set of difference fields contained in the audit clues, determine the specific business objects that need to be rectified. After determining the rectification objects, collect the rectified data and perform standardization processing on the rectified data according to the data governance process described in Step 1. Subsequently, generate a summary fingerprint for the standardized rectified data, and write the entity unique identifier, source system, collection time, and summary fingerprint into a new registration block. Then, write the registration block into the data monotonic chain in an append-only manner, thereby forming a new version record of the rectified data in the data monotonic chain.
[0046] Step 82: After the rectified data is written into the data monotonic chain, the processing flow described in steps 2 to 5 is executed sequentially on the rectified data. Specifically, this includes: regenerating the engineering object association data, performing consistency calculations on the rectified cost data, and reconstructing version difference information based on the data monotonic chain. Based on this, rectified audit trails are regenerated using the rectified data. Through these methods, the rectified data is compared and analyzed with the original data within the same governance and audit framework, thereby verifying the changes in the data level resulting from the rectification.
[0047] Step 83: After generating the rectified audit leads, the evidence package matching and summary processing flow described in Step 6 is further executed on the rectified audit leads to generate new evidence package summary fingerprints and document summary fingerprints. The document summary fingerprints are then written into a new registration block and appended only to the data monotonic chain. Simultaneously, the data dictionary is version-updated using the set of differing fields contained in the rectified audit leads as a limiting condition. Specifically, only the fields involved in the set of differing fields are updated, including the field unit, field precision, field value range, whether the field is required, field order rules, and the correspondence between source system fields and data dictionary fields, forming a new version of the data dictionary. By limiting the update scope, the governance rules for irrelevant fields are avoided, ensuring the controllability and traceability of the data dictionary evolution.
[0048] Through the above technical solution, this invention achieves a closed-loop linkage between audit findings, data rectification, and governance rule evolution. Rectified data is versioned and recorded via a data monotonic chain, ensuring that historical data is not overwritten and the rectification process is traceable. Version updates of the data dictionary are directly driven by the set of differing fields in the audit leads, providing clear data support for adjustments to governance rules. This effectively solves the problems of difficulty in verifying rectification results and timely evolution of governance rules in power engineering audits, enabling the audit process not only to identify problems but also to drive continuous optimization of data governance rules.
[0049] This invention provides a power engineering audit information management system based on data governance, comprising: The data governance module establishes a data dictionary and generates unique entity identifiers. It standardizes the original records to obtain summary fingerprints, writes the unique entity identifier, source system, collection time, and summary fingerprint into the registration block, and appends them to form a data monotonic chain. The project object anchoring module generates project object anchor keys based on the project unique identifier, object domain, and object local identifier, and binds the entity unique identifier with the project object anchor key as project object associated data; The cost consistency aggregation module obtains the total project cost, total material cost, and contract amount based on the project object association data and data dictionary, generates consistent calculation results, and generates process topology aggregation results based on the process topology diagram; The version evolution control module constructs a version graph based on the data monotonic chain to obtain version difference information, and transfers and references the registration block and version difference information according to the audit status set and status transition rules. The audit trail generation module generates audit trails based on consistency calculation results and version difference information. The audit evidence storage module selects the registration block referenced by the audit clues from the data monotonic chain to assemble the evidence package and generate the evidence package summary fingerprint. Based on the structured audit text and the evidence package summary fingerprint, it generates the document summary fingerprint and writes it into the registration block. The rectification closed-loop evolution module identifies rectification targets based on audit clues and collects data after rectification, writes it into the data monotonic chain, aggregates duplicate expenses to the audit evidence storage module, and updates the data dictionary to the new version.
[0050] It should be noted that the interval and threshold sizes are set for ease of comparison. The size of the threshold depends on the amount of sample data and the base number set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless calculations, and the formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0051] The embodiments of the present invention have been described above, but the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of the present embodiments, all of which are within the protection scope of the present embodiments.
Claims
1. A method for managing power engineering audit information based on data governance, characterized in that, Includes the following steps: Step 1: Establish a data dictionary and generate unique entity identifiers. Standardize the original records to obtain summary fingerprints. Write the unique entity identifier, source system, collection time, and summary fingerprint into the registration block and append them to form a data monotonic chain. Step 2: Generate anchor keys for engineering objects based on the project unique identifier, object domain, and object local identifier, and bind the entity unique identifier and the engineering object anchor keys to the engineering object associated data; Step 3: Based on the project object association data and data dictionary, obtain the total project cost, total material cost, and contract amount, generate consistent calculation results, and generate process topology convergence results based on the process topology diagram; Step 4: Construct a version graph based on the data monotonic chain to obtain version difference information, and transfer and reference the registration block and version difference information according to the audit state set and state transition rules; Step 5: Generate audit trails based on the consistency calculation results and version difference information; Step 6: Select the registration block referenced by the audit clue from the data monotonic chain, assemble the evidence package and generate the evidence package summary fingerprint. Generate the document summary fingerprint based on the structured audit text and the evidence package summary fingerprint and write it into the registration block. Step 7: Based on the audit clues, identify the rectification targets and collect the rectified data, write it into the data monotonic chain, repeat steps 3 to 6, and update the data dictionary to the new version.
2. The power engineering audit information management method based on data governance according to claim 1, characterized in that, A data dictionary is established and unique entity identifiers are generated. The original records are standardized to obtain summary fingerprints. The unique entity identifier, source system, collection time, and summary fingerprint are written into the registration block and appended to form a data monotonic chain, including: Step 11: Establish a data dictionary. The data dictionary defines the field names, data types, units, precision, value ranges, mandatory fields, and order rules for each of the following: project, contract, list, materials, process, equipment, personnel, invoice, and on-site evidence. It also records the correspondence between the source system fields and the data dictionary fields. Step 12: Determine the object domain to which the original record belongs based on the data dictionary, and extract the values of the set of fields marked as entity local identifiers from the data dictionary in the original record. After concatenating the values of the object domain and the set of entity local identifier fields according to the field order rules, perform irreversible fixed-length identifier calculation to obtain the unique identifier of the entity. Step 13: Perform standardization processing on the original records according to the data dictionary. The standardization processing includes rearranging fields according to field order rules, converting units according to field units and rounding values according to field precision, validating values according to field value ranges, and writing the missing values of fields that are required and missing into the data dictionary to obtain the original records after standardization. Step 14: Perform irreversible fixed-length identifier calculation on the standardized original record to obtain the summary fingerprint, write the entity unique identifier, source system, collection time, and summary fingerprint into the registration block, and write the registration block into the data monotonic chain in an append-only manner.
3. The power engineering audit information management method based on data governance according to claim 1, characterized in that, Generate anchor keys for project objects, and bind the unique identifier of the entity to the anchor keys of the project objects as associated data for the project objects, including: Step 21: Extract the unique identifier of the project, the object domain and the object local identifier from the original records according to the data dictionary, and perform standardization processing; Step 22: Connect the standardized project unique identifier, the standardized object domain, and the standardized object local identifier according to the field order rules of the data dictionary, and perform irreversible fixed-length identifier calculation to obtain the project object anchor key; Step 23: Locate the registration block containing the entity's unique identifier in the data monotonic chain based on the entity's unique identifier, and establish a binding relationship between the entity's unique identifier and the project object's anchor key to form project object associated data.
4. The power engineering audit information management method based on data governance according to claim 1, characterized in that, Based on the project object association data and data dictionary, the total project cost, total material cost, and contract amount are obtained. Consistent calculation results are generated, and process topology convergence results are generated based on the process topology diagram, including: Step 31: Using the anchor key of the project object as the aggregation key, extract the project quantity field, project unit price field, material quantity field, material unit price field and contract amount field specified by the data dictionary from the project object associated data, and perform unit conversion and numerical rounding on the specified fields according to the field units and field precision specified by the data dictionary. Step 32: For records with the same anchor key in the associated data of engineering objects, multiply the quantity field and the unit price field to obtain the detailed engineering cost and sum them to obtain the total engineering cost. Multiply the quantity field and the unit price field of materials to obtain the detailed material cost and sum them to obtain the total material cost. Determine the contract amount according to the contract amount value rules defined by the data dictionary. Step 33: For each project object anchor key, calculate the difference between the total project cost and the contract amount to obtain the project difference value, calculate the difference between the total material cost and the total project cost to obtain the material difference value, and write the project object anchor key, project difference value, and material difference value into the consistency calculation result. Step 34: Map the associated data of the project objects, the total project cost, the total material cost, the contract amount, and the consistency calculation results to the nodes of the process topology diagram according to the correspondence between the object domains, field names, and node identifiers registered in the data dictionary. Then, perform bottom-up aggregation calculation on each node according to the directed connection relationship of the process topology diagram to obtain the process topology convergence result. The process topology diagram includes node identifiers and directed connection relationships between nodes.
5. The power engineering audit information management method based on data governance according to claim 1, characterized in that, Version difference information is obtained by constructing a version graph based on the data monotonic chain. The version difference information is then transferred and referenced according to the audit state set and state transition rules, including: Step 41: Use the entity's unique identifier as the retrieval key to retrieve the registration block containing the entity's unique identifier in the data monotonic chain, and sort the registration blocks according to the collection time to obtain the version sequence. Use the registration blocks in the version sequence as nodes of the version graph. Step 42: Establish directed edges to form a version graph for the registration blocks that are adjacent in the collection time in the version sequence, and compare the standardized original records corresponding to the starting and ending registration blocks of each directed edge field by field according to the field names and field order rules defined by the data dictionary to generate version difference information containing entity unique identifier, summary fingerprint pair and difference field set. Step 43: Predefine the audit state set and state transition rules, execute the audit state transition according to the state transition rules, and generate a state transition record by referencing the registration block and version difference information corresponding to the audit state transition each time the audit state transition occurs.
6. The power engineering audit information management method based on data governance according to claim 1, characterized in that, Audit trails are generated based on consistency calculation results and version difference information, including: Step 51: Using the anchor key of the project object as the index key, retrieve the corresponding entity unique identifier from the associated data of the project object to form a set of entity unique identifiers, and establish the correspondence between the anchor key of the project object and the set of entity unique identifiers. Step 52: Read the engineering difference and material difference corresponding to the anchor key of the engineering object according to the consistency calculation result, and determine the zero value or non-zero value of the engineering difference and material difference according to the difference judgment rules defined by the data dictionary. When the triggering condition of the difference judgment rule is met, generate a difference class audit clue containing the anchor key of the engineering object and the engineering difference or material difference. Step 53: Based on the correspondence between the anchor key of the engineering object and the set of unique entity identifiers, map the anchor key of the engineering object of the difference-type audit clue to the set of unique entity identifiers, and associate the set of unique entity identifiers with the version difference information to generate an audit clue. The audit clue includes engineering difference, material difference, summary fingerprint pairs and difference field set.
7. The power engineering audit information management method based on data governance according to claim 1, characterized in that, The evidence package is assembled from the registration block referenced by the audit clues selected from the data monotonic chain, and an evidence package summary fingerprint is generated. A document summary fingerprint is generated based on the structured audit text and the evidence package summary fingerprint and written into the registration block, including: Step 61: Read the entity unique identifier and summary fingerprint pair according to the audit clues, retrieve the registration block containing the entity unique identifier in the data monotonic chain using the entity unique identifier as the search key to obtain the registration block set, and filter out the registration block in the registration block set that matches the summary fingerprint and summary fingerprint pair; Step 62: Select the start registration block and the end registration block that correspond to the summary fingerprint and are collected in chronological order as the registration blocks referenced by the audit clues, and assemble the registration blocks referenced by the audit clues into an evidence package; Step 63: Sort the registration blocks in the evidence package according to the collection time and determine the fixed order. Read the summary fingerprint of each registration block in the fixed order and connect them. Perform irreversible fixed-length identifier calculation to obtain the summary fingerprint of the evidence package. Step 64: Connect the structured audit text and the evidence package summary fingerprint in a fixed order, perform irreversible fixed-length identifier calculation to obtain the document summary fingerprint, and write the document summary fingerprint into a new registration block and then write it into the data monotonic chain in an append-only manner.
8. The power engineering audit information management method based on data governance according to claim 7, characterized in that, The evidence package is assembled from the registration blocks referenced by the audit trail selected from the data monotonic chain, and an evidence package summary fingerprint is generated, including: Step 71: From the set of registration blocks obtained in step 61, determine the set of field names of the standardized original records corresponding to each registration block according to the data dictionary; Step 72: Using the set of discrepancies in the audit trail as the coverage target, successively select the registration block in the registration block set that has the largest number of field names belonging to the set of discrepancies in the field name set and add it to the evidence package; Step 73: When the number of fields is equal, the selection order is determined by the lexicographical order of the summary fingerprint of the registration block, and the selection is repeated until the field names of the difference field set are all covered by the registration blocks in the evidence package.
9. A power engineering audit information management method based on data governance according to claim 1, characterized in that, Based on audit clues, identify the targets for rectification and collect data after rectification. Write the data into a data monotonic chain, repeat steps 3 to 6, and version the data dictionary to the new version, including: Step 81: Read the entity unique identifier and difference field set according to the audit clues, determine the rectification object by the entity unique identifier and collect the rectification data, perform standardization processing on the rectification data according to step 1 to obtain the summary fingerprint, write it into the registration block and then write it into the data monotonic chain in an append-only manner. Step 82: Perform steps 2 to 5 sequentially on the rectified data to update the associated data of the engineering object and generate the consistency calculation results and version difference information after rectification. Generate audit clues after rectification based on the consistency calculation results and version difference information. Step 83: After performing step 6 to generate evidence package summary fingerprints and document summary fingerprints for the rectified audit clues and writing them into the new registration block, write them into the data monotonic chain in an append-only manner. Then, limit the update scope with the set of difference fields in the rectified audit clues. Update the corresponding fields in the data dictionary in terms of field units, field precision, field value range, whether the field is required, field order rules, and the correspondence between the source system fields and the data dictionary fields to form a new version of the data dictionary.
10. A power engineering audit information management system based on data governance, characterized in that, The power engineering audit information management method based on data governance as described in any one of claims 1-9 includes: The data governance module establishes a data dictionary and generates unique entity identifiers. It standardizes the original records to obtain summary fingerprints, writes the unique entity identifier, source system, collection time, and summary fingerprint into the registration block, and appends them to form a data monotonic chain. The project object anchoring module generates project object anchor keys based on the project unique identifier, object domain, and object local identifier, and binds the entity unique identifier with the project object anchor key as project object associated data; The cost consistency aggregation module obtains the total project cost, total material cost, and contract amount based on the project object association data and data dictionary, generates consistent calculation results, and generates process topology aggregation results based on the process topology diagram; The version evolution control module constructs a version graph based on the data monotonic chain to obtain version difference information, and transfers and references the registration block and version difference information according to the audit status set and status transition rules. The audit trail generation module generates audit trails based on consistency calculation results and version difference information. The audit evidence storage module selects the registration block referenced by the audit clues from the data monotonic chain to assemble the evidence package and generate the evidence package summary fingerprint. Based on the structured audit text and the evidence package summary fingerprint, it generates the document summary fingerprint and writes it into the registration block. The rectification closed-loop evolution module identifies rectification targets based on audit clues and collects data after rectification, writes it into the data monotonic chain, aggregates duplicate expenses to the audit evidence storage module, and updates the data dictionary to the new version.