Power grid project cost rationality analysis system

Through data fusion, dynamic thresholds and machine learning models, combined with the cost map module and anomaly analysis module, the problems of real-time monitoring and anomaly analysis in power grid project cost analysis are solved, real-time monitoring and anomaly analysis of power grid project costs are realized, and fault diagnosis and decision-making efficiency are improved.

CN120633638AInactive Publication Date: 2025-09-12GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD

Patent Information

Application Number
CN202510742525.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing technology, power grid project cost analysis relies on static reports and manual post-audit, which cannot monitor cost anomalies in real time during project implementation. As a result, irreversible losses have already occurred when problems are discovered. Manual investigation of the causes of cost anomalies is time-consuming and relies on experience.

Method used

The data fusion module is used to fuse multi-source data in real time, dynamic thresholds and machine learning models are used to track project progress, the cost map module is combined to build an influencing factor map, and the anomaly analysis module is used to automatically generate attribution reports to achieve automated analysis of abnormal data.

Benefits of technology

It realizes real-time monitoring and anomaly analysis of power grid project costs, reduces manual intervention, improves fault diagnosis and decision-making efficiency, supports personalized monitoring based on historical behavior and time dynamic changes, and generates highly readable attribution reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633638A_ABST
    Figure CN120633638A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, in particular to a power grid project cost rationality analysis system, which comprises a data fusion module for collecting heterogeneous data sources of a power grid project, fusing multi-source data in real time, and converting the data into standard data through a unified data cleaning engine; the dynamic analysis module is used for setting a dynamic threshold value according to the standard data, tracking a single index of a project progress and the dynamic threshold value based on a machine learning model, and positioning abnormal data; according to entities and relations in the standard data, when the system is used, synonyms or different-name entities in multiple sources are efficiently recognized and unified through the graph matching and embedding technology, manual account checking is reduced, a dynamic threshold value is combined with a machine learning model, the execution condition of each task in a project is continuously tracked, and the construction efficiency is improved. Different from a traditional fixed threshold value, the system supports personalized monitoring according to historical behaviors, target curves and time dynamic changes, and is beneficial to being closer to actual construction fluctuation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular to a system for analyzing the rationality of power grid project costs. Background Art

[0002] Cost rationality analysis for power grid projects is crucial for ensuring reasonable project costs, avoiding waste, and improving investment returns. This analysis typically involves considering multiple aspects, including but not limited to project planning, design, construction, and subsequent operations and maintenance costs.

[0003] Patent publication number CN104077665A states in its specification that “the present invention discloses a system and method for collecting power grid project cost analysis data; comprising a user basic data maintenance module for managing users, departments, roles and resources, the user management module receiving a call from a business data processing module; a Bowei file conversion engine for parsing Bowei files and converting Bowei files into files in DataSet format or Xml format; the Bowei files include cost 2008 version projects and settlement 2012 version projects; the Bowei file conversion engine receiving a call from a business data processing module; an Excel file import and export engine for parsing Excel files in xls and xlsx formats according to a configuration file.” The import analysis can identify Excel's multi-layer column headers, and the export analysis can automatically generate the required Excel format. The Excel file import and export engine receives the call of the business data processing module; the indicator data table of the present invention is very stable during subsequent upgrades and adjustments. Although the above technology realizes flexible adjustment of indicators through configurable Excel templates, solving the problem of frequent code modifications caused by frequent changes in indicators in traditional systems, traditional cost analysis relies on static reports and manual post-audit, and cannot monitor cost anomalies during project implementation in real time, resulting in irreversible losses when problems are discovered. The existing technology only provides data comparison or simple statistics, and manual investigation of the causes of cost anomalies is required, which is time-consuming and relies on experience.

[0004] In summary, developing a power grid project cost rationality analysis system is still a key issue that needs to be solved urgently in the field of data analysis technology. Summary of the Invention

[0005] The purpose of the present invention is to solve the problem in the prior art that although the above-mentioned technology realizes flexible adjustment of indicators through configurable Excel templates and solves the problem of frequent code modification caused by frequent changes in indicators in traditional systems, traditional cost analysis relies on static reports and manual post-audit, and is unable to monitor cost anomalies during project implementation in real time, resulting in irreversible losses when problems are discovered. The existing technology only provides data comparison or simple statistics, and manual investigation of the causes of cost anomalies is required, which is time-consuming and dependent on experience.

[0006] To achieve the above-mentioned object, the present invention provides a power grid project cost rationality analysis system, comprising: a data fusion module, which collects heterogeneous data sources of power grid projects, fuses multi-source data in real time, and converts it into standard data through a unified data cleaning engine; A dynamic analysis module sets dynamic thresholds based on the standard data, tracks individual indicators of project progress and the dynamic thresholds based on a machine learning model, and locates abnormal data; A cost map module constructs a cost influencing factor map based on the entities and relationships in the standard data; The anomaly analysis module uses a graph reasoning engine to associate unstructured data with structured cost data based on the anomaly data and automatically generates an attribution report.

[0007] Furthermore, the operation process of the data fusion module includes: The heterogeneous data sources include budget data sets, settlement data sets, material purchase price data, and historical project libraries. A heterogeneous data tensor constructor is introduced to uniformly model these data stored in different formats, granularities, and semantic structures. The expression formula is: ,in Represents time The target data below, Indicates from arrive Take the union of all subitems, Indicates the Function of elements , Representative Input data, Indicates the Input-related parameters, It represents the function The domain and range of Represents input The dimension is , Indicates the Additional inputs related to the function, Representation function The output result is a dimensional real vector space, using the embedding mapping plus graph matching model and setting the matching threshold , match multiple source entities (equipment code, material category, project name), expression formula: ,in represents the matching confidence after embedding, is an embedding operation, Represent graph structure, text data, node features and historical data respectively, Indicates that these sets To merge, Represents two input and To match, Represents two input data, Respectively represent input and The vector representation after the embedding operation is represents a weight matrix used to calculate the relationship between embedding vectors, represents the transpose operation, Represents the activation function, the matching confidence is higher than the matching threshold Entity pair , considered as the same standard entity .

[0008] Furthermore, the operation process of the data fusion module includes: The real-time fusion of multi-source data is processed using a weighted attention fusion mechanism, which is expressed as: ,in Indicates in Category and The weighted value of the indicator, Indicates that the status Next Category and The original observation value of the indicator, is the weighting coefficient, For all the sets Status Sum, aggregate It is with A collection of all states related to a category, Is an exponential function representing the state Next Category and Characteristic function of the indicator The output value of Is a characteristic function according to the state and contextual information related to it Calculate a value, is an exponential function used to transform eigenvalues ​​into higher weights, For all states Sum, Is an exponential function representing the state Next Category and Characteristic function of the indicator The output value of handles null values, abnormal values ​​and format conflicts. The final standardized data structure is defined as: , in Representative time point The standardized data matrix under Indicates a time point Next 1st project Category No. The weighted value of the indicator, Indicates a time point Next 2nd project Category No. The weighted value of the indicator, Indicates a time point Next Project No. Category No. The weighted value of the indicator, The dimension of the data matrix is , is the number of items is the number of indicators.

[0009] Furthermore, the operation process of the dynamic analysis module includes: From the standard data tensor In [1], we extract the individual progress indicators of the power grid project and construct a multi-dimensional time series, expressing the formula: ,in It's in time The data set at a moment, It's in time Moment The first project The observed values ​​of the progress indicators, The index value range of the project is from 1 to There are a total of projects, The index value range of the progress indicator is , Indicates that the time span of the data is from the first moment to the moment, and for each individual progress indicator , set dynamic threshold , expression formula: ,in is the dynamic threshold of the indicator, is the threshold model, is the history state encoder, The engineering unit at time The target control curve encoding, Its dynamic uncertainty is modeled by a Gaussian process, It follows a normal distribution. is the mean value representing the central expected value of the dynamic threshold, It is Project No. The historical behavior characteristic encoding of each indicator.

[0010] Furthermore, the operation process of the dynamic analysis module includes: The current actual single progress indicator value Compared with the dynamic threshold interval, the residual tensor is constructed and expressed as: ,in It's in time No. The first project The standardized residual value of each indicator, It's in time Moment The first project The observed values ​​of the progress indicators, is the mean value representing the central expected value of the dynamic threshold, Its dynamic uncertainty is modeled by a Gaussian process, It's time All items All indicators The two-dimensional tensor composed of the standardized residuals of , using graph convolution-variational autoencoder, constructs time series feature compression and abnormal probability estimation, expressed as: ,in Indicates the The first project Indicators in time The anomaly score, It is the Softmax function that maps the output to a probability distribution, It is the second half of the variational autoencoder (VAE) used to reconstruct the latent space vector into the output. represents a trained variational autoencoder model, is an anomaly detector, It is the first The indicators from time point 1 to A two-dimensional tensor of Represents vector concatenation operation, It is a project The structured static information vector of When the current engineering unit is determined No. Progress indicators in time Abnormal, locate abnormal data Record to the exception log, and collect the exception data to drive the early warning system and reversely optimize the threshold model With anomaly detector training process.

[0011] Furthermore, the operation process of the cost map module includes: Through the standard data tensor Define entity set and express formula: ,in is a collection of entities, It is a project engineering node. It is a component (equipment), It is raw materials and procurement items, It's the construction process. The external environment (region, season, price) uses a multi-modal embedding and relationship extraction model to embed the influencing factors in the structured data and text into a joint semantic space. The expression formula is: ,in is an entity The vector representation of Indicates that the entity is encoded through an encoder Encoded into a vector, is the set of encoder parameters, is an entity and The vector representation of Representing an entity and The type of predicted relationship between It is based on the relationship between two entity vectors to infer their relationship type. It is The weight matrix corresponding to the relationship type, is an entity and Through the The matching score of the class relationship, is the activation function used to normalize the score into a probability form, Is to find the relationship type index with the largest score , output the structured semantic relationship labels between entity pairs ,in It is a relationship set. When constructing the cost influencing factor map, all the entity-relationship-entity pairs obtained are used to form a map triple set, which is expressed as follows: ,in is a set of triples, Is a triple representing the entity To Entity There is a relationship between , Represents the head entity in the triple and tail entity All belong to entity collection , Representing an entity and The relationship between Belongs to a relation set , each node Incidental attribute vector , forming a heterogeneous graph with attributes, expressed as: ,in It represents a graph structure used to organize and express cost maps related to engineering costs. is a collection of entities, is a set of relations, is a set of triples, is a collection of attributes, is an entity The attribute vector of For all entities in the graph Extract the attribute vector to form the attribute set .

[0012] Furthermore, the operation process of the cost map module includes: The propagation factor path integral model is introduced to characterize the indirect impact of different nodes on the total cost and calculate the cost from the source node. To the cost aggregation node The causal path contribution of is expressed as: ,in Indicates from arrive The set of all paths of is the edge weight propagation coefficient, is the causal weight of the relationship type, is the attention correlation function between attribute vectors, It is from arrive All valid reasoning paths Perform the summation, It is the path All triples (entity-relationship-entity) on are multiplied in sequence. Indicates the entity in the path Through the relationship To Entity The graph triples of is an entity The attribute feature vector of .

[0013] Furthermore, the operation process of the cost map module includes: By using heterogeneous graph convolution and structure-preserving objective functions to calculate node embeddings, subsequent predictions and visualizations are performed, expressed as: ,in is an entity The final embedding vector of It is a neural network used by hypergraph neural network to learn entity vectors from graph structure. It is a cost map. is the graph structure preservation loss function, It is to traverse the triple relationship in all graphs, is the TransE model structure, is the relationship vector, is the weight hyperparameter, is an entity The original attribute vector of is the predicted embedding obtained by inputting entity attributes into a multi-layer perceptron (MLP) neural network. The attribute preservation loss represents the distance between the vector of the current entity and the embedding based on attribute prediction. The final output format of the graph is, ,in It is a cost map. is the set of vector embeddings of all entities, is the graph inference function.

[0014] Furthermore, the operation process of the abnormality analysis module includes: Assume that the input abnormal data is a set of event sequences, and express the formula: ,in represents a set of abnormal data samples, Indicates the Abnormal samples, is the total number of abnormal samples, Indicates the The data ranges from 1 to Perform a collection traversal. Indicates the The numerical feature vector of abnormal data, Indicates the The timestamp of the data. The context label representing the anomaly. The anomaly data is mapped to the entity set in the graph through a soft semantic matching function. and triple sets , the unstructured data is associated with the structured cost data, and the unstructured text set is Each text is aligned with the deep semantic encoder and graph entity embedding to finally select a group of strongly related texts Mapping to abnormal nodes Surrounding influence area .

[0015] Furthermore, the operation process of the abnormality analysis module includes: Based on the heterogeneous graph, a differentiable causal graph reasoning network is constructed to infer the root cause path, and the expression formula is: ,in is the relational reasoning matrix, is the structural path causal influence factor, is the indicator function marking the true root cause, Yes arrive The overall loss function of all abnormal data Sum the values ​​of It is The attribution loss value of abnormal data, For all Candidate causal nodes of anomalies Take the negative logarithm after doing the weighted sum, Is a node in the graph The embedding vector representation of It is abnormal data The embedded representation after mapping to the graph node, is the node similarity score, It is to perform softmax normalization on the total scores of all candidate nodes. is the sum representing the probability value, Represents a causal node For abnormal events The attribution score, For all slave nodes To abnormal node A collection of paths Sum, In each path Perform product operations on each edge in the path, is a graph triple on the path, Represents an edge The confidence level, is the distance decay factor, is the comprehensive confidence score of the attribution path. The automatic generation of the attribution report is based on the templated report generation model, expressed as follows: ,in It is Anomaly attribution reports, is the report generation function, It is the The abnormal data is mapped to the entity node in the graph. is the set of all causal entity nodes in the attribution path, Is related to each causal node A collection of matched unstructured text summaries, is the comprehensive confidence score of the attribution path.

[0016] Beneficial effects Compared with the known public technology, the technical solution provided by the present invention has the following beneficial effects: When in use, the present invention uses graph matching and embedding technology to efficiently identify and unify synonyms or heteronymous entities in multiple sources, reducing manual reconciliation. Dynamic thresholds are combined with machine learning models to continuously track the execution of each task in the project. Unlike traditional fixed thresholds, this system supports personalized monitoring based on historical behavior, target curves and dynamic changes over time, which is conducive to being closer to actual construction fluctuations.

[0017] When used, the present invention is conducive to clearly displaying the factors affecting project costs in the form of a graph, breaking through the limitations of traditional tabular and isolated analysis, efficiently extracting entities and relationships from structured data and text, facilitating the construction of a standard and unified knowledge asset library, and automatically combining attribution analysis with structured and unstructured data to generate a highly readable attribution report. In practical applications, it is convenient to reduce the workload of manually reviewing data and contracts one by one, which is conducive to improving fault diagnosis and decision-making efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a system diagram of a power grid project cost rationality analysis system of the present invention. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0020] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products or devices.

[0021] The present invention is described in further detail below with reference to the accompanying drawings: Example: like Figure 1 As shown, the present invention provides a power grid project cost rationality analysis system, including: a data fusion module, which collects heterogeneous data sources of power grid projects, and fuses multi-source data in real time, and converts it into standard data through a unified data cleaning engine; Furthermore, the operation process of the data fusion module includes: The heterogeneous data sources include budget data sets, settlement data sets, material purchase price data, and historical project libraries. A heterogeneous data tensor constructor is introduced to uniformly model these data stored in different formats, granularities, and semantic structures. The expression formula is: ,in Represents time The target data below, Indicates from arrive Take the union of all subitems, Indicates the Function of elements , Representative Input data, Indicates the Input-related parameters, It represents the function The domain and range of Represents input The dimension is , Indicates the Additional inputs related to the function, Representation function The output result is a dimensional real vector space, using the embedding mapping plus graph matching model and setting the matching threshold , match multiple source entities (equipment code, material category, project name), expression formula: ,in represents the matching confidence after embedding, is an embedding operation, Represent graph structure, text data, node features and historical data respectively, Indicates that these sets To merge, Represents two input and To match, Represents two input data, Respectively represent input and The vector representation after the embedding operation is represents a weight matrix used to calculate the relationship between embedding vectors, represents the transpose operation, Represents the activation function, the matching confidence is higher than the matching threshold Entity pair , considered as the same standard entity .

[0022] Furthermore, the operation process of the data fusion module includes: The real-time fusion of multi-source data is processed using a weighted attention fusion mechanism, which is expressed as: ,in Indicates in Category and The weighted value of the indicator, Indicates that the status Next Category and The original observation value of the indicator, is the weighting coefficient, For all the sets Status Sum, aggregate It is with A collection of all states related to a category, Is an exponential function representing the state Next Category and Characteristic function of the indicator The output value of Is a characteristic function according to the state and contextual information related to it Calculate a value, is an exponential function used to transform eigenvalues ​​into higher weights, For all states Sum, Is an exponential function representing the state Next Category and Characteristic function of the indicator The output value of handles null values, abnormal values ​​and format conflicts. The final standardized data structure is defined as: , in Representative time point The standardized data matrix under Indicates a time point Next 1st project Category No. The weighted value of the indicator, Indicates a time point Next 2nd project Category No. The weighted value of the indicator, Indicates a time point Next Project No. Category No. The weighted value of the indicator, The dimension of the data matrix is , is the number of items is the number of indicators.

[0023] Specifically, the data fusion module of this system is used to uniformly collect, clean, match and fuse data from various sources and structures in power grid projects, and ultimately convert them into standard data with unified structure and analyzability, thereby breaking down data silos and unifying modeling. It can integrate data originally scattered across budgeting, settlement, procurement and other systems from multiple sources, thereby significantly improving data utilization efficiency. Through graph matching and embedding technology, it can efficiently identify and unify synonyms or heteronymous entities in multiple sources, reducing manual reconciliation.

[0024] A dynamic analysis module sets dynamic thresholds based on the standard data, tracks individual indicators of project progress and the dynamic thresholds based on a machine learning model, and locates abnormal data; Furthermore, the operation process of the dynamic analysis module includes: From the standard data tensor In [1], we extract the individual progress indicators of the power grid project and construct a multi-dimensional time series, expressing the formula: ,in It's in time The data set at a moment, It's in time Moment The first project The observed values ​​of the progress indicators, The index value range of the project is from 1 to There are a total of projects, The index value range of the progress indicator is , Indicates that the time span of the data is from the first moment to the moment, and for each individual progress indicator , set dynamic threshold , expression formula: ,in is the dynamic threshold of the indicator, is the threshold model, is the history state encoder, The engineering unit at time The target control curve encoding, Its dynamic uncertainty is modeled by a Gaussian process, It follows a normal distribution. is the mean value representing the central expected value of the dynamic threshold, It is Project No. The historical behavior characteristic encoding of each indicator.

[0025] Furthermore, the operation process of the dynamic analysis module includes: The current actual single progress indicator value Compared with the dynamic threshold interval, the residual tensor is constructed and expressed as: ,in It's in time No. The first project The standardized residual value of each indicator, It's in time Moment The first project The observed values ​​of the progress indicators, is the mean value representing the central expected value of the dynamic threshold, Its dynamic uncertainty is modeled by a Gaussian process, It's time All items All indicators The two-dimensional tensor composed of the standardized residuals of , using graph convolution-variational autoencoder, constructs time series feature compression and abnormal probability estimation, expressed as: ,in Indicates the The first project Indicators in time The anomaly score, It is the Softmax function that maps the output to a probability distribution, It is the second half of the variational autoencoder (VAE) used to reconstruct the latent space vector into the output. represents a trained variational autoencoder model, is an anomaly detector, It is the first The indicators from time point 1 to A two-dimensional tensor of Represents vector concatenation operation, It is a project The structured static information vector of When the current engineering unit is determined No. Progress indicators in time Abnormal, locate abnormal data Record to the exception log, and collect the exception data to drive the early warning system and reversely optimize the threshold model With anomaly detector training process.

[0026] Specifically, the dynamic analysis module of this system is used to monitor the construction progress indicators in power grid engineering projects in real time and detect anomalies. It sets dynamic thresholds that change over time based on standard data, and combines them with machine learning models to continuously track the execution of each task in the project. Once a certain progress is found to deviate from the expected range, it is promptly identified as an anomaly and recorded. Unlike traditional fixed thresholds, this system supports personalized monitoring based on historical behavior, target curves and dynamic changes over time, which is conducive to being closer to actual construction fluctuations. By using graph convolution and variational autoencoder technology, it can simultaneously consider the complex relationships between different projects and indicators, effectively compress time features and accurately identify small but critical anomalies.

[0027] A cost map module constructs a cost influencing factor map based on the entities and relationships in the standard data; Furthermore, the operation process of the cost map module includes: Through the standard data tensor Define entity set and express formula: ,in is a collection of entities, It is a project engineering node. It is a component (equipment), It is raw materials and procurement items, It's the construction process. The external environment (region, season, price) uses a multi-modal embedding and relationship extraction model to embed the influencing factors in the structured data and text into a joint semantic space. The expression formula is: ,in is an entity The vector representation of Indicates that the entity is encoded through an encoder Encoded into a vector, is the set of encoder parameters, is an entity and The vector representation of Representing an entity and The type of predicted relationship between It is based on the relationship between two entity vectors to infer their relationship type. It is The weight matrix corresponding to the relationship type, is an entity and Through the The matching score of the class relationship, is the activation function used to normalize the score into a probability form, Is to find the relationship type index with the largest score , output the structured semantic relationship labels between entity pairs ,in It is a relationship set. When constructing the cost influencing factor map, all the entity-relationship-entity pairs obtained are used to form a map triple set, which is expressed as follows: ,in is a set of triples, Is a triple representing the entity To Entity There is a relationship between , Represents the head entity in the triple and tail entity All belong to entity collection , Representing an entity and The relationship between Belongs to a relation set , each node Incidental attribute vector , forming a heterogeneous graph with attributes, expressed as: ,in It represents a graph structure used to organize and express cost maps related to engineering costs. is a collection of entities, is a set of relations, is a set of triples, is a collection of attributes, is an entity The attribute vector of For all entities in the graph Extract the attribute vector to form the attribute set .

[0028] Furthermore, the operation process of the cost map module includes: The propagation factor path integral model is introduced to characterize the indirect impact of different nodes on the total cost and calculate the cost from the source node. To the cost aggregation node The causal path contribution of is expressed as: ,in Indicates from arrive The set of all paths of is the edge weight propagation coefficient, is the causal weight of the relationship type, is the attention correlation function between attribute vectors, It is from arrive All valid reasoning paths Perform the summation, It is the path All triples (entity-relationship-entity) on are multiplied in sequence. Indicates the entity in the path Through the relationship To Entity The graph triples of is an entity The attribute feature vector of .

[0029] Furthermore, the operation process of the cost map module includes: By using heterogeneous graph convolution and structure-preserving objective functions to calculate node embeddings, subsequent predictions and visualizations are performed, expressed as: ,in is an entity The final embedding vector of It is a neural network used by hypergraph neural network to learn entity vectors from graph structure. It is a cost map. is the graph structure preservation loss function, It is to traverse the triple relationship in all graphs, is the TransE model structure, is the relationship vector, is the weight hyperparameter, is an entity The original attribute vector of is the predicted embedding obtained by inputting entity attributes into a multi-layer perceptron (MLP) neural network. The attribute preservation loss represents the distance between the vector of the current entity and the embedding based on attribute prediction. The final output format of the graph is, ,in It is a cost map. is the set of vector embeddings of all entities, is the graph inference function.

[0030] Specifically, the cost map module of this system converts various standard data in power grid projects (such as equipment, materials, construction processes, etc.) into a map structure with semantic relationships, so as to comprehensively characterize the direct and indirect impact of various factors on the total project cost. It includes multiple complex steps such as entity extraction, relationship recognition, multimodal fusion, map construction and reasoning analysis, and finally forms an intelligent cost knowledge map that can be used for prediction, attribution, and optimization. Then, through the multimodal embedding model, the above structured information (such as database fields) and unstructured descriptions (such as technical specifications and contract texts) are embedded in a shared semantic space, which is conducive to clearly displaying the factors affecting the project cost in the form of a map, breaking through the limitations of traditional tabular and isolated analysis, and efficiently extracting entities and relationships from structured data and text, facilitating the construction of a standard and unified knowledge asset library.

[0031] The anomaly analysis module uses a graph inference engine to correlate unstructured data with structured cost data based on the anomaly data and automatically generates an attribution report; Furthermore, the operation process of the abnormality analysis module includes: Assume that the input abnormal data is a set of event sequences, and express the formula: ,in represents a set of abnormal data samples, Indicates the Abnormal samples, is the total number of abnormal samples, Indicates the The data ranges from 1 to Perform a collection traversal. Indicates the The numerical feature vector of abnormal data, Indicates the The timestamp of the data. The context label representing the anomaly. The anomaly data is mapped to the entity set in the graph through a soft semantic matching function. and triple sets , the unstructured data is associated with the structured cost data, and the unstructured text set is Each text is aligned with the deep semantic encoder and graph entity embedding to finally select a group of strongly related texts Mapping to abnormal nodes Surrounding influence area .

[0032] Furthermore, the operation process of the abnormality analysis module includes: Based on the heterogeneous graph, a differentiable causal graph reasoning network is constructed to infer the root cause path, and the expression formula is: ,in is the relational reasoning matrix, is the structural path causal influence factor, is the indicator function marking the true root cause, Yes arrive The overall loss function of all abnormal data Sum the values ​​of It is The attribution loss value of abnormal data, For all Candidate causal nodes of anomalies Take the negative logarithm after doing the weighted sum, Is a node in the graph The embedding vector representation of It is abnormal data The embedded representation after mapping to the graph node, is the node similarity score, It is to perform softmax normalization on the total scores of all candidate nodes. is the sum representing the probability value, Represents a causal node For abnormal events The attribution score, For all slave nodes To abnormal node A collection of paths Sum, In each path Perform product operations on each edge in the path, is a graph triple on the path, Represents an edge The confidence level, is the distance decay factor, is the comprehensive confidence score of the attribution path. The automatic generation of the attribution report is based on the templated report generation model, expressed as follows: ,in It is Anomaly attribution reports, is the report generation function, It is the The abnormal data is mapped to the entity node in the graph. is the set of all causal entity nodes in the attribution path, Is related to each causal node A collection of matched unstructured text summaries, is the comprehensive confidence score of the attribution path.

[0033] Specifically, the anomaly analysis module of this system aims to use graph reasoning technology to perform automated attribution analysis on abnormal data on power grid costs, and combine structured and unstructured data to generate highly readable attribution reports. In practical applications, it is convenient to reduce the workload of manually reviewing data and contracts one by one, which is conducive to improving fault diagnosis and decision-making efficiency. At the same time, it deeply integrates structured data and unstructured text to facilitate more comprehensive anomaly identification and analysis. The automatically generated Chinese natural language report makes it easier for managers to quickly understand the problem and make adjustment decisions.

[0034] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A power grid project cost rationality analysis system, characterized in that: include: The data fusion module collects heterogeneous data sources of power grid projects, integrates multi-source data in real time, and converts it into standard data through a unified data cleaning engine; A dynamic analysis module sets dynamic thresholds based on the standard data, tracks individual indicators of project progress and the dynamic thresholds based on a machine learning model, and locates abnormal data; A cost map module constructs a cost influencing factor map based on the entities and relationships in the standard data; The anomaly analysis module uses a graph reasoning engine to associate unstructured data with structured cost data based on the anomaly data and automatically generates an attribution report.

2. A power grid project cost rationality analysis system according to claim 1, characterized in that: The operation process of the data fusion module includes: The heterogeneous data sources include budget data sets, settlement data sets, material purchase price data, and historical project libraries. A heterogeneous data tensor constructor is introduced to uniformly model these data stored in different formats, granularities, and semantic structures. The expression formula is: ,in Represents time The target data below, Indicates from arrive Take the union of all subitems, Indicates the Function of elements , Representative Input data, Indicates the Input-related parameters, It represents the function The domain and range of Represents input The dimension is , Indicates the Additional inputs related to the function, Representation function The output result is a dimensional real vector space, using the embedding mapping plus graph matching model and setting the matching threshold , match multiple source entities (equipment code, material category, project name), expression formula: ,in represents the matching confidence after embedding, is an embedding operation, Represent graph structure, text data, node features and historical data respectively, Indicates that these sets To merge, Represents two input and To match, Represents two input data, Respectively represent input and The vector representation after the embedding operation is Represents a weight matrix used to calculate the relationship between embedding vectors, represents the transpose operation, Represents the activation function, the matching confidence is higher than the matching threshold Entity pair , considered as the same standard entity .

3. A power grid project cost rationality analysis system according to claim 2, characterized in that: The operation process of the data fusion module includes: The real-time fusion of multi-source data is processed using a weighted attention fusion mechanism, which is expressed as: ,in Indicates in Category and The weighted value of the indicator, Indicates that the status Next Category and The original observation value of the indicator, is the weighting coefficient, For all the sets Status Sum, aggregate It is with A collection of all states related to a category, Is an exponential function representing the state Next Category and Characteristic function of the indicator The output value of Is a characteristic function according to the state and contextual information related to it Calculate a value, is an exponential function used to transform eigenvalues ​​into higher weights, For all states Sum, Is an exponential function representing the state Next Category and Characteristic function of the indicator The output value of handles null values, abnormal values ​​and format conflicts. The final standardized data structure is defined as: , in Representative time point The standardized data matrix under Indicates a time point Next 1st project Category No. The weighted value of the indicator, Indicates a time point Next 2nd project Category No. The weighted value of the indicator, Indicates a time point Next Project No. Category No. The weighted value of the indicator, The dimension of the data matrix is , is the number of items is the number of indicators.

4. A power grid project cost rationality analysis system according to claim 3, characterized in that: The operation process of the dynamic analysis module includes: From the standard data tensor In [1], we extract the individual progress indicators of the power grid project and construct a multi-dimensional time series, expressing the formula: ,in It's in time The data set at a moment, It's in time Moment The first project The observed values ​​of the progress indicators, The index value range of the project is from 1 to There are a total of projects, The index value range of the progress indicator is , Indicates that the time span of the data is from the first moment to the moment, and for each individual progress indicator , set dynamic threshold , expression formula: ,in is the dynamic threshold of the indicator, is the threshold model, is the history state encoder, The engineering unit at time The target control curve encoding, Its dynamic uncertainty is modeled by a Gaussian process, It follows a normal distribution. is the mean value representing the central expected value of the dynamic threshold, It is Project No. The historical behavior characteristic encoding of each indicator.

5. A power grid project cost rationality analysis system according to claim 4, characterized in that: The operation process of the dynamic analysis module includes: The current actual single progress indicator value Compared with the dynamic threshold interval, the residual tensor is constructed and expressed as: ,in It's in time No. The first project The standardized residual value of each indicator, It's in time Moment The first project The observed values ​​of the progress indicators, is the mean value representing the central expected value of the dynamic threshold, Its dynamic uncertainty is modeled by a Gaussian process, It's time All items All indicators The two-dimensional tensor composed of the standardized residuals of , uses graph convolution-variational autoencoder to construct time series feature compression and abnormal probability estimation, expressed as: ,in Indicates the The first project Indicators in time The anomaly score, It is the Softmax function that maps the output to a probability distribution. It is the second half of the variational autoencoder (VAE) used to reconstruct the latent space vector into the output. represents a trained variational autoencoder model, is an anomaly detector, It is the first The indicators from time point 1 to A two-dimensional tensor of Represents vector concatenation operation, It is a project The structured static information vector of When the current engineering unit is determined No. Progress indicators in time Abnormal, locate abnormal data Record to the exception log, and collect the exception data to drive the early warning system and reversely optimize the threshold model With anomaly detector training process.

6. A power grid project cost rationality analysis system according to claim 5, characterized in that: The operation process of the cost map module includes: Through the standard data tensor Define entity set and express formula: ,in is a collection of entities, It is a project engineering node. It is a component (equipment), It is raw materials and procurement items, It's the construction process. The external environment (region, season, price) uses a multi-modal embedding and relationship extraction model to embed the influencing factors in the structured data and text into a joint semantic space. The expression formula is: ,in is an entity The vector representation of Indicates that the entity is passed through an encoder Encoded into a vector, is the set of encoder parameters, is an entity and The vector representation of Representing an entity and The type of predicted relationship between It is based on the relationship between two entity vectors to infer their relationship type. It is The weight matrix corresponding to the relationship type, is an entity and Through the The matching score of the class relationship, is the activation function used to normalize the score into a probability form, Is to find the relationship type index with the largest score , output the structured semantic relationship labels between entity pairs ,in It is a relationship set. When constructing the cost influencing factor map, all the entity-relationship-entity pairs obtained are used to form a map triple set, which is expressed as follows: ,in is a set of triples, Is a triple representing the entity To Entity There is a relationship between , Represents the head entity in the triple and tail entity All belong to entity collection , Representing an entity and The relationship between Belongs to a relation set , each node Incidental attribute vector , forming a heterogeneous graph with attributes, expressed as: ,in It represents a graph structure used to organize and express cost maps related to engineering costs. is a collection of entities, is a set of relations, is a set of triples, is a collection of attributes, is an entity The attribute vector of For all entities in the graph Extract the attribute vector to form the attribute set .

7. A power grid project cost rationality analysis system according to claim 6, characterized in that: The operation process of the cost map module includes: The propagation factor path integral model is introduced to characterize the indirect impact of different nodes on the total cost and calculate the cost from the source node. To the cost aggregation node The causal path contribution of is expressed as: ,in Indicates from arrive The set of all paths of is the edge weight propagation coefficient, is the causal weight of the relationship type, is the attention correlation function between attribute vectors, It is from arrive All valid reasoning paths Perform the summation, It is the path All triples (entity-relationship-entity) on are multiplied in sequence. Indicates the entity in the path Through the relationship To Entity The graph triples of is an entity The attribute feature vector of .

8. A power grid project cost rationality analysis system according to claim 7, characterized in that: The operation process of the cost map module includes: By using heterogeneous graph convolution and structure-preserving objective functions to calculate node embeddings, subsequent predictions and visualizations are performed, expressed as: ,in is an entity The final embedding vector of It is a neural network used by hypergraph neural network to learn entity vectors from graph structure. It is a cost map. is the graph structure preservation loss function, It is to traverse the triple relationship in all graphs, is the TransE model structure, is the relationship vector, is the weight hyperparameter, is an entity The original attribute vector of is the predicted embedding obtained by inputting entity attributes into a multi-layer perceptron (MLP) neural network. The attribute preservation loss represents the distance between the vector of the current entity and the embedding based on attribute prediction. The final output format of the graph is, ,in It is a cost map. is the set of vector embeddings of all entities, is the graph inference function.

9. A power grid project cost rationality analysis system according to claim 8, characterized in that: The operation process of the abnormality analysis module includes: Assume that the input abnormal data is a set of event sequences, and express the formula: ,in represents a set of abnormal data samples, Indicates the Abnormal samples, is the total number of abnormal samples, Indicates the The data ranges from 1 to Perform a collection traversal. Indicates the The numerical feature vector of abnormal data, Indicates the The timestamp of the data. The context label representing the anomaly. The anomaly data is mapped to the entity set in the graph through a soft semantic matching function. and triple sets , the unstructured data is associated with the structured cost data, and the unstructured text set is Each text is aligned with the deep semantic encoder and graph entity embedding to finally select a group of strongly related texts Mapping to abnormal nodes Surrounding influence area .

10. A power grid project cost rationality analysis system according to claim 9, characterized in that: The operation process of the abnormality analysis module includes: Based on the heterogeneous graph, a differentiable causal graph reasoning network is constructed to infer the root cause path, and the expression formula is: ,in is the relational reasoning matrix, is the structural path causal influence factor, is the indicator function marking the true root cause, Yes arrive The overall loss function of all abnormal data Sum the values ​​of It is The attribution loss value of abnormal data, For all Candidate causal nodes of anomalies Take the negative logarithm after doing the weighted sum, Is a node in the graph The embedding vector representation of It is abnormal data The embedded representation after mapping to the graph node, is the node similarity score, It is to perform softmax normalization on the total scores of all candidate nodes. is the sum representing the probability value, Represents a causal node For abnormal events The attribution score, For all slave nodes To abnormal node A collection of paths Sum, In each path Perform product operations on each edge in the path, is a graph triple on the path, Represents an edge The confidence level, is the distance decay factor, is the comprehensive confidence score of the attribution path. The automatic generation of the attribution report is based on the templated report generation model, expressed as follows: ,in It is Anomaly attribution reports, is the report generation function, It is the The abnormal data is mapped to the entity node in the graph. is the set of all causal entity nodes in the attribution path, Is related to each causal node A collection of matched unstructured text summaries, It is the comprehensive confidence score of the attribution path. The output includes anomaly location, root cause chain, evidence fragment, and confidence score.

Citation Information

Patent Citations

  • Power grid project manufacturing cost analysis data collecting system and method

    CN104077665A

Cited By

  • Project cost auditing method and system based on multi-source data fusion

    CN121279966A