Power grid project investment financial digital intelligent penetrating management and control system and method

By constructing a digital and intelligent penetrating control system for power grid project investment and finance, the problems of fragmented multi-source data and lagging data analysis in the power industry have been solved, realizing automated control and real-time risk warning throughout the entire process, and improving management efficiency and scientific decision-making.

CN122492384APending Publication Date: 2026-07-31STATE GRID GANSU ELECTRIC POWER CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID GANSU ELECTRIC POWER CORP
Filing Date
2026-05-19
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

The existing financial management model in the power industry suffers from problems such as fragmented multi-source data, outdated data analysis and presentation models, reliance on human experience for investment decisions, and lagging process processing and risk monitoring, making it impossible to achieve full-process, multi-dimensional digital intelligent control and data linkage analysis.

Method used

A digital and intelligent, penetrating control system for power grid project investment and finance is constructed. Through a data fusion layer, intelligent analysis layer, business application layer, and risk control layer, it realizes multi-source data fusion, intelligent analysis, visualization report generation, and full-process automated control. It adopts technologies such as pre-trained large language models, natural language to SQL intelligent query, association rule mining, dynamic cost prediction, and risk early warning algorithms.

Benefits of technology

It has achieved full integration of business and financial data, enhanced the ability to gain insights into all aspects of business, enabled scientific investment decisions and real-time risk warnings, improved management efficiency and risk control capabilities, and supported multi-level data penetration queries and visualization displays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492384A_ABST
    Figure CN122492384A_ABST
Patent Text Reader

Abstract

This invention provides a digital and intelligent penetrating control system and method for power grid project investment finance. The system includes: a data fusion layer, which connects multiple business systems to build a business-finance integrated data resource pool; an intelligent analysis layer, based on a pre-trained large language model, including a natural language to SQL intelligent query module; a business application layer, which performs data mining, intelligent document processing, and automatically generates visual reports; and a risk control layer, based on dynamic cost prediction and early warning algorithms, which constructs a closed-loop cost control system for the entire process, enabling multi-level penetrating queries. This system breaks down data silos from multiple sources, achieving full-domain integration and penetrating querying of business and financial data; innovates data analysis models to enhance business insight capabilities; constructs scientific investment decisions to achieve precise cost allocation; and realizes automated control and real-time risk early warning throughout the entire process, improving management efficiency and risk control levels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power technology, and more specifically, to a digital and transparent management system and method for financial control of power grid project investment. Background Technology

[0002] Currently, the power industry generally adopts a traditional digital management system for financial cost control and project investment management. Its core technology relies on a distributed basic information system. Specifically, business data such as finance, materials, engineering, assets, and marketing are stored separately in multiple independent business systems, including a financial middle platform, PMS 3.0, and Marketing 2.0. However, these systems lack standardized data and interconnectivity. Data exchange relies on manual cross-system collection, cleaning, and processing, resulting in low efficiency and a high risk of data errors. Data analysis primarily relies on traditional manual statistics and report summaries. The data presentation is simplistic and lacks intuitiveness, making it difficult to quickly grasp key business information and achieve layer-by-layer analysis from overall business overview to specific project details, and from the financial backend to the business frontend. Project investment decisions and cost resource allocation rely excessively on management personnel's industry experience, lacking data-supported scientific and intelligent judgment logic. Cost budget management can only perform post-event accounting, failing to achieve pre-event forecasting and real-time monitoring. Project review and document processing rely on manual screening and key information extraction, resulting in long review cycles and prominent issues of duplicate applications. Meanwhile, risk identification and early warning response are only at the hourly level, leading to problems of lagging monitoring and untimely risk handling.

[0003] It is evident that existing technologies operate under a model of manual dominance, decentralized management, and post-event control, failing to establish a comprehensive, multi-dimensional digital control and data linkage analysis mechanism, and thus unable to meet the core requirements of penetrating supervision and lean management.

[0004] To address the aforementioned issues, patent CN119809780A provides a blockchain-based system for transparent supervision of fiscal funds. This system constructs a blockchain network encompassing enterprises, banks, finance departments, and regulatory nodes, ensuring data is stored on the blockchain and is tamper-proof. It employs a self-attention supervision model combined with smart contracts for automatic approval, achieving full-process transparency, real-time traceability, and precise risk control of funds. However, the fragmented nature of multi-source data creates silos, resulting in low integration between business and finance, and fails to support the multi-dimensional, multi-indicator transparent analysis required by the power industry. This makes it unsuitable for the power industry's needs for transparent supervision and lean management. Patent CN121526811A discloses a bottom-level asset management system supporting batch transparent analysis. Through data acquisition → batch structuring → indicator monitoring → transparent display, it supports batch transparent analysis of massive amounts of underlying assets, with traceable data for each transaction and drill-down capabilities for indicators. However, this system does not support multi-dimensional intelligent analysis and alerts. Patent CN116150262A provides a penetrating monitoring system. This system adopts a distributed multi-node architecture and achieves real-time monitoring and traceability of funds, expenses, and financial data at the group level through the integration of multi-source financial and business data, hierarchical access control, full-link data penetration query, and anomaly warning. However, this system does not use a large-scale model for comprehensive analysis, making it unable to accurately target investments.

[0005] In summary, considering the actual business scenarios and development needs of the power industry, existing traditional control technologies have the following four core shortcomings, making them unable to meet the requirements of penetrating supervision and the industry's digital transformation needs: 1. Fragmented data from multiple sources creates isolated data silos, resulting in low integration between business and finance: The financial, asset, marketing, and engineering business systems are not integrated, data standards are not uniform, cost project data needs to be collected manually across systems, data connectivity and accuracy are poor, and full-chain data traceability and penetrating query cannot be achieved; 2. Outdated data analysis and presentation methods, and insufficient business insight capabilities: Traditional manual statistical analysis and report summarization methods are not intuitive enough, making it difficult to quickly grasp key business information and failing to achieve a visual display of input and output and multi-dimensional penetrating analysis. 3. Investment decisions and resource allocation are crude and rely heavily on human experience: The lack of intelligent judgment logic in cost budget management and project investment allocation, and the absence of correlation analysis between "cost input and business indicators", easily leads to inefficient input and resource misallocation, making it impossible to achieve precise investment; 4. Lagging process handling and risk monitoring result in low management efficiency: Project document screening and key information extraction rely on manual work, resulting in long review cycles and high rates of duplicate submissions; risk identification and early warning response are at the hourly level, making it impossible to achieve full-process automated monitoring and real-time control, and risk handling is not timely. Summary of the Invention

[0006] The purpose of this invention is to provide a digital and intelligent penetrating control system and method for power grid project investment finance. It solves the problem of multi-source data silos, realizes the full-domain integration and penetrating query of business and financial data, innovates the data analysis and presentation mode, and improves the ability to gain insights into the full dimensions of business; it constructs a scientific investment decision-making system, realizes precise allocation of cost resources, and can carry out full-process automated control and real-time risk warning, thereby improving management efficiency and risk control capabilities.

[0007] To achieve the above objectives, embodiments of the present invention provide a digital and intelligent penetrating control system for financial management of power grid project investment, the system comprising: The data fusion layer is designed to connect multiple core business systems, collect multi-source data, and build a business and finance integrated data resource pool after standardization processing. The intelligent analysis layer is designed to use a pre-trained large language model as a base to perform deep learning on the data in the business and finance integrated data resource pool, and includes at least a natural language to SQL intelligent query module to automatically convert natural language questions into structured query statements. The business application layer is configured to perform in-depth data correlation mining, intelligent processing of multi-source documents, and automatic generation of visual reports based on the results of the intelligent analysis layer. The risk management layer is designed to build a closed loop of cost management that spans the entire process from pre-event prediction, in-event monitoring, and post-event analysis, based on a dynamic cost prediction model and risk early warning algorithm, and enables multi-level penetrating queries.

[0008] Preferably, the pre-trained large language model in the intelligent analysis layer is fine-tuned using a hierarchical differential hybrid fine-tuning method with LoRA as the primary method and IA3 as the secondary method; wherein, According to formula (1), freeze all weights of the first 10 Transformer blocks of the pre-trained large language model to preserve general language capabilities: (1) in, Let W be all the trainable parameters of a certain layer of a neural network, including the weights W. ) and bias b( ), used to represent all unknown parameters that need to be learned in the model; This is the complete backbone weight set for a pre-trained large language model. The layer number of the Transformer layer. This represents the total number of Transformer layers in the pre-trained large language model. It represents the derivative of a multivariable function with respect to one of its variables; For trainable Transformer layers of layer 11 and above, LoRA low-rank adaptation modules are deployed in the projection layers of multi-head attention, and IA3 element-wise scaling modules are embedded in the projection layers of the feedforward network. According to formula (2), the backbone weights of the pre-trained large language model are trained using the standard cross-entropy loss function, and only the adapter parameters of LoRA and IA3 are updated: (2) in, It is a loss function whose value is a non-negative real number, used to measure the degree of prediction error of the model under the current parameters; N The total number of samples in a training batch; T How many tokens each text is split into; (...) represents the probability distribution of the next token output by the large language model; Let be the next token predicted by the model at position t. This is the actual next token, which is the (t+1)th token that actually exists in the training data; The parameters of the backbone model are completely frozen; These are parameters for the trainable Adapter module.

[0009] Preferably, the LoRA low-rank adaptation module achieves incremental adaptation through low-rank decomposition, and the transformation formula is as follows: , in, To output a tensor; For input tensors; Main path weight matrix; This is the transpose of the main path weight; For input application Regularization: During training, some elements are randomly set to 0, and scaling is performed to keep the expected value unchanged during testing; For low-rank fitting matrix ; For low-rank fitting matrix ; The rank of a low-rank fit; This is the scaling factor; The IA3 element-wise scaling module achieves precise adaptation by adjusting the activation value. The transformation formula is as follows: , in, For Hadama accumulation.

[0010] Preferably, the natural language to SQL intelligent query module includes: A vector knowledge base is used to vectorize and store SQL sample data S, data table structure definitions D, and business domain knowledge K; where the i-th training data is denoted as a triple. ,in, For the question text, For SQL statements, Metadata identifier; during the training data vectorization process, according to formula (3) through the embedding model Map text to 3D vector space: (3) The semantic similarity retrieval unit is used to calculate the user question vector. The cosine similarity with vectors in the knowledge base is calculated, and the top-K similar data are returned as context; where the similarity is calculated using cosine similarity according to formula (4): , (4) in, Candidate vectors; For query vector; It is the dot product of vectors; for The L2 norm; for The L2 norm; The SQL generation inference unit is used to assemble the user question, the retrieved context, and the data table structure definition into a prompt text based on the retrieval enhancement generation paradigm, and input it into the large language model to generate SQL statements; wherein, according to formula (5), the SQL generation process is formalized into conditional probability generation: (5) in, For natural language questions input by the user, For the context of the retrieved similar examples, Define the structure of the relevant data tables; Multi-tenant data isolation units are used for data filtering based on tenant identifiers during vector retrieval and SQL execution phases; each tenant is assigned a unique identifier. Both training data and query execution are filtered based on tenant identifiers. The data filtering rules are expressed as follows: , in, For the original dataset DA subset containing only the data that the current tenant has the right to access; {...} is used to define a set in which all elements satisfy the conditions specified in the curly braces; d For the original dataset D Each individual element in; In a multi-tenant system, each data record will have a... This field identifies which tenant this data belongs to; A unique identifier for the current tenant; The training data incremental update unit is used to dynamically and incrementally update the training data, including SQL samples, data table structure definitions, and business documents; among them, for document-type data, file hashing is used to avoid duplicate training of the same content according to formula (6): (6) in, This is a collection of hash values ​​for the stored files.

[0011] Preferably, the data association deep mining performed at the business application layer includes: The association rule mining unit is used to extract association rules between cost input items and business indicator items from business data using the Apriori algorithm and / or FP-Growth algorithm, and to evaluate the strength of the rules through support, confidence, and lift. Here, let the dataset be... Include n Each transaction record T A set of business metrics; association rules X → Y The strength of [the measure] is measured by three metrics: support, confidence, and lift.

[0012]

[0013]

[0014] in, For cost input items, For a set of business metrics; when the degree of improvement A value greater than 1 indicates cost input. With business metrics There is a positive correlation; when A value less than 1 indicates a negative correlation. The input-output efficiency assessment unit is used to build an assessment model, quantify the overall efficiency score of each business unit, and identify units with high input and low output, as well as units with low input and high output; where, let the first... i The input vector of each business unit is =( , ,..., The output vector is =( , ,..., If the overall efficiency score of the business unit is: , in, For the first j The weighting coefficients of output indicators, For the first l Weighting coefficients of input factors; based on efficiency scores The system automatically identifies < High input, low output units and > Low-input, high-output units provide a quantitative basis for optimizing resource allocation; The spatial clustering analysis unit is used to perform spatial clustering analysis on transformer substations using the DBSCAN density clustering algorithm in conjunction with a geographic information system, identifying areas with abnormal regional operation and maintenance efficiency; where a transformer substation is defined. of Neighborhood is defined as When | When |≥MinPts, the area The core object is marked, and all the substations in its neighborhood form a cluster.

[0015] Preferably, when the business application layer performs intelligent multi-source document processing and automatic generation of visual reports, it specifically includes: The OCR recognition unit is used to perform text recognition on document images using an architecture combining convolutional neural networks (CNNs) and recurrent neural networks (RNNs); where the input image is... I Through feature extraction network F Feature maps are extracted and then processed through a sequence modeling network. S The process of generating and recognizing text sequences is as follows: ; The information extraction unit is used to extract key structured information from the recognized text using Named Entity Recognition (NER) technology; wherein, the document text is denoted as... T The goal of named entity recognition is to find the optimal labeled sequence. , ; The natural language parsing module is used to perform word segmentation, part-of-speech tagging, syntactic analysis, and semantic role labeling on document content based on a pre-trained language model; among them, the document semantic representation aggregates contextual information through an attention mechanism. , in, This is the attention output at position i; Attention(Q,K,V) is the scaled dot product attention function. Q For query matrix; K The key matrix; V It is a value matrix; for Q and K The dot product matrix; This is the scaling factor; d for Q and K The dimension; The report generation engine uses a large language model to automatically generate professional reports, including charts and graphs, based on report template instructions, data analysis results, and visualization chart configurations. The report generation process is formalized as a conditional generation task. , in, For report template instructions, For the data analysis results, Configure visualization charts.

[0016] Preferably, the risk management layer includes: A dynamic cost forecasting model is used to predict project costs throughout their entire lifecycle based on time series analysis and machine learning regression models; whereby, assuming the project cost is in the [missing information] phase... t The actual cost of a time is The predicted cost is calculated using a multi-factor regression model based on formula (7): (7) in, ..., Let there be n variables affecting the cost. , ,..., The regression coefficients for each factor are denoted as . This is the random error term; Anomaly detection unit is used to identify anomalous data points and trigger risk warnings based on the anomaly score calculation formula using the isolated forest algorithm. The knowledge graph reasoning engine is used to perform knowledge reasoning based on a knowledge graph that stores industry rules and entity relationships in the form of triples, and to achieve compliance checks. The triples are represented as follows: , in, For the head entity, For the relationship, For tail entities, For a collection of entities, For a set of relations; The inference engine performs knowledge reasoning based on graph neural networks (GNN) and updates entity embeddings according to formula (9): (9) in, N ( i ) is an entity i The set of neighboring nodes, and Here, σ is a learnable parameter, and σ is the activation function. The multi-level penetration query unit is used to respond to user operations, penetrating from the current level of data to the next level of detailed data, down to the underlying business details; where penetration query is from the first... l layer to the first l Drilling operation at +1 level: , in, The core operation of multi-level penetration query is to input the target data node at the current level and output all direct detailed data of its next level. For the first The current data node of the layer; For the first Detailed data nodes of the layer; To store the first In the layer, all with Detailed data nodes with parent-child relationships serve as the data source for drill-down operations.

[0017] On the other hand, the present invention provides a method for intelligent and penetrating financial management and control of power grid project investment based on the above-mentioned intelligent and penetrating financial management and control system for power grid project investment, the method comprising: S1: Data integration, connecting multiple core business systems and collecting multi-source data, and constructing a business and finance integrated data resource pool after standardized processing; S2: Intelligent analysis, based on a pre-trained large language model, performs deep learning on the data in the business and finance integrated data resource pool, and responds to natural language questions by automatically converting it into SQL statements for querying; S3: Business Applications. Based on intelligent analysis results, perform in-depth data correlation mining to identify resource configuration optimization links, and respond to instructions to intelligently process multi-source documents and automatically generate visual reports. S4: Risk Management, based on dynamic cost prediction models and risk early warning algorithms, performs full-cycle prediction and anomaly detection of project costs, and provides multi-level penetrating queries.

[0018] Preferably, the data association deep mining in S3 includes: The Apriori algorithm and / or FP-Growth algorithm are used to extract the association rules between cost input items and business indicator items; Construct an input-output efficiency evaluation model to quantify the overall efficiency score of each business unit; By combining geographic information systems, the DBSCAN density clustering algorithm is used to perform spatial clustering analysis on the transformer substations to identify areas with abnormal efficiency.

[0019] Preferably, the risk warning algorithm in S4 adopts the Isolation Forest algorithm, assuming the dataset... X Include m One sample, each sample x have d Dimensional features, calculate the anomaly score according to formula (8): (8) in, For the sample x Average path length in an isolated tree As a normalization factor, when outlier scores When the value approaches 1, it is identified as an abnormal data point, triggering a risk warning.

[0020] Through the aforementioned technical solutions, this invention connects multi-source business systems, constructs a standardized business and financial integrated data resource pool, solves the data silo problem, and achieves traceability and penetrating query of data across the entire chain. Secondly, based on a large model, data querying and analysis can be achieved through natural language interaction, and combined with association rule mining, it provides data-driven scientific basis for investment decisions, realizing the transformation from "experience-driven" to "data-driven." Simultaneously, the integration of OCR and NLP technologies enables automated document processing, significantly improving review efficiency. Furthermore, based on dynamic prediction models and anomaly detection algorithms, it achieves minute-level risk warnings and closed-loop management throughout the entire process, significantly improving management efficiency and risk control capabilities. In addition, it can automatically generate visual reports and support multi-level data penetration from provincial companies to grassroots work teams, meeting the requirements of penetrating supervision and achieving precise investment control.

[0021] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0022] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1This is a flowchart of the intelligent and penetrating control method for financial investment in power grid projects provided by the present invention; Figure 2 This is the overall architecture diagram of the intelligent and penetrating control system for financial investment in power grid projects provided by the present invention. Detailed Implementation

[0023] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0024] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0025] See Figure 2 This invention provides a digital and intelligent, penetrating control system for power grid project investment finance. The system includes a data fusion layer, designed to connect multiple core business systems, used to collect multi-source data and construct a business-finance integrated data resource pool after standardized processing. By building a unified data acquisition and fusion underlying architecture and deploying standardized data interface, it connects more than 10 core business systems in the power industry, such as the financial middle platform, PMS3.0, and Marketing 2.0. It comprehensively collects multi-dimensional data including financial revenue and expenditure, project investment, asset operation and maintenance, marketing services, geographic information, and ledger data. Through three standardized preprocessing steps—data cleaning, format unification, and tag classification—redundant, erroneous, and duplicate data are eliminated, constructing a structured and standardized business-finance integrated data resource pool covering more than 20 million data points. Simultaneously, a data lifecycle management mechanism is established to achieve real-time synchronization and dynamic updates of data across various business systems, completely abandoning the inefficient model of manual cross-system data collection and organization. This provides a complete, accurate, and unified data foundation for subsequent intelligent analysis and penetrating queries, achieving traceability, correlation, and penetration of the entire data chain.

[0026] The system also includes an intelligent analysis layer, configured to use a pre-trained large language model as a foundation for deep learning on data from the business-finance integrated data resource pool. It includes at least a natural language to SQL intelligent query module to automatically convert natural language questions into structured query statements. Thus, with the large model (Guangming Power's large model) as the core foundation, the massive business and financial data from the business-finance integrated data resource pool are used as training samples to feed artificial intelligence for deep learning, repeatedly training the agent to understand the financial management logic, cost budgeting rules, and project investment judgment standards of the power industry. Simultaneously, addressing the technical problems of existing large language model fine-tuning techniques, such as high computational cost of full-parameter fine-tuning, insufficient generalization ability of single low-rank adaptation methods, and the tendency for professional domain fine-tuning to lead to degradation of the model's general language capabilities, this invention proposes a hierarchical differentiated hybrid fine-tuning method based on LoRA and supplemented by IA3, specifically including: First, according to formula (1), freeze all weights of the first 10 Transformer blocks of the pre-trained large language model to fully preserve the model's basic general language capabilities and avoid degradation of general capabilities caused by professional fine-tuning: (1) in, Let W be all the trainable parameters of a certain layer of a neural network, including the weights W. ) and bias b( ), used to represent all unknown parameters that need to be learned in the model; This is the complete backbone weight set for a pre-trained large language model. The layer number of the Transformer layer. This represents the total number of Transformer layers in the pre-trained large language model. It represents the derivative of a multivariable function with respect to one of its variables; Secondly, for trainable Transformer layers 11 and above, a differentiated module configuration is adopted: LoRA low-rank adaptation modules are deployed in the four projection layers (q_proj, k_proj, v_proj, o_proj) of the multi-head attention network, and IA3 element-wise scaling modules are embedded in the gated projection layer (gate_proj) and up projection layer (up_proj) of the feedforward network. The LoRA low-rank adaptation module achieves incremental adaptation through low-rank decomposition, with the preferred configuration being rank r=8 and scaling factor α=32. The transformation formula is: , in, To output a tensor; For input tensors; Main path weight matrix; This is the transpose of the main path weight; For input application Regularization: During training, some elements are randomly set to 0, and scaling is performed to keep the expected value unchanged during testing; For low-rank fitting matrix ; For low-rank fitting matrix ; The rank of a low-rank fit; This is the scaling factor; The IA3 element-wise scaling module achieves precise adaptation by adjusting the activation value. The transformation formula is as follows: , in, For Hadama accumulation.

[0027] Next, according to formula (2), the backbone weights of the pre-trained large language model are trained using the standard cross-entropy loss function, and only the adapter parameters of LoRA and IA3 are updated: (2) in, It is a loss function whose value is a non-negative real number, used to measure the degree of prediction error of the model under the current parameters; N The total number of samples in a training batch; T How many tokens each text is split into; (...) represents the probability distribution of the next token output by the large language model; Let be the next token predicted by the model at position t. This is the actual next token, which is the (t+1)th token that actually exists in the training data; The parameters of the backbone model are completely frozen; These are the parameters for the trainable Adapter module. This strategy significantly reduces the computational and storage overhead of fine-tuning while maintaining both domain-specific accuracy and the model's native general-purpose capabilities.

[0028] In this implementation, a natural language to SQL intelligent query module based on retrieval enhancement was constructed for the financial data query scenario in the power industry. This module consists of four core components: vector knowledge base construction, semantic similarity retrieval, SQL generation and reasoning, and multi-tenant data isolation. It achieves automatic conversion from user natural language questions to structured query statements. The vector knowledge base is used to vectorize and store SQL sample data S (containing historical query questions and corresponding SQL statement pairs), data table structure definitions D (DDL statements and field semantic descriptions), and business domain knowledge K (financial management systems, indicator calculation rules, and explanations of business terms); where the i-th training data is denoted as a triple. ,in, For the question text, For SQL statements, Metadata identifier; during the training data vectorization process, according to formula (3) through the embedding model Map text to 3D vector space: (3) The semantic similarity retrieval unit is used to calculate the user question vector. The cosine similarity with vectors in the knowledge base is calculated, and the top-K similar data are returned as context; where the similarity is calculated using cosine similarity according to formula (4): , (4) in, Candidate vectors; For query vector; This is the dot product (inner product) of vectors. for L2 norm (Euclidean modulus); for The L2 norm (Euclidean modulus) is used; the search results are sorted from high to low similarity, and the K most similar data are selected to construct the prompt context, providing domain knowledge support for subsequent SQL generation.

[0029] The SQL generation inference unit is used to assemble the user question, the retrieved context, and the data table structure definition into a prompt text based on the Retrieval Enhanced Generation (RAG) paradigm, and input it into the large language model to generate SQL statements; wherein, according to formula (5), the SQL generation process is formalized into conditional probability generation: (5) in, For natural language questions input by the user, For the context of the retrieved similar examples, Define the structure of the relevant data tables; based on the above context information, the large language model generates SQL statements that are semantically correct and syntactically correct.

[0030] Multi-tenant data isolation units are used for data filtering based on tenant identifiers during vector retrieval and SQL execution phases; each tenant is assigned a unique identifier. Both training data and query execution are filtered based on tenant identifiers. The data filtering rules are expressed as follows: , in, For the original dataset D A subset containing only the data that the current tenant has the right to access; {...} is used to define a set in which all elements satisfy the conditions specified in the curly braces; d For the original dataset D Each individual element in; In a multi-tenant system, each data record will have a... This field identifies which tenant this data belongs to; This serves as a unique identifier for the current tenant. The tenant data isolation mechanism ensures that each tenant can only access data within its authorized scope. During the vector retrieval phase, it automatically filters training data that is not from the current tenant, and during the SQL execution phase, it automatically appends tenant filtering conditions, thus achieving secure isolation and access control for data from multiple units.

[0031] The training data incremental update unit is used to dynamically and incrementally update the training data, including SQL samples, data table structure definitions, and business documents; among them, for document-type data, file hashing is used to avoid duplicate training of the same content according to formula (6): (6) in, This is a collection of hash values ​​from stored files. Through this incremental update mechanism of training data, the system can continuously accumulate domain knowledge during operation, constantly improve the accuracy of SQL generation, and adapt to the evolving business scenarios.

[0032] Furthermore, to enhance the user experience, the system can employ Server Send Events (SSE) technology to achieve streaming responses, pushing query results and analysis processes to the front end in real time. Users can instantly view the SQL generation process, data retrieval progress, and analysis results without waiting for the entire process to complete. This mechanism significantly reduces the first-byte response time and improves user satisfaction in complex query scenarios.

[0033] To address the shortcomings of existing technologies, such as the lack of correlation analysis between "cost input" and "business indicators" and the extensive nature of resource allocation, this invention, based on a business-finance integrated data resource pool, integrates multi-source data including geographic information, equipment ledgers, operation and maintenance costs, and project investments in the power industry. Utilizing association rule mining technology, it creates a comprehensive "asset map" and "resource allocation map." This includes integrating geographic information, ledger data, and operation and maintenance costs for all transmission lines, substations, and public transformer areas across the province. Asset information for each regional company is statistically analyzed by voltage level to accurately mine the inherent correlation between "cost input" and "business indicators." This allows for the rapid identification of high-input, low-output and low-input, high-output business processes, precisely locating "key improvement areas" and "inefficient operation and maintenance units." This provides data-driven and visualized analytical basis for resource allocation, promoting concentrated resource investment in key areas and efficient processes, avoiding inefficient investment, and achieving comprehensive asset management and precise resource allocation. Specifically, the system also includes a business application layer, configured to perform deep data correlation mining, intelligent multi-source document processing, and automatic generation of visual reports based on the results of the intelligent analysis layer. The deep data correlation mining performed by the business application layer includes: The association rule mining unit uses the Apriori and FP-Growth algorithms to extract association rules between cost input items and business indicator items from business data, and evaluates the strength of the rules through support, confidence, and lift. The dataset is defined as follows: Include n Each transaction record T A set of business metrics; association rules X → Y The strength of [the measure] is measured by three metrics: support, confidence, and lift.

[0034]

[0035]

[0036] in, For cost input items, For a set of business metrics; when the degree of improvement A value greater than 1 indicates cost input. With business metrics There is a positive correlation; when A value less than 1 indicates a negative correlation. In one specific implementation, the system sets a minimum support threshold min. sup =0.05, minimum confidence threshold min conf =0.6, filtering out strong association rules with business value.

[0037] The input-output efficiency assessment unit is used to build an assessment model, quantify the overall efficiency score of each business unit, and identify units with high input and low output, as well as units with low input and high output; where, let the first... i The input vector of each business unit is =( , ,..., The output vector is =( , ,..., If the overall efficiency score of the business unit is: , in, For the first j The weighting coefficients of output indicators, For the first l Weighting coefficients of input factors; based on efficiency scores The system automatically identifies < High input, low output units and > Low-input, high-output units provide a quantitative basis for optimizing resource allocation; The spatial clustering analysis unit is used to perform spatial clustering analysis on transformer substations using the DBSCAN density clustering algorithm in conjunction with a geographic information system, identifying areas with abnormal regional operation and maintenance efficiency; where a transformer substation is defined. of Neighborhood is defined as When | When |≥MinPts, the area The core object is marked, and all its neighboring transformer substations form a cluster. Through cluster analysis, abnormal areas with significantly lower operation and maintenance efficiency than similar surrounding transformer substations can be identified as candidate objects for "key improvement transformer substations".

[0038] Simultaneously, when the business application layer performs intelligent multi-source document processing and automatic generation of visual reports, it integrates natural language processing technology and OCR optical character recognition technology to build an intelligent document processing module. This module intelligently filters various multi-source documents during project application and review processes, automatically extracting key information and replacing manual document processing, significantly improving the efficiency and accuracy of project reviews. Furthermore, based on large-scale models and real-time data analysis results, an intelligent report generation engine is built, automatically generating professional investment performance reports containing visual charts such as heatmaps, bar charts, and line charts. These reports intuitively display the input and output of various business operations, supporting Word / PDF multi-format output and a streaming interactive experience. This visualizes and intuitively presents data, helping users quickly grasp key business information and gain a comprehensive understanding of business operations, from an overall business overview to specific project details. Specifically, this includes: The OCR recognition unit is used to perform text recognition on document images using an architecture combining convolutional neural networks (CNNs) and recurrent neural networks (RNNs); where the input image is... I Through feature extraction network F Feature maps are extracted and then processed through a sequence modeling network. S The process of generating and recognizing text sequences is as follows: ; The information extraction unit is used to extract key structured information from the recognized text using Named Entity Recognition (NER) technology; wherein, the document text is denoted as... T The goal of named entity recognition is to find the optimal labeled sequence. , ; The natural language parsing module is used to perform word segmentation, part-of-speech tagging, syntactic analysis, and semantic role labeling on document content based on a pre-trained language model; among them, the document semantic representation aggregates contextual information through an attention mechanism. , in, This is the attention output at position i; Attention(Q,K,V) is the scaled dot product attention function. Q For query matrix; K The key matrix; V It is a value matrix; for Q and K The dot product matrix; This is the scaling factor; d for Q and K The dimension; The report generation engine uses a large language model to automatically generate professional reports, including charts and graphs, based on report template instructions, data analysis results, and visualization chart configurations. The report generation process is formalized as a conditional generation task. , in, For report template instructions, For the data analysis results, Configure visualization charts. The system supports various chart types such as heatmaps, bar charts, and line charts, and automatically selects the optimal visualization scheme based on data characteristics.

[0039] In this embodiment, the system also includes a risk control layer, configured to construct a closed-loop cost control system covering the entire process from pre-event prediction, in-event monitoring, and post-event analysis based on a dynamic cost prediction model and risk early warning algorithm, and to enable multi-level, penetrating queries. Specifically, this includes: A dynamic cost forecasting model is used to predict project costs throughout their entire lifecycle based on time series analysis and machine learning regression models; whereby, assuming the project cost is in the [missing information] phase... t The actual cost of a time is The predicted cost is calculated using a multi-factor regression model based on formula (7): (7) in, ..., Let there be n variables affecting the cost. , ,..., The regression coefficients for each factor are denoted as . This is a random error term; the model parameters are obtained through training on historical data and are updated regularly to adapt to market changes.

[0040] The anomaly detection unit is used to identify anomalous data points based on the anomaly score calculation formula using the Isolation Forest algorithm. Let the dataset be... X Include m One sample, each sample x have d Dimensional features, calculate the anomaly score according to formula (8): (8) in, For the sample x Average path length in an isolated tree As a normalization factor, when outlier scores When the value approaches 1, it is identified as an abnormal data point, triggering a risk warning.

[0041] The knowledge graph reasoning engine is used to perform knowledge reasoning based on a knowledge graph that stores industry rules and entity relationships in the form of triples, and to achieve compliance checks. The triples are represented as follows: , in, For the head entity, For the relationship, For tail entities, For a collection of entities, For a set of relations; The inference engine performs knowledge reasoning based on graph neural networks (GNN) and updates entity embeddings according to formula (9): (9) in, N ( i ) is an entity i The set of neighboring nodes, and σ is a learnable parameter and σ is an activation function; by aggregating neighbor information through a multi-layer graph neural network, automatic reasoning and compliance checks of complex business rules are achieved.

[0042] The multi-level penetration query unit is used to respond to user operations, penetrating from the current level of data to the next level of detailed data, down to the underlying business details; where penetration query is from the first... l layer to the first l Drilling operation at +1 level: , in, The core operation of multi-level penetration query is to input the target data node at the current level and output all direct detailed data of its next level. For the first The current data node of the layer; For the first Detailed data nodes of the layer; To store the first In the layer, all with Detailed data nodes with parent-child relationships serve as the data source for drill-down operations. By recursively calling drill-down operations, multi-level penetration can be achieved from the provincial company to the municipal company, county company, power supply station, and work team, directly reaching the original business data, enabling precise problem location and closed-loop rectification.

[0043] Thus, the aforementioned intelligent risk warning and penetrating supervision system has constructed a closed-loop cost control system covering the entire process from "pre-event prediction to in-event monitoring to post-event analysis." Relying on the power industry cost control knowledge graph (accumulating over 2000 industry rules, such as technical upgrade project approval processes and equipment selection standards) and an intelligent decision engine, it automatically generates project input-output analysis reports, improving review efficiency by 60%. Combining historical project data with real-time market information (such as equipment prices and labor costs), it builds a dynamic cost prediction model and risk warning algorithm, achieving a prediction accuracy of 92%. Through association rule mining and anomaly detection technology, it automatically identifies investment risks such as budget overruns, non-compliant expenditures, and duplicate applications, reducing warning response time from hours to minutes, achieving 100% automated monitoring. Simultaneously, it improves the "indicator monitoring - multi-dimensional penetration - resource allocation - closed-loop management" mechanism, creating a full-link governance system of "precise investment - intelligent management - closed-loop assessment - supervision and evaluation." This supports multi-level problem location and rectification from the financial backend to the business frontend, and from provincial companies to grassroots teams, achieving real-time control down to specific data details and underlying business details.

[0044] On the other hand, such as Figure 1 As shown, this invention provides a method for intelligent and penetrating control of financial data in power grid project investment based on the aforementioned power grid project investment financial data-driven control system. The method includes: S1: Data integration, connecting multiple core business systems and collecting multi-source data, and constructing a business and finance integrated data resource pool after standardized processing; S2: Intelligent analysis, based on a pre-trained large language model, performs deep learning on the data in the business and finance integrated data resource pool, and responds to natural language questions by automatically converting it into SQL statements for querying; S3: Business Applications. Based on intelligent analysis results, perform in-depth data correlation mining to identify resource configuration optimization links, and respond to instructions to intelligently process multi-source documents and automatically generate visual reports. S4: Risk Management, based on dynamic cost prediction models and risk early warning algorithms, performs full-cycle prediction and anomaly detection of project costs, and provides multi-level penetrating queries.

[0045] Based on the above methods, traditional professional scenario barriers are broken down, and an intelligent scenario system covering the entire process of cost projects—"feasibility study – review – reserve – execution – evaluation"—is constructed. Differentiated judgment logics are designed for professional scenarios such as document review and output analysis to adapt to the business characteristics of different scenarios. In the full-process management system, the feasibility study stage achieves preliminary feasibility assessment by automatically verifying key information such as feasibility study reports and project estimates; the review stage abandons a single review standard and sets core review dimensions based on project type differences. For example, power transmission projects focus on unit transmission cost and line utilization rate, while distribution network renovation projects focus on the potential for reducing line loss rate and improving power supply reliability; the reserve stage prioritizes projects based on project review scores and the urgency of regional needs through automated algorithms, providing a basis for the subsequent fund disbursement order; the execution stage monitors the progress of project cost expenditures in real time, automatically triggering an early warning mechanism when abnormal deviations occur; the evaluation stage forms a closed-loop management by comparing budgeted values ​​with actual costs and expected goals with actual results, driving continuous upgrades in cost management.

[0046] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0047] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0048] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0049] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0050] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0051] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0052] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0053] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0054] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A digital and intelligent penetrating control system for financial investment in power grid projects, characterized in that, The system includes: The data fusion layer is designed to connect multiple core business systems, collect multi-source data, and build a business and finance integrated data resource pool after standardization processing. The intelligent analysis layer is configured to perform deep learning on the data in the business and finance integrated data resource pool based on a pre-trained large language model, and includes at least a natural language to SQL intelligent query module for automatically converting natural language questions into structured query statements. The business application layer is configured to perform in-depth data correlation mining, intelligent processing of multi-source documents, and automatic generation of visual reports based on the results of the intelligent analysis layer. The risk management layer is designed to build a closed loop of cost management that spans the entire process from pre-event prediction, in-event monitoring, and post-event analysis, based on a dynamic cost prediction model and risk early warning algorithm, and enables multi-level penetrating queries.

2. The power grid project investment financial digitalization and penetration control system according to claim 1, characterized in that, The pre-trained large language model in the intelligent analysis layer is fine-tuned using a hierarchical differential hybrid fine-tuning method, primarily based on LoRA and secondarily on IA3; among which... According to formula (1), all weights of the first 10 Transformer blocks of the pre-trained large language model are frozen to preserve general language capabilities: ,(1) in, Let W be all the trainable parameters of a certain layer of a neural network, including the weights W. ) and bias b( ), used to represent all unknown parameters that need to be learned in the model; This is the complete backbone weight set for a pre-trained large language model. The layer number of the Transformer layer. This represents the total number of Transformer layers in the pre-trained large language model. It represents the derivative of a multivariable function with respect to one of its variables; For trainable Transformer layers of layer 11 and above, LoRA low-rank adaptation modules are deployed in the projection layers of multi-head attention, and IA3 element-wise scaling modules are embedded in the projection layers of the feedforward network. According to formula (2), the backbone weights of the pre-trained large language model are fixed throughout the training process using the standard cross-entropy loss function, and only the adapter parameters of LoRA and IA3 are updated: (2) in, It is a loss function whose value is a non-negative real number, used to measure the degree of prediction error of the model under the current parameters; N The total number of samples in a training batch; T How many tokens each text is split into; (...) represents the probability distribution of the next token output by the large language model; Let be the next token predicted by the model at position t. This is the actual next token, which is the (t+1)th token that actually exists in the training data; The parameters of the backbone model are completely frozen; These are parameters for the trainable Adapter module.

3. The power grid project investment financial digitalization and penetration control system according to claim 2, characterized in that, The LoRA low-rank adaptation module achieves incremental adaptation through low-rank decomposition, and the transformation formula is as follows: , in, To output a tensor; For input tensors; Main path weight matrix; This is the transpose of the main path weight; For input application Regularization: During training, some elements are randomly set to 0, and scaling is performed to keep the expected value unchanged during testing; For low-rank fitting matrix ; For low-rank fitting matrix ; The rank of a low-rank fit; This is the scaling factor; The IA3 element-wise scaling module achieves precise adaptation by adjusting the activation value, and the transformation formula is as follows: , in, For Hadama accumulation.

4. The power grid project investment financial digitalization and penetration control system according to claim 1, characterized in that, The Natural Language to SQL Intelligent Query Module includes: A vector knowledge base is used to vectorize and store SQL sample data S, data table structure definitions D, and business domain knowledge K; where the i-th training data is denoted as a triple. ,in, For the question text, For SQL statements, Metadata identifier; during the training data vectorization process, according to formula (3) through the embedding model Map text to 3D vector space: ,(3) The semantic similarity retrieval unit is used to calculate the user question vector. The cosine similarity with vectors in the knowledge base is calculated, and the top-K similar data are returned as context; where the similarity is calculated using cosine similarity according to formula (4): , (4) in, Candidate vectors; For query vector; It is the dot product of vectors; for The L2 norm; for The L2 norm; The SQL generation inference unit is used to assemble the user question, the retrieved context, and the data table structure definition into a prompt text based on the retrieval enhancement generation paradigm, and input it into the large language model to generate SQL statements; wherein, according to formula (5), the SQL generation process is formalized into conditional probability generation: ,(5) in, For natural language questions input by the user, For the context of the retrieved similar examples, Define the structure of the relevant data tables; Multi-tenant data isolation units are used for data filtering based on tenant identifiers during vector retrieval and SQL execution phases; each tenant is assigned a unique identifier. Both training data and query execution are filtered based on tenant identifiers. The data filtering rules are expressed as follows: , in, For the original dataset D A subset containing only the data that the current tenant has the right to access; {...} is used to define a set in which all elements satisfy the conditions specified in the curly braces; d For the original dataset D Each individual element in; In a multi-tenant system, each data record will have a... This field identifies which tenant this data belongs to; A unique identifier for the current tenant; The training data incremental update unit is used to dynamically and incrementally update the training data, including SQL samples, data table structure definitions, and business documents; among them, for document-type data, file hashing is used to avoid duplicate training of the same content according to formula (6): ,(6) in, This is a collection of hash values ​​for the stored files.

5. The power grid project investment financial intelligent penetration control system according to claim 1, characterized in that, The in-depth data association mining performed at the business application layer includes: The association rule mining unit is used to extract association rules between cost input items and business indicator items from business data using the Apriori algorithm and / or FP-Growth algorithm, and to evaluate the strength of the rules through support, confidence, and lift. Here, let the dataset be... Include n Each transaction record T A set of business metrics; association rules X → Y The strength of [the measure] is measured by three metrics: support, confidence, and lift. in, For cost input items, For a set of business metrics; when the degree of improvement A value greater than 1 indicates cost input. With business metrics There is a positive correlation; when A value less than 1 indicates a negative correlation. The input-output efficiency assessment unit is used to build an assessment model, quantify the overall efficiency score of each business unit, and identify units with high input and low output, as well as units with low input and high output; where, let the first... i The input vector of each business unit is =( , ,..., The output vector is =( , ,..., If the overall efficiency score of the business unit is: , in, For the first j The weighting coefficients of output indicators, For the first l Weighting coefficients of input factors; based on efficiency scores The system automatically identifies < High input, low output units and > Low-input, high-output units provide a quantitative basis for optimizing resource allocation; The spatial clustering analysis unit is used to perform spatial clustering analysis on transformer substations using the DBSCAN density clustering algorithm in conjunction with a geographic information system, identifying areas with abnormal regional operation and maintenance efficiency; where a transformer substation is defined. of Neighborhood is defined as When | When |≥MinPts, the area The core object is marked, and all the substations in its neighborhood form a cluster.

6. The power grid project investment financial intelligent penetration control system according to claim 1, characterized in that, When the business application layer performs intelligent multi-source document processing and automatic generation of visual reports, it specifically includes: The OCR recognition unit is used to perform text recognition on document images using an architecture combining convolutional neural networks (CNNs) and recurrent neural networks (RNNs); where the input image is... I Through feature extraction network F Feature maps are extracted and then processed through a sequence modeling network. S The process of generating and recognizing text sequences is as follows: ; The information extraction unit is used to extract key structured information from the recognized text using Named Entity Recognition (NER) technology; wherein, the document text is denoted as... T The goal of named entity recognition is to find the optimal labeled sequence. , ; The natural language parsing module is used to perform word segmentation, part-of-speech tagging, syntactic analysis, and semantic role labeling on document content based on a pre-trained language model; among them, the document semantic representation aggregates contextual information through an attention mechanism. , in, This is the attention output at position i; Attention(Q,K,V) is the scaled dot product attention function. Q For query matrix; K The key matrix; V It is a value matrix; for Q and K The dot product matrix; This is the scaling factor; d for Q and K The dimension; The report generation engine uses a large language model to automatically generate professional reports, including charts and graphs, based on report template instructions, data analysis results, and visualization chart configurations. The report generation process is formalized as a conditional generation task. , in, For report template instructions, For the data analysis results, Configure visualization charts.

7. The power grid project investment financial digitalization and penetration control system according to claim 1, characterized in that, The risk management layer includes: A dynamic cost forecasting model is used to predict project costs throughout their entire lifecycle based on time series analysis and machine learning regression models; whereby, assuming the project cost is in the [missing information] phase... t The actual cost of a time is The predicted cost is calculated using a multi-factor regression model based on formula (7): ,(7) in, ..., Let there be n variables affecting the cost. , ,..., The regression coefficients for each factor are denoted as . This is the random error term; Anomaly detection unit is used to identify anomalous data points and trigger risk warnings based on the anomaly score calculation formula using the isolated forest algorithm. The knowledge graph reasoning engine is used to perform knowledge reasoning based on a knowledge graph that stores industry rules and entity relationships in the form of triples, and to achieve compliance checks. The triples are represented as follows: , in, For the head entity, For the relationship, For tail entities, For a collection of entities, For a set of relations; The inference engine performs knowledge reasoning based on graph neural networks (GNN) and updates entity embeddings according to formula (9): ,(9) in, N ( i ) is an entity i The set of neighboring nodes, and Here, σ is a learnable parameter, and σ is the activation function. The multi-level penetration query unit is used to respond to user operations, penetrating from the current level of data to the next level of detailed data, down to the underlying business details; where penetration query is from the first... l layer to the first l Drilling operation at +1 level: , in, The core operation of multi-level penetration query is to input the target data node at the current level and output all direct detailed data of its next level. For the first The current data node of the layer; For the first Detailed data nodes of the layer; To store the first In the layer, all with Detailed data nodes with parent-child relationships serve as the data source for drill-down operations.

8. A method for intelligent financial data-driven penetration management of power grid project investment based on the power grid project investment financial data-driven penetration management system according to any one of claims 1-7, characterized in that, The method includes: S1: Data integration, connecting multiple core business systems and collecting multi-source data, and constructing a business and finance integrated data resource pool after standardized processing; S2: Intelligent analysis, based on a pre-trained large language model, performs deep learning on the data in the business and finance integrated data resource pool, and automatically converts it into SQL statements for querying in response to natural language problems; S3: Business Applications. Based on intelligent analysis results, perform in-depth data correlation mining to identify resource configuration optimization links, and respond to instructions to intelligently process multi-source documents and automatically generate visual reports. S4: Risk Management, based on dynamic cost prediction models and risk early warning algorithms, performs full-cycle prediction and anomaly detection of project costs, and provides multi-level penetrating queries.

9. The method for intelligent and transparent financial management and control of power grid project investment according to claim 8, characterized in that, The data association deep mining in S3 includes: The Apriori algorithm and / or FP-Growth algorithm are used to extract the association rules between cost input items and business indicator items; Construct an input-output efficiency evaluation model to quantify the overall efficiency score of each business unit; By combining geographic information systems, the DBSCAN density clustering algorithm is used to perform spatial clustering analysis on the transformer substations to identify areas with abnormal efficiency.

10. The method for intelligent and transparent financial management of power grid project investment according to claim 8, characterized in that, The risk warning algorithm in S4 adopts the isolated forest algorithm, assuming the dataset... X Include m One sample, each sample x have d Dimensional features, calculate the anomaly score according to formula (8): ,(8) in, For the sample x Average path length in an isolated tree As a normalization factor, when outlier scores When the value approaches 1, it is identified as an abnormal data point, triggering a risk warning.