Energy data analysis method and system based on large model and proprietary knowledge base

Through analyzing complex multi-source heterogeneous energy data based on large models and proprietary knowledge bases, the shortcomings in data integration and correlation modeling are solved, and higher analytical accuracy and inference adaptability are achieved.

CN120179869AActive Publication Date: 2025-06-20NANJING DEEPCTRLS TECHNOLOGIES CO LTD

Patent Information

Application Number
CN202510654328.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

When facing complex multi-source heterogeneous energy data, the existing technology has inaccurate data integration, weak correlation and dynamic change modeling capabilities, neglecting the interdependence between energy points, building accurate graph structures and optimizing inference paths, making it difficult for the accuracy and real-time nature of the inference process to meet actual needs.

Method used

Using energy data analysis methods based on large models and proprietary knowledge bases, we collect and preprocess multi-source energy data, and divide them into four types of energy points, calculate features and connect them, build a graph convolution network model, and generate graph structures. Then, subqueries are generated based on the graph structure, and initial answers are obtained using a large language model, directed dependency graphs and dynamic tree inference structures are constructed, inference chains are optimized, and energy data analysis results are generated based on the optimal inference path, and traceability optimization is performed.

Benefits of technology

It improves the accuracy of energy data analysis, improves the transparency and controllability of inference results, and enhances the system's ability to adapt to changes in energy data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179869A_ABST
    Figure CN120179869A_ABST
Patent Text Reader

Abstract

The invention discloses an energy data analysis method and system based on a large model and a special knowledge base, and relates to the technical field of energy data analysis and intelligent reasoning, and the method comprises the steps: collecting and preprocessing multi-source energy data, dividing the multi-source energy data into four types of energy points, calculating the characteristics of the energy points, and carrying out the connection of the energy points, the connectivity is detected and optimized, a two-layer graph convolutional network GCN model is constructed, and a graph structure is generated; generating a sub-query based on a graph structure, obtaining an initial answer by using an LLM model, generating an initial global reasoning chain, interacting a special knowledge base and the initial global reasoning chain, optimizing the initial global reasoning chain, constructing a dynamic tree reasoning structure, performing dynamic branch adjustment based on an artificial potential field, and optimizing a reasoning path; through the combination mechanism of the GCN model, the LLM model and the special knowledge base, the reasoning accuracy is improved, the reasoning path is dynamically adjusted by using the artificial potential field, and the dynamic optimization capability of the reasoning chain is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of energy data analysis and intelligent reasoning, and in particular to an energy data analysis method and system based on a large model and a proprietary knowledge base. Background Art

[0002] With the rapid development of information technology, the collection and analysis of energy data have gradually become an important part of energy management and decision-making. Especially driven by big data, artificial intelligence (AI), and Internet of Things (IoT) technologies, the methods for collecting and analyzing energy data are becoming increasingly diverse. In the early stage, the analysis methods of energy data mainly relied on traditional statistical analysis and simple regression models, with low processing efficiency and poor accuracy, making it difficult to adapt to complex energy systems. With the rise of machine learning and deep learning technologies, data-driven prediction models have gradually replaced traditional methods, providing more accurate predictions and optimization solutions by mining a large amount of historical energy data.

[0003] Despite certain progress in the prior art, there are still multiple problems. When faced with complex multi-source heterogeneous energy data, the prior art has certain limitations, which easily lead to inaccurate data integration, thereby affecting the accuracy of the analysis results. The prior art has weak modeling capabilities for the correlation and dynamic changes between data, often ignoring the interdependent relationships between energy points. The prior art still has certain deficiencies in constructing accurate graph structures, optimizing inference paths, and enhancing dynamic adjustments during the inference process. Especially in the analysis of multi-source energy data, there is a lack of effective inference chain construction and optimization mechanisms, making it difficult for the accuracy and real-time performance of the inference process to meet actual requirements. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an energy data analysis method based on a large model and a proprietary knowledge base to solve the problems that the prior art has certain limitations when faced with complex multi-source heterogeneous energy data, which easily leads to inaccurate data integration, thereby affecting the accuracy of the analysis results. The prior art has weak modeling capabilities for the correlation and dynamic changes between data, often ignoring the interdependent relationships between energy points. The prior art still has certain deficiencies in constructing accurate graph structures, optimizing inference paths, and enhancing dynamic adjustments during the inference process. Especially in the analysis of multi-source energy data, there is a lack of effective inference chain construction and optimization mechanisms, making it difficult for the accuracy and real-time performance of the inference process to meet actual requirements.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: In the first aspect, the present invention provides an energy data analysis method based on a large model and a proprietary knowledge base, which includes Collect multi-source energy data and preprocess it. Divide the multi-source energy data into four types of energy points, calculate the characteristics of the energy points and connect the energy points. Detect and optimize the connectivity, construct a two-layer graph convolutional network (GCN) model, and generate a graph structure. Generate sub-queries based on the graph structure, use the LLM model to obtain initial answers, construct a directed dependency graph, generate an initial global inference chain, interact the proprietary knowledge base with the initial global inference chain, optimize the initial global inference chain, map the optimized inference chain to the nodes of a tree, construct a dynamic tree-shaped inference structure, perform dynamic branch adjustment based on the artificial potential field, and optimize the inference path. Use the LLM model based on the optimal inference path to obtain the energy data analysis results, trace each inference process, and optimize the energy data analysis results based on the trace results. Store the data generated during the inference process and update the system with real-time data.

[0007] As a preferred solution of the energy data analysis method based on the large model and the proprietary knowledge base of the present invention, wherein: the step of collecting multi-source energy data and preprocessing it, constructing a model, and generating a graph structure includes: The multi-source energy data includes power consumption time series, equipment operating status, climate, and energy price data. Divide the preprocessed multi-source energy data into four types of energy points, generate an energy point set, calculate the electricity feature vector using the time feature engineering method, calculate the feature vectors of equipment, weather, and price respectively using the feature combination method, and generate a feature vector set. Use the equipment drive association method to construct energy connections, generate an energy connection set from all the energy connections, construct an initial energy graph, perform connectivity detection using the graph structure balancing method, and perform dynamic adjustment and connection reduction operations based on the detection results to obtain an optimized energy graph. Collect historical multi-source energy data, preprocess it and extract features, and use it as a training set. Construct a two-layer graph convolutional network (GCN) model, define a loss function to supervise the task, use the feature vectors of the training set as input, perform iterative training, stop the iteration after reaching the maximum number of iterations, output the optimized model weight parameters, apply them to the two-layer graph convolutional network (GCN) model, use the feature vector set as input, obtain the final enhanced embedding set, and add it to the optimized energy graph to obtain a graph structure.

[0008] As a preferred solution of the energy data analysis method based on a large model and a proprietary knowledge base according to the present invention, wherein: generating an initial global inference chain, interacting the proprietary knowledge base with the initial global inference chain, optimizing the initial global inference chain, constructing a dynamic tree-shaped inference structure, and performing dynamic branch adjustment based on an artificial potential field, and optimizing the inference path includes: Using the task prompt decomposition method to divide it into sub-queries, generating a list of sub-queries, and inputting it together with the graph structure into a pre-configured large language model LLM to generate an initial answer, calculating the answer confidence, using the fixed threshold screening method to screen the initial answers with the answer confidence greater than the corresponding answer confidence threshold, and generating an initial answer set; Based on the dependency relationship between sub-queries in the sub-query list, constructing a directed dependency graph, generating an initial global inference chain, and using a pre-trained DistilBERT model for each sub-query to generate a question vector; Collecting and preprocessing the document data of the proprietary knowledge base, using FAISS to construct a document index, generating knowledge base vectors through a pre-trained DistilBERT model, calculating the cosine similarity between the question vector and the knowledge base vectors, performing a k-nearest neighbor search on the question vector through the document index constructed by FAISS to obtain a candidate document set, segmenting it by natural paragraphs, and using a pre-trained DistilBERT model to generate paragraph vectors; Calculating the cosine similarity between the paragraph vector and the question vector to obtain candidate answer texts, using a pre-trained SpanBERT question-answering model to extract the answer segment that best matches the sub-query, calculating the confidence score of the answer segment, sorting it in descending order to obtain an optimized answer vector, calculating the answer similarity score between the vector of the initial answer and the vector of the optimized answer, summing the answer similarity score and the confidence score of the highest answer segment vector using the weighted summation method to obtain a comprehensive consistency score, setting a comprehensive consistency score threshold using the empirical method, comparing it with the comprehensive consistency score, and performing operations based on the comparison result to obtain an optimized inference chain; Mapping each sub-query-answer pair in the optimized inference chain to a node of the tree and using a pre-configured large language model LLM to evaluate the sub-query and other sub-queries for causal dependencies, adding a directed edge if there is a dependency and calculating an initial weight for each node of the tree otherwise remaining unchanged. After adding the edges, using depth-first search to detect whether there is a cycle. If there is a cycle, sorting the initial weights of the nodes forming the cycle in descending order and removing the edges of the node with the lowest initial weight to obtain an initial dynamic tree-shaped inference structure ; Calculating an initial weight for each node of the tree Calculate the initial consistency score and compare it with the set initial consistency threshold. Based on the comparison result, use depth - first search (DFS) to backtrack to its parent node , generate a list of low - consistency nodes , generate alternative sub - queries , input the alternative sub - queries into a pre - configured large language model (LLM) to generate new answers , and generate nodes of a new tree with the alternative sub - queries , calculate the new consistency score , compare the new consistency score with the initial consistency score , and based on the difference in the comparison result, perform operations to replace and retain nodes of the tree to obtain an updated dynamic tree - shaped inference structure , traverse all paths from the root to the leaves in the updated dynamic tree - shaped inference structure , each path contains a set of nodes. For each path , calculate the comprehensive score , and sort them in descending order, select the path with the highest score as the optimal path.

[0009] As a preferred solution of the energy data analysis method based on a large model and a proprietary knowledge base according to the present invention, wherein: obtaining the energy data analysis result using the LLM model based on the optimal inference path includes: Collect the answers of each tree node from the optimal path, arrange the answers according to the logical order of energy data analysis to form an analysis answer list; The logical order of the energy data analysis includes checking the source nodes of each answer, determining their logical dependencies, inputting the analysis answer list into a pre - configured large language model (LLM), giving clear instructions based on the logical dependencies, and outputting the energy analysis result and explanations.

[0010] As a preferred solution of the energy data analysis method based on a large model and a proprietary knowledge base according to the present invention, wherein: tracing each inference process and optimizing the energy data analysis result based on the tracing result includes: Extract data from the optimized inference chain CoQ, perform structured processing using natural language processing (NLP), generate corresponding tracing tags for each answer, and form a tracing set; Extract a set of feature vectors from the graph structure, use an exponential decay model to calculate the time - decay weight for each feature, and weight each feature in the set of feature vectors using the time - decay weight to optimize the energy analysis result, associate the final energy analysis result with the tracing set, and generate a structured output.​

[0011] As a preferred solution of the energy data analysis method based on the large model and the proprietary knowledge base of the present invention, wherein: storing the data generated in the inference process includes: Converting the final analysis result, the trace set, the graph structure, and the optimized inference chain into a unified JSON format, adding metadata tags to each piece of data, and using a hash function to generate a unique identifier for the data to obtain the processed data; Sharding the processed data according to the identifier and timestamp, storing the sharded data in a distributed database, ensuring the balanced distribution of data shards among nodes through consistent hashing, recording the shard index at the same time, constructing an inverted index to accelerate retrieval, and establishing an adjacency list for the graph structure to store edge relationships.

[0012] As a preferred solution of the energy data analysis method based on the large model and the proprietary knowledge base of the present invention, wherein: using the real-time data update system includes: Aligning the real-time data and the stored sharded data in terms of time to generate a fusion data set, calculating the fusion confidence of the fused data, storing the fusion data set and the fusion confidence and overwriting the old records to obtain the stored new data; Updating the graph structure and the updated inference chain according to the stored new data, and calculating the data change significance Based on experience, set a data change significance threshold. If the data change significance is greater than the data change significance threshold, enter the analysis stage, otherwise remain unchanged.

[0013] In a second aspect, the present invention provides a file encryption system, including, A collection and construction module for collecting multi-source energy data and preprocessing it, dividing the multi-source energy data into four types of energy points, calculating the characteristics of the energy points and connecting the energy points, detecting and optimizing the connectivity, constructing a two-layer graph convolutional network (GCN) model, and generating a graph structure; A generation and optimization module for generating subqueries based on the graph structure, obtaining an initial answer using an LLM model, constructing a directed dependency graph, generating an initial global inference chain, interacting the proprietary knowledge base with the initial global inference chain, and optimizing the initial global inference chain; An inference chain path module for mapping the optimized inference chain to the nodes of a tree, constructing a dynamic tree-shaped inference structure, dynamically adjusting the branches based on the artificial potential field, and optimizing the inference path; An analysis and trace module for obtaining the energy data analysis result using an LLM model based on the optimal inference path, tracing each inference process, and optimizing the energy data analysis result based on the trace result; A storage and update module for storing the data generated in the inference process and using the real-time data update system.

[0014] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the energy data analysis method based on a large model and a proprietary knowledge base as described in the first aspect of the present invention is implemented.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the energy data analysis method based on a large model and a proprietary knowledge base as described in the first aspect of the present invention is implemented.

[0016] The beneficial effects of the present invention are as follows: By collecting and preprocessing multi-source energy data, the multi-source energy data is divided into four types of energy points, the features of the energy points are calculated and the energy points are connected, the connection degree is detected and optimized, a two-layer graph convolutional network GCN model is constructed to generate a graph structure; based on the graph structure, sub-queries are generated, the LLM model is used to obtain an initial answer, a directed dependency graph is constructed, an initial global inference chain is generated, the proprietary knowledge base and the initial global inference chain are interacted, the initial global inference chain is optimized, the optimized inference chain is mapped to the nodes of a tree, a dynamic tree-shaped inference structure is constructed, dynamic branch adjustment is performed based on the artificial potential field, and the inference path is optimized; based on the optimal inference path, the LLM model is used to obtain the energy data analysis result, traceability is performed for each inference process, and the energy data analysis result is optimized based on the traceability result, improving the accuracy of energy data analysis, enhancing the transparency and controllability of the inference result, and enhancing the adaptability of the system to changes in energy data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0018] Figure 1 It is a flowchart of the energy data analysis method based on a large model and a proprietary knowledge base in Embodiment 1.

[0019] Figure 2 It is a schematic diagram of the energy data analysis system based on a large model and a proprietary knowledge base in Embodiment 1.

[0020] Figure 3 It is a flowchart of the optimization process of the dynamic tree-shaped inference structure in Embodiment 1.

[0021] Figure 4Schematic diagram of the data storage and real-time update mechanism in Embodiment 1. Detailed implementation manners

[0022] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific implementation manners of the present invention in conjunction with the accompanying drawings of the specification.

[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0024] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other with other embodiments.

[0025] Embodiment 1, referring to Figures 1 to 4 , which is the first embodiment of the present invention. This embodiment provides an energy data analysis method based on a large model and a proprietary knowledge base, including the following steps: S1. Collect multi-source energy data and preprocess it. Divide the multi-source energy data into four types of energy points, calculate the features of the energy points and perform energy point connection, detect and optimize the connection degree, construct a two-layer graph convolutional network (GCN) model, and generate a graph structure; Specifically, collect multi-source energy data and preprocess it; The multi-source energy data includes power consumption time series, device operating status (switch status, power), climate (temperature and humidity), and energy price data; The preprocessing includes aligning all data according to a unified timestamp, using statistical tests to check the integrity of the multi-source energy data, using a periodic regression filling method to fill in missing values, using a sliding window Z-score anomaly detection method to detect and delete outliers in the multi-source energy data, and performing standardization processing on the multi-source energy data; Divide the preprocessed multi-source energy data into four types of energy points, including power consumption points, device points, weather points, and price points; Extract the power consumption at each time point from the preprocessed multi-source energy data and use a time feature engineering method to calculate features to obtain a power consumption feature vector; Extract the static information (switch state) of each device, hourly weather data, and hourly price data from the preprocessed multi-source energy data, and use the feature combination method to calculate the features respectively to obtain the feature vectors of the device, weather, and price; Generate an energy point set for all energy points and a feature vector set for all feature vectors; Use the device-driven association method to construct energy connections, including detecting the operating state (0 means off, 1 means on) of each device in the preprocessed multi-source energy data at time If it is on, connect the power consumption - device, and use feature multiplication to calculate the intensity of the power consumption - device; Calculate the power consumption change rate, price change rate, and weather change rate, and use the Pearson correlation coefficient to calculate the synchronization score of power consumption - weather, calculate the synchronization score of power consumption - price, and calculate the synchronization score of device - weather respectively. Use the empirical method to set the corresponding synchronization scores respectively, make a comparison, and screen the energy points greater than the corresponding thresholds for connection; Generate an energy connection set based on all energy connections, construct an initial energy graph from the energy connection set and the energy point set, and use the graph structure balance method to detect the connection degree to obtain the qualified connection degree and the unqualified connection degree (insufficient connection degree); For the unqualified connection degree, use the dynamic threshold adjustment method to dynamically adjust the thresholds of the synchronization score and response coefficient of the initial energy graph, and regenerate the energy connection set and the feature vector set according to the dynamically adjusted thresholds; For the qualified connection degree, use the intensity quantile truncation method to streamline the connections; Obtain the optimized energy graph based on the regenerated energy connection set and the streamlined connections and energy point set; Set each energy point as a node of the graph, and mark each edge of the optimized energy graph as 1 and the non-edge as 0; Collect historical multi - energy data, preprocess it and extract features to be used as the training set; Construct a two - layer graph convolutional network GCN model, define the loss function to supervise the task, and the formula is as follows: , Among them, is the mean square error, which is used to measure the gap between the model prediction value and the true value, is the number of nodes participating in training, is the output embedding of node v in the second layer, is the prediction layer weight matrix, with a dimension of 64×1, is the true value of node v; Calculate the initial embedding based on the feature vectors of the training set, and the formula is as follows: , Among them, is the output embedding of node v of the graph at the 0th layer, fixed to 64 dimensions, is the weight matrix at the 0th layer, with dimensions 64× , is the input feature dimension (1 for power consumption nodes, depending on the number of devices for nodes of the device graph, 2 for weather nodes, and 1 for price nodes), is the feature vector of node v of the graph, is the bias vector, with dimensions 64; Calculate the dynamic edge weight , and the formula is as follows: , Among them, is the cosine similarity, used to measure the correlation between the output embedding of node v of the graph at the 0th layer and the output embedding of node at the 0th layer, is node 's neighbor set, c is the index of the neighbor set, is the node of the graph that is not equal to node v of the graph; The first layer of convolution, the formula is as follows: , Among them, is the output embedding of node v of the graph at the 1st layer, is the ReLU activation function, is the neighbor set of node v of the graph, is the first layer of convolution weight matrix, is the first layer of self-loop weight matrix; The second layer of convolution, the formula is as follows: , Among them, is the second layer weight matrix, is the second layer of self-loop weight matrix; Use the fixed value method to set the maximum number of iterations, use the Adam optimizer, aim to minimize the loss L, perform iterative training, stop the iteration after reaching the maximum number of iterations, and output the optimized model weight parameters; Apply the optimized model weight parameters to the graph convolutional network GCN model, use the feature vector set as the input, and obtain the final enhanced embedding set H; Add the final enhanced embedding set H to the optimized energy graph to obtain the graph structure G.

[0026] By extracting features from preprocessed multi-source data and converting the data into a format suitable for machine learning algorithms, the training efficiency and accuracy of subsequent models can be greatly improved. By building connections between energy points and detecting their correlations through device-driven association methods, synchronization score calculation methods, etc., the relationship network between energy points can be accurately constructed, and highly correlated energy points can be identified, helping the energy management system to more accurately identify key factors in the decision-making process. By constructing a two-layer graph convolutional network (GCN) model and using loss functions for supervised learning and optimizing model weights, it can efficiently process graph structured data and capture the complex relationships and interactions between nodes, which can effectively improve prediction accuracy while maintaining good generalization capabilities.

[0027] S2. Generate subqueries based on the graph structure, use the LLM model to get the initial answer, build a directed dependency graph, generate an initial global reasoning chain, interact the proprietary knowledge base with the initial global reasoning chain, optimize the initial global reasoning chain, map the optimized reasoning chain to the nodes of the tree, build a dynamic tree-shaped reasoning structure, perform dynamic branch adjustment based on the artificial potential field, and optimize the reasoning path; Specifically, according to the node information and feature vector of the graph structure G, the task prompt decomposition method is used to divide it into subqueries , and assign a unique identifier to each subquery to generate a subquery list Q; The subquery , Index for the number of subqueries, including :What is the historical electricity consumption benchmark value? : The impact of climate conditions on electricity consumption, :The impact of equipment efficiency on electricity consumption, : The impact of energy prices on electricity consumption; Input the subquery list Q and the graph structure G into the preconfigured large language model LLM, with the instruction: "Require that each subquery be Generate initial answer ”, for each initial answer Calculate answer confidence , the formula is as follows: , in, For subqueries The number of edges involved, is the total number of edges in the graph structure G, is the data integrity weight (range 0 to 1); Use the empirical method to set the answer confidence threshold, use the fixed threshold screening method to screen out the initial answers whose answer confidence is greater than the corresponding answer confidence threshold, and generate an initial answer set; Based on the dependency relationships among Q subqueries, construct a directed dependency graph D, apply topological sorting to the dependency graph D to determine the subquery execution order, and adjust the order of the initial answer set according to the sorting; The said dependency relationships include For the basic query, there is no dependency. Dependency (Climate impact requires historical benchmark), Dependency (Equipment impact requires historical benchmark), Dependency , and (Price impact requires the results of the previous three); Input the sorted answer set into a pre-configured large language model LLM, with the instruction "Integrate in dependency order", and generate an initial global inference chain from the results and subqueries; For each subquery in the initial global inference chain Use the pre-trained DistilBERT model to generate question vectors; Collect the document data of the proprietary knowledge bases (energy consumption standard database and climate model library), preprocess all the documents in the proprietary knowledge bases, build a document index using FAISS, generate knowledge base vectors through the pre-trained DistilBERT model, calculate the cosine similarity between the question vectors and the knowledge base vectors, perform a k-nearest neighbor search on the question vectors through the document index built by FAISS, and select the l corresponding documents with the highest similarity (l is a value set based on expert experience), and mark them as the candidate document set; Segment each document in the candidate document set by natural paragraphs to generate a paragraph list, and use the pre-trained DistilBERT model to generate paragraph vectors; Calculate the cosine similarity between the paragraph vectors and the question vectors, sort the cosine similarities in descending order, select the top p (p is a value set based on expert experience) paragraphs to form a relevant paragraph set, and splice them in the original order to form a candidate answer text; Use the pre-trained SpanBERT's question-answering model on the candidate answer text to extract the answer fragment that best matches the subquery , Let be the candidate document quantity index, obtain the answer fragment vector, and calculate the confidence score of the answer fragment. The formula is as follows: , where is the confidence score of the answer fragment answering the subquery , , is the answer fragment Similarity with subquery , is the problem vector of the subquery , is the answer fragment vector of the answer fragment , is the answer length ratio is the candidate text of the i-th subquery and the th candidate document , and are the weights and biases of the pre-trained SpanBERT question-answering model (the values are determined by the model); Select the answer fragment vector with the highest confidence score as the optimized answer vector, calculate the answer similarity score between the vector of the initial answer and the vector of the optimized answer, sum the answer similarity score and the confidence score of the highest answer fragment vector using the weighted summation method to obtain the comprehensive consistency score, set the comprehensive consistency score threshold using the empirical method. If the comprehensive consistency score is greater than or equal to the comprehensive consistency score threshold, no adjustment is required, otherwise mark the subquery as "query to be feedback"; For the subquery marked as "query to be feedback" encapsulate it into a feedback data packet, and input the feedback data packet into the pre-configured large language model LLM, with the instruction: "Correct the initial answer according to the proprietary knowledge base answer and confidence, and ensure consistency with the energy data analysis scenario", to obtain the optimized inference chain , where n is the total number of inference chains; The said feedback data packet includes the subquery, the initial answer, the optimized answer, the confidence, the supporting document, and the comprehensive consistency score; Map each subquery-answer pair in the optimized inference chain CoQ to the nodes of a tree and initialize all the nodes of the tree as an isolated node set N; For each node of the tree , analyze the semantics of its subquery , extract keywords (such as "historical electricity consumption", "climate impact"), and use the pre-configured large language model LLM to evaluate its causal dependence on other subqueries with the instruction: "Judge whether '[[' ']]' depends on the result of '[[' ']]', and whether 'climate impact on electricity consumption' requires '[[' ']]' as a premise". If there is a dependence, the node of the tree Calculate the initial weights for each node of the tree The formula is as follows: where is the subquery The importance of the energy analysis, evaluated by the LLM (range [0, 1], e.g., "historical electricity consumption" = 0.8, "secondary factor" = 0.3); After adding the edges, use depth - first search to detect if there is a cycle. If there is a cycle, sort the initial weights of the nodes forming the cycle in descending order and remove the edges of the node with the lowest initial weight; Generate an initial dynamic tree - shaped inference structure is the set of edges is the set of initial weights of the nodes; For each node of the tree Calculate the initial consistency score The formula is as follows: where is a tuning parameter used to balance confidence and conflict impact (set to 0.5 to ensure balance), is the repulsive potential field representing the answer and the support document conflict degree, is the repulsive force coefficient that controls the conflict penalty intensity (a fixed value obtained from experimental experience), is the answer and the support document semantic similarity (calculated using cosine similarity, range [0, 1]); Set the initial consistency threshold using the empirical method. If < the initial consistency threshold, then mark the node of the tree as a "low - consistency node" and use depth - first search DFS to backtrack to its parent node to generate a list of low - consistency nodes ; For each node of the tree in the list of low - consistency nodes subquery and the answer of the parent node to generate an alternative subquery Input the alternative subquery into the pre - configured large language model LLM to generate a new answer ; ​​​​​​Replace the subquery and the new answer to generate nodes of a new tree and calculate the new consistency score Compare the new consistency score with the initial consistency score If add the new node as a child node of the parent node to replace the node of the tree Otherwise, retain the node of the tree to obtain the updated dynamic tree - shaped inference structure ; Traverse all paths from root to leaf in the updated dynamic tree - shaped inference structure Each path contains a set of nodes. Calculate the comprehensive score for each path using the following formula: , where is the path length, i.e., the number of nodes in the path, and are the length penalty coefficient and the heuristic adjustment coefficient respectively, set using the fixed - value method, is the heuristic function, estimating the "distance" from the end - node of the path to the final goal (energy prediction result), is the preliminary prediction value based on the existing answers (estimated by LLM), is the answer of the end - node of the path, is the path number index; Sort all the calculated comprehensive scores in descending order and select the path with the highest score as the optimal path.

[0028] By transforming the energy data analysis task into multiple subqueries and assigning a unique identifier to each subquery, the decomposition of complex tasks is achieved. By combining the subqueries with a graph structure, using a large language model (LLM) to generate preliminary answers and calculate confidence levels, calculating confidence levels can effectively filter out reliable answers and prevent low-quality or irrelevant answers from entering the subsequent analysis process, thereby improving the quality of the analysis results. By constructing a dependency graph and applying topological sorting, the execution order of the subqueries is determined, which helps to reasonably arrange the execution logic between subqueries and ensure that the results of preceding queries can provide the correct basic data for subsequent queries. By generating question vectors and calculating the cosine similarity with vectors in the knowledge base, relevant content related to the subqueries can be accurately selected from a large number of documents, optimizing the retrieval efficiency of information in the knowledge base, and selecting the most relevant documents through similarity sorting, reducing interference from irrelevant information and enhancing the pertinence and efficiency of knowledge application. By using the SpanBERT model to extract the answer fragments that best match the subqueries, the deviation caused by semantic errors or information loss can be effectively reduced. By mapping the optimized inference chain to the nodes of a tree and analyzing the causal relationships between subqueries, a dynamic tree-shaped inference structure is constructed to ensure the flexibility and adaptability of the inference process. It provides strong structural support for the inference process, enabling the inference chain to self-regulate and optimize in complex environments, enhancing the robustness and adaptability of the system. By calculating the consistency scores of the tree nodes and providing feedback on inconsistent nodes to optimize the inference path, this operation can effectively filter out low-quality paths, correct and optimize them, and ultimately ensure the generation of the optimal path. By traversing the updated tree-shaped inference structure, calculating the comprehensive scores of each path, and selecting the path with the highest score as the final output, it is ensured that the analysis results are based on the most valuable inference path, thereby achieving the most optimized analysis results.

[0029] S3. Use the LLM model based on the optimal inference path to obtain the energy data analysis results, trace each inference process, and optimize the energy data analysis results based on the trace results; Specifically, collect the answers of the nodes of each tree from the optimal path, arrange the answers according to the logical order of energy data analysis to form an analysis answer list; The logical order of the energy data analysis includes checking the source nodes of each answer, determining their logical dependencies, placing basic trend answers (such as "Historical electricity consumption has been increasing year by year") at the beginning, followed by influencing factors (such as "Climate impact leads to an increase in electricity consumption", "Equipment efficiency is stable"), and finally adding economic factors (such as "Energy price fluctuations are rising"); Input the analysis answer list into a pre-configured large language model (LLM) and give a clear instruction: "Based on the following sequential information, analyze the main trends and influencing factors of energy data: First, the historical electricity consumption trend, second, the climate impact, third, the equipment efficiency, and finally, the impact of energy prices. Please summarize the energy usage patterns and key driving factors and explain the basis for the analysis." The pre-configured large language model (LLM) outputs the energy analysis results and explanations (such as "Based on historical trends and climate impact, energy demand continues to grow, equipment efficiency is stable and has not significantly changed the pattern, while price fluctuations indicate the need to pay attention to cost control").

[0030] By collecting the answers of each node of each tree from the optimal path, arranging the answers according to the logical order of energy data analysis, forming an analysis answer list, ensuring the structuring and hierarchization of the analysis process, avoiding information chaos or dislocation. By checking the source nodes of each answer and determining their logical dependencies, a strict analysis order is formed, separating the basic trend answers, influencing factors, and economic factors hierarchically, ensuring the clear causal relationship between data during the analysis process. By inputting the answer list arranged in logical order into the pre-configured large language model (LLM) and providing clear instructions to the model, ensuring that the model's analysis and reasoning fully conform to the expected logical order and analysis framework.

[0031] Furthermore, extract each sub-query, answer, and supporting document from the optimized reasoning chain (CoQ), use natural language processing (NLP) to structure each supporting document, generate corresponding traceability tags for each answer, and the traceability tags contain the answer, the confidence level of the answer (calculated based on document authority and data consistency), and the source identifier of the supporting document, forming a traceability set. Use the empirical method to set a confidence threshold for the answers in the traceability set. If the confidence level of an answer is less than the set confidence threshold for the answer, mark the corresponding answer in the energy analysis result as abnormal and delete it. Extract a set of feature vectors from the graph structure G, use the exponential decay model to calculate the time decay weight for each feature, and use the time decay weight to weight each feature in the set of feature vectors to optimize the energy analysis result. The formula is as follows: , where, is the final energy analysis result (quantifying the energy analysis value, the final predicted electricity consumption value), is the energy analysis result after deleting outliers, is the regression coefficient (obtained by fitting historical data), is the set of feature vectors weighted for each feature; Associate the final energy analysis results with the traceability set to generate a structured output.

[0032] By extracting each sub-query, answer, and supporting document from the optimized inference chain CoQ, and using natural language processing (NLP) to structure the supporting documents, it ensures that the source of each answer, the authority of the basis document, and data consistency are fully reflected and quantified. By generating corresponding traceability labels, information such as the confidence level of each answer and the source identifier of the supporting document is associated, providing strong support for the final analysis results. By extracting a set of feature vectors from the graph structure G and combining an exponential decay model to calculate the time decay weight for each feature, it can assign different importance to data at different time points and optimize the features by weighting, enabling the analysis system to better adapt to dynamically changing energy data, improving the system's response speed and accuracy to the latest situation. By associating the finally optimized energy analysis results with the traceability set to generate a structured output, the integrity and systematization of the results can be achieved.

[0033] S4. Store the data generated during the inference process and update the system with real-time data; Specifically, convert the final analysis results , the traceability set, the graph structure G, and the optimized inference chain CoQ into a unified JSON format, add metadata tags to each piece of data, and use a hash function to generate a unique identifier for the data to obtain the processed data; Slice the processed data according to the identifier and timestamp, store the sliced data in a distributed database, ensure the balanced distribution of data slices among nodes through consistent hashing, record the slice index at the same time, build an inverted index to accelerate retrieval, and establish an adjacency list for the graph structure to store edge relationships to ensure that the query latency meets the real-time requirements.

[0034] By uniformly converting the final analysis results, the traceability set, the graph structure G, and the optimized inference chain CoQ into the JSON format, it is convenient for cross-platform storage and data exchange. By slicing according to the identifier and timestamp, a large amount of data can be effectively segmented and managed. By introducing consistent hashing, the problem of uniform distribution of data in a distributed system is solved, ensuring the balanced distribution of data among different nodes, thus effectively optimizing the storage and query performance. Recording the slice index and building an inverted index can accelerate the data retrieval process, ensuring the balanced distribution and fast retrieval of data in a distributed system, avoiding data access bottlenecks and performance degradation problems, and ensuring the efficient operation of the system. By establishing an adjacency list for the graph structure to store edge relationships, the edge relationships between nodes in the graph are efficiently represented, ensuring that adjacent nodes can be quickly accessed during query, thus accelerating the traversal and query of the graph structure.

[0035] Furthermore, align the real-time data and the stored multi-energy data in terms of time to generate a fused dataset, calculate the fusion confidence of the fused data, and store the fused dataset and the fusion confidence to overwrite the old records; Adjust the energy graph according to the type of the stored new data. If it is a new device and scenario, add nodes to the graph; if it reveals new relationships, add edges; Calculate new feature vectors based on the new energy graph, and use the graph neural network GNN to re-encode the subgraphs affected by the new data to generate an updated graph structure; For the sub-queries in the optimized inference chain CoQ that are affected by the new data , use the pre-configured large language model LLM to regenerate new answers, and update the inference chain based on the new answers; Calculate the data change significance , and the formula is as follows: , where is the new fused data, is the historical data value; Set the data change significance threshold based on experience. If the data change significance is greater than the data change significance threshold, enter the analysis phase; otherwise, remain unchanged; The analysis phase includes generating new energy data analysis results based on the updated inference chain and the updated graph structure.

[0036] By aligning the real-time data and historical multi-source energy data, we ensure that data of different time scales and sources can be compared and processed within the same time frame, generate a fused data set and calculate its fusion confidence, further ensuring the reliability and accuracy of the data. By updating old records, we ensure the timeliness of the data and reflect the latest energy conditions. By adjusting the energy graph structure, we can flexibly respond to new data types and newly discovered relationships, ensuring that the graph structure is constantly adaptive and expanded as new data is added, so that the model can always effectively reflect the overall picture of the current energy system. The subgraphs affected by new data are re-encoded through the graph neural network (GNN). Ensure that new data can deeply influence and update the graph structure. By updating the reasoning chain, it can ensure that the generated answers are more timely and targeted, and improve the rationality of the data analysis results. Based on the set significance threshold, only when the data change reaches a certain level will it enter the further analysis stage. This design not only improves the efficiency of data analysis, but also avoids unnecessary computing burden caused by data fluctuations. By combining the updated reasoning chain and graph structure, the final new energy data analysis results are generated, ensuring that all previous data fusion, graph structure update and reasoning chain optimization are fully utilized to produce accurate analysis results that are in line with the current energy system status.

[0037] This embodiment also provides an energy data analysis system based on a large model and a proprietary knowledge base, including: The collection and construction module is used to collect and preprocess multi-source energy data, classify the multi-source energy data into four types of energy points, calculate the characteristics of energy points and connect energy points, detect and optimize the connectivity, build a two-layer graph convolutional network GCN model, and generate a graph structure; Generate an optimization module, which is used to generate subqueries based on the graph structure, use the LLM model to obtain the initial answer, build a directed dependency graph, generate an initial global reasoning chain, interact the proprietary knowledge base with the initial global reasoning chain, and optimize the initial global reasoning chain; The reasoning chain path module is used to map the optimized reasoning chain into tree nodes, build a dynamic tree-shaped reasoning structure, perform dynamic branch adjustment based on the artificial potential field, and optimize the reasoning path; The analysis and tracing module is used to obtain energy data analysis results based on the optimal reasoning path using the LLM model, trace each reasoning process, and optimize the energy data analysis results based on the tracing results; The storage update module is used to store the data generated by the reasoning process and use real-time data to update the system.

[0038] This embodiment also provides a computer device, which is applicable to the case of the energy data analysis method based on a large model and a proprietary knowledge base, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the energy data analysis method based on a large model and a proprietary knowledge base as proposed in the above embodiment.

[0039] The computer device can be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the outer shell of the computer device. It can also be an external keyboard, a touchpad, or a mouse, etc.

[0040] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the energy data analysis method based on a large model and a proprietary knowledge base as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read-Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0041] In summary, the present invention includes the following steps: collecting and preprocessing multi-source energy data, classifying the multi-source energy data into four types of energy points, calculating the characteristics of the energy points and connecting the energy points, detecting and optimizing the connectivity, constructing a two-layer graph convolutional network (GCN) model to generate a graph structure; generating sub-queries based on the graph structure, obtaining an initial answer using an LLM model, constructing a directed dependency graph, generating an initial global inference chain, interacting the proprietary knowledge base with the initial global inference chain, optimizing the initial global inference chain, mapping the optimized inference chain to the nodes of a tree, constructing a dynamic tree-shaped inference structure, performing dynamic branch adjustment based on the artificial potential field, and optimizing the inference path; obtaining the energy data analysis result using the LLM model based on the optimal inference path, tracing each inference process, optimizing the energy data analysis result based on the tracing result, improving the accuracy of the energy data analysis, enhancing the transparency and controllability of the inference result, and strengthening the adaptability of the system to changes in energy data.

[0042] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An energy data analysis method based on a large model and a proprietary knowledge base, characterized in that: include, Collect and preprocess multi-source energy data, classify the multi-source energy data into four types of energy points, calculate the characteristics of energy points and connect energy points, detect and optimize the connectivity, build a two-layer graph convolutional network (GCN) model, and generate a graph structure; Generate subqueries based on the graph structure, use the LLM model to get the initial answer, build a directed dependency graph, generate an initial global reasoning chain, interact with the proprietary knowledge base and the initial global reasoning chain, optimize the initial global reasoning chain, map the optimized reasoning chain to the nodes of the tree, build a dynamic tree-shaped reasoning structure, perform dynamic branch adjustment based on the artificial potential field, and optimize the reasoning path; Based on the optimal reasoning path, the LLM model is used to obtain the energy data analysis results, and each reasoning process is traced back, and the energy data analysis results are optimized based on the traceability results; Store the data generated by the reasoning process and update the system with real-time data.

2. The energy data analysis method based on a large model and a proprietary knowledge base as claimed in claim 1, characterized in that: The method of collecting and preprocessing multi-source energy data, building a model, and generating a graph structure includes: The multi-source energy data includes electricity consumption time series, equipment operation status, climate and energy price data; The preprocessed multi-source energy data is divided into four types of energy points to generate an energy point set. The electricity consumption feature vector is calculated using the time feature engineering method. The feature combination method is used to calculate the feature vectors of equipment, weather and price respectively to generate a feature vector set. Use the device driver association method to build energy connections, generate energy connection sets from all energy connections, build an initial energy graph, and use the graph structure balance method to detect connectivity. Based on the detection results, perform dynamic adjustment and streamline connection operations to obtain an optimized energy graph. Collect historical multi-source energy data, pre-process and extract features as training sets; A two-layer graph convolutional network (GCN) model is constructed, and the loss function supervision task is defined. The feature vector of the training set is used as input, and iterative training is performed. The iteration is stopped after the maximum number of iterations is reached. The optimized model weight parameters are output and applied to the two-layer graph convolutional network (GCN) model. The feature vector set is used as input to obtain the final enhanced embedding set, which is added to the optimized energy graph to obtain the graph structure.

3. The energy data analysis method based on a large model and a proprietary knowledge base as claimed in claim 2, characterized in that: The generating of the initial global reasoning chain, interacting the proprietary knowledge base with the initial global reasoning chain, optimizing the initial global reasoning chain, constructing a dynamic tree-shaped reasoning structure, dynamically adjusting branches based on an artificial potential field, and optimizing the reasoning path include: Use the task prompt decomposition method to divide it into sub-queries, generate a sub-query list, and input it into the pre-configured large language model LLM with the graph structure to generate initial answers, calculate the answer confidence, and use the fixed threshold screening method to screen the initial answers whose answer confidence is greater than the corresponding answer confidence threshold to generate an initial answer set; Based on the dependencies between subqueries in the subquery list, a directed dependency graph is constructed to generate an initial global reasoning chain. The pre-trained DistilBERT model is used for each subquery to generate a question vector. Collect and preprocess the document data of the proprietary knowledge base, build a document index using FAISS, generate knowledge base vectors using the pre-trained DistilBERT model, calculate the cosine similarity between the question vector and the knowledge base vector, perform k-nearest neighbor search on the question vector using the document index built by FAISS, obtain the candidate document set, and segment it into natural paragraphs, and use the pre-trained DistilBERT model to generate paragraph vectors; Calculate the cosine similarity between the paragraph vector and the question vector to obtain the candidate answer text, use the pre-trained SpanBERT question-answering model to extract the answer fragment that best matches the subquery, and calculate the confidence score of the answer fragment, sort it in descending order to obtain the optimized answer vector, calculate the answer similarity score of the initial answer vector and the optimized answer vector, sum the answer similarity score and the confidence score of the highest answer fragment vector using the weighted summation method to obtain the comprehensive consistency score, use the empirical method to set the comprehensive consistency score threshold, compare it with the comprehensive consistency score, perform operations based on the comparison results, and obtain the optimized reasoning chain; Map each subquery-answer pair in the optimized inference chain to a node in the tree , evaluate the subquery using the preconfigured Large Language Model (LLM) With other subqueries If there is a causal dependency, add a directed edge , for each tree node Calculate the initial weight, otherwise it remains unchanged. After adding the edges, use depth-first search to detect whether there is a loop. If there is a loop, sort the initial weights of the nodes that make up the loop in descending order, remove the edges of the nodes with the lowest initial weight, and get the initial dynamic tree inference structure. ; For each tree node Calculate the initial consistency score and compare it with the set initial consistency threshold. Based on the comparison result, use depth-first search DFS to trace back to its parent node , generate a list of low consistency nodes , generating an alternative subquery , will replace the subquery Input into the pre-configured large language model LLM to generate new answers and replace the subquery with Generate new tree nodes , calculate the new consistency score , the new consistency score Consistency score with the initial Perform comparisons, and based on the differences in the comparison results, perform operations to replace tree nodes and retain tree nodes to obtain an updated dynamic tree-shaped reasoning structure , traverse the updated dynamic tree inference structure All the paths from the root to the leaves in Contains a set of nodes, for each path Calculate the composite score , and sort them in descending order, and select the path with the highest score as the optimal path.

4. The energy data analysis method based on a large model and a proprietary knowledge base as claimed in claim 3, characterized in that: The energy data analysis results obtained by using the LLM model based on the optimal reasoning path include: Collect the answers of each tree node from the optimal path, arrange the answers according to the logical order of energy data analysis, and form a list of analysis answers; The logical sequence of the energy data analysis includes checking the source node of each answer, determining its logical dependencies, inputting the analysis answer list into a preconfigured large language model (LLM), giving clear instructions based on the logical dependencies, and outputting energy analysis results and instructions.

5. The energy data analysis method based on a large model and a proprietary knowledge base as claimed in claim 4, characterized in that: The above-mentioned process of tracing back each reasoning process and optimizing the energy data analysis results based on the tracing results include: Extract data from the optimized reasoning chain CoQ, use natural language processing (NLP) for structured processing, generate corresponding traceability labels for each answer, and form a traceability set; A set of feature vectors is extracted from the graph structure, and the time decay weight is calculated for each feature using an exponential decay model. Each feature in the feature vector set is weighted using the time decay weight to optimize the energy analysis results. The final energy analysis results are associated with the traceability set to generate structured output.

6. The energy data analysis method based on a large model and a proprietary knowledge base as claimed in claim 5, characterized in that: The data generated by the storage reasoning process includes: Convert the final analysis results, traceability set, graph structure, and optimized reasoning chain into a unified JSON format, add metadata tags to each data item, use hash functions to generate unique data identifiers, and obtain processed data; The processed data is sharded according to the identifier and timestamp, and the sharded data is stored in a distributed database. Consistent hashing is used to ensure the balanced distribution of data shards among nodes. At the same time, the shard index is recorded, an inverted index is built to speed up retrieval, and an adjacency table is established for the graph structure to store edge relationships.

7. The energy data analysis method based on a large model and a proprietary knowledge base as claimed in claim 6, characterized in that: The real-time data updating system includes: Time-align the real-time data and the stored fragmented data to generate a fused data set, calculate the fusion confidence of the fused data, store the fused data set and the fusion confidence and overwrite the old records to obtain the stored new data; Calculate the significance of data changes based on the graph structure updated by the stored new data and the updated reasoning chain; The data change significance threshold is set based on experience. If the data change significance is greater than the data change significance threshold, the analysis phase is entered; otherwise, it remains unchanged.

8. An energy data analysis system based on a large model and a proprietary knowledge base, based on the energy data analysis method based on a large model and a proprietary knowledge base according to any one of claims 1 to 7, characterized in that: include, The collection and construction module is used to collect and preprocess multi-source energy data, classify the multi-source energy data into four types of energy points, calculate the characteristics of energy points and connect energy points, detect and optimize the connectivity, build a two-layer graph convolutional network GCN model, and generate a graph structure; Generate an optimization module, which is used to generate subqueries based on the graph structure, use the LLM model to obtain the initial answer, build a directed dependency graph, generate an initial global reasoning chain, interact the proprietary knowledge base with the initial global reasoning chain, and optimize the initial global reasoning chain; The reasoning chain path module is used to map the optimized reasoning chain into tree nodes, build a dynamic tree-shaped reasoning structure, perform dynamic branch adjustment based on the artificial potential field, and optimize the reasoning path; The analysis and tracing module is used to obtain energy data analysis results based on the optimal reasoning path using the LLM model, trace each reasoning process, and optimize the energy data analysis results based on the tracing results; The storage update module is used to store the data generated by the reasoning process and use real-time data to update the system.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the energy data analysis method based on a large model and a proprietary knowledge base as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the energy data analysis method based on a large model and a proprietary knowledge base as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Conversational database query method and device based on large language model agent

    CN119494401A

  • Chat integration with grid-based data structure

    US12271707B1

Cited By

  • AI question answering method and system based on video key frame and progress association

    CN120611031A

  • Text reasoning method, electronic equipment and storage medium

    CN121581244A

  • System and method for integrating artificial intelligence assistants with website building systems

    US20250252124A1