Electric power public information model optimization method and system based on large language model

Through the optimization method of power public information model based on large language models, the problem of lack of dynamic and scalability of existing CIM modeling methods is solved, and more efficient knowledge expression and application capabilities in the power field are achieved, reducing computing resources and time overhead.

CN120123490APending Publication Date: 2025-06-10NARI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510170614.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing public information model (CIM) modeling methods lack dynamics and scalability, making it difficult to effectively utilize the logical reasoning and knowledge graph extraction capabilities of large language models, and cannot meet the complex application needs in the power field.

Method used

The power public information model optimization method is adopted based on the large language model. By collecting power field data, building a pre-training corpus, segmenting and cleaning data using sliding window method, building a fine-tuning instruction data set, and fine-tuning the pre-training model is obtained by using the low-rank adaptive method LoRA to obtain the power public information large model.

Benefits of technology

It improves the CIM's knowledge expression ability, enhances the application ability of the model in the power field, significantly reduces computing resources and time overhead, and supports power applications such as fault plan generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123490A_ABST
    Figure CN120123490A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power public information model optimization method and system based on a large language model, and the method comprises the steps: collecting electric power field data, and constructing a pre-training corpus; segmenting the corpus by using a sliding window method, and removing low-quality and repeated text fragments; constructing a fine tuning instruction data set based on CIM / E and CIM / RDF specifications; and carrying out fine tuning on the pre-training model by adopting a low-rank adaptive method LoRA to obtain a large electric power public information model. By means of the logical reasoning ability of the large language model and the knowledge extraction ability of the knowledge graph, the knowledge expression ability of the public information model of the power grid regulation and control basic platform is improved, and power application of fault plan generation can be supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for optimizing a power common information model, and particularly to a method and system for optimizing a power common information model based on a large language model. Background Art

[0002] The traditional Common Information Model (CIM) is an abstract model used to describe the main objects of power enterprises, especially those related to power operation. By providing a standard method to represent power system resources using object classes, attributes, and the relationships between them, CIM aims to facilitate the integration of applications independently developed by different vendors, such as the integration between power generation or distribution management systems. By defining a common language based on CIM, it provides convenience for integration, enabling these applications or systems to access common data and exchange information without relying on the internal representation of information. However, the existing CIM modeling methods are mainly based on static modeling using UML description language, lacking dynamics and scalability. With the rapid development of large language models, it is of great practical significance to combine the logical reasoning ability of large language models and the knowledge extraction ability of knowledge graphs to enhance the knowledge expression ability of CIM. Summary of the Invention

[0003] Object of the Invention: The object of the present invention is to provide a method and system for optimizing a power common information model based on a large language model, which can enhance the knowledge expression ability of the common information model and support power applications for generating fault plans.

[0004] Technical Solution: A method for optimizing a power common information model based on a large language model according to the present invention includes the following steps:

[0005] (1) Collect power domain data and construct a pre-training corpus;

[0006] (2) Use the sliding window method to segment the corpus and remove low-quality and duplicate text segments;

[0007] (3) Based on the CIM / E and CIM / RDF specifications, construct a fine-tuning instruction data set;

[0008] (4) Use the low-rank adaptation method to fine-tune the pre-trained model to obtain a large power common information model.

[0009] Preferably, the removal of low-quality text segments in step (2) specifically means constructing prompt words to enable the large language model to judge whether the text segment belongs to the power domain, whether it contains valuable information, and whether there are logical errors. For texts that are difficult to automatically filter by the large language model, manual review of the text is used.

[0010] Preferably, the text fragments with semantic duplicates removed in step (2) are specifically to convert the text fragments into vector representations using the gte-Qwen2-7B-instruct model, calculate the cosine similarity between the text fragment vectors, and remove the semantically similar text fragments according to the similarity threshold.

[0011] Preferably, the fine-tuning instruction dataset in step (3) includes one-hop relationships, two-hop relationships, three-hop relationships, and community relationships.

[0012] Preferably, the fine-tuning in step (4) is specifically to use the low-rank decomposition method to decompose the weight matrix in the pre-trained model into two small matrices:

[0013] W LoRA = W + ΔW, ΔW = AB T

[0014] where and are trainable matrices, r represents the low-rank dimension, and r << min(d, k);

[0015] Update the parameters of A and B, keep the original model weight W frozen, and optimize using the cross-entropy loss function L:

[0016]

[0017] where y i,j represents the true label value at the position of the j-th word in the true label distribution of the i-th sample, represents the probability distribution predicted by the model, and M is the length of the question or answer;

[0018] Preferably, the one-hop relationship means that there is a direct and explicit relationship between two knowledge points; the two-hop relationship means that there is an indirect and implicit relationship between two knowledge points, connected by an intermediate knowledge point; the three-hop relationship means that there is a relatively long-distance and implicit relationship between two knowledge points, connected by three intermediate knowledge points, and one of the knowledge points is the core knowledge point; the community relationship means that three knowledge points are directly connected to each other pairwise to form a closed-loop triangular relationship.

[0019] Preferably, the knowledge points are represented as classes or object instances, and the relationships between the knowledge points are represented as associations between classes or references between object instances.

[0020] An optimization system for a power public information model based on a large language model according to the present invention includes:

[0021] A data collection module: used to collect power domain data;

[0022] Data cleaning module: used to segment and clean the corpus;

[0023] Fine-tuning instruction generation module: used to generate a diverse dataset of fine-tuning instructions;

[0024] Model training module: used to fine-tune the pre-trained model.

[0025] A computer device, characterized in that it includes one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the program is executed by the processor, the steps of the method for optimizing the power public information model based on the large language model are implemented.

[0026] A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method for optimizing the power public information model based on the large language model are implemented.

[0027] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: by combining the logical reasoning ability of the large language model and the knowledge extraction ability of the knowledge graph, the knowledge expression ability of CIM is improved; by constructing a diverse dataset of fine-tuning instructions, the application ability of the model in the power field is enhanced; the LoRA method is used for fine-tuning, significantly reducing the computational resources and time overhead. Brief Description of the Drawings

[0028] Figure 1 It is the flowchart of the method of the present invention. Detailed Embodiment

[0029] The technical solution of the present invention will be further described below with reference to the drawings.

[0030] As Figure 1 shown, a method for optimizing a power public information model based on a large language model according to the present invention includes the following steps:

[0031] (1) Collect data in the power field and construct a pre-training corpus;

[0032] The pre-training corpus in the power field is crucial for enhancing the domain-specific professional knowledge of the large language model. In this embodiment, a number of Chinese and English corpora are used, including public data from power dispatching agencies of various countries, SCADA / EMS data, academic papers, technical reports, standard documents, and books.

[0033] (2) Use the sliding window method to segment the corpus and remove low-quality and duplicate text segments;

[0034] Specifically, for the collected data, the sliding window method is used to divide the corpus into text segments within the length limit, ensuring that the sentences before and after the sliding window are completely retained to avoid information loss.

[0035] Specifically, low-quality, noisy, and text segments irrelevant to the power field are removed. Using a pre-trained corpus in the power field, prompting words are constructed to let the large language model judge whether the text segment belongs to the power field, whether it contains valuable information, whether there are logical errors, etc. For texts that are difficult to automatically filter through the large language model, manual review of the text is used. In this embodiment, deepseek-v3 is used as the large language model.

[0036] Specifically, text segments with duplicate semantics are removed. The gte-Qwen2-7B-instruct model is used to convert the text segments into vector representations, calculate the cosine similarity between the text segment vectors, and remove semantically similar text segments according to the similarity threshold.

[0037] (3) Based on the CIM / E and CIM / RDF specifications, construct a fine-tuning instruction dataset;

[0038] In this embodiment, the data relationships in the CIM model are divided into four types: one-hop relationship, two-hop relationship, three-hop relationship, and community relationship.

[0039] Specifically, in the CIM model, knowledge points can be understood as classes or object instances, and the relationships between knowledge points can be understood as associations between classes or references between object instances.

[0040] Specifically, a one-hop relationship means that there is a direct and explicit relationship between two knowledge points. In the CIM model, it is manifested as a direct association between two classes or a direct reference between two object instances.

[0041] Specifically, a two-hop relationship means that there is an indirect and implicit relationship between two knowledge points, connected by an intermediate knowledge point. In the CIM model, it is manifested as an indirect association between two classes through another class, or an indirect reference between two object instances through another object instance.

[0042] Specifically, a three-hop relationship means that there is a relatively long-distance and implicit relationship between two knowledge points, connected by three intermediate knowledge points, and one of the knowledge points is the core knowledge point. The core knowledge point refers to the knowledge point with the most connections in the knowledge graph, indicating its high importance and wide applicability in the knowledge system.

[0043] Specifically, the Community relationship means that three knowledge points are directly connected pairwise, forming a closed-loop triangular relationship, representing an explicit and strongly correlated relationship.

[0044] Based on the above four relationship types, this embodiment designs corresponding Prompt templates and examples to guide the large language model (deepseek-v3) to generate high-quality fine-tuning instructions.

[0045] Prompt 1: Generation of QA pairs in the power domain based on One-hop relationship (CIM / RDF)

[0046]

[0047] Prompt 2: Generation of QA pairs in the power domain based on Two-hop relationship (CIM / RDF)

[0048]

[0049]

[0050] Prompt 3: Generation of QA pairs in the power domain based on Three-hop relationship and core knowledge points (CIM / RDF)

[0051]

[0052] Prompt 4: Generation of QA pairs in the power domain based on Community relationship (CIM / RDF)

[0053]

[0054]

[0055] Prompt 5: Generation of device descriptions based on CIM / E data

[0056]

[0057] Prompt 6: Generation of SPARQL query statements based on CIM / RDF data

[0058]

[0059]

[0060] (4) Use the Low-Rank Adaptation method LoRA to fine-tune the pre-trained model to obtain the large power public information model.

[0061] In this embodiment, Qwen2.5-72b-instruct is selected as the pre-trained model, and the low-rank adaptation (LoRA) method is used to fine-tune the Qwen2.5-72b-instruct model. LoRA updates the model parameters by adding low-rank matrices to the weight matrices of the pre-trained model, thus significantly reducing the computational resources and time required for fine-tuning while maintaining the model performance.

[0062] Specifically, denote the power domain dataset as where, q i represents the question, a i represents the corresponding answer, and N is the total number of samples in the dataset.

[0063] Specifically, the fine-tuning method of LoRA updates the model parameters by introducing low-rank matrices on the basis of freezing the large model parameters, which specifically includes the following steps:

[0064] Model layer parameter decomposition: For the weight matrix in the pre-trained model, use the low-rank decomposition method to decompose it into two small matrices:

[0065] W LoRA =W + ΔW, ΔW = AB T

[0066] where, and are trainable matrices, r represents the low-rank dimension, and usually r << min(d, k).

[0067] Weight update method: During the fine-tuning process, only update the parameters of A and B, and keep the original model weight W frozen, thus significantly reducing the computational overhead and storage requirements.

[0068] To ensure that the model generates high-quality answers, the cross-entropy loss function L is used for optimization:

[0069]

[0070] where, y i,j represents the true label value at the position of the j-th word in the true label distribution of the i-th sample, represents the probability distribution predicted by the model, and M is the length of the question or answer.

[0071] After fine-tuning Qwen2.5-72b-instruct on the power domain dataset through LoRA, a large power public information model is obtained.

[0072] This embodiment provides an optimization system for a power public information model based on a large language model, including:

[0073] Data collection module: used to collect data in the power field;

[0074] Data cleaning module: used to segment and clean the corpus;

[0075] Fine-tuning instruction generation module: used to generate a diverse fine-tuning instruction dataset; Model training module: used to fine-tune the pre-trained model.

Claims

1. A method for optimizing a power public information model based on a large language model, characterized in that: The method comprises the following steps: (1) Collect data in the power field and build a pre-training corpus; (2) Use the sliding window method to segment the corpus and remove low-quality and repeated text segments; (3) Construct a fine-tuning instruction dataset based on CIM / E and CIM / RDF specifications; (4) A low-rank adaptive method is used to fine-tune the pre-trained model to obtain a large model of power public information.

2. The method for optimizing a power public information model based on a large language model according to claim 1, characterized in that: The removal of low-quality text fragments in step (2) is specifically to construct prompt words so that the large language model can determine whether the text fragment belongs to the power field, whether it contains valuable information, and whether there are logical errors. For texts that are difficult to be automatically filtered by the large language model, manual review of the text is used.

3. The method for optimizing a power public information model based on a large language model according to claim 1, characterized in that: The step (2) of removing semantically repeated text segments specifically involves using the gte-Qwen2-7B-instruct model to convert the text segments into vector representations, calculating the cosine similarity between the text segment vectors, and removing semantically similar text segments according to a similarity threshold.

4. The method for optimizing a power public information model based on a large language model according to claim 1, characterized in that: The fine-tuning instruction data set in step (3) includes one-hop relations, two-hop relations, three-hop relations and community relations.

5. The method for optimizing a power public information model based on a large language model according to claim 1, characterized in that: The fine-tuning described in step (4) is specifically to adjust the weight matrix in the pre-trained model Use low-rank decomposition to decompose it into two small matrices: W LoRA =W+ΔW,ΔW=AB T in, and is a trainable matrix, r represents the low-rank dimension, r<<min(d,k); Update the parameters of A and B, keep the original model weight W frozen, and use the cross entropy loss function L for optimization: Among them, y i,j Represents the true label value of the position of the jth word in the true label distribution of the i-th sample, represents the probability distribution predicted by the model, and M is the length of the question or answer.

6. The method for optimizing a power public information model based on a large language model according to claim 4, characterized in that: The one-hop relationship refers to a direct and explicit relationship between two knowledge points; the two-hop relationship refers to an indirect and implicit relationship between two knowledge points, which are connected through an intermediate knowledge point; the three-hop relationship refers to a long-distance and implicit relationship between two knowledge points, which are connected through three intermediate knowledge points, and one of the knowledge points is a core knowledge point; the community relationship refers to a triangular relationship in which three knowledge points are directly connected in pairs to form a closed loop.

7. The method for optimizing the electric power public information model based on a large language model according to claim 6, characterized in that: The knowledge points are represented as classes or object instances, and the relationships between the knowledge points are represented as associations between classes or references between object instances.

8. A power public information model optimization system based on a large language model, characterized in that: include: Data collection module: used to collect data in the power field; Data cleaning module: used to segment and clean the corpus; Fine-tuning instruction generation module: used to generate a diverse set of fine-tuning instruction data; Model training module: used to fine-tune the pre-trained model.

9. A computer device, characterized in that: It comprises one or more processors, a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of an electric power public information model optimization method based on a large language model as described in any one of claims 1-7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for optimizing a power public information model based on a large language model as described in any one of claims 1 to 7 are implemented.