Large model retrieval enhancement generation method oriented to policies in electricity price field
By building a tree-structured electricity price policy knowledge base and combining the search and enhancement generation method of the RAG framework, the problem of difficulty for users to obtain and understand electricity price policy information is solved, and an accurate and real-time understanding of electricity price policy is achieved, and technical support is provided for decision-making in the power industry.
Patent Information
- Application Number
- CN202510423281.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
It is difficult for ordinary users to obtain electricity price policy information conveniently and accurately, and due to the professionalism and frequent updates of electricity price policy documents, it is difficult for users to understand and keep up with the latest policy trends.
Build a large-scale retrieval enhancement generation method for the field of electricity prices. By building a professional tree-structured knowledge base, and combining the RAG framework, using keyword matching, semantic understanding and RRF algorithms, realizing instant retrieval and fusion of information.
It improves the understanding and accurate acquisition of electricity price policies, ensures the accuracy and real-timeness of the answers, and provides technical support for decision-making support in the power industry.
Smart Images

Figure CN119940512A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer science and artificial intelligence technology, and specifically relates to a large-model retrieval enhancement generation method for electricity price policies. Background Art
[0002] Electricity price policies play a decisive role in the setting and adjustment of electricity prices. However, due to problems such as information dispersion, access restrictions and technical barriers, it is difficult for ordinary users to obtain electricity price policy information conveniently and accurately. At the same time, since electricity price policy documents often use professional terms and usually contain a lot of background information, technical details and appendices, users also have many difficulties in understanding the documents and accurately calculating electricity prices after obtaining electricity price policy documents. In addition, due to the rapid changes in electricity price policies and market environment, policy documents are often updated, and it is difficult for ordinary users to keep up with the latest policy trends. Therefore, building a detailed and complete electricity price knowledge base and strengthening the accuracy of questions and answers based on this knowledge base has become a key issue in this method.
[0003] By training on massive amounts of text data, large language models have learned deep patterns and rich knowledge of language, and have the ability to understand and generate across domains. However, in specific field applications such as electricity price policies, to achieve professional-level accuracy and practicality, the model needs to have a deep understanding of domain-specific expertise. Researchers have explored a variety of strategies, including domain-adaptive training, continuous learning, and integration with external knowledge bases.
[0004] The development of prompt engineering can help users better use large models in various scenarios and research fields. The traditional prompt engineering method is templated prompts, but for problems with complex domain knowledge, more precise prompts are needed to improve the model's understanding ability and generation quality. In recent years, prompt engineering has gradually entered the data-driven and context-aware stage. Using machine learning algorithms to select the best prompt words, adjust the prompt structure, and use technologies such as reinforcement learning and genetic algorithms, as well as combining contextual information, considering the dynamic adjustment of prompts in different contexts, improves the relevance and accuracy of the generated results.
[0005] RAG (Retrieval-Augmented Generation) is a technology that combines information retrieval and natural language generation. It aims to enhance the generation capability of the model by instantly retrieving relevant information from external databases. In application scenarios in specific fields, the RAG method can retrieve the most relevant document fragments from the pre-built electricity price policy knowledge base, provide prompts for the final answer of the large model, and generate answers or strategic recommendations that are both accurate and in line with the current policy background. RAG not only improves the accuracy of the model's answers to professional questions, but also ensures the real-time and effectiveness of the recommendations, providing technical support for decision-making support in the power industry. Summary of the invention
[0006] The purpose of the present invention is to provide a large-scale model retrieval enhancement generation method for electricity price policy.
[0007] The method of the present invention is specifically as follows: Step (1) constructing a specialized knowledge base; constructing a tree structure based on the hierarchical relationship between file titles to achieve structured storage of documents; Step (2) enhanced retrieval: for a question, relevant information is retrieved from the knowledge base according to the requirements of the large model to integrate internal and external knowledge and improve the accuracy of decision output; Step (3) uses the big model to integrate external knowledge and internal model knowledge to deal with issues related to electricity prices, and the big model answers professional questions related to electricity prices.
[0008] Furthermore, the method of constructing the knowledge base in step (1) is as follows: (1-1) Initialize the knowledge base structure: define the document collection , for each document Set the root node. , The total number of documents, initialized to a dictionary of basic attributes, including the current node name and a collection of child nodes , each child node Same structure as the corresponding root node; (1-2) Document content analysis: As a collection of paragraphs, , Indicates paragraphs, , For Documentation The number of paragraphs; using regular expressions right Make a judgment, if If the match is successful The value of , otherwise : ; The value of When calculating the level of the title, a new structure is created and Same nodes; control of the node hierarchy is done by the stack The push and pop operations are implemented according to the last-in-first-out principle; the top element of the stack is popped up according to the level until the current processing level, and then the new node is added to the child node list of the current top node of the stack, and the new node is pushed into the stack; when The value of and When it is not empty, it means that the current content is the non-text part of the file. It traverses the text of this part and adds the reference list according to the sequence number and the key value of the reference content. In the corresponding root node attributes; after processing, return the tree structured file; After completing the structured storage of knowledge base documents, extract the file content and vectorize the text data; Split into multiple words, map to a high-dimensional vector space, and convert the document All words in are defined as the word set , for each word Find the corresponding row in the predefined embedding matrix E and get The corresponding embedding vector , whose value is , Indicates words, , For Documentation The number of words.
[0009] Furthermore, the method for enhancing the retrieval in step (2) is as follows: (2-1) Measure the importance of a word based on its frequency of occurrence in the document and its inverse document frequency; In the documentation Match score , Expressive words In the documentation The frequency of occurrence in Expressive words The inverse document frequency of , Indicates that it contains words The number of documents; sort by matching score from high to low, and get the top results, and the corresponding words are taken as the first result set , which is the keyword set; (2-2) The cosine similarity method is used to calculate the semantic similarity between two keyword vectors. and keywords The semantic similarity of , and Respectively represent keywords and words The word vector of Represents the vector dimension; sort by semantic similarity from high to low, and get the top results, and the corresponding keywords are used as the second result set ; (2-3) Use the RRF algorithm to analyze the first result set and the second result set Rearrange: Keywords The word vector , if it ranks first in another result set The inverse ranking score is If it is not in the other result set, the inverse ranking score is 0. The inverse ranking scores of each keyword in the two sets are recorded as and ; Calculate for each document Fusion score , and Respectively represent the relative importance of keyword matching and semantic understanding, and ; Finally, sort by fusion score from high to low to get the top result set .
[0010] Further, step (3) is specifically: All word embedding vectors after segmentation and mapping are combined into a document vector , the large model is based on the document vector and the result set Generate best answer , Indicates that after a given question and retrieval, the large model generates an answer The probability of .
[0011] The method of the present invention builds a highly specialized knowledge base, which widely collects legal provisions and policy documents related to energy pricing, providing great convenience for retrieval and in-depth analysis. When solving professional problems in the field of electricity prices, the questions are answered based on the big model, and the RAG framework is used to retrieve the most relevant document fragments from the constructed electricity price policy knowledge base, which prompts the final answer of the big model, generates accurate answers that are in line with the current policy background, and provides technical support for decision-making support in the power industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic diagram of the overall method of the present invention; Figure 2 This is a schematic diagram of the tree structure of the knowledge base; Figure 3 Enhanced schematic diagram for retrieval; Figure 4 This is a diagram of the applet's user interface. DETAILED DESCRIPTION
[0013] A large-scale retrieval enhancement generation method for electricity price policies. This method builds a specialized knowledge base and constructs a tree structure based on the hierarchical relationship between file titles to achieve structured storage of documents. When solving electricity price problems, the RAG framework is introduced. According to the needs of the large model, keyword matching, semantic understanding and RRF algorithms are used to instantly retrieve relevant information from the knowledge base, integrate internal and external knowledge, and give the final answer. The overall method is as follows Figure 1 As shown: Step (1) construct a specialized knowledge base; the knowledge base is constructed by systematically integrating and analyzing relevant policy documents, covering legal provisions and policy guidance. According to the hierarchical relationship between document titles, a tree structure is constructed to achieve structured storage of documents. The method for constructing a knowledge base is as follows: (1-1) Initialize the knowledge base structure: define the document collection , for each document Set the root node. , The total number of documents, initialized to a dictionary of basic attributes, including the current node name and a collection of child nodes , each child node Same structure as the corresponding root node.
[0014] (1-2) Document content analysis: As a collection of paragraphs, , Indicates paragraphs, , For Documentation The number of paragraphs; using regular expressions right Make a judgment, if If the match is successful The value of , otherwise : ; The value of When calculating the level of the title, a new structure is created and Same nodes; control of the node hierarchy is done by the stack The push and pop operations are implemented according to the last-in-first-out principle; the top element of the stack is popped up according to the level until the current processing level, and then the new node is added to the child node list of the current top node of the stack, and the new node is pushed into the stack; when The value of and When it is not empty, it means that the current content is the non-text part of the file. It traverses the text of this part and adds the reference list according to the sequence number and the key value of the reference content. The corresponding root node attributes. After processing, the tree structured file is returned, such as Figure 2 shown.
[0015] After completing the structured storage of knowledge base documents, extract the file content and vectorize the text data; Split into multiple words, map to a high-dimensional vector space, and convert the document All words in are defined as the word set , for each word Find the corresponding row in the predefined embedding matrix E and get The corresponding embedding vector , whose value is , Indicates words, , For Documentation The number of words.
[0016] Step (2) Enhanced retrieval: For a problem, it is necessary to retrieve relevant information from the above knowledge base in real time according to the needs of the large model to integrate internal and external knowledge and improve the accuracy of decision output. Figure 3 As shown, the method of enhancing retrieval is as follows: (2-1) The importance of a word is measured based on its frequency of occurrence in the document and its inverse document frequency. In the documentation Match score , Expressive words In the documentation The frequency of occurrence in Expressive words The inverse document frequency of , Indicates that it contains words The number of documents. Sort by matching score from high to low, and get the top results, and the corresponding words are taken as the first result set , which is the keyword set.
[0017] (2-2) The cosine similarity method is used to calculate the semantic similarity between two keyword vectors. and keywords The semantic similarity of , and Respectively represent keywords and words The word vector of Represents the vector dimension. Arrange from high to low according to the semantic similarity, and get the top results, and the corresponding keywords are used as the second result set .
[0018] (2-3) Use the RRF algorithm to analyze the first result set Rearrange: Keywords The word vector , if it ranks first in another result set The inverse ranking score is If it is not in the other result set, the inverse ranking score is 0. The inverse ranking scores of each keyword in the two sets are recorded as and ; Calculate for each document Fusion score , and Respectively represent the relative importance of keyword matching and semantic understanding, and Finally, sort by fusion score from high to low to get the top result set .
[0019] Step (3) Use the big model to integrate external knowledge and internal model knowledge to deal with issues related to electricity prices. The big model answers professional questions related to electricity prices; All word embedding vectors after segmentation and mapping are combined into a document vector , the large model is based on the document vector and the result set Generate best answer , Indicates that after a given question and retrieval, the large model generates an answer The probability of .
[0020] In order to facilitate user experience, the large model retrieval enhancement generation method for electricity price policy is used in WeChat mini program. Users can log in to the mini program and ask questions related to electricity price. The mini program user interface is as follows: Figure 4 shown.
Claims
1. A large-scale model retrieval and enhanced generation method for electricity price policies, characterized by: Step (1) constructing a specialized knowledge base; constructing a tree structure based on the hierarchical relationship between file titles to achieve structured storage of documents; Step (2) Enhanced retrieval; For a problem, relevant information is retrieved from the knowledge base according to the needs of the big model to integrate internal and external knowledge and improve the accuracy of decision output; Step (3) uses the big model to integrate external knowledge and internal model knowledge to deal with issues related to electricity prices, and the big model answers professional questions related to electricity prices.
2. The large-scale model retrieval enhancement generation method for electricity price policy according to claim 1 is characterized in that: Step (1) The method of constructing the knowledge base is as follows: (1-1) Initialize the knowledge base structure: define the document collection , for each document Set the root node. , The total number of documents, initialized to a dictionary of basic attributes, including the current node name and a collection of child nodes , each child node Same structure as the corresponding root node; (1-2) Document content analysis: As a collection of paragraphs, , Indicates paragraphs, , For Documentation The number of paragraphs; using regular expressions right Make a judgment, if If the match is successful The value of , otherwise : ; The value of When calculating the level of the title, a new structure is created and Same nodes; The node level is controlled by the stack The push and pop operations are implemented according to the last-in-first-out principle; the top element of the stack is popped up according to the level until the current processing level, and then the new node is added to the child node list of the current top node of the stack, and the new node is pushed into the stack; when The value of and When it is not empty, it means that the current content is the non-text part of the file. It traverses the text of this part and adds the reference list according to the sequence number and the key value of the reference content. In the corresponding root node attributes; after processing, return the tree structured file; After completing the structured storage of knowledge base documents, extract the file content and vectorize the text data; Split into multiple words, map to a high-dimensional vector space, and convert the document All words in are defined as the word set , for each word Find the corresponding row in the predefined embedding matrix E and get The corresponding embedding vector , whose value is , Indicates words, , For Documentation The number of words.
3. The large-scale model retrieval enhancement generation method for electricity price policy according to claim 1 is characterized in that: The method of enhancing the retrieval in step (2) is as follows: (2-1) Measure the importance of a word based on its frequency of occurrence in the document and its inverse document frequency; In the documentation Match score , Expressive words In the documentation The frequency of occurrence in Expressive words The inverse document frequency of , Indicates that it contains words The number of documents; sort by matching score from high to low, and get the top results, and the corresponding words are taken as the first result set , which is the keyword set; (2-2) The cosine similarity method is used to calculate the semantic similarity between two keyword vectors. and keywords The semantic similarity of , and Respectively represent keywords and words The word vector of Represents the vector dimension; sort by semantic similarity from high to low, and get the top results, and the corresponding keywords are used as the second result set ; (2-3) Use the RRF algorithm to analyze the first result set and the second result set Rearrange: Keywords The word vector , if it ranks first in another result set The inverse ranking score is If it is not in the other result set, the inverse ranking score is 0. The inverse ranking scores of each keyword in the two sets are recorded as and ; Calculate for each document Fusion score , and Respectively represent the relative importance of keyword matching and semantic understanding, and ; Finally, sort by fusion score from high to low to get the top result set .
4. The large-scale model retrieval enhancement generation method for electricity price policy according to claim 1 is characterized in that: Step (3) is as follows: All word embedding vectors after segmentation and mapping are combined into a document vector , the large model is based on the document vector and the result set Generate best answer , Indicates that after a given question and retrieval, the large model generates an answer The probability of .
Citation Information
Patent Citations
RAG knowledge question-answering method and device based on fusion vector and keyword retrieval
CN117951274A
Large model knowledge question-answering method and device for power field
CN118585626A
Policy knowledge question and answer method and device and storage medium
CN118797021A
Short text query expansion enhancement retrieval method based on knowledge base hierarchical tree structure
CN118861088A
Facility agriculture intelligent question answering method based on big language model retrieval enhancement generation
CN118916457A
Cited By
Multi-level knowledge base construction method for RAG
CN120874989A
Cross-domain official document automatic generation method and system, storage medium and program product
CN120975065A