Intelligent questioning and answering method for knowledge in early stage of power grid engineering project based on knowledge graph and GPT technology
By building an intelligent question-and-answer method based on knowledge graph and GPT technology, the problem of insufficient informatization and intelligence in the early stage of knowledge management in power grid engineering projects is solved, efficient knowledge sharing and accurate intelligent question-and-answer are achieved, and management efficiency and reliability are improved.
Patent Information
- Application Number
- CN202510261510.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-04
AI Technical Summary
The knowledge management in the early stages of power grid engineering projects has insufficient information and intelligence, which leads to problems such as difficulty in obtaining information, difficulty in sharing knowledge, and knowledge loss. The randomness and hallucination of the GPT model answers lead to incorrect answers.
Build an intelligent question-answer method based on knowledge graph and GPT technology. By analyzing the current status of knowledge management business in the early stages of power grid engineering projects, building a knowledge graph pattern layer, forming a logical relationship graph, uploading management files and quickly extracting and sorting them, combining expert experience to obtain entity-relationship-entity triplets, establishing a knowledge graph, and building a large language model based on GPT technology for pre-training and fine-tuning of instructions, optimizing the intelligent question-answer process.
It has improved the level of informatization of knowledge management in the early stages of power grid engineering projects, eliminated the randomness and illusion of model answers, improved the reliability and accuracy of intelligent question-and-answer, promoted knowledge sharing and innovation, and reduced management costs and risks.
Smart Images

Figure CN120258131A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and specifically relates to a knowledge intelligent question-answering method for the early stage of a power grid engineering project based on knowledge graph and GPT technology. Background Art
[0002] Knowledge graph technology is a semantic knowledge base that stores and represents entities and their relationships through a graph structure. It can transform complex data and information into a structured form that is easy to understand and process, and supports a variety of applications such as intelligent search, recommendation systems, natural language processing, etc., thereby improving the accuracy and efficiency of information retrieval.
[0003] The GPT model, or Generative Pre-trained Transformer, is an advanced artificial intelligence model based on the Transformer architecture, developed by OpenAI. It is pre-trained on massive text data through unsupervised learning and can generate coherent and logical text. The core advantage of GPT lies in its pre-training-fine-tuning training strategy, which enables it to have general language capabilities while also being able to flexibly adapt to various downstream tasks. With its powerful natural language processing capabilities, the GPT model has demonstrated outstanding performance in multiple fields such as text generation and dialogue systems.
[0004] With the rapid economic development and the growing demand for electricity, the scale and complexity of power grid construction, as an important part of national infrastructure construction, continues to increase. As the starting point of the entire project, the early stage of the power grid project involves multiple key links such as project planning, feasibility study, and preliminary design. Its management efficiency and quality directly affect the smooth implementation of subsequent projects and the overall benefits of the project. However, there are a series of urgent problems in the current knowledge management of the early stage of the power grid, which seriously restrict the improvement of the informatization and intelligence level of the management work in the early stage. Knowledge graph technology can standardize the storage of complex knowledge, and mine the relationship between knowledge and establish knowledge links. GPT technology can build a large language model based on algorithms and application scenarios, and pre-train through a certain amount of data input to realize intelligent question and answer of knowledge in the early stage of the transmission network project.
[0005] Therefore, knowledge graphs and GPT technologies can be applied to effectively solve problems such as insufficient information management of knowledge in the early stages of power grid projects, difficulty in acquiring knowledge, difficulty in sharing and transferring knowledge, and knowledge loss. Summary of the invention
[0006] The purpose of the present invention is to provide a knowledge intelligent question-answering method for the early stage of power grid engineering projects based on knowledge graph and GPT technology, so as to solve the problem of wrong answers caused by the randomness and hallucination of the model's own answers, and improve the reliability of intelligent question-answering.
[0007] To achieve the above objectives, the technical solution of the present invention is as follows:
[0008] A method for intelligent knowledge Q&A in the early stage of power grid engineering projects based on knowledge graphs and GPT technology, characterized in that the method includes:
[0009] Analyze the current situation of knowledge management business in the early stage of power grid engineering projects, construct a knowledge graph schema layer based on business characteristics to form a logical relationship graph, which includes knowledge entities as business control nodes and other knowledge entities related to business control nodes. In the present invention, the knowledge entities as business control nodes can be grassland protection, cultural relics protection, site selection planning, etc., and the knowledge entities related to grassland protection, cultural relics protection, and site selection planning can be relevant national laws and regulations, State Grid management documents, local management documents, upfront fees, approval departments, responsible departments, etc.;
[0010] Upload the management document materials in the early stage of power grid engineering projects, and use a knowledge extraction algorithm model to quickly extract and organize the document materials. Store documents of the same type under the same knowledge entity. Specifically, all national laws and regulations documents can be stored under the knowledge entity of national laws and regulations, all local management documents can be stored under the knowledge entity of local management documents, and other types of documents can be stored under the corresponding knowledge entities to obtain the knowledge entities for upfront work and a structured knowledge base for upfront work;
[0011] Combine the business control experience of experts in the early stage to obtain entity-relationship-entity triples; establish a related relationship table for knowledge entities, input for modeling, and construct a knowledge graph for the early stage of power grid engineering projects;
[0012] Build a large language model for the early stage of power grid engineering projects based on GPT technology, use the knowledge graph network as the text input of the GPT model for model training, and optimize and improve the accuracy and reliability of the model during the intelligent Q&A process.
[0013] Furthermore, the construction of the knowledge graph schema layer includes determining the knowledge entities in the early stage of power grid engineering projects. The knowledge entities include business control nodes, responsible departments, approval departments, national laws and regulations, unit internal management documents, local management documents, and upfront fees related to the business control nodes. Use data mining technology to identify the text categories of knowledge entities, obtain the association relationships between entities, and construct the knowledge graph schema layer for the early stage of power grid engineering projects.
[0014] Furthermore, the extraction and collation of the said document materials include using the BIOES annotation method to determine the document type; or adopting a solution that combines a Chinese word segmentation model with a Transformer model to train based on manually annotated data, make predictions on unannotated document data, and then input the results into a large language model (LLM). Based on the set prompts including entity types and relationship construction formats, etc., an entity relationship output in standard format is obtained.
[0015] Manual annotation means using the BIOES annotation method to clearly represent the internal and external relationships in the entity, and then setting extraction rules for common entity and relationship types in the document;
[0016] Integrating and extracting business data based on rules, deep learning, and a large language model means using a combination of a Chinese word segmentation model for entities and relationships with relatively high independence and low occurrence frequency.
[0017] The solution of the Transformer model is trained based on manually annotated data, makes predictions on unannotated document data, and finally inputs the results into the large language model (LLM). The large language model (LLM) uses Qwen2.5 - 7B - Instruct of Alibaba. Based on the set prompts including entity types and relationship construction formats, an entity relationship output in standard format is obtained and processed to form a database.
[0018] Furthermore, the pre - work structured knowledge base includes a pre - work business control knowledge base, a system document knowledge base, a system rule knowledge base, and a pre - work cost knowledge base
[0019] Business entities are extracted from the pre - work business control knowledge base: business control nodes, responsible departments, and approval departments. At the same time, with the business control node as the core, the association relationships with other business entities are shown;
[0020] Business entities are extracted from the system document knowledge base, and the extracted business entities include national laws and regulations, State Grid management documents, and local management documents;
[0021] The system rule knowledge base is a refinement of the system document knowledge base in the pre - project phase of the power grid project, that is, a data set of rule entries of national laws and regulations, State Grid management documents, and local management documents related to business control nodes, which is quickly extracted according to the document materials. The structured form includes rule names, rule types, keywords, and source documents;
[0022] The preliminary cost knowledge base is a collection of knowledge data on the management regulations and business processes of the preliminary costs of power grid engineering projects. It is quickly extracted from document materials. The structured form includes the name of the preliminary cost, the principle of cost calculation, and relevant institutional documents. The business entity of the preliminary cost can be extracted from the preliminary cost knowledge base.
[0023] Furthermore, the obtained entity-relationship-entity triple includes connecting related knowledge entities through keywords. The structured knowledge entity database corresponds to the input mode layer to obtain the knowledge graph network of the preliminary stage of power grid engineering projects. In the knowledge graph network, the knowledge structure system and relevance can be clearly and intuitively seen.
[0024] Furthermore, the construction of the model includes pre-training and instruction fine-tuning. The pre-training includes:
[0025] Unify the collected data into a standard format;
[0026] Train a Qwen2.5-7B-Instruct large language model with the collected data. Take the constructed knowledge base for business control in the preliminary work, institutional document knowledge base, institutional rule knowledge base, preliminary cost knowledge base, and knowledge graph network of the preliminary stage as the input of the GPT model for pre-training, so that the model can better understand the professional terms related to knowledge management in the preliminary stage of power grid engineering projects and can generate coherent and accurate answers based on the knowledge graph.
[0027] Instruction fine-tuning is to use the Fine-tuning method to adjust the parameters of some layers of the model.
[0028] Furthermore, the model training includes optimizing the large language model by adopting a comprehensive construction optimization method combining enhanced generation and prompt engineering;
[0029] For files with different contents in different projects, after data cleaning, classify and construct the knowledge base, use the retrieval-augmented generation (RAG) technology for knowledge retrieval, combine RAG and the knowledge graph, form an array of all the information retrieved from the knowledge graph, calculate the two-norm of the array vector, and return all the content exceeding the set threshold to the RAG knowledge base for retrieval to obtain the final result. The threshold is the distance from the user query vector. If it exceeds the threshold, the distance is too large, and it is determined that the retrieved information does not match the user query relationship and is not used as the basis for generating an answer.
[0030] The beneficial effects of the present invention are:
[0031] The present invention uses knowledge graph technology to standardize the knowledge in the early stage of power grid engineering projects, excavate the correlation relationships between knowledge entities in the early stage, and construct a knowledge network. A large language model for knowledge Q&A in the early stage of power grid engineering projects is established using GPT technology. The knowledge graph network is used as the input of the large language model. According to the question, the model algorithm will retrieve relevant knowledge entities and their correlation relationships from the knowledge graph based on keywords. Based on the retrieval results of the knowledge graph and adjusted by the model algorithm, the answer text is generated, which can eliminate the randomness of the model's own answers and the incorrect answers caused by hallucinations, improve the reliability of intelligent Q&A, and enhance the informatization level of knowledge management. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is the flow chart of the knowledge intelligent Q&A method in the early stage of the power grid engineering project of the present invention.
[0033] Figure 2 It is the relationship diagram of the knowledge entities and the relationships between entities of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to make the objectives, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0035] The explanations of the professional terms in the present invention are as follows:
[0036] BIOES annotation method: The BIOES annotation method is a sequence annotation method used for named entity recognition (NER) in natural language processing (NLP). It identifies named entities and their boundaries in the text through a series of tags.
[0037] Fine-tuning method: Fine-tuning is a deep learning technique used to further adjust the model parameters on the basis of a pre-trained model to adapt to specific downstream tasks. This method usually involves pre-training a model on a large-scale dataset to enable it to learn general language features, and then continuing to train on a smaller dataset specific to the task, so that the model can learn specific features related to the task. Through Fine-tuning, the model can retain the extensive knowledge learned in the pre-training stage while adapting to the specific requirements of the new task, thus achieving better performance in various natural language processing tasks.
[0038] Retrieval-Augmented Generation (RAG) Technology: Retrieval-Augmented Generation (RAG) technology is an artificial intelligence technology that combines information retrieval and language generation models. It retrieves relevant information from an external knowledge base and uses it as context input to assist large language models (LLMs) in generating more accurate and richer answers. RAG typically consists of two stages: first, retrieving context-related information, and then using the retrieved knowledge to guide the generation process. This approach enables the model to utilize information from external data sources, improving the accuracy and relevance of answers, which is particularly important when dealing with questions that require domain-specific knowledge.
[0039] A method for intelligent knowledge-based Q&A in the pre-project stage of power grid engineering projects based on knowledge graphs and GPT technology includes the following steps: Analyze the current situation of knowledge management in the pre-project stage of power grid engineering projects, construct a knowledge graph schema layer based on business characteristics to form a logical relationship diagram; Upload the management document materials in the pre-project stage of power grid engineering projects, and use a knowledge extraction algorithm model to quickly extract and organize the documents to obtain pre-project knowledge entities and a pre-project structured knowledge base; Combine the pre-project business control experience precipitated by experts to obtain entity-relationship-entity triples. Figure 2 Embody the triples, for example, the entity is a business control node - the relationship is business approval - the entity is an approval department, or the entity is a business control node - the relationship is relevant local documents - the entity is a local management document, etc., and the entities corresponding to the relationship will be concatenated; Establish an entity-related relationship table, input for modeling, and construct a knowledge graph for the pre-project stage of power grid engineering projects; At the same time, construct a large language model for the pre-project stage of power grid engineering projects based on GPT technology, use the knowledge graph network as the text input of the GPT model for pre-training, and optimize and improve the accuracy and reliability of the model during the intelligent Q&A process. The entire process of the intelligent knowledge-based Q&A method for the pre-project stage of power grid engineering projects is as Figure 1 shown.
[0040] 1. Current Situation of Knowledge Management in the Pre-project Stage of Power Grid Engineering Projects
[0041] The pre-project stage of power grid engineering projects, as the starting point of the entire project, involves multiple key links such as project planning, feasibility study, and preliminary design. Its management efficiency and quality directly affect the smooth implementation of subsequent projects and the overall benefits of the project. However, there are a series of problems in the current knowledge management of power grid pre-project work that urgently need to be solved, which seriously restricts the improvement of the informatization and intelligent level of pre-project management work.
[0042] (1) Trivial work content and information asymmetry: The preliminary work of power grid construction involves multiple professional fields and a large number of management documents. These documents have a large amount of information, and the updates are not standardized and timely, resulting in serious information asymmetry. When dealing with preliminary work, managers often face difficulties in obtaining information and time-consuming in searching for materials, which seriously affects work efficiency.
[0043] (2) Insufficient organization and filing of materials and informatization management: The preliminary materials include national and local laws and regulations, relevant rules and regulations and policies of the State Grid and provincial companies. These materials are important bases for the preliminary stage of power grid engineering projects. However, at present, these materials have not been effectively organized, filed and managed informatization, and it is difficult to form a unified and standardized knowledge system, which brings great inconvenience to access and utilization. In the present invention, the State Grid management documents generally refer to the internal management documents of the unit.
[0044] (3) Lack of a unified knowledge sharing platform: There is a lack of a unified knowledge sharing platform within the power grid enterprise, resulting in unsmooth information transmission channels for knowledge sharing, and it is difficult to promote and popularize knowledge. There is a lack of effective knowledge exchange and sharing mechanisms among different departments and specialties, which restricts the innovation and utilization efficiency of knowledge.
[0045] (4) Management experience not effectively precipitated: There are a large number of power grid projects, but the management experience accumulated during the project implementation has not been effectively precipitated into the company's data assets. These valuable experience and knowledge have not been fully explored and utilized, and have not brought opportunities for the enterprise to reduce costs and create value-added efficiency.
[0046] (5) Non-standardized and unified processing standards: For the handling of the same thing, different documents may have multiple processing standards, resulting in difficulty in standardization and unification in actual operation. This inconsistency in processing standards brings chaos and uncertainty to the work, increasing the management difficulty and risk.
[0047] (6) Knowledge loss caused by personnel flow: During the personnel flow process, the knowledge transfer in the handover process is not in place, resulting in a large amount of valuable knowledge loss. This knowledge loss not only affects the continuity and stability of work, but also brings serious challenges to the enterprise's knowledge accumulation and inheritance.
[0048] In response to the above problems and challenges, by constructing a knowledge graph in the field of power grid engineering and using GPT technology to achieve functions such as natural language processing, the informatization and intelligent level of management work in the preliminary stage can be significantly improved, work efficiency can be increased, knowledge sharing and innovation can be promoted, management experience can be precipitated and costs can be reduced, processing standards can be standardized and risks can be reduced, and knowledge loss can be prevented and inheritance efficiency can be improved. This will help to promote the digital transformation and sustainable development of the power industry and provide support for energy security and economic development.
[0049] 2. Knowledge Graph Construction in the Pre - project Phase of Power Grid Engineering Projects
[0050] 2.1 Construction of the Schema Layer
[0051] According to the knowledge management business process and characteristics in the pre - project phase of power grid engineering projects, the knowledge entities in the pre - project phase of power grid engineering projects are analyzed and obtained: business control nodes, responsible departments, approval departments, national laws and regulations, State Grid management documents, local management documents, and pre - project fees. The text is identified using data mining techniques to obtain the association relationships between entities, and the schema layer of the knowledge graph for the pre - project phase of power grid engineering projects is constructed. In the present invention, the association rule technique in data mining is adopted. Regarding association rules, for example, the business control node is set as grassland protection. Taking grassland protection as a keyword, other knowledge entities related to grassland protection are obtained using data mining techniques, and the business control node grassland protection is connected with other knowledge entities related to grassland protection. The data mining in the present invention is specifically embodied as frequent item mining. By identifying and analyzing the key business knowledge entities and association relationships in the business process documents and related files, and supplemented by expert experience summary, the final entity - relationship - entity triple relationship can be obtained, which is also the schema layer for constructing the knowledge graph.
[0052] 2.2 Construction of the Data Layer
[0053] During the construction of the data layer, the data sources come from a large number of national regulatory documents, management specifications for the pre - project phase of transmission and transformation projects, State Grid management documents, etc., including structured and unstructured documents. In order to quickly extract business entity data from a large number of files, an algorithm model constructed by ourselves is adopted. In the present invention, the extraction standard for extracting business entity data from a large number of files is set according to the file type, including extracting the current statement by keyword, segment extraction, extraction by chapter number, etc.
[0054] The naming of knowledge entities and relationships in documents related to power grid engineering projects is highly professional. For example, some entities and relationships are named as: relevant State Grid documents or relevant upfront costs. The document naming indicates the content involved in the document, and this kind of naming is highly professional. Using traditional keyword extraction algorithms based on public datasets (such as TF-IDF, LDA, TextRank, etc.) has poor effects, and it is also difficult to obtain accurate professional word outputs directly using large models. Therefore, an extraction scheme combining manual annotation + extraction based on rules and deep learning + integration of large language models is adopted. First, a small amount of document data is manually annotated. Since there may be nested relationships between entities such as departments, institutions, and projects in the documents, the BIOES annotation method is adopted to clearly represent the internal and external relationships in the entities. Then, extraction rules are set for common entity and relationship types in the documents. The extraction rules can be, for example, department and institution words, project words, and specification clause words. The BIOES annotation method is used for entity naming recognition here, and it is necessary to mark which are entity names such as project names, cost names, and specification regulations. Regarding the relationship types in the present invention, as Figure 2 shown, the relationship types between business control nodes and other entities include relevant State Grid documents, relevant laws and regulations, relevant local documents, relevant upfront costs, business approvals, business responsibilities, etc. For entities and relationships with high independence and low occurrence frequencies, a scheme of Chinese word segmentation model + Transformer model is used to train based on the manually annotated data and make predictions in the unannotated document data. Finally, the results are input into a large language model (LLM), and based on the set prompts (including entity types and relationship construction formats, etc.), entity relationship outputs in standard formats are obtained, and a database is formed after processing. The scheme of Chinese word segmentation model + Transformer model training based on the manually annotated data can be to perform entity naming recognition through the Chinese word segmentation model + Transformer model, and compare the entity naming recognition results obtained by the Chinese word segmentation model + Transformer model with the manually annotated data to train the Chinese word segmentation model +
[0055] Transformer model and improve the accuracy of entity naming recognition by the Chinese word segmentation model + Transformer model. For the upload of materials containing business entities in the early stage of power grid engineering projects, the algorithm automatically processes the documents through keyword recognition and fixed segmentation rules according to the set rules to achieve entity naming recognition and determine which are entity names such as project names, cost names, and specification regulations. The purpose of automatically processing the documents here is to build a standard vectorized knowledge base for RAG retrieval, form a standardized and structured knowledge base, and at the same time extract data corresponding to entity categories from it to form the data layer of the knowledge graph as the data source of the knowledge graph.
[0056] 2.2.1 Pre - work Business Control Knowledge Base
[0057] The pre - work business control knowledge base is a collection of knowledge data obtained by combining the "Implementation Guide for the Law - based Construction of Transmission and Substation Projects of State Grid Corporation of China (2018 Edition)" with the management experience in the pre - work stage. It is quickly extracted according to relevant documents. The structured form includes pre - work business control nodes, attributes related to business control nodes (business stage, responsible department, approval unit, project category, handling stage, etc.), national laws and regulations related to business control nodes, relevant State Grid management documents, relevant local management documents, relevant pre - work fees, etc. Among them, business entities can be extracted from the pre - work business control knowledge base: business control nodes, responsible departments, approval departments. At the same time, with business control nodes as the core, the association relationships with other business entities are shown.
[0058] Regarding the extraction of business entity data from a large number of documents in the present invention, for example, if the business control node is defined as grassland protection, grassland protection can be used as a keyword, or keywords related to grassland protection can be used to retrieve in the database, and all files related to grassland protection can be extracted.
[0059] 2.2.2 System Document Knowledge Base
[0060] The system document knowledge base is a collection of knowledge data of national laws and regulations, State Grid management documents, and local management documents related to the pre - work stage of power grid engineering projects. It is quickly extracted according to document materials. The structured form includes file name, document number, file type, issuing unit, issuing date, effective date, file status, etc. Among them, business entities can be extracted from the system document knowledge base: national laws and regulations, State Grid management documents, local management documents.
[0061] 2.2.3 System Rule Knowledge Base
[0062] The system rule knowledge base is a refinement of the system document knowledge base in the pre - work stage of power grid engineering projects, that is, a collection of rule entry data of national laws and regulations, State Grid management documents, and local management documents related to business control nodes. It is quickly extracted according to document materials. The structured form includes rule name, rule type, keyword, source file, etc.
[0063] 2.2.4 Pre - work Expense Knowledge Base
[0064] The pre - work expense knowledge base is a collection of knowledge data of the management regulations and business processes of pre - work expenses for power grid engineering projects. It is quickly extracted according to document materials. The structured form includes pre - work expense name, expense calculation principle, relevant system documents, etc. Among them, the business entity that can be extracted from the pre - work expense knowledge base is: pre - work expense.
[0065] 2.3 Knowledge Graph Construction
[0066] To provide an intuitive and highly interactive display method, Echarts is selected as the core graphics library. In terms of interaction, it can not only meet the rendering requirements of large-scale graphs but also provide rich interactive functions such as node click to expand / collapse, drag and zoom, and path highlighting, enabling users to intuitively explore and understand the business control nodes and their associations in the early stage of complex power grid engineering projects. Figure 2 In it, there can be multiple specific entities under each knowledge entity. For example, there can be multiple specific entities under the knowledge entity such as business control nodes: grassland protection, cultural relics protection, site selection planning, etc. There can also be multiple specific departments under the knowledge entity such as approval departments. Multiple specific entities under the same knowledge entity can be selected through a dropdown list or by node click to expand / collapse, etc. Each graph in the core graphics library represents a knowledge entity, encodes each knowledge entity, and defines the boundary of each knowledge entity. The said boundary is also a code. Through the boundary code, each entity is associated, and the related knowledge entities are connected in series.
[0067] Considering that the knowledge graph in the early stage of power grid engineering projects may be very large, an incremental loading strategy is adopted, that is, first load the core nodes and the edges directly associated with the core nodes, or other knowledge entities directly associated with the core nodes, and gradually load more levels of data as the user operates. In addition, for the content in non-critical areas, such as secondary nodes or long-distance connections, a lazy loading mechanism is implemented, and the loading is only carried out when the user clearly requests to view, further optimizing the performance. Provide a global search box to enable users to quickly locate specific entities; at the same time, add various filters (such as by category, level, etc.) to facilitate users to filter out the content they are interested in. These tools help users efficiently find the information they need and improve the user experience. Following modern Web design principles, adopt a simple and intuitive design style, reasonably arrange the positions of components such as buttons, labels, and prompts to ensure a smooth operation process. Especially for power grid engineering projects involving multiple professional fields, a clear interface design helps reduce the learning cost of users and improve work efficiency. Considering the access requirements on different devices and browsers, adopt a responsive layout to ensure that the page can be well displayed on various screen sizes and enhance the user experience.
[0068] To increase the intelligence level of the system, an intelligent Q&A system based on GPT technology is docked and integrated with the knowledge graph. Specifically, under the traditional method of retrieving in the file knowledge base by RAG, the retrieval results in the knowledge graph are added, and all the results are handed over to the large model. Users can directly enter questions on the interface, and the model will generate coherent and accurate answers based on the information in the knowledge graph. The specific procedure is: user's question - vectorization - RAG knowledge base retrieval - knowledge graph retrieval - submission of retrieval results and user input to the large model together - generation of answers. This process combines the advantages of knowledge graph technology and GPT model, eliminates the randomness and hallucination problems of traditional large language model answers, and improves the reliability of intelligent Q&A. Through NLP technology, the natural language input of users is parsed to identify key concepts and intentions therein, so as to more accurately match corresponding knowledge entities and relationships. For example, if a user asks about the risk warning situation of a certain project, the system can provide targeted information based on business control nodes and relevant regulatory documents. Such intelligent interaction not only improves user satisfaction but also promotes the effective transmission of knowledge. The WebSocket real-time communication protocol is established to ensure that the data in the knowledge graph can be synchronized and updated with the backend. Whenever new laws, regulations, and policy documents are uploaded, the system will immediately notify users of the update and prompt to refresh the interface, and update relevant content in the graph. Regarding the newly uploaded files, there may be existing entities in the newly uploaded files. If there is a relationship, it will be directly added; if not, a new relationship will be created.
[0069] The knowledge graph platform for the preliminary stage of power grid engineering projects constructed by the present invention solves the problems existing in the prior art, such as complex knowledge graph display and unreliable GPT model answers, and lays a solid foundation for future expansion and development.
[0070] Regarding the construction of the knowledge graph of the present invention:
[0071] First, knowledge entities are determined. As Figure 2 shown, the knowledge entities include business control nodes, responsible departments, approval departments, national laws and regulations, State Grid management documents, local management documents, preliminary expenses, etc. Data mining technology is used to identify the text to obtain the association relationships between entities, and the schema layer of the knowledge graph for the preliminary stage of power grid engineering projects is constructed. Regarding the relationships between entities, for example, those related to business control nodes include: responsible departments, approval departments, national laws and regulations, State Grid management documents, local management documents, preliminary expenses, work content, business stages, and work results. The business control nodes are associated with responsible departments, approval departments, national laws and regulations, internal unit documents, local management documents, preliminary expenses, work content, business stages, and work results to establish the schema layer related to business control nodes.
[0072] Secondly, classify the documents according to the document categories, such as national laws and regulations documents, State Grid management documents, local management documents, and preliminary expense documents. Since the naming of entities and relationships in the documents related to the preliminary work of the power grid is highly professional. For example, the document is named as relevant State Grid document or relevant preliminary expense, and the document naming indicates the content involved in the document. This kind of naming is highly professional. An extraction scheme combining manual annotation, rule- and deep learning-based extraction, and large language model integration can be used to determine the document type.
[0073] Select Echarts as the core graphics library. In this graphics library, each graphic represents a knowledge entity. For example, when there are seven knowledge entities: business control node, responsible department, approval department, national laws and regulations, State Grid management document, local management document, and preliminary expense, each graphic represents an entity. Documents of the same category are all stored under one graphic. For example, all local management documents are stored under the graphic representing local management documents.
[0074] Next, establish the relationships between each specific knowledge entity. For example, a business control node is a knowledge entity, and the business control node knowledge entity includes multiple specific knowledge entities such as grassland protection, cultural relics protection, and site selection planning. National laws and regulations are a knowledge entity, and the specific knowledge entities included in the national laws and regulations knowledge entity are: national laws and regulations related to grassland protection, national laws and regulations related to cultural relics protection, national laws and regulations related to site selection planning, etc. Through keyword search, establish the relationships between specific knowledge entities. For example, use "grassland protection" as a keyword for search, and associate the grassland protection business control node with the national laws and regulations related to grassland protection, approval department, responsible department, and preliminary expense related to grassland protection.
[0075] 3 Construction of a Large Language Model for the Preliminary Stage of Power Grid Engineering Projects Based on GPT Technology
[0076] Collect data through the construction of the knowledge graph for the preliminary stage of power grid engineering projects in step 2, and use the data collected in step 2 after cleaning through the construction of a large language model for the preliminary stage of power grid engineering projects based on GPT technology in step 3.
[0077] 3.1 Construction of a Large Language Model
[0078] 3.1.1 Data Preparation and Preprocessing
[0079] The data sources include a large number of national regulatory documents, management specifications for the preliminary stage of power transmission and transformation projects, State Grid management documents, etc., including both structured and unstructured documents. Therefore, the collected data is cleaned to remove redundant information, correct incorrect formats, and unify the data from different sources into a standard format, which refers to the html format and markdown format that are convenient for the large model to read. In particular, for some documents containing visual information such as charts and flowcharts, OCR (Optical Character Recognition) technology and image analysis algorithms are used to extract the information therein and associate it with the text information to form a multi-modal knowledge representation. For example, for the processing of picture tables, first record their positions in the text, then separate them, convert them into html format and write them back to the original positions, retaining their relationships with the context. Through these steps, the quality of the data input into the model is ensured, thereby improving the training effect of the model.
[0080] 3.1.2 Model Selection and Training
[0081] In the present invention, model selection and training include model selection, incremental pre-training, and instruction fine-tuning. Both incremental pre-training and instruction fine-tuning belong to the content of model training.
[0082] According to the current situation of knowledge management business in the preliminary stage of power grid engineering projects, a large language model for knowledge Q&A is constructed using GPT technology.
[0083] The so-called pre-training means that the data sources during the training process include a large number of national regulatory documents, management specifications for the preliminary stage of power transmission and transformation projects, State Grid management documents, etc. Since the relationships between business control nodes and other knowledge entities have been determined in the knowledge graph, that is, the relationships between various knowledge entities have been clarified in the knowledge graph, during the process of intelligent Q&A through the large language model, the large language model is pre-trained to make the Q&A results provided by the large language model consistent with the association relationships between knowledge entities in the knowledge graph. For example, when the business control node in the knowledge graph is defined as grassland protection, other knowledge entities related to grassland protection have been determined in the knowledge graph. When asking questions in the large language model with "grassland protection" as the keyword, through multiple model trainings, the reply results of the large language model are made consistent with the specific content recorded in the knowledge entities related to grassland protection provided in the knowledge graph.
[0084] Select Qwen2.5-7B-Instruct of Alibaba as the base model. Due to its large number of parameters and powerful generation ability, it can perform excellently in a variety of natural language processing tasks. In addition, the pre-training dataset of Qwen2.5 covers a wide range of text types, providing a good starting point for subsequent domain-specific fine-tuning. Aiming at the current situation of knowledge management business in the early stage of power grid engineering projects, incremental pre-training is carried out on the general Qwen2.5 model. In this invention, the incremental pre-training of the large language model means that the association relationship between the business control nodes determined in the knowledge graph and other knowledge entities is directly used as the reply result, that is, as the original text input into the large language model. Incremental pre-training + instruction fine-tuning + prompt engineering is the construction method of LLM applications. Different from the training set, test set, and validation set of traditional machine learning model training, it adopts "pre-training - prompt prediction", and guides the model to directly adapt to and execute specific downstream tasks through the designed prompt (Prompt). Specifically, for files with low change possibility, strong content professionalism, high accuracy requirements, and many general scenarios, such as national laws and regulations, State Grid management documents, etc., after being segmented according to the format or specific characters, they are used as a dataset for incremental training of the general large language model. This step aims to enable the model to better understand and generate professional terms and technical details related to power grid engineering projects.
[0085] According to the characteristics of the relevant documents in the early stage of power grid work, for documents with low change possibility, strong content professionalism, high accuracy requirements, and many general scenarios, such as laws and regulations, after being segmented according to the format or specific characters, they are used as a dataset for incremental pre-training, that is, secondary training, of the general large language model; the Fine-tuning method is adopted, and only some layers or specific parameters are adjusted while keeping most of the pre-training parameters unchanged. In this way, without destroying the generalization ability of the original model, it can be made more adaptable to specific tasks in the power grid engineering field.
[0086] 3.2 Algorithm Optimization
[0087] Adopt a comprehensive construction optimization method of retrieval-augmented generation + prompt engineering based on the large language model.
[0088] For documents such as management specifications and system documents that have a high update frequency and different content in different projects, after data cleaning, a knowledge base is classified and constructed, and the retrieval-augmented generation (RAG) technology is used for knowledge retrieval. For entities with a low frequency of occurrence and complex relationships in relevant documents, the traditional RAG technology that retrieves based on document fragments has significant limitations. Therefore, RAG is combined with a knowledge graph, and all the information retrieved from the knowledge graph is formed into an array, and the two-norm calculation of the array vector is performed. All content exceeding the set threshold is returned to the RAG knowledge base for retrieval to obtain the final result. In this article, the threshold is the distance from the user query vector. Exceeding the threshold means the distance is too large, and it is determined that this piece of information has little relationship with the user query and is not used as the basis for generating an answer. Regarding the use of the two-norm calculation in this invention, the norm is the Euclidean distance, which can be used as a similarity measurement method. The traditional two-norm calculation method is the square root of the sum of the squares of each element of the array. In the implementation of this system, in order to facilitate calculation, the square root is not taken, and the sum of the squares of each element of the array is directly calculated.
[0089] This processing method better understands the structured and semantic information in the question and answer, provides richer background information for low-frequency content, and improves the reasoning ability of complex relationships; to improve the accuracy of the model's understanding of user input, a large language model dynamic context awareness mechanism is introduced. When a user asks a question, the model not only considers the current question text but also combines the previous conversation history and context information to generate a more coherent and accurate answer. This mechanism is particularly applicable to continuous conversation scenarios; for common actual usage requirements of users, such as providing risk warnings and other scenarios, an instruction fine-tuning dataset is constructed, and the model is trained with instruction fine-tuning so that it can provide high-quality answers according to user needs in specific task scenarios. The so-called instruction fine-tuning refers to constructing a question-and-answer pair scenario according to the actual usage of users to guide the model on how to answer specific questions. It is mainly constructed manually, considering the common question scenarios of users; the prompt engineering for the questions or instructions input by users is optimized, and more professional terms or a structured way of expression are used to improve the accuracy of the model's understanding.
[0090] 4 Knowledge Intelligent Q&A for the Pre-construction Phase of Power Grid Engineering Projects Based on Knowledge Graph and GPT Technology
[0091] 4.1 Knowledge Pretraining for the Pre-construction Phase of Power Grid Engineering Projects
[0092] In the model pre-training stage, the previously constructed business control knowledge base, institutional document knowledge base, institutional rule knowledge base, preliminary cost knowledge base, and knowledge graph network of the preliminary stage are all used as inputs for the GPT model for pre-training. On the one hand, after segmenting the business control nodes, national regulatory documents, management specifications for the preliminary stage of power transmission and transformation projects, State Grid management documents, etc. in the knowledge base according to formats or specific characters, they are used as datasets to input into the model, enabling the model to better understand the professional terms related to knowledge management in the preliminary stage of power grid engineering projects. On the other hand, based on the constructed knowledge graph network of the preliminary stage, the large language model is trained to accurately identify key knowledge entities and the relationships between them, enabling the model to generate coherent and accurate answers according to the knowledge graph. As the pre-trained knowledge base and knowledge graph are updated to be more and more large and rich, the accuracy and reliability of the model are continuously optimized and improved during the intelligent question-answering process.
[0093] 4.2 Knowledge Intelligent Question Answering for the Preliminary Stage of Power Grid Engineering Projects Based on Knowledge Graph
[0094] After text pre-training, using the large language model, the management personnel in the preliminary stage of power grid engineering projects can input corresponding questions or select common questions according to the actual situation on-site, and the model can quickly achieve intelligent question answering according to the storage in the knowledge graph knowledge base to assist in scientific decision-making. At the same time, users can trace the source of the results in the knowledge base or knowledge graph. Because the large language model may have hallucinations, the so-called tracing the source means that when returning the answer, relevant documents and which specific paragraph or part in the document or which piece of data in the knowledge graph the answer comes from are also returned, and users can conduct manual verification to further confirm its accuracy.
[0095] Finally, it should be noted that the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art. The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A knowledge intelligent question-answering method for the pre-project stage of power grid engineering projects based on knowledge graph and GPT technology, characterized in that, The method includes: Analyze the current situation of knowledge management business in the early stage of power grid engineering projects, determine knowledge entities, construct the schema layer of the knowledge graph based on business characteristics, and form a logical relationship diagram. The logical relationship diagram includes knowledge entities as business control nodes and other knowledge entities related to business control nodes; Upload the management document materials in the early stage of power grid engineering projects, extract and organize the document materials using a knowledge extraction algorithm model, store the same type of documents under the same knowledge entity, and obtain the knowledge entity of the preliminary work and the structured knowledge base of the preliminary work; Combine the business control experience in the early stage to obtain entity-relationship-entity triples; establish a related relationship table of knowledge entities, and construct a knowledge graph for the early stage of power grid engineering projects; Build a large language model for the early stage of power grid engineering projects based on GPT technology, use the knowledge graph as the text input of the GPT model for model training, and optimize the model during the intelligent question-answering process.
2. The knowledge intelligent question answering method for the pre - stage of power grid engineering projects based on knowledge graph and GPT technology according to claim 1, wherein The construction of the schema layer of the knowledge graph includes determining the knowledge entities in the early stage of power grid engineering projects. The knowledge entities include business control nodes, the responsible departments, approval departments, national laws and regulations, internal management documents of the unit, local management documents, and preliminary fees related to the business control nodes. Use data mining technology to identify the text categories of knowledge entities, obtain the association relationships between entities, and construct the schema layer of the knowledge graph for the early stage of power grid engineering projects.
3. The knowledge intelligent question - answering method for the pre - stage of power grid engineering projects based on knowledge graph and GPT technology according to claim 1, characterized in that The extraction and organization of the document materials include using the BIOES annotation method to determine the document type; or using a scheme that combines a Chinese word segmentation model and a Transformer model to train based on manually annotated data, make predictions in unannotated document data, and then input the results into a large language model LLM to obtain the entity relationship output in standard format based on the set prompt including entity types and relationship construction formats.
4. A method for intelligent knowledge-based Q&A in the preliminary stage of a power grid engineering project based on a knowledge graph and GPT technology, characterized in that The obtaining of entity-relationship-entity triples includes connecting related knowledge entities through keywords.
5. A method for intelligent knowledge Q&A in the early stage of power grid engineering projects based on knowledge graph and GPT technology, characterized in that, The construction of the model includes pre-training and instruction fine-tuning. The pre-training includes: Unify the collected data into a standard format; Train a Qwen2.5-7B-Instruct large language model with the collected data, use the constructed knowledge base of business control in the preliminary work, system document knowledge base, system rule knowledge base, preliminary cost knowledge base, and knowledge graph network in the early stage as the input of the GPT model for pre-training, so that the model can better understand the professional terms related to knowledge management in the early stage of power grid engineering projects and can generate coherent and accurate answers based on the knowledge graph. Instruction fine-tuning is to adjust the parameters of some layers of the model using the Fine-tuning method.
6. The knowledge intelligent question-answering method for the pre-project stage of a power grid project based on a knowledge graph and GPT technology according to claim 1, characterized in that The model training includes optimizing the large language model by adopting an enhanced generation combined with prompt engineering; For files with different content in different projects, after data cleaning, classify and construct a knowledge base, use the Retrieval-Augmented Generation (RAG) technology for knowledge retrieval, combine RAG with a knowledge graph, form an array with all the information retrieved from the knowledge graph, calculate the two-norm of the array vector, return all the content that exceeds the set threshold to the RAG knowledge base for retrieval to obtain the final result. The threshold is the distance from the user query vector. If it exceeds the threshold, the distance is too large, and it is determined that the retrieved information does not match the user query relationship and is not used as the basis for generating an answer.