Electric power safety knowledge question-answering method and device and storage medium

By converting power safety knowledge data into text vectors and combining large language models and RAG technology, the problem of inaccurate answers in traditional question-and-answer systems is solved, achieving higher accuracy and scalability.

CN120104779APending Publication Date: 2025-06-06SHENZHEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510026200.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The traditional power safety knowledge Q&A system has the problem of inaccurate Q&A results, which are mainly reflected in the limitations of keyword matching and lack of predictability and correlation, resulting in poor generalization and service complexity.

Method used

A power safety knowledge question and answer method is adopted, and power safety knowledge data is collected and processed, and it is converted into a text vector and stored in a database. It uses large language models and RAG technology to perform intelligent question and answers to generate accurate answers.

Benefits of technology

It achieves more accurate and context-understood answers in complex power safety scenarios, improves the accuracy and scalability of the Q&A system, and adapts to growing data and knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104779A_ABST
    Figure CN120104779A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power security knowledge question-answering method and device and a storage medium, and the method comprises the steps: collecting and processing electric power security knowledge data, converting the electric power security knowledge data into a first text vector, and storing the first text vector in an electric power security knowledge database; obtaining an electric power safety knowledge problem input by a user, converting the electric power safety knowledge problem into a second text vector, and searching background knowledge corresponding to the electric power safety knowledge problem in the electric power safety knowledge database through vector similarity calculation; the background knowledge is embedded into a cue word template to generate cue words, then the cue words are input into the large model, and electric power safety knowledge answers are output. The electric power security knowledge question-answering method has higher accuracy and context understanding ability. The method can be widely applied to the technical field of power production safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power safety technology, and in particular to a power safety knowledge question and answer method, device and storage medium. Background Art

[0002] The power industry is a basic industry. Whether the power production is safe or not is not only related to the benefits and development of the enterprise itself, but also affects the stability of the entire society, economic development and the normal life of the people. The safety knowledge of the power industry is a basic skill that the power industry staff must master. Power safety knowledge has many characteristics - professionalism, complexity, standardization, and dynamism. It is a typical knowledge-intensive field and involves knowledge of electronics, physics, chemistry, mathematics and other disciplines. Therefore, it is necessary to establish a power safety knowledge question and answer system.

[0003] With the continuous advancement of technology, the knowledge question-answering system continues to evolve. The knowledge question-answering system has gradually moved from the initial basic model to a more intelligent and efficient stage. Its development process can be roughly divided into three stages: the basic keyword stage, the rule-driven stage, and the large language model technology-driven stage. The knowledge question-answering in the basic keyword stage is mainly based on database queries and rule bases, used to answer common questions, and provide basic information query functions. The knowledge question-answering in the rule-driven stage gradually introduced rules and templates to handle more types of questions, provide flexible answers, and expand the question-answering capabilities. In the technology-driven stage, with the help of the large language model, the introduction of its powerful contextual semantic understanding and text generation capabilities combined with semantic vector retrieval technology, the knowledge question-answering application has been significantly improved, achieving more accurate and faster answers.

[0004] Traditional power safety knowledge question-and-answer systems usually have technical problems such as inaccurate question-and-answer results, which are mainly reflected in two aspects: First, due to the limitation of keywords, traditional intelligent question-and-answer mainly relies on text content matching and requires complete keyword information. It has poor expression ability for unseen input languages ​​and poor question-and-answer generalization; second, it lacks predictability and relevance, which requires users to interact multiple times to deal with related issues, increasing time costs and service complexity. Summary of the invention

[0005] The present invention aims to solve one of the technical problems in the related art at least to a certain extent. To this end, one object of the present invention is to provide a method, device and storage medium for answering questions about power safety knowledge, which can intelligently answer questions about power safety knowledge, and the answers are highly accurate and can adapt to complex power safety scenarios.

[0006] The technical solution adopted by the present invention is: In a first aspect, the present invention provides a method for answering questions about power safety knowledge, which includes: collecting and processing power safety knowledge data, converting the power safety knowledge data into a first text vector and storing it in a power safety knowledge database; obtaining power safety knowledge questions input by users, converting the power safety knowledge questions into a second text vector, and retrieving background knowledge corresponding to the power safety knowledge questions in the power safety knowledge database through vector similarity calculation; embedding the background knowledge into a prompt word template to generate a prompt word, and then inputting the prompt word into a large model to output the power safety knowledge answer.

[0007] Among them, the collection and processing of power safety knowledge data, converting the power safety knowledge data into a first text vector and storing it in a power safety knowledge database, includes: collecting power safety knowledge data set files, the types of which include: power subject textbooks, national standard documents and instruction manuals, and related power examination question banks and answers; organizing the data set files into power safety knowledge data in the form of question-answer pairs; inputting the organized power safety knowledge data into an M3E text embedding model, performing text vectorization processing, and converting it into a first text vector; and using a Chroma vector database to store the first text vector.

[0008] Among them, the background knowledge is embedded into the prompt word template to generate a prompt word, and then the prompt word is input into the big model. Before outputting the answer to the power safety knowledge, it also includes: loading and fine-tuning the big model, which is the ChatGLM2-6b model, and fine-tuning the big model based on the P-Tuning v2 method.

[0009] Among them, this method is implemented based on the LangChain framework.

[0010] In a second aspect, the present invention provides an electric power safety knowledge question and answer device, which includes: an electric power safety knowledge database construction module, which is used to collect and process electric power safety knowledge data, and convert the electric power safety knowledge data into a first text vector and store it in an electric power safety knowledge database; a background knowledge retrieval module, which is used to obtain electric power safety knowledge questions input by a user, convert the electric power safety knowledge questions into a second text vector, and retrieve the background knowledge corresponding to the electric power safety knowledge questions in the electric power safety knowledge database through vector similarity calculation; an electric power safety knowledge answer generation module, which is used to embed the background knowledge into a prompt word template to generate a prompt word, and then input the prompt word into a large model to output an electric power safety knowledge answer.

[0011] Among them, the power safety knowledge database construction module includes: a data set file acquisition unit, which is used to collect power safety knowledge data set files, and the types of data set files include: electric power subject textbooks, national standard documents and instruction manuals, and related electric power examination question banks and answers; a data set file processing unit, which is used to organize the data set files into power safety knowledge data in the form of question and answer pairs; a text vector conversion unit, which is used to input the organized power safety knowledge data into the M3E text embedding model, perform text vectorization processing, and convert it into a first text vector; a text vector storage unit, which is used to use the Chroma vector database to store the first text vector.

[0012] It also includes: a large model loading and fine-tuning module, which is used to load and fine-tune the large model, the large model is the ChatGLM2-6b model, and the large model is fine-tuned based on the P-Tuning v2 method.

[0013] In a third aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method as described above.

[0014] The beneficial effects of the present invention are: The present invention uses the LangChain framework to load the ChatGLM2-6B base model and combines RAG technology to build an intelligent question-answering system in the power safety knowledge scenario. This method has the following advantages: (1) Higher accuracy and context understanding: The ChatGLM2-6B model can better understand contextual information, thereby providing more appropriate answers in complex power safety scenarios. Combined with RAG technology, it can provide more accurate information.

[0015] (2) Better scalability: Through the Chroma vector database and M3E text embedding technology, the knowledge base can be easily expanded to adapt to the ever-growing data and knowledge. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flow chart of an embodiment of the electric power safety knowledge question-answering method of the present invention; Figure 2 yes Figure 1 A schematic flow chart of an embodiment of step S11; Figure 3 It is a structural schematic diagram of an embodiment of the electric power safety knowledge question-answering device of the present invention; Figure 4 yes Figure 3 A structural diagram of an embodiment of an electric power safety knowledge database construction module 11. DETAILED DESCRIPTION

[0017] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application may be combined with each other.

[0018] The method of this application is implemented based on the LangChain framework. This application uses the LangChain framework to develop a large language model application, and uses the large language model with external data sources through the components and tools provided by LangChain to quickly build a power safety knowledge question-and-answer system. LangChain provides developers with a powerful toolkit to build complex applications based on the large language model (LLM) with a concise and modular architecture design. Embodiment 1

[0019] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of an embodiment of the electric power safety knowledge question-answering method of the present invention. Figure 1 As shown, the method comprises the steps of: S11: Collecting and processing power safety knowledge data, converting the power safety knowledge data into a first text vector and storing it in a power safety knowledge database; Specifically, see Figure 2 , Figure 2 FIG. 4 is a flow chart of an embodiment of step S11. Figure 2 As shown, step S11 includes the following sub-steps: S111: Collect power safety knowledge data set files; The types of data set files include: electrical engineering textbooks, national standard documents and instruction manuals, related electrical engineering test question banks and answers, etc.

[0020] The electrical engineering textbooks are textbook documents related to the electrical engineering discipline, such as "Electrical Engineering", "Basics of Electrical and Information Technology", "Registered Electrical Engineer (Professional Foundation)" and other textbooks.

[0021] National standard documents and instruction manuals include "Regulations on Supervision and Management of Power Safety Hazards", "Power Safety Knowledge and Cases", "Power Engineering Construction Safety Specifications", "Power System Operation Safety Regulations", "GB 50034-2013 Building Lighting Design Standards", "JGJ16-2008 Civil Building Electrical Design Specifications" and other legal and regulatory documents and standard specification documents. Convert these documents into PDF format with text extraction.

[0022] By using crawler technology to crawl questions from the registered electrical engineer professional qualification examination, we obtained relevant power examination question banks and answers such as "Registered Electrical Engineer Public Basic Examination Question Bank", "Registered Electrical Engineer (Power Supply and Distribution Major) Load Classification and Calculation Examination Questions", "Electrical Engineer Power Generation, Transmission and Transformation Professional Examination Papers and Answers", and "Senior Technician Review Question Bank for Power Meter Connection".

[0023] S112: organizing the data set file into power safety knowledge data in the form of question-answer pairs; After obtaining the power safety knowledge dataset file, in order to facilitate the fine-tuning of the large model, GPT-3.5 was used to integrate the contents of the subject textbooks, national standard documents and instruction manuals into the format of {"question":"...","answer":"..."} question-answer pairs. Specifically, this article uses the following prompt: prompt = "Please convert all the above information into the form of fill-in-the-blank question-answering questions. The question-answering format is 'question: (question content)\n answer: (answer content)\n\n'", and the obtained files are integrated into json files as training data.

[0024] For the test question bank and answers, we first perform data cleaning, delete and modify the crawled data with incorrect format, and finally manually proofread it and convert it into high-quality question and answer pairs based on the content of the knowledge corpus.

[0025] LangChain provides a document loader to process documents in various formats. By setting the mode of UnstructuredFileLoader to elements, each question-answer pair is split into independent text blocks, thus ensuring the integrity of the question-answer pair. When constructing the training set and test set, 70% of the annotated corpus and questions and answers are used for model fine-tuning, and 30% are used as a test set to test model performance.

[0026] S113: inputting the sorted power safety knowledge data into the M3E text embedding model, performing text vectorization processing, and converting it into a first text vector; Text vectorization (Embedding) is a technology that maps high-dimensional sparse data to low-dimensional dense space. This technology is widely used in natural language processing, recommendation systems, image processing and other fields to convert discrete and sparse input data (such as vocabulary, user ID, item ID) into dense and continuous vector representation. When using RAG (Retrieval-Augmented Generation) technology in combination with a large language model, there is a question about how to choose a suitable Chinese Embedding model. Chinese Embedding models are very critical in RAG technology because they directly affect the effect of information retrieval and the quality of generated text.

[0027] The selection criteria of different dimensions of text vectorization models mainly include the following aspects: (1) Model performance is one of the most important criteria. An embedding model with excellent performance can provide more accurate vector representation, thereby improving the accuracy of information retrieval and the quality of generated text; (2) Processing speed. The computational efficiency of the model is also very critical. A model with fast processing speed can significantly improve the response speed of the system in practical applications, thereby improving the user experience; (3) Vector dimension size. The dimension size of the embedding vector directly affects the storage and computing cost of the model. Higher dimensions can capture more detailed information, but it will also increase computing overhead, so it is necessary to find a balance between dimension size and performance.

[0028] The text vectorization model of this method is implemented through the M3E text embedding model. M3E is the abbreviation of Moka MassiveMixed Embedding. Moka means that this model is trained, open sourced and evaluated by MokaAI. Massive means that this model is trained through tens of millions (22 million+) Chinese sentence pair data sets. Mixed means that the model supports homogeneous text similarity calculation in Chinese and English, heterogeneous text retrieval and other functions. Embedding means that the model is a text embedding model that can convert natural language into dense vectors. Therefore, the M3E model is a powerful open source Embedding model that has been trained with a large number of sentence pair data sets. In the LangChain framework, you can load the pre-trained model file through the interface and use this embedding model for text vectorization.

[0029] S114: Using the Chroma vector database to store the first text vector.

[0030] A vector database is a database that stores data as high-dimensional vectors, which are mathematical representations of features or attributes. Each vector has a certain number of dimensions, which can range from tens to thousands, depending on the complexity and granularity of the data. Vectors are usually generated by applying some kind of transformation or embedding function to the raw data (such as text, images, audio, video, etc.), and the embedding function can be based on various methods such as machine learning models, word embedding, and feature extraction algorithms. The main advantage of a vector database is that it allows fast and accurate similarity search and retrieval based on the vector distance or similarity of the data, which means that instead of using the traditional method of querying the database based on exact matches or predefined criteria, a vector database can be used to find the most similar or relevant data based on semantic or contextual meaning.

[0031] In order to realize the storage of vector data and the query of semantic vectors, this method uses Chroma vector data to save the vector data of text. LangChain provides a good encapsulation for the Chroma vector database, and developers can use the interface to quickly implement the storage and retrieval of vector data. The vector database Chroma can quickly process and retrieve a large amount of vector data through its efficient data structure and algorithm optimization. Some main features of the Chroma vector database: (1) Efficient vector index: Chroma uses efficient index structures such as inverted index, KD-tree or graph-based index to speed up vector search. (2) Support for multiple similarity metrics: It supports multiple vector similarity metrics, including Euclidean distance, cosine similarity, etc., so that it can be widely used in different application scenarios. (3) Scalability and elasticity: Chroma can support horizontal expansion and adapt to the needs of large-scale data sets. At the same time, it can also effectively handle dynamic changes in data and adapt to rapidly developing storage needs. (4) Easy to integrate and use: Chroma is designed with an easy-to-use API interface and supports access to multiple programming languages, which is convenient for developers to integrate and use in different systems and applications.

[0032] S12: obtaining a power safety knowledge question input by a user, converting the power safety knowledge question into a second text vector, and retrieving background knowledge corresponding to the power safety knowledge question in the power safety knowledge database through vector similarity calculation; In step S12, the power safety knowledge question is processed in the same way as step S113 and converted into a second text vector.

[0033] Vector similarity calculation is a method of vector retrieval. The method of vector retrieval is to process the user query content, extract keywords and convert them into vector form, use knowledge vector index, and use methods such as approximate nearest neighbor retrieval to find the most similar text block vector in the knowledge vector library to obtain the knowledge fragment most relevant to the user query content. Specifically, vector retrieval mainly consists of the following two steps. The first is question vectorization. The system uses the same pre-trained language model as the text data to encode the question and vectorize the question. In this way, the system captures the semantic information of the question and compares the similarity. The second is vector similarity retrieval. Use vector similarity retrieval technology to find the knowledge vector that is most similar to the question vector in the knowledge base, which usually involves calculating the cosine similarity or other similarity measurement methods between vectors. In this way, the knowledge most relevant to the question can be found quickly.

[0034] S13: embed the background knowledge into a prompt word template to generate a prompt word, and then input the prompt word into the large model to output the power safety knowledge answer.

[0035] Preferably, before step S13, the method further includes the steps of: loading the ChatGLM2-6b large language model, and fine-tuning the large language model based on the P-Tuning v2 method.

[0036] ChatGLM2-6b is an advanced reasoning framework. Its efficient dynamic reasoning and memory optimization technology enable it to perform well in tests of multiple Chinese and English public data sets. To load the ChatGLM2-6b large language model in the LangChain framework, you can use the LLM wrapper to construct a ChatGLM model class. To construct your own LLM class in LangChain, you need to use the input and output modules of the LangChain framework to implement a custom LLM class, and complete the following in the custom LLM class: First, load the pre-trained model file of the user-defined LLM model in the initialization method, where the model_path parameter is used to specify the pre-trained file path of the LLM; second, implement the _call method so that it accepts a Prompt string and returns a response string.

[0037] In order to adapt the ChatGLM2-6b model to a specific task or field, in this article’s power field, especially the large language model training combined with RAG technology, the method of combining PEFT (Parameter-Efficient Fine-Tuning) and RAG is usually considered to improve the model performance, that is, on the basis of the original model, additional training is performed using a new, relatively small data set, and the parameters are adjusted to meet the requirements of the power safety field. PEFT is a technique that adapts to new tasks by adjusting only a small number of additional parameters while keeping most of the parameters of the pre-trained model unchanged. These additional parameters can be newly added embedding layers, low-rank matrices, or other types of parameters, which are used to "guide" or "adjust" the output of the pre-trained model to make it more suitable for new tasks.

[0038] The main methods of PEFT include LoRA, Adapter Tuning, Freeze, and P-Tuning series. LoRA (Low-Rank Adaptation) approximates the update of model parameters by adding a low-rank matrix near the original model weight matrix. This method achieves fine-tuning by optimizing this low-rank matrix without modifying the original model parameters. AdapterTuning achieves fine-tuning by inserting small neural networks (called adapters) between each layer of the model. These adapters contain trainable weights, while the original parameters of the model remain unchanged. The freezing method Freeze is to train with only a small number of parameters and freeze most of the parameters of the model. Generally, Freeze only fine-tunes the parameters of the fully connected layers of the last few layers and freezes all other parameters. Fine-tuning only the parameters of the last few layers can retain most of the knowledge of the pre-trained model, while adapting to the specific requirements of specific tasks through fine-tuning. The main problem that Freeze solves is to ensure that the knowledge learned by the model can be fully utilized and to reduce the computing resources and time consumption during the fine-tuning process.

[0039] The P-Tuning (parameter fine-tuning) series is a method for specific tasks that marks and optimizes specific parameters of pre-trained models. There are two schemes: v1 and v2. P-Tuning v1 is a method that replaces discrete labels with differentiable virtual labels. This method only adds virtual labels to the input layer and uses a prompt encoder (BiLSTM+MLP) to encode and learn the virtual labels. However, there are two main problems with the P-Tuning v1 method: lack of model parameter scale and task versatility, and its effectiveness for hard sequence labeling tasks (such as sequence tagging) has not been verified; lack of scale versatility. When the model parameter scale exceeds 10 billion or a smaller model (100M to 1B), there is a big difference between the P-Tuning v1 method and full fine-tuning, which limits the applicability of the P-Tuning v1 method.

[0040] Based on the problems of the P-Tuning v1 method, a subsequent P-Tuning v2 version is proposed. The specific improvements are as follows: a multi-task learning method is used to pre-train on prompts based on a multi-task dataset and adapt to downstream tasks, which can better initialize the pseudo-labels of continuous prompts and provide better initialization effects; the re-parameterized encoder is removed. For smaller models, the improvement of using a re-parameterized encoder is very small and may even affect the performance of the model; prompt word tags are added as fine-tunable parameters at each layer, so that the model can participate in the training and learning process at each layer to adapt to the requirements of specific tasks. In addition, the model parameters are frozen during training, and only the prompt words of each layer are trained; only the PrefixEncoder needs to be saved and loaded to reduce the video memory required for training, which will better understand the prompts of the question and generate accurate answers. This paper uses the P-Tuning v2 method to fine-tune the model.

[0041] For the question-answering task in this step, good prompt words and reasonable question-answering strategies are also the key to solving the problem well. In order to obtain good answering results, the following methods are used in prompt words and question-answering strategies: (1) Further call out the knowledge in the model through role-playing. The large model originally has a lot of basic knowledge and common sense related to power safety during pre-training. However, due to the small number of model parameters and the need to maintain the generalization of the model, the training of the ChatGLM2-6B model uses a lot of knowledge compression, parameter compression and other methods, which makes it difficult to call out relevant knowledge in a certain field. Therefore, it is possible to consider activating the knowledge in related professional fields by letting the large model perform role-playing. Therefore, we first give the model an identity in the prompt words "Now you are an expert in the field related to power safety" to improve the accuracy of question and answer.

[0042] (2) Adjustment of prompt words. The ChatGLM2-6B model is very sensitive to prompt words. Changing some seemingly insignificant words may have a huge impact on the model's performance. Therefore, this method carefully designs and repeatedly adjusts the prompt words. By converting and training multiple expressions of a single sentence, the model can generate the same answer for the same prompt word input but different expressions, thereby improving the model's stability for the prompt words.

[0043] (3) Adjustment of question-answering strategy. Fine-tuning the ChatGLM2-6B model requires a high-quality data set. In the power safety knowledge question-answering task, fine-tuning is generally based on single-discussion dialogues, which will inevitably cause the large model to lose a certain degree of contextual understanding ability. Therefore, the strategy of this method is: ask the fine-tuned ChatGLM model questions, and further use the answers of this model and the results obtained by RAG retrieval as the input of an original model to obtain the results. In this way, the fine-tuning knowledge of the fine-tuned model and the relevant knowledge of the knowledge base are utilized at the same time, and the reasoning ability of the model context is fully utilized to ensure the final good effect. In addition, due to the limitation of the number of parameters of the ChatGLM2-6B model itself, it is easy to have phantom reading, repeated reading, and wrong reading. Therefore, different strategies are adopted for different types of questions and answers: for example, for questions with only one answer, the question and answer are used three times, and the option with the most appearances is finally selected as the final result; for questions with multiple results as answers, the method of taking the intersection of the two questions and answers is used as the final result; and for free question-answering questions, the question and answer are used three times, and the results of the three times are spliced ​​together as the final result. Embodiment 2

[0044] See also Figure 3 , Figure 3 Schematic diagram of the structure of an embodiment of the electric power safety knowledge question-answering device of the present invention. Figure 3 As shown, the device includes: an electric power safety knowledge database construction module 11, a background knowledge retrieval module 12 and an electric power safety knowledge answer generation module 13.

[0045] The power safety knowledge database construction module 11 is used to collect and process power safety knowledge data, and convert the power safety knowledge data into a first text vector and store it in the power safety knowledge database.

[0046] Specifically, please refer to Figure 4 , Figure 4 FIG. 1 is a schematic diagram of a structure of an embodiment of a power safety knowledge database construction module 11. Figure 4As shown, the power safety knowledge database construction module 11 includes a data set file acquisition unit 111, a data set file processing unit 112, a text vector conversion unit 113 and a text vector storage unit 114. The data set file acquisition unit 111 is used to collect power safety knowledge data set files, and the types of data set files include: electric power subject textbooks, national standard documents and instruction manuals, and related electric power examination question banks and answers. The data set file processing unit 112 is used to organize the data set files into power safety knowledge data in the form of question and answer pairs. The text vector conversion unit 113 is used to input the organized power safety knowledge data into the M3E text embedding model, perform text vectorization processing, and convert it into a first text vector. The text vector storage unit 114 is used to use the Chroma vector database to store the first text vector.

[0047] The background knowledge retrieval module 12 is used to obtain the power safety knowledge question input by the user, convert the power safety knowledge question into a second text vector, and retrieve the background knowledge corresponding to the power safety knowledge question in the power safety knowledge database through vector similarity calculation.

[0048] The power safety knowledge answer generation module 13 is used to embed the background knowledge into the prompt word template to generate prompt words, and then input the prompt words into the large model to output the power safety knowledge answer.

[0049] Preferably, the device includes a model loading and fine-tuning module (not shown) for loading and fine-tuning the large model, which is a ChatGLM2-6b model, and fine-tuning the large model based on the P-Tuning v2 method.

[0050] Specifically, the working methods of each module in this embodiment have been described in detail in the first embodiment and will not be repeated here. Embodiment 3

[0051] The present invention further provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method described in the first embodiment.

[0052] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A method for answering questions about power safety knowledge, characterized in that: include: Collecting and processing power safety knowledge data, converting the power safety knowledge data into a first text vector and storing it in a power safety knowledge database; Acquire a power safety knowledge question input by a user, convert the power safety knowledge question into a second text vector, and retrieve background knowledge corresponding to the power safety knowledge question in the power safety knowledge database by vector similarity calculation; The background knowledge is embedded into a prompt word template to generate prompt words, and then the prompt words are input into a large model to output power safety knowledge answers.

2. The method according to claim 1, characterized in that The collecting and processing of the power safety knowledge data, and converting the power safety knowledge data into a first text vector and storing it in a power safety knowledge database, includes: Collecting electric power safety knowledge data set files, the types of which include: electric power subject textbooks, national standard documents and instruction manuals, and related electric power test question banks and answers; Arranging the data set file into power safety knowledge data in the form of question-answer pairs; The sorted power safety knowledge data is input into the M3E text embedding model, and the text is vectorized and converted into the first text vector; The Chroma vector database is used to store the first text vector.

3. The method according to claim 1, characterized in that The process of embedding the background knowledge into a prompt word template to generate a prompt word, and then inputting the prompt word into a large model, before outputting an answer to the power safety knowledge, further includes: The large model is loaded and fine-tuned. The large model is the ChatGLM2-6b model. The large model is fine-tuned based on the P-Tuning v2 method.

4. The method according to any one of claims 1 to 4, characterized in that: The method is implemented based on the LangChain framework.

5. An electric power safety knowledge question and answer device, characterized in that: include: An electric power safety knowledge database construction module is used to collect and process electric power safety knowledge data, and convert the electric power safety knowledge data into a first text vector and store it in the electric power safety knowledge database; A background knowledge retrieval module is used to obtain a power safety knowledge question input by a user, convert the power safety knowledge question into a second text vector, and retrieve the background knowledge corresponding to the power safety knowledge question in the power safety knowledge database through vector similarity calculation; The power safety knowledge answer generation module is used to embed the background knowledge into a prompt word template to generate prompt words, and then input the prompt words into the large model to output the power safety knowledge answer.

6. The device according to claim 5, characterized in that The power safety knowledge database construction module includes: A data set file collection unit is used to collect electric power safety knowledge data set files, the types of which include: electric power subject textbooks, national standard documents and instruction manuals, and related electric power test question banks and answers; A data set file processing unit, used for arranging the data set file into power safety knowledge data in the form of question-answer pairs; A text vector conversion unit, used to input the sorted power safety knowledge data into the M3E text embedding model, perform text vectorization processing, and convert it into a first text vector; The text vector storage unit is used to store the first text vector using a Chroma vector database.

7. The device according to claim 5 or 6, characterized in that Also includes: The large model loading and fine-tuning module is used to load and fine-tune the large model, where the large model is the ChatGLM2-6b model, and the large model is fine-tuned based on the P-Tuning v2 method.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Model training method and question answering method for question answering system

    CN118093841A

  • Knowledge management question-answering system construction method and device

    CN118446296A

  • Digital-intelligent power regulation and control question answering method and system based on large language model

    CN119003742A