Power grid equipment intelligent question and answer optimization method and system based on large language model
By adopting large language models and multiple natural language processing technologies in the grid equipment question and answer system, the problem of inefficient complex semantic understanding and knowledge update is solved, and more accurate and efficient question and answer output is achieved.
Patent Information
- Application Number
- CN202510403150.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing grid equipment Q&A systems have problems of inefficiency and insufficient accuracy in dealing with complex semantic understanding and rapid update of knowledge.
The intelligent Q&A optimization method of power grid equipment based on large language models is adopted, and the parsing ability and answer accuracy of the question and answer system are improved through technical means such as part of speech annotation, semantic role annotation, vector space matching, Transformer architecture, large language model training and logical expression optimization.
It improves the analysis and processing capabilities of power grid equipment problems, outputs more accurate and professional answers, enhances the pertinence and accuracy of information retrieval, and meets users' needs for quick access to information.
Smart Images

Figure CN120197708A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent question - answering tools, and particularly to an optimization method and system for intelligent question - answering of power grid equipment based on large - language models. Background Technique
[0002] Early power grid equipment question - answering systems were mostly constructed based on rule - bases. For questions outside the pre - set rules, especially those involving complex semantic understanding, they could not be accurately parsed. For example, when a user asks "How to handle the abnormal situation of the voltage regulation device in a substation that has intelligent monitoring function and has been newly put into use", it is very difficult for traditional systems to accurately identify the "substation" modified by the complex attributive "with intelligent monitoring function and newly put into use", and the relationship with "the handling of abnormal voltage regulation device", resulting in an inability to give an effective answer. In addition, question - answering systems based on keyword matching can only simply search for answers in the database according to the keywords in the question and cannot understand the deep semantics of the question. For example, for the question "How to ensure the stability of distributed energy access in an intelligent power grid", relying solely on keyword matching may not comprehensively cover various factors related to the stability of distributed energy access, such as equipment compatibility, control strategies, etc., thus giving a one - sided or inaccurate answer.
[0003] With the continuous development of power grid technology, new equipment, technologies, and operation specifications are emerging continuously. For traditional question - answering systems based on fixed rule - bases or static knowledge - bases, the process of updating knowledge is cumbersome and time - consuming. When searching for relevant information on the complex fault diagnosis of large - scale substations, due to the large amount of data such as equipment information and fault cases stored in the database, simple query algorithms may need to traverse a large amount of irrelevant data, resulting in a long retrieval response time and being unable to meet the user's need for quickly obtaining information. At the present stage, there is an optimization method and system for intelligent question - answering of power grid equipment based on large - language models. Summary of the Invention
[0004] In order to solve the problems of low retrieval efficiency and poor retrieval accuracy in intelligent question - answering in the power grid, the present invention provides an optimization method and system for intelligent question - answering of power grid equipment based on large - language models.
[0005] In the first aspect, an optimization method for intelligent question - answering of power grid equipment based on large - language models provided by the present invention adopts the following technical solutions: An optimization method for intelligent question - answering of power grid equipment based on large - language models includes: Obtain question data of power grid equipment and create a dynamically updated vocabulary of power grid equipment for storing question data; Use a part - of - speech tagging model of conditional random fields to tag the part - of - speech of the question data, and assign semantic roles to each word through semantic role annotation; Grid equipment knowledge retrieval is performed on problem data based on annotation and semantic roles, including matching documents in the grid equipment knowledge base and user questions using a vector space; A large language model is constructed with the Transformer architecture as the basic framework, including introducing the GPT architecture and the BERT architecture based on the Transformer decoder and encoder respectively; Parameter updates are performed on the constructed large language model, including updating model parameters using a variant of the stochastic gradient descent algorithm; Answer optimization is performed based on the processing results of the large language model to obtain the final intelligent Q&A output.
[0006] Furthermore, the part-of-speech tagging model using conditional random fields is used to tag the part-of-speech of the problem data, including defining feature functions based on the text sequence in the problem data, constructing a conditional probability model of the conditional random field according to the defined feature functions, using the grid equipment text with tagged part-of-speech as training data, training the conditional probability model by maximizing the log-likelihood function of the training data, and tagging the part-of-speech of the problem data according to the trained conditional probability model.
[0007] Furthermore, the matching of documents in the grid equipment knowledge base and user questions using a vector space includes calculating the term frequency and inverse document frequency using the documents and user question data respectively, calculating the basic weight of the question data in the document based on the term frequency and inverse document frequency, assigning different weight adjustment coefficients to different types of keywords according to the grid equipment knowledge hierarchy to obtain the final weight of the keywords in the document, calculating the similarity between the user question vector and the document vector in the grid equipment knowledge base using cosine similarity, setting a similarity threshold according to the calculated similarity value, and screening out documents with a similarity greater than the threshold.
[0008] Furthermore, the construction of a large language model with the Transformer architecture as the basic framework includes semantic encoding of the input text based on vector data to obtain the semantic representation of each word, introducing a position encoding module and calculating the position encoding in combination with semantic information. The calculation formula of the position encoding after combining with semantic information is: , where, represents the position of the keyword in the document, i represents the dimension index, represents the dimension of the model, represents the result obtained by semantic encoding, represents the function that converts the semantic representation into a numerical value.
[0009] Further, the Transformer decoder and encoder respectively introduce the GPT architecture and the BERT architecture, including building a GPT structure based on the Transformer decoder, using the autoregressive feature of the decoder to generate the output sequence from left to right in turn, and performing pre-training in a self-supervised learning manner. Then, a BERT architecture based on the Transformer encoder is constructed, and bidirectional language representation learning is carried out through the multi-head attention mechanism, and the BERT architecture is trained by randomly masking some words in the input text.
[0010] Further, the method of updating the model parameters using the variant algorithm of stochastic gradient descent includes defining a loss function through a question-and-answer data set in the field of power grid equipment, calculating the gradient of the current model parameters according to the loss function, calculating the first-order moment estimate and the second-order moment estimate in turn using the gradient of the current model parameters, and performing bias correction on the first-order moment estimate and the second-order moment estimate. Finally, the model parameters are updated according to the corrected first-order moment estimate and second-order moment estimate.
[0011] Further, the method of optimizing the answer according to the processing result of the large language model includes using first-order predicate logic to transform the operation process and fault diagnosis logic of power grid equipment into logical expressions, and matching the generated logical expressions with a pre-established logical rule library of power grid equipment operation specifications.
[0012] In a second aspect, an intelligent question-and-answer optimization system for power grid equipment based on a large language model includes: A data acquisition module, configured to: acquire question data of power grid equipment and create a dynamically updated vocabulary of power grid equipment for storing the question data; An analysis module, configured to: label the part-of-speech of the question data using a part-of-speech tagging model of conditional random fields and assign semantic roles to each word through semantic role annotation; A conversion module, configured to: perform knowledge retrieval of power grid equipment based on the question data with labels and semantic roles, including matching the documents in the power grid equipment knowledge base and the user's question using a vector space; A model module, configured to: build a large language model with the Transformer architecture as the basic framework, including respectively introducing the GPT architecture and the BERT architecture based on the Transformer decoder and encoder; A training module, configured to: update the parameters of the built large language model, including using a variant algorithm of stochastic gradient descent to update the model parameters; An optimization module, configured to: optimize the answer according to the processing result of the large language model to obtain the final intelligent question-and-answer output.
[0013] In a third aspect, the present invention provides a computer-readable storage medium storing multiple instructions, which are adapted to be loaded and executed by a processor of a terminal device to implement the method for optimizing intelligent question answering of power grid equipment based on a large language model.
[0014] In a fourth aspect, the present invention provides a terminal device including a processor and a computer-readable storage medium. The processor is configured to implement each instruction; the computer-readable storage medium is used to store multiple instructions, which are adapted to be loaded and executed by the processor to implement the method for optimizing intelligent question answering of power grid equipment based on a large language model.
[0015] In summary, the present invention has the following beneficial technical effects: 1. By performing part-of-speech tagging and semantic role tagging on the power grid equipment question data, the present invention can accurately analyze the structure and semantic information of the questions. For professional terms and complex operation process expressions in the power grid equipment field, it can more accurately identify their grammar and semantic roles. This accurate tagging helps the subsequent large language model to better understand the questions, thereby improving the parsing and processing capabilities of the questions and finally outputting more accurate answers that conform to the professional knowledge of power grid equipment.
[0016] 2. By using the word frequency and inverse document frequency of the documents and user questions, and combining with the knowledge hierarchy structure of power grid equipment to assign different weights to keywords, and then screening the documents through cosine similarity, the present invention can accurately find the documents most relevant to the user questions from a large number of power grid equipment knowledge bases. This method improves the pertinence and accuracy of information retrieval, avoids the interference of irrelevant information, and enables the large language model to generate answers based on more accurate and relevant information.
[0017] 3. By introducing a position encoding formula that combines semantic information, the present invention enables the large language model to better understand the relationship between word order and semantics in the text of the power grid equipment field. For texts describing operation steps and fault analysis of power grid equipment, the model can more deeply grasp the semantic logic therein, thereby improving the understanding ability of operation processes and equipment relationships and outputting answers that conform to professional logic.
[0018] 4. By utilizing the autoregressive feature and self-supervised learning pre-training of the GPT architecture, the present invention enables the model to learn language patterns and knowledge from a large number of power grid equipment texts, which is particularly suitable for generative question-and-answer tasks and can generate smooth and professional operation steps according to the learned language habits and knowledge.
[0019] 5. By utilizing the bidirectional language representation learning and masked language model training of the BERT architecture, the present invention can deeply understand the semantics of words in the context, improve the understanding and processing capabilities of complex semantics, and contribute to answering complex power grid equipment questions.
[0020] 6. Through calculating gradients, first - moment estimation, second - moment estimation, bias correction, and parameter update, the present invention makes the model training process more stable and the convergence speed faster. For large - scale power grid equipment domain data and complex large - language models, this algorithm can adaptively adjust the learning rate according to different parameter gradients, avoiding the problems of slow convergence or getting stuck in local optima that may occur in traditional gradient - descent algorithms, improving the efficiency and performance of model training, enabling the model to learn knowledge and language patterns in the power grid equipment domain faster, and shortening the development and optimization time.
[0021] 7. By transforming the operation process and fault diagnosis logic of power grid equipment into logical expressions and matching them with a pre - established logical rule base, the present invention can perform logical verification and optimization on the answers generated by the large - language model. It ensures that the generated answers are not only accurate in language expression but also logically conform to the operation specifications and principles of power grid equipment, avoiding illogical situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a schematic diagram of the overall process of an intelligent question - answering optimization method for power grid equipment based on a large - language model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The present invention will be further described in detail below with reference to the accompanying drawings.
[0024] Embodiment 1 Referring to Figure 1 , an intelligent question - answering optimization method for power grid equipment based on a large - language model in this embodiment includes: Obtain question data of power grid equipment and create a dynamically updated vocabulary of power grid equipment for storing the question data; Use a part - of - speech tagging model of conditional random fields to tag the part - of - speech of the question data, and assign semantic roles to each word through semantic role annotation; Perform power grid equipment knowledge retrieval based on the question data with tags and semantic roles, including matching documents in the power grid equipment knowledge base and user questions using vector spaces; Build a large - language model with the Transformer architecture as the basic framework, including introducing the GPT architecture and BERT architecture based on the Transformer decoder and encoder respectively; Update the parameters of the built large - language model, including using a variant of the stochastic gradient descent algorithm to update the model parameters; Optimize the answers according to the processing results of the large - language model to obtain the final intelligent question - answering output.
[0025] Specifically, an intelligent question - answering optimization method for power grid equipment based on a large - language model includes the following steps: As Figure 1 shown, S1: Obtain the problem data of grid equipment, and create a dynamically updated vocabulary of grid equipment for storing the problem data; By establishing an interface with the equipment management system of the power grid enterprise, regularly collect the problem records related to the equipment. For example, the equipment management system of a regional power grid company records a large number of problems about transformers, such as "How to handle the too high oil temperature of the transformer" and "What is the reason for the abnormal sound of the transformer". The data collection module, according to the preset time interval (such as 2 am every day), obtains the newly added problem data by calling the API provided by the system, and transmits it to the data storage module for temporary storage. After that, the technical document library contains a large number of operation manuals, technical specifications and training materials for grid equipment. The data collection module uses text mining technology to analyze the documents in the library. First, through keyword matching, such as "Question Answering" and "Common Problems", locate the chapters that may contain problem data. Then, use natural language processing technology to structure the text in these chapters, extract the problem data, preprocess the collected problem data, including removing HTML tags (if the data comes from a web page), converting to lowercase letters, removing stop words (such as words with little impact on semantic understanding like "of", "already", "in", etc.). Set a task for regular update. When updating, re-collect the problem data and repeat the above data preprocessing and word frequency statistics steps. For newly emerging words, if their word frequency reaches a certain threshold, add them to the vocabulary.
[0026] S2: Use the part-of-speech tagging model of conditional random fields to tag the part of speech of the problem data, and assign semantic roles to each word through semantic role labeling; For grid equipment text, analyze the context relationship between words, define a set of feature functions. For each word in the text sequence , the feature function , comprehensively consider the lexical features of the current word itself, the lexical features of its previous word and the next word as well as their corresponding part-of-speech tags and , the feature function specifically includes used to identify that when is a specific grid equipment name and is a noun (equipment name), the value is 1, otherwise it is 0; the feature function used to capture when is an operation verb, is a parameter noun, and is a verb, When it is a noun, the value is 1, and in other cases it is 0. For example, if is "check" (verb), is "voltage" (noun, parameter), and is a verb, is a noun, then the value of this feature function is 1, and in other cases it is 0 According to the defined feature function, construct the conditional probability model of the conditional random field. For the given text sequence x and the corresponding part-of-speech tag sequence y, the conditional probability is expressed as: , where, is the normalization factor, by calculating over all possible part-of-speech tag sequences to perform and obtain, is the feature function weight, representing the importance of this feature function in the model, represents the probability that the part-of-speech tag sequence y appears given the text sequence x. n is the length of the text sequence x, that is, the number of words.
[0027] Collect a large amount of text data of power grid equipment. These data come from various professional documents such as power grid equipment operation manuals, technical reports, and fault diagnosis records. Preprocess the collected text, including removing noise information, standardizing the text format, etc. Then, accurately annotate the part of speech of each word in the text to form a training sample set. Use the gradient descent method to train the conditional random field model to maximize the log-likelihood function of the training data. The log-likelihood function is: , where, is the weight set of the feature function, by calculating the gradient of the log-likelihood function with respect to the weight and updating the weight according to the formula , is the text sequence included in each sample, is the corresponding part-of-speech tag sequence, where, is the learning rate, which controls the step size of each weight update. During the training process, continuously iterate and update the weight until the log-likelihood function converges to a stable value, thereby obtaining a trained part-of-speech tagging model.
[0028] S3. Perform power grid equipment knowledge retrieval based on the labeled and semantic role problem data, including matching the documents in the power grid equipment knowledge base and the user's question using the vector space; For the document collection in the power grid equipment knowledge base, for each document D and each word t in it, calculate the word frequency , and its calculation formula is , where is the number of times the word t appears in the document D. Through this formula, the relative frequency of the word t in the document D is obtained.
[0029] Calculate the inverse document frequency, and the formula is , where is the total number of documents in the document collection , is the number of documents containing the word t, which measures the importance of the word t in the entire document collection. By comprehensively calculating , the basic weight of the word in the document is obtained, laying a foundation for subsequent construction of the document vector.
[0030] Introduce a pre-trained word vector model, such as Word2Vec or BERT embedding, to convert each word t in the document and the user's question into a corresponding semantic vector , so that the vector representation can contain the semantic information of the word. For example, for "transformer" and "substation equipment", the pre-trained word vectors will make their representations in the vector space have similar features, reflecting semantic similarity. According to the power grid equipment knowledge hierarchy, different weight adjustment coefficients are assigned to different types of words. For words of equipment name type, such as "transformer", "circuit breaker", etc., the weight adjustment coefficient is set to , to highlight its core position in the power grid equipment knowledge; for general operation verbs, such as "operate", "inspect", etc., the weight adjustment coefficient is set to . Thus, the final weight of the word t in the document D is obtained: , Based on the final weight of the word, represent the document D as a vector , and the specific method is to perform weighted combination of the weights of all words in the document and the corresponding semantic vectors. For example , thus constructing a document vector representation that can comprehensively reflect word frequency information, semantic information, and knowledge hierarchy. The same method is used to construct the vector representation of the user's question .
[0031] Adopt the cosine similarity calculation method to calculate the user question vector and the document vector in the power grid equipment knowledge base The similarity is calculated, and a similarity threshold is set based on the calculated similarity value. Documents with a similarity greater than the threshold are filtered out. These documents are the ones highly relevant to the user's question. For example, when the user's question is "How to maintain a capacitor", relevant documents such as a capacitor maintenance manual are filtered out through the above vector representation and similarity calculation.
[0032] S4. Build a large language model based on the Transformer architecture, including introducing the GPT architecture and the BERT architecture based on the Transformer decoder and encoder respectively; Establish a multi-head attention mechanism. The traditional attention mechanism focuses on important information in the input sequence by calculating the relationship between the query Q, key K, and value V. The multi-head attention mechanism, on this basis, parallels multiple attention heads. The formula is: , where are the query, key, and value matrices respectively, is the dimension of the key, , and are the linear transformation matrices for different heads, is the linear transformation matrix of the output. Through the parallel processing of multiple heads, the model can simultaneously focus on different aspects of the input sequence and capture richer semantic information. is denoted as the attention head. For example, when processing the text of the power grid equipment operation process, different heads can respectively focus on the equipment name, operation steps, operation conditions, etc. Since the Transformer architecture itself does not have the ability to perceive the sequence order, position encoding is introduced to make up for this defect. The position encoding generates a unique encoding vector for each position through a specific function and is added to the word embedding vector and then input into the model. The common position encoding formula is: , , where represents the position of the keyword in the document, i is the index dimension, represents the dimension of the model. The traditional position encoding is calculated only based on position information and does not consider semantic relevance. This embodiment proposes a method for enhancing semantic-aware position encoding. First, semantic encoding is performed on the input text using the vector data obtained through S3. Semantic encoding is the process of converting the text into a vector representation that the computer can understand and process when processing the input text. The purpose of this process is to enable the computer to capture the semantic information in the text and obtain the semantic representation of each word. Then, for the calculation of position encoding, it no longer depends solely on the position index. , but combined with semantic information, the position encoding calculation formula after combining with speech information is: , , Among them, represents the position of the keyword in the document, i represents the dimension index, represents the dimension of the model, represents the result obtained by semantic encoding, represents the function that converts the semantic representation into a numerical value.
[0033] After that, a GPT structure is built based on the Transformer decoder. Utilizing the autoregressive feature of the decoder, the output sequence is generated sequentially from left to right. During the generation process, the previously generated words are used as input, and through the parametric transformation of the model, the probability distribution of the next word is predicted.
[0034] Self-supervised learning is adopted for pre-training, aiming to minimize the cross-entropy loss function where, is the vector representation of the true next word, is the probability distribution predicted by the model. N represents the number of samples. The model parameters are continuously adjusted to enable it to learn the statistical laws and semantic information of the language, thereby acquiring the ability to generate text that conforms to grammar and semantic logic. Finally, a BERT model structure based on the Transformer encoder is constructed. Through the multi-head attention mechanism, the model can simultaneously focus on the context information of each word in the input text, realizing bidirectional language representation learning.
[0035] Execute the masked language model (MLM) pre-training task. Randomly mask some words in the input text. Using the formula as the loss function, where M is the set of indices of the masked words, is the sequence of unmasked words, is the probability that the model predicts the masked word . Train the model's ability to predict masked words based on the context. Execute the next sentence prediction (NSP) pre-training task. Take two sentences as input and use the binary cross-entropy loss function to train the model to judge whether the second sentence is the next sentence of the first sentence, thereby learning the logical relationship and coherence between sentences.
[0036] S5. Update the parameters of the constructed large language model, including using a variant of the stochastic gradient descent algorithm to update the model parameters; First, we need to determine the appropriate loss function according to the specific fine-tuning task and the Q&A datasets in the field of power grid equipment (these datasets are sourced from texts such as operation manuals, maintenance records, fault reports, technical specifications, etc.) , different loss functions are used for different tasks. In this embodiment, the cross-entropy loss function is used in the text classification task, and the negative log-likelihood loss function is used in the language generation task.
[0037] At each iteration t, calculate the loss function with respect to the current model parameters gradient . This calculation is completed through the backpropagation algorithm. Starting from the loss function, calculate the contribution of each parameter to the loss function layer by layer. Input a text sample in the field of power grid equipment into the model, the model will generate a prediction result, and then compare this prediction result with the true result (from our Q&A dataset) to obtain the difference between them, and this difference is the loss. Through backpropagation, trace how this difference is affected by each parameter of the model, and finally calculate the gradient of each parameter with respect to this loss , and use the gradient to make the prediction result of the model closer to the true result.
[0038] Having obtained the gradient , next calculate the first moment estimate . Here, the formula is used: , where is the decay rate of the first moment estimate, usually taking 0.9. is the first moment estimate of the gradient .
[0039] The purpose of this step is to smooth the change of the gradient. We combine the currently calculated gradient with the historical first moment estimate . Since , we assign a higher weight (0.9) to the historical information and a lower weight (0.1) to the current gradient . This is in the fine-tuning in the field of power grid equipment. We need to consider both the gradient information obtained from this sample in the current iteration and the gradient information accumulated in previous iterations. This can reduce the fluctuation of the gradient because the data in the field of power grid equipment may have noise. For example, non-standard expressions in operation manuals or recording errors in fault reports may lead to abnormal values in some gradients. Calculate the first moment estimate , we can avoid making excessive adjustments to the model parameters due to the gradients of individual noisy data, making the parameter updates more robust and thus making the training of the model more stable.
[0040] While calculating the first-moment estimate, calculate the second-moment estimate , use the formula: , where, is the second-moment estimate of the gradient , is the decay rate of the second-moment estimate, usually taking 0.999.
[0041] The second-moment estimate is the exponential moving average of the squared gradient , which reflects the second moment (uncentered variance) of the gradient. means paying more attention to the historical squared gradient information and only giving the current squared gradient a weight of 0.001. This can measure the change amplitude of the gradient. For words or concepts that appear rarely but are very important in the field of power grid equipment (such as some rare fault types), their gradients may suddenly become large in some iterations, will capture this large gradient change. This helps us adjust different learning rates for different parameters according to the change amplitudes of different gradients in subsequent steps, because different parameters may have different degrees of influence on the model performance in different iterations, and the second-moment estimate can better detect these changes.
[0042] In the initial stage of iteration, since and are calculated through exponential moving average, they are initialized to 0 and rely more on less historical information at the beginning, so they tend to be 0, which may lead to inaccurate estimates. To correct this bias, perform bias correction and use the formula: , , where, represents the t-th power of the decay rate of the first-moment estimate, t is the iteration number, represents the t-th power of the decay rate of the second-moment estimate. Perform bias correction on the first-moment estimate and the second-moment estimate . As the iteration number t increases, and It will gradually approach 0, so the impact of deviation correction will gradually decrease. For fine-tuning in the field of power grid equipment, this deviation correction is very important because when starting fine-tuning with different pre-trained models and datasets, it is necessary to make the model adapt to the new dataset and tasks as soon as possible in the initial stage to avoid slow or inaccurate training due to initial deviations.
[0043] Finally, use the corrected first-order moment estimate and second-order moment estimate to update the model parameters. Through the formula: , for parameter update, where, and represent the model parameters at iterations t and t - 1 respectively, is the learning rate, is a very small constant, such as , used to prevent the denominator from being zero.
[0044] The learning rate controls the step size of parameter update, while the part adaptively adjusts the update step size of each parameter according to the first-order and second-order moment estimates of the gradient. In the field of power grid equipment, different professional terms and concepts have different importance and occurrence frequencies. For example, general operation terms such as "inspection" and "maintenance" will appear frequently, while rare professional fault terms such as "insulator flashover" will appear less frequently. Through this adaptive learning rate adjustment, the model can update parameters at different speeds according to the gradient characteristics of different parameters, and learn the characteristics and relationships of these terms more effectively. In this way, the model can better learn the professional terms, language patterns and knowledge logic in the field of power grid equipment, and then show better performance in various natural language processing tasks of power grid equipment (such as question answering, text generation, information extraction, etc.).
[0045] S6. Optimize the answer according to the processing results of the large language model to obtain the final intelligent question answering output.
[0046] In the field of power grid equipment, the operation process has a clear sequence and logical relationship. For example, for an operation process such as "power off first, then repair", use first-order predicate logic to convert it into a logical expression. Suppose we define the predicate P(x) to represent "perform a power-off operation on device x", and M(x) to represent "perform a repair operation on device x", then this operation process can be represented as , where the " " represents the logical "implication" relationship, that is, if the power-off operation on device x is completed, then the repair operation should be carried out next.
[0047] For more complex operation processes, such as "first cut off the power, then test for voltage, then ground, and finally repair", it can be expressed as , where represents "performing a voltage test operation on device x", represents "performing a grounding operation on device x". In this way, converting the actual operation steps into logical expressions helps to perform formal logical analysis on the operation process.
[0048] For "when a short circuit occurs in the device, it will cause excessive current, which in turn triggers the action of the protection device", we can define the predicate represents "a short circuit occurs in device x", represents "the current in device x is excessive", represents "the protection device of device x acts". Then it can be converted into .
[0049] For the case of multiple causes leading to one effect, such as "device overload or device short circuit will cause the device to overheat", it can be expressed as , where represents "device x is overloaded", represents "device x overheats". " " represents the logical "or", indicating that as long as one of the conditions is met, it will cause the result of the device overheating.
[0050] First, a logical rule library for the operation specifications of power grid equipment needs to be established. This library contains various standard operation processes and logical rules for fault diagnosis in the field of power grid equipment. These rules are summarized based on information such as industry standards, operation manuals, and technical specifications. For example, for the operation of transformers, the rule library may contain rules such as "before performing on-load voltage regulation operation on a transformer, it is necessary to first check the operating status of the transformer". Converting it into a logical expression, such as , where represents "performing on-load voltage regulation operation on transformer x", represents "checking the operating status of transformer x", " " represents the logical "and", that is, the prerequisite for performing on-load voltage regulation operation is to first check the operating status of the transformer. The logical rule library can be stored in a database or a file system and can be updated and expanded according to new information and knowledge in the field of power grid equipment.
[0051] After the large language model processes the user's question and generates the corresponding logical expressions for operation processes or fault diagnosis, we match these generated logical expressions with the expressions in the logical rule library. For example, if the logical expression generated by the large language model is ),search in the logic rule library to check if there is the same or conflicting logical expression. If there is and is inconsistent with , it indicates that there may be problems with the generated answer and optimization is required.
[0052] For more complex situations, it may involve the equivalence judgment of logical expressions and logical reasoning. For example, the generated expression is , while in the rule library it is and . Through the equivalence judgment of logic, it can be found that the generated expression may be redundant or incorrect, and the answer needs to be adjusted to conform to the standard operation process and logical rules.
[0053] By converting the processing results of the large language model into logical expressions and matching them with the logical rule library, the logical accuracy and consistency of the generated answer can be ensured. It can avoid situations that violate the operation specifications of power grid equipment and fault diagnosis logic, improve the quality and reliability of the answer. For complex operation processes and fault diagnosis explanations, it can avoid logical errors such as reversed operation order and incorrect causal relationship, making the final intelligent Q&A output more in line with professional requirements.
[0054] In the operation and maintenance platform of power grid equipment, when the user asks about operation steps or fault reasons, the large language model will generate corresponding answers. This step can ensure the logical correctness of these answers. For example, when the user asks "how to perform maintenance operations on a circuit breaker", the large language model may generate a series of operation steps. After converting these steps into logical expressions and matching them with the logical rule library, it can ensure that the sequence and logical relationship of the operation steps are correct. By converting the processing results of the large language model into logical expressions and matching them with the logical rule library, the answer can be optimized from a logical perspective, ensuring the professionalism and reliability of the intelligent Q&A system in the field of power grid equipment, and providing high-quality answers that conform to operation specifications and logic for users. This method utilizes the rigor of first-order predicate logic, converts the information in natural language form into formal logical expressions, and then combines the pre-established logical rule library to check and optimize the answer, which is an important link to ensure the performance of the intelligent Q&A system.
[0055] Embodiment 2 The difference between this embodiment and Embodiment 1 is that this embodiment provides a power grid equipment intelligent Q&A optimization system based on a large language model, including: A data acquisition module, configured to: acquire problem data of power grid equipment and create a dynamically updated power grid equipment vocabulary for storing the problem data; An analysis module, configured to: label the part-of-speech of the question data using a part-of-speech tagging model of conditional random field, and assign semantic roles to each word through semantic role labeling; A conversion module, configured to: perform power grid equipment knowledge retrieval based on the question data with labeled part-of-speech and semantic roles, including matching the documents in the power grid equipment knowledge base and the user's question using a vector space; A model module, configured to: build a large language model with the Transformer architecture as the basic framework, including introducing the GPT architecture and the BERT architecture based on the Transformer decoder and encoder respectively; A training module, configured to: update the parameters of the built large language model, including updating the model parameters using a variant algorithm of stochastic gradient descent; An optimization module, configured to: optimize the answer according to the processing result of the large language model to obtain the final intelligent question and answer output.
[0056] A computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded and executed by a processor of a terminal device for the above-mentioned power grid equipment intelligent question and answer optimization method based on a large language model.
[0057] A terminal device, including a processor and a computer-readable storage medium, the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are suitable for being loaded and executed by the processor for the above-mentioned power grid equipment intelligent question and answer optimization method based on a large language model.
[0058] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.
Claims
1. A method for optimizing intelligent question and answer of power grid equipment based on a large language model, characterized in that: include: Obtain problem data of power grid equipment and create a dynamically updated power grid equipment vocabulary to store problem data; The part-of-speech tagging model of conditional random fields is used to tag the part-of-speech of the question data, and semantic roles are assigned to each word through semantic role tagging; Retrieve knowledge about power grid equipment based on question data with annotations and semantic roles, including matching documents in the power grid equipment knowledge base with user questions using vector space; Build a large language model based on the Transformer architecture, including the introduction of the GPT architecture and the BERT architecture based on the Transformer decoder and encoder respectively; Update the parameters of the constructed large language model, including using a variant of the stochastic gradient descent algorithm to update the model parameters; The answer is optimized based on the processing results of the large language model to obtain the final intelligent question and answer output.
2. According to claim 1, a method for optimizing intelligent question and answer of power grid equipment based on a large language model is characterized in that: The method uses a conditional random field part-of-speech tagging model to tag the parts of speech of the problem data, including defining a feature function based on a text sequence in the problem data, and constructing a conditional probability model of the conditional random field according to the defined feature function, using the power grid equipment text with the parts of speech annotated as training data, training the conditional probability model to maximize the log-likelihood function of the training data, and tagging the parts of speech of the problem data according to the trained conditional probability model.
3. The method for optimizing intelligent question and answer of power grid equipment based on a large language model according to claim 1, characterized in that: The method uses vector space to match documents and user questions in a power grid equipment knowledge base, including using document and user question data to respectively calculate word frequency and inverse document frequency, calculating the basic weight of question data in the document based on the word frequency and inverse document frequency, assigning different weight adjustment coefficients to different types of keywords according to the power grid equipment knowledge hierarchy, obtaining the final weight of the keyword in the document, using cosine similarity to calculate the similarity between the user question vector and the document vector in the power grid equipment knowledge base, setting a similarity threshold according to the calculated similarity value, and screening out documents with a similarity greater than the threshold.
4. The method for optimizing intelligent question and answer of power grid equipment based on a large language model according to claim 1, characterized in that: The Transformer architecture is used as the basic framework to build a large language model, including semantic encoding of the input text according to the vector data to obtain the semantic representation of each word, introducing the position encoding module and calculating the position encoding in combination with the semantic information. The position encoding calculation formula after combining the voice information is: , in, It represents the position of the keyword in the document, i represents the dimension index, Expressed as the dimension of the model, It is represented as the result of semantic encoding. Represented as a function that converts a semantic representation into a numerical value.
5. The method for optimizing intelligent question and answer of power grid equipment based on a large language model according to claim 1, characterized in that: The method introduces the GPT architecture and the BERT architecture based on the Transformer decoder and encoder respectively, including building a GPT structure based on the Transformer decoder, using the autoregressive characteristics of the decoder to generate output sequences from left to right, using self-supervised learning to perform pre-training, and then building a BERT architecture based on the Transformer encoder, performing bidirectional language representation learning through a multi-head attention mechanism, and training the BERT architecture by randomly masking some words in the input text.
6. The method for optimizing intelligent question and answer of power grid equipment based on a large language model according to claim 1, characterized in that: The method of updating the model parameters by using a variant of the stochastic gradient descent algorithm includes defining a loss function through a question-and-answer data set in the field of power grid equipment, calculating the gradient of the current model parameters according to the loss function, calculating the first-order moment estimate and the second-order moment estimate in sequence by using the gradient of the current model parameters, performing deviation correction on the first-order moment estimate and the second-order moment estimate, and finally updating the model parameters according to the corrected first-order moment estimate and the second-order moment estimate.
7. The method for optimizing intelligent question and answer of power grid equipment based on a large language model according to claim 1, characterized in that: The answer optimization based on the processing results of the large language model includes using first-order predicate logic to convert the operation process and fault diagnosis logic of the power grid equipment into logical expressions, and matching the generated logical expressions with a pre-established logical rule base of the power grid equipment operation specifications.
8. An intelligent question-answering optimization system for power grid equipment based on a large language model, characterized in that: include: The data acquisition module is configured to: acquire problem data of power grid equipment and create a dynamically updated power grid equipment vocabulary to store the problem data; The analysis module is configured to: tag the part of speech of the question data using a part-of-speech tagging model of a conditional random field, and assign a semantic role to each word through semantic role tagging; A conversion module, configured to: perform power grid equipment knowledge retrieval based on the question data with annotations and semantic roles, including matching documents in the power grid equipment knowledge base with user questions using a vector space; The model module is configured to: build a large language model based on the Transformer architecture, including introducing the GPT architecture and BERT architecture based on the Transformer decoder and encoder respectively; The training module is configured to: update the parameters of the constructed large language model, including updating the model parameters using a stochastic gradient descent variant algorithm; The optimization module is configured to optimize the answer according to the processing results of the large language model to obtain the final intelligent question and answer output.
9. A computer-readable storage medium storing a plurality of instructions, characterized in that: The instructions are suitable for being loaded by a processor of a terminal device and executing a method for optimizing intelligent question and answer of power grid equipment based on a large language model as described in claim 1.
10. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is used to implement each instruction; and the computer-readable storage medium is used to store multiple instructions, characterized in that: The instructions are suitable for being loaded by a processor and executing a method for optimizing intelligent question and answer of power grid equipment based on a large language model as described in claim 1.
Citation Information
Cited By
Keyword library and vector-based hybrid retrieval knowledge base construction system and method
CN121031758A
Multi-hop reasoning method and device based on masking knowledge activation
CN121998104A