Power supply question and answer method, device and equipment based on pseudo feedback and attention mechanism and medium

Through the powered question-and-answer method of pseudo-feedback and attention mechanism, combined with BM25 and reading comprehension model, the problem of inaccurate answer generation of multi-hop questions in the existing technology is solved, and the accuracy and reasoning ability of answer generation are improved, especially in application in power documents.

CN120371956APending Publication Date: 2025-07-25STATE GRID BEIJING ELECTRIC POWER CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510440689.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing reading comprehension models are not effective when dealing with multi-hop user questions, making it difficult to effectively answer multiple questions raised by users or questions that require reasoning to answer.

Method used

The powered question-and-answer method based on pseudo-feedback and attention mechanism is adopted. Through the combination of BM25 model and reading comprehension model, combined with the BERT model, parallel convolutional self-attention model and LSTM model, single-hop and multi-hop problems are handled, and the pseudo-associated feedback and attention mechanism are used to improve the accuracy of answer generation.

Benefits of technology

Improve the accuracy of answer generation of multi-hop questions, improve the effect of reading comprehension, enhance the ability to reason across sentences in power documents, and reduce the return of non-related answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371956A_ABST
    Figure CN120371956A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of power questions and answers, and aims to solve the problems that most power supply problems are multi-hop problems, and an existing reading understanding model is poor in power supply multi-hop problem answering effect. Inputting the article and the initial question into a BM25 model for correlation sorting, and outputting a plurality of TOP-K1 sentences related to the initial question; inputting the initial question and a TOP-K1 sentence into a reading understanding model, and outputting a first answer; splicing the first answer and the initial question as a second question; according to the second question, inputting the article into a BM25 model for correlation sorting, and outputting a plurality of related TOP-K2 sentences; inputting the second question and the TOP-K2 sentence into the reading understanding model, and outputting a second answer; the second answer is a final answer. The false correlation feedback is introduced into the BERT model, the initial question query is expanded, and the accuracy of answer generation is improved by screening initial answers and reducing noise sentences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power Q&A, and particularly relates to a power supply Q&A method, device, equipment and medium based on pseudo-feedback and attention mechanism. Background Art

[0002] At present, natural language processing, especially text reading comprehension technology, has shown unique advantages in text Q&A of online Q&A services based on knowledge bases. The text reading comprehension technology can accurately analyze the user's question description, quickly identify the type of user questions, such as inquiries about electricity handling procedures or questions about electricity prices, and instantly provide effective solutions according to the knowledge base. In addition, it can deeply understand the potential needs of customers, such as energy-saving suggestions or consultations on electrical safety knowledge, and ensure that the most relevant content is quickly extracted and provided from a large number of prefabricated speech templates. The application of machine reading comprehension technology not only speeds up the service response speed, improves the accuracy of the service, but also significantly improves the professional level of seat personnel when answering user questions, and improves customer satisfaction and trust.

[0003] The current research work on dedicated small models for reading comprehension is divided according to the coherence of the text required to generate answers. User questions that only require a single text segment to generate answers are called single-hop user questions, and user questions that require multiple discontinuous text segments to answer are called multi-hop user questions. In recent years, more research has focused on single-hop English datasets, and most of the dataset settings are based on knowledge graphs or fixed forms. After statistics, it is found that in the existing intelligent online Q&A service dialogue tasks, the answers and the number of hops corresponding to each user question are uncertain, and there are three types of user questions in total: (1) single-hop user questions, multi-hop reasoning user questions, and single-time multi-questions. In the actual online Q&A service scenario, users will simultaneously ask multiple user questions or ask user questions that require reasoning in one input. However, the existing dedicated small models for reading comprehension are only applicable to single-hop user questions and have poor effects when answering multi-hop questions, while most of the power supply questions asked by users are multi-hop questions. Summary of the Invention

[0004] The purpose of the present invention is to provide a power supply Q&A method, device, equipment and medium based on pseudo-feedback and attention mechanism, so as to solve the defect that the existing reading comprehension model in the background art is difficult to handle multi-hop questions.

[0005] In order to achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect of the present invention, a power supply Q&A method based on pseudo-feedback and attention mechanism is provided, which is characterized by including: Obtain an article and an initial question of a power supply user; Classify the initial question; among them, the categories of the initial question include single-hop questions and multi-hop questions; When the initial question is a single-hop question: input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question; When the initial question is a multi-hop question: input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the first answer; splice the first answer and the initial question to obtain the second question; input the second question and the article into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K2 sentences related to the second question; input the second question and the TOP-K2 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question.

[0006] Preferably, the step of classifying the initial question includes: Format the initial question into the format of [CLS]+initial question text+[SEP]; where [CLS] represents the classification flag and [SEP] represents the separation flag; Input the formatted initial question into the BERT-base model. In the BERT-base model, perform a linear transformation on the formatted initial question to obtain eigenvalues, calculate the eigenvalues using the softmax function to obtain the probabilities of each category, and classify the initial question into the category with the highest probability.

[0007] Preferably, in the step of inputting the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question: The reading comprehension model includes a BERT model, a parallel convolutional self-attention model, and an LSTM model; Use the BERT model to encode the initial question and the TOP-K1 sentences to obtain the encoded initial question and TOP-K1 sentences; Input the encoded initial question and TOP-K1 sentences into the parallel convolutional self-attention model, and the parallel convolutional self-attention model calculates the answer start position feature and the answer end position feature of the first answer; Input the answer start position feature and the answer end position feature of the first answer into the LSTM model to identify the start position and the end position of the first answer in the TOP-K1 sentences; Intercept the text between the start position and the end position in the TOP-K1 sentences as the first answer.

[0008] Preferably, in the step of inputting the second question and the TOP-K2 sentences into the reading comprehension model and the reading comprehension model outputting the final answer to the initial question: The reading comprehension model includes a BERT model, a parallel convolutional self-attention model, and an LSTM model; Use the BERT model to encode the second question and the TOP-K2 sentences to obtain the encoded second question and TOP-K2 sentences; Input the encoded second question and TOP-K2 sentences into the parallel convolutional self-attention model, and the parallel convolutional self-attention model calculates the answer start position feature and the answer end position feature of the second answer; Input the answer start position feature and the answer end position feature of the second answer into the LSTM model to identify the start position and the end position of the second answer in the TOP-K2 sentences; Intercept the text between the start position and the end position in the TOP-K2 sentences as the second answer.

[0009] Preferably, the BERT model includes: First, project the TOP-K1 sentences and the initial question Q respectively using and three matrices to obtain the projected matrices and , where W Q 1 is the query vector matrix, W K 1 is the key vector matrix, W V 1 is the value vector matrix; Calculate the similarity matrix S of the TOP-K1 sentences and the initial question Q:

[0010] where d k is the spatial dimension; The attention of the TOP-K1 sentences relative to the initial question Q is :

[0011] The attention of the initial question Q relative to the TOP-K1 sentences is :

[0012] Among them, represents stacking m times, max col represents taking the maximum value of the similarity matrix S column by column; The extracted features , and are concatenated to form an aligned feature matrix X 1:

[0013] The aligned feature matrix X 1 is enhanced by a fully connected feed-forward network FFN, and the original value vector V c is added to obtain the integrated feature matrix X 2:

[0014] The integrated feature matrix X 2 is projected h times differently to obtain h sets of vectorized representation feature matrices X 3:

[0015] Among them, d = h * d k .

[0016] Preferably, the parallel convolutional self-attention model includes: Construct a convolutional neural network, set three parallel one-dimensional convolutions CNN i , i = 1, 2, 3, concatenate the outputs of the three one-dimensional convolutions, and then pass through another one-dimensional convolution CNN 4 ; Input the vectorized representation feature matrix X 3 into the convolutional neural network, and add the vectorized representation feature matrix X 3 as a residual connection. The output at this time is Z:

[0017] After outputting Z, use the multi-head self-attention mechanism to calculate the answer start and end position features O :

[0018]

[0019]

[0020] Among them, the projection matrix ; Perform two independent linear mappings on the head and tail position features of the answer to separately generate the head position feature of the answer O and the tail position feature of the answer O 11 O 21

[0021] Preferably, the LSTM model includes: The probability of predicting the head position feature of the answer O 11 as the probability of the head pointer of an answer interval:

[0022] Input the head position feature of the answer O 11 and the tail position feature of the answer O 21 , and output and generate the probability of the tail position feature of the answer O 21 as the probability of the tail pointer of an answer interval:

[0023] The final training loss is the cross-entropy loss of the prediction probabilities of the head and tail pointer index positions:

[0024] Minimize the training loss, and the obtained a S is the head position of the first answer, a E is the tail position of the first answer, and the text between the head position and the tail position of the first answer is the first answer.

[0025] In the second aspect of the present invention, a power supply question and answer device based on pseudo-feedback is provided, including: An acquisition module for obtaining an article and an initial question of a power supply user; A classification module for classifying the initial question; A single-hop question module for when the initial question is a single-hop question: input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question; ​​The multi-hop question module is used when the initial question is a multi-hop question: input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the first answer; splice the first answer and the initial question to obtain the second question; input the second question and the article into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K2 sentences related to the second question; input the second question and the TOP-K2 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question.

[0026] In the third aspect of the present invention, an electronic device is provided, including a processor and a memory, and the processor is used to execute a computer program stored in the memory to implement the power supply question-answering method based on pseudo-feedback and attention mechanism.

[0027] In the fourth aspect of the present invention, a computer-readable storage medium is provided, and the computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the power supply question-answering method based on pseudo-feedback and attention mechanism is implemented.

[0028] Compared with the prior art, the beneficial effects of the present invention are as follows: Introduce pseudo-association feedback into the BERT model, expand the initial question query, reduce noise sentences through initial answer screening, improve the accuracy of answer generation, and improve the reading comprehension effect; Introduce a parallel convolutional attention module to capture local and global semantic dependencies simultaneously, solve the cross-sentence reasoning problem in long text power documents to promote the interaction of local and global information, and thus improve the reasoning ability; The parallel convolutional module effectively processes the complex sentence structure in power documents, the accuracy rate on the power data set, and the attention mechanism strengthens the matching of the question and power terms in the document, reducing the return of irrelevant answers. Description of the Drawings

[0029] The specification drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 It is a flowchart of the power supply question-answering method based on pseudo-feedback and attention mechanism in Embodiment 1 of the present invention; Figure 2 It is a framework diagram of the power supply question-answering method based on pseudo-feedback and attention mechanism in Embodiment 1 of the present invention; Figure 3 It is a structural diagram of the BERT model of the present invention; Figure 4 This is the structural block diagram of the power supply question-answering device based on the pseudo-feedback and attention mechanism in Embodiment 2 of the present invention; Figure 5 This is the structural block diagram of an electronic device in Embodiment 3 of the present invention. Detailed implementation manners

[0030] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0031] The following detailed descriptions are all exemplary descriptions, aiming to provide further detailed descriptions of the present invention. Unless otherwise specified, all technical terms adopted by the present invention have the same meaning as commonly understood by those of ordinary skill in the art to which the present application belongs. The terms used in the present invention are only for the purpose of describing specific implementation manners, and are not intended to limit the exemplary implementation manners according to the present invention.

[0032] Embodiment 1 A question-answering method based on a pseudo-feedback attention mechanism, comprising: S1. Obtain an article and an initial question Q of a power supply user: The article referred to in this solution may be "Technical Management Regulations for Electric Energy Measurement Devices" and "General Design Specifications for Electric Energy Measurement Devices".

[0033] S2. Classify the initial question; wherein, the categories of the initial question include single-hop questions and multi-hop questions The samples of single-shot multiple questions correspond to 2 answers, and single-hop and multi-hop initial questions correspond to 1 answer. Therefore, classifying the initial question is regarded as a binary classification task, and the label is determined by the number of answers.

[0034] According to the input form of the BERT model, the initial question Q is concatenated as [CLS] + initial question + [SEP], where [CLS] is the classification marker of the Bert-base model, and [SEP] serves as a separator, and then it is sent to the BERT-base model for encoding. The encoded initial question is input into a linear layer, and the output is the eigenvalue after linear transformation. Then, the softmax function is used to calculate the eigenvalue after linear transformation to obtain the probability of each category, and the initial question is classified into the category with the highest probability:

[0035] In the formula, is the label category, is the predicted value.

[0036] The initial question is divided into single-hop questions and multi-hop questions according to the number of answers; For single - turn multiple questions: First, preferentially use the preset rules to split and generate sub - questions, and then classify the sub - questions into single - hop questions and multi - hop questions; if the rules fail, classify them as multi - hop questions; the rule is string matching, such as splitting according to question marks.

[0037] When the initial question is a single - hop question: S31. Input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP - K1 sentences related to the initial question: Adopt the BM25 ranking algorithm to score according to the relevance degree between each sentence Ci in the article and the initial question of the power - supply user Q and record it as , select the first m sentences with the highest scores to obtain the TOP - K1 sentences C:

[0038] Among them, c m is the m th relevant sentence.

[0039] S32. Input the initial question and the TOP - K1 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question; First, input the TOP - K1 sentences C and the initial question Q into the BERT model: Record the TOP - K1 sentences as C :

[0040] Record the initial question Q as:

[0041] Among them, n is the number of initial questions, m is the number of TOP - K1 sentences, d is the output dimension of BERT, is the vector space of the BERT model.

[0042] Alignment is a necessary step for information fusion and interaction between the TOP - K1 sentences and the initial question. In this application, the alignment between the TOP - K1 sentences and the initial question is achieved through the attention in two directions: the TOP - K1 sentences to the initial question (C2Q) and the initial question to the TOP - K1 sentences (Q2C), and finally, the representation of the TOP - K1 sentences perceived by the initial question is obtained. The specific steps are as follows: First, project the TOP - K1 sentences and the initial question Q respectively using and three matrices to obtain the projected matrices and , among which, W Q1 is the query vector matrix, W K 1 is the key vector matrix, W V 1 is the value vector matrix.

[0043] Calculate the similarity matrix S between the TOP-K1 sentences and the initial question Q: (1) where d k is the dimension.

[0044] The attention of the TOP-K1 sentence C relative to the initial question Q is , corresponding to Figure 3 Align(C,Q) in (2) The attention of the initial question Q relative to the TOP-K1 sentence C is , corresponding to Figure 3 Align(Q,C) in (3) where means stacking m times, and max col means taking the maximum value by column for the similarity matrix S.

[0045] Concatenate the extracted features , and to form the aligned feature matrix X 1: (4) Enhance the aligned feature matrix X 1 through the fully connected feed-forward network FFN, and add the original value vector, i.e., the matrix V c , as cross-layer connection, to obtain the integrated feature matrix X 2: (5) Perform h different projections on the integrated feature matrix X 2 to obtain h groups of vectorized representation feature matrices X 3: (6) where d = h * d k .

[0046] The encoded initial question and the encodings of the TOP-K1 sentences are input into the parallel convolutional self-attention model, and the parallel convolutional self-attention model calculates the answer start position feature and the answer end position feature of the first answer: The parallel convolutional self-attention model includes: Construct a convolutional neural network: Parallel three one-dimensional convolutions CNN i (i=1,2,3) , where the window sizes are 3, 5, and 7. Concatenate the outputs of the three convolutions, then pass through another one-dimensional convolution, and then add the feature matrix of the vectorized representation X 3 is used as the residual connection for easy training. The output at this time is: (7) For all the above convolutions, set the number of output channels to be equal to the number of input channels, and then use the multi-head self-attention mechanism to calculate the answer start and end position features after the output Z O : (8) (9) (10) Among them, the projection matrix is ; The answer start and end position features O are subjected to two independent linear mappings to respectively generate the answer start position feature O 11 and the answer end position feature O 21 : (11) (12) Among them, B is the batch size and L is the sequence length, both of which are set manually.

[0047] Set the LSTM model, input the features of the start and end positions of the first answer into the LSTM model, predict the probability of the answer, and output the start and end positions of the first answer: First, predict the probability of the answer start position feature O 11 as the probability of a pointer to the start of an answer interval: (13) Then, input the answer start position feature O 11 and the answer end position feature O 21 , and generate the answer end position feature through the LSTM model outputO 21 Probability as the tail pointer of an answer interval: (14) The final training loss is the probability of the answer tail position feature O 21 and the probability of the answer tail position feature O 21 cross-entropy loss: (15) Minimize the training loss, and the obtained a S is the start position of the first answer, a E is the end position of the first answer. The text between the start position and the end position of the first answer is the first answer, and the first answer is the final answer.

[0048] S4. When the initial question is a multi-hop question: S41. Input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; S42. Input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the first answer; The generation method of the first answer is the same as the steps in formulas (1) to (15), which will not be elaborated here.

[0049] S43. Connect the first answer Q and the initial question with [SEP] to obtain the second question Q':

[0050] S44. Input the second question and the article into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K2 sentences related to the second question; Retrieve the sentences related to the second question from the article, and record the final sentence retrieved in the second step as: , for each sub-question of the second question, select the first m sentences with the highest scores to obtain the TOP-K2 sentences.

[0051] S45. Input the second question Q' and the TOP-K2 sentences into the reading comprehension model, and the reading comprehension model outputs the second answer; The second answer is the final answer to the initial question: Connect the TOP-K2 sentences and the second question Q' with and Three matrices are projected. The method for generating the second answer is the same as the steps in formulas (1) to (15), which will not be elaborated here.

[0052] Embodiment 2 As Figure 4 shown, based on the same inventive concept as the above embodiment, the present invention also provides a power supply question-answering device based on pseudo-feedback and attention mechanism, including: An acquisition module for obtaining an article and an initial question of a power supply user; A classification module for classifying the initial question; A single-hop question module, when the initial question is a single-hop question: input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question; A multi-hop question module, when the initial question is a multi-hop question: input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the first answer; splice the first answer and the initial question to obtain the second question; input the second question and the article into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K2 sentences related to the second question; input the second question and the TOP-K2 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question.

[0053] Embodiment 3 As Figure 5 shown, the present invention also provides an electronic device 100 for implementing the power supply question-answering method based on pseudo-feedback and attention mechanism; The electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on at least one processor 102, and at least one communication bus 104.

[0054] The memory 101 can be used to store the computer program 103. The processor 102 realizes the steps of the power supply question-answering method based on pseudo-feedback and attention mechanism in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101.

[0055] The memory 101 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0056] At least one processor 102 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or the processor 102 may also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 through various interfaces and lines.

[0057] The memory 101 in the electronic device 100 stores multiple instructions to implement a power supply question-and-answer method based on a pseudo-feedback and attention mechanism. The processor 102 can execute the multiple instructions to implement: Obtain an article and an initial question of the power supply user; Classify the initial question; among them, the categories of the initial question include single-hop questions and multi-hop questions; When the initial question is a single-hop question: input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question; When the initial problem is a multi-hop problem: The article and the initial problem are input into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial problem; the initial problem and the TOP-K1 sentences are input into the reading comprehension model, and the reading comprehension model outputs the first answer; the first answer and the initial problem are concatenated to obtain the second problem; the second problem and the article are input into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K2 sentences related to the second problem; the second problem and the TOP-K2 sentences are input into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial problem.

[0058] Embodiment 4 If the modules / units integrated in the electronic device 100 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory and read-only memory (ROM, Read-Only Memory).

[0059] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, system, or computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0060] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for realizing the processesFigure 1 one process or multiple processes and / or blocks Figure 1 means for the functions specified in one block or multiple blocks.

[0061] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 in one block or multiple blocks.

[0062] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 in one block or multiple blocks.

[0063] In the description of this specification, the descriptions referring to the terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A power supply question-answering method based on pseudo-feedback and attention mechanism, characterized in that Including: Obtain the initial questions of the article and the power supply users; Classify the initial questions; wherein, the categories of the initial questions include single-hop questions and multi-hop questions; When the initial question is a single-hop question: Input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; Input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question; When the initial question is a multi-hop question: Input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; Input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the first answer; Concatenate the first answer and the initial question to obtain the second question; Input the second question and the article into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K2 sentences related to the second question; Input the second question and the TOP-K2 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question.

2. The power supply question and answer method based on pseudo-feedback and attention mechanism according to claim 1, characterized in that The step of classifying the initial questions includes: Format the initial question into the format of [CLS]+initial question text+[SEP]; where [CLS] represents the classification flag and [SEP] represents the separation flag; Input the formatted initial question into the BERT-base model. In the BERT-base model, perform a linear transformation on the formatted initial question to obtain eigenvalues, calculate the eigenvalues using the softmax function to obtain the probabilities of each category, and classify the initial question into the category with the highest probability.

3. The power supply question-answering method based on pseudo-feedback and attention mechanism according to claim 1, characterized in that, In the step of inputting the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question: The reading comprehension model includes a BERT model, a parallel convolutional self-attention model, and an LSTM model; Use the BERT model to encode the initial question and the TOP-K1 sentences to obtain the encoded initial question and TOP-K1 sentences; Input the encoded initial question and TOP-K1 sentences into the parallel convolutional self-attention model, and the parallel convolutional self-attention model calculates the answer start position feature and the answer end position feature of the first answer; Input the answer start position feature and the answer end position feature of the first answer into the LSTM model to identify the start position and the end position of the first answer in the TOP-K1 sentences; Intercept the text between the start position and the end position in the TOP-K1 sentences as the first answer.

4. The power supply question-answering method based on pseudo-feedback and attention mechanism according to claim 1, characterized in that In the step of inputting the second question and the TOP-K2 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question: The reading comprehension model includes a BERT model, a parallel convolutional self-attention model, and an LSTM model; Use the BERT model to encode the second question and the TOP-K2 sentences to obtain the encoded second question and TOP-K2 sentences; Input the encoded second question and TOP-K2 sentences into the parallel convolutional self-attention model, and the parallel convolutional self-attention model calculates the answer start position feature and answer end position feature of the second answer; Input the answer start position feature and answer end position feature of the second answer into the LSTM model to identify the start position and end position of the second answer in the TOP-K2 sentences; Intercept the text between the start position and the end position in the TOP-K2 sentences as the second answer.

5. The power supply question-answering method based on pseudo-feedback and attention mechanism according to claim 3, characterized in that, The BERT model includes: First, project the TOP-K1 sentences and the initial question Q respectively using and onto three matrices to obtain the projected matrices and , where W Q 1 is the query vector matrix, W K 1 is the key vector matrix, W V 1 is the value vector matrix; Calculate the similarity matrix S of the TOP-K1 sentences and the initial question Q: Among them, d k is the spatial dimension; The attention of the TOP-K1 sentences to the initial question Q is : The attention of the initial question Q to the TOP-K1 sentences is :[[-END]] Among them, represents stacking m times, max col represents taking the maximum value of the similarity matrix S column by column; The extracted features , and are concatenated to form an aligned feature matrix X 1: The aligned feature matrix X 1 is enhanced by a fully-connected feed-forward network FFN, and a matrix V c is added to obtain the integrated feature matrix X 2: For the integrated feature matrix X Perform h different projections on 2 to obtain h sets of feature matrices represented in vector form X 3: Among them, d = h * d k 。 6. The power supply question-answering method based on pseudo-feedback and attention mechanism according to claim 5, wherein The parallel convolutional self-attention model includes: Build a convolutional neural network and set three parallel one-dimensional convolutions CNN i , i = 1, 2, 3, concatenate the outputs of the three one-dimensional convolutions, and then pass through another one-dimensional convolution CNN 4 ; The feature matrix represented in vector form X 3 is input into the convolutional neural network, and then the feature matrix represented in vector form X 3 is used as a residual connection, and the output at this time is Z: Calculate the feature of the start and end positions of the answer using the multi-head self-attention mechanism after the output Z O : Among them, the projection matrix ; Perform two independent linear mappings on the start and end position features of the answer O to generate the start position feature of the answer O 11 and the end position feature of the answer O 21 .

7. The power supply question-answering method based on pseudo-feedback and attention mechanism according to claim 6, characterized in that The LSTM model includes: Probability of the first position feature of the predicted answer O 11 Probability as the pointer to the start of an answer span: The starting position feature of the input answer O 11 and the ending position feature of the answer O 21 , generate the ending position feature of the answer through the LSTM model O 21 The probability of serving as the pointer to the end of an answer interval: The final training loss is the cross-entropy loss of the prediction probability of the start and end pointer index positions: Minimize the training loss, and the obtained a S is the starting position of the first answer, a E is the ending position of the first answer. The text between the starting position and the ending position of the first answer is the first answer.

8. The power supply question-and-answer device based on pseudo-feedback is characterized in that, Include: An acquisition module for obtaining the initial questions of the article and the power supply user; A classification module for classifying the initial questions; A single-hop question module, when the initial question is a single-hop question: input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question; A multi-hop question module, when the initial question is a multi-hop question: input the article and the initial question into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K1 sentences related to the initial question; input the initial question and the TOP-K1 sentences into the reading comprehension model, and the reading comprehension model outputs the first answer; Concatenate the first answer and the initial question to obtain the second question; Input the second question and the article into the BM25 model for relevance ranking, and the BM25 model outputs several TOP-K2 sentences related to the second question; input the second question and the TOP-K2 sentences into the reading comprehension model, and the reading comprehension model outputs the final answer to the initial question.

9. An electronic device, characterized in that, It includes a processor and a memory. The processor is used to execute the computer program stored in the memory to implement the power supply question and answer method based on pseudo-feedback and attention mechanism as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, it implements the power supply question and answer method based on pseudo-feedback and attention mechanism as described in any one of claims 1 to 7.