An answer detection method and device

By integrating the glyph information of the to-processed documents into the pre-trained model, a richer coding vector is generated, which solves the problem of insufficient answer detection accuracy in the prior art, and achieves higher answer detection accuracy.

CN114416952BActive Publication Date: 2025-07-01BEIJING KINGSOFT DIGITAL ENTERTAINMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210068692.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-13
Filing Date
2022-01-20
Publication Date
2025-07-01
Estimated Expiration
2042-01-20

AI Technical Summary

Technical Problem

When performing reading comprehension tasks, the existing pre-trained model determines whether there is an answer in the text to be detected and what is the specific answer, resulting in insufficient accuracy of answer detection.

Method used

Obtain the glyph information corresponding to each word unit in the pending document and the question to be queried, and input the answer detection model as the input set. The vector encoding module and the probability prediction module generate the coding vector to determine the answer detection result.

Benefits of technology

By integrating glyph features, the accuracy of the answer detection results of the answer detection model in the pending documents is improved, and more features are used for answer detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114416952B_ABST
    Figure CN114416952B_ABST
Patent Text Reader

Abstract

The present application provides an answer detection method and apparatus, wherein the answer detection method includes: obtaining glyph information corresponding to each word unit in a document to be processed and a question to be queried, inputting the document to be processed, the question to be queried, and the glyph information as an input set into an answer detection model to obtain an encoded vector of the input set, and determining and outputting an answer detection result corresponding to the question to be queried in the document to be processed according to the encoded vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence in computer technology, and particularly to an answer detection method and apparatus, a computing device, and a computer-readable storage medium. Background Art

[0002] Artificial intelligence (AI) refers to the ability of an engineered (i.e., designed and manufactured) system to perceive the environment, as well as the ability to acquire, process, apply, and represent knowledge. The development status of key technologies in the field of artificial intelligence includes key technologies such as machine learning, knowledge graphs, natural language processing, computer vision, human-computer interaction, biometric recognition, virtual reality / augmented reality, etc. Natural Language Processing (NLP) refers to the use of a computer to process information such as the form, sound, and meaning of natural language, that is, operations and processing on the input, output, recognition, analysis, understanding, generation, etc. of words, phrases, sentences, and texts. NLP is an important direction in the field of computer science and artificial intelligence, and it studies various theories and methods that can achieve effective communication between humans and computers in natural language.

[0003] For natural language processing tasks, pre-trained models are usually selected for processing. The current conventional method for machine reading comprehension is to input the question and text into a pre-trained model, and the model processes the question and text accordingly to obtain the start and end positions of the answer corresponding to the question in the text. It can be seen that when the existing pre-trained model performs a reading comprehension task, it only determines whether there is an answer in the text to be detected and what the specific answer is by the start position and end position of the answer. The accuracy of the answer output in this way needs to be improved. Summary of the Invention

[0004] In view of this, embodiments of the present application provide an answer detection method and apparatus, a computing device, and a computer-readable storage medium to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of the present application, an answer detection method is provided, including:

[0006] Obtain the glyph information corresponding to each word unit in the document to be processed and the question to be queried;

[0007] Take the document to be processed, the question to be queried, and the glyph information as an input set and input it into an answer detection model to obtain an encoded vector of the input set;

[0008] Determine an answer detection result corresponding to the question to be queried in the document to be processed according to the encoded vector and output it.

[0009] Optionally, the answer detection model includes a vector encoding module and a probability prediction module;

[0010] Correspondingly, inputting the to-be-processed document, the to-be-query question, and the glyph information as an input set into the answer detection model to obtain an encoded vector of the input set includes:

[0011] Inputting the to-be-processed document, the to-be-query question, and the glyph information as an input set into the vector encoding module for encoding processing to generate an encoded vector of the input set.

[0012] Optionally, inputting the to-be-processed document, the to-be-query question, and the glyph information as an input set into the vector encoding module for encoding processing to generate an encoded vector of the input set includes:

[0013] Inputting the to-be-processed document, the to-be-query question, and the glyph information as an input set into the vector encoding module;

[0014] Wherein, the vector encoding module performs encoding processing on the to-be-processed document and the to-be-query question to generate a first encoded sub-vector of the input set, performs encoding processing on the glyph information to generate a second encoded sub-vector of the input set, and sums the first encoded sub-vector and the second encoded sub-vector to generate an encoded vector of the input set.

[0015] Optionally, inputting the to-be-processed document, the to-be-query question, and the glyph information as an input set into the vector encoding module in the answer detection model for encoding processing to generate an encoded vector of the input set includes:

[0016] Inputting the to-be-processed document and the to-be-query question into the vector encoding module for encoding processing to generate a word vector and a segmentation vector corresponding to each word unit in the to-be-processed document and the to-be-query question, and summing the word vector and the segmentation vector to generate a first encoded sub-vector; and,

[0017] Inputting the glyph information into the vector encoding module for encoding processing to generate a second encoded sub-vector of the to-be-processed document and the to-be-query question;

[0018] Summing the first encoded sub-vector and the second encoded sub-vector to generate an encoded vector of the input set.

[0019] Optionally, determining and outputting an answer detection result corresponding to the to-be-query question in the to-be-processed document according to the encoded vector includes:

[0020] Input the encoded vector into the probability prediction module to obtain the probability prediction results corresponding to each word unit in the input set;

[0021] Determine the answer detection result corresponding to the query problem in the document to be processed according to the probability prediction results and output it.

[0022] Optionally, the determining the answer detection result corresponding to the query problem in the document to be processed according to the probability prediction results and outputting it includes:

[0023] Take the position of the word unit with the highest probability in the probability distribution of the start position in the probability prediction results in the document to be processed as the start position of the answer detection result;

[0024] Take the position of the word unit with the highest probability in the probability distribution of the end position in the probability prediction results in the document to be processed as the end position of the answer detection result;

[0025] Take the word units between the start position and the end position as the answer detection result and output it.

[0026] Optionally, the obtaining the glyph information corresponding to each word unit in the document to be processed and the query problem includes:

[0027] Query the glyph information corresponding to each word unit in the document to be processed and the query problem in the glyph information query library of the document.

[0028] Optionally, the glyph information includes at least one of font, font size, character color, background color, and position coordinates in the document to be processed.

[0029] Optionally, the inputting the document to be processed, the query problem, and the glyph information into the answer detection model as an input set includes:

[0030] Concatenate the document to be processed, the query problem, and the glyph information, and input the concatenated result into the answer detection model as an input set, where in the concatenated result, the target word units in the document to be processed and the query problem correspond to the glyph information of the target word units, and the target word unit is any word unit in the document to be processed and the query problem.

[0031] Optionally, the answer detection model includes a vector encoding module;

[0032] Correspondingly, the inputting the concatenated result into the answer detection model as an input set includes:

[0033] Input the concatenated result into the vector encoding module for encoding processing to generate the encoded vector of the input set.

[0034] According to a second aspect of the embodiments of the present application, an answer detection device is provided, including:

[0035] An acquisition module, configured to acquire glyph information corresponding to each word unit in a document to be processed and a question to be queried;

[0036] An input module, configured to input the document to be processed, the question to be queried, and the glyph information as an input set into an answer detection model to obtain an encoded vector of the input set;

[0037] A determination module, configured to determine an answer detection result corresponding to the question to be queried in the document to be processed according to the encoded vector and output the result.

[0038] According to a third aspect of the embodiments of the present application, a computing device is provided, including a memory, a processor, and computer instructions stored in the memory and executable on the processor. When the processor executes the instructions, the steps of the answer detection method are implemented.

[0039] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores computer instructions. When the instructions are executed by a processor, the steps of the answer detection method are implemented.

[0040] According to a fifth aspect of the embodiments of the present application, a chip is provided, which stores computer instructions. When the instructions are executed by the chip, the steps of the answer detection method are implemented.

[0041] In the embodiments of the present application, by acquiring glyph information corresponding to each word unit in a document to be processed and a question to be queried, inputting the document to be processed, the question to be queried, and the glyph information as an input set into an answer detection model to obtain an encoded vector of the input set, and determining an answer detection result corresponding to the question to be queried in the document to be processed according to the encoded vector and outputting the result.

[0042] Based on the existing document to be processed and question to be queried, the embodiments of the present application incorporate glyph features for documents (rich texts such as word, pdf, web pages, etc.), encode word-related features into tensors, and integrate them into the model. Compared with the existing technical means, more features can be used for answer detection, which is beneficial to improving the accuracy of the answer detection result obtained by the answer detection model when detecting answers in the document to be processed. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a structural block diagram of a computing device provided by the embodiments of the present application;

[0044] Figure 2 is a flowchart of the answer detection method provided by the embodiments of the present application;

[0045] Figure 3 It is a schematic diagram of an answer detection process provided by an embodiment of the present application;

[0046] Figure 4 It is a schematic diagram of another answer detection process provided by an embodiment of the present application;

[0047] Figure 5 It is a schematic diagram of an answer detection method applied to a reading comprehension scenario provided by an embodiment of the present application;

[0048] Figure 6 It is a schematic structural diagram of an answer detection device provided by an embodiment of the present application. Detailed implementation manners

[0049] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.

[0050] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the", and "said" used in one or more embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more of the associated listed items.

[0051] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "in response to determining".

[0052] First, the noun terms related to one or more embodiments of the present invention are explained.

[0053] information extraction: Information extraction.

[0054] BERT: Bidirectional Encoder Representations from Transformers, which are bidirectional encoding representations based on Transformers.

[0055] character feature: Relevant features of words in a document, including but not limited to font, font size, word color, background color, position coordinates of words, etc.

[0056] embedding: Encoding that can encode numbers into multi-dimensional tensors for input into neural network learning.

[0057] token: Before any actual processing of the input text, it needs to be segmented into language units such as words, punctuation marks, numbers, or letters. These units are called tokens. For English text, a token can be a word, a punctuation mark, a number, etc. For Chinese text, the smallest token can be a phrase, a character, a punctuation mark, a number, etc.

[0058] In this application, an answer detection method, device, computing device, and computer-readable storage medium are provided, which will be described in detail one by one in the following embodiments.

[0059] Figure 1 FIG. shows a structural block diagram of a computing device 100 according to an embodiment of the present application. The components of the computing device 100 include but are not limited to a memory 110 and a processor 120. The processor 120 is connected to the memory 110 through a bus 130, and a database 150 is used to store data.

[0060] The computing device 100 further includes an access device 140, and the access device 140 enables the computing device 100 to communicate via one or more networks 160. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 140 may include one or more of any type of wired or wireless network interfaces (e.g., Network Interface Card (NIC)), such as IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.

[0061] In an embodiment of the present application, the above components of the computing device 100 and Figure 1 other components not shown in the figure may also be connected to each other, for example, through a bus. It should be understood that Figure 1The block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.

[0062] The computing device 100 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 100 can also be a mobile or stationary server.

[0063] Among them, the processor 120 can execute Figure 2 the steps in the shown answer detection method.

[0064] Currently, the technologies for completing document text extraction using neural networks are mainly language models such as BERT. The vector layer of the BERT model consists of three different vector representations, namely word vectors, segment vectors, and position vectors. Word vectors represent the vector representation of the word, segment vectors are used to distinguish different sentences in a sentence pair, and position vectors enable BERT to learn the order information of each word in the input sentence. The corresponding elements of these three vectors are added together for synthetic representation to obtain the input vector of the BERT encoding layer. The text is encoded into these three vectors respectively, and the label (answer) corresponding to each word is input to the model, and a model for solving specific tasks can be trained.

[0065] However, although BERT-like language models are widely used, they are mainly used to solve problems related to pure text. The documents to be processed in real scenarios have multiple features compared to pure text, and these features can play an important role in solving problems. BERT-like language models do not utilize these features. The present application encodes these features and integrates them into the BERT model so that the model can encode richer content.

[0066] Based on this, the answer detection method provided by the embodiments of the present application obtains the glyph information corresponding to each word unit in the document to be processed and the query question, takes the document to be processed, the query question, and the glyph information as an input set and inputs them into the answer detection model to obtain the encoded vector of the input set, determines the answer detection result corresponding to the query question in the document to be processed according to the encoded vector, and outputs it. The flowchart of the specific answer detection method is as Figure 2 shown.

[0067] Based on the existing document to be processed and the question to be queried, the embodiment of the present application incorporates the glyph features of the document (rich text such as word, pdf, web page, etc.), encodes the features related to the characters into tensors, and integrates them into the model. Compared with the existing technical means, it can utilize more features for answer detection, which is beneficial to improving the accuracy of the answer detection results obtained by the answer detection model when detecting answers in the document to be processed.

[0068] Figure 2 FIG. shows a flowchart of an answer detection method according to an embodiment of the present application, including steps 202 to 206.

[0069] Step 202, obtain the glyph information corresponding to each word unit in the document to be processed and the question to be queried.

[0070] Specifically, the document to be processed includes, but is not limited to, rich text such as word documents, pdf documents, web pages, etc.; since plain text is text without any text decoration, without any bold, underline, italic, graphics, symbols or special characters and special printing formats, relatively speaking, the document to be processed can be understood as a document containing any one, two or more of text decoration, bold, underline, italic, graphics, symbols, special characters, special printing formats; the glyph information is the relevant features of the word units in the document to be processed, including but not limited to the font, font size, character color, background color, position coordinates in the document to be processed, etc.

[0071] Since each word unit included in the document to be processed and each word unit included in the question to be queried correspond to different glyph information, after the embodiment of the present application obtains the document to be processed and the question to be queried, it can also obtain the glyph information corresponding to each word unit in the document to be processed and the glyph information corresponding to each word unit in the question to be queried, and perform answer detection in combination with the glyph information to improve the accuracy of the answer detection results.

[0072] Among them, the content of the glyph information can include the font, font size, character color, background color, position coordinates in the document, etc. of each word unit in the document to be processed, and can also include the font, font size, character color, etc. of each word unit in the question to be queried.

[0073] Since after obtaining the glyph information, the document to be processed, the question to be queried, and the glyph information need to be used as an input set and input into the answer detection model for processing, and before inputting into the answer detection model, the document to be processed, the question to be queried, and the glyph information can be concatenated to construct the input set, which can be specifically implemented in the following ways:

[0074] Concatenate the to-be-processed document, the to-be-query problem, and the glyph information, and use the concatenation result as the input set to input into the answer detection model. Among them, in the concatenation result, the target word units in the to-be-processed document and the to-be-query problem correspond to the glyph information of the target word units, and the target word unit is any word unit in the to-be-processed document and the to-be-query problem.

[0075] Specifically, the method of concatenating up and down can be adopted, that is, the to-be-processed document and the to-be-query problem are on the top, and the glyph information is on the next line of the to-be-processed document and the to-be-query problem, and the glyph information corresponds to its respective word unit to generate the corresponding concatenation result.

[0076] Therefore, if the input set contains two lines, the first line is the to-be-processed document and the to-be-query problem, and the glyph information is placed on the second line. The input set composed of the to-be-processed document, the to-be-query problem, and the glyph information has a linear data structure, that is, there is a one-to-one mutual relationship between the elements in the data structure (each word unit in the to-be-processed document and the to-be-query problem corresponds to its glyph information).

[0077] Specifically, when implementing, to obtain the glyph information corresponding to each word unit in the to-be-processed document and the to-be-query problem, specifically, query the glyph information corresponding to each word unit in the to-be-processed document and the to-be-query problem in the glyph information query library of the document.

[0078] Specifically, for the to-be-processed document and the to-be-query problem, first, the glyph information of each word unit (each character) in the to-be-processed document and the to-be-query problem can be queried from the glyph information query library of the document, and the glyph information, the to-be-processed document, and the to-be-query problem are input into the answer detection model for processing.

[0079] In practical applications, the glyph information query library can be the PDFMiner tool library. PDFMiner is a tool for extracting information from PDF documents. Different from other PDF-related tools, it focuses entirely on obtaining and analyzing text data. PDFMiner allows users to obtain the exact position of the text on the page, as well as other information such as fonts or lines. It includes a PDF converter that can convert PDF files into other text formats (such as HTML). It has an extensible PDF parser that can be used for other purposes besides text analysis. In addition to extracting the text in the to-be-processed document, this tool library can also extract the position coordinates, fonts, font sizes, character colors, background colors, colors, etc. of each word unit in the text in the to-be-processed document, as well as extract data such as tables and pictures in the to-be-processed document.

[0080] When the document to be processed is a PDF document, the PDFMiner can be directly used to obtain the glyph information corresponding to each word unit in the document to be processed; when the document to be processed is a Word document or a web page, the Word document or the web page can be first stored as a PDF document, and then the PDFMiner is used to obtain the glyph information corresponding to each word unit in the document to be processed.

[0081] Step 204: Input the document to be processed, the question to be queried, and the glyph information as an input set into the answer detection model to obtain the encoded vector of the input set.

[0082] Specifically, after obtaining the question to be queried, the document to be processed, and the glyph information, the document to be processed, the question to be queried, and the glyph information can be used as an input set and input into the answer detection model for encoding processing to obtain the encoded vector of the input set, so as to perform answer detection based on the encoded vector and obtain the answer detection result corresponding to the question to be queried in the document to be processed.

[0083] Among them, in the process of constructing the input set based on the document to be processed, the question to be queried, and the glyph information, the document to be processed, the question to be queried, and the glyph information can be concatenated, and specifically, the up-and-down concatenation method can be adopted. As shown in the input layer of Figure 4 that is, the document to be processed and the question to be queried are on the top, and the glyph information (CF represents the glyph information of the word unit. For example, CF 这 represents the glyph information of the word unit "this",) is on the next line of the document to be processed and the question to be queried, and each word unit in the document to be processed and the question to be queried corresponds to its glyph information one by one to construct the input set.

[0084] Specifically in implementation, the answer detection model includes a vector encoding module and a probability prediction module;

[0085] Correspondingly, inputting the document to be processed, the question to be queried, and the glyph information as an input set into the answer detection model to obtain the encoded vector of the input set can specifically input the document to be processed, the question to be queried, and the glyph information as an input set into the vector encoding module for encoding processing to generate the encoded vector of the input set.

[0086] Further, inputting the document to be processed, the question to be queried, and the glyph information as an input set into the vector encoding module for encoding processing to generate the encoded vector of the input set can be specifically implemented in the following way:

[0087] Input the document to be processed, the question to be queried, and the glyph information as an input set into the vector encoding module;

[0088] Among them, the vector encoding module encodes the document to be processed and the query question to generate a first encoded sub-vector of the input set, encodes the glyph information to generate a second encoded sub-vector of the input set, and sums the first encoded sub-vector and the second encoded sub-vector to generate an encoded vector of the input set.

[0089] Specifically, the answer detection model described in the embodiments of the present application includes a vector encoding module and a probability prediction module. The vector encoding module can be specifically implemented by a BERT model, and the vector encoding module can be used to encode the query question, the document to be processed, and the glyph information.

[0090] Specifically, the vector encoding module can encode the document to be processed and the query question to generate a first encoded sub-vector, and the vector encoding module can encode the glyph information to generate a second encoded sub-vector, and perform a summation operation on the first encoded sub-vector and the second encoded sub-vector to obtain an encoded vector of the input set, so as to comprehensively perform answer detection according to the first encoded sub-vector and the second encoded sub-vector, thereby improving the accuracy of the answer detection result.

[0091] In addition, taking the document to be processed, the query question, and the glyph information as an input set and inputting them into the vector encoding module in the answer detection model for encoding processing to generate an encoded vector of the input set can be specifically implemented in the following manner:

[0092] Input the document to be processed and the query question into the vector encoding module for encoding processing to generate a word vector and a segmentation vector corresponding to each word unit in the document to be processed and the query question, sum the word vector and the segmentation vector to generate a first encoded sub-vector; and,

[0093] Input the glyph information into the vector encoding module for encoding processing to generate a second encoded sub-vector of the document to be processed and the query question;

[0094] Sum the first encoded sub-vector and the second encoded sub-vector to generate an encoded vector of the input set.

[0095] Specifically, the word vector is a numerical vector representation corresponding to each word unit in the document to be processed; the segmentation vector is the sentence vector to which each word unit in the document to be processed belongs.

[0096] When inputting the query question and the document to be processed into the answer detection model, the following format can be adopted: [[CLS], query question, [SEP], document to be processed, [SEP]].

[0097] As described above, the vector encoding module can be implemented by a BERT model. The schematic diagram of the architecture of the BERT model is as Figure 3 shown, including an embedding layer and an encoder. The encoder contains n encoding layers, and these n encoding layers are connected in sequence. In practical applications, the number of encoding layers is determined according to actual requirements and is not limited here.

[0098] Specifically, the query problem and the document to be processed are used as an input set and input into the BERT model. The embedding layer of the BERT model performs word segmentation on the input set to obtain the word units of the input set, and performs pre-embedding processing on the word units to obtain the character vectors and segmentation vectors corresponding to the word units. Then, the character vectors and segmentation vectors are added together, that is, the corresponding positions of the character vectors and segmentation vectors are added to generate the first text sub-vector corresponding to the word units of the input set. The first text sub-vector is input into the first encoding layer in the encoder, and the output vector of the first encoding layer is input into the second encoding layer... and so on. Finally, the output vector of the last encoding layer is obtained, and the output vector of the last encoding layer is used as the first encoded sub-vector of the query problem and the document to be processed.

[0099] In addition, since the glyph information contains the position coordinate information of each word unit in the text obtained by PDFMiner extracting text from the document to be processed in the text, during the process of encoding the glyph information, the position coordinate information can be extracted from the glyph information, and pre-embedding processing is performed on the position coordinate information to obtain the coordinate vectors corresponding to each word unit, and pre-embedding processing is performed on other information in the glyph information except the position coordinate information to obtain the glyph vectors corresponding to each word unit. Then, the coordinate vectors and glyph vectors are added together, that is, the corresponding positions of the coordinate vectors and glyph vectors are added to generate the second text sub-vector corresponding to the word units of the input set. The second text sub-vector is input into the first encoding layer in the encoder, and the output vector of the first encoding layer is input into the second encoding layer... and so on. Finally, the output vector of the last encoding layer is obtained, and the output vector of the last encoding layer is used as the second encoded sub-vector of the query problem and the document to be processed.

[0100] After generating the first encoded sub-vector and the second encoded sub-vector, a summation operation is performed on the first encoded sub-vector and the second encoded sub-vector to generate the encoded vector of the document to be processed and the query problem.

[0101] Alternatively, in the case of concatenating the document to be processed, the question to be queried, and the glyph information, and using the concatenated result as the input set to input into the answer detection model, specifically, the concatenated result can be input into the vector encoding module of the answer detection model, and the vector encoding module encodes the input set to generate the encoding vector of the input set.

[0102] Since the glyph information is a unique attribute of the document to be processed, it can effectively distinguish different modules of the document and can reflect information such as the coordinate information of word units, the font color of word units, and the background color in different modules. Combining the glyph information for answer detection is beneficial to improving the accuracy of the answer detection result. Therefore, in the embodiments of the present application, on the basis of the BERT model, the glyph information for documents (rich texts such as word, pdf, web pages, etc.) is incorporated. At the same time, the coordinate vector of the word unit in the document is used to replace the position vector of the word unit in the sentence, so that the information or features used in the answer detection process are more diverse, thereby improving the accuracy of the answer detection result obtained by the answer detection model when detecting answers in the document to be processed.

[0103] Step 206, determine the answer detection result corresponding to the question to be queried in the document to be processed according to the encoding vector and output it.

[0104] Specifically, after obtaining the encoding vectors of the document to be processed and the question to be queried, the answer detection can be performed according to the encoding vectors to obtain the answer detection result of the question to be queried in the document to be processed.

[0105] When specifically implemented, determining the answer detection result corresponding to the question to be queried in the document to be processed according to the encoding vector and outputting it can be specifically implemented in the following manner:

[0106] Input the encoding vector into the probability prediction module to obtain the probability prediction result corresponding to each word unit in the input set;

[0107] Determine the answer detection result corresponding to the question to be queried in the document to be processed according to the probability prediction result and output it.

[0108] Further, determining the answer detection result corresponding to the question to be queried in the document to be processed according to the probability prediction result and outputting it can be specifically implemented in the following manner:

[0109] Take the position of the word unit with the highest probability in the probability distribution of the start position in the probability prediction result as the start position of the answer detection result in the document to be processed;

[0110] Take the position of the word unit with the highest probability in the probability distribution of the end position in the probability prediction result as the end position of the answer detection result in the document to be processed;

[0111] Take the word units between the starting position and the ending position as the answer detection result and output it.

[0112] Specifically, as mentioned above, the answer detection model includes a probability prediction module. Therefore, the encoded vector can be input into the probability prediction module, and the probability prediction module can predict the probability of each word unit in the document to be processed as the answer detection result according to the encoded vector. Moreover, the probability prediction module can specifically be used to predict the probability of each word unit in the input set as the starting position of the predicted answer, the probability of each word unit in the input set as the ending position of the predicted answer, or the probability of each word unit in the input set as the predicted answer.

[0113] In practical applications, the probability prediction module can be implemented by an LSTM model. Specifically, the encoded vector can be processed by the fully connected layer of the LSTM model, and the processing result can be normalized to generate the probability that each word unit in the input set is the starting position of the answer to the question to be queried, the ending position of the answer to the question to be queried, and / or the probability of being the answer to the question to be queried. The schematic diagram of the specific implementation process is as Figure 3 shown.

[0114] Figure 3 The answer detection model in [reference] includes a vector encoding module and a probability prediction module. After the document to be processed, the question to be queried, and the glyph information are input into the answer detection model as the input set, the embedding layer in the vector encoding module of the answer detection model performs embedding processing on them, and then the n encoders in the vector encoding module perform encoding processing on the embedding processing result, that is, process and generate the word vector, segmentation vector, coordinate vector, and glyph vector corresponding to the input set. Then, sum the word vector, segmentation vector, coordinate vector, and the glyph vector, and input the summation result into the probability prediction module of the answer detection model. The probability prediction module predicts the probability of each word unit in the document to be processed as the answer to the question to be queried to generate the corresponding probability prediction result (label). Then, the answer detection result corresponding to the question to be queried can be determined according to the probability prediction result.

[0115] Figure 4 If the text information contained in the document to be processed in [reference] is "This is an example", then the glyph information of the document to be processed can be obtained from the PDFMiner tool library. CF represents the glyph information of the word unit, such as CF 这That is, it represents the glyph information of the word unit "this", and then inputs the subsequent "This is an example" and the glyph information into the input layer of the answer detection model. The encoding layer of the answer detection model encodes "This is an example" to generate corresponding word vectors and segmentation vectors. The encoding layer of the answer detection model encodes the position coordinate information of each word unit in the glyph information in the document to be processed to generate corresponding coordinate vectors; and the encoding layer of the answer detection model encodes the other information in the glyph information except the position coordinate information to generate corresponding glyph vectors. Then, the word vectors, segmentation vectors, coordinate vectors, and the glyph vectors are summed, and the probability prediction result (label) corresponding to each word in the document to be processed is determined according to the summation result, so as to determine the answer detection result corresponding to the query question according to the probability prediction result.

[0116] Based on the BERT model, the embodiment of the present application incorporates the glyph information of the document to be processed (rich text such as word, pdf, web page, etc.), encodes the glyph information related to the word unit in the document to be processed into glyph vectors and coordinate vectors, and incorporates them into the answer detection process; at the same time, uses the coordinate vectors of the word unit in the document to be processed to replace the position vectors of the word unit in the sentence, and by combining the glyph vectors, realizes answer detection using more dimensional features or information, which is beneficial to improving the accuracy of the answer detection result obtained by the answer detection model in the document to be processed.

[0117] Figure 5 The figure shows a schematic diagram of an answer detection method according to an embodiment of the present application applied to a reading comprehension scenario, including steps 502 to 520.

[0118] Step 502, obtain the question and the document.

[0119] Step 504, query the glyph information corresponding to each word unit in the document and the question in the glyph information query library.

[0120] Step 506, input the question and the document into the vector encoding module of the answer detection model for encoding processing, and generate word vectors, segmentation vectors, and coordinate vectors corresponding to each word unit in the question and the document.

[0121] Step 508, sum the word vectors, segmentation vectors, and coordinate vectors to generate a first encoded sub-vector.

[0122] Step 510, input the glyph information into the vector encoding module of the answer detection model for encoding processing, and generate a second encoded sub-vector of the question and the document.

[0123] Step 512, sum the first encoded sub-vector and the second encoded sub-vector to generate an encoded vector of the question and the document.

[0124] Step 514: Input the encoded vector into the probability prediction module of the answer detection model to obtain the probability prediction results corresponding to each word unit in the document.

[0125] Step 516: Use the position of the word unit with the highest probability in the probability distribution of the start position in the probability prediction results as the start position of the answer detection result in the to-be-processed document.

[0126] Step 518: Use the position of the word unit with the highest probability in the probability distribution of the end position in the probability prediction results as the end position of the answer detection result in the to-be-processed document.

[0127] Step 520: Use the word units between the start position and the end position as the answer detection result and output it.

[0128] Since the glyph information is a unique attribute of the document, it can effectively distinguish different modules of the document and can reflect information such as the coordinate information of word units, the font color of word units, and the background color in different modules. Combining the glyph information for answer detection is beneficial to improving the accuracy of the answer detection result. Therefore, in the embodiments of the present application, on the basis of the BERT model, the glyph information for documents (rich texts such as word, pdf, web pages, etc.) is incorporated. At the same time, the coordinate vector of the word unit in the document is used to replace the position vector of the word unit in the sentence, so that the information or features used in the answer detection process are more diversified, thereby achieving the purpose of improving the accuracy of the answer detection result obtained by the answer detection model in the document for answer detection.

[0129] Corresponding to the above method embodiments, the present application also provides embodiments of an answer detection device. Figure 6 The structural schematic diagram of the answer detection device according to an embodiment of the present application is shown. As Figure 6 shown, the device 600 includes:

[0130] An acquisition module 602, configured to acquire the glyph information corresponding to each word unit in the to-be-processed document and the to-be-query question;

[0131] An input module 604, configured to input the to-be-processed document, the to-be-query question, and the glyph information as an input set into the answer detection model to obtain the encoded vector of the input set;

[0132] A determination module 606, configured to determine and output the answer detection result corresponding to the to-be-query question in the to-be-processed document according to the encoded vector.

[0133] Optionally, the answer detection model includes a vector encoding module and a probability prediction module;

[0134] Correspondingly, the input module 604 includes:

[0135] A first generation sub-module configured to input the document to be processed, the question to be queried, and the glyph information as an input set into the vector encoding module for encoding processing to generate an encoding vector of the input set.

[0136] Optionally, the first generation sub-module includes:

[0137] A first input unit configured to input the document to be processed, the question to be queried, and the glyph information as an input set into the vector encoding module;

[0138] Wherein, the vector encoding module encodes the document to be processed and the question to be queried to generate a first encoded sub-vector of the input set, encodes the glyph information to generate a second encoded sub-vector of the input set, and sums the first encoded sub-vector and the second encoded sub-vector to generate an encoding vector of the input set.

[0139] Optionally, the first generation sub-module includes:

[0140] A second input unit configured to input the document to be processed and the question to be queried into the vector encoding module for encoding processing to generate a word vector and a segmentation vector corresponding to each word unit in the document to be processed and the question to be queried, and sum the word vector and the segmentation vector to generate a first encoded sub-vector;

[0141] A third input unit configured to input the glyph information into the vector encoding module for encoding processing to generate a second encoded sub-vector of the document to be processed and the question to be queried;

[0142] A generation unit configured to sum the first encoded sub-vector and the second encoded sub-vector to generate an encoding vector of the input set.

[0143] Optionally, the determination module 606 includes:

[0144] An acquisition sub-module configured to input the encoding vector into the probability prediction module to obtain a probability prediction result corresponding to each word unit in the input set;

[0145] An output sub-module configured to determine an answer detection result corresponding to the question to be queried in the document to be processed according to the probability prediction result and output it.

[0146] Optionally, the output sub-module includes:

[0147] The first processing unit is configured to use the position of the word unit with the highest probability in the probability distribution of the start position in the probability prediction result as the start position of the answer detection result in the to-be-processed document;

[0148] The second processing unit is configured to use the position of the word unit with the highest probability in the probability distribution of the end position in the probability prediction result as the end position of the answer detection result in the to-be-processed document;

[0149] The output unit is configured to use the word units between the start position and the end position as the answer detection result and output it.

[0150] Optionally, the obtaining module 602 includes:

[0151] The query sub-module is configured to query the glyph information corresponding to each word unit in the to-be-processed document and the to-be-query question in the glyph information query library of the document.

[0152] Optionally, the glyph information includes at least one of font, font size, font color, background color, and position coordinates in the to-be-processed document.

[0153] Optionally, the input module 604 includes:

[0154] The second generation sub-module is configured to splice the to-be-processed document, the to-be-query question, and the glyph information, and use the splicing result as an input set to input into the answer detection model. In the splicing result, the target word units in the to-be-processed document and the to-be-query question correspond to the glyph information of the target word units, and the target word unit is any word unit in the to-be-processed document and the to-be-query question.

[0155] Optionally, the answer detection model includes a vector encoding module;

[0156] Correspondingly, the second generation sub-module is further configured to:

[0157] Input the splicing result into the vector encoding module for encoding processing to generate an encoded vector of the input set.

[0158] Based on the BERT model, the embodiment of the present application incorporates the glyph information of the to-be-processed document (rich text such as word, pdf, web page, etc.), encodes the glyph information related to the word units in the to-be-processed document into glyph vectors and coordinate vectors, and integrates them into the answer detection process; at the same time, uses the coordinate vector of the word unit in the to-be-processed document to replace the position vector of the word unit in the sentence, and by combining the glyph vectors, realizes answer detection using more dimensional features or information, which is beneficial to improving the accuracy of the answer detection result obtained by the answer detection model in the to-be-processed document.

[0159] The above is a schematic solution of an answer detection device according to this embodiment. It should be noted that the technical solution of this answer detection device and the technical solution of the above answer detection method belong to the same concept. For the details not described in detail in the technical solution of the answer detection device, reference can be made to the description of the technical solution of the above answer detection method.

[0160] It should be noted that each component in the apparatus claim should be understood as a functional module that must be established to implement each step of the program flow or each step of the method. Each functional module is not an actual functional division or separation limitation. The apparatus claim defined by such a set of functional modules should be understood as mainly implementing the functional module architecture of the solution through the computer program recorded in the specification, rather than understanding it as an entity apparatus mainly implemented by hardware.

[0161] An embodiment of the present application further provides a computing device, including a memory, a processor, and computer instructions stored in the memory and executable on the processor. When the processor executes the instructions, the steps of the above-mentioned answer detection method are implemented.

[0162] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above answer detection method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above answer detection method.

[0163] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions. When the instructions are executed by a processor, the steps of the answer detection method as described above are implemented.

[0164] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above answer detection method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above answer detection method.

[0165] An embodiment of the present application discloses a chip, which stores computer instructions. When the instructions are executed by a processor, the steps of the answer detection method as described above are implemented.

[0166] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0167] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0168] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0169] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0170] The preferred embodiments of the present application disclosed above are only used to help illustrate the present application. The alternative embodiments do not elaborate on all the details and do not limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the present application. The present application selects and specifically describes these embodiments to better explain the principles and practical applications of the present application, so that those skilled in the art can understand and utilize the present application well. The present application is only limited by the claims and their full scope and equivalents.

Claims

1. A method for answer detection, characterized in that, Including: Obtain the glyph information corresponding to each word unit in the document to be processed and the question to be queried. The glyph information includes at least one of font, font size, character color, background color, and position coordinates in the document to be processed. Among them, each word unit included in the document to be processed and each word unit included in the question to be queried correspond to different glyph information; Concatenate the document to be processed, the question to be queried, and the glyph information, and use the concatenation result as an input set to input into an answer detection model to obtain the encoded vector of the input set. Among them, in the concatenation result, the target word unit in the document to be processed and the question to be queried corresponds to the glyph information of the target word unit, and the target word unit is any word unit in the document to be processed and the question to be queried; Determine the answer detection result corresponding to the question to be queried in the document to be processed according to the encoded vector and output it.

2. The answer detection method according to claim 1, wherein The answer detection model includes a vector encoding module and a probability prediction module; Correspondingly, the step of using the document to be processed, the question to be queried, and the glyph information as an input set to input into the answer detection model to obtain the encoded vector of the input set includes: Use the document to be processed, the question to be queried, and the glyph information as an input set to input into the vector encoding module for encoding processing to generate the encoded vector of the input set.

3. The answer detection method according to claim 2, wherein The step of using the document to be processed, the question to be queried, and the glyph information as an input set to input into the vector encoding module for encoding processing to generate the encoded vector of the input set includes: Use the document to be processed, the question to be queried, and the glyph information as an input set to input into the vector encoding module; Among them, the vector encoding module encodes the document to be processed and the question to be queried to generate the first encoded sub-vector of the input set, encodes the glyph information to generate the second encoded sub-vector of the input set, and sums the first encoded sub-vector and the second encoded sub-vector to generate the encoded vector of the input set.

4. The answer detection method according to claim 2, wherein The step of using the document to be processed, the question to be queried, and the glyph information as an input set to input into the vector encoding module in the answer detection model for encoding processing to generate the encoded vector of the input set includes: Input the document to be processed and the question to be queried into the vector encoding module for encoding processing to generate the word vector and segmentation vector corresponding to each word unit in the document to be processed and the question to be queried, and sum the word vector and the segmentation vector to generate the first encoded sub-vector. Among them, the word vector is the numerical vector representation corresponding to each word unit in the document to be processed, and the segmentation vector is the sentence vector to which each word unit in the document to be processed belongs; and, Input the glyph information into the vector encoding module for encoding processing to generate the second encoded sub-vector of the document to be processed and the question to be queried; Sum the first encoded sub-vector and the second encoded sub-vector to generate the encoded vector of the input set.

5. The answer detection method according to claim 2, wherein Determining and outputting the answer detection result corresponding to the query problem in the to-be-processed document according to the encoding vector includes: Inputting the encoding vector into the probability prediction module to obtain the probability prediction result corresponding to each word unit in the input set; Determining and outputting the answer detection result corresponding to the query problem in the to-be-processed document according to the probability prediction result.

6. The answer detection method according to claim 5, characterized in that, Determining and outputting the answer detection result corresponding to the query problem in the to-be-processed document according to the probability prediction result includes: Taking the position of the word unit with the highest probability in the probability distribution of the start position in the probability prediction result in the to-be-processed document as the start position of the answer detection result; Taking the position of the word unit with the highest probability in the probability distribution of the end position in the probability prediction result in the to-be-processed document as the end position of the answer detection result; Taking the word units between the start position and the end position as the answer detection result and outputting it.

7. The answer detection method according to claim 1, characterized in that Obtaining the glyph information corresponding to each word unit in the to-be-processed document and the query problem includes: Querying the glyph information query library of the document for the glyph information corresponding to each word unit in the to-be-processed document and the query problem.

8. The answer detection method according to claim 1, wherein, The answer detection model includes a vector encoding module; Correspondingly, inputting the splicing result as an input set into the answer detection model includes: Inputting the splicing result into the vector encoding module for encoding processing to generate the encoding vector of the input set.

9. An answer detection device, characterized in that, Including: An acquisition module configured to acquire the glyph information corresponding to each word unit in the to-be-processed document and the query problem, where the glyph information includes at least one of font, font size, font color, background color, and position coordinates in the to-be-processed document, and each word unit included in the to-be-processed document and each word unit included in the query problem correspond to different glyph information; An input module configured to splice the to-be-processed document, the query problem, and the glyph information, and input the splicing result as an input set into the answer detection model to obtain the encoding vector of the input set, where in the splicing result, the target word unit in the to-be-processed document and the query problem corresponds to the glyph information of the target word unit, and the target word unit is any word unit in the to-be-processed document and the query problem; A determination module configured to determine and output the answer detection result corresponding to the query problem in the to-be-processed document according to the encoding vector.

10. A computing device, comprising a memory, a processor, and computer instructions stored on the memory and executable on the processor, characterized in that, When the processor executes the instructions, the steps of the method according to any one of claims 1-8 are implemented.

11. A computer-readable storage medium storing computer instructions, characterized in that, When the instructions are executed by the processor, the steps of the method according to any one of claims 1-8 are implemented.

Citation Information

Patent Citations

  • Answer detection method and device

    CN112328777A