Question and answer data extraction method and device

By combining text recognition and layout analysis, and utilizing deep learning models to extract mathematical question-and-answer data from complex documents, this method solves the problems of low efficiency and poor accuracy of traditional methods, and achieves efficient and accurate question-and-answer data extraction.

CN120892522APending Publication Date: 2025-11-04IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510974750.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and accurately extract high-quality mathematical question-and-answer data from complex documents. Traditional methods are inefficient, costly, or subject to copyright issues, and they also struggle to handle complex formulas and charts.

Method used

By combining text recognition and layout analysis, a question-and-answer data extraction method is constructed by using a deep learning model to identify the type and location of text blocks, extracting questions and answers by combining semantic content, and using layout information to help distinguish text block types and establish logical relationships.

Benefits of technology

It improves the accuracy and efficiency of question-and-answer data extraction, and can handle documents containing complex formulas and charts, thus solving the shortcomings of traditional methods in terms of recognition accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892522A_ABST
    Figure CN120892522A_ABST
Patent Text Reader

Abstract

The invention provides a question and answer data extraction method and device, and the method comprises the steps: carrying out the text recognition of a data source image, and obtaining a to-be-extracted text; the data source image is subjected to layout analysis, layout information is determined, and the layout information is used for representing physical attributes of all text blocks in the to-be-extracted text and spatial arrangement characteristics of all the text blocks in the to-be-extracted text; and extracting a question text and an answer text matched with the question text from the to-be-extracted text based on the semantic content of the to-be-extracted text and the layout information. According to the method, the structure and semantics of the text can be more accurately understood, documents containing complex formulas and charts are effectively processed, the accuracy and efficiency of question and answer data extraction are improved, and the problems of low recognition precision and low extraction efficiency when complex documents are processed by a traditional method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a question-and-answer data extraction method and apparatus. Background Technology

[0002] With the rapid development of artificial intelligence, especially the breakthroughs in large-scale language models across various fields, it has become possible to build mathematical question-answering models using AI. Training high-performance mathematical question-answering models requires a large amount of high-quality mathematical question-answering data.

[0003] Currently, acquiring large-scale, accurate, and high-quality mathematical question-answering data faces challenges. Traditional methods for extracting mathematical question-answering data, such as manual editing and crawling online question banks, have limitations. Manual editing is time-consuming and labor-intensive, making it difficult to produce on a large scale; crawling online question banks may encounter copyright issues, and the data formats are inconsistent, resulting in varying quality. Summary of the Invention

[0004] This invention provides a question-and-answer data extraction method and apparatus to address the deficiencies in the prior art.

[0005] This invention provides a question-and-answer data extraction method, comprising the following steps: Perform text recognition on the source image to obtain the text to be extracted; The image from the data source is analyzed to determine the layout information. The layout information is used to characterize the physical properties of each text block in the text to be extracted and the spatial arrangement characteristics of each text block in the text to be extracted. The physical properties of each text block refer to the features that each text block has itself and can be directly measured or observed. The spatial arrangement characteristics of each text block in the text to be extracted refer to the position of each text block in the text to be extracted and its relative positional relationship with other text blocks. Based on the semantic content of the text to be extracted and the layout information, the question text and the answer text matching the question text are extracted from the text to be extracted.

[0006] According to a question-and-answer data extraction method provided by the present invention, the step of extracting question text and answer text matching the question text from the text to be extracted based on the semantic content of the text to be extracted and the layout information includes: Based on the semantic content of the text to be extracted, the question text and multiple candidate answer texts corresponding to the question text are determined from the text to be extracted; Based on the layout information, the answer text that matches the question text is determined from each candidate answer text.

[0007] According to a question-and-answer data extraction method provided by the present invention, the step of determining the question text and multiple candidate answer texts corresponding to the question text from the text to be extracted based on the semantic content of the text to be extracted includes: Based on the question text retrieval formula, the question text is determined by searching the text to be extracted. Based on the position of the question text in the text to be extracted, the retrieval range of the answer text is determined; Based on the answer text retrieval formula, the text to be extracted is retrieved within the retrieval range to determine the multiple candidate answer texts.

[0008] According to a question-and-answer data extraction method provided by the present invention, the step of determining the answer text matching the question text from each candidate answer text based on the layout information includes: Based on the layout information, determine the content area to which the question text belongs; Based on the content area to which the question text belongs, the answer text that matches the question text is determined from each candidate answer text.

[0009] According to a question-and-answer data extraction method provided by the present invention, the step of performing layout analysis on the data source image to determine layout information, and extracting question text and answer text matching the question text from the text to be extracted based on the semantic content of the text to be extracted and the layout information, includes: Based on the image from the data source, a prompt text is constructed, which is used to indicate the extraction approach for the question-and-answer data; Based on the data extraction model, the prompt text is applied to obtain the text to be extracted from the data source image, and the layout analysis of the data source image is performed to determine the layout information. Based on the semantic content of the text to be extracted and the layout information, the question text and the answer text matching the question text are extracted from the text to be extracted.

[0010] According to a question-and-answer data extraction method provided by the present invention, the step of performing layout analysis on the data source image to determine the layout information of the text to be extracted includes: Determine the type of each text block in the image from which the data is sourced; The layout information is determined based on the logical order and type of each text block.

[0011] According to a question-and-answer data extraction method provided by the present invention, the step of performing layout analysis on the data source image to determine the layout information of the text to be extracted includes: Using a pre-trained layout analysis model, preliminary layout segmentation is performed on the data source image to obtain multiple text blocks; For each text block, extract visual and textual features; The extracted visual and text features are input into the fine-tuned layout analysis model, which outputs the physical attributes of each text block. The layout analysis model is fine-tuned based on the layout annotation data. The layout information is constructed based on the physical properties of each text block and the spatial arrangement characteristics of each text block.

[0012] According to a question-and-answer data extraction method provided by the present invention, the step of performing layout analysis on the data source image to determine the layout information of the text to be extracted includes: The source image of the data is scanned to obtain multiple text blocks; For each text block, the rules in the layout analysis rule base are applied sequentially to determine whether each text block satisfies any rule; the layout analysis rule base includes layout information analysis rules for different types of text blocks; If each text block satisfies any rule, then the layout information of the corresponding text block is determined according to the rule. If any text block does not meet any rules, the layout information of that text block is determined based on the default rules; Integrate the layout information of all text blocks to construct the layout information of the text to be extracted.

[0013] According to a question-and-answer data extraction method provided by the present invention, the step of performing text recognition on the data source image to obtain the text to be extracted includes: Perform region detection on the data source image to determine the text region and formula region in the data source image; Text recognition is performed on the text region and the formula region respectively to obtain the text to be extracted.

[0014] According to a question-and-answer data extraction method provided by the present invention, the step of performing text recognition on the data source image to obtain the text to be extracted includes: The image from the data source is subjected to text recognition to obtain initial recognized text; Based on the text correction model, the initial identified text is corrected to obtain the text to be extracted.

[0015] The present invention also provides a question-and-answer data extraction device, comprising the following modules: The determining unit is used to perform text recognition on the source image of the data to obtain the text to be extracted; The analysis unit is used to perform layout analysis on the data source image to determine layout information. The layout information is used to characterize the physical attributes of each text block in the text to be extracted and the spatial arrangement characteristics of each text block in the text to be extracted. The physical attributes of each text block refer to the features that each text block itself has and can be directly measured or observed. The spatial arrangement characteristics of each text block in the text to be extracted refer to the position of each text block in the text to be extracted and its relative positional relationship with other text blocks. The extraction unit is used to extract question text and answer text matching the question text from the text to be extracted based on the semantic content of the text to be extracted and the layout information.

[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the question-and-answer data extraction method as described above.

[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the question-and-answer data extraction method as described above.

[0018] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the question-and-answer data extraction method as described above.

[0019] The question-and-answer data extraction method and apparatus provided by this invention, by combining the semantic content and layout information of the text to be extracted, achieves accurate and efficient extraction of questions and answers from complex data source images. Since layout information can provide physical attributes such as the type, position, size, and font of text blocks, as well as the spatial arrangement relationships between text blocks, it can help distinguish different types of text blocks such as questions, answers, and titles, and establish the logical relationships between them. Even in the presence of complex formulas and charts, it can more accurately understand the structure and semantics of the text, effectively process documents containing complex formulas and charts, improve the accuracy and efficiency of question-and-answer data extraction, and solve the problems of low recognition accuracy and low extraction efficiency of traditional methods when processing complex documents. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0021] Figure 1This is one of the flowcharts illustrating the question-and-answer data extraction method provided by the present invention.

[0022] Figure 2 This is a flowchart illustrating an implementation method for step 130 in the question-and-answer data extraction method provided by the present invention.

[0023] Figure 3 This is a flowchart illustrating an implementation method for step 131 of the question-and-answer data extraction method provided by the present invention.

[0024] Figure 4 This is a flowchart illustrating an implementation method for step 132 of the question-and-answer data extraction method provided by the present invention.

[0025] Figure 5 This is a flowchart illustrating an implementation method for step 120 in the question-and-answer data extraction method provided by the present invention.

[0026] Figure 6 This is one of the flowcharts illustrating step 110 of the question-and-answer data extraction method provided by the present invention.

[0027] Figure 7 This is a second flowchart illustrating an implementation method for step 110 in the question-and-answer data extraction method provided by the present invention.

[0028] Figure 8 This is the second flowchart of the question-and-answer data extraction method provided by the present invention.

[0029] Figure 9 This is a schematic diagram of the question-and-answer data extraction device provided by the present invention.

[0030] Figure 10 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0032] Currently, building mathematical question-and-answer models requires a large amount of high-quality question-and-answer data, but acquiring this data faces numerous challenges. Traditional methods, such as manual editing, while ensuring data quality, are inefficient and costly, failing to meet the data volume requirements for training large models. Web crawlers can quickly acquire data, but may face copyright issues, and the quality and format of the crawled data vary greatly, requiring significant effort for cleaning and organization. Furthermore, automatically extracting question-and-answer data from documents containing complex mathematical formulas, charts, and other elements is also difficult. Traditional OCR technology has low recognition accuracy when processing such documents, struggling to accurately distinguish between question and answer areas, and unable to understand the logical relationships between questions and answers. Therefore, there is an urgent need for a technical solution that can efficiently and accurately extract high-quality mathematical question-and-answer data from complex documents.

[0033] To address this issue, the present invention provides a question-and-answer data extraction method. This method can be applied to the extraction of mathematical question-and-answer data, as well as to the extraction of question-and-answer data in other fields, such as legal question-and-answer data extraction, physics, chemistry, and biology question-and-answer data extraction, etc. The embodiments of the present invention do not specifically limit this application. For ease of understanding of the technical solution of the present invention, the following embodiments are all illustrated using the application to question-and-answer data extraction as an example.

[0034] in, Figure 1 This is one of the flowcharts illustrating the question-and-answer data extraction method provided by the present invention, such as... Figure 1 As shown, the method includes steps 110, 120 and 130.

[0035] Step 110: Perform text recognition on the source image to obtain the text to be extracted.

[0036] Here, the data source image refers to an image containing the question-and-answer data to be extracted, which includes text, formulas, charts, etc. The data source image can be a scanned or photographed image from printed documents such as textbooks, exercise books, and exam papers, or it can be a screenshot of an electronic document containing question-and-answer content.

[0037] The text to be extracted refers to the text content that can be used for subsequent steps to extract questions and answers after text recognition. It contains the document content of questions and answers. The text to be extracted can be the text content of the source image of the text to be extracted after text recognition.

[0038] The source image may contain text and formula regions. If traditional OCR methods are used for text recognition, traditional OCR engines, typically designed for general text, may struggle with complex mathematical formulas and symbols. Furthermore, formulas have diverse formatting, making traditional rule-based methods difficult to handle effectively. This can lead to errors or loss of formula regions, inaccurate recognition of mathematical symbols and formula structures, misidentification of formula regions as text, resulting in garbled characters, and failure to effectively distinguish between text and formula regions, leading to inaccurate text recognition results and affecting overall accuracy. To address this, this invention employs a deep learning-based formula recognition method, such as a dedicated formula recognition model (e.g., a Transformer-based model), to identify formula regions and fuse the results with the OCR recognition results for text regions, thereby improving overall recognition accuracy. Alternatively, region detection can be performed on the source image to distinguish between text and formula regions before processing them using different OCR methods.

[0039] Furthermore, considering that the text to be extracted may contain a large amount of redundant information or format noise, which may affect the subsequent extraction accuracy, the text to be extracted can be preprocessed, such as removing watermarks, correcting skew, removing noise, and unifying the encoding format, in order to improve the accuracy of subsequent layout analysis and semantic understanding.

[0040] Step 120: Perform layout analysis on the source image to determine the layout information. The layout information is used to characterize the physical attributes of each text block in the text to be extracted and the spatial arrangement characteristics of each text block in the text to be extracted. The physical attributes of each text block refer to the features that each text block itself has and can be directly measured or observed. The spatial arrangement characteristics of each text block in the text to be extracted refer to the position of each text block in the text to be extracted and its relative positional relationship with other text blocks.

[0041] Specifically, layout analysis refers to the structured analysis of source images, aiming to identify elements such as paragraphs, headings, lists, tables, and images within the source images and determine their logical order and hierarchical structure. Here, logical order refers to the reading order between text blocks, such as reading headings first, then paragraphs, and finally chart descriptions. Hierarchical structure refers to the hierarchical relationships between text blocks, such as the relationship between headings and subheadings, or between chapters and paragraphs. For example, the layout of a book can be analyzed into elements such as cover, table of contents, chapters, paragraphs, and images, and the inclusion relationships between chapter headings and paragraphs, as well as the sequential reading order of each chapter, can be determined.

[0042] Page layout information is used to characterize the physical properties of each text block in the text to be extracted and the spatial arrangement characteristics of each text block in the text to be extracted. The physical properties of each text block refer to the features that the text block itself has that can be directly measured or observed. The physical properties of each text block can include the type of text block (such as heading, paragraph, list item, table cell, etc.), the size of the text block (such as width and height), the font of the text block (such as font name, font size, weight, etc.), and the color of the text block.

[0043] The spatial arrangement features of text blocks in the text to be extracted refer to the position of the text block on the page and its relative positional relationship with other text blocks. Spatial arrangement features may include the coordinates of the text block on the page, the distance between text blocks, the alignment of text blocks, and the connection relationship between text blocks (such as whether they belong to the same paragraph, whether they are adjacent, or whether they are in the same column).

[0044] As an alternative implementation, rule-based methods can be used, such as identifying headings, paragraphs, and lists based on features like font size, position, and alignment of text blocks, to determine the layout information of the text to be extracted.

[0045] As an alternative embodiment, machine learning methods, such as using deep learning models (e.g., LayoutLM, Detectron2), can be used to identify the type and location of text blocks and determine the layout information of the text to be extracted.

[0046] In addition, considering that the layout of the text to be extracted may contain various noises and irregularities, such as text tilting due to poor scanning quality or layout chaos due to document editing errors, the layout of the text to be extracted can be corrected and cleaned before the layout analysis is performed. For example, image processing technology can be used for tilt correction, and rule or machine learning methods can be used for layout error correction, so as to improve the accuracy of layout analysis.

[0047] Step 130: Based on the semantic content of the text to be extracted and the layout information, extract the question text and the answer text that matches the question text from the text to be extracted.

[0048] Specifically, layout information is used to characterize the physical attributes and spatial arrangement features of each text block in the text to be extracted. It can help distinguish the types of text blocks (such as questions, answers, titles, paragraphs, etc.) and determine the relationships between text blocks (such as the correspondence between questions and answers, the hierarchical relationship between titles and body text, etc.). Since the source image contains text and formula regions, when performing text recognition on the source image, formula recognition errors or omissions may lead to incorrect extraction of formula content, or the inability to effectively distinguish between text and formula regions may result in inaccurate text recognition results. Consequently, the extracted text may contain a large amount of erroneous or incomplete information, affecting subsequent question and answer extraction.

[0049] If the question and answer texts are extracted based solely on the semantic content of the text to be extracted, the layout information of the text may not be effectively utilized, resulting in inaccurate extraction of questions and answers. For example, the title or paragraph content may be misidentified as a question or answer, or questions and answers may not be correctly matched.

[0050] It should be noted that layout information is used to characterize the physical attributes of each text block in the text to be extracted and the spatial arrangement characteristics of each text block within the text. Since the physical attributes of each text block provide its visual characteristics (such as font, font size, color, etc.), they can help distinguish the type of text block, such as headings, paragraphs, list items, etc. The spatial arrangement characteristics of each text block within the text to be extracted can infer the reading order and logical structure of the text blocks. For example, adjacent text blocks may belong to the same paragraph, and aligned text blocks may belong to the same list. Therefore, combining the physical attributes of each text block with its spatial arrangement characteristics within the text to be extracted allows for a more comprehensive understanding of the text's layout structure and semantic information, thereby further improving the accuracy and efficiency of question and answer extraction.

[0051] Therefore, embodiments of the present invention utilize layout information to assist semantic understanding. For example, based on information such as the font, font size, and position of text blocks, it determines whether a text block is a title or a question; based on the distance and alignment between text blocks, it determines whether text blocks belong to the same question or answer; and based on the type of text block, it classifies and filters text blocks, thereby improving the accuracy and efficiency of question and answer extraction. For instance, based on the layout feature that question text usually precedes answer text, combined with semantic analysis results, questions and answers can be identified and extracted more accurately.

[0052] In this context, the answer text that matches the question text refers to the text content that can answer or explain the question text. A question text may correspond to one answer text; for example, a multiple-choice question may have only one correct answer. A text may also correspond to multiple answer texts; for example, an open-ended question may have multiple different solutions and answers, or a question may have both a detailed explanation and a concise answer.

[0053] As an optional embodiment, natural language processing techniques (e.g., named entity recognition, keyword extraction, etc.) can be used to identify the question text and multiple candidate answer texts in the text to be extracted. Then, layout information (e.g., the distance between the question and the candidate answer texts, alignment, font, etc.) can be used to match the candidate answer texts. Finally, the best matching result can be selected based on the matching score, and the question text and the answer text that matches the question text can be extracted from the text to be extracted.

[0054] It is understood that each of the above steps can be implemented independently using a Large Language Model (LLM). For example, step 110 can directly use an LLM to perform Visual Question Answering (VQA) on the image; step 120 can directly use an LLM for layout analysis and understanding; and step 130 can directly use an LLM to extract questions and answers without explicit feature engineering and rule design. Alternatively, multiple steps can be implemented simultaneously using a Large Language Model (LLM). For instance, the data source image containing text, formulas, and charts can be directly input into the LLM, and the LLM can be required to directly output the extracted question-and-answer pairs (i.e., question text and answer text). Here, a Large Language Model refers to a deep learning model with a large number of parameters and powerful natural language processing capabilities. It can be a pre-trained Transformer model, such as BERT or LLaMA, or a model specifically trained for question-and-answer data extraction. This embodiment of the invention does not specifically limit this.

[0055] The question-and-answer data extraction method provided in this invention combines the semantic content and layout information of the text to be extracted to accurately and efficiently extract questions and answers from complex data source images. Since layout information provides physical attributes such as the type, position, size, and font of text blocks, as well as the spatial arrangement between text blocks, it can help distinguish different types of text blocks such as questions, answers, and titles, and establish logical relationships between them. Even in the presence of complex formulas and charts, it can more accurately understand the structure and semantics of the text, effectively process documents containing complex formulas and charts, improve the accuracy and efficiency of question-and-answer data extraction, and solve the problems of low recognition accuracy and low extraction efficiency of traditional methods when processing complex documents.

[0056] Based on the above embodiments, Figure 2 This is a flowchart illustrating an implementation method for step 130 of the question-and-answer data extraction method provided by the present invention, as shown below. Figure 2 As shown, step 130, based on the semantic content of the text to be extracted and the layout information, extracts the question text and the answer text matching the question text from the text to be extracted, including: Step 131: Based on the semantic content of the text to be extracted, determine the question text and multiple candidate answer texts corresponding to the question text from the text to be extracted; Step 132: Based on the layout information, determine the answer text that matches the question text from each candidate answer text.

[0057] Specifically, candidate answer text refers to text fragments that may form a correct question-answer pair with the question text. These text fragments are semantically related to the question text but have not yet undergone precise matching and screening. The semantic content of the text to be extracted is used to represent the meaning and information of the text, including keywords, entities, syntactic structures, semantic relationships, etc., so as to identify potential question and answer texts based on the semantic content of the text to be extracted and using natural language processing techniques (such as named entity recognition, keyword extraction, semantic similarity calculation, etc.). For example, keyword extraction algorithms can be used to extract keywords from the text to be extracted, and then text fragments related to the question text can be selected as candidate answer texts based on the semantic similarity of the keywords.

[0058] As an optional implementation, the question text (e.g., text fragments starting with keywords such as "please ask:") can be determined from the text to be extracted using regular expressions or keyword matching. Then, a pre-trained language model (such as BERT, RoBERTa, etc.) is used to encode the text to be extracted. Next, the semantic similarity between text fragments is calculated, and text fragments with high similarity to the question text are used as candidate answer texts. Finally, the question text and multiple candidate answer texts corresponding to the question text are extracted from the text to be extracted.

[0059] Furthermore, considering that multiple candidate answer texts may contain redundant information, noise interference, or have low actual relevance to the question text, it is necessary to filter and sort the candidate answer texts to improve the accuracy of answer matching. For example, the question text is "What is the chemical formula and density of water?", candidate answer text 1 is "Water is composed of two hydrogen atoms and one oxygen atom, and the density of water is 1 g / cm³", and candidate answer text 2 is "The density of water is 1 g / cm³". Although candidate answer text 2 is semantically similar to the question text, it is located in a paragraph of another chapter and is clearly unrelated to the question text in terms of layout. In other words, candidate answer text 2 is not a matching answer text for the question text.

[0060] Because layout information includes physical attributes such as the position, size, font, and alignment of text blocks, as well as the spatial relationships between text blocks, it can help identify text segments among multiple candidate answer texts that are more closely related to the question text in terms of layout, and eliminate text segments that are obviously unrelated in layout, thereby improving the accuracy of answer matching. For example, if the question text and a candidate answer text are in the same column and close to each other, the candidate answer text is more likely to be the correct answer.

[0061] As an optional embodiment, layout features (such as distance, alignment, font differences, etc.) between each candidate answer text and the question text can be calculated. Then, a machine learning model (such as SVM, Random Forest, etc.) can be used to classify the candidate answer texts. The candidate answer texts with a classification result of "match" are taken as the final answer. The answer text that matches the question text is determined from multiple candidate answer texts.

[0062] Based on any of the above embodiments Figure 3 This is a flowchart illustrating an implementation method for step 131 of the question-and-answer data extraction method provided by the present invention, as shown below. Figure 3 As shown, step 131, based on the semantic content of the text to be extracted, determines the question text and multiple candidate answer texts corresponding to the question text from the text to be extracted, including: Step 1311: Based on the question text retrieval formula, search in the text to be extracted to determine the question text; Step 1312: Determine the retrieval scope of the answer text based on the position of the question text in the text to be extracted; Step 1313: Based on the answer text retrieval formula, search the text to be extracted within the retrieval scope to determine multiple candidate answer texts.

[0063] Specifically, a question text retrieval formula refers to the rules or patterns used to find question text within a text. These can be constructed using regular expressions, keyword lists, or machine learning-based models. For example, for multiple-choice questions, a regular expression containing keywords such as "which of the following options" or "which of the following statements is correct" can be used as the question text retrieval formula; for true / false questions, a regular expression containing keywords such as "whether" or "correct / incorrect" can be used.

[0064] Because the question text search query contains the feature information of the question text, when searching in the text to be extracted based on the question text search query, potential question texts can be quickly and accurately located and obtained. For example, the question text search query is "[1-9]+、.*[?]", which can match text fragments that start with a number and end with a question mark, such as "1、What is the chemical formula of water?".

[0065] The search scope for the answer text refers to the area within the text to be extracted from which to find the answer text. Typically, the answer text appears near the question text, such as within a certain range after the question text. After identifying the question text, a search scope can be defined based on its position within the text to be extracted, reducing the search space for the answer text and improving search efficiency. For example, if the question text "What is the chemical formula of water?" is located at the 20th character of the 10th line in the text to be extracted, then the corresponding search scope could be the area from the 21st character of the 10th line to the end of the 20th line. The size of this scope can be adjusted according to the specific situation. For example, for multiple-choice questions, the scope can be limited to one or two lines after the question text; for open-ended questions, the scope can be expanded to the entire paragraph or chapter.

[0066] Furthermore, answer text retrieval formulas refer to rules or patterns used to find answer text within text. These can be constructed using regular expressions, keyword lists, or machine learning-based models. For example, a regular expression containing keywords such as "Answer:" or "Solution:" can be used as an answer text retrieval formula. Alternatively, a pre-trained language model can be used to encode the question and answer texts, then their similarity can be calculated, and text fragments with similarity exceeding a certain threshold can be considered answer texts. Because answer text retrieval formulas contain feature information about the answer text, searching the text to be extracted within the defined retrieval range based on the answer text retrieval formula can quickly and accurately locate potential answer texts and use them as candidate answer texts. For example, a regular expression containing options such as "A., B., C., D." can be used as an answer text retrieval formula to extract the options for multiple-choice questions as candidate answer texts.

[0067] Based on any of the above embodiments Figure 4 This is a flowchart illustrating an implementation method for step 132 of the question-and-answer data extraction method provided by the present invention, as shown below. Figure 4 As shown, step 132, based on the layout information, determines the answer text that matches the question text from each candidate answer text, including: Step 1321: Based on the layout information, determine the content area to which the question text belongs; Step 1322: Based on the content area to which the question text belongs, determine the answer text that matches the question text from each candidate answer text.

[0068] Specifically, the content area to which the question text belongs refers to a specific layout area within the source image where the question text is located. Text blocks within this area typically share similar layout features and are semantically related to the question text. Content areas can include paragraph areas, heading areas, list areas, formula areas, figure areas, table areas, etc. Since layout information includes physical attributes such as the position, size, font, alignment, line spacing, and paragraph spacing of each text block in the source image, as well as the spatial arrangement between text blocks, the content area to which the question text belongs can be determined based on the layout features and spatial relationships of the text blocks.

[0069] As an optional implementation, the content area to which the question text belongs can be determined based on the layout characteristics of the text blocks surrounding the question text. For example, if the question text is adjacent to multiple text blocks with the same indentation and bullet points, it can be determined that the question text belongs to a list area; if the question text is located below a large, bold text block, it can be determined that the question text belongs to the heading area represented by that text block. For example, if the question text is located above or below a table and has a clear alignment with the table's header row or data rows, it can be determined that the question text belongs to a table area.

[0070] After determining the content area to which the question text belongs, considering that the answer text is usually located in the same or related page area as the question text and has a certain spatial proximity to the question text, the answer text that matches the question text can be determined from the candidate answer texts based on the content area to which the question text belongs. For example, if the question text belongs to a list area, candidate answer texts with similar formatting and indentation to the question text can be searched within that list area; if the question text belongs to a table area, candidate answer texts can be searched in the table rows or columns related to the question text.

[0071] As an optional embodiment, the page distance (e.g., Euclidean distance, Manhattan distance, etc.) between each candidate answer text and the question text can be calculated, and the candidate answer text with the smallest page distance can be selected as the answer text that matches the question text. The answer text that matches the question text can be determined from each candidate answer text.

[0072] Based on any of the above embodiments, layout analysis is performed on the data source image to determine layout information. Based on the semantic content of the text to be extracted and the layout information, question text and answer text matching the question text are extracted from the text to be extracted, including: Based on the images from the data source, prompt text is constructed to indicate the extraction approach for the question-and-answer data. Based on the data extraction model, the application of prompt text extracts the text to be extracted from the source image. The layout of the source image is analyzed to determine the layout information. Based on the semantic content of the text to be extracted and the layout information, the question text and the answer text that matches the question text are extracted from the text to be extracted.

[0073] Specifically, prompt text refers to a set of natural language descriptions or instructions used to guide the question-and-answer data extraction task of the data extraction model. It indicates the extraction strategy, which can be understood as how to identify questions and answers from the source images and how to associate them. Prompt text is constructed based on the source images. For example, different prompt texts can be designed and constructed according to the type (e.g., textbooks, exam papers) and content (e.g., topics, chapters) of the source images. For textbook-type source images, prompt text can be constructed to instruct the model to extract definitions, theorems, examples, and other knowledge points, as well as related questions and answers. For exam paper-type source images, prompt text can be constructed to instruct the model to extract different types of questions, such as multiple-choice, fill-in-the-blank, and problem-solving questions, and their corresponding answers.

[0074] The example prompt text is as follows: {Please extract all multiple-choice questions and their options from this exam paper image and return them in JSON format. Exam paper image: …}.

[0075] After constructing the prompt text, it can be input into the data extraction model so that the model can perform question and answer data extraction tasks according to the instructions in the prompt text, such as identifying the types of questions and answers, extracting the content of questions and answers, and establishing the relationship between questions and answers, to obtain question text and answer text.

[0076] The data extraction model can be a deep learning-based multimodal model that can process both image and text inputs and generate text outputs, or it can be a specially trained model for question-and-answer data extraction, such as a Transformer-based model that can generate structured data containing questions and answers based on the input prompt text and the image from the data source.

[0077] Based on any of the above embodiments Figure 5 This is a flowchart illustrating an implementation method for step 120 of the question-and-answer data extraction method provided by the present invention, as shown below. Figure 5 As shown, step 120 performs layout analysis on the data source image to determine the layout information of the text to be extracted, including: Step 121: Determine the type of each text block in the image from which the data is sourced; Step 122: Determine the layout information based on the logical order and type of each text block.

[0078] Specifically, the type of a text block refers to the role or category it plays in the document structure. Text block types can include headings, paragraphs, list items, formulas, images, tables, etc. The type can be determined using rule-based methods, such as font size, weight, and position, to identify the type of each text block in the source image. Alternatively, machine learning methods can be used, such as training a classifier to classify text blocks. Visual features (e.g., image pixel values) and textual features (e.g., words, parts of speech) can be used as input to the classifier to determine the type of each text block in the source image. This embodiment of the invention does not impose specific limitations on this approach.

[0079] The logical order of text blocks refers to the reading order or dependency relationship between text blocks, which is used to characterize the information organization structure between text blocks. For example, if text block 1 is the title of a chapter and text block 2 is a paragraph under the same chapter, then the logical order of text block 1 and text block 2 is that text block 1 comes before text block 2, and text block 2 depends on text block 1.

[0080] Since the logical order of text blocks reflects the organization of the text, and the type of text blocks reflects the function and role of the text blocks, combining the logical order and type of text blocks can help to better understand the structure and content of the document, thereby extracting question and answer data more accurately and obtaining layout information.

[0081] As an alternative implementation, a graph structure can be used to represent the logical order and type between text blocks. For example, each text block can be treated as a node, and the logical relationships between text blocks (such as inclusion, dependency, etc.) can be treated as edges. The type of the edge can represent the type of the text block. Then, graph algorithms (such as depth-first search, breadth-first search, etc.) can be used to traverse the graph structure to determine the relationships between text blocks and obtain the layout information.

[0082] Based on any of the above embodiments, layout analysis is performed on the data source image to determine the layout information of the text to be extracted, including: Using a pre-trained layout analysis model, preliminary layout segmentation is performed on the source image to obtain multiple text blocks; For each text block, extract visual and textual features; The extracted visual and text features are input into the fine-tuned layout analysis model, which outputs the physical properties of each text block. The layout analysis model is then fine-tuned based on the layout annotation data. Page layout information is constructed based on the physical properties of each text block and the spatial arrangement characteristics of each text block.

[0083] Specifically, a pre-trained layout analysis model refers to a deep learning model that has been pre-trained on a large-scale general dataset and possesses general layout analysis capabilities. It is typically trained using a large amount of document image data, combined with self-supervised or supervised learning methods. A pre-trained layout analysis model has the ability to recognize basic elements such as text regions, tables, and images in an image; that is, it can roughly locate the layout structure in an image and extract preliminary visual and textual features. Models such as LayoutLM and Detectron2 can be used as pre-trained models.

[0084] Preliminary page layout segmentation refers to dividing the source image into multiple non-overlapping regions, each containing some text or non-text content. The pre-trained page layout analysis model can use object detection algorithms (such as Faster R-CNN, Mask R-CNN) or semantic segmentation algorithms (such as U-Net) to perform preliminary page layout segmentation on the source image, obtaining multiple text blocks. These multiple text blocks refer to regions in the image that have certain semantic meaning; for example, text blocks can include title regions, paragraph regions, list item regions, table regions, image regions, etc.

[0085] The text block may contain different types of content such as text, formulas, and images. Based on this, visual features and text features are extracted for each text block. Visual features refer to the appearance attributes of the text block in the image, which are used to characterize information such as the shape, size, position, and color of the text block. Text features refer to the attributes of the text inside the text block, which are used to characterize information such as the font, font size, alignment, and line spacing of the text inside the text block.

[0086] As an alternative implementation, a convolutional neural network (CNN) can be used to extract visual features from text blocks, while a recurrent neural network (RNN) or Transformer model can be used to extract text features. Considering that different types of features have different dimensions and scales, after obtaining the visual and text features, these features can be normalized and fused to better input them into the fine-tuned layout analysis model.

[0087] Given that pre-trained layout analysis models are typically trained on large-scale datasets, initial segmentation using these models can remove background noise and highlight important layout elements (such as text and image regions). However, to accurately identify specific layout types (such as questions and answers in an exam paper), pre-trained layout analysis models usually require targeted training or fine-tuning. Training a separate model to determine the physical properties of each text block would likely require a large amount of labeled data and computational resources, and would not fully utilize the existing knowledge of the pre-trained model.

[0088] Based on this, embodiments of the present invention fine-tune a pre-trained layout analysis model using layout annotation data, enabling the layout analysis model to learn the characteristics of specific layouts. The resulting fine-tuned layout analysis model can then more accurately identify various layout types and obtain the physical properties of each text block. In other words, embodiments of the present invention not only effectively utilize the existing knowledge of the pre-trained model but also eliminate the need to train a separate model, thus reducing the cost of model training.

[0089] The physical properties of each text block refer to the features that each text block possesses and can be directly measured or observed. The spatial arrangement features of each text block in the text to be extracted refer to the position of each text block in the text to be extracted and its relative positional relationship with other text blocks. Combining the two can comprehensively and accurately describe the layout structure of the document, providing key information for subsequent question and answer data extraction.

[0090] As an optional implementation, the layout information of the text to be extracted can be constructed based on rules or heuristic algorithms, combined with the physical attributes and spatial arrangement features of each text block. For example, the positional relationship between titles and paragraphs, font size, and other information can be determined based on the physical attributes and spatial arrangement features of each text block, and the hierarchical relationship between titles and paragraphs can be determined based on this information, thereby constructing the chapter structure of the document and obtaining the layout information.

[0091] As an alternative embodiment, a graph neural network (GNN) can be used to treat each text block as a node in a graph. The spatial relationships between text blocks (e.g., adjacent, aligned, contained, etc.) are determined based on the spatial arrangement features of the text blocks, and the spatial relationships between text blocks are treated as edges in the graph. Then, the graph neural network is used to learn the feature representation of each node, and finally, the layout information is constructed based on these feature representations.

[0092] This invention combines a pre-trained model and a fine-tuned model, and makes full use of the physical properties and spatial arrangement features of text blocks to achieve accurate analysis of the layout information of the source image. This provides a reliable structured information foundation for subsequent question-and-answer data extraction, and significantly improves the accuracy and efficiency of question-and-answer data extraction.

[0093] Based on any of the above embodiments, layout analysis is performed on the data source image to determine the layout information of the text to be extracted, including: The source image of the data is scanned to obtain multiple text blocks; For each text block, the rules in the layout analysis rule base are applied sequentially to determine whether each text block satisfies any of the rules; the layout analysis rule base includes layout information analysis rules for different types of text blocks; If each text block satisfies any rule, then the layout information of the corresponding text block is determined according to that rule; If any text block does not meet any rules, the layout information of any text block is determined based on the default rules; Integrate the layout information of all text blocks to construct the layout information of the text to be extracted.

[0094] Specifically, a text block refers to the smallest unit with complete semantic meaning in an image, typically containing a continuous piece of text or a single image element. By scanning the source image, different scanning strategies and image processing techniques can be used to identify text regions, image regions, table regions, etc., and obtain multiple text blocks.

[0095] The system can scan the source image in a preset order, such as from left to right or from top to bottom, line by line to obtain multiple text blocks. Alternatively, a sliding window technique can be used, moving across the image in steps and recognizing the text block within the current window with each movement. Optical Character Recognition (OCR) technology can also be employed to identify each text region in the image and use the recognition result for each region as the corresponding text block.

[0096] A layout analysis rule base can be understood as a predefined set of rules used to describe the characteristics of different types of text blocks. This rule base stores multiple layout information analysis rules, each corresponding to a specific layout type and its characteristics. For example, the layout analysis rule base includes heading rules, paragraph rules, list rules, and table rules. Heading rules describe the font size, weight, and position of heading text; paragraph rules describe the alignment, line spacing, and first-line indentation of paragraph text; list rules describe the starting symbols and indentation formats of list items; and table rules describe table borders and row / column separators.

[0097] This can be achieved by manually analyzing a large number of document images to summarize the characteristics of different layout types, and then transforming these characteristics into executable rules to build a pre-defined layout analysis rule library. For example, it can be observed that headings usually have larger fonts and are bold, so a heading rule can be created that specifies that the font size is greater than or equal to a certain threshold and the font weight is bold.

[0098] For each text block, the rules in the layout analysis rule base are applied sequentially to determine whether each text block satisfies any of the rules. For example, for any text block, its font size, alignment, position, and other features can be extracted and compared with the conditions of each rule in the layout analysis rule base to determine whether it satisfies the rule.

[0099] If any text block satisfies any rule, it indicates that the text block conforms to the layout type characteristics described by that rule. In this case, the layout information of the corresponding text block can be determined according to any rule. For example, if a text block satisfies a heading rule (such as a font size greater than or equal to 20 and located at the beginning of a line), then the text block can be determined to be a heading, and its layout information can be determined as a first-level heading or a second-level heading based on its position and hierarchical relationship.

[0100] If any text block does not meet any rules, it means that the text block does not conform to the layout type characteristics described by any rule in the layout analysis rule base. In this case, the layout information of any text block can be determined based on the default rules. The default rules here refer to pre-defined general rules that apply to all text blocks. For example, all text blocks that do not meet any rules are considered paragraphs. For example, the default rule can be set to "body paragraph" and the font size can be set to the default size.

[0101] In other words, the embodiments of the present invention take into account that the layout analysis rule base may not exhaust all types of text block layout analysis rules, and there may be some unknown or special types of text blocks. Therefore, if any text block does not meet any rules, the default rules are used for processing to avoid omissions, thereby ensuring the completeness of the layout analysis.

[0102] After obtaining the layout information of all text blocks, the layout information of all text blocks is integrated to construct the layout information of the text to be extracted. Specifically, based on the layout information of each text block, the positional and hierarchical relationships of each text block can be determined, and the logical structure of the document, such as chapter structure or list structure, can be constructed based on this information to obtain the layout information of the text to be extracted.

[0103] This invention, by combining a layout analysis rule base with default rules, can effectively identify various types of text blocks and construct accurate layout information, providing a high-quality data foundation for subsequent question-and-answer data extraction, thereby improving the accuracy and efficiency of question-and-answer data extraction.

[0104] Based on any of the above embodiments Figure 6 This is one of the flowcharts illustrating step 110 of the question-and-answer data extraction method provided by the present invention, as shown below. Figure 6 As shown, step 110 performs text recognition on the source image to obtain the text to be extracted, including: Step 111a: Perform region detection on the data source image to determine the text region and formula region in the data source image; Step 112a: Perform text recognition on the text region and the formula region respectively to obtain the text to be extracted.

[0105] Specifically, text regions refer to areas in the source image containing natural language text, characterized by neatly arranged lines, uniform character spacing, and semantic coherence. Formula regions refer to areas in the source image containing mathematical formulas, chemical formulas, and other symbols, characterized by various mathematical symbols, subscripts, superscripts, parentheses, etc., exhibiting complex structures and diverse layouts. Region detection of the source image refers to the process of segmenting the source image into different regions and identifying the type of each region (e.g., text region, formula region, etc.).

[0106] As an alternative implementation, deep learning-based object detection models (such as Faster R-CNN, YOLO, etc.) can be used to determine text and formula regions in the source image.

[0107] Because text regions and formula regions have different characteristics, traditional OCR engines have difficulty processing both types of regions simultaneously, which can easily lead to formula recognition errors or loss. In this embodiment of the invention, after determining the text region and formula region in the data source image, text recognition is performed on the text region and formula region separately. Thus, a dedicated formula recognition engine can be used to process the formula region, improving the accuracy of formula recognition. At the same time, a traditional OCR engine can be used to process the text region, improving the efficiency and accuracy of text recognition.

[0108] Based on any of the above embodiments Figure 7 This is a second flowchart illustrating an implementation method for step 110 of the question-and-answer data extraction method provided by the present invention, as shown below. Figure 7 As shown, step 110 performs text recognition on the source image to obtain the text to be extracted, including: Step 111b: Perform text recognition on the source image to obtain the initial recognized text; Step 112b: Based on the text correction model, perform text correction on the initial identified text to obtain the text to be extracted.

[0109] Specifically, the initial recognized text refers to the text content directly identified from the data source image without error correction processing; it is obtained after performing text recognition on the data source image. Furthermore, considering that the data source image may contain text and formulas, during text recognition of the data source image, errors may occur due to the OCR engine's weak ability to recognize formulas or poor image quality, resulting in a large number of errors or noise in the initial recognized text. For example, "x" in a formula might be recognized as the letter "X", "√" as "V", and "sin" as "sim", etc. In other words, there is a difference between the initial recognized text and the actual text, affecting subsequent question-and-answer data extraction.

[0110] Based on this, the embodiments of the present invention input the initial recognition text into the text correction model, which automatically detects and corrects errors in the text, such as spelling errors, grammatical errors, semantic errors, etc., thereby improving the accuracy and readability of text recognition and obtaining the text to be extracted. This text to be extracted has higher accuracy and completeness, and can provide a better foundation for subsequent question and answer data extraction.

[0111] Among them, text error correction models can be rule-based error correction models, such as using edit distance algorithms to correct spelling errors; they can also be statistical error correction models, such as using N-gram models to correct grammatical errors; or they can be deep learning-based error correction models, such as using Transformer models to learn the contextual information of text and predict the correct text content.

[0112] Based on any of the above embodiments Figure 8 This is the second flowchart of the question-and-answer data extraction method provided by the present invention, as shown below. Figure 8 As shown, the method includes: The data source image can be obtained by converting a PDF document into an image format or by scanning a paper document. The image format can be TIFF, JPEG, etc. After obtaining the data source image, image classification models or rules can be used to determine whether the questions and answers are together. If the questions and answers are together (i.e., the question and answer texts are on the same page or adjacent pages, and there is a clear layout relationship between them, such as the question and answer being under the same question number, or the answer immediately following the question), it can be marked as QA not separated. If the questions and answers are not together (i.e., the question and answer texts are in different chapters or documents, and there is no clear layout relationship between them, such as the question being in the exercises section and the answer being in the answer explanation section at the end of the book), it can be marked as QA separated. This can then serve as a basis for selecting different data extraction strategies. For example, for data with QA not separated, a local search strategy can be used to quickly locate the answer; for data with QA separated, a global search strategy can be used to find the answer throughout the entire document.

[0113] Furthermore, considering the potential quality issues of the original source images, such as blurriness, tilt, or noise, these images can be converted into high-resolution image sequences. These sequences can then undergo preprocessing, including automatic rotation correction, denoising, binarization, and page segmentation. Given that data extraction models may experience performance issues when processing long texts, using four pages as input generally helps to better balance accuracy and efficiency, facilitating subsequent question-and-answer data extraction by the data extraction model.

[0114] Next, OCR recognition is performed on the acquired image sequence to obtain the text to be extracted. Simultaneously, layout analysis is performed on the source image to determine the layout information. OCR recognition can be performed using one or any combination of the following methods: OCR recognition of image sequences is performed using a specially trained OCR model.

[0115] An end-to-end scene text recognition method is adopted to recognize both text and formulas.

[0116] The text region and formula region are detected separately, and the corresponding OCR models are used for recognition respectively. Then the results are merged.

[0117] OCR recognition is performed on the image sequence to obtain preliminary recognition results. These preliminary results are then input into a text correction model, which uses its language context understanding capabilities to correct and improve any fuzzy or incorrect OCR recognition.

[0118] Based on the text to be extracted and the layout information, a prompt text is constructed and input into the data extraction model, which then extracts the question text and the answer text that matches the question text.

[0119] Furthermore, the output of the data extraction model (question text, answer text, the location of the question text, the location of the answer text, the relationship between the question text and the answer text, etc.) is converted into a unified structured data format, such as JSON, XML, or a relational database table. The structured data may contain the following information: a unique question-answer pair ID, question text (LaTeX representation containing mathematical formulas), answer text (LaTeX representation containing mathematical formulas), the document ID and page number of the question document, the document ID and page number of the answer text, and the confidence score of the data extraction.

[0120] After obtaining the structured data, it undergoes format cleaning to remove redundant spaces and line breaks, and standardizes the representation of mathematical formulas. Furthermore, rules, templates, or another (potentially smaller) model can be used for preliminary validation of the extracted structured data. For example, checking if both the question and answer contain at least one mathematical expression, or if the answer contains key variables from the question. For multiple-choice questions, it can be verified that the answer is in the option list. A larger model can then be used again, inputting the question and answer, to determine if the answer is reasonable or answers the question, and output a validation score. Data with low automated validation scores or proportionally sampled data undergoes manual review to further improve data quality.

[0121] The cleaned and verified structured data is stored in a database or data warehouse to build a mathematical question-and-answer dataset that can be used for training a large mathematical question-and-answer model.

[0122] The following is an example of the output results from the data extraction model: Chapter 3 Quadratic Equations, Example 3.1: Calculate the equation $x$. 2 Find the solution to -4x + 4 = 0. Solution: The original equation is (x - 2). 2 = 0$ Therefore, $x - 2 = 0$, so $x = 2$. Answer: $x=2$}.

[0123] The following is an example of cleaned and validated structured data: {Issue ID": "doc_X_page_Y_prob_Z", Problem Text: "Calculate equation $x" 2 Find the solution to -4x + 4 = 0. Answer text: The original equation is $(x-2)$. 2 = 0$ Therefore, $x - 2 = 0$, which gives $x = 2$. "Source Page Number": "Y", "Document ID": "doc_X", "Extract confidence level": 0.98}.

[0124] The question-and-answer data extraction device provided by the present invention is described below. The question-and-answer data extraction device described below can be referred to in correspondence with the question-and-answer data extraction method described above.

[0125] Based on any of the above embodiments Figure 9 This is a schematic diagram of the question-and-answer data extraction device provided by the present invention, as shown below. Figure 9 As shown, the device includes: The determining unit 910 is used to perform text recognition on the data source image to obtain the text to be extracted; Analysis unit 920 is used to perform layout analysis on the data source image to determine layout information. The layout information is used to characterize the physical attributes of each text block in the text to be extracted and the spatial arrangement characteristics of each text block in the text to be extracted. The physical attributes of each text block refer to the features that each text block itself has and can be directly measured or observed. The spatial arrangement characteristics of each text block in the text to be extracted refer to the position of each text block in the text to be extracted and its relative positional relationship with other text blocks. Extraction unit 930 is used to extract question text and answer text matching the question text from the text to be extracted based on the semantic content of the text to be extracted and the layout information.

[0126] Based on any of the above embodiments, based on the semantic content of the text to be extracted and the layout information, the question text and the answer text matching the question text are extracted from the text to be extracted, including: Based on the semantic content of the text to be extracted, the question text and multiple candidate answer texts corresponding to the question text are determined from the text to be extracted; Based on the layout information, the answer text that matches the question text is determined from each candidate answer text.

[0127] Based on any of the above embodiments, based on the semantic content of the text to be extracted, the question text and multiple candidate answer texts corresponding to the question text are determined from the text to be extracted, including: Based on the question text retrieval formula, the question text is determined by searching within the text to be extracted; Based on the position of the question text in the text to be extracted, the retrieval scope of the answer text is determined; Based on the answer text retrieval formula, the text to be extracted is retrieved within the retrieval scope to determine multiple candidate answer texts.

[0128] Based on any of the above embodiments, determining the answer text that matches the question text from each candidate answer text based on layout information includes: Based on the layout information, determine the content area to which the question text belongs; Based on the content area to which the question text belongs, the answer text that matches the question text is determined from each candidate answer text.

[0129] Based on any of the above embodiments, layout analysis is performed on the data source image to determine layout information. Based on the semantic content of the text to be extracted and the layout information, question text and answer text matching the question text are extracted from the text to be extracted, including: Based on the images from the data source, prompt text is constructed to indicate the extraction approach for the question-and-answer data. Based on the data extraction model, the application of prompt text extracts the text to be extracted from the source image. The layout of the source image is analyzed to determine the layout information. Based on the semantic content of the text to be extracted and the layout information, the question text and the answer text that matches the question text are extracted from the text to be extracted.

[0130] Based on any of the above embodiments, layout analysis is performed on the data source image to determine the layout information of the text to be extracted, including: Determine the type of each text block in the image from which the data is sourced; The layout information is determined based on the logical order and type of each text block.

[0131] Based on any of the above embodiments, layout analysis is performed on the data source image to determine the layout information of the text to be extracted, including: Using a pre-trained layout analysis model, preliminary layout segmentation is performed on the source image to obtain multiple text blocks; For each text block, extract visual and textual features; The extracted visual and text features are input into the fine-tuned layout analysis model, which outputs the physical properties of each text block. The layout analysis model is then fine-tuned based on the layout annotation data. Page layout information is constructed based on the physical properties of each text block and the spatial arrangement characteristics of each text block.

[0132] Based on any of the above embodiments, layout analysis is performed on the data source image to determine the layout information of the text to be extracted, including: The source image of the data is scanned to obtain multiple text blocks; For each text block, the rules in the layout analysis rule base are applied sequentially to determine whether each text block satisfies any of the rules; the layout analysis rule base includes layout information analysis rules for different types of text blocks; If each text block satisfies any rule, then the layout information of the corresponding text block is determined according to that rule; If any text block does not meet any rules, the layout information of any text block is determined based on the default rules; Integrate the layout information of all text blocks to construct the layout information of the text to be extracted.

[0133] Based on any of the above embodiments, text recognition is performed on the data source image to obtain the text to be extracted, including: Perform region detection on the source image to identify text and formula regions within the source image; Text recognition is performed on the text area and the formula area respectively to obtain the text to be extracted.

[0134] Based on any of the above embodiments, text recognition is performed on the data source image to obtain the text to be extracted, including: Perform text recognition on the source image to obtain the initial recognized text; Based on the text correction model, text correction is performed on the initially identified text to obtain the text to be extracted.

[0135] Figure 10 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 10As shown, the electronic device may include a processor 1010, a communications interface 1020, a memory 1030, and a communication bus 1040, wherein the processor 1010, the communications interface 1020, and the memory 1030 communicate with each other via the communication bus 1040. The processor 1010 can call logical instructions in the memory 1030 to execute a question-and-answer data extraction method. This method includes: performing text recognition on a data source image to obtain text to be extracted; performing layout analysis on the data source image to determine layout information, wherein the layout information is used to characterize the physical attributes of each text block in the text to be extracted and the spatial arrangement features of each text block in the text to be extracted; and extracting question text and answer text matching the question text from the text to be extracted based on the semantic content of the text to be extracted and the layout information.

[0136] Furthermore, the logical instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0137] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the question-and-answer data extraction method provided by the above methods. The method includes: performing text recognition on a data source image to obtain text to be extracted; performing layout analysis on the data source image to determine layout information, wherein the layout information is used to characterize the physical attributes of each text block in the text to be extracted and the spatial arrangement features of each text block in the text to be extracted; and extracting question text and answer text matching the question text from the text to be extracted based on the semantic content of the text to be extracted and the layout information.

[0138] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the question-and-answer data extraction method provided by the above methods. The method includes: performing text recognition on a data source image to obtain text to be extracted; performing layout analysis on the data source image to determine layout information, the layout information being used to characterize the physical properties of each text block in the text to be extracted and the spatial arrangement features of each text block in the text to be extracted; and extracting question text and answer text matching the question text from the text to be extracted based on the semantic content of the text to be extracted and the layout information.

[0139] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0140] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A question-and-answer data extraction method, characterized in that, include: Perform text recognition on the source image to obtain the text to be extracted; The image from the data source is analyzed to determine the layout information. The layout information is used to characterize the physical properties of each text block in the text to be extracted and the spatial arrangement characteristics of each text block in the text to be extracted. The physical properties of each text block refer to the features that each text block has itself and can be directly measured or observed. The spatial arrangement characteristics of each text block in the text to be extracted refer to the position of each text block in the text to be extracted and its relative positional relationship with other text blocks. Based on the semantic content of the text to be extracted and the layout information, the question text and the answer text matching the question text are extracted from the text to be extracted.

2. The question-and-answer data extraction method according to claim 1, characterized in that, The step of extracting question text and answer text matching the question text from the text to be extracted based on the semantic content of the text to be extracted and the layout information includes: Based on the semantic content of the text to be extracted, the question text and multiple candidate answer texts corresponding to the question text are determined from the text to be extracted; Based on the layout information, the answer text that matches the question text is determined from each candidate answer text.

3. The question-and-answer data extraction method according to claim 2, characterized in that, The step of determining the question text and multiple candidate answer texts corresponding to the question text from the text to be extracted based on the semantic content of the text to be extracted includes: Based on the question text retrieval formula, the question text is determined by searching the text to be extracted. Based on the position of the question text in the text to be extracted, the retrieval range of the answer text is determined; Based on the answer text retrieval formula, the text to be extracted is retrieved within the retrieval range to determine the multiple candidate answer texts.

4. The question-and-answer data extraction method according to claim 2, characterized in that, The step of determining the answer text that matches the question text from each candidate answer text based on the layout information includes: Based on the layout information, determine the content area to which the question text belongs; Based on the content area to which the question text belongs, the answer text that matches the question text is determined from each candidate answer text.

5. The question-and-answer data extraction method according to any one of claims 1 to 4, characterized in that, The step of performing layout analysis on the data source image to determine layout information, and extracting question text and answer text matching the question text from the text to be extracted based on the semantic content of the text to be extracted and the layout information, includes: Based on the image from the data source, a prompt text is constructed, which is used to indicate the extraction approach for the question-and-answer data; Based on the data extraction model, the prompt text is applied to obtain the text to be extracted from the data source image, and the layout analysis of the data source image is performed to determine the layout information. Based on the semantic content of the text to be extracted and the layout information, the question text and the answer text matching the question text are extracted from the text to be extracted.

6. The question-and-answer data extraction method according to any one of claims 1 to 4, characterized in that, The step of performing layout analysis on the data source image to determine the layout information of the text to be extracted includes: Determine the type of each text block in the image from which the data is sourced; The layout information is determined based on the logical order and type of each text block.

7. The question-and-answer data extraction method according to any one of claims 1 to 4, characterized in that, The step of performing layout analysis on the data source image to determine the layout information of the text to be extracted includes: Using a pre-trained layout analysis model, preliminary layout segmentation is performed on the data source image to obtain multiple text blocks; For each text block, extract visual and textual features; The extracted visual and text features are input into the fine-tuned layout analysis model, which outputs the physical attributes of each text block. The layout analysis model is fine-tuned based on the layout annotation data. The layout information is constructed based on the physical properties of each text block and the spatial arrangement characteristics of each text block.

8. The question-and-answer data extraction method according to any one of claims 1 to 4, characterized in that, The step of performing layout analysis on the data source image to determine the layout information of the text to be extracted includes: The source image of the data is scanned to obtain multiple text blocks; For each text block, the rules in the layout analysis rule base are applied sequentially to determine whether each text block satisfies any rule; the layout analysis rule base includes layout information analysis rules for different types of text blocks; If each text block satisfies any rule, then the layout information of the corresponding text block is determined according to the rule. If any text block does not meet any rules, the layout information of that text block is determined based on the default rules; Integrate the layout information of all text blocks to construct the layout information of the text to be extracted.

9. The question-and-answer data extraction method according to any one of claims 1 to 4, characterized in that, The step of performing text recognition on the source image to obtain the text to be extracted includes: Perform region detection on the data source image to determine the text region and formula region in the data source image; Text recognition is performed on the text region and the formula region respectively to obtain the text to be extracted.

10. The question-and-answer data extraction method according to any one of claims 1 to 4, characterized in that, The step of performing text recognition on the source image to obtain the text to be extracted includes: The image from the data source is subjected to text recognition to obtain initial recognized text; Based on the text correction model, the initial identified text is corrected to obtain the text to be extracted.

11. A question-and-answer data extraction device, characterized in that, include: The determining unit is used to perform text recognition on the source image of the data to obtain the text to be extracted; The analysis unit is used to perform layout analysis on the data source image to determine layout information. The layout information is used to characterize the physical attributes of each text block in the text to be extracted and the spatial arrangement characteristics of each text block in the text to be extracted. The physical attributes of each text block refer to the features that each text block itself has and can be directly measured or observed. The spatial arrangement characteristics of each text block in the text to be extracted refer to the position of each text block in the text to be extracted and its relative positional relationship with other text blocks. The extraction unit is used to extract question text and answer text matching the question text from the text to be extracted based on the semantic content of the text to be extracted and the layout information.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the question-and-answer data extraction method as described in any one of claims 1 to 10.

13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the question-and-answer data extraction method as described in any one of claims 1 to 10.

14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the question-and-answer data extraction method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Visual question and answer method and device, electronic equipment and storage medium

    CN114445826A

  • Question answering method and device based on document image, equipment, storage medium and program product

    CN118586403A