Method, device and system for generating question and answer pairs based on portable document
By constructing a structured dataset of portable documents and dividing text blocks based on heading levels, the problems of poor segmentation quality and information loss when generating question-and-answer pairs from portable documents are solved, resulting in high-quality question-and-answer pairs that ensure the semantic integrity and relevance of text blocks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LAUNCH TECH CO LTD
- Filing Date
- 2026-03-17
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies for generating question-and-answer pairs based on portable documents suffer from problems such as poor segmentation quality, information loss, insufficient metadata association, and poor quality stability, resulting in low-quality generated question-and-answer pairs.
Construct a structured dataset of portable documents, including heading level and non-text information, divide the text into blocks by heading level, generate high-quality question-answer pairs, and perform quantitative and qualitative validation.
It achieves semantic integrity and relevance of text blocks, generates high-quality question-answer pairs that are strongly associated with the document context, and avoids problems such as information loss and unstable quality.
Smart Images

Figure CN122019729A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing, and specifically relates to a method, apparatus, and system for generating question-and-answer pairs based on portable documents. Background Technology
[0002] Currently, many scenarios require generating question-and-answer pairs based on portable documents (e.g., portable document format (PDF) documents). How to generate high-quality question-and-answer pairs based on portable documents is a problem that needs to be solved. Summary of the Invention
[0003] This application provides a method, apparatus, and system for generating question-and-answer pairs based on portable documents, which can generate high-quality question-and-answer pairs based on portable documents.
[0004] The first aspect of this application provides a method for generating question-answer pairs from a portable document. The method includes: constructing a structured dataset corresponding to the portable document, wherein the structured dataset includes the title level of the portable document; dividing the structured dataset based on the title level to obtain a first set of text blocks; and generating a question-answer pair corresponding to each text block in the first set of text blocks, wherein each question-answer pair consists of a question and an answer corresponding to the question.
[0005] In some implementations, the structured dataset includes non-textual information from the portable document.
[0006] In some implementations, the step of dividing the structured dataset based on the title level to obtain a first set of text blocks includes: dividing the structured dataset with the title level as the boundary to obtain a basic set of text blocks; calculating the semantic similarity of the text within each text block in the basic set of text blocks; if the basic set of text blocks does not include text blocks that satisfy a first condition, then the basic set of text blocks is determined as the first set of text blocks; if the basic set of text blocks includes text blocks that satisfy the first condition, then the structured dataset is re-divided to obtain the first set of text blocks; wherein, the first condition includes: the semantic similarity corresponding to the text block includes semantic similarity below a preset threshold.
[0007] In some implementations, the step of dividing the structured dataset based on the heading level to obtain a first set of text blocks includes: dividing the structured dataset with the heading level as the boundary to obtain a basic set of text blocks; merging the first text block and the second text block in the basic set of text blocks to obtain the first set of text blocks; wherein the first text block and the second text block are adjacent, and the first text block or the second text block is an empty text block.
[0008] In some implementations, generating a question-and-answer pair corresponding to each text block in the first set of text blocks includes: for each text block in the first set of text blocks, generating a question-and-answer pair associated with the metadata of the text block based on the text block and the metadata of the text block, wherein the metadata of the text block includes information for indicating the source of the text block and / or information for indicating the attributes of the text block.
[0009] In some implementations, generating the question-and-answer pairs corresponding to the text block includes one or more of the following: when the text length of the text block is greater than or equal to a preset length, dividing the text block into multiple sub-text blocks, and generating a question-and-answer pair corresponding to each of the multiple sub-text blocks; generating a first number of question-and-answer pairs corresponding to the text block, wherein the first number satisfies: the first number is related to the text length of the text block, and / or the first number is greater than or equal to a preset number.
[0010] In some implementations, after generating the question-answer pair corresponding to the text block, the method further includes performing quantitative and qualitative verification on the question-answer pair.
[0011] In some implementations, after generating the question-answer pairs corresponding to the text block, the method further includes: determining question-answer pairs with similar questions among the question-answer pairs corresponding to the text block; calculating the semantic similarity of the answers in the question-answer pairs with similar questions; determining question-answer pairs with similar answers from the question-answer pairs with similar questions based on the semantic similarity; and removing one or more question-answer pairs from the question-answer pairs with similar answers.
[0012] A second aspect of this application provides an apparatus for generating question-answer pairs based on a portable document. The apparatus includes: a construction module for constructing a structured dataset corresponding to the portable document, wherein the structured dataset includes a title level of the portable document; a partitioning module for partitioning the structured dataset based on the title level to obtain a first set of text blocks; and a generation module for generating a question-answer pair corresponding to each text block in the first set of text blocks, wherein each question-answer pair consists of a question and an answer corresponding to the question.
[0013] A third aspect of this application provides a system for generating question-and-answer pairs based on portable documents. The system includes: a memory for storing code; and a processor for executing the code stored in the memory to perform the method described in the first aspect.
[0014] In this embodiment, a structured dataset corresponding to the portable document is constructed, providing a data foundation for subsequent text block segmentation. Furthermore, the structured dataset is segmented based on the header level of the portable document to obtain the first set of text blocks, ensuring the semantic integrity and relevance of the text blocks. Using the method provided in this application, high-quality question-answer pairs can be generated based on the portable document. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0016] Figure 1 This is a schematic flowchart illustrating a method for generating question-and-answer pairs based on portable documents, provided in an embodiment of this application.
[0017] Figure 2 Provided for the embodiments of this application Figure 1 Example flowchart of step S110.
[0018] Figure 3 Provided for one embodiment of this application Figure 1 Example flowchart of step S120.
[0019] Figure 4 Provided for another embodiment of this application Figure 1 Example flowchart of step S120.
[0020] Figure 5 Provided for another embodiment of this application Figure 1 Example flowchart of step S120.
[0021] Figure 6 Provided for one embodiment of this application Figure 1 Example flowchart of step S130.
[0022] Figure 7 This is a schematic diagram of the device for generating question-and-answer pairs based on portable documents provided in an embodiment of this application.
[0023] Figure 8 This is a schematic diagram of the structure of a system for generating question-and-answer pairs based on portable documents, provided in an embodiment of this application. Detailed Implementation
[0024] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0025] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0026] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0027] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0028] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0029] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0030] Currently, many scenarios require the generation of question-and-answer pairs based on portable documents (e.g., PDF documents). For example, during enterprise digital transformation, a large amount of core knowledge, such as product manuals and technical manuals, is stored in PDF format. How to efficiently extract knowledge from these PDF documents and generate high-quality question-and-answer pairs is key to building enterprise knowledge bases and intelligent customer service systems.
[0031] Traditional methods of manually compiling question-and-answer pairs from documents suffer from high costs, low efficiency, and delayed updates. With the development of large language models, automated question-and-answer pair generation has become a mainstream trend. Related technologies have proposed schemes for automatically generating question-and-answer pairs based on portable documents. Taking PDF documents as an example, these schemes typically include the following steps: First, the PDF document is converted to plain text format using a tool, discarding non-text elements such as images and tables. Then, the plain text is mechanically segmented based on a fixed word count threshold. For example, each 500 characters is divided into a text block. By segmenting the portable document, multiple text blocks can be obtained. These text blocks can be understood as the basic processing units obtained from segmenting the portable document, forming the basis for subsequent question-and-answer pair generation. Next, a single large language model is used to generate question-and-answer pairs based on the text blocks. Finally, simple deduplication is performed using string matching before outputting the results.
[0032] However, the proposed solutions for automatically generating question-and-answer pairs based on portable documents have many problems and cannot guarantee the quality of the generated question-and-answer pairs.
[0033] For example, solutions proposed in related technologies often suffer from poor segmentation quality because their segmentation logic deviates from the inherent structure of the document. These technologies typically mechanically segment portable documents based on a fixed word count or simple paragraphs, without considering the document's text level and content relevance. This can lead to splitting coherent content from the same chapter into multiple text blocks, or merging unrelated content from different chapters into a single text block. In other words, fragmented segmentation or mixed content is prone to occur, resulting in semantically incomplete or logically inconsistent question-and-answer pairs.
[0034] For example, solutions proposed in related technologies lose a significant amount of information because they fail to consider non-text elements such as images and tables when parsing documents. Portable documents typically contain multiple elements, including text, images, and tables, but related technologies only extract text during document parsing, discarding non-text elements like images and tables. However, a large amount of information in portable documents (e.g., numerous key technical parameters and operating procedures in PDF product manuals) is usually presented in the form of non-text elements such as images and tables. Discarding non-text elements causes the generated question-and-answer pairs to omit core knowledge, resulting in a question-and-answer pair that fails to fully cover the value of the document.
[0035] For example, the solutions proposed in related technologies fail to establish a link between question-and-answer pairs and metadata. Metadata can include information indicating the source of a text block and / or information indicating the attributes of the text block. Metadata can include one or more of the following: the storage path of the portable document, the name of the portable document, the source page number of the text block, the source title of the text block, the product name of the text block, the product category of the text block, etc. In related technologies, metadata such as the chapter title, page number, and product name are not fully utilized when generating question-and-answer pairs. Therefore, the generated question-and-answer pairs are disconnected from the document context. The questions in the question-and-answer pairs cannot clearly point to a specific document or a specific product, and the answers in the question-and-answer pairs cannot be accurately linked to the document's source. This makes it difficult to locate the document location corresponding to the answer during subsequent knowledge retrieval.
[0036] For example, the solutions proposed in related technologies suffer from poor quality stability due to the lack of a multi-level, quantitative quality control system. These solutions rely on a single model to generate question-and-answer pairs and employ simple string matching checks to deduplicate them. Since question-and-answer pairs generated by a single model are prone to containing unfounded, inaccurate, or completely fabricated content, and the simple string deduplication scheme cannot identify semantic repetitions and lacks quantitative quality assessment standards, it is difficult to effectively filter out error rates and redundant question-and-answer pairs. This results in poor quality stability of the generated question-and-answer pairs and makes them unsuitable for direct application in high-reliability scenarios.
[0037] As described above, the proposed methods for generating question-and-answer pairs based on portable documents have many problems, resulting in low-quality generated question-and-answer pairs. Due to the multi-element nature and hierarchical complexity of portable documents, end-to-end automation "from portable document parsing to high-quality question-and-answer pair generation" still faces many bottlenecks. The current core pain point in the industry is concentrated on "structured parsing and precise segmentation of portable documents," that is, how to achieve unified structured processing of multiple elements in portable documents, achieve precise segmentation based on the inherent hierarchy of the document, and then generate high-quality question-and-answer pairs that are strongly related to the document context.
[0038] In conclusion, how to generate high-quality question-answer pairs based on portable documents is a technical problem that needs to be solved.
[0039] To address the aforementioned technical problems, this application provides a method for generating question-and-answer pairs based on portable documents. In this application, a structured dataset corresponding to the portable document is constructed, providing a data foundation for subsequent text block segmentation. Furthermore, the structured dataset is segmented based on the header level of the portable document to obtain a first set of text blocks, ensuring the semantic integrity and relevance of the text blocks. Using the method provided in this application, high-quality question-and-answer pairs can be generated based on portable documents.
[0040] The method provided in the embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0041] like Figure 1 As shown, the method provided in this application embodiment includes the following steps S110 to S130.
[0042] In step S110, a structured dataset corresponding to the portable document is constructed, wherein the structured dataset includes the title hierarchy of the portable document.
[0043] Step S110 can be performed by the model. For example, a portable document is input into the model, and the model outputs a structured dataset corresponding to the portable document. Step S110 can be referred to as the structured parsing of the portable document.
[0044] The portable documents mentioned in this application embodiment can be understood as documents with good visual structure but lacking machine-friendly and semantically clear data structures. Therefore, the portable documents in this application embodiment can be understood as unstructured documents. In this application embodiment, the portable document serves as the basis for generating question-answer pairs; therefore, the portable document can also be called the original document or source document. The portable documents mentioned in this application embodiment can include documents of various formats, such as PDF documents. This application embodiment does not specifically limit the type of portable document. The portable documents in this application embodiment include, but are not limited to, product manuals, technical manuals, etc.
[0045] The structured datasets mentioned in this application can be understood as data stored in a predefined format. These predefined formats include, but are not limited to, JavaScript Object Notation (JSON) format and Extensible Markup Language (XML) format. Data stored in JSON format can be called a JSON structured dataset, and data stored in XML format can be called an XML structured dataset. JSON structured datasets offer better model reading efficiency, while XML structured datasets have advantages in hierarchical nesting. In practical applications, the appropriate structured dataset format can be selected based on requirements. When using an XML structured dataset, subsequent partitioning of the XML structured dataset can be achieved based on XML tag parsing.
[0046] For example, in step S110, a JSON structured dataset corresponding to the portable document is constructed; this process can be called JSON structured modeling. As another example, in step S110, an XML structured dataset corresponding to the portable document is constructed; this process can be called XML structured modeling.
[0047] The structured dataset includes the heading hierarchy of portable documents. The heading hierarchy can be understood as the chapter titles (source_title) of the portable document. The heading hierarchy is used to construct the hierarchical structure of the portable document. Portable documents have a certain hierarchical structure, which is an inherent structure of the portable document. The hierarchical structure of a portable document reflects its organizational framework and the logical relationships between its content. For example, the hierarchical structure of a portable document can be constructed by setting different levels of headings (first-level heading, second-level heading, third-level heading, etc.). As a concrete example, the structured dataset includes the chapter titles of portable documents, where number 1 corresponds to a first-level heading, number 2 to a second-level heading, and number 3 to a third-level heading.
[0048] By constructing a structured dataset corresponding to portable documents, unstructured portable documents can be converted into structured data, providing a data foundation for subsequent accurate segmentation.
[0049] In some implementations, step S110 constructs a structured dataset corresponding to the portable document, including: converting the portable document into a document of a specific format, and constructing a structured dataset corresponding to the portable document based on the document of the specific format.
[0050] For example, you can first convert a portable document to a specific format, then extract and store the text content and other elements from that formatted document, and finally use those extracted elements to build a structured dataset according to certain rules. The specific format mentioned here can include Markdown. As a concrete example, you can first convert a portable document to Markdown format, then extract and store the text content and other elements from that Markdown document, and finally use those extracted elements to build a JSON or XML structured dataset according to certain rules.
[0051] Since portable documents are not easily identifiable, converting them into documents of a specific format can help in the construction of structured datasets.
[0052] In some implementations, the structured dataset in step S110 contains non-textual information from portable documents.
[0053] Information in portable documents can be divided into text information and non-text information. Text information can be understood as text content that can be selected, copied, and searched; essentially, text information is a series of "character codes" with coordinates and font information. Non-text information can be understood as visual elements that cannot be directly selected and edited; non-text information is "drawn" on the page in the form of images, tables, or paths. Both text and non-text information may contain important information in the portable document; therefore, non-text information is also a key element of portable documents. In the embodiments of this application, non-text information may include one or more of the following: images, tables, positional relationships between text and images, positional relationships between text and tables, and positional relationships between images and tables.
[0054] For example, in step S110, the portable document is first converted to a specific format, and then a structured dataset is constructed based on the specific format document. When converting the portable document to the specific format document, in addition to preserving the text, images, tables, and the original positional relationships between the text, images, and tables are also preserved. When constructing the structured dataset based on the specific format document, in addition to extracting and storing the text content from the document, it is also necessary to extract and store the image files and table files from the document. In this way, the structured dataset constructed based on the specific format document can include the non-textual information of the portable document.
[0055] By employing a unified, multi-element structured parsing scheme, the resulting structured dataset retains key elements from the portable document, such as text, images, and tables, and associates them with the text content through the `Img_path` field. This approach solves the information loss problem caused by discarding non-text elements in traditional schemes, ensuring that the generated question-and-answer pairs cover all types of core knowledge within the portable document.
[0056] Therefore, preserving non-textual information in structured datasets can avoid the loss of important information in the final question-answer pairs, thus providing a foundation for generating high-quality question-answer pairs in the future.
[0057] In addition to the heading hierarchy and non-textual information mentioned above, the structured dataset corresponding to a portable document can also include other information.
[0058] In some implementations, the structured dataset corresponding to the portable document may further include one or more of the following information: information indicating the storage path of the portable document, information indicating the file name of the portable document, information indicating the storage path of images in the portable document, and information indicating the storage path of tables in the portable document. In the embodiments of this application, this information can be collectively referred to as basic metadata. Basic metadata belongs to the metadata described above.
[0059] The filename of a portable document can be referred to as the source document name. The storage paths for images and tables can be collectively referred to as the image / table path (Img_path). Taking a PDF document as an example, the structured dataset corresponding to a PDF document can include: PDF document path (PDF_path), source document name (Source_PDF), chapter and page numbers (Source_page), chapter title (Source_title), and image / table path (Img_path).
[0060] To better understand step S110, see below. Figure 2 The steps of S110 will be explained in more detail with specific examples. Figure 2 The example shown includes steps S201 to S206.
[0061] In step S201, the PDF document is read.
[0062] In this step, PDF documents such as product manuals or technical brochures are imported into the model. The model then reads the PDF documents.
[0063] In step S202, the PDF document is converted to a different format.
[0064] In this step, the PDF document is converted to Markdown format, preserving the original positional relationships between text, images, tables, and other elements.
[0065] In step S203, extract text information and non-text information.
[0066] In this step, text content, image files, table files, and other elements are extracted from the Markdown document, and these elements are stored separately.
[0067] In step S204, a JSON structured dataset is constructed.
[0068] In this step, the extracted elements are used to construct a JSON structured dataset according to "element type (text / image / table), content / path, heading level (text_level), page number (page_idx)". The elements containing the text_level field are the heading text, and the numbers 1 / 2 / 3... correspond to the first-level / second-level / third-level... headings, respectively.
[0069] In step S205, basic metadata is recorded.
[0070] In this step, basic metadata such as the storage path and file name of the PDF document are recorded.
[0071] In step S206, a JSON structured dataset is output.
[0072] This step outputs a JSON structured dataset containing text, image / table paths, heading levels, page numbers, and basic metadata.
[0073] Steps S201 to S206 convert unstructured PDF documents into structured datasets, providing a data foundation for subsequent precise segmentation. Furthermore, all key elements are preserved, thus preventing subsequent information loss.
[0074] In the example above, step S204 constructed a JSON structured dataset based on the extracted elements. In other examples, an XML structured dataset can be constructed based on the extracted elements. If an XML structured dataset is constructed, attributes such as element type, heading level, and page number can be defined using XML tags.
[0075] The above provides a detailed description of step S110. Step S120 will be described below.
[0076] In step S120, a structured dataset is divided based on the title hierarchy of the portable document to obtain the first set of text blocks.
[0077] Step S110 mentions constructing a structured dataset corresponding to the portable document, wherein the structured dataset includes the header level of the portable document. In step S120, the structured dataset can be divided based on the header level to obtain a first set of text blocks. Step S120 can be executed by a model. For example, the structured dataset obtained in step S110 can be input into the model, and the model can output the first set of text blocks.
[0078] In this embodiment, the first text block set can be understood as a collection of multiple text blocks. The first text block set can also be called a standardized text block set. Each text block in the first text block set can be viewed as a basic processing unit obtained by segmenting the structured dataset corresponding to the portable document. The text block is the basis and foundation for subsequent question-and-answer pair generation. In other words, the text block is the core input unit for subsequent question-and-answer pair generation.
[0079] Dividing a structured dataset by heading level can be understood as grouping the content between one heading level and the next into the same text block. The content within this block corresponds to the content of the preceding heading level. This structured segmentation based on heading level ensures the semantic integrity of each text block.
[0080] Structured datasets based on heading levels in portable documents allow for the association of text blocks in the first set of text blocks with unique chapter titles. Different portable documents may have different heading levels; the heading level recognition rules can be adjusted to suit different types of portable documents.
[0081] The structured dataset based on the title hierarchy of portable documents can achieve structured segmentation of the dataset, avoiding the semantic breaks and contextual disconnects caused by traditional mechanical segmentation. This ensures the semantic integrity and relevance of text blocks, thereby guaranteeing the quality of the subsequently generated question-and-answer pairs.
[0082] Step S110 mentions that the structured dataset, in addition to containing the text information of the portable documents, may also contain the non-text information of the portable documents. Therefore, the text blocks obtained by partitioning the structured dataset based on the heading level may contain not only the text information of the portable documents but also the non-text information of the portable documents. For example, in addition to the body text, the text blocks may also contain image information and / or table information. Image information can be presented in the text block as the storage path of the image, such as an image link. Table information can be presented in the text block as the storage path of the table, such as a table link. The image paths and table paths mentioned here can be used to indicate the attributes of the text blocks, and therefore, image paths and table paths are also part of the metadata of the text blocks. The inclusion of image paths and / or table paths in the text blocks can improve the metadata of the text blocks, thereby providing contextual support for the subsequent generation of question-answer pairs and helping to generate high-quality question-answer pairs.
[0083] Furthermore, text blocks obtained from structured datasets partitioned based on heading levels can also include one or more of the following: the name of the portable document corresponding to the text block (i.e., the source document name), the chapter title corresponding to the text block, and the chapter page number corresponding to the text block. The source document name, chapter title, and chapter page number mentioned here can be used to indicate the source of the text block and all belong to the metadata described above. The inclusion of one or more of the source document name, chapter title, and chapter page number in a text block can further improve the metadata of the text block and provide contextual support for subsequent question-and-answer pair generation.
[0084] In some implementations, the text blocks in the first set of text blocks can have a preset structure or preset form. A text block can include multiple fields, each used to carry different elements contained within the text block. Fields used to carry metadata can be called metadata fields. For example, a text block includes one or more of the following fields: chapter title (source_title), chapter page number (source_page), body text (text), image / table path (Img_path), and text block (Chunk). The chapter title field carries the chapter title corresponding to the text block. The chapter page number field carries the chapter page number corresponding to the text block. For example, the chapter page number page_idx+1 corresponding to the text block is stored in the chapter page number field. The body text field carries the body text of the text block. For example, the body text of the text block is stored in the body text field concatenated with newline characters. The image / table path field stores the storage path of images and / or tables. For example, the storage paths of images and / or tables are stored in the image / table path field in the format "", where different paths are connected by semicolons. The text block field is used to carry the content of both the body text field and the image path field. For example, the content of the body text field and the content of the image path field are connected by newlines in the order they appear in a portable document and stored in the text block field.
[0085] It should be understood that in some cases, the set of text blocks obtained by partitioning the structured dataset based on the heading level is the final set of text blocks, i.e., the first set of text blocks. In other cases, the set of text blocks obtained by partitioning the structured dataset based on the heading level is not the final set of text blocks; the first set of text blocks can only be obtained after processing the set of text blocks obtained by partitioning the structured dataset based on the heading level. See the description below for details.
[0086] In some implementations, such as Figure 3 As shown, step S120, which divides the structured dataset based on the title level to obtain the first set of text blocks, may include the following steps S1210 to S1230.
[0087] In step S1210, the structured dataset is divided with the title level as the boundary to obtain a set of basic text blocks.
[0088] For example, a structured dataset can be divided into basic text blocks by using the heading level as the boundary.
[0089] In step S1220, for each text block in the basic text block set, the semantic similarity of the text within the text block is calculated.
[0090] In other words, for each text block in the basic text block set, the semantic similarity corresponding to that text block can be calculated. The semantic similarity of text within a text block can be understood as the semantic similarity between a part of the content within the text block and the content of the entire text block. For each text block, multiple semantic similarities can be calculated. For example, if a text block is divided into multiple segments, the semantic similarity between each segment and the text block can be calculated. The number of semantic similarities corresponding to the text block is the same as the number of segments contained in the text block.
[0091] In step S1230, if the basic text block set does not include a text block that satisfies the first condition, then the basic text block set is determined as the first text block set.
[0092] The first condition here includes: the semantic similarity of the text blocks includes semantic similarity below a preset threshold. After calculating the semantic similarity of each text block in step S1220, if the semantic similarity of a certain text block includes semantic similarity below the preset threshold, then the text block is considered to meet the first condition. If the basic text block set does not include text blocks that meet the first condition, then the basic text block set is determined as the first text block set.
[0093] In step S1240, if the basic text block set includes text blocks that satisfy the first condition, the structured dataset is re-divided to obtain the first text block set.
[0094] After calculating the semantic similarity of each text block in step S1220, if the semantic similarity of a text block includes a semantic similarity below a preset threshold, then the text block is considered to meet the first condition. If the basic text block set includes text blocks that meet the first condition, the structured dataset is re-divided to obtain the first text block set. For example, if the basic text block set includes text blocks that meet the first condition, the partitioning boundaries are readjusted, and the structured dataset is re-divided based on the adjusted partitioning boundaries to obtain the first text block set.
[0095] By performing steps S1210 to S1240 above, the semantic relevance of text blocks can be further improved. Especially for portable documents with unclear heading levels, performing steps S1210 to S1240 to divide the structured dataset can further ensure the quality of the subsequently generated question-answer pairs.
[0096] In some implementations, such as Figure 4 As shown, step S120 divides the structured dataset based on the title level to obtain the first set of text blocks, which may include the following steps S1210 and S1250.
[0097] In step S1210, the structured dataset is divided with the title level as the boundary to obtain a set of basic text blocks.
[0098] For an explanation of step S1210, please refer to the previous description; it will not be repeated here.
[0099] In step S1250, the first text block and the second text block in the basic text block set are merged to obtain the first text block set.
[0100] Here, the first and second text blocks are adjacent, and either the first or second text block is an empty text block. In this embodiment, an empty text block can be understood as a text block that does not contain any substantial content. For example, if a text block contains neither body text nor image or table information, it can be considered an empty text block. Generally, the content between a first-level heading and the first second-level heading under that heading may correspond to an empty text block, the content between a second-level heading and the first third-level heading under that heading may correspond to an empty text block, and so on.
[0101] As a concrete example, the text blocks in the first set of text blocks obtained by partitioning the structured dataset include the following fields: chapter title (source_title), chapter page number (source_page), body text (text), image / table path (Img_path), and text block (Chunk). If the text block (Chunk) field is empty, the text block is considered an empty text block.
[0102] In step S1250, an empty text block is merged with its adjacent text blocks into a single text block. Taking an empty text block as an example, when merging text blocks, the text block following the first text block (or the next text block after the first text block) can be used as the second text block. When merging the first and second text blocks, the chapter titles of the first and second text blocks can be merged. Furthermore, when merging the first and second text blocks, all content in the first text block except for the chapter titles can be discarded. In other words, for the text block formed by merging the first and second text blocks, all content in that text block except for the chapter titles is based on the content of the second text block.
[0103] As a specific example, if the text block (Chunk) field of a certain text block in the basic text block set obtained in step S1210 is empty, then the chapter title (source_title) field of that text block is merged with the chapter title (source_title) field of the next text block, and the contents of other fields are based on the contents of the next text block to form a new text block.
[0104] The two implementation methods of step S120 have been described in detail above. It should be understood that the two implementation methods described above can be used alone or in combination.
[0105] In some implementations, prior to step S120, the method provided in this application embodiment may further include: removing redundant information from the structured dataset. This redundant information includes one or more of the following: homepage copyright notice, endpage notice, header, footer, and table of contents. This process can be referred to as redundancy cleaning.
[0106] In some implementations, prior to step S120, the method provided in this application embodiment may further include: validating the content of the structured dataset. Validating the content of the structured dataset may include one or more of the following: verifying the accuracy of text recognition in the structured dataset, and verifying the consistency between the order of elements in the structured dataset and the order of elements in the portable document.
[0107] If the identified text in the structured dataset is inaccurate, the method provided in this application embodiment may further include: correcting the inaccurate text in the structured dataset. If the identified order of elements in the structured dataset is inconsistent with or deviates from the order of elements in the portable document, the method provided in this application embodiment may further include: correcting the order of elements in the structured dataset.
[0108] For better understanding, please refer to the following. Figure 5 The steps of S120 will be explained in more detail with specific examples. Figure 5 The example shown includes steps S501 to S508.
[0109] In step S501, input a JSON structured dataset.
[0110] In this step, the JSON structured dataset is input into the model.
[0111] In step S502, redundant information is cleaned.
[0112] In this step, redundant information is removed from the JSON structured dataset. Redundant information includes one or more of the following: homepage copyright notice, endpage notice, header, footer, and table of contents.
[0113] In step S503, content accuracy verification and sequence verification are performed.
[0114] In this step, the JSON structured dataset is validated for content accuracy. If inaccurate text is detected in the JSON structured dataset, the content accuracy validation fails; otherwise, it passes. Additionally, the JSON structured dataset is validated for order. If the order of elements in the JSON structured dataset is found to be inconsistent with or deviates from the order of elements in the source document, the order validation fails; otherwise, it passes.
[0115] If both the content accuracy check and the sequence check pass, proceed to step S504. If either the content accuracy check or the sequence check fails, proceed to step S505.
[0116] In step S504, the data in the JSON structured dataset is reprocessed.
[0117] In this step, if the content accuracy check in step S503 fails, then in step S504, inaccurate text in the JSON structured dataset is corrected. If the order check in step S503 fails, then in step S504, the order of elements in the JSON structured dataset is corrected.
[0118] In step S505, the JSON structured dataset is divided into text blocks based on the header level.
[0119] In this step, the JSON structured dataset is divided along the header level to obtain a set of basic text blocks.
[0120] In step S506, the metadata fields of the text block are completed.
[0121] In this step, the title text corresponding to the text block is stored in the chapter title (source title) field. The chapter page number corresponding to the text block, page_idx+1, is stored in the chapter page number (source page) field. The body text of the text block is concatenated with newline characters and stored in the body text (text) field. The storage paths of images and / or tables are stored in the image / table path (Img_path) field in the format "", where different paths are connected with semicolons. The content of the body text (text) field and the content of the image / table path (Img_path) field are concatenated with newline characters in the order of the source document and stored in the text block (Chunk) field.
[0122] In step S507, empty text blocks are merged.
[0123] In this step, if the Chunk field of a text block in the base text block set is empty, the source_title field of that text block is merged with the source_title field of the next text block, and the contents of other fields are based on the contents of the next text block to form a new text block.
[0124] In step S508, a set of standardized text blocks is output.
[0125] In this step, the model outputs a standardized set of text blocks, divided by chapter level, consisting of the source document name (source_PDF), chapter title (source_title), chapter page number (source_page), main text (text), image path (Img_path), and text chunk (Chunk) fields.
[0126] In steps S501 to S508, redundant information is removed through cleaning, and the semantic integrity and relevance of text blocks are ensured through structured segmentation. Simultaneously, the metadata is improved, providing contextual support for the subsequent generation of question-answer pairs.
[0127] The above text has provided a detailed description of step S120. Step S130 will be described below.
[0128] See Figure 1 In step S130, based on each text block in the first text block set, a question-answer pair corresponding to the text block is generated, wherein each question-answer pair consists of a question and the answer corresponding to the question.
[0129] For each text block in the first set of text blocks, one or more question-answer pairs can be generated. The question-answer pairs generated for the first set of text blocks can form a question-answer pair set, which can be called the initial question-answer pair set.
[0130] Step S130 can be performed by the model. That is, the first set of text blocks can be input into the model, and the model outputs the question-answer pairs corresponding to the first set of text blocks. The model here can include a large language model, such as a locally deployed large language model. This application embodiment does not specifically limit the model used to generate question-answer pairs, and large language models include, but are not limited to: GPT model, Llama model, Qwen model, etc.
[0131] In the embodiments of this application, other operations can be performed before generating question-answer pairs, which will be described below.
[0132] In some implementations, prior to step S130, the method provided in this application embodiment may further include: completing the metadata of each text block.
[0133] The supplementary metadata may include, but is not limited to, metadata related to portable documents and metadata related to products. Metadata related to portable documents may include the storage path of the portable document, such as the storage path of a PDF (PDF_path). Metadata related to products may include, but is not limited to, product name (product_name) and product category (product_category). As a concrete example, each text block in the first text block set obtained in step S120 can be labeled with its corresponding PDF path (PDF_path), product name (Product_name), and product category (Product_Category). Supplementing the metadata of text blocks, especially product-related metadata, can provide product-specific information for question-and-answer generation.
[0134] In some implementations, question-answer pairs can be generated directly based on the text blocks in the first set of text blocks obtained in step S120.
[0135] For example, after obtaining the first set of text blocks in step S120, the first set of text blocks is directly input into the large language model. The large language model generates question-answer pairs corresponding to each text block.
[0136] In some other implementations, the text blocks in the first set of text blocks obtained in step S120 can be converted into a specific format, and then question-answer pairs can be generated based on the text blocks in the specific format.
[0137] The specific format here can include comma-separated values (CSV) format. For example, after converting the first set of text blocks in JSON format obtained in step S120 into a first set of text blocks in CSV format, the first set of text blocks in CSV format is input into the large language model. The large language model generates question-and-answer pairs corresponding to each text block in the first set of text blocks in CSV format. Converting the text blocks in the first set of text blocks into a specific format and then generating question-and-answer pairs based on the first set of text blocks facilitates subsequent model reading and processing of data, thereby improving subsequent processing efficiency.
[0138] Before converting the text blocks in the first set of text blocks to a specific format, the operation mentioned above, which completes the atom count of each text block, can be performed first. This converts the completed original data into a specific format as well, further facilitating subsequent model reading and processing, thereby improving subsequent processing efficiency.
[0139] For example, the text blocks in the first set of text blocks in JSON format obtained in step S120 are input into the model. The large language model first labels each text block with the complete PDF path, product name, and product category, and then converts the first set of text blocks containing complete metadata from JSON format to CSV format. The model outputs a set of text blocks in CSV format containing complete metadata such as PDF path, product name, and product category.
[0140] The above describes the operations that may be performed before generating question-and-answer pairs. The following describes how to generate question-and-answer pairs.
[0141] In some implementations, step S130, which generates a question-and-answer pair corresponding to each text block in the first text block set, may include: for each text block in the first text block set, generating a question-and-answer pair associated with the metadata of the text block based on the text block and the metadata of the text block, wherein the metadata of the text block includes information indicating the source and / or attributes of the text block.
[0142] For details regarding metadata, please refer to the preceding introduction; it will not be repeated here. In this embodiment, prompt words can be constructed to associate the generated question-and-answer pairs for each text block with the metadata of that text block. The metadata of the text block can be input into the model, and prompt words can be constructed based on the requirement that the generated question-and-answer pairs are associated with the metadata of the text block. For example, prompt words can be designed based on requirements such as "the question contains the product name, the answer faithfully reflects the original text, and the associated chapter title." As a specific example, prompt words are designed based on the requirement that "the question contains the product name, the answer faithfully reflects the original text, and the associated chapter title." The large language model is called, and the prompt words and text blocks containing fields such as chapter title, product name, and text block are input. For each text block, the large language model generates question-and-answer pairs based on the input information and the prompt words, and the generated question-and-answer pairs are associated with the product name and chapter title.
[0143] By incorporating metadata such as product names and chapter titles into the generation of question-and-answer pairs, questions can be clearly linked to specific products and chapters, making it easier to accurately trace the original source of the answers.
[0144] In some implementations, generating question-and-answer pairs corresponding to text blocks may include: dividing the text block into multiple sub-text blocks if the text length of the text block is greater than or equal to a preset length, and generating question-and-answer pairs corresponding to each sub-text block.
[0145] In this embodiment, "the text length of a text block is greater than or equal to a preset length" may include one or more of the following: the number of characters in the main text of the text block is greater than or equal to a preset number; the number of words in the main text of the text block is greater than or equal to a preset number; the number of lines in the main text of the text block is greater than or equal to a preset number; the number of paragraphs in the main text of the text block is greater than or equal to a preset number. When the text length of the main text of a text block in the first text block set is greater than or equal to the preset length, the text block is further divided into multiple sub-text blocks, and a corresponding question-and-answer pair is generated for each of the multiple sub-text blocks. In this way, semantic coherence can be guaranteed, which is beneficial to ensuring the quality of the generated question-and-answer pairs.
[0146] As mentioned earlier, the body text of a text block can be contained in a text field. In this case, "the number of characters contained in the body text of the text block" can be understood as the number of characters contained in the text field. "the number of words contained in the body text of the text block" can be understood as the number of words contained in the text field. "the number of lines contained in the body text of the text block" can be understood as the number of lines contained in the text field. "the number of paragraphs contained in the body text of the text block" can be understood as the number of paragraphs contained in the text field.
[0147] For example, if the main text of a text block contains a number of characters greater than or equal to a preset number, the text length of the text block is considered to be greater than or equal to the preset length. As a concrete example, if the preset number of characters is set to 4000, a text block containing a main text of 4000 characters or more is considered to have a text length greater than or equal to the preset length. This text block is then divided into two sub-text blocks, and a question-and-answer pair is generated for each sub-text block.
[0148] When the length of a text block is greater than or equal to a preset length, two sub-text blocks obtained by dividing the text block can include the same content, thus ensuring semantic coherence. In other words, when dividing a text block into sub-text blocks, the later sub-text block can trace back to part of the content of the earlier sub-text block. For example, with a preset word count of 4000, for a text block containing 7000 words of the main text, the text block is divided into two sub-text blocks, and a question-and-answer pair is generated for each sub-text block. The first sub-text block contains the first 4000 words of the 7000 words, and the second sub-text block contains the last 4000 words of the 7000 words. The last 1000 words in the first sub-text block are also the first 1000 words in the second sub-text block. That is, the second sub-text block traces back to 1000 words in the first text block.
[0149] In some implementations, generating question-and-answer pairs corresponding to text blocks may include: generating a first number of question-and-answer pairs corresponding to text blocks, wherein the first number satisfies the following conditions: the first number is related to the text length of the text block, and / or the first number is greater than or equal to a preset number.
[0150] In some embodiments, generating question-and-answer pairs corresponding to a text block may include generating a first number of question-and-answer pairs corresponding to the text block, wherein the first number satisfies the condition that the first number is related to the text length of the text block. That is, the number of question-and-answer pairs generated for a text block is related to the text length of the text block. For example, for question-and-answer pairs with longer text lengths, a larger number of question-and-answer pairs are generated; for question-and-answer pairs with shorter text lengths, a smaller number of question-and-answer pairs are generated. As a specific example, for a text block containing 50 or fewer characters, one question-and-answer pair is generated; for a text block containing more than 50 characters but less than or equal to 100 characters, two question-and-answer pairs are generated; and for a text block containing more than 100 characters, 3-8 question-and-answer pairs are generated.
[0151] In other embodiments, generating question-and-answer pairs corresponding to text blocks may include generating a first number of question-and-answer pairs corresponding to the text blocks, wherein the first number satisfies the condition that the first number is greater than or equal to a preset number. That is, the number of question-and-answer pairs generated for each text block is at least a preset number. For example, if the preset number is 2, then at least 2 question-and-answer pairs are generated for each text block.
[0152] In some implementations, the method provided in this application embodiment may further include: performing verification on question-answer pairs generated for text blocks; if the number of question-answer pairs that pass the verification is less than a preset number, then regenerating question-answer pairs for the text block until the number of question-answer pairs that pass the verification reaches the preset number or the number of times question-answer pairs are generated for the text block reaches the preset number.
[0153] In other words, a "generation-verification-regeneration" mechanism can be adopted when generating question-answer pairs. In this embodiment, to distinguish it from subsequent verification, the verification here can be called preliminary verification. This verification may include one or more of the following: relevance verification and fidelity verification. Relevance verification can be understood as verifying the relevance between the questions and answers in the generated question-answer pairs. If the relevance between the questions and answers in a question-answer pair is high, the question-answer pair passes the relevance verification; if the relevance between the questions and answers in a question-answer pair is low, the question-answer pair fails the relevance verification. Fidelity verification can be understood as verifying whether the answers in the generated question-answer pairs are faithful to the content of the text block. If the answers in a question-answer pair are faithful to the content of the text block, the question-answer pair passes the fidelity verification; if the answers in a question-answer pair are not faithful to the content of the text block, the question-answer pair fails the fidelity verification. Through this verification mechanism, it can be ensured that question-answer pairs meeting the quantity and quality requirements are generated for each text block.
[0154] For example, the initial validation includes relevance and fidelity checks, with a preset quantity of 3 and a preset number of iterations of 5. After generating question-and-answer pairs for a text block, relevance and fidelity checks are performed on the pairs. If the number of generated question-and-answer pairs that pass both relevance and fidelity checks is greater than or equal to 3, then those pairs are used as the corresponding question-and-answer pairs for that text block. If the number of generated question-and-answer pairs that pass both relevance and fidelity checks is less than 3, then question-and-answer pairs are regenerated for that text block until the number of generated question-and-answer pairs that pass both relevance and fidelity checks reaches 3 or the number of times question-and-answer pairs are generated for that text block reaches 5.
[0155] In the embodiments of this application, the "preset quantity" and "preset number of times" mentioned above can be adjusted according to actual needs to adapt to different types or formats of portable documents or different application scenarios.
[0156] The initial validation process of "generation-validation-regeneration" can filter out model illusions and incorrect answers, which helps to improve the quality and stability of the final output question-answer pairs, so that the output question-answer pairs can be applied to high-reliability scenarios (such as enterprise knowledge bases and intelligent customer service).
[0157] To better understand step S130, see below. Figure 6 Step S130 will be explained in more detail with more specific examples. Figure 6 The example shown includes steps S601 to S6010.
[0158] In step S601, input a set of text blocks in CSV format.
[0159] The JSON-formatted text block set is converted into a CSV-formatted text block set and then input into the locally deployed large language model.
[0160] In step S602, the text length of the text block is determined.
[0161] For a text block, if the number of characters in the body text field is greater than or equal to 4000, the text block is determined to be a long text block; if the number of characters in the body text field is less than 50, the text block is determined to be a short text block; if the number of characters in the body text field is between 50 and 4000, the text block is determined to be a regular text block. For long text blocks, proceed to step S603; for short text blocks, proceed to step S604.
[0162] In step S603, the text block is split twice.
[0163] In this step, long text blocks are further segmented according to the rule of "tracing back the first 1000 words to ensure semantic coherence" to form sub-text blocks.
[0164] In step S604, it is determined that a single question-and-answer pair is generated.
[0165] In this step, it is determined that a question-and-answer pair is generated for each short text block, i.e., a single question-and-answer pair.
[0166] In step S605, a prompt word containing metadata is constructed.
[0167] In this step, prompts are designed based on requirements such as "the question contains the product name, the answer is a faithful copy of the original text, and the relevant chapter title".
[0168] In step S606, the large model is invoked to generate question-answer pairs.
[0169] In this step, the locally deployed large language model is invoked, and the metadata of prompt words and text blocks (chapter titles, product names, text blocks, etc.) is input into the model to generate question-answer pairs.
[0170] In steps S607 to S6010, the “generate-verify-regenerate” mechanism is executed.
[0171] In step S607, the quantity and quality of question-answer pairs are initially verified.
[0172] In this step, for each text block, preliminary relevance and fidelity checks are performed on the question-and-answer pairs corresponding to that text block.
[0173] In step S608, it is determined whether the question-answer pair meets the requirements.
[0174] In this step, for each text block, it is determined whether the corresponding question-and-answer pair meets the requirements. If the number of question-and-answer pairs that pass the relevance and fidelity checks for a given text block is less than 3, then the question-and-answer pair for that text block does not meet the requirements. If the number of question-and-answer pairs that pass the relevance and fidelity checks for a given text block is greater than or equal to 3, then the question-and-answer pair for that text block meets the requirements.
[0175] For a text block whose question-answer pair meets the requirements, execute step S6010, which outputs the question-answer pair generated for that text block.
[0176] For text blocks where the question-answer pair does not meet the requirements, first return to step S606, that is, regenerate the question-answer pair for the text block; then, execute step S608 again to determine whether the newly generated question-answer pair for the text block meets the requirements. If it meets the requirements, execute step S6010; if it does not meet the requirements, execute step S609.
[0177] In step S609, it is determined whether the number of retries is less than 5.
[0178] In this step, it is determined whether the number of question-answer pairs generated for a certain text block is less than 5. If the number of question-answer pairs generated for the text block is less than 5, then return to execute steps S606 and S608; if the number of question-answer pairs generated for the text block is not less than 5, then execute step S6010.
[0179] By going through steps S607 to S609, it can be guaranteed that at least 2 question-answer pairs will be generated for each text block and 3-8 question-answer pairs will be generated for regular text blocks.
[0180] In step S6010, the initial question-answer pair set is output.
[0181] In this step, the initial set of question-answer pairs, grouped by text blocks, is output, along with the associated complete metadata.
[0182] Through steps S601 to S6010, metadata can be used to improve the product relevance and chapter relevance of question-and-answer pairs, and text length adaptation and preliminary verification can be used to ensure the quantity and basic quality of question-and-answer pairs.
[0183] The above provides a detailed description of step S130.
[0184] In some implementations, after generating the question-and-answer pairs corresponding to the text blocks, the method provided in this application embodiment may further include: performing quantitative and qualitative verification on the question-and-answer pairs corresponding to the text blocks.
[0185] Performing quantitative verification on question-and-answer pairs can be understood as evaluating them using quantifiable, pre-defined metrics. For example, multiple quantifiable metrics can be set, and for each question-and-answer pair, a score can be calculated for each of these metrics. A comprehensive score is then determined based on these scores. This comprehensive score serves as the result of the quantitative verification. For instance, the average score across multiple metrics can be used as the comprehensive score. This average can be an arithmetic mean or a weighted average. Furthermore, the results of the quantitative verification can be categorized according to the comprehensive score. For example, the results can be divided into four levels: excellent, good, average, and poor, based on the comprehensive score.
[0186] After performing quantitative verification on question-and-answer pairs, those that pass the verification can be retained, while those that fail can be removed. The pass / fail status of a question-and-answer pair can be determined based on the comprehensive score mentioned above. For example, a scoring threshold can be set; if the comprehensive score of the question-and-answer pair is greater than or equal to the threshold, the pair passes the quantitative verification; if the comprehensive score is less than the threshold, the pair fails. Alternatively, a passing level can be set; if the level corresponding to the comprehensive score of the question-and-answer pair is higher than the preset level, the pair passes the verification; if the level corresponding to the comprehensive score is equal to or lower than the preset level, the pair fails.
[0187] As a concrete example, three indicators—"question-answer relevance," "answer-original text fidelity," and "answer completeness"—are used as preset indicators. For each question-answer pair, scores for these three indicators are calculated separately, and the average of the three scores is used as the comprehensive score for quantitative verification. Furthermore, a quantitative verification result with a comprehensive score greater than or equal to 0.85 is classified as excellent; a comprehensive score less than 0.85 but greater than or equal to 0.7 is classified as good; a comprehensive score less than 0.7 but greater than or equal to 0.5 is classified as average; and a comprehensive score less than 0.5 is classified as poor. Question-answer pairs with quantitative results classified as excellent, good, or average are retained, while those with a poor quantitative result are removed.
[0188] The quantization verification described above can be performed by a model. This model includes semantic similarity models, such as pre-trained semantic similarity models. Semantic similarity models include, but are not limited to: the Qwen3 32B model (version 3, 32 billion parameters), the BERT series of models, and the GPT-4 generative pre-trained transformer model. As a concrete example, by inputting question-answer pairs into a pre-trained semantic similarity model, the model outputs scores for three metrics: "question-answer relevance," "answer-text fidelity," and "answer completeness."
[0189] Performing qualitative validation on question-and-answer pairs can be understood as evaluating their quality in a non-quantitative way. Quality grades can be used to represent the quality of question-and-answer pairs determined through qualitative validation. After determining the quality of question-and-answer pairs through qualitative validation, qualified pairs can be retained and unqualified pairs can be discarded. For example, multiple quality grades can be set, with a specific quality grade designated as qualified and other quality grades as unqualified. If the qualitative validation determines that the quality of a question-and-answer pair corresponds to the specific quality grade, then that pair is retained; otherwise, it is discarded. As a concrete example, three quality grades can be set: Good, Uncertain, and Poor, with Good being the qualified grade and Uncertain and Poor being the unqualified grades. If the qualitative test results indicate that a question-and-answer pair is of good quality, then the pair is retained; if the qualitative test results indicate that a question-and-answer pair is of undetermined or poor quality, then the pair is removed.
[0190] Qualitative validation can be performed by a large language model. This large language model can include one deployed locally. A large language model can include the Qwen3 32B model. For example, question-answer pairs are input into a locally deployed large language model, which outputs the qualitative validation result (e.g., quality grade) for each question-answer pair.
[0191] In some embodiments, after determining the quality of question-answer pairs through qualitative verification, the method provided in this application may further include: manually reviewing the results of the qualitative verification. After manual review, based on the quality of the question-answer pairs, retaining those that meet the quality standards and removing those that do not.
[0192] In the embodiments of this application, performing quantitative and qualitative verification on question-answer pairs can be referred to as multi-level verification. Through multi-level verification, erroneous or non-compliant question-answer pairs can be filtered out, which helps to ensure the quality of the final question-answer pairs.
[0193] When performing multi-level verification, quantitative verification can be performed first, followed by qualitative verification. In this case, quantitative verification can be referred to as the first round of verification, and qualitative verification as the second round. Alternatively, qualitative verification can be performed first, followed by quantitative verification. In this case, qualitative verification can be referred to as the first round of verification, and quantitative verification as the second round.
[0194] As a concrete example, the initial set of question-answer pairs mentioned above is first input into the pre-trained semantic similarity model, which performs the first round of quantization verification. For each question-answer pair, the semantic similarity model calculates scores for three metrics: "question-answer relevance," "answer-text fidelity," and "answer completeness," and calculates the average of these scores as the comprehensive quantization verification score. A comprehensive score greater than or equal to 0.85 corresponds to an excellent quantization verification result; a comprehensive score less than 0.85 but greater than or equal to 0.7 corresponds to a good quantization verification result; a comprehensive score less than 0.7 but greater than or equal to 0.5 corresponds to a fair quantization verification result; and a comprehensive score less than 0.5 corresponds to a poor quantization verification result. The semantic similarity model discards question-answer pairs with a poor quantization rating and outputs question-answer pairs with excellent, good, and fair quantization ratings. Then, the question-answer pairs output by the semantic similarity model are input into a locally deployed large language model, which performs the second round of qualitative verification. After performing qualitative validation, the large language model outputs evaluation results for each input question-answer pair. The evaluation results are categorized into three classes: Class 1 (good quality), Class 2 (quality undetermined), and Class 3 (poor quality). After the large language model outputs the evaluation results, they are manually reviewed, and unqualified evaluation results are removed.
[0195] Through multi-level verification of "quantitative verification + qualitative verification", model illusions and incorrect answers are effectively filtered out, which helps to improve the quality and stability of question-answer pairs. This makes it easier to apply question-answer pairs directly to high-reliability scenarios (such as enterprise knowledge bases and intelligent customer service).
[0196] In the embodiments of this application, the multi-level verification of "quantitative verification + qualitative verification" can be used alone or in combination with the preliminary verification described above. For example, for the initial question-answer pair set, the preliminary verification is performed first, followed by the multi-level verification of "quantitative verification + qualitative verification". In this way, the quality stability of question-answer pairs is further improved through a dual quality control system, further promoting the application of question-answer pairs in high-reliability scenarios.
[0197] The multi-level verification provided in the embodiments of this application has been described above. In the embodiments of this application, semantic deduplication can also be performed on the generated question-answer pairs.
[0198] In some implementations, after generating question-answer pairs corresponding to text blocks, the method provided in this application embodiment may further include: determining question-answer pairs with similar questions among the question-answer pairs corresponding to text blocks; calculating the semantic similarity of answers among question-answer pairs with similar questions; determining question-answer pairs with similar answers from question-answer pairs with similar questions based on the semantic similarity of answers; and removing one or more question-answer pairs from question-answer pairs with similar answers.
[0199] The process of "determining similar question-answer pairs within a question-answer pair corresponding to a text block" can be called question clustering. In some embodiments, "determining similar question-answer pairs within a question-answer pair corresponding to a text block" may include: determining similar question-answer pairs within a question-answer pair corresponding to a text block based on the feature vectors of the questions in the question-answer pair. For example, based on the feature vectors, question-answer pairs corresponding to the same text block are divided into a predetermined number of clusters, where the questions in the question-answer pairs within each cluster are similar.
[0200] Question clustering can be implemented by a model based on clustering algorithms. For example, questions and answers corresponding to the same text block are input into the model, which extracts feature vectors for each question and uses a clustering algorithm to cluster similar questions based on these feature vectors. This model can include semantic embedding models. Clustering algorithms include, but are not limited to, density-based spatial clustering of applications with noise (DBSCAN) and K-Means clustering. DBSCAN is superior in noise filtering, especially suitable for handling massive amounts of question-answer pairs. When the number of question-answer pairs is large, DBSCAN should be preferred. K-Means can divide question-answer pairs corresponding to the same text block into a predetermined number of clusters, reducing the difficulty of parameter tuning in the question clustering process. K-Means is particularly suitable for scenarios with a small number of question-answer pairs.
[0201] After identifying question-answer pairs with similar questions, the semantic similarity of the answers in these pairs is calculated. Based on this semantic similarity, question-answer pairs with similar answers are then selected from those pairs. A semantic similarity threshold can be set, and these thresholds can be used to determine which pairs have similar answers. Alternatively, question-answer pairs with answers whose semantic similarity is greater than the threshold can be identified as having similar answers. In practical applications, the semantic similarity threshold can be adjusted to accommodate different types of portable documents.
[0202] After identifying question-answer pairs with similar answers, remove one or more such pairs. For example, you can keep the question-answer pair with the longest and most complete answer among the similar pairs and remove the rest.
[0203] By semantic deduplication, redundant question-answer pairs are eliminated, which can improve the purity and quality stability of question-answer pairs, making it easier to apply question-answer pairs to high-reliability scenarios (such as enterprise knowledge bases and intelligent customer service).
[0204] In the embodiments of this application, the multi-level validation and semantic deduplication described above can be used individually or in combination. For example, multi-level validation can be performed first, followed by semantic deduplication. This effectively filters out model illusions and incorrect answers, while also eliminating semantically redundant content, thus improving the quality stability of question-answer pairs. Alternatively, preliminary validation can be performed on the initial question-answer pair, followed by multi-level validation, and then semantic deduplication. This triple quality control mechanism effectively filters out model illusions and incorrect answers while eliminating semantically redundant content, maximizing the quality stability of question-answer pairs and enabling them to be directly applied to high-reliability scenarios (such as enterprise knowledge bases and intelligent customer service).
[0205] The following will illustrate this with more specific examples.
[0206] As a concrete example, a first round of quantitative verification and a second round of qualitative verification are performed. A more detailed description of the first and second rounds of quantitative and qualitative verification can be found in the previous example, and will not be repeated here. After the second round of qualitative verification, semantic deduplication is performed. First, the question-answer pairs that have passed the multi-level verification are grouped according to text blocks, with question-answer pairs corresponding to the same text block grouped together. Then, the verified question-answer pairs are input into the semantic embedded model. The semantic embedded model extracts the feature vectors of the questions for the same group of question-answer pairs and uses the DBSCAN algorithm to cluster question-answer pairs with similar questions within the same group, where question-answer pairs within the same cluster have similar questions. Next, the semantic similarity of the answers for question-answer pairs within the same cluster is calculated. Finally, the question-answer pair corresponding to the longest and most complete answer is retained. In this example, multi-level verification filters out erroneous question-answer pairs, and semantic deduplication improves the purity of the question-answer pairs, thus ensuring the quality of the question-answer pairs.
[0207] Step S130 mentions that for each text block in the first text block set, a question-and-answer pair corresponding to the text block is generated. In the embodiments of this application, in addition to performing the secondary checksum and / or semantic deduplication described above on the generated question-and-answer pairs, other operations can also be performed on the generated question-and-answer pairs.
[0208] In some implementations, the method provided in this application embodiment may further include: classifying the question-answer pairs corresponding to the first text block set.
[0209] It should be understood that the meaning of the question-and-answer pairs corresponding to the first text block mentioned here differs depending on the implementation. For example, if no secondary validation or semantic deduplication is performed after generating the initial question-and-answer pair set, then the "question-and-answer pairs corresponding to the first text block set" mentioned here can be understood as the question-and-answer pairs in the initial question-and-answer pair set. As another example, if secondary validation is performed after generating the initial question-and-answer pair set but semantic deduplication is not performed, then the "question-and-answer pairs corresponding to the first text block set" mentioned here can be understood as the question-and-answer pairs retained after performing secondary validation on the initial question-and-answer pair set. Again, if semantic deduplication is performed after generating the initial question-and-answer pair set but secondary validation is not performed, then the "question-and-answer pairs corresponding to the first text block set" mentioned here can be understood as the question-and-answer pairs retained after performing semantic deduplication on the initial question-and-answer pair set. Finally, if secondary validation and semantic deduplication are performed after generating the initial question-and-answer pair set, then the "question-and-answer pairs corresponding to the first text block set" mentioned here can be understood as the question-and-answer pairs retained after performing secondary validation and semantic deduplication on the initial question-and-answer pair set.
[0210] Question-answer pairs can be categorized by labeling them. These labels are called category tags. Multiple levels of tags can be preset and used to label each question-answer pair. For example, multiple first-level tags and corresponding second-level tags can be preset. First-level tags include, but are not limited to: intellectual property, main Android application package (APK), hardware information, product information, vehicle software, service guide, and safety information. Furthermore, the maximum number of tags per level for a question-answer pair can be set. For example, the maximum number of second-level tags per question-answer pair can be set to three.
[0211] A model can be used to classify question-answer pairs. For example, inputting question-answer pairs into the Qwen3 32B model will output the classification and labeling results. After the model classifies the question-answer pairs, the classification results can be manually verified. For unclassified question-answer pairs, they can be manually classified. For incorrectly classified question-answer pairs, the classification results can be manually corrected. For example, after the model outputs the classification and labeling results for the question-answer pairs, the classification and labeling results can be manually verified. For question-answer pairs without labels, labels can be manually added. For question-answer pairs with incorrectly labeled labels, the labels can be manually corrected.
[0212] By categorizing question-and-answer pairs, classified management of these pairs is achieved, facilitating subsequent precise retrieval and application. For example, categorizing question-and-answer pairs can meet various application needs such as enterprise knowledge base construction, intelligent question-and-answer system training, and customer service robot deployment.
[0213] In some implementations, after classifying the question-answer pairs corresponding to the first text block set, the method provided in this application embodiment may further include: performing semantic deduplication on question-answer pairs of a specific type in the question-answer pairs corresponding to the first text block set.
[0214] The meaning of "question-answer pairs corresponding to the first set of text blocks" can be found in the previous description, and will not be repeated here.
[0215] The specific types of question-answer pairs mentioned here can include question-answer pairs labeled with specific tags. For example, specific types of question-answer pairs include one or more of the following: question-answer pairs labeled with "Intellectual Property", question-answer pairs labeled with "Service Guide", and question-answer pairs labeled with "Security Information". To distinguish it from the semantic deduplication described above, the semantic deduplication performed on specific types of question-answer pairs in this embodiment can be referred to as secondary deduplication.
[0216] In some implementations, the method provided in this application embodiment may further include: storing question-answer pairs corresponding to the first text block set.
[0217] The meaning of "question-answer pairs corresponding to the first set of text blocks" can be found in the previous description, and will not be repeated here.
[0218] The question-and-answer pairs corresponding to the first set of text blocks can be stored in a specific database, such as MySQL or MongoDB. Storing the question-and-answer pairs in a specific database enables standardized storage.
[0219] When storing question-and-answer pairs, an index can be created for each pair. The index can use a specific format; for example, a question-and-answer pair index might include the following fields: question-and-answer pair ID, question, answer, storage path of the portable document (e.g., PDF path), source document name, chapter title, chapter page number, text block, product name, product analogy, primary category, secondary category, and generation time. Indexing question-and-answer pairs improves subsequent retrieval efficiency and can meet various application needs such as enterprise knowledge base construction, intelligent question-and-answer system training, and customer service robot deployment.
[0220] In some implementations, the method provided in this application embodiment may further include: exporting question-answer pairs corresponding to the first text block set in multiple formats.
[0221] The meaning of "question-answer pairs corresponding to the first set of text blocks" can be found in the previous description, and will not be repeated here.
[0222] In this embodiment, exporting question-and-answer pairs in multiple formats is supported. For example, question-and-answer pairs can be exported in JSON format. Another example is exporting question-and-answer pairs in CSV format. Exporting question-and-answer pairs in multiple formats allows them to be applied to different application scenarios, meeting various application needs such as enterprise knowledge base construction, intelligent question-and-answer system training, and customer service robot deployment.
[0223] The above text combined Figures 1 to 6 This application describes in detail embodiments of a method for generating question-answer pairs based on portable documents. In these embodiments, the entire process of "portable document parsing - structured segmentation - question-answer pair generation - multi-level validation - semantic deduplication - question-answer pair storage" can be automated using a model. Compared to traditional solutions, the end-to-end automated architecture in these embodiments replaces a significant amount of manual document organization, segmentation, and question-answer pair writing, greatly reducing labor and time costs and improving knowledge extraction efficiency.
[0224] Furthermore, in this embodiment, during the structured parsing of portable documents, non-text content such as images and tables is retained while retaining the text content, thus avoiding the loss of key information in the final generated question-and-answer pairs. Precise segmentation based on the heading hierarchy and content relevance of portable documents ensures the semantic integrity of text blocks. Linking metadata to question-and-answer pairs during generation enhances their document relevance and traceability. A multi-level verification mechanism combining quantitative and qualitative checks, along with / or semantic deduplication, significantly improves the accuracy, diversity, and stability of question-and-answer pairs, reduces manual intervention costs, and meets the high reliability requirements of enterprise knowledge base construction.
[0225] The apparatus embodiments of this application are described in detail below. It should be understood that the description of the apparatus embodiments corresponds to the description of the method embodiments. Therefore, any parts not described in detail can be referred to the foregoing description of the method embodiments.
[0226] This application also provides an apparatus 700 for generating question-and-answer pairs based on portable documents, which is used to perform the aforementioned method for generating question-and-answer pairs based on portable documents. Figure 7 A schematic diagram of the structure of the apparatus 700 for generating question-and-answer pairs based on portable documents provided in an embodiment of this application is shown.
[0227] like Figure 7 As shown, the apparatus 700 for generating question-and-answer pairs based on portable documents includes a construction module 710, a segmentation module 720, and a generation module 730.
[0228] The construction module 710 is used to construct a structured dataset corresponding to the portable document, wherein the structured dataset includes the title hierarchy of the portable document.
[0229] The segmentation module 720 is used to segment the structured dataset based on the title level to obtain a first set of text blocks.
[0230] The generation module 730 is used to generate a question-answer pair corresponding to each text block in the first text block set, wherein each question-answer pair consists of a question and the answer corresponding to the question.
[0231] This application also provides a system 800 for generating question-and-answer pairs based on portable documents. For example... Figure 8 As shown, the system 800 for generating question-and-answer pairs based on portable documents includes: a memory 810 for storing code; and a processor 820 for executing the code stored in the memory to perform the method for generating question-and-answer pairs based on portable documents as described in any of the preceding embodiments.
[0232] This application also provides a computer-readable storage medium. The computer-readable storage medium stores program code that can be used to execute the method for generating question-answer pairs based on portable documents in any of the preceding embodiments.
[0233] This application also provides a computer program product. This computer program product includes a method for performing the portable document-based question-and-answer pair generation method described in any of the preceding embodiments.
[0234] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0235] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above-described functional configurations can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0236] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, disks, or optical discs. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0237] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0238] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0239] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0240] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may fall into one place or be distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0241] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for generating question-answer pairs based on portable documents, characterized in that, include: Construct a structured dataset corresponding to the portable document, wherein the structured dataset includes the title hierarchy of the portable document; The structured dataset is divided based on the title hierarchy to obtain a first set of text blocks; Based on each text block in the first set of text blocks, a question-and-answer pair corresponding to the text block is generated, wherein each question-and-answer pair consists of a question and the answer corresponding to the question.
2. The method according to claim 1, characterized in that, The structured dataset contains non-textual information from the portable document.
3. The method according to claim 1, characterized in that, The first text block set is obtained by partitioning the structured dataset based on the title level, including: The structured dataset is divided with the title level as the boundary to obtain a set of basic text blocks; For each text block in the basic text block set, calculate the semantic similarity of the text within that text block; If the basic text block set does not include a text block that satisfies the first condition, then the basic text block set is determined as the first text block set. If the basic text block set includes text blocks that satisfy the first condition, then the structured dataset is repartitioned to obtain the first text block set; The first condition includes: the semantic similarity of the text block includes semantic similarity below a preset threshold.
4. The method according to claim 1, characterized in that, The first text block set is obtained by partitioning the structured dataset based on the title level, including: The structured dataset is divided with the title level as the boundary to obtain a set of basic text blocks; The first text block and the second text block in the basic text block set are merged to obtain the first text block set; wherein the first text block and the second text block are adjacent, and the first text block or the second text block is an empty text block.
5. The method according to claim 1, characterized in that, The step of generating a question-answer pair corresponding to each text block in the first text block set includes: For each text block in the first set of text blocks, a question-answer pair associated with the metadata of the text block is generated based on the text block and the metadata of the text block, wherein the metadata of the text block includes information indicating the source of the text block and / or information indicating the attributes of the text block.
6. The method according to claim 1, characterized in that, The generation of the question-answer pair corresponding to the text block includes one or more of the following: If the text length of the text block is greater than or equal to a preset length, the text block is divided into multiple sub-text blocks, and a question-and-answer pair corresponding to each of the multiple sub-text blocks is generated; Generate a first number of question-answer pairs corresponding to the text block, wherein the first number satisfies the following conditions: the first number is related to the text length of the text block, and / or the first number is greater than or equal to a preset number.
7. The method according to claim 1, characterized in that, After generating the question-answer pair corresponding to the text block, the method further includes: Quantitative and qualitative checks are performed on the question-and-answer pairs.
8. The method according to any one of claims 1-7, characterized in that, After generating the question-answer pair corresponding to the text block, the method further includes: Identify question-answer pairs with similar questions within the question-answer pairs corresponding to the text block; Calculate the semantic similarity of answers in question-answer pairs that are similar to the given question; Based on the semantic similarity, determine question-answer pairs with similar answers from question-answer pairs with similar questions; Remove one or more question-answer pairs from those with similar answers.
9. A device for generating question-and-answer pairs based on portable documents, characterized in that, The device includes: A construction module is used to construct a structured dataset corresponding to the portable document, wherein the structured dataset includes the title hierarchy of the portable document; The segmentation module is used to segment the structured dataset based on the title level to obtain a first set of text blocks; The generation module is used to generate a question-answer pair corresponding to each text block in the first text block set, wherein each question-answer pair consists of a question and the answer corresponding to the question.
10. A system for generating question-and-answer pairs based on portable documents, characterized in that, include: Memory, used to store code; A processor for executing the code stored in the memory to perform the method as described in any one of claims 1 to 8.