A method for extracting question-and-answer pairs, a question-and-answer method, a system, a device, and a medium

By determining keywords in the document for paragraph segmentation and strong association processing, combining narrow domain and general question-and-answer pair extraction model, the problem of predetermined vocabulary list limitation in the prior art is solved, and the generation of full coverage high-quality question-and-answer pairs is achieved.

CN120104719BActive Publication Date: 2025-07-29CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510585307.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-29
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing Q&A pair extraction methods rely on the predetermined vocabulary list, resulting in the inability to extract Q&A pairs that are not related to the predetermined vocabulary list, which has limitations.

Method used

By determining the keywords of the original document, paragraph segmentation and keyword strong correlation processing are performed, the Q&A pair extraction model is used to generate a full-coverage high-quality Q&A pair pair, and the narrow domain and general Q&A pair extraction model are combined to process different types of text paragraphs.

Benefits of technology

It realizes the full coverage of high-quality Q&A pair extraction of documents, ensures that each paragraph has corresponding keywords, and improves the accuracy and coverage of Q&A pairs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104719B_ABST
    Figure CN120104719B_ABST
Patent Text Reader

Abstract

The present application discloses a method for extracting question-and-answer pairs, a question-and-answer method, a system, a device, and a medium, which relate to the technical field of intelligent question-and-answer of large models. The method includes: determining keywords corresponding to the original document; performing paragraph segmentation on the original document with corresponding keywords to obtain each first text paragraph; performing keyword strong association processing on the first text paragraphs according to the keywords corresponding to the original document to obtain target text paragraphs; extracting question-and-answer pairs from the target text paragraphs through a question-and-answer pair extraction model to obtain corresponding question-and-answer pairs. Through this method, high-quality question-and-answer pair extraction covering the entire document is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent question - answering of large models, and particularly to a method for extracting question - answer pairs, a question - answering method, a system, a device, and a medium. Background Art

[0002] With the breakthrough of natural language processing (NLP) technology, the document question - answering system based on large models has gradually become the core tool for improving information retrieval efficiency. The most important core part of this question - answering system is the determination of question - answer pairs. The existing methods for extracting question - answer pairs usually associate with a pre - defined vocabulary related to the business and the business documents to narrow the extraction scope of question - answering knowledge for the purpose of accurately extracting question - answer pairs. However, due to the existence of the pre - defined vocabulary, the extraction effect has limitations, and it can only extract question - answer pairs strongly associated with the pre - defined vocabulary, and it is impossible to extract question - answer pairs for some business documents without a pre - defined vocabulary. Summary of the Invention

[0003] In view of this, the present invention provides a method for extracting question - answer pairs, a question - answering method, a system, a device, and a medium, aiming to perform high - quality extraction of question - answer pairs with full coverage of documents.

[0004] To achieve the above object, in the first aspect of the embodiments of the present application, a method for extracting question - answer pairs is provided, and the method includes:

[0005] Determine the keywords corresponding to the original document;

[0006] Perform paragraph segmentation on the original document with corresponding keywords to obtain each first text paragraph;

[0007] Perform strong keyword association processing on the first text paragraphs according to the keywords corresponding to the original document to obtain target text paragraphs;

[0008] Extract question - answer pairs for the target text paragraphs through a question - answer pair extraction model to obtain corresponding question - answer pairs.

[0009] Optionally, performing strong keyword association processing on the first text paragraphs according to the keywords corresponding to the original document to obtain target text paragraphs includes:

[0010] Extract the exact keywords in the first text paragraphs;

[0011] Perform strong keyword association processing on the first text paragraphs through the exact keywords of the first text paragraphs to obtain corresponding target text paragraphs;

[0012] Perform strong keyword association processing on the first text paragraphs that have not obtained corresponding exact keywords through the keywords corresponding to the original document to obtain corresponding target text paragraphs.

[0013] Optionally, the method further includes:

[0014] Performing paragraph segmentation on the original document without corresponding keywords to obtain respective target text paragraphs;

[0015] Performing Q&A pair extraction on the target text paragraphs by the Q&A pair extraction model to obtain corresponding Q&A pairs, including:

[0016] Determining keyword features of the target text paragraphs;

[0017] Performing Q&A pair extraction on the target text paragraphs with keyword features by a narrow-domain Q&A pair extraction model to obtain corresponding first Q&A pairs;

[0018] Performing Q&A pair extraction on the target text paragraphs without keyword features by a general Q&A pair extraction model to obtain corresponding second Q&A pairs.

[0019] Optionally, the method further includes:

[0020] Determining a first similarity between the first Q&A pairs and the keyword features in the corresponding target text paragraphs;

[0021] Filtering the first Q&A pairs according to the first similarity;

[0022] Determining a second similarity between the second Q&A pairs and historical second Q&A pairs;

[0023] Determining second Q&A pairs to be filtered from the second Q&A pairs and the historical second Q&A pairs according to the second similarity.

[0024] Optionally, performing paragraph segmentation on the original document with corresponding keywords to obtain respective first text paragraphs, including:

[0025] Performing paragraph segmentation on the original document with corresponding keywords to obtain respective initial text paragraphs;

[0026] Performing recognition processing on the obtained initial text paragraphs to determine whether type features corresponding to various paragraph types are present in the initial text paragraphs;

[0027] Determining the paragraph type corresponding to the type features possessed by the initial text paragraphs as the type of the initial text paragraphs, where the various paragraph types at least include enumerated paragraphs, table paragraphs, and picture paragraphs;

[0028] Performing content extraction on the initial text paragraphs by an extraction rule corresponding to the type of the initial text paragraphs to determine first text paragraphs corresponding to the initial text paragraphs.

[0029] Optionally, content extraction is performed on the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph to determine a first text paragraph corresponding to the initial text paragraph, including:

[0030] Content extraction is performed on the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph to obtain a text paragraph to be filtered corresponding to the initial text paragraph;

[0031] According to the relationship between the number of tokens of the text paragraph to be filtered and the token threshold, irrelevant field filtering is performed on the text paragraph to be filtered to obtain a first text paragraph corresponding to the initial text paragraph.

[0032] Optionally, content extraction is performed on the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph to obtain a text paragraph to be filtered corresponding to the initial text paragraph, including:

[0033] In the case where the type of the initial text paragraph includes an enumerated paragraph, each enumerated text is extracted, and each extracted enumerated text is associated with the main body of the initial text paragraph;

[0034] In the case where the type of the initial text paragraph includes a table-like paragraph, the row data of the table is extracted, and the column name of the table is determined as the attribute of the corresponding row data;

[0035] The extracted row data and row data attributes are associated with the main body of the initial text paragraph;

[0036] In the case where the type of the initial text paragraph includes a picture-like paragraph, the network address of the picture is extracted and associated with the main body of the initial text paragraph;

[0037] After the extraction and main body association of each type of data in the initial text paragraph are completed, a text paragraph to be filtered corresponding to the initial text paragraph is obtained.

[0038] Optionally, training the narrow-domain Q&A pair extraction model includes:

[0039] The to-be-annotated sample paragraph is annotated with the standard Q&A pair of the to-be-annotated sample paragraph to obtain a training sample paragraph, and the to-be-annotated sample paragraph is the target text paragraph after strong association processing;

[0040] The initial model is trained with the training sample paragraph, and the performance of the trained initial model is evaluated with the standard Q&A pair of the training sample paragraph;

[0041] According to the evaluation result, it is determined whether a qualified narrow-domain Q&A pair extraction model is trained.

[0042] A second aspect of the embodiments of the present application provides a question-answering method based on question-answer pairs, the method comprising:

[0043] Extract keywords from the question text;

[0044] Determining, based on the extracted keywords, a first target question-answer pair corresponding to the question text from a first question-answer pair set, where the first question-answer pair in the first question-answer pair set is extracted using the question-answer pair extraction method described in the first aspect of the present application;

[0045] Answering the question text based on the first target question-answer pair;

[0046] Determining, based on a question text from which corresponding keywords have not been extracted, a second target question-answer pair corresponding to the question text from a second set of question-answer pairs, wherein the second question-answer pair in the second set of question-answer pairs is extracted using the method for extracting question-answer pairs described in the first aspect of the present application;

[0047] Based on the second target question-answer pair, answer the question text.

[0048] Optionally, before determining a second target question-answer pair corresponding to the question text from the second question-answer pair set based on the question text from which no corresponding keywords are extracted, the method further includes:

[0049] Determining, based on the question text from which no corresponding keyword is extracted, a target keyword corresponding to the terminal that initiated the question text;

[0050] Constructing a new question text according to the question text and the determined target keywords;

[0051] According to the new question text, determining a first target question-answer pair corresponding to the question text from a first question-answer pair set;

[0052] The step of determining, based on the question text from which no corresponding keywords are extracted, a second target question-answer pair corresponding to the question text from the second question-answer pair set includes:

[0053] If the target keyword corresponding to the terminal that initiated the question text is not determined, a second target question-answer pair corresponding to the question text is determined from the second question-answer pair set based on the question text from which the corresponding keyword is not extracted.

[0054] A third aspect of an embodiment of the present application provides a system for extracting question-answer pairs, the system comprising:

[0055] A first keyword determination module is used to determine keywords corresponding to the original document;

[0056] A paragraph segmentation module, configured to segment an original document with corresponding keywords to obtain each first text paragraph;

[0057] A strong association module, configured to perform strong keyword association processing on the first text paragraph according to the keywords corresponding to the original document to obtain a target text paragraph;

[0058] A question-and-answer pair extraction module, configured to extract question-and-answer pairs from the target text paragraph through a question-and-answer pair extraction model to obtain corresponding question-and-answer pairs.

[0059] A fourth aspect of the embodiments of the present application provides a question-and-answer system based on question-and-answer pairs. The system includes:

[0060] A second keyword extraction module, configured to extract keywords in the question text;

[0061] A first target question-and-answer pair determination module, configured to determine a first target question-and-answer pair corresponding to the question text from a first question-and-answer pair set according to the extracted keywords. The first question-and-answer pairs in the first question-and-answer pair set are extracted by a question-and-answer pair extraction system provided in the third aspect of the present application;

[0062] A first answer module, configured to answer the question text based on the first target question-and-answer pair;

[0063] A second target question-and-answer pair determination module, configured to determine a second target question-and-answer pair corresponding to the question text from a second question-and-answer pair set according to the question text for which no corresponding keywords are extracted. The second question-and-answer pairs in the second question-and-answer pair set are extracted by the question-and-answer pair extraction system described in the third aspect of the present application;

[0064] A second answer module, configured to answer the question text based on the second target question-and-answer pair.

[0065] A fifth aspect of the embodiments of the present application provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and running on the processor. When the computer program is executed by the processor, the steps in a method for extracting question-and-answer pairs as described in any item of the first aspect are implemented.

[0066] A sixth aspect of the embodiments of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in a method for extracting question-and-answer pairs as described in any item of the first aspect are implemented.

[0067] Beneficial effects

[0068] Adopt a method for extracting question - answer pairs provided by the present application. The method includes: determining keywords corresponding to the original document; performing paragraph segmentation on the original document with corresponding keywords to obtain each first text paragraph; performing strong keyword association processing on the first text paragraph according to the keywords corresponding to the original document to obtain target text paragraphs; and extracting question - answer pairs from the target text paragraphs through a question - answer pair extraction model to obtain corresponding question - answer pairs.

[0069] By performing paragraph segmentation on the original document with keywords according to a preset segmentation method and strongly associating the obtained first text paragraphs with the keywords, it is ensured that each paragraph in the original document has corresponding keywords. Furthermore, when extracting questions and answers from paragraphs, corresponding question extraction can be performed on each paragraph. That is, high - quality question - answer pair extraction covering the entire original document is achieved through the question extraction method of the present application. Brief Description of the Drawings

[0070] Figure 1 is a flowchart of a method for extracting question - answer pairs proposed in an embodiment of the present application;

[0071] Figure 2 is another flowchart of a method for extracting question - answer pairs provided in an embodiment of the present application;

[0072] Figure 3 is a flowchart of a question - answering method based on question - answer pairs provided in an embodiment of the present application;

[0073] Figure 4 is a schematic diagram of a system for extracting question - answer pairs provided in an embodiment of the present application;

[0074] Figure 5 is a schematic diagram of a question - answering system based on question - answer pairs provided in an embodiment of the present application. Detailed Embodiments

[0075] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for explaining the present invention, rather than for limiting the protection scope of the present invention.

[0076] Refer to Figure 1 and Figure 2 , Figure 1 is a flowchart of a method for extracting question - answer pairs proposed in an embodiment of the present application, Figure 2It is another flowchart of a method for extracting question-and-answer pairs, and the method includes the following steps:

[0077] S11: Determine the keywords corresponding to the original document.

[0078] In this embodiment, the original document refers to a series of technical documents about products written by engineers, such as product manuals, technical specifications, product instructions, etc., and the original document may include product models, versions, etc. When the product is a vehicle, the original document includes the vehicle series, model, and vehicle version, etc.

[0079] In order to generate a large number of accurate question-and-answer pairs based on the original document, this embodiment sets keywords for the original document. By determining whether the original document contains the keyword, more accurate question-and-answer pairs can be generated based on the original document. The keyword can be specified by the user or pre-inserted by the user in the original document. When the product in the original document is a vehicle, the keyword corresponding to the original document can be the vehicle series representing a larger scope in the original document.

[0080] S12: Segment the original document with corresponding keywords to obtain each first text paragraph.

[0081] In this embodiment, if the original document has corresponding keywords, it indicates that the content in the original document is related to the keywords, and the original document can be further processed to generate question-and-answer pairs related to the keywords subsequently. Specifically, in the case where the original document has corresponding keywords, the original document is segmented. The segmentation method can be determined by the staff according to different types of original documents for targeted division. After segmentation, multiple paragraphs corresponding to the original document are obtained, and each paragraph is the first text paragraph in this embodiment.

[0082] S13: Perform keyword strong association processing on the first text paragraph according to the keyword corresponding to the original document to obtain the target text paragraph.

[0083] In this embodiment, after a large number of first text paragraphs are obtained by segmenting the original document with corresponding keywords, each first text paragraph is respectively subjected to strong association processing with the keyword corresponding to the original document, that is, the keyword corresponding to the original document is inserted into each first text paragraph. After each first text paragraph is strongly associated with the keyword, a respective target text paragraph is obtained, and the target text paragraph corresponding to the first text paragraph is the text paragraph with the corresponding keyword inserted.

[0084] S14: Extract question-and-answer pairs from the target text paragraph through a question-and-answer pair extraction model to obtain the corresponding question-and-answer pairs.

[0085] In this embodiment, the question-and-answer pair includes a question and an answer, which is used to answer questions raised by users in the large model question-and-answer scenario. The generation of the question-and-answer pair mainly depends on the question-and-answer pair extraction model and the target text paragraph. Specifically, the target text paragraph is vectorized using Transformers, and the embedding feature representation of the target paragraph text is input into the question-and-answer pair extraction model. The question-and-answer pair extraction model outputs multiple question-and-answer pairs corresponding to the target text paragraph according to the keywords corresponding to the target text paragraph and the text content in the paragraph. Preferably, the questions in the question-and-answer pair can be generated according to the keywords.

[0086] Through the question-and-answer pair extraction method provided in this embodiment, the original document with keywords is segmented into multiple first text paragraphs, and then the first text paragraphs are strongly associated with the keywords to generate target text paragraphs, so that the obtained target text paragraphs cover the content of the original document completely, and the question-and-answer pairs generated through the target text paragraphs cover the entire original document while improving the quality of the extracted question-and-answer pairs, that is, ensuring that the extracted question-and-answer pairs can accurately point to the accurate vehicle series, etc.

[0087] Combined with the above embodiments, in one implementation manner, the embodiment of the present invention also provides a method for extracting question-and-answer pairs. In this method for extracting question-and-answer pairs, according to the keywords corresponding to the original document, the first text paragraphs are subjected to strong keyword association processing to obtain target text paragraphs, including steps S21 to S23:

[0088] S21: Extract the exact keywords in the first text paragraph.

[0089] S22: Perform strong keyword association processing on the first text paragraph through the exact keywords of the first text paragraph to obtain the corresponding target text paragraph.

[0090] S23: Perform strong keyword association processing on the first text paragraphs for which the corresponding exact keywords have not been extracted through the keywords corresponding to the original document to obtain the corresponding target text paragraphs.

[0091] In this embodiment, in order to extract more accurate question-and-answer pairs from the target text paragraph, this embodiment also performs more accurate keyword strong association processing on the first text paragraph. Specifically, after obtaining the first text paragraph through paragraph segmentation processing of the original document, accurate keyword extraction is performed on each first text paragraph. For example, when the keyword of the original document with a corresponding keyword is Series A, it is determined that the original document has a keyword and is cut into multiple first text paragraphs; if the single first text paragraph obtained by segmentation includes a keyword of a vehicle version smaller in scope than Series A, which is the luxury version, then "the luxury version of Series A" is used as the accurate keyword of this first text paragraph; if the single first text paragraph obtained by segmentation includes a keyword of a vehicle version smaller in scope than Series A, which is the basic version, then "the basic version of Series A" is used as the accurate keyword of this first text paragraph.

[0092] In this embodiment, the vehicle series is a larger range level compared to the vehicle version. For example, under the 2i vehicle series, there can be a pure electric type of the 2i vehicle series and a hybrid gasoline type of the 2i series; the vehicle model is a larger range level compared to the vehicle version. For example, under the "pure electric type", there can be a luxury version of the pure electric type and a basic version of the pure electric type. Preferably, when the keyword corresponding to the original document is specified by the user, a larger range of vehicle series or vehicle model can be specified as the keyword corresponding to the original document; when determining the accurate keyword of the first text paragraph obtained by segmenting the original document with a corresponding keyword, a keyword smaller in scope compared to the keyword corresponding to the original document can be extracted as the accurate keyword of this first text paragraph, and this accurate keyword is strongly associated with this first text paragraph, so that better-quality question-and-answer pairs can be extracted based on the first text paragraph after strong association. For example, when the keyword corresponding to the original document is a vehicle series keyword, both the vehicle version and the vehicle model can be used as accurate keywords. When the keyword corresponding to the original document is a vehicle model keyword, the vehicle version can be used as the accurate keyword.

[0093] In this embodiment, if a more accurate keyword smaller in scope than the keyword corresponding to the original document can be extracted from the first text paragraph, then this first text paragraph is subjected to keyword strong association processing with the extracted more accurate keyword, and the corresponding target text paragraph is obtained after the strong association processing.

[0094] In this embodiment, if no precise keywords with a smaller range than the keywords corresponding to the original document are extracted from the first text paragraph, the keywords corresponding to the original document are strongly associated with the first text paragraph to obtain the corresponding target text paragraph, so as to ensure that each first text paragraph obtained by segmenting the original document can undergo strong keyword association processing, thereby improving the quality of the extracted question-and-answer pairs while achieving full coverage of the original document, that is, ensuring that the extracted question-and-answer pairs can accurately point to the vehicle series, or the vehicle model, or the vehicle version, etc.

[0095] Combined with the above embodiments, in one implementation, the embodiments of the present invention further provide a method for extracting question-and-answer pairs. In this method for extracting question-and-answer pairs, the method further includes steps S31 to S34:

[0096] S31: Segment the original document without corresponding keywords to obtain each target text paragraph.

[0097] S32: Extract question-and-answer pairs from the target text paragraph through a question-and-answer pair extraction model to obtain the corresponding question-and-answer pairs, including: determining the keyword features of the target text paragraph.

[0098] S33: Extract question-and-answer pairs from the target text paragraph with keyword features through a narrow-domain question-and-answer pair extraction model to obtain the corresponding first question-and-answer pairs.

[0099] S34: Extract question-and-answer pairs from the target text paragraph without keyword features through a general question-and-answer pair extraction model to obtain the corresponding second question-and-answer pairs.

[0100] In this embodiment, for the original document without corresponding keywords, correspondingly, the content in this original document will not have specific directional keywords (such as keywords pointing to the vehicle series, keywords pointing to the vehicle model, keywords pointing to the vehicle version). This application will directly segment this type of original document without any strong keyword association processing, and the text paragraphs obtained after segmentation are the target text paragraphs that can be directly used for question-and-answer pair extraction.

[0101] In this embodiment, since the text paragraphs for question-and-answer pair extraction include target text paragraphs with strong keyword association processing and target text paragraphs without any strong keyword association processing. Therefore, in order to ensure the extraction effect of question-and-answer pairs, this application provides two question-and-answer pair extraction models. One is a trained and qualified narrow-domain question-and-answer pair extraction model obtained by training based on a large number of target text paragraphs with strong keyword association processing; the other is a trained and qualified general question-and-answer pair extraction model obtained by training based on a large number of target text paragraphs without any strong keyword association processing. When the target text paragraph belongs to the text paragraph with strong keyword association processing, it is determined that the target text paragraph belongs to the target text paragraph with keyword features. At this time, the narrow-domain question-and-answer pair extraction model is used to extract question-and-answer pairs for the target text paragraph with keyword features, and a plurality of first question-and-answer pairs corresponding to the target text paragraph are obtained. When the target text paragraph belongs to the text paragraph without strong keyword association processing, it is determined that the target text paragraph belongs to the target text paragraph without keyword features. At this time, the general question-and-answer pair extraction model is used to extract question-and-answer pairs for the target text paragraph without keyword features, and a plurality of second question-and-answer pairs corresponding to the target text paragraph are obtained. For example, if the target text paragraph X is about the adjustment of vehicle seats without special model and version requirements, then for the target text paragraph X without keyword features, the general question-and-answer pair extraction model can directly generate relevant second question-and-answer pairs such as "How is the seat adjusted?" from the target text paragraph X.

[0102] Combined with the above embodiments, in one implementation manner, the embodiment of the present invention further provides a method for extracting question-and-answer pairs. In this method for extracting question-and-answer pairs, the method further includes steps S41 to S44:

[0103] S41: Determine the first similarity between the first question-and-answer pair and the keyword features in the corresponding target text paragraph.

[0104] S42: Filter the first question-and-answer pair according to the first similarity.

[0105] S43: Determine the second similarity between the second question-and-answer pair and the historical second question-and-answer pair.

[0106] S44: Determine the second question-and-answer pair to be filtered from the second question-and-answer pair and the historical second question-and-answer pair according to the second similarity.

[0107] In this embodiment, in order to ensure that the first question-and-answer pair generated by the narrow-domain question-and-answer pair extraction model can more accurately fit the content in the target text paragraph, the present application will calculate the similarity between the first question-and-answer pair and the keyword features in the corresponding target text paragraph to determine the quality of the generated first question-and-answer pair. If the similarity is high, it indicates that the quality of the first question-and-answer pair is good, and the first question-and-answer pair can well point to the keyword feature. If the similarity is low, it indicates that the quality of the first question-and-answer pair is poor, and the first question-and-answer pair cannot well point to the keyword feature.

[0108] Specifically, for the generated first question-and-answer pair, calculate the similarity between the first question-and-answer pair and the keyword features in the target text paragraph to which the first question-and-answer pair belongs, and obtain the first similarity corresponding to the first question-and-answer pair. In the case where the first similarity is less than the preset similarity threshold, it is determined that the quality of the first question-and-answer pair is poor, and the first question-and-answer pair is filtered out; in the case where the first similarity is greater than or equal to the preset similarity threshold, it is determined that the quality of the first question-and-answer pair is good, and the first question-and-answer pair is retained. The similarity threshold can be set according to the actual scenario and will not be specifically limited here.

[0109] In this embodiment, for the second question-and-answer pair randomly generated by the general question-and-answer pair extraction model, in order to prevent a large number of duplicate second question-and-answer pairs from being recorded, after the second question-and-answer pair is generated, the generated second question-and-answer pair is calculated for similarity with the historical second question-and-answer pairs generated in the past, and the corresponding second similarity is obtained, and then the second question-and-answer pair to be filtered is determined according to the second similarity.

[0110] Specifically, in the case where the second similarity between the currently generated second question-and-answer pair and the historical second question-and-answer pairs generated in the past is less than the preset second similarity threshold, it is determined that the similarity between the second question-and-answer pair and the historical second question-and-answer pairs is low, and the second question-and-answer pair is retained; in the case where the second similarity between the currently generated second question-and-answer pair and the historical second question-and-answer pairs generated in the past is greater than or equal to the preset second similarity threshold, it is determined that the similarity between the second question-and-answer pair and the corresponding historical second question-and-answer pair is high. At this time, one of the second question-and-answer pair and the historical second question-and-answer pair needs to be filtered. In this embodiment, the confidence level when generating the second question-and-answer pair and the confidence level when generating the historical second question-and-answer pair will be compared, and the one with the lower confidence level among the second question-and-answer pair and the historical second question-and-answer pair will be filtered. The second similarity threshold can be set according to the actual scenario and will not be specifically limited here.

[0111] Combining the above embodiments, in one implementation manner, the embodiments of the present invention also provide a method for extracting question-and-answer pairs. In this method for extracting question-and-answer pairs, the original document with corresponding keywords is segmented into paragraphs to obtain each first text paragraph, including steps S51 to S53:

[0112] S51: Segment the original document with corresponding keywords into paragraphs to obtain each initial text paragraph.

[0113] S52: Perform recognition processing on the obtained initial text paragraphs to determine whether the initial text paragraphs have type characteristics corresponding to various paragraph types.

[0114] S53: Determine the paragraph type corresponding to the type characteristics of the initial text paragraph as the type of the initial text paragraph, and the various paragraph types at least include enumerated paragraphs, table paragraphs, and picture paragraphs.

[0115] S54: Extract the content of the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph to determine the first text paragraph corresponding to the initial text paragraph.

[0116] In this embodiment, since the content in the original document is usually complex and diverse, and the presentation forms of the content may be different for different types of content. For example, the explanatory content about automotive parts may be presented in the form of pictures, and the operation content about different buttons of the car may be presented in the form of tables. Therefore, in this embodiment, in order to extract the content more accurately and effectively, extraction rules are set for different types of content.

[0117] Specifically, after segmenting the original document with corresponding keywords according to a certain segmentation method to obtain each initial text paragraph, then determine the type of the initial text paragraph, such as a table paragraph or a picture paragraph, according to the presentation form of the content involved in the initial text paragraph. Determine the extraction rules corresponding to each type for different types of initial text paragraphs, and extract the content of the corresponding initial text paragraph based on the extraction rules. Generate the first text paragraph corresponding to the initial text paragraph through the extracted content.

[0118] In this embodiment, an alternative implementation manner for segmenting the original document with corresponding keywords is as follows: Parse the original document, read the content stream in the original document to identify the grid data, text content, graphic elements, and page information (such as paragraph indentation, blank lines, and delimiters, etc.) of each page in the original document. According to the recognition result, segment the original document through preset rules. The preset rules at least include dividing the content from the continuous text with a first-line indent of 2 characters to the position before the next first-line indent of two characters into a paragraph; dividing the content from the continuous text with a first-line indent of 2 characters to before a blank line or a delimiter into a paragraph; dividing the content between two identical specific text intervals into a paragraph; the title characters before the first-line indent of 2 characters (such as "1.", "Two", " The content after ” etc.) is divided into the text with a first-line indent of 2 characters for paragraph division.

[0119] In this embodiment, an optional implementation manner for determining the type of the initial text paragraphs obtained by paragraph segmentation is as follows: The type determination method for each initial text paragraph obtained by paragraph segmentation is the same. Here, one initial text paragraph is used for illustration. The paragraph types involved in the initial text paragraphs in this application at least include enumerated paragraphs, table paragraphs, and picture paragraphs. For each paragraph type, this application pre-defines the type features corresponding to the paragraph type. Then, the initial text paragraphs obtained by paragraph segmentation are identified to determine whether there are type features corresponding to a certain or certain paragraph types in the initial text paragraph. When it is identified and determined that there are type features corresponding to a certain or certain paragraph types in the initial text paragraph, the certain or certain paragraph types are determined as the type of the initial text paragraph. For example, when it is identified and determined that there are type features corresponding to the table paragraph type in the initial text paragraph, the type of the initial text paragraph is determined to be a table paragraph; when it is identified and determined that the initial text paragraph not only has type features corresponding to the table paragraph type but also has type features corresponding to the picture paragraph type, it is determined that the type of the initial text paragraph belongs to both the table paragraph type and the picture paragraph type.

[0120] In this embodiment, when the type of an initial text paragraph includes multiple types, the content of the initial text paragraph is extracted based on the multiple extraction rules corresponding to the multiple types, and the first text paragraph corresponding to the initial text paragraph is generated through the extracted content. For example, in the case where the type of the initial text paragraph belongs to both the table paragraph type and the picture paragraph type at the same time, the extraction rules corresponding to the table paragraph type and the extraction rules corresponding to the picture paragraph type are simultaneously adopted to extract the content of the initial text paragraph. The content of the table part in the initial text paragraph is extracted through the extraction rules corresponding to the table paragraph type, and the content of the picture part in the initial text paragraph is extracted through the extraction rules corresponding to the picture paragraph type.

[0121] In this embodiment, for an enumerated paragraph, its corresponding type feature is that any one of various enumeration symbols appears in the text paragraph, and the texts where the simultaneously appearing enumeration symbols are located are adjacent to each other and have the same font size. Among them, various enumeration symbols at least include bullet points (such as “•”, “ ” etc.), numerical numbers (such as “1.”, “(1)” etc.), and alphabetical numbers (such as “a”, “A” etc.).

[0122] For a table - type paragraph, its corresponding type feature is that there is a mixture of vertical and horizontal lines with a preset number of similar average intervals in the text paragraph, and some texts with the same font size and the same font distance are inserted therein. Wherein, the preset number can be set according to the actual scenario and is not specifically limited herein.

[0123] For a picture - type paragraph, when the original document with corresponding keywords is in PDF format, its corresponding type feature is that the Do operator appears in the text paragraph, and the object type operated by the Do operator is Image. While for the original document with corresponding keywords in word, html format, etc., for a picture object there is a corresponding tag, and it is possible to directly determine whether the initial text paragraph obtained after paragraph segmentation has the type feature corresponding to the picture - type paragraph by identifying whether the initial text paragraph has the tag corresponding to the picture object, that is, as long as it is identified and determined that the initial text paragraph has the tag corresponding to the picture object, it is determined that the initial text paragraph has the type feature corresponding to the picture - type paragraph, and further it is determined that the type of the initial text paragraph includes the picture - type paragraph.

[0124] Combined with the above - mentioned embodiments, in one implementation manner, the embodiments of the present invention further provide a method for extracting question - answer pairs. In this method for extracting question - answer pairs, content extraction is performed on the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph to determine a first text paragraph corresponding to the initial text paragraph, including steps S61 to S62:

[0125] S61: Perform content extraction on the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph to obtain a text paragraph to be filtered corresponding to the initial text paragraph.

[0126] S62: Filter out irrelevant fields from the text paragraph to be filtered according to the relationship between the number of tokens of the text paragraph to be filtered and the token threshold, and obtain a first text paragraph corresponding to the initial text paragraph.

[0127] In this embodiment, since the amount of data that the question - answer pair extraction model can process at one time is limited, and for the text of a single paragraph, inputting it into the model without further segmentation for question - answer pair extraction has a better effect. Therefore, this application further filters out invalid information for the text paragraphs with a large amount of data obtained after segmentation, so as to ensure the effect of the model for question - answer pair extraction.

[0128] Specifically, according to the extraction rules corresponding to the type of the initial text paragraph, the content of the initial text paragraph is extracted. After the extraction is completed, the text paragraph to be filtered corresponding to the initial text paragraph is initially obtained. Determine the number of tokens in the text paragraph to be filtered. The number of tokens can refer to the basic unit Token when a natural language processing model processes text data. Then, compare the number of tokens in the text paragraph to be filtered with the set token threshold. When the number of tokens in the text paragraph to be filtered is greater than the token threshold, the irrelevant fields in the text paragraph to be filtered are filtered. The irrelevant fields at least include stop words, modal particles, punctuation marks, etc., so as to reduce the number of tokens in the text paragraph to be filtered below the token threshold. The initial text paragraph remaining after the irrelevant field filtering forms the first text paragraph.

[0129] Combined with the above embodiments, in one implementation manner, the embodiments of the present invention further provide a method for extracting question-and-answer pairs. In this method for extracting question-and-answer pairs, through the extraction rules corresponding to the type of the initial text paragraph, the content of the initial text paragraph is extracted to obtain the text paragraph to be filtered corresponding to the initial text paragraph, including steps S71 to S74:

[0130] S71: When the type of the initial text paragraph includes an enumeration paragraph, extract each enumeration text and associate each extracted enumeration text with the main body of the initial text paragraph.

[0131] S72: When the type of the initial text paragraph includes a table paragraph, extract the row data of the table and determine that the column name of the table is the attribute of the corresponding row data.

[0132] S73: Associate the extracted row data and row data attributes with the main body of the initial text paragraph.

[0133] S74: When the type of the initial text paragraph includes a picture paragraph, extract the network address of the picture and associate it with the main body of the initial text paragraph.

[0134] S75: When the extraction of each type of data in the initial text paragraph and the main body association are completed, the text paragraph to be filtered corresponding to the initial text paragraph is obtained.

[0135] In this embodiment, the present application provides respective corresponding extraction rules for different types of initial text paragraphs, including the extraction rules for enumeration paragraphs, the extraction rules for table paragraphs, and the extraction rules for picture paragraphs. Among them, the enumeration type refers to an expression that clearly lists all possible options. For example, the device status has only three types: "power on, power off, standby". When filling in the device status, only one of these three can be selected for filling, and filling in others is invalid. These three answers are the enumeration type regarding the device status.

[0136] When the type of the initial text paragraph includes an enumerated paragraph, it indicates that the initial text paragraph involves enumerated content, such as examples of several situations. In the case where the initial text paragraph involves enumerated content, in this embodiment, each enumerated text is associated with the main body of the initial text paragraph. The main body of the initial text paragraph refers to the other text content in the initial text paragraph excluding the enumerated content, list content, and picture content. Associating the extracted content with the main body of the initial text paragraph means splicing the extracted content into the main body of the initial text paragraph.

[0137] When the type of the initial text paragraph includes a table paragraph, it indicates that the initial text paragraph involves table content, including the row content of the table and the column content corresponding to the rows. In the case where the initial text paragraph involves table content, in this embodiment, first, the content in each row of the table is extracted to obtain row data, and then, according to the corresponding relationship between the rows and columns, the column corresponding to the row is determined, and the column name of this column is determined as the attribute corresponding to the row data of this row. After obtaining the row data and the corresponding row data attributes, the row data and the row data attributes are associated with the main body of the initial text paragraph.

[0138] When the type of the initial text paragraph includes a picture paragraph, it indicates that the initial text paragraph involves picture content. In the case where the initial text paragraph involves picture content, in this embodiment, the network address of the picture is extracted and associated with the main body of the initial text paragraph. In the special case where the picture has no network address, this application can also extract the picture, store it at a specified network address, and associate the network address with the main body of the initial text paragraph.

[0139] In this embodiment, after completing the extraction of each data type in a single initial text paragraph and associating all the extracted data with the main body of this single initial text paragraph, the text paragraph to be filtered corresponding to this single initial text paragraph is obtained. For example, if the type of a single initial text paragraph includes an enumerated paragraph and a picture paragraph, then each enumerated text in the enumerated content of this single initial text paragraph is extracted and associated with the main body of this single initial text paragraph, and at the same time, the network address in the picture content of this single initial text paragraph is extracted and associated with the main body of this single initial text paragraph. After completing all the associations, the text paragraph to be filtered corresponding to this single initial text paragraph is obtained.

[0140] Combined with the above embodiments, in one implementation manner, the embodiment of the present invention also provides a method for extracting question-answer pairs. In this method for extracting question-answer pairs, training a narrow-domain question-answer pair extraction model includes steps S81 to S83:

[0141] S81: Annotate the to-be-annotated sample passage using the standard question-and-answer pairs of the to-be-annotated sample passage to obtain a training sample passage, where the to-be-annotated sample passage is the target text passage after strong association processing.

[0142] S82: Train an initial model using the training sample passage, and evaluate the performance of the trained initial model using the standard question-and-answer pairs of the training sample passage.

[0143] S83: Determine whether a qualified narrow-domain question-and-answer pair extraction model is obtained through training according to the evaluation result.

[0144] The training method of a narrow-domain question-and-answer pair extraction model provided in this embodiment is as Figure 2 shown. Specifically, the staff sets corresponding standard question-and-answer pairs based on the to-be-annotated sample passage, and annotates the to-be-annotated sample passage using the standard question-and-answer pairs. Among them, the to-be-annotated sample passage refers to the target text passage obtained after strong association processing with keywords. After the to-be-standard sample passage is annotated by the standard question-and-answer pairs, a training sample passage is generated, and this training sample passage is used for subsequent training of the initial model.

[0145] Input the training sample passage into the initial model for model training, and evaluate the performance of the trained initial model according to the training question-and-answer pairs output by the initial model and the standard question-and-answer pairs corresponding to the training sample passage. Specifically, calculate the training question-and-answer pairs output by the initial model and the standard question-and-answer pairs, and evaluate the performance of the initial model according to the calculation result. If the calculation result is that the calculated value approaches the preset value, the evaluation result is that the training is qualified, and this initial model has been trained into a qualified narrow-domain question-and-answer pair extraction model. Otherwise, the evaluation result is that the training is unqualified, and a qualified narrow-domain question-and-answer pair extraction model cannot be obtained.

[0146] In this embodiment, for the narrow-domain Q&A pair extraction model, the full-training method is used. The training parameter train_type of the model is set to full, and at the same time, the set keywords are limited to a smaller range to make the training data more accurate. The narrow-domain Q&A pair extraction model trained in this way is more suitable for Q&A scenarios with a smaller scope and stronger professionalism. For the general Q&A pair extraction model, the incremental fine-tuning form is used. The training parameter train_type of the model is set to lora, and no keywords are specified for fine-tuning. The model trained by this method is more adaptable to the extraction of Q&A pairs from general-purpose documents, has a weaker dependence on keywords, and at the same time cannot extract accurate questions and answers strongly related to related words. This application provides an evaluation function during and after model training. During training, the model will randomly select a preset percentage (such as 30%) of the data as the validation set to evaluate the model, and at the same time generate an evaluation file to display the accuracy and loss in each epoch. After training is completed, the user can also upload the validation set and call the evaluation method to manually evaluate the model.

[0147] Reference Figure 3 , Figure 3 FIG. is a flowchart of a Q&A method based on Q&A pairs proposed in another embodiment of the present application. The method includes the following steps:

[0148] S91: Extract keywords from the question text.

[0149] S92: According to the extracted keywords, determine the first target Q&A pair corresponding to the question text from the first Q&A pair set. The first Q&A pairs in the first Q&A pair set are extracted by the Q&A pair extraction method described in any of the above embodiments.

[0150] S93: Answer the question text based on the first target Q&A pair.

[0151] S94: According to the question text for which no corresponding keywords are extracted, determine the second target Q&A pair corresponding to the question text from the second Q&A pair set. The second Q&A pairs in the second Q&A pair set are extracted by the Q&A pair extraction method described in any of the above embodiments.

[0152] S95: Answer the question text based on the second target Q&A pair.

[0153] Based on the question-answer pair provided in the above embodiment, this embodiment applies the question-answer pair to the question scenario input by the user. Specifically, keyword extraction is performed on the question text input by the user. If the keyword corresponding to the question text can be extracted, the first target question-answer pair corresponding to the question text is determined from the first question-answer pair set based on the keyword. Among them, the first question-answer pair set records a large number of first question-answer pairs extracted by the question-answer pair extraction method in the above embodiment. Since the first question and answer is associated with a corresponding keyword, the first question-answer pair corresponding to the keyword can be accurately found through the keyword in the question text, and the first question-answer pair is determined as the first target question-answer pair that fits the question text. Based on the text content in the first target question-answer pair, a corresponding answer is given to the question text input by the user.

[0154] If the corresponding keywords cannot be extracted from the question text, this embodiment will determine a second target question-answer pair that can answer the question text from the second question-answer pair set. Among them, the second question-answer pair set records a large number of second question-answer pairs extracted by the question-answer pair extraction method in the above embodiment. Since there are no corresponding keywords in the second question-answer pairs and the question text, the second question-answer pair that is more consistent with the question text is screened from the multiple second question-answer pairs based on the text content in the question text, and is determined as the second target question-answer pair. Based on the text content in the second target question-answer pair, a corresponding answer is given to the question text input by the user.

[0155] In conjunction with the above embodiments, in one implementation, an embodiment of the present invention further provides a question-answering method based on question-answer pairs. In this question-answering method based on question-answer pairs, before determining a second target question-answer pair corresponding to the question text from a second set of question-answer pairs based on the question text from which no corresponding keywords have been extracted, the method further includes steps S101 to S104:

[0156] S101: Determine, based on a question text from which no corresponding keyword is extracted, a target keyword corresponding to a terminal that initiates the question text.

[0157] S102: constructing a new question text according to the question text and the determined target keywords;

[0158] S103: According to the new question text, determine a first target question-answer pair corresponding to the question text from a first question-answer pair set.

[0159] S104: The method of determining a second target question-answer pair corresponding to the question text from the second question-answer pair set based on the question text from which the corresponding keywords have not been extracted includes: in the case where the target keyword corresponding to the terminal that initiated the question text has not been determined, determining a second target question-answer pair corresponding to the question text from the second question-answer pair set based on the question text from which the corresponding keywords have not been extracted.

[0160] In this embodiment, the present application takes into account that users usually ask questions related to themselves when asking technical questions. Therefore, when the corresponding keywords are not extracted from the question text entered by the user, the corresponding target keywords can be determined by the terminal from which the user initiated the question text. Specifically, when the user issues the question text through a mobile terminal, the model of the vehicle bound to the mobile terminal is determined and the model of the vehicle bound to the mobile terminal is determined as the target keyword; when the user issues the question text through an in-vehicle terminal, the model of the vehicle in which the in-vehicle terminal is located is determined as the target keyword. Based on the determined target keywords and the constructed new question text, the new question text is a question text including the target keywords. Then, according to the above steps S91 to S92, a first target question-answer pair that matches the new question text is determined from the first question-answer pair set. If the corresponding target keyword cannot be determined based on the terminal from which the question text was initiated, a second target question-answer pair corresponding to the question text is still determined from the second question-answer pair set based on the content of the question text from which the corresponding keywords were not extracted.

[0161] Based on the same inventive concept, an embodiment of the present application provides a system for extracting question and answer pairs, referring to Figure 4 , Figure 4 : is a schematic diagram of a system for extracting question-answer pairs proposed in one embodiment of the present application, the system comprising:

[0162] A first keyword determination module 401 is used to determine keywords corresponding to the original document;

[0163] The paragraph segmentation module 402 is used to segment the original document having the corresponding keywords into paragraphs to obtain each first text paragraph;

[0164] A strong association module 403 is configured to perform keyword strong association processing on the first text paragraph based on the keywords corresponding to the original document to obtain a target text paragraph;

[0165] The question-answer pair extraction module 404 is configured to extract question-answer pairs from the target text paragraph using a question-answer pair extraction model to obtain corresponding question-answer pairs.

[0166] In this embodiment, the system further includes a Mysql database, a server, and a Milvus vector database. After determining the product-related instruction manual as the original document for which Q&A pairs are to be generated, the format of the original document is converted to the markdown format, and then the original document is imported into the Mysql database, which is connected to the server background to facilitate the staff to perform visual operations on the original document in the system. The obtained target text paragraphs are saved in the Milvus vector database in the form of embedding data.

[0167] The server is used for the management of the original document and Q&A pairs. Users can use it to upload, save, download, and preprocess the original document, and at the same time can set the keywords of the original document. Also, after the document is generated into Q&A pairs, operations such as screening, modifying, and deleting Q&A pairs can be performed. At the same time, the system also includes a front-end page for interacting with users, and the keywords corresponding to the original document can be specified by the user through the front-end page. Users can import the original document to be extracted through the import button on the page, use the edit button to modify or edit a certain original document individually, and select to configure keywords. After completion, the original document is added to the extraction queue by single selection or multiple selection and extracted in order. Users can select a pre-trained large model to generate Q&A pairs. It is recommended that users select the most suitable model according to the type of the uploaded original document to improve the quality of the generated Q&A pairs. If the generated document is highly professional, use a narrow-domain Q&A pair extraction model and provide keywords to get more accurate answers. On the contrary, for documents with strong generalization and versatility, a general Q&A pair extraction model can be used, and keywords do not need to be provided to generate generalized Q&A pairs. The extracted Q&A pairs are saved in the storage device, and at the same time, the corresponding original document id, used model, parameters, submitting user, generation time, status, etc. are saved. The generated Q&A pairs are saved in the Mysql database and will also be displayed on the front-end page for users to view, modify, change the status, etc.

[0168] Optionally, the strong association module 403 includes:

[0169] An exact keyword extraction module for extracting exact keywords in the first text paragraph;

[0170] A strong association first sub-module for performing keyword strong association processing on the first text paragraph through the exact keywords of the first text paragraph to obtain the corresponding target text paragraph;

[0171] A strong association second sub-module for performing keyword strong association processing on the first text paragraph for which no corresponding exact keywords are extracted through the keywords corresponding to the original document to obtain the corresponding target text paragraph.

[0172] Optionally, the Q&A pair extraction system further includes:

[0173] A paragraph segmentation module, configured to segment the original document without corresponding keywords into paragraphs to obtain respective target text paragraphs;

[0174] A keyword feature determination module, configured to extract Q&A pairs from the target text paragraphs through the Q&A pair extraction model, including: determining the keyword features of the target text paragraphs;

[0175] A first Q&A pair determination module, configured to extract Q&A pairs from the target text paragraphs with keyword features through a narrow-domain Q&A pair extraction model to obtain corresponding first Q&A pairs;

[0176] A second Q&A pair determination module, configured to extract Q&A pairs from the target text paragraphs without keyword features through a general Q&A pair extraction model to obtain corresponding second Q&A pairs.

[0177] Optionally, the Q&A pair extraction system further includes:

[0178] A first similarity determination module, configured to determine a first similarity between the first Q&A pair and the keyword features in the corresponding target text paragraph;

[0179] A first filtering module, configured to filter the first Q&A pair according to the first similarity;

[0180] A second similarity determination module, configured to determine a second similarity between the second Q&A pair and the historical second Q&A pair;

[0181] A second filtering module, configured to determine the second Q&A pair to be filtered from the second Q&A pair and the historical second Q&A pair according to the second similarity.

[0182] Optionally, the paragraph segmentation module 402 includes:

[0183] An initial text paragraph determination module, configured to segment the original document with corresponding keywords into paragraphs to obtain respective initial text paragraphs;

[0184] An identification processing module, configured to perform identification processing on the obtained initial text paragraphs to determine whether the initial text paragraphs have type features corresponding to various paragraph types;

[0185] A type determination module, configured to determine the paragraph type corresponding to the type features possessed by the initial text paragraphs as the type of the initial text paragraphs, where the various paragraph types at least include enumerated paragraphs, table paragraphs, and picture paragraphs;

[0186] A paragraph determination module, configured to extract the content of the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph, so as to determine a first text paragraph corresponding to the initial text paragraph.

[0187] Optionally, the paragraph determination module includes:

[0188] A first sub-module for paragraph determination, configured to extract the content of the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph, and obtain a text paragraph to be filtered corresponding to the initial text paragraph;

[0189] A second sub-module for paragraph determination, configured to filter out irrelevant fields from the text paragraph to be filtered according to the relationship between the number of tokens in the text paragraph to be filtered and the token threshold, so as to obtain a first text paragraph corresponding to the initial text paragraph.

[0190] Optionally, the first sub-module for paragraph determination includes:

[0191] An enumeration type paragraph extraction module, configured to extract each piece of enumeration text and associate each extracted piece of enumeration text with the main body of the initial text paragraph when the type of the initial text paragraph includes an enumeration type paragraph;

[0192] A first extraction module for table type paragraphs, configured to extract the row data of the table and determine that the column name of the table is the attribute of the corresponding row data when the type of the initial text paragraph includes a table type paragraph;

[0193] A second extraction module for table type paragraphs, configured to associate the extracted row data and row data attributes with the main body of the initial text paragraph;

[0194] A picture type paragraph extraction module, configured to extract the network address of the picture and associate it with the main body of the initial text paragraph when the type of the initial text paragraph includes a picture type paragraph;

[0195] A text paragraph to be filtered determination module, configured to obtain a text paragraph to be filtered corresponding to the initial text paragraph when the extraction and main body association of each type of data in the initial text paragraph are completed.

[0196] Optionally, the system for extracting question-answer pairs further includes:

[0197] A first model training module, configured to label the to-be-labeled sample paragraph with standard question-answer pairs of the to-be-labeled sample paragraph to obtain a training sample paragraph, where the to-be-labeled sample paragraph is the target text paragraph after strong association processing;

[0198] A second model training module is used to train the initial model using the training sample paragraphs and evaluate the performance of the trained initial model using standard question-answer pairs of the training sample paragraphs;

[0199] The third module of model training is used to determine whether a qualified narrow-domain question-answer pair extraction model has been trained based on the evaluation results.

[0200] Based on the same inventive concept, an embodiment of the present application provides a question-answering system based on question-answer pairs, referring to Figure 5 , Figure 5 : is a schematic diagram of a question-answering system based on question-answer pairs proposed in one embodiment of the present application, the system comprising:

[0201] The second keyword extraction module 501 is used to extract keywords from the question text;

[0202] A first target question-answer pair determination module 502 is configured to determine, based on the extracted keywords, a first target question-answer pair corresponding to the question text from a first question-answer pair set, wherein the first question-answer pair in the first question-answer pair set is extracted by a question-answer pair extraction system provided in the third aspect of the present application;

[0203] A first answering module 503, configured to answer the question text based on the first target question-answer pair;

[0204] A second target question-answer pair determination module 504 is configured to determine, based on a question text for which no corresponding keywords have been extracted, a second target question-answer pair corresponding to the question text from a second set of question-answer pairs, wherein the second question-answer pairs in the second set of question-answer pairs are extracted by a question-answer pair extraction system provided in the third aspect of the present application;

[0205] The second answering module 505 is configured to answer the question text based on the second target question-answer pair.

[0206] Optionally, the second target question-answer pair determination module 504 includes:

[0207] A target keyword determination module, configured to determine a target keyword corresponding to a terminal that initiated the question text based on a question text from which no corresponding keyword has been extracted;

[0208] A question text construction module, configured to construct a new question text based on the question text and the determined target keywords;

[0209] A first target question-answer pair determination submodule is configured to determine, based on the new question text, a first target question-answer pair corresponding to the question text from the first question-answer pair set;

[0210] The second target question-answer pair determination submodule is used to determine the second target question-answer pair corresponding to the question text from the second question-answer pair set based on the question text from which the corresponding keywords have not been extracted, including: in the case where the target keyword corresponding to the terminal that initiated the question text has not been determined, determining the second target question-answer pair corresponding to the question text from the second question-answer pair set based on the question text from which the corresponding keywords have not been extracted.

[0211] Based on the same inventive concept, another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps in the method for extracting question and answer pairs as described in any of the above embodiments are implemented.

[0212] Based on the same inventive concept, another embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the steps of the question-answer pair extraction method as described in any of the above embodiments.

[0213] This application extracts corresponding question-answer pairs from original documents with corresponding keywords, and also extracts corresponding question-answer pairs from question-answer pairs without corresponding keywords, thereby achieving high-quality question-answer pair extraction with full coverage of original documents. As a result, in the question-answer scenario, this application can provide more diverse question-answer pairs, and can also better select questions and answers that better meet user needs.

[0214] Those skilled in the art will appreciate that embodiments of the present invention may provide methods, devices, electronic devices, storage media, or computer program products. Therefore, embodiments of the present invention may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0215] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0216] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising said element.

[0217] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0218] In addition, in the specification and claims, "and / or" means at least one of the connected objects, and the character " / " generally indicates that the related objects before and after are in an "or" relationship.

[0219] The above has introduced in detail a method, a question-and-answer method, a system, a device, and a medium for extracting question-and-answer pairs provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for extracting question-and-answer pairs, characterized in that, The method includes: Determine the keywords corresponding to the original document; Perform paragraph segmentation on the original document with corresponding keywords to obtain each first text paragraph; Perform strong keyword association processing on the first text paragraphs according to the keywords corresponding to the original document to obtain target text paragraphs; Perform paragraph segmentation on the original document without corresponding keywords to obtain each target text paragraph; Extract question-and-answer pairs from the target text paragraphs through a question-and-answer pair extraction model to obtain corresponding question-and-answer pairs; Among them, performing strong keyword association processing on the first text paragraphs according to the keywords corresponding to the original document to obtain target text paragraphs includes: extracting the exact keywords in the first text paragraphs; performing strong keyword association processing on the first text paragraphs through the exact keywords of the first text paragraphs to obtain corresponding target text paragraphs; performing strong keyword association processing on the first text paragraphs that have not been extracted to obtain corresponding exact keywords through the keywords corresponding to the original document to obtain corresponding target text paragraphs; Among them, extracting question-and-answer pairs from the target text paragraphs through a question-and-answer pair extraction model to obtain corresponding question-and-answer pairs includes: determining the keyword features of the target text paragraphs; extracting question-and-answer pairs from the target text paragraphs with keyword features through a narrow-domain question-and-answer pair extraction model to obtain corresponding first question-and-answer pairs; extracting question-and-answer pairs from the target text paragraphs without keyword features through a general question-and-answer pair extraction model to obtain corresponding second question-and-answer pairs; Extract the keywords in the user-entered question text; determine the first target question-and-answer pair corresponding to the question text from the first question-and-answer pair set according to the extracted keywords; answer the question text based on the first target question-and-answer pair; determine the target keywords corresponding to the terminal that initiated the question text according to the question text for which no corresponding keywords have been extracted, where the terminal that initiated the question text includes a mobile terminal bound to a vehicle and an in-vehicle terminal of the vehicle, and the target keywords include the model of the vehicle bound to the terminal; construct a new question text according to the question text and the determined target keywords; determine the first target question-and-answer pair corresponding to the question text from the first question-and-answer pair set according to the new question text; in the case where the target keywords corresponding to the terminal that initiated the question text have not been determined, determine the second target question-and-answer pair corresponding to the question text from the second question-and-answer pair set according to the question text for which no corresponding keywords have been extracted; answer the question text based on the second target question-and-answer pair.

2. The method for extracting question-and-answer pairs according to claim 1, wherein The method further includes: Determine the first similarity between the first question-and-answer pair and the keyword features in the corresponding target text paragraph; Filter the first question-and-answer pair according to the first similarity; Determine the second similarity between the second question-and-answer pair and the historical second question-and-answer pair; Determine the second question-and-answer pair to be filtered from the second question-and-answer pair and the historical second question-and-answer pair according to the second similarity.

3. The method for extracting question-and-answer pairs according to claim 1, characterized in that Performing paragraph segmentation on the original document with corresponding keywords to obtain each first text paragraph includes: Perform paragraph segmentation on the original document with corresponding keywords to obtain each initial text paragraph; Perform recognition processing on the obtained initial text paragraphs to determine whether there are type characteristics corresponding to various paragraph types in the initial text paragraphs; Determine the paragraph type corresponding to the type characteristics of the initial text paragraph as the type of the initial text paragraph, and the various paragraph types at least include enumerated paragraphs, table paragraphs, and picture paragraphs; Perform content extraction on the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph to determine a first text paragraph corresponding to the initial text paragraph.

4. The method for extracting question-and-answer pairs according to claim 3, characterized in that, Performing content extraction on the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph to determine a first text paragraph corresponding to the initial text paragraph includes: Perform content extraction on the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph to obtain a text paragraph to be filtered corresponding to the initial text paragraph; Perform irrelevant field filtering on the text paragraph to be filtered according to the relationship between the number of tokens in the text paragraph to be filtered and the token threshold to obtain a first text paragraph corresponding to the initial text paragraph.

5. The method for extracting question-and-answer pairs according to claim 4, wherein Performing content extraction on the initial text paragraph through an extraction rule corresponding to the type of the initial text paragraph to obtain a text paragraph to be filtered corresponding to the initial text paragraph includes: In the case where the type of the initial text paragraph includes an enumerated paragraph, extract each enumerated text and associate each extracted enumerated text with the main body of the initial text paragraph; In the case where the type of the initial text paragraph includes a table paragraph, extract the row data of the table and determine that the column name of the table is the attribute of the corresponding row data; Associate the extracted row data and row data attributes with the main body of the initial text paragraph; In the case where the type of the initial text paragraph includes a picture paragraph, extract the network address of the picture and associate it with the main body of the initial text paragraph; In the case where the extraction and main body association of each type of data in the initial text paragraph are completed, obtain a text paragraph to be filtered corresponding to the initial text paragraph.

6. The method for extracting question-and-answer pairs according to claim 1, wherein Train the narrow-domain question-answer pair extraction model, including: Label the to-be-labeled sample paragraph with the standard question-answer pair of the to-be-labeled sample paragraph to obtain a training sample paragraph, and the to-be-labeled sample paragraph is the target text paragraph after strong association processing; Train the initial model with the training sample paragraph and evaluate the performance of the trained initial model with the standard question-answer pair of the training sample paragraph; According to the evaluation result, determine whether a qualified narrow-domain question-answer pair extraction model is trained.

7. A system for extracting question-and-answer pairs, characterized in that The system includes: A first keyword determination module for determining the keyword corresponding to the original document; A paragraph segmentation module for performing paragraph segmentation on the original document with corresponding keywords to obtain each first text paragraph; A strong association module for performing keyword strong association processing on the first text paragraph according to the keyword corresponding to the original document to obtain a target text paragraph; A paragraph segmentation module for performing paragraph segmentation on the original document without corresponding keywords to obtain each target text paragraph; A question-answer pair extraction module is used to extract question-answer pairs from the target text paragraph using a question-answer pair extraction model to obtain corresponding question-answer pairs; The strong association module includes: an accurate keyword extraction module for extracting accurate keywords from a first text paragraph; a strong association first submodule for performing keyword strong association processing on the first text paragraph using the accurate keywords of the first text paragraph to obtain a corresponding target text paragraph; a strong association second submodule for performing keyword strong association processing on the first text paragraph from which the corresponding accurate keywords have not been extracted using the keywords corresponding to the original document to obtain a corresponding target text paragraph; The question-answer pair extraction module includes: a keyword feature determination module for determining the keyword features of the target text paragraph; a first question-answer pair determination module for extracting question-answer pairs from the target text paragraph with keyword features using a narrow-domain question-answer pair extraction model to obtain a corresponding first question-answer pair; a second question-answer pair determination module for extracting question-answer pairs from the target text paragraph without keyword features using a general question-answer pair extraction model to obtain a corresponding second question-answer pair; The second keyword extraction module is used to extract keywords from the question text input by the user; a first target question-answer pair determination module, configured to determine, from a first question-answer pair set, a first target question-answer pair corresponding to the question text based on the extracted keywords; A first answering module, configured to answer the question text based on the first target question-answer pair; a target keyword determination module, configured to determine, based on a question text from which no corresponding keyword has been extracted, a target keyword corresponding to a terminal that initiated the question text, wherein the terminal that initiated the question text includes a mobile terminal bound to a vehicle and an in-vehicle terminal of the vehicle, and the target keyword includes a model of the vehicle bound to the terminal; A question text construction module, configured to construct a new question text based on the question text and the determined target keywords; A first target question-answer pair determination submodule is configured to determine, based on the new question text, a first target question-answer pair corresponding to the question text from the first question-answer pair set; A second target question-answer pair determination submodule is configured to, if a target keyword corresponding to the terminal that initiated the question text has not been determined, determine a second target question-answer pair corresponding to the question text from the second question-answer pair set based on the question text from which no corresponding keyword has been extracted; The second answering module is used to answer the question text based on the second target question-answer pair.

8. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and running on the processor, wherein when the computer program is executed by the processor, the steps in the method for extracting question and answer pairs as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the method for extracting question-answer pairs as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Question and answer matching method and device, electronic equipment and storage medium

    CN117633180A

  • Question and answer pair generation method, electronic equipment, storage medium and computer program product

    CN119271774A

  • Real-time voice stream and text dialogue interaction system based on large language model

    CN119539076A

  • Document-based intelligent question and answer method and device, electronic equipment and storage medium

    CN119917615A