Question-and-answer data generation method and apparatus, and computer device and storage medium

By combining cutting document blocks and target models, Q&A data is automatically extracted and generated, and the problems of low efficiency and poor quality of Q&A data mining in the existing technology are solved, and efficient and automated high-quality Q&A data generation is achieved.

WO2025092056A1PCT designated stage expired Publication Date: 2025-05-08DOUYIN VISION CO LTD

Patent Information

Application Number
PCT/CN2024/107859
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-07-26
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The prior art is inefficient and poor data quality when mining Q&A data from existing text data, and usually relies on manual annotation to lead to high costs, poor timeliness and uneven data quality.

Method used

By acquiring multiple cut document blocks, using the target question model and answer model, Q&A data is automatically extracted and generated, including the target question data and the target answer data, and optimized data quality through the scoring model.

Benefits of technology

It realizes efficient and automated mining of high-quality Q&A data from existing text data, reduces labor costs and time costs, and improves data quality and timeliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024107859_08052025_PF_FP_ABST
    Figure CN2024107859_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of data processing, and in particular to a question-and-answer (QA) data generation method and apparatus, and a computer device and a storage medium. The method comprises: acquiring a plurality of segmented document blocks, wherein each segmented document block includes a first preset number of pieces of text data; on the basis of the text data in the segmented document blocks and a target questioning model, obtaining target question data corresponding to the segmented document blocks; acquiring full data, wherein the full data is content information which is associated with the target question data and is included in a complete document composed of the plurality of segmented document blocks; on the basis of the segmented document blocks, the target question data, the full data and a target answering model, obtaining target answer data corresponding to the target question data; and generating QA data on the basis of the target question data and the target answer data. The present disclosure solves the problems in the related art of low efficiency and poor data quality in QA data mining for a document.
Need to check novelty before this filing date? Find Prior Art

Description

Question and answer data generation method, device, computer equipment and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese invention patent application number 202311434678.6, entitled “Method, device, computer equipment and storage medium for generating question and answer data” and filed on October 31, 2023, and the entire application is incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of data processing, and in particular to a method, apparatus, computer equipment, and storage medium for generating question and answer data. Background Art

[0004] Most companies have large amounts of existing text data, accumulated over long periods of business operations. Companies looking to implement intelligent large-scale models need to leverage this existing data. However, this data is not cleansed, organized, or labeled, making it inefficient and low-quality for technicians to mine the required QA (Question and Answer) data from existing documents.

[0005] Currently, in order to ensure that high-quality QA data can be mined from existing documents, manual labeling is usually used. However, this method incurs high labor and time costs, and timeliness cannot be guaranteed. At the same time, manual labeling relies on the professional qualities of the staff themselves, resulting in uneven quality of the mined data, and even a large amount of poor quality data.

[0006] Therefore, related technologies have problems of low efficiency and poor data quality when performing QA data mining on documents.

[0007] Summary of the Invention

[0008] In view of this, the present disclosure provides a method, apparatus, computer device and storage medium for generating question and answer data to solve the problems of low efficiency and poor data quality in related technologies when performing QA data mining on documents.

[0009] In a first aspect, the present disclosure provides a method for generating question-answer data, the method comprising:

[0010] Acquire a plurality of cut document blocks, wherein each cut document block contains a first preset number of text data;

[0011] According to the text data in the cut document block and the target question model, the target question data corresponding to the cut document block is obtained;

[0012] Acquire full data, wherein the full data is content information associated with the target question data contained in a complete document composed of multiple segmented document blocks;

[0013] Obtain target answer data corresponding to the target question data based on the segmented document blocks, target question data, full data, and target answer model;

[0014] Generate question and answer data based on target question data and target answer data.

[0015] In an optional embodiment, before obtaining target question data corresponding to the segmented document block based on the text data in the segmented document block and the target question model, the method further includes:

[0016] Determine the questioning strategy based on the text data in the cut document block;

[0017] Process the text data according to the questioning strategy and extract multiple questions;

[0018] Obtaining first scores for multiple questions according to a question scoring model;

[0019] determining a plurality of candidate questions from the plurality of questions according to the first scores;

[0020] The initial question model is optimized according to the candidate questions to obtain the target question model.

[0021] In an optional embodiment, determining a questioning strategy based on text data in a segmented document block includes:

[0022] Determine the text type of the text data;

[0023] Determine the corresponding questioning strategy based on the text type.

[0024] In an optional embodiment, before obtaining target answer data corresponding to the target question data based on the segmented document blocks, the target question data, the full data, and the target answer model, the method further includes:

[0025] Determine the answer strategy based on the text data in the segmented document block and the target question data;

[0026] Processing the text data according to the answer strategy to obtain multiple answer data;

[0027] obtaining a second score for the plurality of answer data according to the answer scoring model;

[0028] determining a plurality of candidate answer data from the plurality of answer data according to the second score;

[0029] The initial answer model is optimized based on the candidate answer data to obtain the target answer model.

[0030] In an optional embodiment, determining an answer strategy based on the text data in the segmented document block and the target question data includes:

[0031] Determine the text type of the text data;

[0032] Determine the answer strategy based on the text type and target question data.

[0033] In an optional embodiment, obtaining target answer data corresponding to the target question data based on the segmented document blocks, the target question data, the full data, and the target answer model includes:

[0034] Obtain multiple answer data based on the segmented document blocks, target question data, full data, and target answer model;

[0035] Split the answer data into the smallest units to obtain multiple target fields;

[0036] comparing the description contents of the target field in every second predetermined number of answer data;

[0037] The description content corresponding to the answer data with the second highest score is retained;

[0038] Integrate the description content to obtain the target answer data.

[0039] In an optional embodiment, after obtaining the question-answer data about the segmented document blocks, the method further includes:

[0040] Use the question-answer data scoring model to score the question-answer data to obtain a third score for the question-answer data;

[0041] Perform data cleaning on the question and answer data based on the third score to obtain target question and answer data that meets the preset requirements;

[0042] The data sample is expanded according to the target question and answer data to obtain a third preset number of question and answer data.

[0043] In an optional embodiment, the cut document block includes text data after image recognition.

[0044] In a second aspect, the present disclosure provides a device for generating question-answer data, the device comprising:

[0045] A first acquisition module is configured to acquire a plurality of cut document blocks, wherein each cut document block contains a first preset number of text data;

[0046] The first obtaining module is used to obtain target question data corresponding to the cut document block according to the text data in the cut document block and the target question model;

[0047] A second acquisition module is configured to acquire full data, wherein the full data is content information associated with the target question data contained in a complete document composed of a plurality of segmented document blocks;

[0048] The second obtaining module is used to obtain target answer data corresponding to the target question data according to the segmented document blocks, the target question data, the full data and the target answer model;

[0049] The third module is used to generate question-answer data based on the target question data and the target answer data.

[0050] In a third aspect, the present disclosure provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to execute the method for generating question and answer data of the first aspect or any corresponding embodiment thereof.

[0051] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method for generating question and answer data according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0053] FIG1 is a flow chart of a method for generating question-answer data according to some embodiments of the present disclosure;

[0054] FIG2 is a schematic diagram of a complete flow chart of a method for generating question-answer data according to some embodiments of the present disclosure;

[0055] FIG3 is a structural block diagram of a device for generating question-answer data according to some embodiments of the present disclosure;

[0056] FIG4 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0057] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0058] Most companies already have large amounts of existing text data, accumulated over a long period of business operations. Companies looking to implement intelligent large-scale models need to leverage this existing data. However, this data has not been cleaned, organized, or labeled, requiring only pre-training. Pre-training is very expensive, and most companies lack the capacity to accumulate sufficient data, resulting in limited results.

[0059] Therefore, the best strategy for companies is to fine-tune the model. This requires cleaning and quality assurance (QA) of past business data. However, due to concerns about data security, companies find it difficult to hand over data to third-party annotation companies for processing. Therefore, to ensure high-quality QA data can be mined from existing documents, manual annotation is often used. However, this method is labor-intensive and time-consuming, with uncertain timeliness. Furthermore, manual annotation relies on the professional expertise of the staff involved, resulting in inconsistent data quality and even a high incidence of poor quality data.

[0060] In order to solve the above problems, the embodiment of the present disclosure proposes an embodiment of a method for generating question and answer data. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0061] In this embodiment, a method for generating question-and-answer data is provided. FIG1 is a flow chart of a method for generating question-and-answer data according to some embodiments of the present disclosure. As shown in FIG1 , the method can be applied to a server side. The method flow includes the following steps:

[0062] Step S101 : obtaining a plurality of cut document blocks, wherein each cut document block contains a first preset number of text data.

[0063] Optionally, in an embodiment of the present disclosure, the server side obtains a document for question-and-answer data (i.e., QA) to be extracted (or mined), and segments the document, for example, into 10 equal-proportion portions, to obtain a plurality of segmented document blocks. If the segmentation is equal-proportion, each segmented document block contains the same first preset amount of text data. For example, if a document contains 100,000 words of text data, after being segmented into 10 portions, each segmented document block contains the same 10,000 words of text data.

[0064] In addition, before the document is segmented, there may be many problems with the text in the acquired document. In this case, these texts need to be preprocessed to obtain correct text in a unified format. As shown in Table 1, Table 1 records some possible problems with the text and the corresponding solutions:

[0065] Table 1

[0066] In addition to the issues listed above, other text issues may also arise. For example, when automatically extracting system text, misunderstandings may occur due to grammatical errors and typos in the document. Corresponding solutions include error detection and correction, semantic ambiguity resolution, etc. It should be noted that the above content is merely an example, and the embodiments of the present disclosure include, but are not limited to, the above text issues and data preprocessing methods.

[0067] Step S102 : obtaining target question data corresponding to the cut document block according to the text data in the cut document block and the target question model.

[0068] Optionally, the text data in each cut document block is input into a trained target question model to obtain target question data corresponding to the cut document block.

[0069] Step S103 : obtaining full data, wherein the full data is content information associated with the target question data contained in a complete document composed of a plurality of segmented document blocks.

[0070] Alternatively, some answers may require context to fully understand, and simply extracting the question and answer may result in missing important contextual information. In this case, we can expand the scope of context during processing to include contextual information closely related to the question and answer. This allows the model to capture the necessary contextual information to better understand the question and answer. We can also encode contextual information using sentence or paragraph representations. This approach helps the model capture long-range contextual information and incorporate it into the prediction process.

[0071] However, if only context is used to obtain question and answer data, there may be a lack of comprehensiveness. For example, the answer to a question in a document may depend on the content of another document. In this case, full data is needed. Among them, the full data is stored in a vector database (a complete document database composed of multiple cut document blocks). After the question is vectorized, it is matched with the data in the vector database. The similarity of the vector is used to filter out the content information most relevant to the target question data, such as 10 data. Then, the target question data and the 10 related data fragments are given to the question selection answer strategy related model. After applying the answer strategy, the corresponding target answer data is obtained.

[0072] At this point, the full data set can be determined in the following ways: 1. Using an embedding model and a vector database, cross-document search and information extraction can be achieved. 2. Document linking: If there are clear links between documents, documents can be crawled and processed sequentially according to the links to form a contextual understanding. 3. Compound document modeling: Build a multi-input model that simultaneously inputs multiple related documents for understanding and reasoning. 4. Answer extraction and synthesis: When the answer spans a large range, answer extraction is first performed to extract all possible answer fragments. Then, answer synthesis is performed to appropriately stitch these fragments together. During the extraction phase, multiple answer fragments can be set and each answer fragment is assigned a score. During the synthesis phase, the highest-scoring answer is selected, or high-scoring answers are appropriately spliced ​​together. 5. Paragraph-level processing: When the answer may span multiple paragraphs, the paragraph is used as the processing unit, and information from multiple paragraphs is considered simultaneously to generate the answer. Documents can be segmented using sliding windows or other methods to capture a wider range of context.

[0073] It is understandable that the full data already contains context information associated with the target question data.

[0074] Step S104 , obtaining target answer data corresponding to the target question data according to the segmented document blocks, the target question data, the full data, and the target answer model.

[0075] Optionally, when determining the target answer data corresponding to the target question data of each cut document block, the embodiment of the present disclosure needs to consider the text data of each cut document block, the target question data output by the target question model, and all the full data associated with the target question data, and then input these data that need to be considered into the trained target answer model to obtain the target answer data corresponding to the target question data.

[0076] Step S105: Generate question and answer data based on the target question data and the target answer data.

[0077] Optionally, each segmented document block corresponds to one target question data and one target answer data. Then, the question-answer data consisting of one target question data and one target answer data is the final question-answer data mined from each segmented document block.

[0078] In addition, the final question and answer data mined from all the segmented document blocks can be used to form a QA library for storing these question and answer data.

[0079] In an embodiment of the present disclosure, a plurality of cut document blocks are obtained, wherein each cut document block contains a first preset number of text data; target question data corresponding to the cut document block is obtained based on the text data in the cut document block and the target question model; full data is obtained, wherein the full data is content information associated with the target question data contained in a complete document composed of a plurality of cut document blocks; target answer data corresponding to the target question data is obtained based on the cut document block, the target question data, the full data, and the target answer model; and question and answer data is generated based on the target question data and the target answer data. In this way, the embodiment of the present disclosure can efficiently mine high-quality question and answer data from existing documents, achieve the purpose of automation, greatly reduce labor costs and time costs, and solve the problems of low efficiency and poor data quality in manual mining when performing question and answer data mining on documents in related technologies.

[0080] In some optional implementations, before obtaining target question data corresponding to the segmented document block based on the text data in the segmented document block and the target question model, the method further includes:

[0081] Determine the questioning strategy based on the text data in the cut document block;

[0082] Process the text data according to the questioning strategy and extract multiple questions;

[0083] Obtaining first scores for multiple questions according to a question scoring model;

[0084] determining a plurality of candidate questions from the plurality of questions according to the first scores;

[0085] The initial question model is optimized according to the candidate questions to obtain the target question model.

[0086] Optionally, in the disclosed embodiment, before generating a trained target question model, the initial question model needs to be continuously optimized to obtain the final target question model. Before optimizing the initial question model, it is necessary to screen the QA structured data. Among them, the text suitable for extracting QA structured data has the following characteristics:

[0087] a. Rich content: The text should contain rich information points and descriptions

[0088] b. Clear logic: The text should be well-structured and accurately expressed to ensure that the questions and answers extracted from it are accurate.

[0089] c. Factual and objective: The text is objective and factual, which helps to generate accurate and reference-worthy questions and answers.

[0090] d. Domain and task relevance: The ideal text type should be highly relevant to the specific domain and task to ensure that the extracted QA data has practical application value.

[0091] At this time, based on the above-mentioned text characteristics suitable for extracting QA structured data, the questioning strategy for each cut document block is determined, so as to extract multiple question questions (hereinafter also referred to as "questions") from each cut document block.

[0092] In order to dig out high-quality questions so as to construct relevant answers in the subsequent process, the embodiment of the present disclosure sets up a question scoring model in advance to score the multiple questions obtained. Among them, the scoring criteria of the question scoring model are designed as follows: 1. Difficulty of the question: select challenging questions. 2. Diversity of questions: ensure that the questions come from different fields and knowledge points. 3. Accuracy of the question: measure whether the question can accurately describe a specific concept or fact. 4. Quality of the statement of the question: ensure that the question has good grammar and clarity. 5. Clarity of the question: whether the question is clear and easy to understand. 6. Relevance of the question: whether the question is closely related to the given topic or data set. 7. The amount of information in the question: whether the question involves meaningful information. 8. Complexity of the question: whether the question is of a certain difficulty and the answer cannot be obtained by simple search.

[0093] Multiple questions are input into the question scoring model to obtain a score for each question. Questions are sorted according to the score to select high-quality candidate questions. It is understandable that there may be multiple high-quality candidate questions. In this case, multiple cycles are required, and a scoring threshold is set at the same time to retain only questions above the threshold. In addition, the embodiment of the present disclosure can also adjust the scoring threshold as needed to obtain questions of different quantities and qualities.

[0094] Based on the feedback from the question scoring model, the retained questions are optimized, including modifying the question wording and adding key information.

[0095] The selected and optimized candidate questions are used as training samples and input into the initial question model for training. This continuously optimizes the initial question model and improves the quality of questions. After multiple rounds of iteration, the system will generate a high-quality question set and ultimately obtain the target question model.

[0096] In addition, in the process of training the initial question model and obtaining the target question model, the embodiment of the present disclosure also proposes some optimization directions:

[0097] a. Multi-factor scoring: In addition to the aforementioned scoring criteria, multiple factors can be combined to score questions. For example, an NLP model can be used to perform semantic analysis on questions and then comprehensively score them based on factors such as relevance, information content, and complexity.

[0098] b. Adaptive Threshold: During the question screening phase, we use an adaptive threshold strategy to dynamically adjust the scoring threshold. For example, we dynamically adjust the scoring threshold based on factors such as the number of questions, question difficulty, and question category. This allows us to more accurately screen high-quality questions and improve the quality of our question library.

[0099] c. Balanced sample weights: During the return training phase, we try to use a balanced sample weighting approach to avoid model overfitting. We assign different weights to questions based on factors such as their importance and difficulty, balancing the impact of different types of questions during training.

[0100] d. Incorporating Expert Knowledge: During the problem optimization phase, we attempt to incorporate domain expert knowledge and optimize the problem through expert feedback. Experts revise and modify the problem to improve its quality. Furthermore, based on expert feedback, we can further optimize the problem scoring model.

[0101] e. Cross-validation: During the question collection and preprocessing phase, cross-validation can be used to increase the diversity of the question pool. For example, when processing different datasets on the same topic, a subset of questions can be extracted from each dataset and then scored and filtered. This approach allows for a broader set of questions, helping to improve the quality of the question pool.

[0102] In the disclosed embodiments, a question scoring model is used to screen high-quality questions, allowing the initial question model to continuously adapt to the mining strategies and styles of specific businesses or industries, thereby improving the accuracy of the trained target question model. Dynamic optimization is also performed using relevant strategies and algorithms, allowing for learning and improvement through the continuous generation of high-quality and authentic synthetic data.

[0103] In some optional implementations, determining a questioning strategy based on text data in a segmented document block includes:

[0104] Determine the text type of the text data;

[0105] Determine the corresponding questioning strategy based on the text type.

[0106] Optionally, feature analysis of the text type of the text data is performed. The following text types are well suited for extracting QA structured data:

[0107] a. Educational materials: Textbooks, handouts, tutorials, etc. contain rich knowledge and explanatory information.

[0108] b. Technical documentation: Product manuals, API documentation, development guides, etc. provide detailed technical descriptions and operating methods.

[0109] c. Research Papers: Research texts such as academic papers and industry reports usually contain rich analysis and conclusions, especially the abstract, which often describes the research questions and conclusions.

[0110] d. Regulations and Policy Documents: Official regulations, policy documents and legal documents are usually well-structured and factual.

[0111] e. News reports and articles: These contain a large number of descriptions of facts, events, and opinions, which helps to extract relevant QA data such as current affairs.

[0112] f. Q&A communities and forums: The Q&A and discussions on the platform usually have clear questions and answers, which can be directly extracted to build a QA library.

[0113] g. Manuals and Guides: These documents usually contain specific questions and corresponding answers, such as user guides, FAQs, product manuals, etc.

[0114] According to the above text types, the corresponding questioning strategy is determined, as shown in Table 2 (Table 2 includes both question extraction strategies and answer extraction strategies corresponding to the questions):

[0115] Table 2

[0116] Table 2 only determines the corresponding questioning strategies and answer extraction strategies for educational materials, such as textbooks, handouts, tutorials, and other text types; technical documents, such as product manuals, API documents, development guides, and research papers, such as academic papers and industry reports. Regulations and policy documents, news reports, and articles also obtain corresponding questioning strategies and answer extraction strategies based on the current text type. The embodiments of the present disclosure will not be repeated here.

[0117] When determining the corresponding questioning strategy based on the text type, there are some auxiliary strategies:

[0118] a. Use different questioning models: Depending on the text type, you may need to use different questioning models to extract targeted questions. For example, when processing technical documents, use a pre-trained model for the technical field; when processing regulatory documents, use a pre-trained model for the legal field.

[0119] b. Dynamically adjust questioning strategies: Dynamically adjust questioning strategies based on the quality of the text, domain, and background knowledge. For example, adjust the threshold for keyword extraction, use different entity types and relationship types, etc.

[0120] c. Increased explainability: Maintaining explainability during question generation allows users to track and understand the question generation strategy. To improve the explainability of the question generation process, the question generation strategy is visualized to provide the reasons for question generation and indicate the model's focus on different parts of the text.

[0121] d. Adaptive Strategy Adjustment: When processing large amounts of text, the model may need to automatically adapt to the characteristics and difficulty of different texts. To this end, we develop adaptive questioning strategies, such as dynamically adjusting the questioning model based on the complexity and domain of the text, to achieve higher-quality question mining.

[0122] e. User Interaction: To improve the effectiveness and flexibility of questioning strategies, user interaction is introduced to enable users to participate in adjusting questioning strategies, provide feedback, and receive correction suggestions. For example, after questions are generated, users can be asked to evaluate them and the questioning strategy can be adjusted based on user feedback.

[0123] f. Combining unsupervised and supervised methods: Combining unsupervised and supervised methods in the question-asking strategy selection process improves the robustness of the model. Unsupervised methods help extract the underlying structure in the text, while supervised methods enable more accurate question generation based on existing annotated data.

[0124] g. Weighing question difficulty and diversity: When selecting a questioning strategy, it's important to balance question difficulty and diversity. Simple and direct questions may be easier for users to understand, but they may not cover the entire text. Complex and diverse questions, on the other hand, cover more dimensions but may pose challenges for users. When building a QA database, it's recommended to comprehensively consider both question difficulty and diversity to avoid overly emphasizing any particular type of question.

[0125] In the embodiments of the present disclosure, questioning strategies, answering strategies, etc. are flexibly selected and adjusted according to different texts and industry requirements to make them more suitable for actual application scenarios.

[0126] In some optional implementations, before obtaining target answer data corresponding to the target question data based on the segmented document blocks, the target question data, the full data, and the target answer model, the method further includes:

[0127] Determine the answer strategy based on the text data in the segmented document block and the target question data;

[0128] Processing the text data according to the answer strategy to obtain multiple answer data;

[0129] obtaining a second score for the plurality of answer data according to the answer scoring model;

[0130] determining a plurality of candidate answer data from the plurality of answer data according to the second score;

[0131] The initial answer model is optimized based on the candidate answer data to obtain the target answer model.

[0132] Alternatively, based on the text data within the segmented document blocks and the target question data, an answer strategy can be determined (see Table 2), resulting in multiple answer data sets. Collecting these answer data sets requires a series of preprocessing steps, including removing irrelevant vocabulary, correcting spelling errors, and optimizing the presentation of the answers.

[0133] Set up an answer scoring model in advance, score the generated answers, retain high-quality answers, and feed the scoring information back to the trained initial answer model to optimize the performance of the initial answer model. At this time, you can choose a common natural language processing answer scoring model. The goal of the answer scoring model is to rank the quality of multiple answer data according to specific criteria. The scoring criteria include: Answer accuracy: Evaluate whether the content of the answer is accurate and how well it matches the question. Answer completeness: Evaluate whether the answer has obtained complete information and whether it can fully answer the question. Answer presentation quality: Check the language presentation of the answer, including grammatical accuracy, comprehensibility, etc.

[0134] Then use the answer scoring model to score each answer, which can reflect the overall quality of the answer.

[0135] The second scores of multiple answers are sorted, a threshold is set, and candidate answers with scores exceeding the threshold are retained. The filtered and optimized candidate answer data is used as training samples and input into the initial answer model for further training. This iterative cycle optimizes the performance of the initial answer model.

[0136] Repeat the above steps as needed. During this iteration, the system will continuously generate high-quality answer sets, preparing for the construction of the QA library and subsequent processes, and ultimately obtaining a trained target answer model. The target answer data ultimately generated by the target answer model should be accurate, complete, and clearly expressed.

[0137] During training of the initial answer model, there are some auxiliary strategies:

[0138] a. Multi-task learning: During training, a multi-task learning strategy is used to simultaneously optimize the accuracy, completeness, and presentation quality of responses. This is achieved through techniques such as sharing hidden layer parameters and using soft parameter sharing between tasks.

[0139] b. Use reinforcement learning strategies: Introduce reinforcement learning strategies, such as using the Actor-Critic algorithm or Deep Q-Learning to determine the best answer. This further improves the quality of answer generation.

[0140] c. Integrate with knowledge graphs: In the answer scoring model, reference external knowledge graphs to enhance the accuracy and trust of answers. By leveraging knowledge graphs, the correctness of answers can be further verified and supported, and additional references can be provided where possible.

[0141] d. Real-time fine-tuning: For real-time application scenarios, a real-time fine-tuning strategy is adopted, that is, updating the answer scoring model in real time based on newly generated answers. This will enable the model to adapt to the changing data distribution and improve the quality of answers.

[0142] e. Explainability: Incorporate explainability mechanisms into the answer scoring model, such as attention mechanisms or model sensitivity analysis, to ensure the quality of answers and provide a basis for analysis and adjustment.

[0143] In the disclosed embodiments, iterative optimization and the use of advanced machine learning methods can lay a solid foundation for building a high-quality question-answer library. At the same time, the questioning strategy and answering strategy can be flexibly selected and adjusted to make them more suitable for actual application scenarios.

[0144] In some optional implementations, determining an answer strategy based on the text data in the segmented document block and the target question data includes:

[0145] Determine the text type of the text data;

[0146] Determine the answer strategy based on the text type and target question data.

[0147] Alternatively, as shown in Table 2, an answer strategy can be determined based on the text type and target question data extraction strategy. Based on the answer strategy, multiple answers relevant to the target question data are obtained. These answers are then scored to obtain a high-quality answer set consisting of multiple candidate answer data. The initial answer model is then iteratively optimized to obtain a trained target answer model. It should be noted that the target answer model should only output one answer, which is the target answer data.

[0148] In some optional implementations, obtaining target answer data corresponding to the target question data based on the segmented document blocks, the target question data, the full data, and the target answer model includes:

[0149] Obtain multiple answer data based on the segmented document blocks, target question data, full data, and target answer model;

[0150] Split the answer data into the smallest units to obtain multiple target fields;

[0151] comparing the description contents of the target field in every second predetermined number of answer data;

[0152] The description content corresponding to the answer data with the second highest score is retained;

[0153] Integrate the description content to obtain the target answer data.

[0154] Optionally, for each question, n answers may be generated. These answers may be generated by different initial answer models, or by the same initial answer model under slightly different input conditions. Break the filtered answers into smaller information units (i.e., target fields). These units may be sentences or phrases describing specific facts. Compare each two information units. If they describe the same facts or details, only keep the answer data with the higher score. If they describe different facts or details, keep both.

[0155] The selected information units are combined to form a new response. A supporting strategy is used to determine the order of the information units, including but not limited to their order in the original response or a logical sequence. Finally, the reconstructed response is post-processed, including checking grammar, adjusting word order, and correcting spelling errors, to obtain the final target response data.

[0156] In the disclosed embodiment, by synthesizing multiple answer data, the quality and comprehensiveness of the answers are efficiently guaranteed, and the quality of the QA library construction is greatly improved compared to a single answer.

[0157] In some optional implementations, after obtaining the question-answer data about the segmented document blocks, the method further includes:

[0158] Use the question-answer data scoring model to score the question-answer data to obtain a third score for the question-answer data;

[0159] Perform data cleaning on the question and answer data based on the third score to obtain target question and answer data that meets the preset requirements;

[0160] The data sample is expanded according to the target question and answer data to obtain a third preset number of question and answer data.

[0161] Optionally, the embodiment of the present disclosure sets up a question-answering data scoring model in advance to check the quality of the obtained QA. Among them, the core indicators of QA data are:

[0162] a. Completeness of questions and answers: Ensure that the extracted questions and answers are complete, without being truncated or missing key parts. Ensure that the answers provided are meaningful to users.

[0163] b. Consistency: Check whether similar or repeated questions and answers extracted from different sections or different documents are consistent. Ensure that the answers provided are consistent across documents or document sections.

[0164] c. Relevance: Are the extracted questions and answers relevant to the context or topic? Ensure that the QA data is relevant to the specific query or task.

[0165] d. Readability: Are the questions and answers easy to read and understand? Ensure that the final model can understand the extracted content well.

[0166] Based on the above core indicators, the question and answer data of each cut document block is scored to obtain the third score of the question and answer data. The scoring criteria are:

[0167] Rule-based checking: Use regular expressions or other rules to check the format of questions and answers. Automatically check the completeness of answers, such as ensuring that answers are not truncated.

[0168] Sentence Embedding Comparison: Use a pre-trained sentence embedding model to convert questions and answers into vectors. Compare the embeddings of questions and answers to assess their relevance or similarity.

[0169] Feedback loop: Automatically generate questions using the language model, validate the answers against the document, and compare the automatically generated questions with the actual extracted questions to assess their quality.

[0170] Diversity and duplication check: Use Jaccard similarity, cosine similarity, or other text similarity methods to check for duplication between the extracted QA data.

[0171] Statistical analysis: Automatically count the frequency of certain keywords or words to determine if there are excessive repetitions or missing topics. Use TF-IDF (term frequency–inverse document frequency, a common weighting technique used in information retrieval and data mining) or other techniques to identify unusual or rare words in questions or answers.

[0172] Contextual consistency: Use a pre-trained language model to evaluate the consistency of the answer in context, ensuring that the answer is relevant to the surrounding text.

[0173] Error analysis: Use automated tools, such as grammar checkers or text classifiers, to identify possible text errors or inconsistencies. Also design an automated system to continuously fine-tune and optimize the data extraction model based on feedback and errors.

[0174] Reference dataset comparison: Comparison with a seed dataset of verified quality, using an automated scoring system to assess the quality of the QA data.

[0175] The disclosed embodiment then cleans the Q&A data based on the third score, removing low-quality Q&A data and sensitive privacy data, to obtain target Q&A data that meets preset requirements and ensures high quality. Furthermore, data samples are expanded based on the cleaned target Q&A data to ensure sufficient data samples for high-quality QA data.

[0176] At the same time, after generating multiple QA data and obtaining the QA library, the embodiment of the present disclosure also proposes some auxiliary strategies to support the application of the QA library in more scenarios.

[0177] a. Incremental Update Strategy: As business grows, enterprises may generate new document data. In this case, an incremental update mechanism can be designed to regularly update and optimize the existing QA library to maintain its timeliness and relevance.

[0178] b. User Feedback Integration: If the QA library will be used in scenarios such as customer support, user feedback on answers to questions can be collected to continuously optimize and update the QA library. For example, users can provide feedback on the relevance, accuracy, and understanding of answers, and the system will continuously optimize based on this feedback.

[0179] c. Data hierarchical annotation: Questions and answers in the QA library are annotated hierarchically, such as by category and difficulty level, to facilitate more accurate matching and display of relevant information in user queries, thereby improving the user experience.

[0180] d. Introducing topic models: Use topic models to cluster questions to maintain a good diversity of questions in the QA database. This helps avoid excessive concentration on a single topic and ensures that the QA database covers multiple fields and knowledge points.

[0181] e. Visual Data Assessment: Provides visualization tools to present the QA scoring and screening process and results. Key indicators such as the distribution of questions and answers, and quality score trends can be displayed to help companies gain insights into data quality.

[0182] f. Integrate external knowledge base: Combine the QA library with the external knowledge base to expand the coverage of questions and answers and improve adaptability to user queries.

[0183] In the embodiments of the present disclosure, strategies such as multi-task learning and reinforcement learning are applied to improve data quality. At the same time, the robustness of the model is improved through methods such as model integration and data enhancement, reducing the impact of uncertainty on model performance.

[0184] In some optional implementations, the cut document block includes text data after image recognition.

[0185] Optionally, in the disclosed embodiment, after the server side obtains the document of the question and answer data (i.e., QA) to be extracted (or mined), it can first determine whether the document is an image document or a non-image document. If it is an image document, it is necessary to identify the image document, identify the corresponding text content, and then cut the document, for example, into 10 equal parts, etc., to obtain multiple cut document blocks, so that each cut document block includes the text data after image recognition.

[0186] As shown in FIG2 , FIG2 is a schematic diagram of a complete process of a method for generating question-answer data according to some embodiments of the present disclosure. The specific process is as follows:

[0187] Obtain unstructured data, such as industry books, articles, papers, etc.;

[0188] If the unstructured data is an image, the image is processed and visual encoder feature extraction is used. Then, OCR technology is used to extract text from the image and obtain the context information of the extracted text. The extracted text and context information are used together for text preprocessing and document segmentation, such as segmenting into n document blocks of less than 10k. If the unstructured data is not an image, the text in the document is preprocessed and the document is segmented into n document blocks of less than 10k. The questioning strategy is selected according to the text type of the text in the document block (such as regular documents, special professional documents, news reports, literary works, etc. in Figure 2). Multiple questioning strategies and questioning models (i.e., initial questioning models) are applied to extract n questions, and the n questions are input into the question scoring model for scoring. This cycle is repeated N times to screen out high-quality questions to optimize the questioning model, and finally an optimal question is output to input into the QA library and determine the answer strategy module.

[0189] When determining the answer strategy, the answer strategy is determined based on the text type within each document block (such as general documents, special professional documents, news reports, literary works, etc. in Figure 2) and the input optimal question. Then, multiple answer strategies, each document block, and the full amount of data associated with the optimal question obtained through vector processing of unstructured data are input into the answer model to obtain multiple answer data. These multiple answer data are then input into the answer scoring model. After N rounds of cycles, high-quality answers are screened and the answer model is learned and optimized. The n answers are integrated to form an optimal answer (also known as the best answer), which is input into the QA library. As can be seen, the QA library stores QA data corresponding to one optimal question and one optimal answer.

[0190] Then, the QA scoring model is used to score each QA data in the QA library and perform quality checks. At the same time, the QA data can be cleaned to remove low-quality data and sensitive privacy data. The cleaned QA data is then expanded. Finally, all models are fine-tuned based on the expanded and cleaned QA data.

[0191] In the above embodiments, data processing is completed locally, and enterprise data will not be exposed to third parties, which greatly increases data security; through multiple rounds of training, including the initial question model, the initial answer model, the question scoring model, and the answer scoring model, these models are continuously adapted to the mining strategies and styles of specific enterprises or industries, thereby realizing knowledge transfer and building enterprise-specific mining models; at the same time, based on the balance of hardware resources and performance requirements, the appropriate model architecture and parameter settings are selected, which can reduce the demand for computing resources without affecting the quality of the synthetic data.

[0192] In this embodiment, a device for generating question-and-answer data is also provided. The device is used to implement the above-mentioned embodiments and preferred implementations, and the details that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0193] This embodiment provides a device for generating question-answer data, as shown in FIG3 , including:

[0194] A first acquisition module 301 is configured to acquire a plurality of cut document blocks, wherein each cut document block contains a first preset number of text data;

[0195] The first obtaining module 302 is used to obtain target question data corresponding to the cut document block based on the text data in the cut document block and the target question model;

[0196] The second acquisition module 303 is used to acquire full data, wherein the full data is content information associated with the target question data contained in a complete document composed of multiple segmented document blocks;

[0197] The second obtaining module 304 is configured to obtain target answer data corresponding to the target question data based on the segmented document blocks, the target question data, the full data, and the target answer model;

[0198] The third obtaining module 305 is used to generate question-answer data according to the target question data and the target answer data.

[0199] In some optional embodiments, the device further comprises:

[0200] A first determining module is configured to determine a questioning strategy based on the text data in the cut document block before obtaining target question data corresponding to the cut document block based on the text data in the cut document block and the target question model;

[0201] An extraction module is used to process text data according to the questioning strategy and extract multiple questions;

[0202] A third acquisition module is used to obtain first scores for the multiple questions according to the question scoring model;

[0203] A second determining module is configured to determine a plurality of candidate questions from the plurality of questions according to the first scores;

[0204] The fourth obtaining module is used to optimize the initial question model according to the candidate questions to obtain the target question model.

[0205] In some optional implementations, the first determining module includes:

[0206] a first determining unit, configured to determine a text type of the text data;

[0207] The second determining unit is used to determine a corresponding questioning strategy according to the text type.

[0208] In some optional embodiments, the device further comprises:

[0209] a third determination module for determining an answer strategy based on the text data in the segmented document blocks and the target question data before obtaining target answer data corresponding to the target question data based on the segmented document blocks, the target question data, the full data, and the target answer model;

[0210] a fifth obtaining module, configured to process the text data according to the answer strategy to obtain a plurality of answer data;

[0211] A fourth acquisition module, configured to acquire a second score for the plurality of answer data according to the answer scoring model;

[0212] a fourth determining module, configured to determine a plurality of candidate answer data from the plurality of answer data according to the second score;

[0213] The sixth obtaining module is used to optimize the initial answer model based on the candidate answer data to obtain the target answer model.

[0214] In some optional implementations, the third determining module includes:

[0215] a third determining unit, configured to determine a text type of the text data;

[0216] The fourth determining unit is used to determine the answer strategy according to the text type and the target question data.

[0217] In some optional implementations, the second obtaining module 304 includes:

[0218] An obtaining unit, configured to obtain a plurality of answer data according to the segmented document blocks, the target question data, the full amount of data, and the target answer model;

[0219] Split unit, used to split the answer data into the smallest unit to obtain multiple target fields;

[0220] a comparing unit, configured to compare description contents of the target field in each second preset number of answer data;

[0221] A retaining unit, configured to retain the description content corresponding to the answer data with the second highest score;

[0222] The integration unit is used to integrate the description content to obtain target answer data.

[0223] In some optional embodiments, the device further comprises:

[0224] A scoring module is used to score the question and answer data using the question and answer data scoring model after obtaining the question and answer data about the cut document block, so as to obtain a third score for the question and answer data;

[0225] A cleaning module is used to clean the question and answer data according to the third score to obtain target question and answer data that meets preset requirements;

[0226] The expansion module is used to expand the data sample according to the target question and answer data to obtain a third preset number of question and answer data.

[0227] In some optional implementations, the cut document block includes text data after image recognition.

[0228] The question and answer data generation device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0229] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0230] The embodiment of the present disclosure further provides a computer device having the apparatus for generating question and answer data shown in FIG3 .

[0231] Please refer to Figure 4, which is a structural diagram of a computer device provided by an optional embodiment of the present disclosure. As shown in Figure 4, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 takes a processor 10 as an example.

[0232] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0233] The memory 20 stores instructions that can be executed by at least one processor 10, so as to enable at least one processor 10 to execute the method shown in the above embodiment.

[0234] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created based on the use of a computer device for displaying a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0235] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0236] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0237] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0238] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for generating question-answer data, comprising: Acquire a plurality of cut document blocks, wherein each of the cut document blocks contains a first preset number of text data; Obtaining target question data corresponding to the cut document block according to the text data in the cut document block and the target question model; Acquire full data, wherein the full data is content information associated with the target question data and contained in a complete document composed of a plurality of the cut document blocks; Obtaining target answer data corresponding to the target question data according to the cut document blocks, the target question data, the full amount of data and the target answer model; The question and answer data is generated according to the target question data and the target answer data.

2. The method according to claim 1, wherein Before obtaining the target question data corresponding to the cut document block according to the text data in the cut document block and the target question model, the method further includes: Determining a questioning strategy based on the text data in the cut document block; Processing the text data according to the questioning strategy to extract multiple questions; Obtaining first scores for the plurality of questions according to a question scoring model; Determining a plurality of candidate questions from the plurality of questions according to the first scores; The initial question model is optimized according to the candidate question to obtain the target question model.

3. The method according to claim 2, wherein: The step of determining a questioning strategy based on the text data in the cut document block includes: Determining a text type of the text data; The corresponding questioning strategy is determined according to the text type.

4. The method according to claim 1, wherein: Before obtaining target answer data corresponding to the target question data according to the cut document blocks, the target question data, the full amount of data and the target answer model, the method further includes: Determining an answer strategy based on the text data in the cut document block and the target question data; Processing the text data according to the answer strategy to obtain a plurality of answer data; Obtaining a second score for the plurality of answer data according to the answer score model; determining a plurality of candidate answer data from the plurality of answer data according to the second score; The initial answer model is optimized according to the candidate answer data to obtain the target answer model.

5. The method according to claim 4, wherein: The step of determining the answer strategy according to the text data in the cut document block and the target question data includes: Determining a text type of the text data; The answer strategy is determined according to the text type and the target question data.

6. The method according to claim 4, wherein: The step of obtaining target answer data corresponding to the target question data according to the cut document blocks, the target question data, the full amount of data and the target answer model includes: Obtaining a plurality of answer data according to the cut document blocks, the target question data, the full amount of data and the target answer model; Splitting the answer data into minimum units to obtain multiple target fields; comparing the description contents of the target fields in each second preset number of the answer data; retaining the description content corresponding to the second highest-scoring answer data; The description contents are integrated to obtain the target answer data.

7. The method according to any one of claims 1 to 6, wherein: After obtaining the question-answer data about the cut document block, the method further includes: Scoring the question and answer data using a question and answer data scoring model to obtain a third score for the question and answer data; Performing data cleaning on the question and answer data according to the third score to obtain target question and answer data that meets preset requirements; Data samples are expanded according to the target question and answer data to obtain a third preset number of question and answer data.

8. The method according to any one of claims 1 to 6, wherein: The cut document block includes text data after image recognition.

9. A device for generating question and answer data, comprising: The first acquisition module is used to acquire a plurality of cut document blocks, wherein each of the cut document blocks contains Contains a first preset number of text data; A first obtaining module is used to obtain target question data corresponding to the cut document block according to the text data in the cut document block and the target question model; A second acquisition module is used to acquire full data, wherein the full data is content information associated with the target question data contained in a complete document composed of a plurality of the cut document blocks; A second obtaining module is used to obtain target answer data corresponding to the target question data according to the cut document blocks, the target question data, the full amount of data and the target answer model; The third obtaining module is used to generate the question and answer data according to the target question data and the target answer data.

10. A computer device comprising: The memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for generating question and answer data of any one of claims 1 to 7 by executing the computer instructions.

11. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the method for generating question and answer data according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Question and answer pair generation method and device, and server

    CN110532369A

  • Text-based question and answer method and device, computer equipment and storage medium

    CN114817478A

  • Question and answer pair generation method and device, electronic equipment and computer storage medium

    CN115114416A

  • Question and answer data generation method and device, computer equipment and storage medium

    CN117493508A

Cited By

  • Policy knowledge question and answer method and device, electronic equipment and storage medium

    CN120744069A

  • Synthetic data set construction method and electronic equipment

    CN120975247A

  • Strategy model training method and device, medium and equipment

    CN120996205A