Construction method and device of voice question and answer data set, medium and product

By constructing the OKE-SQA dataset and combining it with a large language model and a retrieval-enhanced generation module, we address the problem of insufficient real semantics and complexity in existing voice question-answering datasets, and achieve efficient evaluation of voice question-answering systems and improved practical application capabilities.

CN120708598APending Publication Date: 2025-09-26TRUE SPACE (ZHUHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511039488.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing voice question-answering datasets lack real semantics and have limited question complexity, making them unable to effectively evaluate knowledge enhancement and temporal reasoning capabilities, limiting their application in practical scenarios.

Method used

The OKE-SQA dataset is constructed. By acquiring various speech clips, original questions, and key questions, it combines a large language model with a retrieval-enhanced generation module to generate target answers. It covers multi-source and interactive question-answering, encompassing multiple categories such as daily life, sports, history, and culture, and acquires external knowledge through web search.

Benefits of technology

It provides a more realistic and challenging evaluation benchmark, improves the evaluation effect of the voice question answering system in knowledge reasoning and temporal reasoning capabilities, and improves the applicability and accuracy of the model in actual scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708598A_ABST
    Figure CN120708598A_ABST
Patent Text Reader

Abstract

The invention provides a voice question and answer data set construction method and device, a medium and a product, and the method comprises the following steps: obtaining voice segments of preset lengths of various types; obtaining an original question, a key question and a question answer corresponding to each voice segment; wherein in each voice segment, the key problem is obtained based on the original problem and the voice segment; and for each voice segment, storing the voice segment and the original question, the key question and the question answer corresponding to the voice segment as a data set sample to obtain a voice question and answer data set. According to the invention, the method achieves the construction of the data set of voices in the real world, questions needing to be answered by external knowledge and answers of the questions, provides a more realistic and challenging reference for evaluating open fields, multi-source and interactive questions and answers, and effectively solves the limitation of an existing SQA data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice question answering technology, and in particular to a method, device, medium and product for constructing a voice question answering dataset. Background Art

[0002] Traditional text-based question answering (Text QA) focuses solely on understanding and extracting answers from clean text input. Spoken Question Answering (SQA) aims to integrate speech and text modalities, enabling models to use speech input to perform question answering tasks in more natural and realistic scenarios. Compared to Text QA, SQA, by introducing raw speech signals, can be closer to real-world applications, providing higher practical value and research significance. Therefore, the development of SQA contributes to the advancement of multimodal information processing systems and improves speech understanding, reasoning, and real-time response capabilities in speech question answering. See Figure 1 In the SQA task, only questions about the audio content and a piece of audio enter the large language model. The large language model understands the audio content and answers the questions only about the audio content, giving the target answer.

[0003] Existing datasets for SQA tasks also show shortcomings, limiting their application in real-world scenarios. They mainly face two limitations: (1) Lack of real semantics: Many widely used SQA datasets rely on synthetic speech generated by text-to-speech (TTS) systems. Although the quality of TTS has improved significantly, it still cannot capture the diversity and complexity of speech. (2) Limited question complexity: Most existing questions can be answered directly from the speech content and rarely involve real-time or knowledge-intensive reasoning. However, real-world SQA usually requires commonsense reasoning, contextual understanding, and domain expertise, which are mostly missing in current datasets. The above limitations make existing SQA benchmarks unable to fully evaluate key real-world capabilities such as knowledge enhancement and temporal reasoning. Summary of the Invention

[0004] The first purpose of the present invention is to provide a method for constructing a voice question answering dataset, which can improve the training and evaluation effects on voice question answering tasks.

[0005] The second object of the present invention is to provide a computer device for implementing the above-mentioned method for constructing a voice question and answer dataset.

[0006] The third object of the present invention is to provide a readable storage medium for implementing the above-mentioned method for constructing the voice question and answer dataset.

[0007] A fourth object of the present invention is to provide a computer program product that implements the method for constructing the above-mentioned voice question and answer dataset.

[0008] In order to achieve the above-mentioned first purpose, the present invention provides a method for constructing a voice question and answer dataset, which includes the following steps: obtaining voice segments of preset lengths for various categories; obtaining the original question, key question, and question answer corresponding to each voice segment; wherein, in each voice segment, the key question is obtained based on the original question and the voice segment; for each voice segment, the voice segment and the original question, key question, and question answer corresponding to the voice segment are stored as a dataset sample to obtain a voice question and answer dataset.

[0009] As can be seen from the above scheme, the present invention realizes the construction of a dataset of real-world speech and questions and their answers that require external knowledge, provides a more realistic and challenging benchmark for evaluating open-domain, multi-source and interactive question answering, and effectively addresses the limitations of existing SQA datasets.

[0010] A further scenario is that the original problem is of real-time type or non-real-time type.

[0011] A further solution is to preset the length to be 3 to 10 seconds.

[0012] A further solution is that the categories include daily life, sports and competitions, history and culture, companies and products, professional fields, and unclassified fields.

[0013] A further solution is to obtain a voice clip by intercepting an online video, where the voice clip includes the speaker's accent and background noise.

[0014] A further solution is that, in the voice question-answering dataset, among the dataset samples with a preset proportion, at least one network search is required to obtain the answer to the target question based on the target question.

[0015] A further solution is that each dataset sample also includes a score, where the score is used to evaluate the number of web search steps required to answer the original question of the dataset sample.

[0016] In order to achieve the above-mentioned second purpose, the present invention provides a computer device, including a processor and a memory, wherein: a computer program is stored on the memory, and when the computer program is executed by the processor, the above-mentioned method for constructing a voice question and answer dataset is implemented.

[0017] In order to achieve the third purpose mentioned above, the present invention provides a computer-readable storage medium on which a computer program is stored, wherein: when the computer program is executed by a processor, the method for constructing the above-mentioned voice question and answer dataset is implemented.

[0018] In order to achieve the fourth objective mentioned above, the present invention provides a computer program product, including computer instructions, wherein: when the computer instructions are executed by a processor, the above-mentioned method for constructing a voice question and answer dataset is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a framework diagram of a voice question answering system based on retrieval enhancement generation in an embodiment of the method for constructing a voice question answering dataset of the present invention.

[0020] Figure 2 It is a flowchart of the voice question answering method implemented by the voice question answering system based on retrieval enhancement generation in the embodiment of the method for constructing a voice question answering dataset of the present invention.

[0021] Figure 3 It is a schematic diagram of a voice question answering method implemented by a voice question answering system based on retrieval enhancement generation in an embodiment of a method for constructing a voice question answering dataset of the present invention.

[0022] Figure 4 2 is a schematic diagram of the construction process of the OKE-SQA dataset in an embodiment of the method for constructing a speech question answering dataset of the present invention.

[0023] Figure 5 2 is a schematic diagram of the categories, proportions, and examples of questions in each category of the OKE-SQA dataset in an embodiment of the method for constructing a voice question answering dataset of the present invention.

[0024] Figure 6 3. This is a schematic diagram of the score distribution of different categories of real-time and non-real-time questions in the OKE-SQA dataset in an embodiment of the method for constructing a voice question answering dataset of the present invention.

[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments. DETAILED DESCRIPTION

[0026] The voice question-answering system based on retrieval enhancement generation of the present invention can understand the audio content and answer the questions only about the audio content when it obtains audio content and questions only about the audio content, and output the corresponding target answers; when it obtains audio content and questions that require external knowledge or are time-sensitive, it can answer the questions through the detected additional knowledge and output the corresponding target answers.

[0027] Example of a method for constructing a voice question answering dataset: See also Figure 1The voice question answering system based on retrieval enhancement generation of this embodiment includes an input acquisition module 11, a text transcription module 12, a large language model module 13, and a retrieval enhancement generation module 14. The input acquisition module 11 is connected to the text transcription module 12 and the large language model module 13 respectively, the text transcription module 12 is connected to the large language model module 13, and the large language model module 13 is connected to the retrieval enhancement generation module 14.

[0028] The input acquisition module 11 is used to acquire the audio to be processed and the original question corresponding to the audio to be processed. The original question corresponding to the audio to be processed is a question that requires understanding of the audio to be processed.

[0029] The text transcription module 12 is used to transcribe the acquired data to be processed into text based on a pre-selected ASR model to obtain a text transcription result.

[0030] The ASR model in this embodiment is preferably the Paraformer-zh model. As a non-autoregressive ASR model, Paraformer-zh efficiently converts speech to text. Compared to traditional autoregressive methods, it offers advantages such as high transcription speed, strong robustness, and reduced error propagation. In other embodiments, other ASR models may be selected based on actual needs.

[0031] The large language model module 13 is used to generate key questions corresponding to the original questions, and generate target answers corresponding to the key questions by combining multiple paragraphs semantically related to the key questions obtained from the retrieval enhancement generation module 14.

[0032] Specifically, the large language model module 13 includes a key question generation unit and a question solving unit. The key question generation unit generates a key question corresponding to the original question based on the first large language model, combining the text transcription results, the original question, and the first preset prompt word; the question solving unit generates an answer corresponding to the key question based on the second large language model, combining multiple paragraphs semantically related to the key question, the key question, and the second preset prompt word. The key question corresponding to the audio to be processed is a question that does not rely on the understanding of the audio to be processed.

[0033] The key question is a well-structured and semantically complete question generated by the first language model based on the text transcription results, the original question, and the first preset prompt word. By generating the key question, colloquial, ambiguous, or fragmented input can be converted into a standardized, search-friendly query basis, ensuring that multiple paragraphs semantically relevant to the key question can be accurately retrieved in the search enhancement generation module 14.

[0034] Optionally, in different embodiments, the first large language model and the second large language model can be different large language models, or can be the same large language model. In this embodiment, they are preferably the same large language model.

[0035] Optionally, in different embodiments, the first large language model and / or the second large language model may be directly set in the text transcription module 12, or may be an external large language model called through an API call.

[0036] The retrieval enhancement generation module 14 is based on the retrieval enhancement generation (RAG) technology and is used to search the key question by accessing an external search engine to obtain multiple paragraphs semantically related to the key question.

[0037] Based on the above-mentioned voice question answering system based on retrieval enhancement generation, a voice question answering method based on retrieval enhancement generation is implemented. Figure 2 and Figure 3 , which is implemented based on the execution of a computer program. Each step will be introduced below.

[0038] First, step S11 is performed to obtain the audio to be processed and the original question corresponding to the audio to be processed.

[0039] Then, step S12 is executed to transcribe the audio to be processed using a preset ASR model to obtain a text transcription result.

[0040] Then, step S13 is executed to input the text transcription result, the original question, and the first preset prompt word into the first large language model to obtain the key question output by the first large language model.

[0041] Then, step S14 is executed to search for the key question by accessing an external search engine to obtain multiple paragraphs semantically related to the key question.

[0042] Finally, step S15 is executed to input multiple paragraphs semantically related to the key question, the key question, and the second preset prompt word into the second largest language model to obtain the target answer corresponding to the key question output by the second largest language model.

[0043] To validate the aforementioned retrieval-enhanced speech question answering system and method, the following describes a detailed experimental process. Specifically, we first constructed a speech question answering dataset (hereinafter referred to as the OKE-SQA dataset) for evaluating the performance of RAGs in large language models. We then compared the OKE-SQA dataset with other approaches to SQA tasks.

[0044] The OKE-SQA dataset is specifically designed for speech tasks that combine LLM and RAG, and has the following properties: (1) Questions are not limited to the explicit content of the speech. Instead, the questions require the model to understand the speech content and then utilize external knowledge to derive answers; (2) Each question has a single, well-defined answer to facilitate reliable evaluation, and both questions and answers are concise and clear; (3) To enhance applicability in real-world scenarios, the speech is collected from real sources, so that realistic speech characteristics, including accents and background noise, can be captured.

[0045] See also Figure 4 ,The construction process of the OKE-SQA dataset includes four steps, each of which will be introduced below.

[0046] First, step S21 is executed to classify the data set.

[0047] Among them, see Figure 4 , the questions are divided into six different categories, including daily life, sports and competitions, history and culture, companies and products, professional fields, and others (i.e., unclassified fields). Each category contains real-time questions (the answers change over time) and non-real-time questions (the answers do not change over time). Regardless of whether the question is real-time or non-real-time, it may require external knowledge to answer the question. Distinguishing between real-time and non-real-time questions also facilitates targeted updates to the answers to the questions in the dataset. Human annotators were instructed to maintain a roughly even ratio between these two question types.

[0048] Then, step S22 is executed: task definition and allocation.

[0049] Each manual annotator is assigned one or two categories and is accompanied by sample question prompts. Before creating a question, the manual annotator needs to determine the category of their main interest, search for online videos related to their category, and record the start and end timestamps of the online video, as well as the video link. In addition, the manual annotator needs to ensure that the length of the extracted start and end timestamps is 3 to 10 seconds, and the main topic mentioned in the question must clearly appear in the video clip. The online video corresponding to the video link should include both audio and video information. The online video can be obtained from an existing online video website. In this embodiment, all online videos are online videos.

[0050] Then, step S23 is executed to determine the original question, the answer to the question, and the key question.

[0051] Among them, the corresponding original questions, answers and key questions designed by the human annotators based on the online videos they searched are obtained, and their expressions are ensured to be clear and unambiguous. When drafting questions and answers, the human annotators need to provide an original question and a key question. The original question cannot reveal the key speech content to ensure that the model must interpret the speech clip. Instead, the key question is replaced with the key terms explicitly mentioned in the speech, making it directly applicable to the retrieval of external knowledge. For example, for an example of a real-time question, the original question is "How many major caves are there in the grotto mentioned in the audio?", and the corresponding key question is "How many major caves are there in the Yungang Grottoes?"; for an example of a non-real-time question, the original question is "How many dead skin cells does the object mentioned in the audio shed per second?", and the corresponding key question is "How many dead skin cells does the human body shed per second?".

[0052] Furthermore, a variety of expression styles can be employed in constructing original questions of different categories to avoid overly homogenous generated content and thus cover a wider range of possible scenarios. This diversity enables the LLM fine-tuned on the OKE-SQA dataset to better understand human questioning and answering styles, thereby generating output that is closer to actual user needs.

[0053] Finally, step S24 is executed to organize the data and obtain the OKE-SQA dataset.

[0054] This involves extracting a speech clip from the start and end timestamps of an online video. This speech clip, along with its corresponding original question, answer, and key question, is stored as a data sample. All acquired data samples are categorized and stored to form the OKE-SQA dataset. The OKE-SQA dataset provides a more challenging and realistic SQA environment, requiring external knowledge to effectively answer questions.

[0055] The OKE-SQA dataset is the first speech-based dataset specifically designed to evaluate the performance of RAGs in large language models. See Table 1 for statistics of the OKE-SQA dataset.

[0056] Table 1. Statistics of the OKE-SQA dataset

[0057] The OKE-SQA dataset contains approximately 1,600 original questions, each with a corresponding key question and answer. The OKE-SQA dataset covers six categories, of which approximately 28% are real-time questions (453 questions), and the remaining 72% are non-real-time questions (1,174 questions). The OKE-SQA dataset includes a variety of categories: Daily Life (6%), Sports and Competitions (17%), History and Culture (22%), Companies and Products (18%), Professional Fields (12%), and Other (i.e., unclassified fields, which account for 25%). Question length averages 19.45 words, with the longest question reaching 45 words. Answer length averages 4.97 words, with the longest being 68 words. The total speech duration in the OKE-SQA dataset is 11,856.35 seconds. These statistics demonstrate that the dataset is well-balanced, focusing on real-time questions while ensuring sufficient coverage of all categories. The wide range in question lengths highlights the diversity of content, which is crucial for assessing LLM performance in relation to the RAG.

[0058] See also Figure 5 , Figure 5 Examples of categories and proportions in the OKE-SQA dataset and examples of questions in each category are shown.

[0059] To further assess the difficulty of the OKE-SQA dataset, we also annotated the scores of each original question by more than 20 human annotators. Each original question was scored based on the number of retrieval steps required to answer it, using the following criteria: (1) score 1: can be answered without any external knowledge; (2) score 5: requires one web search; (3) score 8: requires two web searches; (4) score 10: requires three or more web searches. See Table 2, which shows a summary of the scores of real-time and non-real-time questions in the OKE-SAQ dataset.

[0060] Table 2. Summary of all OKE-SAQ scores

[0061] In the overall distribution of difficulty scores, 30,330 questions received a score of 5, representing 85% of the 35,794 questions, highlighting the significant emphasis placed on retrieval-based reasoning in the OKE-SQA dataset. Furthermore, 7.4% of questions received a high score (8 or 10), indicating the presence of multiple reasoning-intensive queries requiring complex information integration and reasoning.

[0062] Combined with Table 2 and Figure 6, which shows the score distribution of each question type in different categories. Score 5 dominates all categories, confirming that the OKE-SQA dataset is mainly composed of retrieval-based questions. Importantly, Table 2 and Figure 6 The distribution of real-time and non-real-time questions at different difficulty levels is also shown. Temporal dynamics also plays a key role in task complexity. As shown in Table 2, 28% of all questions are real-time, appearing across all difficulty levels. For example, among the highest difficulty questions (scored 10), 31% are real-time, highlighting the challenges of dynamic knowledge retrieval. However, 69% of the 10-score questions are not time-dependent, indicating that static questions still require advanced domain expertise, multi-step reasoning, and comprehensive expertise. In summary, the OKE-SQA dataset presents diverse challenges in terms of retrieval difficulty, real-time dynamics, and domain diversity. Since the OKE-SQA dataset ensures that the majority of samples in the speech question answering dataset require at least one web search to obtain the answer to the key questions, it improves the evaluation of speech question answering datasets and effectively assesses the model's capabilities in temporal reasoning and knowledge integration. Therefore, as a rigorous and comprehensive benchmark, the OKE-SQA dataset helps advance the development and evaluation of speech question answering systems.

[0063] This example introduces several baseline methods for comparison. Specifically, two tasks with different objectives are defined: (1) Generate key questions. This task is used to evaluate whether the system can accurately generate "key questions". Tasks related to RAG require the model to interpret the input question and speech content, and then infer the key information required for knowledge retrieval, and refer to this retrieval-oriented query as a "key question". Whether the system can accurately generate this question is the first key indicator for measuring performance. (2) Generate target answers corresponding to key questions. This task is used to evaluate whether the model can correctly provide target answers corresponding to key questions, and the target answers need to be as close as possible to the answer to the question corresponding to the key question. In this method, the speech is first transcribed into text and then input into a large language model. The large language model uses its contextual reasoning ability to autoregressively generate answers, predicting one word at a time. For methods that do not require automatic speech recognition, the speech is directly input into the speech language model, which captures multiple layers of features such as phonemes, rhythm, and intonation to interpret the underlying semantics. The answer is then generated in the same autoregressive manner, directly completing the conversion process from speech to text.

[0064] The baselines are divided into two parts: (1) ASR-based baselines: The first category of baselines follows an ASR-based pipeline. In these methods, ASR is performed using a Paraformer model to convert the user's voice questions into text. The recognized text is then combined with predefined prompt words and input into a language model to generate answers. Specifically, the recognized text and task-specific prompts are input into the language model together. The model embeds each token (word or subword) into a high-dimensional semantic space and models contextual dependencies through multiple Transformer layers. This enables the model to understand the context, grammar, and intent of the question. Finally, the large language model generates the answer in an autoregressive manner, predicting one token at a time until a complete natural language response is formed. (2) No-ASR baselines: The second category of baselines does not use ASR. To alleviate the performance bottleneck caused by ASR errors, an end-to-end voice question answering system is adopted, which uses a speech language model Speech-LLM, which is a large language model that can directly process raw speech input. These models allow for more natural and fluent question answering by performing semantic understanding and answer generation directly from speech.

[0065] The speech language model first uses a pre-trained speech encoder to extract speech representations from the input signal, capturing intonation and pragmatic features such as speaking rate, rhythm, and pauses. These speech features are then mapped into the language modeling space. When the task includes additional textual context, the model jointly encodes this information with the speech features to achieve cross-modal semantic integration within a unified representation space. Once the speech is processed and the text prompt and question are received, the speech language model autoregressively generates the answer. This generation process is similar to that of text-based language models, but its foundation lies in the speech encoding space.

[0066] In the experimental setup, the pre-configured ASR model for this example is Paraformer-zh. For the ASR baseline, three text-only LLMs were selected: Qwen2.5-7B Instruct, GLM-Zero-Preview, and ChatGLM3-6B. For baselines that do not require ASR, Qwen2-Audio-7B-Instruct and GLM-4-Voice were selected. These two language models can handle both speech and text. For the web search used in the benchmark, the web-searchpro tool was selected.

[0067] Optionally, the GLM-4-PLUS used in the voice question answering system of this embodiment can be fine-tuned on the OKE-SQA dataset. If the GLM-4-PLUS used in the voice question answering system of this embodiment is fine-tuned on the OKE-SQA dataset, the fine-tuning is done on a portion of the OKE-SQA dataset, and the portion of the OKE-SQA dataset used in the following comparison with other baseline methods will not overlap with this portion.

[0068] In determining the evaluation metrics, four metrics were used to comprehensively evaluate the quality of the generated text: (1) BLEU: This is used to measure vocabulary accuracy by evaluating local matching through n-gram precision. (2) ROUGE-N: This emphasizes recall to measure the coverage of key content relative to the reference text, thereby evaluating content completeness. (3) F1-Recall: This balances precision and recall, reducing the bias caused by a single metric, making it very suitable for question answering tasks. (4) BERTScore: This captures deeper semantic similarity by considering lexical substitutions and syntactic changes, complementing the surface-level focus of the n-gram metric.

[0069] By combining traditional statistical methods (BLEU, ROUGE) and deep learning-based metrics (BERTScore), and incorporating the F1 balance mechanism, a more robust evaluation of both formal accuracy and semantic fidelity is performed.

[0070] See Table 3, which shows the performance of the method of this embodiment and other baseline methods in generating key questions.

[0071] Table 3. Performance of the method in this example and other baseline methods in generating key questions

[0072] The method of this embodiment (ours) consistently achieved the best performance across all key evaluation metrics. Specifically, the F1-Recall metric achieved a 47.3% improvement over Qwen2.5-7B-Instruct, highlighting its superior ability to preserve and replace key factual elements during question rewriting. The ROUGE-1 metric saw a 21.8% improvement, reflecting enhanced lexical fidelity and surface form alignment. Compared to GLM-Zero-Preview, the BLEU metric score of this embodiment's method increased by 159.6%, from 3.2756 to 8.5040. For the F1-Recall metric, the method of this embodiment achieved a score of 37.34%, while GLM-4-Voice achieved a mere 2.14%. These significant improvements demonstrate that the method of this embodiment is not only able to identify key entities in spoken queries, but also accurately rephrase them into standard form in the rewritten question.

[0073] See Table 4, which shows the performance of the method of this embodiment and other baseline methods in generating target answers.

[0074] Table 4. Performance of the method in this example and other baseline methods in generating target answers

[0075] Table 4 shows the results for generating target answers to key questions. Our solution (our RagSQA) achieves top performance across all evaluation metrics, particularly in terms of ground truth accuracy and content coverage. Compared to Qwen2.5-7B-Instruct, the F1-Recall metric improves by 167.1%, from 11.99% to 32.03%. The ROUGE-1 metric also improves significantly, from 6.80% to 20.49%, equivalent to a 201.3% increase, indicating greater lexical overlap with the correct answer. Compared to GLM-Zero-Preview, the BLEU metric improves by over eightfold, from 1.3300 to 12.0619. This demonstrates that RagSQA is not only factually accurate but also more fluent and grammatically consistent. GLM-4-Voice's F1-Recall metric is only 0.06%, indicating a near-inability to retrieve the correct information. In contrast, our approach achieves 32.03%, over 530 times higher than the previous approach, demonstrating a stronger ability to identify relevant content from the input.

[0076] To isolate the effect of retrieval, we compare the performance of GLM-4-Plus without RAG with RagSQA of this example. After adding RAG, the F1-Recall metric improved by 119.1%, the ROUGE-1 metric improved by 165.7%, and the BLEU metric improved by 227.6%. These results show that without retrieval, the model has difficulty handling time-sensitive or domain-specific queries, often resulting in inaccurate or incomplete answers. However, by introducing RAG, the model is able to answer high-complexity questions that require this dynamic external information. This improvement emphasizes the importance of retrieval-enhanced large language models in handling knowledge-intensive SQA tasks in the real world.

[0077] In order to intuitively compare the difficulty of the OKE-SQA dataset with existing datasets, an experiment is also conducted to evaluate the performance of the baseline model on multiple datasets, as shown in Table 5. Table 5. Performance of different models on the OKE-SQA dataset and other baseline datasets

[0078] Table 5 shows performance on different datasets, with lower values ​​indicating greater dataset difficulty. The performance of each baseline model on the OKE-SQA dataset drops significantly across multiple metrics, highlighting its increased complexity and the limitations of current models' generalization. While exceeding 50% on LibriSQA and Spoken-SQuAD, these metrics drop by over 70% on OKE-SQA, with some models experiencing reductions of up to 90%. This indicates serious challenges in vocabulary alignment and information matching. The degradation of the F1-Score and F1-Recall metrics is even more pronounced: models that achieve F1-Score and F1-Recall on Spoken-SQuAD drop below 1% on OKE-SQA, indicating a critical failure in both answer accuracy and recall.

[0079] Two main factors contribute to this challenge. First, OKE-SQA contains a large number of real-time questions, whose answers change dynamically over time. Second, many questions require external knowledge, raising the bar for understanding and reasoning. Compared to existing Chinese SQA datasets, OKE-SQA imposes higher requirements in terms of both time and knowledge base. It reveals the limitations of current SQA models in real-world applications and provides a new benchmark for developing more cognitively capable and knowledge-adaptive systems.

[0080] In addition, although the OKE-SQA dataset was constructed to evaluate the performance of complex tasks in the field of voice question answering that require retrieving external knowledge, the OKE-SQA dataset sets standard key questions (in text form) as a reference, so it can also be extended to tasks such as text question answering, and is not limited to the field of voice question answering.

[0081] In summary, the present invention realizes the construction of a dataset of real-world speech and questions and their answers that require external knowledge, provides a more realistic and challenging benchmark for evaluating open-domain, multi-source, and interactive question answering, and effectively addresses the limitations of existing SQA datasets.

[0082] Computer readable storage medium embodiment: If the modules integrated into the computer device of the above embodiment are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the process of the embodiment of the method for constructing a voice question and answer dataset can also be completed by a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a controller, it can implement the steps of the embodiment of the method for constructing a voice question and answer dataset. Among them, the computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0083] Computer program product embodiment: The computer program product of this embodiment includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform each step of the embodiment of the method for constructing a voice question-answering dataset described above.

[0084] Finally, it should be emphasized that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for constructing a voice question answering dataset, characterized in that: The following steps are involved: Get speech segments of preset length for each category; Obtaining the original question, key question, and answer corresponding to each voice segment; wherein, in each voice segment, the key question is obtained based on the original question and the voice segment; For each of the voice segments, the voice segment and the original question, the key question, and the answer to the question corresponding to the voice segment are stored as a data set sample to obtain a voice question and answer data set.

2. The method for constructing a voice question-answering dataset according to claim 1, wherein: The original question is of real-time type or non-real-time type.

3. The method for constructing a voice question-answering dataset according to claim 1, wherein: The preset length is 3 to 10 seconds.

4. The method for constructing a voice question-answering dataset according to claim 1, wherein: The categories include daily life, sports and competitions, history and culture, companies and products, professional fields, and unclassified fields.

5. The method for constructing a voice question-answering dataset according to claim 1, wherein: The voice clip is obtained by intercepting an online video, and the voice clip includes the speaker's accent and background noise.

6. The method for constructing a voice question-answering dataset according to claim 1, wherein: In the dataset samples with a preset proportion of the voice question and answer dataset, at least one network search is required to obtain corresponding answers to key questions in the dataset samples.

7. The method for constructing a voice question-answering dataset according to claim 6, wherein: Each of the dataset samples further includes a score, which is used to evaluate the number of web search steps required to answer the original question of the dataset sample.

8. A computer device comprising a processor and a memory, characterized in that: The memory stores a computer program, which, when executed by the processor, implements the method for constructing a voice question and answer dataset as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a voice question and answer dataset according to any one of claims 1 to 7 is implemented.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the method for constructing a voice question and answer dataset as described in any one of claims 1 to 7 is implemented.