Intelligent question and answer method, device, medium and electronic equipment based on large model
By collecting audio data in a voice-based human-computer interaction system and using a question extraction model to remove noise, accurate response information is generated, solving the problem of inaccurate responses caused by environmental noise interference and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING VOLCANO ENGINE TECH CO LTD
- Filing Date
- 2026-04-16
- Publication Date
- 2026-07-10
AI Technical Summary
In voice-based human-computer interaction systems, environmental noise interference can reduce the accuracy of generated response information, thus affecting the user experience.
By collecting audio data for text recognition, historical dialogue information is obtained, and a question extraction model is used to remove noise from the initial text to generate target content for the large model to generate response information.
It improved the accuracy of response information and enhanced the usability of the human-computer interaction system.
Smart Images

Figure CN122369459A_ABST
Abstract
Description
Technical Field
[0001] This content relates to the field of artificial intelligence technology, specifically to an intelligent question-answering method, device, medium, and electronic device based on a large model. Background Technology
[0002] With the development of artificial intelligence technology, voice-based human-computer interaction systems have become indispensable tools in people's lives. These systems primarily recognize the speaker's speech as text and feed it into a large model, which then responds to the text. However, noise interference in the speech content can significantly affect the accuracy of the generated response. Summary of the Invention
[0003] This content section is provided to briefly introduce the concepts, which will be described in detail in the examples section later. This content section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] Firstly, a large-model-based intelligent question-answering method is provided, the method comprising: Acquire audio data and perform text recognition on the audio data to obtain initial text; The historical dialogue information corresponding to the audio data is obtained, and the historical dialogue information and the initial text are input into the question extraction model to obtain the target content. The question extraction model is used to remove noise content in the initial text. The amount of historical dialogue information corresponding to the historical dialogue data is determined based on the environmental noise complexity. The response information corresponding to the target content is generated based on the large model.
[0005] Secondly, a large-model-based intelligent question-answering device is provided, the device comprising: The recognition module is configured to collect audio data and perform text recognition on the audio data to obtain initial text; The processing module is configured to acquire historical dialogue information corresponding to the audio data, and input the historical dialogue information and the initial text into the question extraction model to obtain the target content. The question extraction model is used to remove noise content in the initial text. The amount of historical dialogue information corresponding to the historical dialogue data is determined based on the environmental noise complexity. The generation module is configured to generate response information corresponding to the target content based on the large model.
[0006] Thirdly, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0007] Fourthly, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0008] Fifthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0009] By employing the above technical solution, when collecting audio data, noise removal can be performed on the initial text corresponding to the audio data to ensure the accuracy of the generated response information. Specifically, after performing text recognition on the audio data to obtain the initial text, historical dialogue information corresponding to the audio data is acquired. Based on this, the historical dialogue information and the initial text are input into the question extraction model. The question extraction model then filters out noise information in the initial text based on the historical dialogue information to obtain clean target content. Since the target content does not contain noise interference information, more accurate response information can be generated based on this target content.
[0010] Other features and advantages of the technical solution will be described in detail in the following examples section. Attached Figure Description
[0011] The above and other features, advantages, and aspects of the technical solution will become more apparent when considered in conjunction with the accompanying drawings and the following examples. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is an example diagram illustrating a voice-based human-computer interaction system according to an exemplary embodiment.
[0012] Figure 2 This is an example diagram illustrating ambient sound collected by a radio according to an exemplary embodiment.
[0013] Figure 3 This is a flowchart illustrating an intelligent question-answering method based on a large model, according to one embodiment.
[0014] Figure 4 This is a specific example diagram illustrating the generation of response information in a large-model-based intelligent question-answering method according to an embodiment.
[0015] Figure 5This is a schematic diagram illustrating the working principle of an embedded model in a large-model-based intelligent question answering method according to an embodiment.
[0016] Figure 6 This is a specific example diagram illustrating data cleaning in a large-model-based intelligent question answering method according to one embodiment.
[0017] Figure 7 This is an example diagram illustrating LORA fine-tuning in a large-model-based intelligent question answering method according to one embodiment.
[0018] Figure 8 This is an example diagram illustrating parameter updates during model training in a large-model-based intelligent question answering method according to one embodiment.
[0019] Figure 9 This is a block diagram illustrating a large-model-based intelligent question-answering device according to one embodiment.
[0020] Figure 10 This is a schematic diagram of the structure of an electronic device according to one embodiment. Detailed Implementation
[0021] The technical solution will now be described in more detail with reference to the accompanying drawings. Although certain scenarios are shown in the drawings, it should be understood that the technical solution can be implemented in various forms and should not be construed as limited to the scenarios described herein. Rather, these scenarios are provided to provide a more thorough and complete understanding of the technical solution. It should be understood that the accompanying drawings and the scenarios described are for illustrative purposes only and are not intended to limit the scope of protection of the technical solution.
[0022] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of the technical solution is not limited in this respect.
[0023] The term "comprising" and its variations as used herein can be open-ended, meaning "including but not limited to". The term "based on" can mean "at least partially based on". The term "one case" means "at least one case"; the term "another case" means "at least one additional case"; the term "some cases" means "at least some cases". Definitions of other terms will be given in the following description.
[0024] It should be noted that the concepts of "first" and "second" mentioned here are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0025] It should be noted that the terms "one" and "more" used here are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0026] The names of messages or information exchanged between the multiple devices in the implementation are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0027] It is understandable that before using the technical solutions provided here, users should be informed of the type, scope of use, and usage scenarios of the personal information involved in accordance with relevant laws and regulations, and their permission should be obtained.
[0028] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations described herein.
[0029] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0030] It is understood that the above notification and user permission acquisition process are merely illustrative and do not constitute a limitation on the implementation of the technical solution. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the technical solution.
[0031] At the same time, it is understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and related provisions.
[0032] In related technologies, for regular text input (Query), large language models (large models) can generate high-quality answers. It can be seen that language models have gradually become an indispensable auxiliary tool in people's daily lives. Through this large language model, people's work efficiency, such as the efficiency of storytelling, writing and manuscript writing, can be greatly improved.
[0033] The input to a large text model is text; however, text input typically requires input devices such as a display screen, keyboard, and input method, making it user-unfriendly. To improve the convenience of interaction, a voice-based human-computer interaction system integrating a large language model has emerged. This system allows voice input to control the large language model to generate response content. The structure of such a system can be as follows: Figure 1 As shown, based on Figure 1 It can be seen that a voice-based human-computer interaction system may include a voice understanding module and a voice synthesis module.
[0034] Here, the speech understanding module can be used to understand speech input and generate a text answer (text reply); the speech synthesis module can be used to convert the text answer (text reply) into audio output, which can then be played through a speaker, i.e., the converted audio data is output to the user through an audio output device (audio playback device). Alternatively, the voice human-computer interaction system may not have a speech synthesis module. In this case, the text answer generated by the speech understanding module can be output through a display device, such as displaying the result text on a screen.
[0035] In some implementations, the speech understanding module can achieve speech understanding through a large multimodal speech model, or by cascading speech recognition with a large language model. However, training a large multimodal speech model is complex and requires massive amounts of audio data, making it difficult to obtain. The method of cascading speech recognition with a large language model allows the speech recognition model to understand the speech input and extract text information, which is then input into the large language model to generate a text answer. This approach fully leverages the generative capabilities of the large language model. Furthermore, because the speech recognition model only performs recognition tasks and has fewer parameters, its requirements for training data are lower than those of a large model, thus ensuring both low latency and effective training.
[0036] The input to a voice-based human-computer interaction system is an audio signal, which typically carries ambient sound (environmental signal). This ambient sound varies depending on the usage scenario, such as... Figure 2 The ambient sounds shown can include background music, noise, background speakers, television sounds, and other multimedia sounds. Based on Figure 2 It is known that the audio information collected by the microphone is very likely to contain redundant information (noise). The text result obtained by the speech recognition model from the audio data collected by the microphone not only includes the speaker's speech, but also a lot of environmental noise.
[0037] The presence of environmental noise can prevent large language models from recognizing the specific question posed by the current user (the speaker), resulting in generated responses that do not meet the user's actual needs—that is, irrelevant answers—which negatively impacts the user experience. In summary, since most voice-based human-computer interaction systems do not collect audio in recording studios or noise-free environments, the audio captured by the microphone is highly likely to contain noise, and the presence of noisy text affects the response performance of large language models.
[0038] To address the aforementioned issues, a large-model-based intelligent question-answering method, device, medium, and electronic device are proposed. This large-model-based intelligent question-answering method can filter the initial text corresponding to the audio data through historical dialogue information to obtain noise-free target content. The target content can then be used to generate accurate response information, thereby improving the accuracy of responses and enhancing the usability of the human-computer interaction system.
[0039] Figure 3 This is a flowchart illustrating a large-model-based intelligent question-answering method according to an exemplary embodiment. This large-model-based intelligent question-answering method can be applied to electronic devices with processing capabilities, such as terminals or servers. Furthermore, this large-model-based intelligent question-answering method can be executed by a large-model-based intelligent question-answering device, which can be implemented by software and / or hardware, and the software and / or hardware can be configured in the electronic device. (Refer to...) Figure 3 This intelligent question-answering method based on a large model may include the following steps.
[0040] In step S310, audio data is acquired and text recognition is performed on the audio data to obtain the initial text.
[0041] In the technical solution, the audio data can be an audio signal acquired by an audio acquisition device. The audio data can include a target signal and a noise signal. The target signal can be the audio signal of the main speaker; the noise signal can be a background speaker signal, background music signal, noise signal, television sound signal, or other multimedia external sound signal, etc.
[0042] In other words, the audio data collected by the audio acquisition device can be the audio data collected when the target user inputs the question information at the current moment. That is, the audio data includes not only the target user's voice signal (target signal) but also ambient sound (noise signal).
[0043] As an alternative approach, after acquiring the initial audio data, it can be preprocessed to obtain the target audio data. Based on this, text recognition is performed on the target audio data to obtain the initial text.
[0044] For example, during the preprocessing of initial audio data, the initial audio data can be input into an audio separation model to separate the mixed audio data. Here, the audio separation model can classify the mixed audio data according to the wavelength and / or frequency of the audio data to achieve the separation of different audio data.
[0045] In the technical solution, the classification results output by the audio classification model can include a first category and a second category. The first category data can be audio signals generated by speakers in the environment where the audio acquisition device is located, such as audio signals generated by the main speaker or background speakers. The second category data can be audio signals output by multimedia devices, such as signals output by audio playback devices, like televisions, mobile phones, speakers, computers, etc. Since the second category audio data is fundamentally different from the first category audio data, and its correlation with the target user's audio data at the current moment is extremely small, to ensure the accuracy and efficiency of response information generation, the initial audio data can be classified and filtered using the audio classification model to obtain the target audio data.
[0046] When obtaining the classification results output by the audio classification model, the audio data of the second category can be deleted from the initial audio data, and text recognition can be performed on the audio data of the first category. This can not only improve the efficiency of subsequent text recognition, but also ensure the accuracy of the generated response information.
[0047] It should be noted that before inputting the initial audio data into the audio classification model, a white noise filtering operation can be performed on the initial audio data to remove background noise from the initial audio data. For example, the audio data in a specified frequency range in the initial audio data can be filtered. Then, the filtered initial audio data can be input into the audio classification model.
[0048] After filtering the target audio data from the audio data using an audio classification model, a speech-to-text operation can be performed, that is, text recognition is performed on the target audio data to obtain the initial text. For example, the target audio data can be input into a speech recognition model to obtain the initial text, which may include the target content corresponding to the target user and the noise content corresponding to the ambient sound.
[0049] In other words, a speech recognition model can extract noisy text containing ambient sounds from the audio data input in the current round. This noisy text can consist of target content and noise content. Specifically, the speech recognition model can be used to convert speech (audio data) into text (initial text).
[0050] As an example, the initial text could be "Oppenheimer is a good movie, shall we go out to dinner tomorrow? Are there any other assistants for Holmes besides Watson?", where "Oppenheimer is a good movie, shall we go out to dinner tomorrow?" is the noise content, and "Are there any other assistants for Holmes besides Watson?" is the target content.
[0051] In step S320, the historical dialogue information corresponding to the audio data is obtained, and the historical dialogue information and the initial text are input into the question extraction model to obtain the target content.
[0052] As an alternative approach, after acquiring the audio data, the corresponding historical dialogue information can be obtained. Here, the historical dialogue information can be the record of the target user's previous N interactions with the voice-based human-computer interaction system in the current session window. Since the target user's input data is speech, the historical speech can be converted into text using a speech recognition model.
[0053] In other words, historical dialogue information can be historical dialogue text, which can be obtained through speech recognition of historical speech (historical audio data). The historical dialogue information can be contextual information of the initial text, meaning that the two are related in content.
[0054] For example, the historical dialogue information of the initial text "Oppenheimer is a good movie, shall we go out for dinner tomorrow? Besides Watson, are there any other assistants for Sherlock Holmes?" could be: "user: Can you tell me a story? assistant: Sure, what kind of story would you like to hear? Science fiction, fairy tale, or detective story? user: Detective story." Here, user refers to the user and assistant refers to the machine.
[0055] It should be noted that before obtaining the historical dialogue information corresponding to the audio data, the complexity of the audio data can be determined, which can be the complexity of the environmental noise. The amount of historical dialogue information corresponding to this complexity is then determined, and the historical dialogue information corresponding to the audio data is obtained based on this amount of information. Here, the complexity of the audio data is positively correlated with the amount of historical dialogue information, which represents the number of dialogue rounds; the more dialogue rounds, the greater the amount of historical dialogue information. Therefore, different levels of environmental noise complexity correspond to different amounts of historical dialogue information (dialogue rounds), meaning there is a corresponding relationship between environmental noise complexity and the amount of historical dialogue information, and the two can be positively correlated.
[0056] In other words, the more complex the environmental noise, the greater the amount of historical dialogue information (dialogue rounds); conversely, the simpler the environmental noise, the smaller the amount of historical dialogue information. For example, when the environmental complexity of the audio data is level three, 10 rounds of historical dialogue information can be obtained, that is, the dialogue content of 10 rounds. As another example, when the environmental complexity of the audio data is level one, 3 rounds of historical dialogue information can be obtained, that is, the dialogue content of 3 rounds.
[0057] In summary, the higher the complexity level of the audio data, the more historical dialogue information can be obtained; conversely, the lower the complexity level of the audio data, the less historical dialogue information can be obtained. This can improve the efficiency of obtaining target content while ensuring the accuracy of target content acquisition.
[0058] As an alternative approach, after obtaining the initial text and historical dialogue information, these can be input into a question extraction model to obtain the target content. The question extraction model can be used to remove noise from the initial text to obtain the target content. The amount of historical dialogue information corresponding to the historical dialogue data is determined based on the complexity of the environmental noise. Here, the input to the question extraction model can be the noisy text (initial text) recognized by the speech recognition model and the historical dialogue information. Its output can be the target user's query filtered from the speech recognition text (initial text) of the current round, combined with contextual information. The output of the question extraction model can be in text format.
[0059] As can be seen, the input data for the question extraction model includes historical dialogue information text and initial text. The initial text can be noisy text containing ambient sound extracted from the audio data of the current round using a speech recognition model. It should be noted that when inputting historical dialogue information and initial text into the question extraction model, the initial text can be appended to the end of the historical dialogue information (historical conversation information) to obtain the target input text. Based on this, the target input text is then input into the question extraction model to obtain the target content.
[0060] In the process of acquiring target content, the question extraction model can perform intent recognition on historical dialogue information to obtain intent recognition results. Based on these results, the target content is extracted from the initial text. In other words, the question extraction model can obtain content in the initial text that is highly relevant to historical dialogue information and use that content as the target content.
[0061] Continuing with the example above, in the initial text "Oppenheimer is a good movie, shall we go out to dinner tomorrow? Are there any other assistants for Sherlock Holmes besides Watson?", the relevance between "Oppenheimer is a good movie, shall we go out to dinner tomorrow?" and the historical dialogue information is relatively weak, while the relevance between "Are there any other assistants for Sherlock Holmes besides Watson?" and the historical dialogue information is relatively strong. Therefore, "Are there any other assistants for Sherlock Holmes besides Watson?" can be taken as the target content. Since the user's previous question and answer were about detective stories, which have a strong relevance to "Are there any other assistants for Sherlock Holmes besides Watson?", "Oppenheimer is a good movie, shall we go out to dinner tomorrow?" is environmental noise, not the real user query, while "Are there any other assistants for Sherlock Holmes besides Watson?" is the question the user really wants to ask, so it can be the output of the question extraction model (query extraction model).
[0062] Additionally, the input format for the question extraction model can be a message list, where the last message can be the recognition result of the current input speech (initial text), and the preceding messages can be historical dialogue messages. For example, when the length of the message list is N, the historical dialogue text can be the first N-1 messages, while the noise text of the current speech recognition (initial text) can be the Nth message.
[0063] By combining contextual text information (historical dialogue information), the true question query of the current user can be accurately extracted from the noisy initial text. By filtering out other redundant information in the environment, the accuracy of the input can be guaranteed, thereby effectively improving the response accuracy of the large language model in the voice human-computer interaction system.
[0064] As described above, the problem extraction model, in determining the target content from the initial text based on historical information, can extract content from the initial text that is highly correlated with historical information and use that content as the target content. In other words, the initial text can include multiple subtexts, and the correlation between each subtext and historical information can be obtained separately. Subtexts with a correlation greater than a preset correlation score are then selected as the target content.
[0065] In some implementations, when there are multiple subtexts with a relevance greater than a preset relevance, these subtexts can be visualized to prompt the user to select the subtext that meets their needs. Then, the user's selection command is received, and the subtext corresponding to that command is used as the target content.
[0066] As an example, the initial text received by the question extraction model is: "user: When did you change your work card strap? Hey, Siri! Your Majesty, I am innocent! Can we ski in Beijing now? The weather in Beijing is so beautiful today, everything is white." The historical dialogue information is: "user: I heard it snowed in Beijing, is that true? assistant: Let me check for you, yes, the temperature in Beijing is close to zero degrees now, it did snow."
[0067] By extracting the initial text, we identified two subtexts with a strong relevance to historical information: the first subtext is "So, can you ski in Beijing now?", and the second subtext is "The weather in Beijing is so beautiful today, everything is covered in white." Comparison shows that both subtexts have a high correlation with historical dialogue information. Therefore, these two subtexts can be displayed to the user, prompting them to select the target content. If the user selects the first subtext, then "So, can you ski in Beijing now?" can be used as the target content.
[0068] Optionally, when there are multiple subtexts with a relevance greater than a preset relevance, the audio data corresponding to each subtext can be obtained separately, and the sound features of the corresponding audio data can be analyzed. Based on this, the sound features are matched with the voice of the target user corresponding to the historical dialogue information. If the matching degree between the two is higher than the preset matching degree, the subtext can be used as the target content. Otherwise, the subtext can be removed.
[0069] Continuing with the example above, when detection determines that both the first and second sub-texts have a high degree of correlation with historical dialogue information, the sound features of the audio data corresponding to the first and second sub-texts can be analyzed. This involves comparing the sound features of the two sub-texts with target sound features, which can be obtained through analysis of the audio data corresponding to the historical dialogue information. For example, if analysis determines that the matching degree between the sound features of the first sub-text's audio data and the target sound features is greater than that between the second sub-text and the target sound features, then the first sub-text can be considered the target content.
[0070] Alternatively, the target user's emotional characteristics can be analyzed using historical information, which can include historical dialogue text and / or historical audio data. The target content can then be selected from the first and second sub-texts based on these emotional characteristics. For example, if sentiment analysis determines that the user's emotion in the audio data corresponding to the first sub-text is happy, and the user's emotion in the audio data corresponding to the second sub-text is sad, while the target user's emotional characteristic is happy, then by comparison, the emotional characteristics of the first sub-text match the target emotional characteristics. Therefore, the first sub-text can be selected as the target content.
[0071] It should be noted that when multiple subtexts meet the criteria identified through the aforementioned voice feature analysis and emotion feature analysis, and it is confirmed that all of them were input by the target user, and the content of the multiple subtexts is inconsistent, these multiple subtexts can be merged and integrated to obtain the target content. For example, if the first subtext is "Wow, it's really snowing in Beijing," and the second subtext is "Can we ski in Beijing now?", both of which have a high correlation with historical dialogue information and were both input by the target user, then these two subtexts can be integrated to obtain the target content "Wow, it's really snowing in Beijing, can we ski now?"
[0072] In step S330, response information corresponding to the target content is generated based on the large model.
[0073] As an alternative approach, after obtaining the target content, the corresponding response information can be generated based on the large model. For example, the target content (extracted query) output by the question extraction model (Query extraction model) can be input into the large language model to generate the corresponding response information (text Answer) through the large language model.
[0074] To better illustrate the process of target content extraction and response information generation, the following is given: Figure 4 The example diagram shown is based on Figure 4 As can be seen, the speech recognition model can receive audio input signals (audio data) and convert them from speech to text to obtain the recognized noisy text (initial text). This noisy text is then input into the query extraction model. Additionally, the query extraction model can also receive historical dialogue information. Based on this historical dialogue information and the noisy text, the query extraction model can output its extracted question (target content). Furthermore, the question extracted by the query extraction model is input into a large language model to generate the corresponding text Answer (response information) for the target content.
[0075] By using a pre-trained lightweight question extraction model, combined with context and text extracted from the input audio, the user's question text (Query text) can be accurately extracted. The filtered-out query text is then fed into a large language model for answer generation, and the final answer is output to the user. Since the input data for the large language model is obtained by filtering the initial text through the question extraction model, the accuracy of the large model's response can be improved. In other words, it can extract the user's query from complex, noisy data in a simple and efficient way, filtering out redundant information, thus improving the accuracy of the generated response.
[0076] As an alternative approach, a problem extraction model can be trained before collecting audio data. This involves performing model design and initialization. The model design and initialization process selects a base model, which can then be trained using the training dataset. Here, the base model can be a large language model, such as a Transformer Decoder. Since pre-trained models have strong text generation capabilities on large-scale training corpora, they can be used as the base model, thus reducing training costs to some extent.
[0077] Additionally, a training dataset can be obtained, which may include multiple training samples. Each training sample may include system prompts, historical session messages, first query information, and second query information. The first query information may include the second query information and noise information. Based on this, the base model is trained using the training dataset to obtain the problem extraction model.
[0078] To improve the training efficiency of the question extraction model, a system prompt is introduced into each training sample. This system prompt can be a system prompt word specifically for question extraction, which can be located at the beginning of the training sample. For example, the system prompt "sys" could be "You are a query extraction expert who can accurately extract the user's true query from noisy text based on the context. Please extract the user's query information from the last sentence of the input message and output the text from the user's perspective."
[0079] In the process of acquiring the training dataset, an initial dataset can be obtained. This initial dataset may include multiple subsets of dialogue data, and each subset may include multi-turn dialogue data. For example, the initial dataset W may include 60,000 daily multi-turn dialogue data entries, where one multi-turn dialogue data entry corresponds to one subset of dialogue data. That is, each subset of dialogue data may include multiple turns of dialogue, such as ten turns of dialogue content.
[0080] Based on this, the last message s=(q, a) of each multi-turn dialogue data (a subset of dialogue data) can be obtained, and this last message can be deleted from the multi-turn dialogue data to obtain the historical conversation message r. The last message includes input information, which can be the user's query, i.e., the user's q (input information) can be extracted from s. At the same time, noise information can be randomly generated. This noise information can be text randomly extracted from the initial dataset, which does not include the last message mentioned above. For example, five sentences can be randomly sampled from the remaining data P (P=Ws), and these five sentences can be f1, f2, f3, f4, and f5, respectively.
[0081] Finally, noise information can be added to the input information to obtain the first query information. This involves randomly combining the noise information with the input information q to obtain the first query information, such as (f1, f4, q, f2, f3, f5). As you can see, the first query information v is a sample with added noise, where sentences are separated by commas. Additionally, the input information q can be used as the second query information, thus obtaining the training samples corresponding to each subset of dialogue data.
[0082] For example, the training sample format can be (sys, r, v, q), where sys is the system prompt message, r is the historical conversation message (equivalent to historical dialogue information), v is the query with noise added (equivalent to the initial text), and q is the original clean query (equivalent to the target content). In the training sample (sys, r, v, q), (sys, r, v) can be used as the input to the base model, while q can be used as the target for training the base model.
[0083] The above steps can be repeated for each piece of data in W to obtain the training dataset W'. That is, for each subset of the dialogue data, the steps of obtaining the last message of each of the multi-turn dialogue data, adding the noise information to the input information to obtain the first query information, and using the input information as the second query information can be executed respectively.
[0084] As an example, the first subset of dialogue data is: "user: Can you tell me a story? assistant: Sure, what kind of story do you want to hear? Science fiction, fairy tale, or detective story? user: Then tell me a detective story. assistant: Sure, have you heard of Sherlock Holmes? It's a very good detective novel. user: Yes, I have. Besides Watson, are there any other assistants for Holmes? assistant: There are many, like his brother."
[0085] The last message in the first subset of dialogue data is "user: I've heard of it. Besides Watson, are there any other assistants for Sherlock Holmes? assistant: There are many, like his brother." Therefore, the input information (second query information) could be "user: I've heard of it. Besides Watson, are there any other assistants for Sherlock Holmes?". Deleting the last message yields the historical conversation message r.
[0086] Additionally, five randomly selected sentences are: "f1: What is this movie about? f2: Oranges are the best. f3: That movie is pretty good, let's go see it together. f4: What is the highest mountain in the world? f5: What is the name of this music?". Adding this noise information to the above input information, the resulting first query information could be: "What is this movie about? I've heard of it. Besides Watson, are there any other assistants for Sherlock Holmes? Oranges are the best. What is the highest mountain in the world? What is the name of this music? That movie is pretty good, let's go see it together." The second query information is: "I've heard of it. Besides Watson, are there any other assistants for Sherlock Holmes?". The first query information, the second query information, historical conversation messages, and system prompts can be used to form the training sample for the above first dialogue data subset.
[0087] By modifying the original multi-turn dialogue data (initial dataset), adding random noise to the tail query, designing reasonable system prompts, using the original context as historical information, and using the original query as the training target, a large number of training samples that can improve the model's query extraction capabilities can be automatically and efficiently generated. Because the above training sample acquisition method is highly versatile, it has significant advantages in most scenarios where the key information extraction capabilities of large models need to be optimized.
[0088] After obtaining the training dataset through the above method, the training dataset can be input into the base model to train the base model and obtain the problem extraction model.
[0089] Optionally, to filter out noise data that is similar to the real query and accelerate training convergence, the training dataset can be filtered using a data cleaning model to obtain a target dataset. The data cleaning model can be used to remove target samples (redundant samples) from the training dataset where the similarity between noise information and the second query information exceeds a preset threshold. The base model is then trained using the filtered target dataset.
[0090] As an example, a target sample can be selected from multiple training samples based on a first cleaning model, where the noise information of the first query in the target sample is similar to that of the second query. Based on this, the target sample is removed from the training dataset to obtain the target dataset. The first cleaning model can be a large model that determines whether the noise (noise information) of the target sample is similar to the question text (input information / second query text). If they are determined to be similar, the training sample (target sample) can be deleted from the training dataset. Otherwise, the target sample can be retained.
[0091] As another example, a second cleaning model can be used to obtain the first feature vector corresponding to the noise information in the target sample and the second feature vector of the second query information. Then, the cosine similarity between the first and second feature vectors is obtained, and it is determined whether this cosine similarity is greater than a preset similarity. If the cosine similarity between the first and second feature vectors is determined to be greater than the preset similarity, the target sample is removed from the training dataset, resulting in the target dataset. The second cleaning model can be a representation model, such as an embedding model, whose working principle can be as follows: Figure 5 As shown, based on Figure 5 As can be seen, the input data of the embedding model is text, and the output data is a feature vector. Conversely, the target sample can be preserved.
[0092] As another example, a first cleaning model can be used to determine whether the noise (noise information) of the target sample is similar to the question text (second query text). If they are determined to be similar, feature vectors can be extracted based on the second cleaning model to calculate the cosine similarity between the noise feature vector (first feature vector) and the question text (query text) feature vector (second feature vector). If the cosine similarity is greater than or equal to a preset similarity, the target sample is removed from the training dataset, resulting in the target dataset. Conversely, if the cosine similarity is less than the preset similarity, the target sample can be retained.
[0093] As a specific implementation method, such as Figure 6 As shown, a first cleaning model (large model) and a second cleaning model (embedding model) can be used to clean the training data. The first cleaning model can determine whether the noise and the problem text are similar. If they are similar, the second cleaning model can be used to extract feature vectors for similar samples and calculate the cosine similarity between the noise feature vector and the problem text feature vector. For samples with a similarity greater than the preset value, they can be deleted from the training dataset.
[0094] Continuing with the example above, during the cleaning of the training dataset, a target sample can be extracted from the training dataset W'. This target sample has the data (sys, r, v, q). Its first query information (noise text) is v, and its second query information (question text) is q, where the first query information v includes the second query information q. Then, the noise text g can be extracted from the first query information. This noise text can then be used as the aforementioned noise information; that is, by deleting the question text q from v, we obtain the pure noise text g.
[0095] Based on this, the first cleaning model can be used to determine the similarity between the noisy text g (noise information) and the question text q (second query information). If the first cleaning model determines that they are not similar, the target sample is retained, and the next training sample is then checked. If the first cleaning model determines that the noise information is similar to the second query information, the second cleaning model can be used to extract features from the noisy text g and the question text q respectively, and the cosine similarity of the feature vectors is calculated. If the cosine similarity is greater than 0.75, the data is deleted. Conversely, if the similarity is less than 0.75, the target sample is retained. The above steps are repeated until every data point (training sample) in W' has been checked, resulting in a cleaned dataset W'', which can then be used as the target dataset.
[0096] For example, the training dataset includes 60,000 training samples. By cleaning it using the method described above, 123 samples with high similarity are found. After deleting these highly similar samples, the final target dataset includes 59,877 training samples.
[0097] It should be noted that during the cleaning process of the training samples in the training dataset, the input data of the first cleaning model can include not only noisy text (noise information) and question text (second query information), but also filtering prompts, such as "Do these two sentences contain the same user query? If yes, please output Yes; if not, please output No".
[0098] By using the data cleaning model described above, the technical solution can obtain a high-quality query extraction training dataset. First, a large model is used to determine whether the noisy text and the question text are similar. Then, a representation model is used to extract feature vectors from the samples that the large model determines are similar. After calculating the cosine similarity, samples with extremely high similarity between noise and the target are filtered out by setting a reasonable threshold. This not only improves the quality of the training dataset but also indirectly improves the stability of model training.
[0099] As an alternative approach, after obtaining the target dataset, the base model can be fine-tuned using the aforementioned high-quality target dataset to improve the model's query extraction accuracy. For example, the efficient and concise LORA (Low-Rank Adaptation) method can be used to fine-tune the base model to obtain the question extraction model. For instance, the learning rate of the base model can be set to 0.0005, and the model can be trained for 5 epochs to obtain the question extraction model.
[0100] Please refer to the above LORA fine-tuning process. Figure 7 For each weight matrix in all Transformer Decoder modules of the base model, a corresponding Adapter can be added. For example, the matrix can first be reduced in rank; if the matrix size is M×N, it can be reduced to M×32, i.e., right-multiplied by an N×32 matrix, which can be denoted as P. Then, the matrix can be increased in rank, i.e., M×32 is increased to M×N. Here, the core logic is right-multiplied by a 32×N matrix, which can be denoted as Q.
[0101] During the training of the base model, the pretrained weights of the LLM (Large Language Model) can be frozen, and P and Q can be updated. The pretrained weights can be the original weights. In addition, when fine-tuning the base model, W = W + PQ, that is, the product of P and Q can be added to the original W to achieve inference.
[0102] It should be noted that during the training of the base model based on the training dataset, each data point (sys, r, v, q) can be split into input and target. As mentioned above, the inputs can be (sys, r, v), and the targets can be q. The technical solution can use the targets to supervise the training process and use backpropagation to update the parameters of the low-rank adaptive adapters (Lora Adapters). After training, the parameters of the Lora Adapters can be merged with the original LLM to obtain the final problem extraction model. The parameter update process described above can be as follows: Figure 8 As shown, Figure 8 In this context, CrossEntropy can be the loss calculated by comparing the outputs with the targets.
[0103] Furthermore, after obtaining the question extraction model through training, it can be deployed and integrated into the system. For example, an open-source inference framework can be used to deploy the question extraction model as an HTTP (Hypertext Transfer Protocol Service) service, and then integrate this service into the voice-based human-computer interaction system. Subsequently, for each user input speech, after the speech recognition model extracts the text, it is no longer directly sent to the large language model for result generation. Instead, the question extraction model first filters redundant information to extract the true user query (target content) from the noisy text, and then sends it to the LLM for response generation to obtain the reply information.
[0104] It should be noted that after generating the reply information, it can be displayed directly to the user, or it can be converted into audio data first and then output to the user through an audio output device. There are no specific restrictions on how the reply information is output, and the choice can be made according to the actual situation.
[0105] By employing a trained question extraction model, user queries can be accurately extracted from noisy text. This avoids the illusions caused by invalid input information in large language models, thus solving the problem of low accuracy in large model responses due to ambient sound reception issues in human-computer dialogue scenarios.
[0106] When audio data is collected, noise removal is performed on the initial text corresponding to the audio data to ensure the accuracy of the generated response information. That is, after text recognition of the audio data to obtain the initial text, the historical dialogue information corresponding to the audio data is obtained. Based on this, the historical dialogue information and the initial text are input into the question extraction model. The question extraction model uses the historical dialogue information to filter the noise information in the initial text to obtain the clean target content. Since the target content does not contain noise interference information, more accurate response information can be generated based on the target content.
[0107] Based on the same inventive concept, a smart question-answering device based on a large model is also provided. Figure 9 This is a block diagram illustrating a large-model-based intelligent question-answering device 900 according to an exemplary embodiment, such as... Figure 9 As shown, the intelligent question-answering device 900 based on a large model may include a recognition module 910, a processing module 920, and a generation module 930.
[0108] The recognition module 910 is configured to collect audio data and perform text recognition on the audio data to obtain initial text; The processing module 920 is configured to acquire historical dialogue information corresponding to the audio data, and input the historical dialogue information and the initial text into the question extraction model to obtain the target content. The question extraction model is used to remove noise content in the initial text. The amount of historical dialogue information corresponding to the historical dialogue data is determined based on the environmental noise complexity. The generation module 930 is configured to generate response information corresponding to the target content based on the large model.
[0109] In some implementations, the processing module 920 is further configured to perform intent recognition on the historical dialogue information based on the question extraction model to obtain intent recognition results; and to extract the target content from the initial text based on the intent recognition results.
[0110] In some implementations, the large-model-based intelligent question-answering device 900 may further include: The data acquisition module is configured to acquire a training dataset, which includes multiple training samples. Each training sample includes system prompt information, historical session messages, a first query information, and a second query information. The first query information includes the second query information and noise information. The training module is configured to train the base model based on the training dataset to obtain the problem extraction model.
[0111] In some implementations, the training module includes: The cleaning submodule is configured to filter the training dataset based on a data cleaning model to obtain a target dataset. The data cleaning model is used to filter out target samples in the training dataset. The similarity between the noise information in the target sample and the second query information exceeds a preset threshold. The training submodule is configured to train the base model using the target dataset.
[0112] In some implementations, the cleaning submodule is further configured to filter the target sample from the plurality of training samples based on a first cleaning model, wherein the noise information of the first query information in the target sample is similar to the second query information; and to remove the target sample from the training dataset to obtain the target dataset.
[0113] In some implementations, the cleaning submodule is further configured to obtain a first feature vector corresponding to the noise information in the target sample and a second feature vector of the second query information based on a second cleaning model; when the cosine similarity between the first feature vector and the second feature vector is greater than a preset similarity, the target sample is filtered out from the training dataset to obtain the target dataset.
[0114] In some implementations, the data acquisition module may also be configured to acquire an initial dataset, the initial dataset comprising multiple subsets of dialogue data, each subset of dialogue data comprising multiple rounds of dialogue data; acquire the last round message of each set of multiple rounds of dialogue data, and delete the last round message from the set of multiple rounds of dialogue data to obtain the historical session message, the last round message comprising input information; The noise information is generated randomly; The noise information is added to the input information to obtain the first query information, and the input information is used as the second query information to obtain the training sample.
[0115] When audio data is collected, noise removal is performed on the initial text corresponding to the audio data to ensure the accuracy of the generated response information. That is, after text recognition of the audio data to obtain the initial text, the historical dialogue information corresponding to the audio data is obtained. Based on this, the historical dialogue information and the initial text are input into the question extraction model. The question extraction model uses the historical dialogue information to filter the noise information in the initial text to obtain the clean target content. Since the target content does not contain noise interference information, more accurate response information can be generated based on the target content.
[0116] The following is for reference. Figure 10 The diagram illustrates a structural schematic of an electronic device 1000 suitable for implementing the above-described technical solution. Terminal devices may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs (Televisions), desktop computers, etc. Figure 10 The electronic device shown is merely an example and should not be construed as limiting its functionality or scope of use.
[0117] like Figure 10As shown, the electronic device 1000 may include a processing unit 1001, such as a central processing unit or a graphics processing unit, which can perform various appropriate actions and processes according to a program stored in the read-only memory 1002 (ROM) or a program loaded from the storage device 1008 into the random access memory 1003 (RAM). The random access memory 1003 also stores various programs and data required for the operation of the electronic device 1000. The processing unit 1001, the read-only memory 1002, and the random access memory 1003 are interconnected via a bus 1004. An input / output interface 1005 (I / O interface) is also connected to the bus 1004.
[0118] Typically, the following devices can be connected to the input / output interface 1005: input devices 1006 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1008 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic device 1000 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 An electronic device 1000 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0119] In particular, depending on certain circumstances, the processes described in the flowchart above can be implemented as computer software programs. For example, a computer program product is provided, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. This computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1008, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the above-described methods.
[0120] It should be noted that the aforementioned computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In one case, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In another case, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.
[0121] In some implementations, clients can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), the internet (e.g., the Internet), and peer-to-peer networks (e.g., ad-hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0122] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0123] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes one or more of these programs, the electronic device causes the following actions: to collect audio data and perform text recognition on the audio data to obtain initial text; to obtain historical dialogue information corresponding to the audio data and input the historical dialogue information and the initial text into a question extraction model to obtain target content, wherein the question extraction model is used to remove noise content from the initial text, and the amount of historical dialogue information corresponding to the historical dialogue data is determined based on the environmental noise complexity; and to generate response information corresponding to the target content based on the large model.
[0124] Computer program code for performing the above operations can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages, as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0125] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0126] The modules mentioned above can be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself; for example, a processing module can also be described as "a module that acquires historical dialogue information corresponding to the audio data, and inputs the historical dialogue information and the initial text into a question extraction model to obtain the target content."
[0127] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Parts (ASSPs), Systems on Chips (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0128] In this context, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0129] The above description is merely illustrative and explains the technical principles employed. Those skilled in the art should understand that the scope of the technical solution is not limited to specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features provided herein that have similar functions.
[0130] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limitations on the scope of the technical solution. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.
Claims
1. An intelligent question-answering method based on a large model, the method comprising: Acquire audio data and perform text recognition on the audio data to obtain initial text; The historical dialogue information corresponding to the audio data is obtained, and the historical dialogue information and the initial text are input into the question extraction model to obtain the target content. The question extraction model is used to remove noise content in the initial text. The amount of historical dialogue information corresponding to the historical dialogue data is determined based on the environmental noise complexity. The response information corresponding to the target content is generated based on the large model.
2. The intelligent question-answering method based on a large model according to claim 1, wherein inputting the historical dialogue information and the initial text into the question extraction model to obtain the target content includes: Based on the question extraction model, intent recognition is performed on the historical dialogue information to obtain intent recognition results; The target content is extracted from the initial text based on the intent recognition result.
3. The intelligent question answering method based on a large model according to claim 1, the method further includes: Obtain a training dataset, which includes multiple training samples. Each training sample includes system prompt information, historical session messages, first query information, and second query information. The first query information includes the second query information and noise information. The base model is trained based on the training dataset to obtain the problem extraction model.
4. The intelligent question answering method based on a large model according to claim 3, wherein training the basic model based on the training dataset includes: The training dataset is filtered based on the data cleaning model to obtain the target dataset. The data cleaning model is used to filter out target samples in the training dataset. The similarity between the noise information in the target sample and the second query information exceeds a preset threshold. The base model is trained using the target dataset.
5. The intelligent question answering method based on a large model according to claim 4, wherein the step of filtering the training dataset based on the data cleaning model to obtain the target dataset includes: The target sample is selected from the plurality of training samples based on the first cleaning model, wherein the noise information of the first query information in the target sample is similar to that of the second query information; The target samples are filtered out from the training dataset to obtain the target dataset.
6. The intelligent question answering method based on a large model according to claim 4, wherein the step of filtering the training dataset based on the data cleaning model to obtain the target dataset includes: Based on the second cleaning model, the first feature vector corresponding to the noise information in the target sample and the second feature vector of the second query information are obtained. When the cosine similarity between the first feature vector and the second feature vector is greater than a preset similarity, the target sample is removed from the training dataset to obtain the target dataset.
7. The intelligent question answering method based on a large model according to claim 3, wherein obtaining the training dataset includes: Obtain an initial dataset, which includes multiple subsets of dialogue data, each subset of dialogue data including multiple rounds of dialogue data; Obtain the last message of each of the multi-turn dialogue data, and delete the last message from the multi-turn dialogue data to obtain the historical session message, wherein the last message includes input information; The noise information is generated randomly; The noise information is added to the input information to obtain the first query information, and the input information is used as the second query information to obtain the training sample.
8. A large-model-based intelligent question-answering device, the device comprising: The recognition module is configured to collect audio data and perform text recognition on the audio data to obtain initial text; The processing module is configured to acquire historical dialogue information corresponding to the audio data, and input the historical dialogue information and the initial text into the question extraction model to obtain the target content. The question extraction model is used to remove noise content in the initial text. The amount of historical dialogue information corresponding to the historical dialogue data is determined based on the environmental noise complexity. The generation module is configured to generate response information corresponding to the target content based on the large model.
9. A computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method according to any one of claims 1-7.
10. An electronic device, comprising: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement the steps of the method according to any one of claims 1-7.