Cross-modal question and answer processing method and device based on large model and storage medium
Through the cross-modal question-and-answer processing method based on large models, the single modal problem of traditional interaction methods is solved, the accuracy and efficiency of cross-modal question-and-answer processing is improved, and more efficient information acquisition and interaction experience is provided.
Patent Information
- Application Number
- CN202510336695.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-08-01
AI Technical Summary
Traditional interaction methods are limited to a single mode, which is difficult to meet users' diverse information acquisition and interaction needs in complex scenarios.
A cross-modal question-and-answer processing method based on a large model is adopted. By performing activity detection of user input speech, text before the speech pause moment is obtained, text response processing is performed in combination with a pre-trained speech question-and-answer system, and cross-modal information processing is performed using a speech encoder and a cross-attention large language model.
It improves the accuracy and efficiency of question-and-answer processing, and achieves a faster, more accurate, richer, more proactive and smoother interactive experience.
Smart Images

Figure CN120407726A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, specifically to artificial intelligence technology fields such as speech interaction processing, large models, machine learning, and natural language processing. In particular, it relates to a cross-modal question and answer processing method, device, and storage medium based on a large model. Background Art
[0002] With the continuous progress of technology, in the field of human-computer interaction, people's requirements for the interaction experience are also increasing day by day.
[0003] Traditional interaction methods are often limited to a single modality, such as text input or simple visual feedback, which can no longer meet users' diverse needs for information acquisition and interaction in complex scenarios. Therefore, in order to better simulate human interaction and handle complex problems in the real world, cross-modal processing has become crucial. Summary of the Invention
[0004] The present disclosure provides a cross-modal question and answer processing method, device, and storage medium based on a large model.
[0005] According to one aspect of the present disclosure, there is provided a cross-modal question and answer processing method based on a large model, including:
[0006] Performing activity detection on the target speech input by the user;
[0007] When it is detected that the input of the target speech pauses, obtaining the first text corresponding to the first input speech before the speech pause moment in the target speech;
[0008] Based on the first text and the first input speech, using a pre-trained speech question and answer processing system to perform text response processing.
[0009] According to another aspect of the present disclosure, there is provided a training method for a speech question and answer system, including:
[0010] Obtaining training data, where the training data includes training text questions, annotated text answers, and annotated feature expressions of the training text questions;
[0011] Using a pre-trained mixing system to mix the training text questions to obtain training speech questions;
[0012] Based on the training speech questions, the training text questions, the annotated feature expressions of the training text questions, and the annotated text answers, adjusting the parameters in the speech question and answer system to obtain the trained speech question and answer processing system, and the speech question and answer system is used in the method of the above aspect and any possible implementation manner.
[0013] According to another aspect of the present disclosure, there is provided a cross-modal question-answering processing device based on a large model, including:
[0014] An activity detection module for detecting the activity of the target voice input by the user;
[0015] A text acquisition module for acquiring the first text corresponding to the first input voice before the voice pause moment in the target voice in response to detecting the pause of the input of the target voice;
[0016] A question-answering processing module for performing text response processing based on the first text and the first input voice by using a pre-trained voice question-answering processing system.
[0017] According to yet another aspect of the present disclosure, there is provided a training device for a voice question-answering system, including:
[0018] An acquisition module for acquiring training data, where the training data includes training text questions, annotated text answers, and annotated feature expressions of the training text questions;
[0019] A mixing processing module for mixing the training text questions by using a pre-trained mixing system to obtain training voice questions;
[0020] A training module for adjusting the parameters in the voice question-answering system based on the training voice questions, the training text questions, the annotated feature expressions of the training text questions, and the annotated text answers to obtain the trained voice question-answering processing system, and the voice question-answering system is used in the device in the above aspect and any possible implementation manner.
[0021] According to still another aspect of the present disclosure, there is provided an electronic device, including:
[0022] At least one processor; and
[0023] A memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods in the above aspect and any possible implementation manner.
[0025] According to still yet another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the methods in the above aspect and any possible implementation manner.
[0026] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the methods of the above-described aspects and any possible implementation manners.
[0027] According to the technology of the present disclosure, the accuracy and efficiency of question-and-answer processing can be effectively improved.
[0028] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0030] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;
[0031] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;
[0032] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure;
[0033] Figure 4 is a schematic architecture diagram of a voice encoder provided by an embodiment of the present disclosure;
[0034] Figure 5 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0035] Figure 6 is a schematic diagram according to the fifth embodiment of the present disclosure;
[0036] Figure 7 is a training principle diagram of a voice question-and-answer system according to an embodiment of the present disclosure;
[0037] Figure 8 is a schematic diagram according to the sixth embodiment of the present disclosure;
[0038] Figure 9 is a schematic diagram according to the seventh embodiment of the present disclosure;
[0039] Figure 10 is a schematic diagram according to the eighth embodiment of the present disclosure;
[0040] Figure 11 is a schematic diagram according to the ninth embodiment of the present disclosure;
[0041] Figure 12 is a block diagram of an electronic device for implementing the method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0043] Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0044] It should be noted that the terminal devices involved in the embodiments of the present disclosure may include, but are not limited to, intelligent devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers (Tablet Computers); display devices may include, but are not limited to, devices with display functions such as personal computers and televisions.
[0045] In addition, the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0046] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure; as Figure 1 shown, this embodiment provides a cross-modal question-and-answer processing method based on a large model, which may specifically include the following steps:
[0047] S101. Perform activity detection on the target voice input by the user;
[0048] S102. When it is detected that the input of the target voice pauses, obtain the first text corresponding to the first input voice before the voice pause moment in the target voice;
[0049] S103. Based on the first text and the first input voice, use a pre-trained voice question-and-answer processing system to perform text response processing.
[0050] The execution subject of the question-and-answer processing method based on the cross-modal large model in this embodiment is a question-and-answer processing device based on the cross-modal large model. This device can be an electronic entity or an application integrated with software.
[0051] Specifically, perform activity detection on the target voice input by the user, that is, detect whether the voice input by the user exists. Since the activity detection is a process of continuously detecting the voice input by the user, through this activity detection, the pause input of the target voice input by the user can be detected.
[0052] In practical applications, when the user inputs the target voice, there will be normal breathing pauses. The duration of the breathing pause is very short, and it can be considered that the user is still inputting the voice and there is no voice pause. However, if the pause duration is long, greater than the normal breathing pause duration, such as reaching or being greater than the first preset duration, at this time, it can be considered that a pause input by the user has occurred. Based on this principle, performing activity detection on the target voice input by the user can detect whether the target voice has a pause input.
[0053] In this embodiment, when it is detected that the target voice has a pause input, obtain the first input voice before the pause moment, and perform speech recognition on the first input voice, and the first text can be obtained. Then, a pre-trained semantic question-answering processing system can be used to perform question-answering processing according to the first input voice and the first text.
[0054] The voice question-answering processing system of this embodiment can adopt a large model and can process cross-modal information such as speech to text. That is to say, in specific implementation, the input question is speech, and the output answer is in text response, that is, cross-modal question-answering processing can be finally achieved.
[0055] The cross-modal question-answering processing method based on a large model in this embodiment can, when there is a pause input during the voice pause, based on the user's first input voice and the first text corresponding to the first input voice, use a voice question-answering processing system to perform question-answering processing. Since the question input by the user is in the voice modality, and the response is in text response processing, therefore, the technical solution in this embodiment can achieve cross-modal question-answering processing. Moreover, when performing text response processing, both the first text and the first input voice are referred to, which can effectively improve the accuracy and efficiency of question-answering processing.
[0056] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure; the cross-modal question-answering processing method based on a large model in this embodiment further describes the technical solution of the present disclosure in more detail on the basis of the technical solution of the above Figure 1 illustrated embodiment. As Figure 2 shown, the cross-modal question-answering processing method based on a large model in this embodiment may specifically include the following steps:
[0057] S201. Perform activity detection on the target voice input by the user;
[0058] S202. When detecting that the target voice pauses input, obtain the first text corresponding to the first input voice before the pause moment in the target voice;
[0059] S203. Use a pre-trained integrity detection model to obtain the speech integrity of the first text;
[0060] Specifically, use this integrity detection model to perform semantic integrity detection on the first text to obtain the semantic integrity of the first text.
[0061] This integrity detection model is implemented based on a neural network model. When in use, input the first text into this integrity detection model, and this integrity detection model can predict the semantic integrity of the first text based on the first text. For example, this semantic integrity can be a value between 0 and 1. The larger the value, the more complete the semantics; otherwise, the smaller the value, the less complete the semantics.
[0062] S204. Based on the semantic integrity, the first text, and the first input voice, use a voice question-and-answer processing system to perform text response processing.
[0063] For example, when specifically implementing this step S204, the following steps may be included:
[0064] (1) Based on the first input voice and the first text, use the speech encoder (Speech Encoder) in the voice question-and-answer processing system to obtain the feature expression of the first input voice, where the length of the feature expression of the first input voice is equal to the length of the first text;
[0065] In this embodiment, the feature expression finally obtained by the speech encoder encoding the first input voice with reference to the first text is in the form of a matrix, and the length of the matrix can be considered as the length of the feature expression of the first input voice. In this embodiment, when the speech encoder encodes the first input voice, it needs to refer to the first text corresponding to the first input voice. Therefore, the length of the feature expression of the first input voice obtained is equal to the length of the first text, that is, the obtained feature expression of the first input voice is a unified acoustic representation of the same length as the first text.
[0066] The commonly used speech encoders in the industry generally adopt an Encoder structure to process the input audio and extract the hidden layer representation at the granularity of audio frames. Since it is a representation at the granularity of audio frames, its granularity is much smaller than that of text tokens, so its sequence length is much longer than that of text sequences, which brings many difficulties to modality alignment. The speech encoder in this embodiment can adopt an Encoder-Decoder structure, perform historical abstraction on the output of the Encoder based on the output of the Decoder, and finally obtain an equi-length acoustic unified representation at the granularity of text tokens; it is easier to perform modality alignment based on this equi-length unified acoustic representation, and the model inference speed is also faster.
[0067] (2) Based on the semantic integrity and the feature expression of the first input speech, use the Cross Attention Large Language Model (LLM) in the speech question-answering processing system to perform text response processing.
[0068] Specifically, when implementing this step (2), it may include the following steps:
[0069] (a1) Detect whether the semantic integrity of the first input speech is less than the preset integrity; if so, execute step (b1); if not, execute step (d1);
[0070] In this embodiment, if the semantic integrity is greater than or equal to the preset integrity, it is considered that the semantic of the first input speech is complete; otherwise, if the semantic integrity is less than the preset integrity, it is considered that the semantic of the first input speech is incomplete.
[0071] (b1) Use the cross-attention large language model to generate question information based on the feature expression of the first input speech; execute step (c1);
[0072] Optionally, the obtained semantic integrity can also be input into the cross-attention large language model. Based on this semantic integrity and the preset integrity, it can be known that the hesitation posting mechanism needs to be triggered.
[0073] Or when it is determined in step (a1) that the semantics is incomplete, directly generate question information as a prompt and input it together with the feature expression of the first input speech into the cross-attention large language model. The cross-attention large language model generates question information according to the input information.
[0074] Of course, optionally, only the feature representation of the first input speech can be input into the cross-attention large language model. This cross-attention large language model itself is a very intelligent large language model. Based on the feature representation of the first input speech, it can recognize that the first input speech is incomplete and further generate question information to prompt the hesitating user to continue completing the question. Therefore, this question information can also be called hesitant question information, which is used when the user is hesitating in inputting to ask a question and prompt the user to better complete the question.
[0075] (c1), perform a text response based on the question information.
[0076] For example, during the process of the user inputting speech, when the speech pauses, the first text corresponding to the user's first input speech can be "I want to listen". In the manner of the above embodiment, when it is detected that the semantics of the first text is incomplete, the question information generated by the cross-attention large language model can be "Do you want to listen to music" and respond to the user. The user can continue to input speech based on the question information to make the input speech more complete.
[0077] (d1), use the cross-attention large language model to generate a first text answer based on the feature representation of the first input speech; perform step (e1);
[0078] (e1), store the feature representation of the first input speech and the first text answer in the cache.
[0079] At this time, since the premise of detection is a pause in input, not the end of speech, directly responding based on the generated first text answer will not only result in bad phenomena such as interrupting the user, but also since the user's speech input has not stopped, the generated first text answer may not be the result the user wants. Therefore, at this time, the feature representation of the first input speech and the first text answer are stored in the cache first and not responded to temporarily. Subsequently, if the user's speech input ends and the user does not add other speech inputs, at this time, a quick response can be directly made based on the first text answer to improve the efficiency of question and answer processing.
[0080] The cross-modal question and answer processing method based on a large model in this embodiment performs text response processing using a speech question and answer processing system based on semantic integrity, the first text, and the first input speech, and can effectively improve the accuracy of question and answer processing.
[0081] Moreover, in this embodiment, a voice encoder in the voice question-and-answer processing system can be used to perform voice encoding based on the first input voice and the first text to obtain a feature representation of the first input voice, such that the length of the feature representation of the first input voice is the same as the length of the first text. In this way, when the Cross Attention LLM in the subsequent voice question-and-answer processing system performs text response processing, the length of the feature representation of the first input voice for reference is the same as the length of the first text, which can reduce the difficulty of modality alignment of different modality information in the text response process and further effectively accelerate the model inference speed.
[0082] Moreover, in this embodiment, when the semantics is incomplete, question information can also be generated to further assist the user in completing the input voice. When the semantics is complete, a text answer can be generated in a timely manner and stored in the cache, so that when the subsequent voice ends, a direct response can be made, effectively shortening the response time, and thus effectively improving the efficiency of question-and-answer processing.
[0083] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure; the cross-modal question-and-answer processing method based on a large model in this embodiment further describes the technical solution of the present disclosure in more detail on the basis of the technical solution of the above Figure 2 illustrated embodiment. As Figure 3 shown, the cross-modal question-and-answer processing method based on a large model in this embodiment may specifically include the following steps:
[0084] S301. Perform liveness detection on the target voice input by the user;
[0085] S302. When it is detected that the input of the target voice pauses, obtain the first input voice before the pause moment in the target voice; and perform speech recognition on the first input voice to obtain the first text;
[0086] S303. Use the voice encoder in the voice question-and-answer processing system to perform voice encoding on the first input voice based on the first text to obtain a feature representation of the first input voice;
[0087] In this embodiment, the length of the feature representation of the first input voice is equal to the length of the first text.
[0088] For example, Figure 4 is a schematic diagram of the architecture of the voice encoder provided by the embodiment of the present disclosure; as Figure 4As shown, the speech encoder in this embodiment adopts an Encoder-Decoder structure. For example, the encoder can adopt a Streaming Multi-Layer Truncated Attention (SMLTA) encoder, and the corresponding decoder can adopt an SMLTA decoder. The equal-length unified representation is based on the text of the input speech, encoding the features of the input speech into features of the same length as the text of the input speech. Then, based on the features of the equal-length unified representation and the features decoded by the decoder, feature fusion processing is performed to obtain the final feature expression of the same length as the text corresponding to the input speech. When in use, directly input the speech and the corresponding text to this speech encoder, and this speech encoder can adopt Figure 4 the architecture shown, and output a feature expression of the input speech that is of the same length as the text.
[0089] S304. Use a pre-trained integrity detection model to detect the semantic integrity of the first text to obtain the semantic integrity;
[0090] S305. Detect whether the semantic integrity is greater than or equal to the preset integrity; if it is greater than or equal to, determine that the semantics is complete and execute step S306; otherwise, determine that the semantics is incomplete and execute step S308;
[0091] S306. Use the cross-attention large language model in the speech question-and-answer processing system to generate the first text answer based on the feature expression of the first input speech; execute step S307;
[0092] In this embodiment, the input of the cross-attention large language model can be in the form of key-value pairs (key, Value), where the values of both key and Value are the feature expressions of the first input speech.
[0093] S307. Store the feature expression of the first input speech and the first text answer in the cache; execute step S310;
[0094] The feature expression of the first input speech and the first text answer can also be stored in the form of key-Value. The pre-storage function that steps S306 - S307 can achieve is to pre-store the first text answer and the corresponding feature expression of the first input speech in the cache, facilitating quick response later.
[0095] S308. Use the cross-attention large language model in the speech question-and-answer processing system to generate question information based on the feature expression of the first input speech; execute step S309;
[0096] S309. Make a text response based on the question information; execute step S310;
[0097] The generation of question information and corresponding Q&A responses in this embodiment belongs to an active questioning mechanism, which can effectively assist users to complete their input voice more actively and smoothly.
[0098] S310. Continue to perform activity detection on the target voice input by the user; execute step S311;
[0099] S311. In response to detecting the end of the target voice input, obtain the feature expression of the user's second input voice before the end of the voice input;
[0100] In this embodiment, during the process of performing activity detection on the target voice input by the user, detecting that the voice input pauses specifically may refer to detecting that the duration of the user's input voice stop reaches a first preset duration, and at this time, it is determined that the user's voice input pauses. Detecting the end of the voice input specifically may refer to detecting that the duration of the user's input voice stop reaches a second preset duration; the second preset duration is greater than the first preset duration.
[0101] That is to say, during the voice input process of the user, first, it is detected that the user's input voice pauses. If the user does not further input voice, then it can be detected that the user's input voice ends. The lengths of the first preset duration and the second preset duration can be set according to the actual scenario and are not limited herein.
[0102] For example, when this step is specifically implemented, it may include the following steps:
[0103] (a2) In response to detecting the end of the voice input, obtain the user's second input voice before the end of the voice input;
[0104] Combined with the above scenario, it can be known that the second input voice in this embodiment may only include the first input voice; or it may also include the first input voice and other voices subsequently supplemented and input by the user.
[0105] (b2) Perform speech recognition on the second input voice to obtain a second text;
[0106] (c2) Based on the second input voice and the second text, use the speech encoder in the speech Q&A processing system to obtain the feature expression of the second input voice, so that the length of the feature expression of the second input voice is equal to the length of the second text.
[0107] The implementation manners of steps (a2)-(c2) are the same as those of steps S302-S303 above.
[0108] S312. Detect whether the similarity between the feature expression of the second input voice and the feature expression of the first input voice in the cache is greater than or equal to a preset similarity; if so, execute step S313; otherwise, execute step S315;
[0109] S313. Obtain the first text answer corresponding to the feature expression of the first input speech from the cache; execute step S314;
[0110] S314. Perform a text response based on the first text answer; end.
[0111] Step S313 is to obtain from the cache the text answer pre-stored in steps S306 - S307. This function can be a prefetch function. The combination of the two realizes the pre-storage and prefetch function of the embodiments of the present disclosure.
[0112] S315. Use the cross-attention large language model in the voice question-and-answer processing system to generate a second text answer based on the feature expression of the second input speech; execute step S316;
[0113] S316. Perform a text response based on the second text answer; end.
[0114] Specifically, in the actual application scenario, since the length of the second preset duration is greater than the length of the first preset duration, it is inevitable to first detect the user's voice pause input and then detect the end of the user's voice. That is to say, before the end of the user's voice, there must be a voice pause input first. Therefore, in the normal scenario, the above pre-storage and prefetch function of this embodiment is adopted, which can achieve a fast response.
[0115] For the integrity of the solution, to avoid the situation where the correct feature expression of the first input speech and the corresponding first text answer are not stored in the cache, when the similarity between the feature expression of the second input speech and the feature expression of the first input speech in the cache does not reach the preset similarity threshold, the cross-attention large language model can also be used to generate a second text answer based on the feature expression of the second input speech; and perform a question-and-answer response to ensure the accuracy of the question and answer.
[0116] The voice question-and-answer processing system of this embodiment using a voice encoder and a cross-attention large language model is an end-to-end model, which can avoid the high latency caused by the serial processing of the traditional cascade scheme and the performance degradation caused by error propagation, and effectively improve the question-and-answer processing efficiency.
[0117] Moreover, in this embodiment, voice pause detection at the audio frame level can be achieved, and based on this, a semantic integrity judgment and an active questioning mechanism are triggered to assist the user to complete the interaction more actively and smoothly.
[0118] Moreover, in this embodiment, when the speech pauses and the semantics are complete, it can also trigger the response of the cross-attention large language model, cache the response, and complete the pre-storage operation; when the speech ends, through the prefetch mechanism, query the corresponding response from the cache to reduce the system response latency, thereby achieving a faster and lower-latency interaction experience and effectively improving the efficiency of question-and-answer processing.
[0119] In summary, the cross-modal question-and-answer processing method based on a large model in this embodiment can achieve a faster, more accurate, richer, and more actively fluent interaction experience by adopting the above technical solutions.
[0120] Figure 5 It is a schematic diagram according to the fourth embodiment of the present disclosure; as Figure 5 shown, this embodiment provides a training method for a voice question-and-answer system, which may specifically include the following steps:
[0121] S501. Obtain training data, where the training data includes training text questions, annotated text answers, and annotated feature expressions of the training text questions;
[0122] Among them, the annotated text answer refers to the answer corresponding to the training text question. The annotated feature expression of the training text question can be vectorized and identified by using a pre-trained feature expression model to obtain the annotated feature expression. For example, the annotated feature expression of the training text question can also be a vector representation obtained by encoding the training text question in the form of one-hot encoding.
[0123] S502. Use a pre-trained mixing system to mix the training text questions to obtain training voice questions;
[0124] In order to implement cross-modal information processing for voice question-and-answer, in actual application scenarios, it is much easier to collect text information than to collect voice information. Therefore, in this embodiment, when collecting training samples, the collected are text questions and annotated text answers. Then, a pre-trained mixing system can be used to mix the training text to generate training voice questions.
[0125] In actual applications, other traditional mixing techniques can also be used to mix the training text questions to obtain training voice questions.
[0126] S503. Based on the training voice questions, training text questions, annotated feature expressions of the training text questions, and annotated text answers, adjust the parameters in the voice question-and-answer system to obtain a trained voice question-and-answer processing system.
[0127] The execution subject of the training method for the voice question-and-answer system in this embodiment can be a training device for the voice question-and-answer system, and this device can be an electronic entity or can also be a software-integrated application.
[0128] The voice question-and-answer system trained in this embodiment is the voice question-and-answer system of any one of the above Figures 1 - 3 embodiments.
[0129] In this embodiment, the training voice questions and training texts can be used as input data, and the labeled feature expressions of the training text questions and the labeled text answers can be used as supervision data. Based on the training voice questions, training text questions, labeled feature expressions of the training text questions, and labeled text answers, the parameters in the voice question-and-answer system can be adjusted to implement the training of the voice question-and-answer system.
[0130] The training method of the voice question-and-answer system in this embodiment can implement end-to-end training of the voice question-and-answer system by adopting the above method, effectively improving the accuracy of the trained voice question-and-answer system.
[0131] Figure 6 is a schematic diagram according to the fifth embodiment of the present disclosure; the training method of the voice question-and-answer system in this embodiment further describes the technical solution of the present disclosure in more detail on the basis of the technical solution of the above Figure 5 illustrated embodiment. As Figure 6 shown, the training method of the voice question-and-answer system in this embodiment may specifically include the following steps:
[0132] S601. Obtain training data, where the training data includes training text questions, labeled text answers, and labeled feature expressions of the training text questions;
[0133] S602. Use a pre-trained mixing system to mix the training text questions to obtain training voice questions;
[0134] The mixing system in this embodiment can be modeled based on a large amount of network audio and video data, user-authorized online real data, and training data of the speech synthesis system to achieve the ability to generate speech spectra from text. This system can accurately generate the speech spectrum corresponding to the input text, and can precisely and flexibly control speech-related attributes such as timbre, rhythm, emotion, and acoustic environment to generate diverse audio data covering real scenarios.
[0135] Based on this mixing system, a large number of audio input requests can be produced on a large scale; thus solving the problem of large-scale labeled audio text data required for cross-modal model training.
[0136] S603. Use the voice encoder in the voice question-and-answer system to obtain the feature expression of the training voice question based on the training text question, where the feature expression of the training voice question is equal to the length of the training text question;
[0137] The implementation manner of this step S603 is the same as the aboveFigure 2 Or Figure 3 The encoding principles of the voice encoders in the illustrated embodiments are the same and will not be elaborated here.
[0138] S604. Use the cross-attention large language model in the voice question-answering system to generate a predicted text answer based on the feature expression of the training voice question.
[0139] S605. Adjust the parameters in the voice question-answering system based on the feature expression of the training voice question, the labeled feature expression of the training text question, the labeled text answer, and the predicted text answer to obtain the trained voice question-answering processing system.
[0140] For example, when this step is specifically implemented, it may include the following steps:
[0141] (a3) Construct a first loss function, namely Loss1, based on the feature expression of the training voice question and the labeled feature expression of the training text question.
[0142] (b3) Construct a second loss function Loss2 based on the labeled text answer and the predicted text answer.
[0143] (c3) Adjust the parameters of the voice encoder and the cross-attention large language model based on the first loss function Loss1 and the second loss function Loss2 to obtain the trained voice question-answering processing system.
[0144] For example, the sum of the first loss function and the second loss function can be taken as the integrated loss function.
[0145] Adjust the parameters of the voice encoder and the cross-attention large language model based on the integrated loss function to make the integrated loss function converge.
[0146] For example, Figure 7 is the training principle diagram of the voice question-answering system of the present disclosure embodiment. As Figure 7 shown, the voice question-answering system composed of the voice encoder and the cross-attention large language model in this embodiment is an end-to-end model. During training, by adopting the above technical solution of this embodiment, the voice encoder and the cross-attention large language model are trained together, which can avoid the high latency caused by the serial processing of the traditional cascade scheme and the performance degradation caused by error propagation, and make the accuracy of the trained voice question-answering system higher.
[0147] It should be noted that considering the scarcity of existing high-quality voice Q&A data, the above technical solution of this embodiment takes the construction of training voice data using a mixing system as an example. In an actual application scenario, training data can also be constructed based on high-quality voice Q&A data. At this time, each piece of training data includes: a training voice question, an annotated text answer, and an annotated feature expression of the training text question corresponding to the training voice question; the training text question is obtained by performing speech recognition on the training voice question. The annotated feature expression of the training text question is obtained in the same way as in the above embodiment. For example, it can be encoded by using the One-Hot encoding method to obtain a vector representation of the training text question. And the training of the voice Q&A system is correspondingly performed in the above S603-S605 manner. In an actual application scenario, the training data constructed by the above two methods of this embodiment can be jointly used to train the voice Q&A system.
[0148] In this embodiment, a large amount of text Q&A data can be mixed to generate a large amount of voice Q&A data, and real voice data can be mixed for end-to-end training, greatly reducing the demand for voice Q&A data and being able to solve the problem that the commonly used solutions in the industry rely heavily on high-quality voice Q&A data.
[0149] In this embodiment, dual Loss training can also be adopted. Loss1 constrains the parameters of the voice encoder so that it retains a certain text space ability. At the same time, the hidden layer features of the voice encoder are input into the cross-attention large language model, enabling the cross-attention large language model to fuse other information outside the text, such as emotion, paralinguistic, speaker, etc., for comprehensive decision-making and providing rich responses.
[0150] The training method of the voice Q&A system in this embodiment can, by adopting the above method, achieve end-to-end training of the voice Q&A system and effectively improve the accuracy of the trained voice Q&A system.
[0151] Figure 8 is a schematic diagram according to the sixth embodiment of the present disclosure; as Figure 8 shown, this embodiment provides a cross-modal Q&A processing device 800 based on a large model, including:
[0152] An activity detection module 801 for detecting the activity of the target voice input by the user;
[0153] A text acquisition module 802 for acquiring the first text corresponding to the first input voice before the voice pause moment in the target voice in response to detecting the pause of the target voice input;
[0154] A Q&A processing module 803 for performing text response processing based on the first text and the first input voice by using a pre-trained voice Q&A processing system.
[0155] The cross-modal question-answering processing device 800 based on a large model in this embodiment realizes the implementation principle and technical effects of cross-modal question-answering processing based on a large model by adopting the above modules, which are the same as those in the above related method embodiments. For details, reference can be made to the descriptions in the above related method embodiments and will not be elaborated here.
[0156] Figure 9 is a schematic diagram according to the seventh embodiment of the present disclosure; the cross-modal question-answering processing device 900 based on a large model in this embodiment further describes the technical solution of the present disclosure in more detail on the basis of the technical solution of the above Figure 8 illustrated embodiment. As Figure 9 shown, the cross-modal question-answering processing device 900 based on a large model in this embodiment includes the modules with the same name and the same functions as those in the above Figure 8 shown: an activity detection module 901, a text acquisition module 902, and a question-answering processing module 903.
[0157] As Figure 9 shown, in this embodiment, the question-answering processing module 903 includes:
[0158] a completeness detection unit 9031, configured to obtain the semantic completeness of the first text by using a pre-trained completeness detection model;
[0159] a question-answering processing unit 9032, configured to perform text response processing by using the speech question-answering processing system based on the semantic completeness, the first text, and the first input speech.
[0160] Further optionally, in an embodiment of the present disclosure, the question-answering processing unit 9032 is configured to:
[0161] Obtain a feature representation of the first input speech by using a speech encoder in the speech question-answering processing system based on the first input speech and the first text, where the length of the feature representation of the first input speech is equal to the length of the first text;
[0162] Perform text response processing by using a cross-attention large language model in the speech question-answering processing system based on the semantic completeness and the feature representation of the first input speech.
[0163] Further optionally, in an embodiment of the present disclosure, the question-answering processing unit 9032 is configured to:
[0164] In response to determining that the semantic completeness is less than a preset completeness, generate question information based on the feature representation of the first input speech by using the cross-attention large language model;
[0165] Perform a text response based on the hesitant question information.
[0166] Further optionally, in an embodiment of the present disclosure, the question and answer processing unit 9032 is configured to:
[0167] In response to determining that the semantic integrity is greater than or equal to a preset integrity, adopt the cross-attention large language model, and generate a first text answer based on the feature expression of the first input voice; and perform a text response based on the first text answer.
[0168] Further optionally, as Figure 9 shown, the cross-modal question and answer processing device 900 based on a large model in this embodiment further includes:
[0169] A storage module 904, configured to store the feature expression of the first input voice and the first text answer in a cache.
[0170] Further optionally, in an embodiment of the present disclosure, the activity detection module 901 is further configured to continue to perform activity detection on the voice input by the user;
[0171] The question and answer processing unit 9032 is further configured to:
[0172] In response to detecting the end of the target voice input, obtain the feature expression of the second input voice of the user before the input ends;
[0173] In response to determining that the similarity between the feature expression of the second input voice and the feature expression of the first input voice in the cache is greater than or equal to a preset similarity, obtain the first text answer corresponding to the feature expression of the first input voice from the cache;
[0174] Perform a text response based on the first text answer.
[0175] Further optionally, in an embodiment of the present disclosure, the question and answer processing unit 9032 is configured to:
[0176] In response to detecting the end of the target voice input, obtain the second input voice of the user before the input ends;
[0177] Perform speech recognition on the second input voice to obtain a second text;
[0178] Based on the second input voice and the second text, adopt the speech encoder in the speech question and answer processing system to obtain the feature expression of the second input voice, wherein the length of the feature expression of the second input voice is equal to the length of the second text.
[0179] Further optionally, in an embodiment of the present disclosure, the detecting that the input of the target voice pauses includes: detecting that the duration of the stopped voice input reaches a first preset duration;
[0180] The detecting that the input of the target voice ends includes: detecting that the duration of the stopped voice input reaches a second preset duration; the second preset duration is greater than the first preset duration. The second preset duration is greater than the first preset duration.
[0181] The cross-modal question-answering processing device 900 based on a large model in this embodiment realizes the implementation principle and technical effects of cross-modal question-answering processing based on a large model by adopting the above modules, which are the same as those in the above related method embodiments. For details, reference can be made to the records in the above related method embodiments and will not be elaborated here.
[0182] Figure 10 is a schematic diagram according to the eighth embodiment of the present disclosure; this embodiment provides a training device 1000 for a voice question-answering system, and this voice question-answering system is the one used in the above Figure 8 or Figure 9 the voice question-answering system shown in the above embodiments, and specifically may include:
[0183] An acquisition module 1001, configured to acquire training data, where the training data includes training text questions, annotated text answers, and annotated feature expressions of the training text questions;
[0184] A mixing processing module 1002, configured to use a pre-trained mixing system to mix the training text questions to obtain training voice questions;
[0185] A training module 1003, configured to adjust the parameters in the voice question-answering system based on the training voice questions, the training text questions, the annotated feature expressions of the training text questions, and the annotated text answers, to obtain the trained voice question-answering processing system. The voice question-answering system in this embodiment is used in the above Figure 8 or Figure 9 the cross-modal question-answering processing device based on a large model shown above.
[0186] The training device 1000 for the voice question-answering system in this embodiment realizes the implementation principle and technical effects of training the voice question-answering system by adopting the above modules, which are the same as those in the above related method embodiments. For details, reference can be made to the records in the above related method embodiments and will not be elaborated here.
[0187] Figure 11 is a schematic diagram according to the ninth embodiment of the present disclosure; the training device 1100 for the voice question-answering system in this embodiment further describes the technical solution of the present disclosure in more detail on the basis of the technical solution of the above Figure 10 shown embodiments. AsFigure 11 As shown, the training device 1100 of the voice question-answering system in this embodiment includes the above-mentioned Figure 10 modules with the same name and function: an acquisition module 1101, a mixing processing module 1102, and a training module 1103.
[0188] As Figure 11 shown, in this embodiment, the training module 1103 includes:
[0189] An encoding unit 11031, configured to use the voice encoder in the voice question-answering system to obtain the feature representation of the training voice question based on the training text question, where the feature representation of the training voice question is equal in length to the training text question;
[0190] A text generation unit 11032, configured to use the cross-attention large language model in the voice question-answering system to generate a predicted text answer based on the feature representation of the training voice question;
[0191] An adjustment unit 11033, configured to adjust the parameters in the voice question-answering system based on the feature representation of the training voice question, the labeled feature representation of the training text question, the labeled text answer, and the predicted text answer, to obtain the trained voice question-answering processing system.
[0192] Further optionally, in an embodiment of the present disclosure, the adjustment unit 11033 is configured to:
[0193] Construct a first loss function based on the feature representation of the training voice question and the labeled feature representation of the training text question;
[0194] Construct a second loss function based on the labeled text answer and the predicted text answer;
[0195] Adjust the parameters of the voice encoder and the cross-attention large language model based on the first loss function and the second loss function to obtain the trained voice question-answering processing system.
[0196] Further optionally, in an embodiment of the present disclosure, the adjustment unit 11033 is configured to: take the sum of the first loss function and the second loss function as an integrated loss function;
[0197] Adjust the parameters of the voice encoder and the cross-attention large language model based on the integrated loss function to make the integrated loss function converge.
[0198] The training device 1100 of the voice question-answering system in this embodiment realizes the principle and technical effects of training the voice question-answering system by adopting the above modules, which are the same as those in the above-related method embodiments. For details, reference can be made to the records in the above-related method embodiments, and details will not be repeated here.
[0199] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0200] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0201] Figure 12 FIG. shows a schematic block diagram of an exemplary electronic device 1200 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0202] As Figure 12 shown, the device 1200 includes a computing unit 1201, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0203] A plurality of components in the device 1200 are connected to the I / O interface 1205, including: an input unit 1206, such as a keyboard, a mouse, etc.; an output unit 1207, such as various types of displays, speakers, etc.; a storage unit 1208, such as a magnetic disk, an optical disc, etc.; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 allows the device 1200 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0204] The computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 executes the various methods and processes described above, such as the above methods of the present disclosure. For example, in some embodiments, the above methods of the present disclosure can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the above methods of the present disclosure described above can be executed. Alternatively, in other embodiments, the computing unit 1201 can be configured to execute the above methods of the present disclosure by any other suitable means (e.g., by means of firmware).
[0205] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0206] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0207] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0208] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0209] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0210] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0211] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0212] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A cross-modal question-answering processing method based on a large model, comprising: Performing liveness detection on the target speech input by the user; When it is detected that the input of the target speech pauses, obtaining a first text corresponding to the first input speech before the speech pause moment in the target speech; Based on the first text and the first input speech, using a pre-trained speech question-answering processing system to perform text response processing.
2. The method according to claim 1, wherein The performing text response processing based on the first text and the first input speech by using a pre-trained speech question-answering processing system includes: Using a pre-trained integrity detection model to obtain the semantic integrity of the first text; Based on the semantic integrity, the first text and the first input speech, using the speech question-answering processing system to perform text response processing.
3. The method according to claim 2, wherein, The performing text response processing based on the semantic integrity, the first text and the first input speech by using the speech question-answering processing system includes: Based on the first input speech and the first text, using a speech encoder in the speech question-answering processing system to obtain a feature representation of the first input speech, wherein the length of the feature representation of the first input speech is equal to the length of the first text; Based on the semantic integrity and the feature representation of the first input speech, using a cross-attention large language model in the speech question-answering processing system to perform text response processing.
4. The method according to claim 3, wherein, The performing text response processing based on the semantic integrity and the feature representation of the first input speech by using a cross-attention large language model in the speech question-answering processing system includes: When it is determined that the semantic integrity is less than a preset integrity, using the cross-attention large language model to generate question information based on the feature representation of the first input speech; Performing a text response based on the question information.
5. The method according to claim 3, wherein, The performing text response processing based on the semantic integrity and the feature representation of the first input speech by using a cross-attention large language model in the speech question-answering processing system includes: When it is determined that the semantic integrity is greater than or equal to the preset integrity, using the cross-attention large language model to generate a first text answer based on the feature representation of the first input speech; and performing a text response based on the first text answer.
6. The method according to claim 5, wherein, After generating the first text answer based on the feature representation of the first input speech by using the cross-attention large language model, the method further includes: Storing the feature representation of the first input speech and the first text answer in a cache.
7. The method according to claim 6, wherein, The method further includes: When it is detected that the input of the target speech ends, obtaining a feature representation of the second input speech of the user before the input ends; When it is determined that the similarity between the feature representation of the second input speech and the feature representation of the first input speech in the cache is greater than or equal to a preset similarity, obtaining the first text answer from the cache; Performing a text response based on the first text answer.
8. The method according to claim 7, wherein, The obtaining a feature representation of the second input speech of the user before the input ends when it is detected that the input of the target speech ends includes: In response to detecting the end of the target voice input, obtain the second input voice of the user before the input ends; Perform speech recognition on the second input voice to obtain a second text; Based on the second input voice and the second text, use the speech encoder in the speech Q&A processing system to obtain the feature representation of the second input voice, where the length of the feature representation of the second input voice is equal to the length of the second text.
9. The method according to claim 7, wherein The detection of the pause of the target voice input includes: Detecting that the duration of the stopped voice input reaches a first preset duration; The detection of the end of the target voice input includes: Detecting that the duration of the stopped voice input reaches a second preset duration; the second preset duration is greater than the first preset duration.
10. A training method for a speech Q&A system, including: Obtain training data, where the training data includes training text questions, annotated text answers, and annotated feature representations of the training text questions; Use a pre-trained mixing system to mix the training text questions to obtain training voice questions; Based on the training voice questions, the training text questions, the annotated feature representations of the training text questions, and the annotated text answers, adjust the parameters in the speech Q&A system to obtain the trained speech Q&A processing system, and the speech Q&A system is used in any of the methods described in claims 1-9 above.
11. The method according to claim 10, wherein, Based on the training voice questions, the training text questions, the annotated feature representations of the training text questions, and the annotated text answers, adjusting the parameters in the speech Q&A system includes: Use the speech encoder in the speech Q&A system to obtain the feature representation of the training voice question based on the training text question, where the feature representation of the training voice question is equal to the length of the training text question; Use the cross-attention large language model in the speech Q&A system to generate a predicted text answer based on the feature representation of the training voice question; Based on the feature representation of the training voice question, the annotated feature representation of the training text question, the annotated text answer, and the predicted text answer, adjust the parameters in the speech Q&A system.
12. The method according to claim 11, wherein, Based on the feature representation of the training voice question, the annotated feature representation of the training text question, the annotated text answer, and the predicted text answer, adjusting the parameters in the speech Q&A system includes: Construct a first loss function based on the feature representation of the training voice question and the annotated feature representation of the training text question; Construct a second loss function based on the annotated text answer and the predicted text answer; Based on the first loss function and the second loss function, adjust the parameters of the speech encoder and the cross-attention large language model.
13. The method according to claim 12, wherein, Based on the first loss function and the second loss function, adjusting the parameters of the speech encoder and the cross-attention large language model includes: Take the sum of the first loss function and the second loss function as the integrated loss function; Based on the integrated loss function, adjust the parameters of the speech encoder and the cross-attention large language model so that the integrated loss function converges.
14. A cross-modal question-answering processing device based on a large model, comprising: An activity detection module for detecting the activity of the target speech input by the user; A text acquisition module for acquiring the first text corresponding to the first input speech before the speech pause moment in the target speech when it is detected that the input of the target speech pauses; A question-answering processing module for performing text response processing based on the first text and the first input speech by using a pre-trained speech question-answering processing system.
15. A training device for a speech question-answering system, comprising: An acquisition module for acquiring training data, where the training data includes training text questions, annotated text answers, and annotated feature expressions of the training text questions; A mixing processing module for mixing the training text questions by using a pre-trained mixing system to obtain training speech questions; A training module for adjusting the parameters in the speech question-answering system based on the training speech questions, the training text questions, the annotated feature expressions of the training text questions, and the annotated text answers to obtain the trained speech question-answering processing system, and the speech question-answering system is used in the device described in claim 14 above.
16. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of claims 1-13.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method described in any one of claims 1-13.
18. A computer program product, comprising a computer program which, when executed by a processor, implements the method described in any one of claims 1-13.
Citation Information
Patent Citations
Tiredness detection early warning method and system and computing device
CN116013373A
Call center voice processing system and method based on artificial intelligence
CN116778927A
Full-duplex dialogue system and method, electronic equipment and storage medium
CN118366458A
End-to-end multi-modal large model training method and system for spoken language questions and answers
CN118782040A