Training method and device of voice multi-mode interaction model

By training the speech multimodal interaction model and aligning the text and audio features, the problem of processing after converting the speech text in the prior art is solved, and the direct response and efficient recognition of the speech interaction model for user speech is realized.

CN120071898APending Publication Date: 2025-05-30BAIRONG ZHIXIN (BEIJING) TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510151385.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing voice interaction model cannot directly identify and process the voice input by the user. It needs to convert the voice into text before processing, resulting in inaccurate semantic expression and affecting the quality of the response.

Method used

A training method for speech multimodal interaction model is proposed. By obtaining a training sample set containing prompt text, prompt audio and reply text, and combining text and audio features for model training, the alignment of audio features and text features is achieved, and semantic understanding is enhanced.

Benefits of technology

The voice interaction model is realized to directly respond to user voice, improve the accuracy and efficiency of the response, and enhance the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071898A_ABST
    Figure CN120071898A_ABST
Patent Text Reader

Abstract

The invention discloses a voice multi-mode interaction model training method and device, and the method comprises the steps: obtaining a training sample set which comprises a plurality of prompt texts, and a prompt audio and a sample reply text corresponding to each prompt text; inputting the training sample set into a to-be-trained voice multi-mode interaction model for model training to obtain a prompt text feature corresponding to each prompt text, and predicting prompt audio features corresponding to the reply text and the prompt audio; determining a first loss value of the trained voice multi-mode interaction model based on the prompt text feature and the prompt audio feature corresponding to each prompt text, and determining a second loss value of the trained voice multi-mode interaction model based on the predicted reply text and the sample reply text corresponding to each prompt text; and if it is determined that the trained voice multi-modal interaction model converges according to the first loss value and the second loss value, determining the trained voice multi-modal interaction model as a trained voice multi-modal interaction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of voice interaction, and in particular, to a method and device for training a voice multi-modal interaction model. Background Art

[0002] With the development of artificial intelligence technology, voice interaction models are widely used in voice interaction products to identify and process the content input by users through the voice interaction model, and output corresponding responses to achieve voice interaction between users and products.

[0003] Currently, voice interaction models usually rely on samples of only one modality, i.e., prompt texts, and are trained and generated in combination with the corresponding reply texts of the prompt texts. Voice interaction models cannot directly identify and process the voices input by users. Instead, the voices input by users need to be first converted into texts, and then the converted texts are used as the inputs of the models. Only then can the voice interaction models identify and process the input texts and output corresponding responses. However, the texts obtained by converting voices may not truly express the semantics actually expressed by the voices, resulting in the voice interaction models being unable to accurately output responses adapted to the voices.

[0004] Therefore, how to train a voice multi-modal interaction model that can accurately output responses adapted to users' voices has become an urgent problem to be solved currently. Summary of the Invention

[0005] The present application proposes a method and device for training a voice multi-modal interaction model, and the main purpose is to train a voice multi-modal interaction model that can accurately output responses adapted to users' voices.

[0006] To achieve the above object, the present application mainly provides the following technical solutions:

[0007] In a first aspect, the present application provides a method for training a voice multi-modal interaction model. The method for training the voice multi-modal interaction model provided in this embodiment may include:

[0008] Obtain a training sample set, where the training sample set includes multiple prompt texts and a corresponding prompt audio and a sample reply text for each prompt text, and the prompt audio is the audio expression corresponding to the prompt text;

[0009] Input the training sample set into a voice multi-modal interaction model to be trained for model training, and obtain a prompt text feature, a predicted reply text, and a prompt audio feature corresponding to the prompt audio for each prompt text;

[0010] Determine the first loss value of the trained speech multi-modal interaction model based on the prompt text features and prompt audio features corresponding to each of the said prompt texts, and determine the second loss value of the trained speech multi-modal interaction model based on the predicted response text and the sample response text corresponding to each of the said prompt texts;

[0011] If it is determined that the trained speech multi-modal interaction model converges according to the first loss value and the second loss value, then determine the trained speech multi-modal interaction model as the trained speech multi-modal interaction model.

[0012] In a second aspect, the present application provides a multi-modal processing method. The multi-modal processing method provided in this embodiment may include: obtaining information to be processed, where the information to be processed includes content of at least one modality among audio and text; inputting the information to be processed into a speech multi-modal interaction model to obtain an identification result output by the speech multi-modal interaction model, and the speech multi-modal interaction model is generated by the training method of the speech multi-modal interaction model in the first aspect.

[0013] In a third aspect, the present application provides a training device for a speech multi-modal interaction model. The training device for the speech multi-modal interaction model provided in this embodiment may include:

[0014] An acquisition module, configured to acquire a training sample set, where the training sample set includes a plurality of prompt texts and a prompt audio and a sample response text corresponding to each prompt text, and the prompt audio is an audio expression corresponding to the prompt text;

[0015] A training module, configured to input the training sample set into a speech multi-modal interaction model to be trained for model training, and obtain prompt text features, predicted response texts, and prompt audio features corresponding to the prompt audio corresponding to each of the said prompt texts;

[0016] A first determination module, configured to determine the first loss value of the trained speech multi-modal interaction model based on the prompt text features and prompt audio features corresponding to each of the said prompt texts, and determine the second loss value of the trained speech multi-modal interaction model based on the predicted response text and the sample response text corresponding to each of the said prompt texts;

[0017] A second determination module, configured to, if a judgment module determines that the trained speech multi-modal interaction model converges according to the first loss value and the second loss value, then obtain the trained speech multi-modal interaction model as the trained speech multi-modal interaction model.

[0018] Fourthly, the present application provides a computer-readable storage medium, where the storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the training method of the voice multi-modal interaction model described in the first aspect, and / or the multi-modal processing method described in the second aspect.

[0019] Fifthly, the present application provides an electronic device, where the electronic device includes: a memory for storing a program; a processor coupled to the memory for running the program to execute the training method of the voice multi-modal interaction model described in the first aspect, and / or the multi-modal processing method described in the second aspect.

[0020] The training method and device for the voice multi-modal interaction model provided by this application, when it is necessary to train the voice multi-modal interaction model, first obtain a training sample set. The training sample set includes multiple prompt texts, as well as the sample response texts and prompt audios corresponding to each prompt text, and input the training sample set into the voice multi-modal interaction model to be trained for model training, so as to obtain the prompt text features, predicted response texts, and prompt audio features corresponding to each prompt text. Then, based on the prompt text features and prompt audio features corresponding to each prompt text, determine the first loss value of the trained voice multi-modal interaction model, and based on the predicted response texts and sample response texts corresponding to each prompt text, determine the second loss value of the trained voice multi-modal interaction model. Finally, if it is determined that the trained voice multi-modal interaction model converges according to the first loss value and the second loss value, then determine the trained voice multi-modal interaction model as the trained voice multi-modal interaction model. It can be seen that the solution provided by this application has at least the following invention points: First, when training the voice multi-modal interaction model, samples of two different modalities, text and audio, are used. In this way, when applying the trained voice multi-modal interaction model to respond to the user's input voice, the corresponding audio and text of the input voice can be combined to more comprehensively and accurately identify the semantics actually expressed in the user's voice. Second, during the training process of the voice multi-modal interaction model, through the loss calculation between the prompt text features and prompt audio features obtained during the training process, the voice multi-modal interaction model can enhance the semantic understanding of the prompt audio based on the prompt text features and retain the understanding of the unique features of the prompt audio by the voice multi-modal interaction model. That is, during the training process of the voice multi-modal interaction model, through loss calculation, the alignment of audio features and text features can be achieved, thereby strengthening the semantic information of audio features and retaining the unique features of audio. Then, based on the aligned audio features and text features, multi-modal features for training the voice multi-modal interaction model are obtained, so that the trained voice multi-modal interaction model can directly respond to the user's input voice, rather than responding after converting the voice to text as in the background technology. Since the solution provided by this embodiment directly generates a response corresponding to the voice when responding to the user's voice, saving the process of "converting the voice to text and then responding to the converted text", the solution provided by this embodiment has high response efficiency; and because the solution provided by this embodiment aligns audio features and text features, strengthens the semantic information of audio features, and retains the unique features of audio during the training of the voice multi-modal interaction model, the solution provided by this embodiment can accurately understand the meaning expressed in the user's audio and has good response quality.Thirdly, during the training process of the voice multi-modal interaction model, by calculating the loss between the predicted response text obtained from training the voice multi-modal interaction model and the sample response text, the fitting degree of the voice multi-modal interaction model on the training data is continuously improved, so that the voice multi-modal interaction model can accurately predict the response text. Based on the above three inventive points, it can be seen that the voice multi-modal interaction model trained according to the solution provided in this application can use the voice input by the user and / or the text corresponding to the voice as the model input. The voice multi-modal interaction model can accurately recognize the semantics actually expressed by the user based on the model input, so as to accurately output a response adapted to the input voice and / or input text, thereby improving the voice interaction effect with the user and bringing a good voice interaction experience to the user.

[0021] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the specific embodiments of this application are specifically described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0023] Figure 1 The flowchart of a method for training a voice multi-modal interaction model provided by an embodiment of this application is shown;

[0024] Figure 2 The schematic diagram of the process of training a voice multi-modal interaction model provided by an embodiment of this application is shown;

[0025] Figure 3 The flowchart of a multi-modal processing method provided by an embodiment of this application is shown;

[0026] Figure 4 The application scenario architecture diagram of a voice multi-modal interaction model provided by an embodiment of this application is shown;

[0027] Figure 5 The structural schematic diagram of a training device for a voice multi-modal interaction model provided by an embodiment of this application is shown;

[0028] Figure 6 The structural schematic diagram of a training device for a voice multi-modal interaction model provided by another embodiment of this application is shown;

[0029] Figure 7 The figure shows a schematic structural diagram of a multimodal processing device provided by an embodiment of the present application. Detailed implementation manners

[0030] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0031] Currently, the voice interaction model used to implement voice interaction between a user and a voice interaction product usually depends on samples of only one modality, i.e., prompt text, and is trained in combination with the response text corresponding to the prompt text. The model training process lacks samples of modalities such as audio. In this way, when the trained voice interaction model conducts voice interaction, it cannot directly use the voice input by the user as the input of the voice interaction model. Instead, it is necessary to first convert the voice input by the user into text of this modality, and then use the converted text as the input of the voice interaction model, so that the voice interaction model can recognize and process the input text and output a corresponding response. However, when converting voice to text, the text obtained by converting the voice may not accurately represent the semantics actually expressed by the voice due to inaccurate conversion, resulting in the voice interaction model being unable to accurately output a response adapted to the voice, thereby affecting the voice interaction effect between the voice interaction product and the user and bringing a bad experience to the user.

[0032] Based on research findings, if a voice multimodal model is trained by relying on samples of both prompt text and the prompt audio corresponding to the prompt text and combining the response texts corresponding to both the prompt text and the prompt audio, then when applying the trained voice multimodal interaction model for voice interaction, the voice input by the user and / or the text corresponding to the voice can be used as the model input, and the voice multimodal interaction model can output a response corresponding to the input voice and / or the input text. In some application scenarios based on the above research, the voice multimodal interaction model can recognize inputs of at least two modalities and give responses corresponding to the input modalities, that is: both voice and text can be used as the model input of the voice multimodal interaction model at the same time, or voice / text alone can be used as the model input of the voice multimodal interaction model, and the voice multimodal interaction model can accurately recognize the semantics actually expressed by the user based on the model input, so as to accurately output a response adapted to the model input.

[0033] Based on the above findings, the embodiments of the present application specifically provide a technical solution for training a voice multi-modal interaction model, specifically as follows: Obtain a training sample set, where the training sample set includes multiple prompt texts, as well as the corresponding prompt audio and sample response texts for each prompt text, and the prompt audio is the audio expression corresponding to the prompt text. Input the training sample set into the voice multi-modal interaction model to be trained for model training to obtain the prompt text features, predicted response texts, and prompt audio features corresponding to the prompt audio for each prompt text. Then, based on the prompt text features and prompt audio features corresponding to each prompt text, determine the first loss value of the trained voice multi-modal interaction model, and based on the predicted response texts and sample response texts corresponding to each prompt text, determine the second loss value of the trained voice multi-modal interaction model. If it is determined that the trained voice multi-modal interaction model converges according to the first loss value and the second loss value, then determine the trained voice multi-modal interaction model as the trained voice multi-modal interaction model. The voice multi-modal interaction model trained according to the solution provided in this embodiment can use the voice input by the user and / or the text corresponding to the voice as the model input. The voice multi-modal interaction model can accurately recognize the semantics actually expressed by the user based on the model input, so as to accurately output a response adapted to the input voice and / or input text, and further improve the voice interaction effect with the user, bringing a good voice interaction experience to the user.

[0034] The technical solution for training the voice multi-modal interaction model provided in this embodiment can train an adapted voice multi-modal interaction model for any business scenario, and this embodiment does not limit the business scenario. Exemplarily, the business scenario can include, but is not limited to, any one of the following: intelligent customer service scenarios for specific services, companion scenarios, healing scenarios, chat scenarios, conversation scenarios, consultation scenarios, etc. The specific service can be flexibly selected based on the needs of the business field. For example, the specific service can include, but is not limited to, bank debt collection and product promotion in the fintech field, digital human guidance, AI education, and psychological counseling in the education field, intelligent robots in the smart home field, and can also include, but is not limited to, live broadcast services in the multimedia field, office tools (such as IM communication tools, knowledge bases, programming, and documents) in the automated office field, etc.

[0035] Based on the above technical solution for training the voice multi-modal interaction model, this embodiment specifically provides a training method and device for the voice multi-modal interaction model. The following specifically describes the training method and device for the voice multi-modal interaction model provided in this embodiment.

[0036] As Figure 1 shown, the embodiments of the present application provide a training method for a voice multi-modal interaction model. The training method for the voice multi-modal interaction model provided in this embodiment can at least include the following steps 101 to 104.

[0037] 101. Obtain a training sample set, where the training sample set includes multiple prompt texts, the prompt audio corresponding to each prompt text, and sample response texts, and the prompt audio is the audio expression corresponding to the prompt text.

[0038] Before training a speech multi-modal interaction model, it is necessary to obtain a training sample set to train the speech multi-modal interaction model based on the training sample set. In order to enable the speech multi-modal interaction model to accurately respond to the speech input by the user, when training the speech multi-modal interaction model, samples of two different modalities, text and audio, need to be used, so that when applying the trained speech multi-modal interaction model to respond to the speech input by the user, the audio and text corresponding to the speech can be combined simultaneously to achieve a more comprehensive and accurate recognition of the semantics actually expressed by the user and accurately output a response adapted to the user's speech. Based on this, the training sample set should at least include multiple prompt texts, the prompt audio corresponding to each prompt text, and sample response texts. The prompt audio is the audio expression corresponding to the prompt text. The prompt text is the text that needs to be responded to, and the type of the prompt text is related to the business scenario. Based on this, the type of the prompt text can include but is not limited to at least one of the following: question text, consultation text, and instruction text.

[0039] The methods for obtaining the training sample set can at least include the following Method 1 and Method 2.

[0040] Method 1, the specific process of obtaining the training sample set can include the following steps 101A to 101C.

[0041] 101A. Extract the prompt texts and the response texts corresponding to the prompt texts for constructing the training sample set from the user interaction data corresponding to the business scenario where the speech multi-modal interaction model is to be applied.

[0042] The speech multi-modal interaction model is usually applied to a specific business scenario. Based on this, it is necessary to obtain the training sample set according to the user interaction data corresponding to the business scenario where the speech multi-modal interaction model is to be applied, so as to train a speech multi-modal interaction model that can accurately respond to the speech input by the user in the business scenario. In some embodiments, the business scenario can include but is not limited to any one of the following: intelligent customer service scenarios for specified services, companion scenarios, healing scenarios, chat scenarios, conversation scenarios, consultation scenarios, etc. The user interaction data can include but is not limited to: question-and-answer interaction data with the user, consultation interaction data with the user, and instruction interaction data with the user. The user can include but is not limited to devices, devices, virtual humans, and real humans that output speech data.

[0043] After obtaining the user interaction data corresponding to the business scenario, extract the prompt text for constructing the training sample set and the sample response text corresponding to the prompt text from the user interaction data. This extraction process may include the following steps: input the user interaction data into a preset extraction model, and the preset extraction model extracts the prompt text and the sample response text corresponding to the prompt text from the user interaction data. Among them, the preset extraction model is used to extract the prompt text and the corresponding sample response text from the user interaction data. The preset extraction model is trained based on multiple groups of data, and each group of data includes user interaction data, the prompt text included in the user interaction data, and the sample response text corresponding to the prompt text. The model type of the preset extraction model can be flexibly selected based on business needs, and the model type of the preset extraction model is not limited in this embodiment. In principle, any model that has the ability to extract prompt text and response text from user interaction data after training can be used as the preset extraction model.

[0044] 101B. Set corresponding emotional features and timbre features for each prompt text.

[0045] The voice multi-modal interaction model needs to be trained based on samples of two modalities, audio and text. Therefore, after extracting the target prompt text, it is necessary to convert the prompt text into the corresponding prompt audio based on the audio to effectively utilize the text modality and audio modality of the same question to train the voice multi-modal interaction model.

[0046] Considering that when the user inputs voice, it may have its own emotional features and timbre features, and these emotional features and timbre features are unique features of audio, which affect the voice multi-modal interaction model's understanding of the semantics of audio. Based on this, in order to enable the voice multi-modal interaction model to accurately recognize and understand the semantics of audio features in the audio modality, it is necessary to set corresponding emotional features and timbre features for each prompt text before converting the prompt text into the corresponding prompt audio.

[0047] To implement the setting of emotional features and timbre features, in some embodiments, the business scenario of the speech multi-modal interaction model to be applied may be determined first, and the emotional features and timbre features of each user that may be involved in this business scenario are evaluated. The emotional features and timbre features are clustered based on user categories (this category can be represented by, but not limited to, the following factors: gender, age, education level, household register, etc.) to obtain the emotional features and timbre features corresponding to each user category. Based on this, for each prompt text, the following is performed respectively: determining the user category whose probability of inputting the current prompt text is greater than the probability threshold; if the determined user category is one, setting the emotional features and timbre features corresponding to this user category as the emotional features and timbre features of the current prompt text; if the determined user category is multiple, creating a copy of the current prompt text, and after creation, the number of the current prompt texts is the same as the number of the determined user categories, and setting the corresponding emotional features and timbre features for each current prompt text based on the emotional features and timbre features corresponding to each user category. When the user category is multiple, in the obtained training sample set, the current prompt text will exist repeatedly, and the prompt audio corresponding to the repeatedly existing current prompt text is converted based on the same prompt text and different emotional features and / or timbre features.

[0048] The specific process of determining the user category whose probability of inputting the current prompt text is greater than the probability threshold may include: obtaining multiple user categories summarized in advance; identifying the probability of each user category inputting the current prompt text through a preset target model; detecting whether there is a user category whose probability of inputting the current prompt text is greater than the probability threshold; if there is, determining the user category whose probability of inputting the current prompt text is greater than the probability threshold; if not, issuing a prompt to inform the model trainer to determine the user category for the current prompt text based on the prompt. The preset target model is trained based on multiple sets of data and is used to determine the probability of a user category inputting a prompt text. Each set of data includes a prompt text, a user category, and the probability of the user category inputting the prompt text.

[0049] If the determined user category is one, it indicates that the current prompt text is probably input by the user of this user category. Therefore, the emotional features and timbre features corresponding to this user category are set as the emotional features and timbre features of the current prompt text, so as to convert the current prompt text into the corresponding prompt audio based on the emotional features and timbre features, and enable the speech multi-modal interaction model to accurately learn and process the speech of the current prompt text input by this user category.

[0050] If multiple user categories are determined, it means that multiple user categories may all input the current prompt text. Therefore, a copy of the current prompt text is created. After creation, the number of copies of the current prompt text is the same as the number of determined user categories. Then, based on the emotional characteristics and voice characteristics corresponding to each determined user category, the corresponding emotional characteristics and voice characteristics are set for each copy of the current prompt text, so as to convert the current prompt text into voices with different emotional characteristics and voice characteristics, enabling the voice multi-modal interaction model to accurately learn and process the voices corresponding to the current prompt text input by different user categories. Exemplarily, the current prompt text is "Prompt Text 1", and the user categories for which the probability of inputting "Prompt Text 1" is greater than the probability threshold are the three user categories of "User Category 1, User Category 2, and User Category 3". A copy of "Prompt Text 1" is created. After creating the copy, the number of "Prompt Text 1" is three. Based on the emotional characteristic 1 and voice characteristic 1 corresponding to User Category 1, the corresponding emotional characteristics and voice characteristics are set for the first "Prompt Text 1". Based on the emotional characteristic 2 and voice characteristic 2 corresponding to User Category 2, the corresponding emotional characteristics and voice characteristics are set for the second "Prompt Text 1". Based on the emotional characteristic 3 and voice characteristic 3 corresponding to User Category 3, the corresponding emotional characteristics and voice characteristics are set for the third "Prompt Text 1". In this way, in the obtained training sample set, "Prompt Text 1" will exist three times, and the prompt audios corresponding to the three repeated "Prompt Text 1" are converted based on different emotional characteristics and voice characteristics.

[0051] 101C. Perform audio conversion on the prompt text based on the emotional characteristics and voice characteristics corresponding to the prompt text to obtain the prompt audio corresponding to each prompt text.

[0052] For each piece of prompt text, perform the following steps: If in step 101B, the corresponding sentiment feature and tone feature are set for the current prompt text based on only one user category, then input the current prompt text, the corresponding sentiment feature and tone feature of the current prompt text into the audio synthesis model, and the audio synthesis model synthesizes the prompt audio corresponding to the current prompt text based on the corresponding input. If in step 101B, the corresponding sentiment feature and tone feature are set for the current prompt text based on multiple user categories, then for each user category, perform the following on the current prompt text: Based on the sentiment feature and tone feature set for the current prompt text based on the current user category, convert the current prompt text into the corresponding prompt audio under the current user category through the audio synthesis model. Exemplarily, the current prompt text is "prompt text 1", and the user categories determined to have a probability greater than the probability threshold for inputting "prompt text 1" are the three user categories of "user category 1, user category 2, and user category 3". Create copies of "prompt text 1", and after creating the copies, the number of "prompt text 1" is three. Set the corresponding sentiment feature and tone feature for the first "prompt text 1" based on the sentiment feature 1 and tone feature 1 corresponding to user category 1. Set the corresponding sentiment feature and tone feature for the second "prompt text 1" based on the sentiment feature 2 and tone feature 2 corresponding to user category 2. Set the corresponding sentiment feature and tone feature for the third "prompt text 1" based on the sentiment feature 3 and tone feature 3 corresponding to user category 3. When determining the prompt audio corresponding to "prompt text 1", based on the sentiment feature 1 and tone feature 1 set for "prompt text 1" based on user category 1, convert "prompt text 1" into the corresponding prompt audio 1 under user category 1 through the audio synthesis model, and based on the sentiment feature 2 and tone feature 2 set for "prompt text 2" based on user category 2, convert "prompt text 2" into the corresponding prompt audio 2 under user category 2 through the audio synthesis model, and based on the sentiment feature 3 and tone feature 3 set for "prompt text 3" based on user category 3, convert "prompt text 3" into the corresponding prompt audio 3 under user category 3.

[0053] It should be noted that the audio synthesis model is used to convert the prompt text into a prompt audio with corresponding sentiment features and tone features based on the prompt text and the corresponding sentiment features and tone features of the prompt text. The audio synthesis model is trained based on multiple groups of data, and each group of data includes a prompt audio and the corresponding prompt text, sentiment feature, and tone feature of the prompt audio. The model type of the audio synthesis model can be flexibly selected based on business needs, and this embodiment does not limit this.

[0054] After the above 101A - 101C, a training sample set is obtained, and this training sample set includes prompt text and the corresponding sample reply text and prompt audio of the prompt text.

[0055] It should be noted that if there are multiple user categories determined in step 101B for some prompt texts, the obtained training sample set will have the following characteristics: the training sample set includes at least multiple target prompt texts, the target response texts corresponding to each target prompt text, and the target prompt audio. The contents of the multiple target prompt texts are the same. The prompt text based on which the target prompt audio corresponding to each target prompt text is obtained is the same, and the target features are different. The target features include at least one of the following: emotional feature and timbre feature.

[0056] Method 2. The specific process of obtaining the training sample set may include the following step 101D.

[0057] 101D. Obtain a real corpus set corresponding to the business scenario of the speech multi-modal interaction model to be applied. The real corpus set includes multiple prompt audios, and the prompt audios are the audios emitted by natural persons; convert each prompt audio in the real corpus set into the corresponding prompt text, and obtain the response text adapted to the prompt text; summarize the prompt audio and the response text corresponding to each prompt text to obtain the training sample set.

[0058] Considering that there are certain differences between the prompt audio corresponding to the prompt text converted by the audio synthesis model in Method 1 and the audio expression of natural persons, and the existence of these differences may affect the speech recognition and processing effects of the speech multi-modal interaction model. Based on this, in order to reduce the influence of these differences, the prompt audios in the training sample set of this embodiment are all real prompt audios emitted by natural persons collected from the real corpus set. The real prompt audios emitted by natural persons carry corresponding emotional features and timbre features. The speech multi-modal interaction model trained based on such a training sample set can more accurately recognize and process the real speech of natural persons.

[0059] After obtaining the prompt audio from the real corpus set, convert the prompt audio into the corresponding prompt text through a preset model for converting audio into text (such as, but not limited to, a large language model). Then, based on the preset empirical mapping relationship data between the prompt text and the response text, obtain the response text adapted to the prompt text. Finally, summarize the prompt text, the prompt audio and the response text corresponding to each prompt text to obtain the training sample set.

[0060] The above two methods for obtaining the training sample set can be used alone or in combination, and this embodiment does not make any limitation in this regard.

[0061] In some improved embodiments, to improve the generalization of the speech multi-modal interaction model, the training sample set includes prompt texts in at least two languages, and the languages of the prompt audio and the sample response text are consistent with the languages of the corresponding prompt texts. The types of languages here may include, but are not limited to: the languages of countries, dialects of regions, etc.

[0062] After obtaining the training sample set, that is, having the prerequisite for training the speech multi-modal interaction model with samples in two different modalities of text and audio, based on this, step 102 can be continued to train the speech multi-modal interaction model.

[0063] 102. Input the training sample set into the speech multi-modal interaction model to be trained for model training, and obtain the prompt text features, predicted response texts, and prompt audio features corresponding to each prompt text.

[0064] The speech multi-modal interaction model usually undergoes multiple iterations of training to accurately output responses adapted to the model input. Based on this, the speech multi-modal interaction model to be trained can include the following two situations:

[0065] In one case, when using the training sample set for the first model training, the speech multi-modal interaction model to be trained is a model that has not been trained based on the training sample set.

[0066] In another case, when using the training sample set for non-first model training, the speech multi-modal interaction model to be trained is a model that has not converged after the previous training.

[0067] The model training process is related to the specific structure of the speech multi-modal interaction model. The following explains the model training process in combination with the structure of the speech multi-modal interaction model. The speech multi-modal interaction model includes an audio encoder and a large language model including at least a self-attention layer and an embedding layer. Then, the specific process of inputting the training sample set into the speech multi-modal interaction model to be trained for model training and obtaining the prompt text features, predicted response texts, and prompt audio features corresponding to each prompt text can include the following steps 102A to 102D.

[0068] 102A. Process the input prompt audio through the audio encoder to obtain initial audio features, and process the input prompt text and sample response text through the embedding layer to obtain prompt text features and sample response text features.

[0069] When training the speech multi-modal interaction model, such as Figure 2 ( Figure 2As shown in the schematic diagram of the process of training the speech multi-modal interaction model, the prompt audio corresponding to each prompt text is input into the audio encoder, and the input prompt audio is processed by the audio encoder to obtain the initial audio features. The audio encoder is used to encode the initial audio features of the prompt audio. The audio encoder can be flexibly selected based on business needs, and this embodiment does not limit it. In principle, any audio encoder that can encode audio into audio features can be used as the audio encoder in this embodiment. It should be noted that the selected audio encoder should minimize the loss of the original features of the prompt audio for the initial audio features, and avoid the situation where the initial audio features cannot reflect the true semantics of the prompt audio.

[0070] In this embodiment, the prompt text and the sample response text corresponding to each prompt text are also input into Figure 2 the embedding layer (i.e., the embedding layer) in, and the input prompt text and sample response text are processed by the embedding layer to obtain the prompt text features and sample response text features corresponding to each prompt text.

[0071] 102B. Through the self-attention layer, the initial audio features are semantically aligned with the prompt text features of the corresponding prompt text, and the aligned initial audio features are dimensionally reduced to obtain the prompt audio features corresponding to each prompt text.

[0072] After obtaining the initial audio features and the prompt text features of the corresponding prompt text, as Figure 2As shown in the figure, the initial audio features and the prompt text features of the corresponding prompt text are input into the self-attention layer (i.e., the Self-Attention layer). The self-attention layer has the ability of semantic alignment and dimensionality reduction. The self-attention layer has a selection mechanism, and its selection mechanism can retain the audio features unique to the prompt audio and the audio features that enhance the audio semantics. The working process of the self-attention layer is as follows: Semantically align the initial audio features with the prompt text features of the corresponding prompt text through the self-attention layer, so as to enhance the semantic information of the prompt audio through the semantics of the prompt text features and retain the unique audio features of the prompt audio; Considering that the data volume of the audio features is large, and a large data volume is not conducive to the training efficiency of model training, and the data volume of the audio features is caused by the relatively large number of dimensions of the audio features themselves, it is necessary to perform dimensionality reduction processing on the aligned initial audio features to reduce the data volume of the audio features. The methods of dimensionality reduction processing can include the following two: One is to determine the dimension of the prompt text features and reduce the dimension of the aligned initial audio features to the dimension of the prompt text, that is, the dimension of the initial audio features is the same as the dimension of the prompt text; The other is to determine a preset dimension, and the preset dimension is lower than the original dimension of the initial audio features, and reduce the dimension of the aligned initial audio to the preset dimension, that is, the dimension of the initial audio features is the preset dimension. The above two dimensionality reduction processing methods can be flexibly selected based on business needs. This embodiment can also adopt other dimensionality reduction methods, and this embodiment does not make any limitations in this regard.

[0073] After performing dimensionality reduction processing on the aligned initial audio features, the prompt audio features of each prompt text are obtained, that is, the prompt audio features corresponding to the prompt text are the reduced-dimensional initial audio features corresponding to the prompt text.

[0074] 102C. Concatenate the prompt audio features and the prompt text features corresponding to each prompt text to obtain the multi-modal features corresponding to each prompt text respectively.

[0075] After obtaining the prompt audio features and the prompt text features, as Figure 2 shown in the figure, it is necessary to concatenate the prompt audio features and the prompt text features corresponding to each prompt text to obtain the multi-modal features corresponding to each prompt text respectively, so as to train the speech multi-modal interaction model based on the samples that fuse the two modalities of text and audio, that is, train the speech multi-modal interaction model based on the multi-modal features that fuse the two modalities of text and audio.

[0076] When concatenating the prompt audio features and the prompt text features, whether to place the prompt audio features in the front or the prompt text features in the front can be flexibly determined based on business needs. This embodiment does not make any limitations on the feature concatenation method.

[0077] 102D. Train the large language model using the multimodal features corresponding to each prompt text and the sample response text features to obtain the predicted response text corresponding to each prompt text.

[0078] After obtaining the multimodal features, the large language model will be trained using the multimodal features corresponding to each prompt text and the sample response text features. In some embodiments, the training process can be, as Figure 2 shown, the large language model further includes a prediction layer, which concatenates the multimodal features corresponding to each prompt text and the sample response text features respectively ( Figure 2 the symbol "+" in it represents concatenation), and inputs the concatenated features of each prompt text into the prediction layer to train the prediction layer through these concatenated features.

[0079] Through steps 102A to 102D, the following data for evaluating the training effect of the speech multimodal interaction model after the current training can be obtained: the prompt text features corresponding to each prompt text, the predicted response text, and the prompt audio features corresponding to the prompt audio. Through these data, it can be evaluated whether the speech multimodal interaction model has the ability to accurately output responses adapted to the model input after the current training is completed.

[0080] 103. Determine the first loss value of the trained speech multimodal interaction model based on the prompt text features and prompt audio features corresponding to each prompt text, and determine the second loss value of the trained speech multimodal interaction model based on the predicted response text and sample response text corresponding to each prompt text.

[0081] To evaluate whether the speech multimodal interaction model has the ability to accurately output responses adapted to the user's speech, the following two data are used: one is the first loss value of the trained speech multimodal interaction model obtained based on the prompt text features and prompt audio features corresponding to each prompt text; the other is the second loss value of the trained speech multimodal interaction model obtained based on the predicted response text and sample response text corresponding to each prompt text. The methods for obtaining the first loss value and the second loss value will be specifically described below.

[0082] First, for the first loss value.

[0083] The degree of semantic difference between the prompt text features and the prompt audio features recognized by the speech multi-modal interaction model affects the ability of the speech multi-modal interaction model to accurately output responses adapted to the user's speech. Therefore, it is necessary to determine the first loss value of the trained speech multi-modal interaction model based on the prompt text features and the prompt audio features corresponding to each prompt text. The first loss value is used to reflect the degree of semantic difference between the prompt text features and the prompt audio features recognized by the speech multi-modal interaction model. Referencing the first loss value in the judgment of whether the speech multi-modal interaction model converges can control the trained speech multi-modal interaction model to align the prompt text features and the prompt audio features, thereby avoiding the loss of audio features that affect the response accuracy during the recognition process of the speech multi-modal interaction model.

[0084] To enable the calculation of the first loss value, as Figure 2 shown, the speech multi-modal interaction model is provided with a first loss calculation method for calculating the semantic difference between the prompt text features and the prompt audio features, and the first loss calculation method includes at least one calculation method. Based on this, the specific process of determining the first loss value of the trained speech multi-modal interaction model based on the prompt text features and the prompt audio features corresponding to each prompt text may include the following steps 103A to 103B:

[0085] 103A. Use each first loss calculation method to perform arithmetic processing on the prompt text features and the prompt audio features corresponding to each prompt text, and obtain the loss value obtained by each first loss calculation method.

[0086] The first loss calculation method can select one or more calculation methods based on business needs. Considering that a single first loss calculation method may have occasional errors, resulting in inaccurate calculation of the first loss value, multiple first loss calculation methods are used to reduce the impact of occasional errors of individual first loss calculation methods on the correctness of the first loss value. The first loss calculation method can be flexibly selected based on business needs, and this embodiment does not limit it. Exemplarily, the first loss calculation method is deployed with a corresponding first loss function, and the first loss function can include but is not limited to at least one of the following: mean square error loss function, mean absolute error loss function, cross-entropy loss function, Huber loss function.

[0087] 103B. Based on the first calculation method weight corresponding to each first loss calculation method, perform weighted arithmetic on the loss values obtained by each first loss calculation method to obtain the first loss value of the trained speech multi-modal interaction model.

[0088] Each first loss calculation method has a corresponding first calculation method weight, which is used to indicate the credibility of the first loss calculation method for calculating the semantic difference between the prompt text feature and the prompt audio feature. The first calculation method weight corresponding to the first loss calculation method can be preset in advance according to the historical credibility of the first loss calculation method for calculating the semantic difference between the prompt text feature and the prompt audio feature during historical model training.

[0089] After obtaining the corresponding loss value through each first loss calculation method, based on the first calculation method weight corresponding to each first loss calculation method, a weighted operation is performed on the loss value obtained by each first loss calculation method. The specific process of the weighted operation can be: respectively determine the product between the loss value corresponding to each first loss calculation method and the first calculation method weight, and add up the products corresponding to each first loss calculation method to obtain the first loss value of the trained speech multi-modal interaction model.

[0090] Second, for the second loss value.

[0091] The degree of difference between the predicted response text recognized by the speech multi-modal interaction model and the sample response text affects the ability of the speech multi-modal interaction model to accurately output a response adapted to the model input. Therefore, it is necessary to determine the second loss value of the trained speech multi-modal interaction model based on the predicted response text and the sample response text corresponding to each prompt text. The second loss value is used to reflect the degree of difference between the predicted response text recognized by the speech multi-modal interaction model and the true response text, that is, the "sample response file".

[0092] To implement the calculation of the second loss value, as Figure 2 shown, the speech multi-modal interaction model is provided with a second loss calculation method for calculating the difference between the predicted response text and the sample response text, and the second loss calculation method includes at least one calculation method. Based on this, the specific process of determining the second loss value of the trained speech multi-modal interaction model based on the predicted response text and the sample response text corresponding to each prompt text can include the following steps 103C to 103D:

[0093] 103C. Use each second loss calculation method to perform arithmetic processing on the predicted response text and the sample response text corresponding to each prompt text to obtain the loss value obtained by each second loss calculation method.

[0094] The second loss calculation method can select one or more calculation methods based on business needs. Considering that accidental errors may occur in a single second loss calculation method, resulting in inaccurate calculation of the second loss value, based on this, multiple second loss calculation methods are used to reduce the impact of accidental errors of individual second loss calculation methods on the correctness of the second loss value. The second loss calculation method can be flexibly selected based on business needs, and this embodiment does not limit this. Exemplarily, the second loss calculation method is deployed with a corresponding second loss function, and the second loss function may include but is not limited to at least one of the following: mean square error loss function, mean absolute error loss function, cross-entropy loss function, Huber loss function.

[0095] 103D. Based on the second calculation method weight corresponding to each second loss calculation method, perform a weighted operation on the loss value obtained by each second loss calculation method to obtain the second loss value of the trained voice multi-modal interaction model; the second calculation method weight is used to indicate the credibility of the second loss calculation method for the difference between the predicted response text and the sample response text.

[0096] Each second loss calculation method has a corresponding second calculation method weight, and the second calculation method weight is used to indicate the credibility of the second loss calculation method for calculating the difference between the predicted response text and the sample response text. The second calculation method weight corresponding to the second loss calculation method can be preset in advance according to the historical credibility of the second loss calculation method for calculating the difference between the predicted response text and the sample response text in historical model training.

[0097] After obtaining the corresponding loss value through each second loss calculation method, based on the second calculation method weight corresponding to each second loss calculation method, perform a weighted operation on the loss value obtained by each second loss calculation method. The specific process of the weighted operation can be: respectively determine the product between the loss value corresponding to each second loss calculation method and the second calculation method weight, and add up the products corresponding to each second loss calculation method to obtain the second loss value of the trained voice multi-modal interaction model.

[0098] 104. If it is determined that the trained voice multi-modal interaction model converges according to the first loss value and the second loss value, then determine the trained voice multi-modal interaction model as the trained voice multi-modal interaction model.

[0099] The first loss value and the second loss value are the basis for determining whether the trained speech multi-modal interaction model converges. When it is determined that the trained speech multi-modal interaction model converges, it indicates that the trained speech multi-modal interaction model has the ability to accurately output a response adapted to the model input, and the trained speech multi-modal interaction model can be put into use. When it is determined that the trained speech multi-modal interaction model does not converge, it indicates that the trained speech multi-modal interaction model does not have the ability to accurately output a response adapted to the model input, and the trained speech multi-modal interaction model still needs iterative training. Based on this, after determining the first loss value and the second loss value, it is necessary to judge whether the trained speech multi-modal interaction model converges according to the first loss value and the second loss value. The methods for judging whether the trained speech multi-modal interaction model converges according to the first loss value and the second loss value can at least include the following four:

[0100] Method 1. The specific process of judging whether the trained speech multi-modal interaction model converges according to the first loss value and the second loss value can include the following steps: Based on the first loss value and the second loss value, determine the total loss value; judge whether the total loss value is not greater than the first threshold; if it is not greater than, determine that the trained speech multi-modal interaction model converges; if it is greater, determine that the trained speech multi-modal interaction model does not converge.

[0101] Both the first loss value and the second loss value are the basis for determining whether the trained speech multi-modal interaction model converges. Therefore, the first loss value and the second loss value can be combined to determine whether the trained speech multi-modal interaction model converges. In one embodiment, the process of combining the first loss value and the second loss value is the process of determining the total loss value based on the first loss value and the second loss value. This process can at least include the following two: One is to obtain the loss weights corresponding to the first loss value and the second loss value respectively. The loss weights are used to indicate the influence degree of the corresponding loss value on the convergence of the speech multi-modal interaction model; based on the obtained loss weights, perform a weighted operation on the first loss value and the second loss value to obtain the total loss value. Among them, the loss weights can be preset based on the historical influence degree of the first loss value and the second loss value on the convergence of the speech multi-modal interaction model in historical model training. The other is to determine the sum of the first loss value and the second loss value as the total loss value.

[0102] The first threshold is a boundary value used to indicate whether the speech multi-modal interaction model converges. When it is determined that the total loss value is greater than the first threshold, it indicates that the trained speech multi-modal interaction model has a poor ability to accurately output a response adapted to the model input and still needs iterative training. Therefore, it is determined that the trained speech multi-modal interaction model does not converge. If it is determined that the total loss value is not greater than the first threshold, it indicates that the trained speech multi-modal interaction model has a better ability to accurately output a response adapted to the model input, and the trained speech multi-modal interaction model can be put into use. Therefore, it is determined that the trained speech multi-modal interaction model converges.

[0103] Method 2. The specific process of determining whether the trained speech multi-modal interaction model converges based on the first loss value and the second loss value may include the following steps: Based on the first loss value and the second loss value, determine the total loss value; judge whether the difference between the total loss value and the total loss value obtained from the previous training of the speech multi-modal interaction model is within the first threshold range; if not, determine that the trained speech multi-modal interaction model does not converge; if it is, judge whether the total loss values obtained from the previous first number of trainings of the speech multi-modal interaction model show a downward trend, and the difference between the total loss value of any adjacent two times, with the latter being compared to the former, is within the first threshold range. If so, determine that the trained speech multi-modal interaction model converges; if not, determine that the trained speech multi-modal interaction model does not converge.

[0104] For the specific process of determining the total loss value based on the first loss value and the second loss value, reference can be made to the above Method 1, which will not be elaborated here.

[0105] The first threshold range is the threshold range to which the total loss value belongs when the speech multi-modal interaction model converges. The first threshold range can be preset based on the historical total loss value when the model converges in historical model training. When it is determined that the difference between the total loss value and the total loss value obtained from the previous training of the speech multi-modal interaction model (if the current training is the first training, the total loss value of the previous training is defaulted to 0) is not within the first threshold range, it indicates that the trained speech multi-modal interaction model has a poor ability to accurately output a response adapted to the model input and still needs iterative training. Therefore, it is determined that the trained speech multi-modal interaction model does not converge. When it is determined that the difference between the total loss value and the total loss value obtained from the previous training of the speech multi-modal interaction model is within the first threshold range, it indicates that the trained speech multi-modal interaction model has the possibility of convergence. Therefore, it is necessary to continue to judge whether the total loss values obtained from the previous first number of trainings of the speech multi-modal interaction model show a downward trend, and the difference between the total loss value of any adjacent two times, with the latter being compared to the former, is within the first threshold range. The first number here can be preset in advance based on the empirical data of historical model training.

[0106] If it is determined that the total loss value obtained from training the speech multi-modal interaction model in the previous first number of adjacent times does not show a downward trend, or shows a downward trend but there is a situation where the difference between the total loss value of any adjacent two times and the total loss value of the previous time is not within the first threshold range, it indicates that the situation where the difference between the total loss value of the current training and the total loss value obtained from training the speech multi-modal interaction model in the previous time is within the first threshold range is only an accidental occurrence. The ability of the trained speech multi-modal interaction model to accurately output responses adapted to the user's speech is poor, and iterative training is still required. Therefore, it is determined that the trained speech multi-modal interaction model has not converged.

[0107] If it is determined that the total loss value obtained from training the speech multi-modal interaction model in the previous first number of adjacent times shows a downward trend, and the difference between the total loss value of any adjacent two times and the total loss value of the previous time is within the first threshold range, it indicates that the situation where the difference between the total loss value of the current training and the total loss value obtained from training the speech multi-modal interaction model in the previous time is within the first threshold range is not an accidental occurrence. The trained speech multi-modal interaction model has a better ability to accurately output responses adapted to the model input. The trained speech multi-modal interaction model can be put into use. Therefore, it is determined that the trained speech multi-modal interaction model has converged.

[0108] Method 3: The specific process of determining whether the trained speech multi-modal interaction model has converged based on the first loss value and the second loss value may include the following steps: Obtain the second thresholds corresponding to the first loss value and the second loss value respectively; Determine whether the first loss value and the second loss value are respectively not greater than their corresponding second thresholds; If so, determine that the trained speech multi-modal interaction model has converged; If not, determine that the trained speech multi-modal interaction model has not converged.

[0109] The first loss value is obtained by calculating the loss between the prompt text features and the prompt audio features obtained during the training process, and it can reflect the alignment of the trained speech multi-modal interaction model between the prompt audio features and the prompt text features. The second loss value is obtained by calculating the loss between the predicted response text and the sample response text obtained from training the speech multi-modal interaction model, and it can reflect the accuracy of the predicted response text of the trained speech multi-modal interaction model. Based on this, both the first loss value and the second loss value are the basis for determining whether the trained speech multi-modal interaction model has converged. Therefore, the first loss value and the second loss value can be used as separate factors to determine whether the trained speech multi-modal interaction model has converged.

[0110] The first loss value and the second loss value respectively have corresponding second thresholds. The second threshold corresponding to the first loss value is the boundary value at which the speech recognition model converges under the factors of the prompt text feature and the prompt audio feature. The second threshold corresponding to the second loss value is the boundary value at which the speech recognition model converges under the factors of the predicted response text and the sample response text.

[0111] If it is determined that the first loss value is greater than the corresponding second threshold, and / or the second loss value is greater than the corresponding second threshold, it indicates that the trained speech multi-modal interaction model is still not ideal in recognizing the prompt text feature and the prompt audio feature, and / or predicting the response text. The ability of the trained speech multi-modal interaction model to accurately output a response adapted to the user's speech is poor, and iterative training is still required. Therefore, it is determined that the trained speech multi-modal interaction model has not converged.

[0112] If it is determined that the first loss value is not greater than the corresponding second threshold and the second loss value is also not greater than the corresponding second threshold, it indicates that the trained speech multi-modal interaction model has reached an ideal state in recognizing the prompt text feature and the prompt audio feature, and in predicting the response text. The trained speech multi-modal interaction model has a better ability to accurately output a response adapted to the model input, and the trained speech multi-modal interaction model can be put into use. Therefore, it is determined that the trained speech multi-modal interaction model has converged.

[0113] Method 4, the specific process of determining whether the trained speech multi-modal interaction model has converged according to the first loss value and the second loss value may include the following steps: Determine whether the differences between the first loss value and the second loss value and the first loss value and the second loss value obtained from the previous training of the speech multi-modal interaction model are respectively within the corresponding second threshold ranges; if they are not respectively within, it is determined that the trained speech multi-modal interaction model has not converged; if they are respectively within, determine whether the first loss values and the second loss values obtained from the previous second number of trainings of the speech multi-modal interaction model both show a downward trend, and the differences between the latter first loss value and the previous first loss value and the differences between the latter second loss value and the previous second loss value in any two adjacent times are respectively within the corresponding second threshold ranges. If so, it is determined that the trained speech multi-modal interaction model has converged; if not, it is determined that the trained speech multi-modal interaction model has not converged.

[0114] The first loss value and the second loss value respectively have corresponding second threshold ranges. The second threshold range corresponding to the first loss value is the threshold range to which the first loss value belongs when the speech multi-modal interaction model converges. The second threshold range corresponding to the second loss value is the threshold range to which the second loss value belongs when the speech multi-modal interaction model converges. The second threshold range can be preset based on the corresponding historical loss values when the model converges in the historical model training.

[0115] If it is determined that the differences between the first loss value and the second loss value and the first loss value and the second loss value obtained from the previous training of the voice multi-modal interaction model are not respectively within the corresponding second threshold ranges, it indicates that the ability of the trained voice multi-modal interaction model to accurately output responses adapted to the user's voice is poor, and iterative training is still required. Therefore, it is determined that the trained voice multi-modal interaction model has not converged.

[0116] If it is determined that the differences between the first loss value and the second loss value and the first loss value and the second loss value obtained from the previous training of the voice multi-modal interaction model are respectively within the corresponding second threshold ranges, it indicates that the trained voice multi-modal interaction model has the possibility of convergence. Therefore, it is necessary to continue to determine whether the first loss value and the second loss value obtained from the previous second number of times of training the voice multi-modal interaction model both show a downward trend, and the differences between the first loss value of the latter time and the first loss value of the previous time and the second loss value of the latter time and the second loss value of the previous time in any two adjacent times are respectively within the corresponding second threshold ranges. The second number here can be set in advance based on the empirical data of historical model training. The second number corresponding to the first loss value and the second number corresponding to the second loss value can be different or the same, which is not limited in this embodiment.

[0117] If it is determined that at least one of the first loss value and the second loss value obtained from the previous second number of times of training the voice multi-modal interaction model does not show a downward trend, or the first loss value and the second loss value obtained from the previous second number of times of training the voice multi-modal interaction model both show a downward trend, and at least one of the differences between the first loss value of the latter time and the first loss value of the previous time and the second loss value of the latter time and the second loss value of the previous time in any two adjacent times is not within the corresponding second threshold range, it indicates that the situation where the differences between the first loss value and the second loss value and the first loss value and the second loss value obtained from the previous training of the voice multi-modal interaction model are respectively within the corresponding second threshold ranges in the current training is only an accidental case, and the ability of the trained voice multi-modal interaction model to accurately output responses adapted to the user's voice is poor, and iterative training is still required. Therefore, it is determined that the trained voice multi-modal interaction model has not converged.

[0118] If it is determined that both the first loss value and the second loss value obtained from the adjacent second-to-last training of the speech multi-modal interaction model show a downward trend, and the difference between the first loss value of the latter time and the first loss value of the previous time, as well as the difference between the second loss value of the latter time and the second loss value of the previous time, are respectively within the corresponding second threshold ranges among any two adjacent times, it indicates that the situation where the differences between the first loss value and the second loss value obtained from the current training and the first loss value and the second loss value obtained from the previous training of the speech multi-modal interaction model are respectively within the corresponding second threshold ranges is not accidental. It shows that the trained speech multi-modal interaction model has the ability to accurately output responses adapted to the model input, and the trained speech multi-modal interaction model can be put into use. Therefore, it is determined that the trained speech multi-modal interaction model converges.

[0119] At least one of the above four methods for determining whether the trained speech multi-modal interaction model converges based on the first loss value and the second loss value can be selected according to business requirements, and this embodiment does not limit this. When two or more methods are selected, the judgment results of various methods can be used for mutual verification. For example: when all the selected methods obtain the result that the trained speech multi-modal interaction model converges, it can be determined that the speech multi-modal interaction model converges; when any one of the selected methods obtains the result that the trained speech multi-modal interaction model does not converge, it is determined that the speech multi-modal interaction model does not converge.

[0120] After the method for determining whether the trained speech multi-modal interaction model converges based on the first loss value and the second loss value as described above, if it is determined that the trained speech multi-modal interaction model converges according to the first loss value and the second loss value, the trained speech multi-modal interaction model is determined as the trained speech multi-modal interaction model to put the speech multi-modal interaction model into use.

[0121] After the method for determining whether the trained speech multi-modal interaction model converges based on the first loss value and the second loss value as described above, if it is determined that the trained speech multi-modal interaction model does not converge according to the first loss value and the second loss value, the model parameters of the trained speech multi-modal interaction model are adjusted, and the adjusted speech multi-modal interaction model is used as the speech multi-modal interaction model to be trained, and the above step 102 is returned to input the training sample set into the speech multi-modal interaction model to be trained for model training, obtaining the prompt text features, predicted reply texts corresponding to each prompt text, and prompt audio features corresponding to the prompt audio, so as to continue the iterative training of the speech multi-modal interaction model.

[0122] If it is determined that the trained voice multi-modal interaction model has not converged based on the first loss value and the second loss value, it indicates that the trained voice multi-modal interaction model still has not converged. At this time, the model parameters of the trained voice multi-modal interaction model can be adjusted. The model parameters here are related to the ability of the voice multi-modal interaction model to accurately output responses adapted to the user's voice, and can include but are not limited to at least one of the following: model parameters and hyperparameters.

[0123] Adjust the model parameters of the trained voice multi-modal interaction model to determine at least one of the following change trends: the change trend between the first loss value and the first loss value of the voice multi-modal interaction model after the previous training, the change trend between the second loss value and the second loss value of the voice multi-modal interaction model after the previous training, and the change trend between the total loss value obtained based on the first loss value and the second loss value and the total loss value of the voice multi-modal interaction model after the previous training. For each model parameter, perform the following process respectively: search for the preset adjustment table, determine the adjustment value corresponding to the current model parameter and the determined change trend, and adjust the current model parameter based on the adjustment value. The preset adjustment table records multiple change trends and the adjustment values corresponding to the current model parameter under each change trend. The multiple change trends of the current model parameter recorded in the preset adjustment table and the corresponding adjustment values under each change trend are summarized based on the situations in the training processes of multiple successfully deployed voice multi-modal interaction models. Exemplarily, the preset adjustment table records change trend 1 (the total loss value increases) and change trend 2 (the total loss value decreases) corresponding to the model parameter "model weight". The adjustment value corresponding to change trend 1 is value 1 (value 1 is used to achieve the change of the total loss value of the voice multi-modal interaction model after adjusting the model weight towards the convergence of the voice multi-modal interaction model), and the adjustment value corresponding to change trend 2 is value 2 (value 2 is used to achieve the change of the total loss value of the voice multi-modal interaction model after adjusting the model weight towards the convergence of the voice multi-modal interaction model).

[0124] After the adjustment is completed, use the adjusted voice multi-modal interaction model as the voice multi-modal interaction model to be trained, and return to execute step 102 above to input the training sample set into the voice multi-modal interaction model to be trained for model training, obtain the prompt text features, predicted reply texts corresponding to each prompt text, and prompt audio features corresponding to the prompt audio, so as to re-iterate the training of the voice multi-modal interaction model until the voice multi-modal interaction model converges.

[0125] The training method of the voice multi-modal interaction model provided by the embodiments of the present application, when it is necessary to train the voice multi-modal interaction model, first obtains a training sample set. The training sample set includes multiple prompt texts, and the corresponding sample response texts and prompt audios for each prompt text, and inputs the training sample set into the voice multi-modal interaction model to be trained for model training, so as to obtain the prompt text features, predicted response texts, and prompt audio features corresponding to the prompt audios for each prompt text. Then, based on the prompt text features and prompt audio features corresponding to each prompt text, the first loss value of the trained voice multi-modal interaction model is determined, and based on the predicted response text and the sample response text corresponding to each prompt text, the second loss value of the trained voice multi-modal interaction model is determined. Finally, if it is determined that the trained voice multi-modal interaction model converges according to the first loss value and the second loss value, the trained voice multi-modal interaction model is determined as the trained voice multi-modal interaction model. It can be seen that the solution provided by the present application has at least the following invention points: First, when training the voice multi-modal interaction model, samples of two different modalities, text and audio, are used. In this way, when applying the trained voice multi-modal interaction model to respond to the user's input voice, the corresponding audio and text of the input voice can be combined to more comprehensively and accurately identify the semantics actually expressed in the user's voice. Second, during the training process of the voice multi-modal interaction model, through the loss calculation between the prompt text features and prompt audio features obtained during the training process, the voice multi-modal interaction model can enhance the semantic understanding of the prompt audio based on the prompt text features and retain the understanding of the unique features of the prompt audio by the voice multi-modal interaction model. That is, during the training process of the voice multi-modal interaction model, through loss calculation, the alignment of audio features and text features can be achieved, thereby strengthening the semantic information of audio features and retaining the unique features of audio. Then, based on the aligned audio features and text features, multi-modal features for training the voice multi-modal interaction model are obtained, so that the trained voice multi-modal interaction model can directly respond to the user's input voice, rather than having to convert the voice to text and then respond as in the background technology. Since the solution provided by this embodiment directly generates a response corresponding to the voice when responding to the user's voice, saving the process of "converting the voice to text and then responding to the converted text", the solution provided by this embodiment has high response efficiency; and because the solution provided by this embodiment aligns audio features and text features, strengthens the semantic information of audio features, and retains the unique features of audio during the training of the voice multi-modal interaction model, the solution provided by this embodiment can accurately understand the meaning expressed in the user's audio and has good response quality.Thirdly, during the training process of the voice multi-modal interaction model, by calculating the loss between the predicted response text obtained from training the voice multi-modal interaction model and the sample response text, the fitting degree of the voice multi-modal interaction model on the training data is continuously improved, enabling the voice multi-modal interaction model to accurately predict the response text. Based on the above three inventive points, it can be seen that the voice multi-modal interaction model trained according to the solution provided in this application can use the voice input by the user and / or the text corresponding to the voice as the model input. The voice multi-modal interaction model can accurately recognize the semantics actually expressed by the user based on the model input, and thus accurately output a response adapted to the input voice and / or input text, thereby improving the voice interaction effect with the user and bringing a good voice interaction experience to the user.

[0126] In some embodiments of this application, before training the voice multi-modal interaction model, the voice multi-modal interaction model can also be pre-trained, so as to perform model training on the basis of a voice multi-modal interaction model that initially has voice recognition capabilities, thereby improving the training efficiency of the voice multi-modal interaction model. Based on this, the training method of the voice multi-modal interaction model provided in this embodiment may further include the following steps: obtaining a pre-training sample set; inputting the pre-training sample set into the voice multi-modal interaction model to be trained for model pre-training; and obtaining the pre-trained voice multi-modal interaction model as the voice multi-modal interaction model that needs to be trained based on the training sample model.

[0127] The pre-training sample set includes multiple target prompt audios and the response text corresponding to each target prompt audio. The samples used as model inputs in the pre-training sample set include samples of only one modality, namely the target prompt audio, which can enable the voice multi-modal interaction model to quickly have initial voice recognition capabilities. In practical applications, the training sample set can be constructed from the Q&A voices corresponding to the business scenarios of the voice multi-modal interaction model to be applied. The data volume of the pre-training sample set can be determined based on business needs, and this embodiment does not limit this. Exemplarily, the target prompt audios in the pre-training sample set can reach approximately 700 hours of Chinese prompt audios and approximately 1300 hours of English prompt audios.

[0128] The process of pre-training the voice multi-modal interaction model can be as follows: performing iterative training for a target number of times on the initial voice multi-modal interaction model based on the pre-training sample set, and obtaining the pre-trained voice multi-modal interaction model as the voice multi-modal interaction model to be trained that needs to be trained based on the training sample model. Based on this pre-trained voice multi-modal interaction model, use the training sample set to adjust and train the voice multi-modal interaction model, so that the fine-tuned voice multi-modal interaction model can accurately recognize the intention actually expressed by the user by combining voice and text, and thus accurately output a response adapted to the voice.

[0129] Further, an embodiment of the present application also provides a multimodal processing method. As Figure 3 shown, the multimodal processing method provided in this embodiment may at least include the following steps 201 to 202:

[0130] 201. Obtain the information to be processed, where the information to be processed includes at least one modality content of audio and the text corresponding to the audio.

[0131] The speech multimodal model generated by the above speech multimodal interaction model training method can use the speech input by the user and / or the text corresponding to the speech as the model input. The speech multimodal interaction model can accurately recognize the semantics actually expressed by the user based on the model input and accurately output a response adapted to the input speech and / or input text. Based on this, after the speech multimodal interaction model is applied, at least one modality content of audio and the text corresponding to the audio can be obtained as the information to be processed according to specific business needs and input into the speech multimodal interaction model, so that the speech multimodal model can accurately recognize the semantics actually expressed by the user and accurately output a response adapted to the model input.

[0132] 202. Input the information to be processed into the speech multimodal interaction model to obtain the recognition result output by the speech multimodal interaction model.

[0133] The speech multimodal interaction model in this embodiment is generated by the speech multimodal interaction model training method. After inputting the information to be processed into the speech multimodal interaction model, the speech multimodal interaction model recognizes and processes according to the specific content included in the information to be processed, and finally obtains a recognition result adapted to the model input.

[0134] The information to be processed only includes the content of the audio modality. The speech multimodal interaction model first recognizes the initial audio features of the audio through an audio encoder, obtains audio features based on the initial audio features through a self-attention layer, and then obtains a response adapted to the audio included in the information to be processed based on the audio features through a prediction layer.

[0135] The information to be processed only includes the content of the text modality. The speech multimodal interaction model first recognizes the text features of the text through an embedding layer, and then obtains a response adapted to the audio included in the information to be processed based on the text features through a prediction layer.

[0136] The information to be processed includes the content of both the audio and text modalities. The speech multimodal interaction model recognizes the initial audio features of the audio through an audio encoder, obtains audio features based on the initial audio features through a self-attention layer, and at the same time, the speech multimodal interaction model first recognizes the text features of the text through an embedding layer. Then, the audio features and text features are concatenated to obtain multimodal features, and a response adapted to the audio and text included in the information to be processed is obtained based on the multimodal features through a prediction layer.

[0137] The multimodal processing method provided in this embodiment obtains the information to be processed including content of at least one modality among audio and text, and inputs the information to be processed into a speech multimodal interaction model to obtain the recognition result output by the speech multimodal interaction model. It can be seen that in the solution provided in this embodiment, the speech input by the user and / or the text corresponding to the speech can be used as the model input, and the speech multimodal interaction model can accurately recognize the semantics actually expressed by the user based on the model input, so as to accurately output a response adapted to the input speech and / or input text, and further improve the speech interaction effect with the user, bringing a good speech interaction experience to the user.

[0138] In some embodiments of the present application, hereinafter Figure 4 taking the application scenario architecture diagram of the speech multimodal interaction model shown as an example, the multimodal processing method provided in this embodiment will be specifically described. As Figure 4 shown, Figure 4 it includes a client and a server, and the speech multimodal interaction model is deployed in the server. When the client enables a voice interaction with the user, the client collects the voice input by the user, and then based on the specific business needs, the client can use at least one modality of the audio corresponding to the voice and the text corresponding to the audio as the information to be processed. After obtaining the information to be processed, the client sends the information to be processed to the server. After receiving the information to be processed sent by the client, the server uses the information to be processed as the model input and inputs it into the speech multimodal interaction model to obtain the recognition result output by the speech interaction model, and this recognition result is based on the reply text adapted to the voice input by the user. The server feeds back the recognition result to the client, and the client plays the corresponding response voice based on the recognition result to respond to the voice input by the user through the response voice, so as to realize the voice interaction between the client and the user.

[0139] In some embodiments of the present application, hereinafter a specific example will be used to illustrate the training of the speech multimodal interaction model and the application process after training. This embodiment specifically involves the following three processes: the model pre-training process, the model fine-tuning process, and the model application process.

[0140] The model pre-training process is to train a speech multi-modal interaction model with preliminary speech recognition ability. Based on the speech multi-modal interaction model with preliminary speech recognition ability, fine-tuning training is performed to obtain a speech multi-modal interaction model, so as to improve the training efficiency of the speech multi-modal interaction model. During the model pre-training process, first, a pre-training sample set needs to be obtained. The pre-training sample set includes multiple target prompt audios and the corresponding response text for each target prompt audio. Then, the pre-training sample set is input into the speech multi-modal interaction model to be trained for model pre-training, so as to pre-train a speech multi-modal interaction model with preliminary speech recognition ability. Finally, the pre-trained speech multi-modal interaction model is obtained as the speech multi-modal interaction model that needs to be trained based on the training sample model.

[0141] The model fine-tuning process is to train a speech multi-modal interaction model that can use the user-input speech and / or the text corresponding to the speech as model inputs, accurately recognize the semantics actually expressed by the user based on the model inputs, and accurately output a response adapted to the input speech and / or input text. The model micro-service process can refer to the Figure 2 process schematic diagram of the speech multi-modal interaction model training shown above, which will not be elaborated here.

[0142] The model application process is to apply the trained speech multi-modal interaction model to a business scenario to implement a voice interaction with the user through the speech multi-modal interaction model. During the model application process, the speech multi-modal interaction model can be deployed on a server. When a client with a communication connection to the server enables a voice interaction with the user, the client can provide at least one type of content from the user-input speech and the text corresponding to the audio collected by it to the server as the information to be processed. The server can use the information to be processed as a model input and input it into the speech multi-modal interaction model to obtain the recognition result output by the speech interaction model, and the recognition result is based on the response text adapted to the user-input speech. The server feeds back the recognition result to the client, and the client plays the corresponding response speech based on the recognition result to respond to the user-input speech through the response speech, thereby realizing the voice interaction between the client and the user.

[0143] Furthermore, an embodiment of the present application also provides a training device for a speech multi-modal interaction model, as Figure 5 shown. The training device for the speech multi-modal interaction model provided in this embodiment may include:

[0144] An acquisition module 31, configured to acquire a training sample set, where the training sample set includes multiple prompt texts, the corresponding prompt audio for each prompt text, and a sample response text, and the prompt audio is the audio expression corresponding to the prompt text;

[0145] A training module 32, configured to input the training sample set into a speech multi-modal interaction model to be trained for model training, so as to obtain a prompt text feature, a predicted response text, and a prompt audio feature corresponding to the prompt audio for each of the prompt texts;

[0146] A first determination module 33, configured to determine a first loss value of the trained speech multi-modal interaction model based on the prompt text feature and the prompt audio feature corresponding to each of the prompt texts, and determine a second loss value of the trained speech multi-modal interaction model based on the predicted response text and the sample response text corresponding to each of the prompt texts;

[0147] A second determination module 34, configured to, if a judgment module 35 determines that the trained speech multi-modal interaction model converges according to the first loss value and the second loss value, obtain the trained speech multi-modal interaction model as a trained speech multi-modal interaction model.

[0148] The training device for the voice multi-modal interaction model provided by the embodiments of the present application, when it is necessary to train the voice multi-modal interaction model, first obtains a training sample set. The training sample set includes multiple prompt texts, as well as the sample response texts and prompt audio corresponding to each prompt text, and inputs the training sample set into the voice multi-modal interaction model to be trained for model training, so as to obtain the prompt text features, predicted response texts, and prompt audio features corresponding to the prompt audio for each prompt text. Then, based on the prompt text features and prompt audio features corresponding to each prompt text, the first loss value of the trained voice multi-modal interaction model is determined, and based on the predicted response texts and sample response texts corresponding to each prompt text, the second loss value of the trained voice multi-modal interaction model is determined. Finally, if it is determined that the trained voice multi-modal interaction model converges according to the first loss value and the second loss value, the trained voice multi-modal interaction model is determined as the trained voice multi-modal interaction model. It can be seen that the solution provided by the present application has at least the following invention points: First, when training the voice multi-modal interaction model, samples of two different modalities, text and audio, are used. In this way, when applying the trained voice multi-modal interaction model to respond to the voice input by the user, the corresponding audio and text of the input voice can be combined to more comprehensively and accurately identify the semantics actually expressed in the user's voice. Second, during the training process of the voice multi-modal interaction model, through the loss calculation between the prompt text features and prompt audio features obtained during the training process, the voice multi-modal interaction model can enhance the semantic understanding of the prompt audio based on the prompt text features and retain the understanding of the unique features of the prompt audio by the voice multi-modal interaction model. That is, during the training process of the voice multi-modal interaction model, through loss calculation, the alignment of audio features and text features can be achieved, thereby strengthening the semantic information of audio features and retaining the unique features of audio. Then, based on the aligned audio features and text features, multi-modal features for training the voice multi-modal interaction model are obtained, so that the trained voice multi-modal interaction model can directly respond to the voice input by the user, rather than having to convert the voice into text and then respond as in the background technology. Since the solution provided by this embodiment directly generates a response corresponding to the voice when responding to the user's voice, saving the process of "converting the voice into text and then responding to the converted text", the solution provided by this embodiment has high response efficiency; and because the solution provided by this embodiment aligns audio features and text features, strengthens the semantic information of audio features, and retains the unique features of audio through loss calculation during the training of the voice multi-modal interaction model, the solution provided by this embodiment can accurately understand the meaning expressed in the user's audio and has good response quality.Thirdly, during the training process of the voice multi-modal interaction model, by calculating the loss between the predicted response text obtained from training the voice multi-modal interaction model and the sample response text, the fitting degree of the voice multi-modal interaction model on the training data is continuously improved, enabling the voice multi-modal interaction model to accurately predict the response text. Based on the above three inventive points, it can be seen that the voice multi-modal interaction model trained according to the solution provided in this application can use the voice input by the user and / or the text corresponding to the voice as the model input. The voice multi-modal interaction model can accurately recognize the semantics actually expressed by the user based on the model input, so as to accurately output a response adapted to the input voice and / or input text, thereby improving the voice interaction effect with the user and bringing a good voice interaction experience to the user.

[0149] In some embodiments of the present application, as Figure 6 shown, the training device of the voice multi-modal interaction model provided in this embodiment may further include: a third determination module 36, configured to, if the determination module 35 determines that the trained voice multi-modal interaction model has not converged according to the first loss value and the second loss value, adjust the model parameters of the trained voice multi-modal interaction model, and use the adjusted voice multi-modal interaction model as the voice multi-modal interaction model to be trained, and trigger the training module 32 to return to execute the step of inputting the training sample set into the voice multi-modal interaction model to be trained for model training to obtain the prompt text features, predicted response texts, and prompt audio features corresponding to each prompt text.

[0150] In some embodiments of the present application, as Figure 6 shown, the determination module 35 may include: a first determination unit 351, configured to perform an addition process on the first loss value and the second loss value to obtain a total loss value; determine whether the total loss value is not greater than a first threshold; if not, determine that the trained voice multi-modal interaction model has converged; if greater, determine that the trained voice multi-modal interaction model has not converged.

[0151] In some embodiments of the present application, as Figure 6 shown, the first determination unit 351 is specifically configured to obtain the loss weights corresponding to the first loss value and the second loss value respectively, where the loss weights are used to indicate the influence degree of the corresponding loss value on the convergence of the voice multi-modal interaction model; based on the obtained loss weights, perform a weighted operation on the first loss value and the second loss value to obtain a total loss value; or, determine the sum of the first loss value and the second loss value as the total loss value.

[0152] In some embodiments of the present application, as Figure 6As shown, the determination module 35 may include: a second determination unit 352, configured to determine a total loss value based on the first loss value and the second loss value; determine whether a difference between the total loss value and a total loss value obtained from the previous training of the speech multi-modal interaction model is within a first threshold range; if not, determine that the trained speech multi-modal interaction model has not converged; if so, determine whether the total loss values obtained from the previous first number of times of training the speech multi-modal interaction model show a downward trend, and the difference between the total loss value of any adjacent two times in the latter and the total loss value of the previous time is within the first threshold range. If so, determine that the trained speech multi-modal interaction model has converged; if not, determine that the trained speech multi-modal interaction model has not converged.

[0153] In some embodiments of the present application, as Figure 6 shown, the second determination unit 352 is specifically configured to obtain loss weights corresponding to the first loss value and the second loss value respectively, where the loss weights are used to indicate the influence degree of the corresponding loss value on the convergence of the speech multi-modal interaction model; based on the obtained loss weights, perform a weighted operation on the first loss value and the second loss value to obtain a total loss value; or, determine the sum of the first loss value and the second loss value as the total loss value.

[0154] In some embodiments of the present application, as Figure 6 shown, the determination module 35 may include: a third determination unit 353, configured to obtain second thresholds corresponding to the first loss value and the second loss value respectively; determine whether the first loss value and the second loss value are respectively not greater than the corresponding second thresholds; if so, determine that the trained speech multi-modal interaction model has converged; if not, determine that the trained speech multi-modal interaction model has not converged.

[0155] In some embodiments of the present application, as Figure 6 shown, the determination module 35 may include: a fourth determination unit 354, configured to determine whether differences between the first loss value and the second loss value and the first loss value and the second loss value obtained from the previous training of the speech multi-modal interaction model are respectively within corresponding second threshold ranges; if not respectively within, determine that the trained speech multi-modal interaction model has not converged; if respectively within, determine whether the first loss value and the second loss value obtained from the previous second number of times of training the speech multi-modal interaction model both show a downward trend, and the difference between the first loss value of any adjacent two times in the latter and the first loss value of the previous time and the difference between the second loss value of the latter and the second loss value of the previous time are respectively within the corresponding second threshold ranges. If so, determine that the trained speech multi-modal interaction model has converged; if not, determine that the trained speech multi-modal interaction model has not converged.

[0156] In some embodiments of the present application, the speech multi-modal interaction model includes an audio encoder and a large language model including at least a self-attention layer and an embedding layer. Then, as Figure 6 shown, the training module 32 may include:

[0157] A first processing unit 321, configured to process the input prompt audio through the audio encoder to obtain initial audio features, and process the input prompt text and sample response text through the embedding layer to obtain prompt text features and sample response text features;

[0158] A second processing unit 322, configured to semantically align the initial audio features with the prompt text features of the corresponding prompt text through the self-attention layer, and perform dimensionality reduction processing on the aligned initial audio features to obtain prompt audio features corresponding to each prompt text;

[0159] A splicing unit 323, configured to splice the prompt audio features corresponding to each prompt text and the prompt text features to respectively obtain multi-modal features corresponding to each prompt text;

[0160] A training unit 324, configured to train the large language model through the multi-modal features corresponding to each prompt text and the sample response text features to obtain predicted response texts corresponding to each prompt text.

[0161] In some embodiments of the present application, the speech multi-modal interaction model is provided with a first loss calculation method for calculating the semantic difference between the prompt text features and the prompt audio features, and the first loss calculation method includes at least one calculation method. Then, as Figure 6 shown, the first determination module 33 may include:

[0162] A first determination unit 331, configured to respectively perform arithmetic processing on the prompt text features and the prompt audio features corresponding to each prompt text by using each first loss calculation method to obtain loss values obtained by each first loss calculation method; based on the first calculation method weights corresponding to each first loss calculation method, perform weighted arithmetic on the loss values obtained by each first loss calculation method to obtain a first loss value of the trained speech multi-modal interaction model; the first calculation method weights are used to indicate the credibility of the first loss calculation method for calculating the semantic difference between the prompt text features and the prompt audio features.

[0163] In some embodiments of the present application, the speech multi-modal interaction model is provided with a second loss calculation method for predicting the difference between the response text and the sample response text, and the second loss calculation method includes at least one calculation method. Then, as Figure 6 shown, the second determination module 33 may include:

[0164] A second determination unit 332 is configured to perform arithmetic processing on the predicted response text and the sample response text corresponding to each of the prompt texts respectively by using each second loss calculation method, so as to obtain a loss value obtained by the operation of each second loss calculation method; based on the second calculation method weights corresponding to each second loss calculation method, perform weighted arithmetic on the loss values obtained by each second loss calculation method to obtain a second loss value of the trained speech multi-modal interaction model; the second calculation method weights are used to indicate the credibility of the second loss calculation method for the difference between the predicted response text and the sample response text.

[0165] In some embodiments of the present application, as Figure 6 shown, the acquisition module 31 is specifically configured to extract, from the user interaction data corresponding to the business scenario of the speech multi-modal interaction model to be applied, the prompt text and the sample response text corresponding to the prompt text for constructing the training sample set; set corresponding emotion features and tone features for each of the prompt texts; perform audio conversion on the prompt texts based on the emotion features and tone features corresponding to the prompt texts to obtain a prompt audio corresponding to each of the prompt texts.

[0166] In some embodiments of the present application, as Figure 6 shown, the training device for the speech multi-modal interaction model provided in this embodiment may further include:

[0167] A pre-training module 37 is configured to obtain a pre-training sample set, where the pre-training sample set includes a plurality of target prompt audios and the response text corresponding to each target prompt audio; input the pre-training sample set into the initial speech multi-modal interaction model for model pre-training, and the pre-trained speech multi-modal interaction model is the speech multi-modal interaction model to be trained.

[0168] In some embodiments of the present application, as Figure 6 shown, the training sample set obtained by the acquisition module 31 at least includes prompt texts that exist repeatedly; the prompt audios corresponding to the repeatedly existing prompt texts are based on the same prompt text and are obtained by audio conversion based on different target feature audios, where the target features include at least one of the following: emotion features and tone features.

[0169] In some embodiments of the present application, as Figure 6 shown, the training sample set obtained by the acquisition module 31 includes prompt texts in at least two languages, and the languages of the prompt audios and the sample response texts are consistent with the languages of the corresponding prompt texts.

[0170] In the training device for the speech multi-modal interaction model provided in the embodiments of the present application, the detailed explanations adopted during the operation of each functional module may refer to the corresponding detailed explanations in the embodiments of the training method for the speech multi-modal interaction model described above, and will not be elaborated herein.

[0171] Further, an embodiment of the present application further provides a multimodal processing device. As Figure 7 shown, the multimodal processing device provided in this embodiment may include:

[0172] An acquisition module 41, configured to acquire information to be processed, where the information to be processed includes content of at least one modality in audio and text corresponding to the audio;

[0173] A processing module 42, configured to input the information to be processed into a voice multimodal interaction model, and obtain an identification result output by the voice multimodal interaction model, where the voice multimodal interaction model is generated by the above-mentioned training method of the voice multimodal interaction model.

[0174] For the multimodal processing device provided in this embodiment, the information to be processed including content of at least one modality in audio and text is acquired, and the information to be processed is input into the voice multimodal interaction model to obtain the identification result output by the voice multimodal interaction model. It can be seen that in the solution provided in this embodiment, the voice input by the user and / or the text corresponding to the voice can be used as the model input, and the voice multimodal interaction model can accurately identify the semantics actually expressed by the user based on the model input, so as to accurately output a response adapted to the input voice and / or input text, thereby improving the voice interaction effect with the user and bringing a good voice interaction experience to the user.

[0175] In the multimodal processing device provided in the embodiment of the present application, the detailed explanations adopted during the operation of each functional module can refer to the corresponding explanations in the above-mentioned multimodal processing method embodiment, which will not be elaborated here.

[0176] Further, an embodiment of the present application further provides a computer-readable storage medium. The storage medium includes a stored program. When the program runs, it controls the device where the storage medium is located to execute the above-mentioned training method of the voice multimodal interaction model, and / or, the above-mentioned multimodal processing method.

[0177] Further, an embodiment of the present application further provides an electronic device. The electronic device includes: a memory for storing a program; a processor coupled to the memory for running the program to execute the above-mentioned training method of the voice multimodal interaction model, and / or, the above-mentioned multimodal processing method.

[0178] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0179] It can be understood that the relevant features in the above methods and devices can be referred to each other. Additionally, the "first", "second", etc. in the above embodiments are used to distinguish the embodiments, and do not represent the superiority or inferiority of each embodiment.

[0180] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0181] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. Based on the above description, the structure required to construct such systems is obvious. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the descriptions of specific languages above are for disclosing the preferred embodiments of the present application.

[0182] In addition, the memory may include non-permanent memory in computer-readable media, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one storage chip.

[0183] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0184] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data splicing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data splicing devices generate means for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0185] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction apparatus that implements the functions specified in one or more of the processes Figure 1 a process or processes and / or boxes Figure 1 a box or boxes.

[0186] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the processes Figure 1 a process or processes and / or boxes Figure 1 a box or boxes.

[0187] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0188] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0189] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0190] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0191] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0192] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for training a speech multimodal interaction model, characterized in that: The method comprises: Acquire a training sample set, wherein the training sample set includes a plurality of prompt texts and a prompt audio and a sample reply text corresponding to each prompt text, wherein the prompt audio is an audio expression corresponding to the prompt text; Inputting the training sample set into the speech multimodal interaction model to be trained for model training, and obtaining prompt text features corresponding to each prompt text, predicted reply text, and prompt audio features corresponding to the prompt audio; Determine a first loss value of the trained speech multimodal interaction model based on the prompt text features and the prompt audio features corresponding to each of the prompt texts, and determine a second loss value of the trained speech multimodal interaction model based on the predicted reply text and the sample reply text corresponding to each of the prompt texts; If it is determined that the trained speech multimodal interaction model converges according to the first loss value and the second loss value, the trained speech multimodal interaction model is determined as the trained speech multimodal interaction model.

2. The method according to claim 1, characterized in that The method further comprises: If it is determined according to the first loss value and the second loss value that the trained speech multimodal interaction model has not converged, the model parameters of the trained speech multimodal interaction model are adjusted, and the adjusted speech multimodal interaction model is used as the speech multimodal interaction model to be trained, and the step of inputting the training sample set into the speech multimodal interaction model to be trained for model training is returned to obtain the prompt text features corresponding to each of the prompt texts, the predicted reply text, and the prompt audio features corresponding to the prompt audio.

3. The method according to claim 1, characterized in that The method further comprises: Based on the first loss value and the second loss value, determine the total loss value; determine whether the total loss value is not greater than a first threshold; if not, determine that the trained speech multimodal interaction model has converged; if greater, determine that the trained speech multimodal interaction model has not converged; and / or, Based on the first loss value and the second loss value, determine the total loss value; determine whether the difference between the total loss value and the total loss value obtained by the previous training of the speech multimodal interaction model is within the first threshold range; if not, determine that the trained speech multimodal interaction model has not converged; if so, determine whether the total loss values ​​obtained by the first number of adjacent trainings of the speech multimodal interaction model show a downward trend, and the difference between the latter total loss value and the previous total loss value in any two adjacent times is within the first threshold range. If so, determine that the trained speech multimodal interaction model converges, if not, determine that the trained speech multimodal interaction model has not converged; and / or, Obtaining second thresholds corresponding to the first loss value and the second loss value respectively; determining whether the first loss value and the second loss value are not greater than the corresponding second thresholds respectively; if so, determining that the trained speech multimodal interaction model converges; if not, determining that the trained speech multimodal interaction model does not converge; and / or, Determine whether the differences between the first loss value and the second loss value and the first loss value and the second loss value obtained from the previous training of the speech multimodal interaction model are respectively within the corresponding second threshold range; if not, it is determined that the trained speech multimodal interaction model has not converged; if respectively, it is determined whether the first loss value and the second loss value obtained from the first second number of adjacent trainings of the speech multimodal interaction model both show a downward trend, and the difference between the second loss value of the latter and the first loss value of the former in any two adjacent times and the difference between the second loss value of the latter and the second loss value of the former are respectively within the corresponding second threshold range. If so, it is determined that the trained speech multimodal interaction model converges; if not, it is determined that the trained speech multimodal interaction model has not converged.

4. The method according to claim 3, characterized in that Determining a total loss value based on the first loss value and the second loss value includes: Obtain loss weights corresponding to the first loss value and the second loss value, respectively, where the loss weights are used to indicate the degree of influence of the corresponding loss values ​​on the convergence of the speech multimodal interaction model; based on the obtained loss weights, perform a weighted operation on the first loss value and the second loss value to obtain a total loss value; Or, the sum of the first loss value and the second loss value is determined as the total loss value.

5. The method according to claim 1, characterized in that The speech multimodal interaction model includes an audio encoder and a large language model including at least a self-attention layer and an embedding layer. Then, the training sample set is input into the speech multimodal interaction model to be trained for model training, including: Processing the input prompt audio through the audio encoder to obtain initial audio features, and processing the input prompt text and sample reply text through the embedding layer to obtain prompt text features and sample reply text features; Semantically aligning the initial audio features with the prompt text features of the corresponding prompt text through the self-attention layer, and performing dimensionality reduction processing on the aligned initial audio features to obtain prompt audio features corresponding to each prompt text; Concatenate the prompt audio features and prompt text features corresponding to each prompt text to obtain multimodal features corresponding to each prompt text; The large language model is trained using the multimodal features corresponding to each prompt text and the sample response text features to obtain the predicted response text corresponding to each prompt text.

6. The method according to any one of claims 1 to 5, characterized in that The speech multimodal interaction model is provided with a first loss calculation method for calculating the semantic difference between the prompt text feature and the prompt audio feature, and the first loss calculation method includes at least one calculation method. Then, based on the prompt text feature and the prompt audio feature corresponding to each of the prompt texts, the first loss value of the trained speech multimodal interaction model is determined, including: using each first loss calculation method to perform calculation processing on the prompt text feature and the prompt audio feature corresponding to each of the prompt texts, respectively, to obtain the loss value calculated by each first loss calculation method; based on the first calculation method weight corresponding to each first loss calculation method, the loss value obtained by each first loss calculation method is weighted to obtain the first loss value of the trained speech multimodal interaction model; the first calculation method weight is used to indicate the credibility of the first loss calculation method for calculating the semantic difference between the prompt text feature and the prompt audio feature; and / or, The speech multimodal interaction model is provided with a second loss calculation method for predicting the difference between a reply text and a sample reply text, and the second loss calculation method includes at least one calculation method. Then, based on the predicted reply text and the sample reply text corresponding to each of the prompt texts, the second loss value of the trained speech multimodal interaction model is determined, including: using each second loss calculation method to perform calculations on the predicted reply text and the sample reply text corresponding to each of the prompt texts, respectively, to obtain the loss value calculated by each second loss calculation method; based on the second calculation method weight corresponding to each second loss calculation method, weighted calculation is performed on the loss value obtained by each second loss calculation method to obtain the second loss value of the trained speech multimodal interaction model; the second calculation method weight is used to indicate the credibility of the second loss calculation method for the difference between the predicted reply text and the sample reply text.

7. The method according to any one of claims 1 to 5, characterized in that Acquiring a training sample set includes: extracting prompt texts and sample reply texts corresponding to the prompt texts for constructing the training sample set from user interaction data corresponding to the business scenario of the voice multimodal interaction model to be applied; setting corresponding emotional features and timbre features for each of the prompt texts; performing audio conversion on the prompt texts based on the emotional features and timbre features corresponding to the prompt texts to obtain prompt audios corresponding to each of the prompt texts; and / or, The method further includes: obtaining a pre-training sample set, the pre-training sample set including a plurality of target prompt audios and a reply text corresponding to each target prompt audio; inputting the pre-training sample set into an initial speech multimodal interaction model for model pre-training, and the speech multimodal interaction model obtained by pre-training is the speech multimodal interaction model to be trained; and / or, The training sample set at least includes repeated prompt texts; the prompt audio corresponding to the repeated prompt texts is obtained by converting the same prompt texts and different target feature audios, wherein the target feature includes at least one of the following: emotional features and timbre features; and / or, The training sample set includes prompt texts in at least two languages, and the languages ​​of the prompt audio and sample reply text are consistent with the languages ​​of the corresponding prompt texts.

8. A multimodal processing method, characterized in that: The method comprises: Acquire information to be processed, where the information to be processed includes content in at least one mode of audio and text corresponding to the audio; The information to be processed is input into a speech multimodal interaction model to obtain a recognition result output by the speech multimodal interaction model, wherein the speech multimodal interaction model is generated by the training method for a speech multimodal interaction model described in any one of claims 1 to 7.

9. A training device for a speech multimodal interaction model, characterized in that: The device comprises: An acquisition module, used to acquire a training sample set, wherein the training sample set includes a plurality of prompt texts and a prompt audio and a sample reply text corresponding to each prompt text, wherein the prompt audio is an audio expression corresponding to the prompt text; A training module, used for inputting the training sample set into the speech multimodal interaction model to be trained for model training, and obtaining prompt text features corresponding to each prompt text, predicted reply text, and prompt audio features corresponding to the prompt audio; A first determination module is used to determine a first loss value of the trained speech multimodal interaction model based on a prompt text feature and a prompt audio feature corresponding to each of the prompt texts, and to determine a second loss value of the trained speech multimodal interaction model based on a predicted reply text and a sample reply text corresponding to each of the prompt texts; The second determination module is used to obtain the trained speech multimodal interaction model as the trained speech multimodal interaction model if the judgment module determines that the trained speech multimodal interaction model converges according to the first loss value and the second loss value.

10. A computer-readable storage medium, characterized in that: The storage medium includes a stored program, wherein, when the program is running, the device where the storage medium is located is controlled to execute the training method of the speech multimodal interaction model described in any one of claims 1 to 7, and / or the multimodal processing method described in claim 8.

11. An electronic device, characterized in that: The electronic device comprises: Memory, used to store programs; A processor, coupled to the memory, is used to run the program to execute the training method of the speech multimodal interaction model described in any one of claims 1 to 7, and / or the multimodal processing method described in claim 8.

Citation Information

Cited By

  • Method and apparatus for training multi-modal speech interaction model

    WO2026171149A1