Voice interaction method and device
Multimodal information is collected through voice interaction devices and voice style prompt information is automatically generated using preset style generation models, which solves the problem of repeatedly specifying the generation style when users switch in the prior art, and improves user experience and efficiency.
Patent Information
- Application Number
- CN202510573038.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
The existing voice interaction system needs to repeatedly specify the generation style when switching users, resulting in cumbersome user experience, especially inefficient in multi-person scenarios.
Multimodal information is collected through voice interaction devices, and a preset style generation model is used to automatically generate voice style prompt information matching to the user, including a Transformer architecture and deep learning model based on multimodal information, to generate voice style prompt information to simplify the user interaction process.
It simplifies voice interaction when users switch users, improves user experience and efficiency, and realizes automatic matching and natural processes with users.
Smart Images

Figure CN120496499A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice interaction technology, and in particular to a voice interaction method and device. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, the application of intelligent assistants on personal computers (PCs) is becoming increasingly widespread. Voice plays a key role in intelligent assistant interaction systems, greatly improving the user experience. Text-to-speech (TTS) in common voice interaction systems can generate various styles of voices, accents, laughter, and other sounds based on user needs. However, this requires users to set the relevant style information in advance. In the actual AIPC end-side experience, switching between different users requires repeated specification of different generation styles. This is especially true in multi-person scenarios where user switching is more frequent. Repeated specification of generation styles makes voice interaction operations cumbersome and inefficient. Summary of the Invention
[0003] In view of this, the present application provides a voice interaction method and device.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present invention provides a method for voice interaction, which includes:
[0006] In response to the collected voice to be replied corresponding to the first object, obtaining multimodal information corresponding to the first object;
[0007] Generate voice style prompt information corresponding to the first object based on the multimodal information using a preset style generation model;
[0008] Generate a reply voice corresponding to the voice to be replied based on the voice style prompt information.
[0009] The embodiment of the present application provides a voice interaction device, including: a sensor and a processor;
[0010] A sensor, configured to obtain multimodal information corresponding to the first object in response to the collected speech to be replied corresponding to the first object;
[0011] The processor is configured to generate voice style prompt information corresponding to the first object based on multimodal information using a preset style generation model; and to generate a reply voice corresponding to the voice to be replied based on the voice style prompt information.
[0012] It should be understood that the above general description and the detailed description below are merely exemplary and explanatory, and do not limit the technical solutions provided in the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which:
[0014] Figure 1 A flowchart of a voice interaction method provided in the related art;
[0015] Figure 2 A flowchart of an exemplary voice interaction method provided in an embodiment of the present application;
[0016] Figure 3 A schematic diagram of an exemplary process for generating voice style prompt information provided in an embodiment of the present application Figure 1 ;
[0017] Figure 4 A schematic diagram of an exemplary process for generating voice style prompt information provided in an embodiment of the present application Figure 2 ;
[0018] Figure 5 A schematic diagram of an exemplary process for generating voice style prompt information provided in an embodiment of the present application Figure 3 ;
[0019] Figure 6 An exemplary process for generating user tags provided in an embodiment of the present application is shown in FIG. Figure 1 ;
[0020] Figure 7 A flowchart of an exemplary feature extraction method provided in an embodiment of the present application;
[0021] Figure 8 An exemplary process for generating user tags provided in an embodiment of the present application is shown in FIG. Figure 2 ;
[0022] Figure 9 A schematic diagram of a network structure for generating user tags according to an embodiment of the present application;
[0023] Figure 10 A schematic diagram of an exemplary process for generating a reply voice provided in an embodiment of the present application Figure 1 ;
[0024] Figure 11 A schematic diagram of a network structure for generating user tags according to an embodiment of the present application;
[0025] Figure 12 A schematic diagram of an exemplary process for generating a reply voice provided in an embodiment of the present application Figure 2 ;
[0026] Figure 13 A schematic diagram of an exemplary process for matching voice style prompt information provided in an embodiment of the present application;
[0027] Figure 14 A flowchart of an exemplary model training method provided in an embodiment of the present application;
[0028] Figure 15 A flowchart of an exemplary voice interaction method provided in an embodiment of the present application;
[0029] Figure 16 A schematic diagram of the structure of a voice interaction device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It should be understood that the specific embodiments described herein are only used to explain the related applications and are not intended to limit the applications. It should also be noted that for ease of description, only the parts relevant to the related applications are shown in the drawings.
[0031] In related technologies, commonly used voice interaction systems are implemented through the paradigm of speech recognition (Automatic Speech Recognition, ASR), large language model (Large Language Model, LLM), and TTS. The exemplary paradigm process is as follows: Figure 1 As shown, ASR10 recognizes the input speech (Speech input) 11 as transcription text (transcription text) 12, LLM13 generates a response text (prompt text) 14 based on the transcription text 12, and TTS15 combines the instruction prompt (instruct text) 16 and the reference timbre (speech prompt) 17 to generate a speech output (Speech out) 18 from the response text 14.
[0032] In terms of TTS, related technologies can generate various styles of voices, including accents, laughter, and other information based on user needs. This is achieved by inserting tokens at different locations in the LLM's output reply text (prompt text). The instruction prompt (instruct text) specifies the dialect, timbre, tone, dialect, and other style requirements specified by the user, and acoustic feature information such as reference timbre is provided through the reference timbre (speech prompt). However, the related method of controlling the TTS generation style requires the user to set the relevant instruction prompt (instruct text) in advance. In the actual AIPC end-side experience, different generation styles need to be repeatedly specified when switching between different users. Especially in home use scenarios, users switch more frequently, and repeatedly specifying the generation style greatly affects the user experience. There is currently no solution to this problem.
[0033] The present invention provides a method for voice interaction, which is implemented by a voice interaction device. Figure 2 As shown, the following steps S201 to S203 are included:
[0034] Step S201: In response to the collected voice to be replied corresponding to the first object, obtain multimodal information corresponding to the first object.
[0035] In the embodiments of the present application, the voice interaction device is an electronic device with voice interaction function, which can be a tablet computer, a laptop computer, a handheld computer, a personal digital assistant (PDA), a desktop computer, etc. The specific voice interaction device is not limited here.
[0036] In an embodiment of the present application, the voice interaction device may collect the voice to be replied corresponding to the first object through a sound collection sensor. Exemplarily, the sound collection sensor may be a microphone or a microphone array.
[0037] In an embodiment of the present application, if the voice interaction device collects the voice to be replied corresponding to the first object, the voice interaction device can collect multimodal information such as image information, identification information, voice information, distance, etc. corresponding to the first object through various sensors.
[0038] Exemplarily, the voice interaction device may collect multimodal information corresponding to the first object through a microphone, a camera, an infrared sensor, a radar, a text input, or a thermal imaging sensor.
[0039] In an embodiment of the present application, the first object may be any object within the range where the voice interaction device can receive the voice, or an object within a preset distance of the voice interaction device, or an object using the voice interaction device.
[0040] Step S202: Generate voice style prompt information corresponding to the first object based on the multimodal information using a preset style generation model.
[0041] In an embodiment of the present application, after obtaining the multimodal information of the first object, the voice interaction device uses a preset style generation model to generate voice style prompt information corresponding to the first object based on the multimodal information.
[0042] In an embodiment of the present application, the preset style generation model can be a generative pre-trained transformer (GPT) model based on a decoder-only transformer architecture. Of course, the preset style generation model can also be other deep learning generation models, such as autoregressive models, deep learning models based on self-attention mechanisms, or other models that can generate speech style cues based on multimodal information.
[0043] In the embodiments of this application, the preset style generation model is a model that can automatically generate matching voice style prompt information based on multimodal information. For example, the voice style prompt information can be "Please speak this sentence using XX tone, XX dialect, and XX speed," or "Please speak this sentence using XX timbre, XX accent, and XX tone," etc. The information involved in the voice style prompt information can be set according to actual needs and application scenarios, and this application does not impose any restrictions on this.
[0044] For example, if the reply voice of the first subject uses accent A, the first subject is currently in a very happy mood, and the first subject is between 60 and 70 years old, then the corresponding voice style prompt information can be "Please speak this sentence in a happy tone and accent A, and speak slower for clarity."
[0045] For example, if the first subject's reply voice uses a B accent, the first subject's current tone is very cute, and the age is between 3 and 5 years old, then the corresponding voice style prompt information can be "Please say this sentence with an S-type timbre, a B accent, and a cute tone."
[0046] Step S203: Generate a reply voice corresponding to the voice to be replied based on the voice style prompt information.
[0047] In an embodiment of the present application, after obtaining the voice style prompt information, the voice interaction device can generate a reply voice corresponding to the voice to be replied.
[0048] In an embodiment of the present application, the voice interaction device can perform voice recognition on the voice to be replied, generate a transcribed text, and then generate a corresponding reply text based on the transcribed text. Finally, the reply text is converted into a voice output in the voice style of the voice style prompt information.
[0049] Compared with the related art that requires repeated specification of the generation style when switching users, this application automatically generates voice style prompt information matching the user based on multimodal information, which can simplify the user's voice interaction process and improve the user's voice interaction experience.
[0050] In some embodiments, the preset style generation model includes a label generation sub-model and a style prompt information generation sub-model. When the voice interaction device performs the above step S202, Figure 3 As shown, the following steps S301 and S302 may also be performed:
[0051] Step S301: Generate a sub-model using the label to determine the user label corresponding to the first object based on multimodal information.
[0052] In an embodiment of the present application, the label generation sub-model can be an intelligent perception model. Using the intelligent perception model, based on multimodal information, tasks such as identity recognition, distance perception, dialect prediction, and age prediction can be performed on the first object, and corresponding user labels can be generated.
[0053] For example, user tags may be saved in a JSON format, such as {"user_id": 1, "emotion": happy, "dialect": Chinese accent, "distance": 1000, "age": 20-25}.
[0054] Step S302: Generate a sub-model using the style prompt information, and generate speech style prompt information corresponding to the first object based on the user tag.
[0055] In an embodiment of the present application, user tags are relatively discrete prompt tag words, and the voice interaction device will use style prompt information to generate a sub-model. Based on the relatively discrete prompt tag words, the text content of the output voice style prompt information will be richer, more diverse and more natural.
[0056] Exemplarily, the style hint information generation sub-model may be a language model (LM), for example, a language model based on a neural network: a recurrent neural network (RNN), a long short-term memory network (LSTM), a transformer, etc.
[0057] like Figure 4 As shown, an exemplary flow chart for generating voice style prompt information is provided, wherein a user tag 41 is input into a language model 42 to generate voice style prompt information 43 corresponding to a first object.
[0058] like Figure 5 As shown, a language model network structure is provided. Relatively discrete prompt tag words {keywords), namely user tags 41, are input into a language model 42 to generate voice style prompt information 43 corresponding to the first object. Language model 42 employs an autoregressive Transformer structure, where S represents the start point of a sequence, T represents a sequence transition point, and E represents the end point of a sequence. Crosshairs represent ignore tokens, and slashes represent text tokens.
[0059] In the embodiments of the present application, the voice style prompt information generated based on multimodal information, the style prompt sentences input by the user, and the LLM generated content in the voice interaction are all natural languages with high consistency, which facilitates the understanding and reasoning of the TTS model.
[0060] In some embodiments, the label generation sub-model includes: a feature extraction network and a label generation network; when the voice interaction device performs the above step S301, Figure 6 As shown, the following steps S601 and S602 may also be performed:
[0061] Step S601: Using a feature extraction network, perform feature extraction and modality alignment on each modal information in the multimodal information to obtain corresponding multiple feature information.
[0062] In an embodiment of the present application, the label generation sub-model includes a feature extraction network and a label generation network, and a voice interaction device performs feature extraction and modality alignment on each multimodal information to obtain corresponding multiple feature information.
[0063] In an embodiment of the present application, the voice interaction device will perform feature extraction on each multimodal information in the multimodal information. For example, the multimodal information may include text information, voice information, and image information. The feature extraction network can obtain multiple extracted feature information by performing feature extraction. For example, text information corresponds to text extraction feature information, voice information corresponds to voice extraction feature information, and image information corresponds to image extraction feature information. Taking into account the different dimensions of text extraction feature information, voice extraction feature information, and image extraction feature information, the multiple extracted feature information obtained will be modally aligned to obtain multiple feature information.
[0064] Step S602: Generate user tags based on multiple feature information using a tag generation network.
[0065] In an embodiment of the present application, the voice interaction device generates user tags based on multiple feature information using a tag generation network.
[0066] In this way, based on the feature extraction network and the label generation network, multimodal information can be used to generate a user label, which matches the first object. Then, the voice style prompt information determined based on the user label also matches the first object, which is more flexible.
[0067] In some embodiments, when the voice interaction device performs the above step S601, Figure 7 As shown, the following steps S701 to S703 may also be performed:
[0068] Step S701: extract features from each modal information in the multimodal information to obtain corresponding multiple extracted feature information.
[0069] For example, multimodal information may include text information, voice information, and image information. The text information, voice information, and image information are respectively represented as: Then, features are extracted through three feature extraction networks:
[0070] For example, the text information is extracted as shown in formula (1):
[0071] T f =Text_Encoder(T) (1);
[0072] Among them, T f is the extracted feature information corresponding to the text information, Text_Encoder is the text extraction network, T f The feature dimension is
[0073] For example, the method for extracting voice information is as follows:
[0074] A f =Audio_Encoder(A) (2);
[0075] Among them, A f is the extracted feature information corresponding to the voice information, Audio_Encoder is the voice extraction network, A f The feature dimension is
[0076] For example, the image information is extracted as shown in formula (3):
[0077] I f =Image_Encoder(I) (3);
[0078] Among them, I f is the extracted feature information corresponding to the image information, Image_Encoder is the image extraction network, I f The feature dimension is
[0079] Among them, N is mini_batch_size, which can be understood as the number of classification labels and users, and d is the feature dimension of each batch.
[0080] Based on the different dimensions of the text-extracted feature information, image-extracted feature information, and speech-extracted feature information discussed above, in order to achieve better feature fusion, these extracted feature information will be modally aligned: linearized and normalized.
[0081] Step S702: linearize each of the plurality of extracted feature information to obtain corresponding plurality of linearized feature information.
[0082] In an embodiment of the present application, the voice interaction device linearizes each of the multiple extracted feature information: text extracted feature information, image extracted feature information, and voice extracted feature information, to obtain corresponding multiple linearized feature information.
[0083] For example, the linearization formula is shown in formula (4):
[0084] e=Linear T (X) (4);
[0085] Among them, X can be text extraction feature information, image extraction feature information, and voice extraction feature information. Linear T is the linearization algorithm, and e is the linearization result.
[0086] Step S703: Normalize each piece of linearized feature information in the plurality of linearized feature information to obtain corresponding plurality of feature information.
[0087] In an embodiment of the present application, the voice interaction device normalizes each linearized feature information in the multiple linearized feature information: text linearized feature information, image linearized feature information, and voice linearized feature information, to obtain corresponding multiple normalized feature information.
[0088] For example, the implementation method of performing modal alignment on multiple extracted feature information can refer to the following formulas (5) to (7):
[0089] For example, the method of performing modal alignment on text extraction feature information is as follows:
[0090] T e =L2_norm(Linear T (T f ))=L2_norm(W T T f +b T ) (5);
[0091] Among them, T e is the text feature information, W T 、b T are model parameters.
[0092] For example, the method of performing modal alignment on speech extraction feature information is as follows:
[0093] A e =L2_norm(Linear A (A f ))=L2_norm(W A A f +b A ) (6);
[0094] Among them, A e is the speech feature information, W A 、b A are model parameters.
[0095] For example, the method of performing modal alignment on the feature information extracted from the image is as follows:
[0096] I e =L2_norm(Linear I (I f ))=L2_norm(W I I f +b I ) (7);
[0097] Among them, I e is the speech feature information, W I 、b I are model parameters.
[0098] In this way, the feature dimensions between the above multiple feature information are
[0099] In this way, a unified processing method between multiple modal information can be ensured, the fusion between multiple modal information can be facilitated, the relationship information between each information can be extracted, and the accuracy of the voice style prompt information determined based on multimodal information can be improved.
[0100] In some embodiments, the multimodal information includes identification information, and at least one of image information and voice information; the label generation network includes a feature fusion network layer and a label classification network layer; when the voice interaction device performs the above step S602, if Figure 8 As shown, the following steps S801 and S802 may also be performed:
[0101] Step S801: Using a feature fusion network layer, feature information corresponding to information other than identification information in a plurality of feature information is fused to obtain fused feature information.
[0102] In an embodiment of the present application, the voice interaction device utilizes a feature fusion network layer to fuse feature information corresponding to information other than identification information in a plurality of feature information to obtain fused feature information.
[0103] For example, the identification information exists in the form of text information. Using the feature fusion network layer, the fusion can be image feature information and voice feature information, or alternatively, one of the two. If a feature is not included, a preset value is entered for that feature. The preset value indicates that the feature is not currently included.
[0104] For example, among the multiple feature information, the feature information corresponding to the information other than the identification information may include image feature information and voice feature information. Then, the method of performing feature fusion (Feature Fusion) on the image feature information and the voice feature information is as follows:
[0105] F e =A e +I e (8);
[0106] Among them, F e To fuse feature information.
[0107] Step S802: Using the label classification network layer, perform matrix multiplication on the fused feature information and the feature information corresponding to the identification information in the plurality of feature information to obtain a user label.
[0108] In an embodiment of the present application, the voice interaction device utilizes a label classification network layer to perform matrix multiplication on the fused feature information and the feature information corresponding to the identification information to obtain a user label.
[0109] For example, the method for determining the user tag is as follows:
[0110]
[0111] Among them, C is the similarity matrix, The diagonal of the matrix is the category label (user label) corresponding to the multimodal information.
[0112] like Figure 9 As shown in FIG. 1 , an exemplary implementation method for generating user labels is provided. The voice interaction device inputs text information (identification information) 91 into a text feature extraction network (Text Encoder) 92 to obtain text feature information 93. The image information 94 is input into an image feature extraction network (Image Encoder) 95 to obtain image feature information 96. The voice information 97 is input into an audio feature extraction network (Audio Encoder) 98 to obtain voice feature information 99. The image feature information 96 and the audio feature extraction information 98 are then input into a feature fusion network layer 910 to obtain fused feature information 911. Finally, the fused feature information 911 and the text feature information 93 are input into a label classification network layer 912 to obtain a user label 913. In the figure, the diagonal lines (grey squares) of the matrix represent positive samples, i.e., user labels.
[0113] In some embodiments, as Figure 10 As shown, the voice interaction device may further perform the following steps S1001 to S1003:
[0114] Step S1001: perform identity recognition on a first object to determine identification information corresponding to the first object.
[0115] In an embodiment of the present application, the voice interaction device performs identity recognition on the first object and determines the identification information corresponding to the first object. At this time, the voice interaction device can obtain the voice to be replied corresponding to the first object, and then perform identity recognition based on the voice to be replied.
[0116] like Figure 11As shown, the voice interaction device is a Clip-Clap Net (CCNet) network structure for image-text (Clip) and voice-text (Clap), comprising a preprocessing network layer 111, a feature extraction network layer 112, and a classifier 113. Preprocessing network layer 111 includes resampling 1111, frame windowing 1112, and MFCC features (Mel Frequency Cepstral Coefficients) 1113; feature extraction network layer 112 includes one-dimensional convolution 1121, depthwise separable convolution 1122, and a transformer encoder 1123; and classifier 103 includes a similarity matrix 1131, a contrastive loss 1132, and a normalization function (Softmax) 1133. The speech to be replied 114 is input into this structure to obtain identification information 115 of the first subject, which also includes information such as the first subject's current emotion, accent, and age.
[0117] like Figure 11 The example shows how to process voice input, using contrastive learning, which has better feature representation and generalization capabilities, and can more effectively achieve user identity classification tasks. The CCNet network structure is as follows Figure 11 As shown, the Text Encoder, Audio Encoder, and Image Encoder are all based on the Transformer network structure. Each encoder uses a specific feature extraction network for each modality to improve consistency during subsequent feature fusion and facilitate model training. Here, the processing method for voice input is shown, and there will also be a processing method for image input. Therefore, the preprocessing network layer is set up with the network layer involved in image preprocessing, and the feature extraction network layer will be set up with a network layer for image feature extraction such as a two-dimensional convolution layer. The classifier is also set up in an image processing method, and comparative learning can also be used to determine the user label corresponding to the image information.
[0118] Step S1002: In response to the preset database storing voice style prompt information corresponding to the identification information, the voice style prompt information corresponding to the identification information is obtained from the preset database.
[0119] In an embodiment of the present application, if the voice style prompt information corresponding to the identification information is stored in a preset database, the voice interaction device can directly obtain the corresponding voice style prompt information from the preset database.
[0120] Step S1003: Generate a reply voice corresponding to the voice to be replied based on the voice style prompt information.
[0121] In an embodiment of the present application, if there is voice style prompt information in a preset database, a reply voice corresponding to the voice to be replied is generated directly based on the voice style prompt information.
[0122] In this way, if the voice style prompt information corresponding to the first object is stored in the preset database, it can be used directly without being generated, thereby reducing the time for generating the voice style prompt information and improving the efficiency of voice interaction.
[0123] In some embodiments, after the voice interaction device executes the above step S1002 of "obtaining voice style prompt information corresponding to the identification information from the preset database", Figure 12 As shown, the following steps S1201 and S1202 are also included:
[0124] Step S1201: In response to the fact that the emotional state represented by the speech to be replied does not match the emotional state represented by the generated speech style prompt information, speech style prompt information corresponding to the first object is generated based on multimodal information, and a reply speech corresponding to the speech to be replied is generated based on the speech style prompt information.
[0125] In an embodiment of the present application, the voice style prompt information contains the user's current emotional state. After obtaining the voice style prompt information, the voice interaction device will continue to determine whether the emotional state represented by the voice to be replied matches the emotional state represented by the generated voice style prompt information. If they do not match, the voice interaction device will generate voice style prompt information corresponding to the first object based on the multimodal information, and generate a reply voice corresponding to the voice to be replied based on the voice style prompt information.
[0126] In this way, by further considering the matching of the user's emotional state, the matching degree between the voice style prompt information and the first object is improved, thereby making the determined voice style prompt information more accurate, and the reply voice determined based on the voice style prompt information will also be more accurate.
[0127] Step S1202: In response to the emotional state represented by the to-be-replied speech matching the emotional state represented by the generated speech style prompt information, a reply speech corresponding to the to-be-replied speech is generated based on the speech style prompt information.
[0128] In an embodiment of the present application, if the emotional state represented by the to-be-replied voice matches the emotional state represented by the generated voice style prompt information, a reply voice corresponding to the to-be-replied voice will be generated directly based on the voice style prompt information.
[0129] In this way, after finding the voice style prompt information corresponding to the identification information, it will be further determined whether the emotional state is consistent with the voice style prompt information. In this way, the obtained voice style prompt information will be more accurate and the voice interaction will be more intelligent.
[0130] In some embodiments, after executing the above step of "generating voice style prompt information corresponding to the first object", the voice interaction device may also execute the following steps: establishing a correspondence between the identification information corresponding to the first object and the voice style prompt information, and storing the established correspondence in a preset database; wherein the identification information includes the identity information of the first object.
[0131] In an embodiment of the present application, the identification information includes identity information of the first object, such as: user name, ID number, number assigned to represent identity, etc. Of course, it can also be other information representing the user's identity, user biometrics, or other.
[0132] In an embodiment of the present application, the above-mentioned step S1001 performs user identity recognition on the first object. In addition to recognition based on the voice to be replied, it can also be based on the face of the first object, the collected user biometrics, or the identity information entered by the user.
[0133] In an embodiment of the present application, the voice interaction device establishes a correspondence between the identification information corresponding to the first object and the voice style prompt information, and stores the established correspondence in a preset database.
[0134] For example, Figure 13 As shown, conditional judgment is performed based on the label results (identification information) recognized by intelligent perception. When it is recognized that the speaker (first object) 131 has been registered in the preset database (database) 132, the corresponding prompt instruction (voice style prompt information) and reference timbre (instruction prompt output) 133 are directly extracted. When it is recognized that the speaker 132 has not been registered in the preset database 132, the dialect, age, timbre and other information (user label) 41 are input into LM42 to generate prompt text (voice style prompt information) 43.
[0135] In an embodiment of the present application, the voice interaction device can also set multiple corresponding voice style prompt information for the same user. For example, a corresponding voice style prompt information can be set for different emotional states of the same user, or corresponding voice style prompt information can be set for different emotional accents of the same user, or different languages. This will be more accurate.
[0136] In some embodiments, as Figure 14 As shown, the voice interaction device may further perform the following steps S1401 to S1404:
[0137] Step S1401: Obtain sample data.
[0138] In an embodiment of the present application, the sample data may be multimodal sample data.
[0139] Step S1402: Generate style prompt information for the sample data using the style generation model to be trained to obtain style prompt information corresponding to the sample data.
[0140] In an embodiment of the present application, the style generation model to be trained has a consistent network structure with the preset style generation model, and the voice interaction device generates style prompt information for the sample data using the style generation model to be trained to obtain style prompt information corresponding to the sample data.
[0141] Step S1403: Calculate the loss information between the style prompt information of the sample data and the target style prompt information preset for the sample data.
[0142] In an embodiment of the present application, the voice interaction device may calculate loss information between the style prompt information of the sample data and the voice style prompt information preset for the sample data.
[0143] Step S1404: Based on the loss information, adjust the model parameters of the style generation model to be trained to obtain a preset style generation model.
[0144] In an embodiment of the present application, network parameters in the style generation model to be trained are adjusted based on the loss information to obtain a preset style generation model.
[0145] In an embodiment of the present application, the style generation model to be trained includes a label generation sub-model to be trained and a style prompt information generation sub-model to be trained, wherein the label generation sub-model to be trained and the style prompt information generation sub-model to be trained can be trained together as discussed in steps S1401 to S1404, and of course can also be trained separately, that is, the label generation sub-model to be trained and the style prompt information generation sub-model to be trained are trained separately.
[0146] Exemplarily, when sample data is input into the sub-model for label generation to be trained, user sample labels will be obtained. The voice interaction device will determine the cross-entropy loss information between the user sample labels and the user target labels preset for the sample data, and then adjust the model parameters of the sub-model for label generation to be trained based on the cross-entropy loss information to obtain the preset label generation sub-model.
[0147] For example, the method of determining the loss information between the user sample label and the user target label preset for the sample data is shown in formula (10):
[0148]
[0149] Among them, l text and l fusionRepresents the cross entropy loss information of the similarity matrix C at different dimensions. The mathematical expression is as follows:
[0150]
[0151] In this way, the label generation sub-model to be trained can be adjusted based on the cross entropy loss information to adjust the model parameters and obtain the preset label generation sub-model.
[0152] For example, Figure 15 As shown in FIG, a flowchart of an exemplary voice interaction method is provided. Figure 15 As shown, the voice interaction method can be implemented on a voice interaction device. The voice interaction device 15 may include an intelligent perception module 151, a prompt generation module 152, and a voice interaction module 153. The voice interaction method is discussed herein in conjunction with the voice interaction device 150. Exemplarily, the voice interaction method includes the following steps S1501 to S1503:
[0153] Step S1501: Determine user tags.
[0154] Here, the voice interaction device 15 obtains multimodal information about the user (first object) through sensors 1511: camera 1512, microphone 1513, radar 1514, and infrared 1515. This multimodal information is input into the perception algorithm 1516 to obtain the user tag 41. That is, the intelligent perception module 151 collects the user's multimodal information and generates the user tag 1517. The multimodal scene perception system includes multiple AIPC-equipped sensors, such as cameras, microphones, and radars. Based on the multimodal information extracted by these sensors, identity confirmation and scene perception can be combined.
[0155] Step S1502: Determine voice style prompt information.
[0156] Here, the voice interaction device 150 matches the corresponding voice style prompt information from the database (preset database) 132 based on the user tag 1517. If there is no voice style prompt information, the voice style prompt information is generated through LM42. Among them, the voice style prompt information (instruction set) 1521 involving the reference timbre 1522 also needs to obtain the corresponding target reference timbre 1523, such as S-type timbre; the matching method from the preset database can be based on the identification information in the user tag, or it can be matched together with the identification information, emotional state, etc. in the user tag, or the information involved in the entire user tag, or part of the information in the user tag, which can be determined based on the corresponding relationship in the preset database. That is, the prompt generation module 152 generates a prompt instruction (voice style prompt information) 43 based on the user tag 41. According to the tag output result of the intelligent perception module 151, the model can generate a corresponding style prompt sentence. If the user identity has been registered, the matching style prompt is extracted from the database. If the user requires a reference timbre, the database provides the corresponding timbre. By adaptively generating style hints, the operations required to change the speech synthesis style when the user changes can be effectively reduced, thereby improving the usability of the voice interaction system.
[0157] Step S1503: Output reply voice.
[0158] Here, the voice to be replied 1531 is input to the voice recognition model 1532 through the voice interaction module 153 to generate a transcribed text, which is then input into the large language model 1533. The large language model 1533 generates a corresponding reply text based on the transcribed text. The voice synthesis model 1534 generates a voice reply 1535 based on the reply text, the prompt instruction (voice style prompt information) 43 and the target timbre reference 1523. The voice interaction module includes ASR, LLM, and TTS. Among them, TTS can generate a voice synthesis style that meets the user's expectations based on the style prompt text and the reference timbre. When the conversation user changes, the style can be switched seamlessly. It is independent of the existing voice interaction system and can be used without retraining the original voice interaction system. It has high adaptability and a wider range of application scenarios.
[0159] An exemplary implementation of voice interaction is: perception algorithm, through Figure 11The provided network structure performs user identity recognition, combines voiceprint and facial multimodal information for detection, and assigns different ID tags (identification information) to different users. When the ID is not registered, the subsequent prompt generation module generates a style prompt sentence and binds it to the ID and saves it in the database (preset database) as an instruction set (voice style prompt information). When the ID is registered, the style prompt sentence ((voice style prompt information)) corresponding to the ID is directly extracted from the database, and no other label prediction is performed to reduce latency and improve user experience. Secondly, emotion recognition, dialect prediction, age prediction and other tasks are performed through deep learning network models, and corresponding labels are generated. Different labels can control different style dimensions. For example, the dialect controls the generated accent, which can be easier for users to understand and the generated content is more friendly. The user's age can control the speed of the generated voice, and the distance can control the volume of the voice playback through post-processing.
[0160] The prompt generation module consists of a language model (LM) and a database. The LM, based on a decoder-only Transformer architecture, generates natural language prompt text based on input labels. The database contains registered and stored instruction sets and reference timbre. Compared to discrete prompt labels, the prompt sentences generated by the LM are richer, more diverse, and more natural. They are highly consistent with user-defined style prompts and LLM-generated content during voice interaction, all in natural language, making them easier for the TTS model to understand and reason about.
[0161] In the embodiments of the present application, the intelligent perception module and the style prompt generation module are both trained and inferred locally on the AIPC, and the user's personal information such as face and voice are stored locally. Even if a cloud-based voice interaction module is used, the privacy and security of the user's information can be ensured.
[0162] After collecting the reply voice corresponding to the first object, the module first performs conditional judgment on the label results recognized by the intelligent perception module. When it is recognized that the speaker has been registered in the database, the corresponding prompt instructions and reference timbre are directly extracted and output to the downstream TTS model. When it is recognized that the speaker has not been registered in the database, the dialect, age, timbre and other information are input into the LM to generate the prompt text. This module includes a speech recognition model, a large language model and a speech synthesis model. The speech recognition model recognizes the user's speech input as text, the large language model generates a reply based on the text, and the speech synthesis model generates different styles of speech based on the reply text, the prompt instructions provided by the prompt generation module and the reference speech. The specific input of the speech synthesis model can be expressed as:
[0163] <soi>Please use an S-type voice, a B-speak accent, and speak this sentence slowly. <sot>Nice to meet you! <laughter> <sosp>speech prompt <sos>speech out <eos>;
[0164] in, <soi>Indicates start of instruction, <sot>Indicates the start of text, <laughter>Indicates the inserted emotion tag, <sosp>represents the reference voice, which is the S type timbre in this embodiment, <sos>Indicates start of speech, <eos>Indicates the end of speech.
[0165] Thus, the embodiment of the present application proposes a voice interaction method based on intelligent scene perception, which aims to adjust the TTS generation style in real time and adaptively according to the user identity and usage scenario. The user only needs to set the style once to register, and then each conversation will be automatically called to improve the cross-user experience. In addition, for unregistered users, the intelligent perception system will also automatically generate appropriate voice style prompt information. The embodiment of the present application can be widely used in intelligent voice interaction scenarios, simplify the user's voice interaction process, and improve the user experience of AIPC end-side voice interaction, which has high practical value.
[0166] An embodiment of the present application provides a voice interaction method, comprising: obtaining multimodal information corresponding to a first object in response to a collected voice message to be replied to corresponding to the first object; generating voice style prompt information corresponding to the first object based on the multimodal information using a preset style generation model; and generating a reply voice message corresponding to the voice message to be replied to based on the voice style prompt information. The voice interaction method provided in this application automatically generates voice style prompt information matching the user based on the multimodal information, thereby simplifying the user's voice interaction process and improving the user's voice interaction experience.
[0167] The embodiment of the present application provides a voice interaction device 16, such as Figure 16 As shown, it includes: a sensor 161 and a processor 162;
[0168] The sensor 161 is configured to obtain multimodal information corresponding to the first subject in response to the collected speech to be replied corresponding to the first subject;
[0169] The processor 162 is configured to generate voice style prompt information corresponding to the first object based on the multimodal information using a preset style generation model; and generate a reply voice corresponding to the voice to be replied based on the voice style prompt information.
[0170] In one embodiment of the present application, the preset style generation model includes a label generation sub-model and a style prompt information generation sub-model. The processor 162 is further used to use the label generation sub-model to determine the user label corresponding to the first object based on multimodal information; and use the style prompt information generation sub-model to generate voice style prompt information corresponding to the first object based on the user label.
[0171] In one embodiment of the present application, the processor 162 is further configured to establish a correspondence between the identification information corresponding to the first object and the voice style prompt information, and store the established correspondence in a preset database; wherein the identification information includes the identity information of the first object.
[0172] In one embodiment of the present application, the label generation sub-model includes: a feature extraction network and a label generation network; the processor 162 is further used to use the feature extraction network to perform feature extraction and modality alignment on each modal information in the multimodal information to obtain corresponding multiple feature information; and use the label generation network to generate user labels based on multiple feature information.
[0173] In one embodiment of the present application, the processor 162 is further used to perform feature extraction on each modal information in the multimodal information to obtain corresponding multiple extracted feature information; linearize each extracted feature information in the multiple extracted feature information to obtain corresponding multiple linearized feature information; and normalize each linearized feature information in the multiple linearized feature information to obtain corresponding multiple feature information.
[0174] In one embodiment of the present application, the multimodal information includes identification information, and at least one of image information and voice information; the label generation network includes a feature fusion network layer and a label classification network layer; the processor 162 is further used to use the feature fusion network layer to fuse feature information corresponding to information other than identification information in multiple feature information to obtain fused feature information; and use the label classification network layer to perform matrix multiplication on the fused feature information and the feature information corresponding to the identification information in the multiple feature information to obtain a user label.
[0175] In one embodiment of the present application, the processor 162 is also used to identify the first object and determine the identification information corresponding to the first object; in response to the voice style prompt information corresponding to the identification information stored in the preset database, obtain the voice style prompt information corresponding to the identification information from the preset database; and generate a reply voice corresponding to the voice to be replied based on the voice style prompt information.
[0176] In one embodiment of the present application, the processor 162 is further used to, in response to a mismatch between the emotional state represented by the voice to be replied and the emotional state represented by the generated voice style prompt information, generate voice style prompt information corresponding to the first object based on multimodal information, and generate a reply voice corresponding to the voice to be replied based on the voice style prompt information; in response to a match between the emotional state represented by the voice to be replied and the emotional state represented by the generated voice style prompt information, generate a reply voice corresponding to the voice to be replied based on the voice style prompt information.
[0177] In one embodiment of the present application, the processor 162 is further used to obtain sample data; generate style prompt information for the sample data using the style generation model to be trained to obtain style prompt information corresponding to the sample data; calculate the loss information between the style prompt information of the sample data and the target style prompt information preset for the sample data; and adjust the model parameters of the style generation model to be trained based on the loss information to obtain a preset style generation model.
[0178] An embodiment of the present application provides a voice interaction device that, in response to a collected speech to be replied corresponding to a first object, obtains multimodal information corresponding to the first object; generates voice style prompt information corresponding to the first object based on the multimodal information using a preset style generation model; and generates a reply speech corresponding to the speech to be replied based on the voice style prompt information. The voice interaction device provided in this application automatically generates voice style prompt information matching the user based on the multimodal information, thereby simplifying the user's voice interaction process and improving the user's voice interaction experience.
[0179] An embodiment of the present application provides a computer-readable storage medium, which stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement the above-mentioned voice interaction method. The computer-readable storage medium can be a volatile memory (volatile memory), such as a random-access memory (RAM); or a non-volatile memory (non-volatile memory), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or it can be a respective device including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.
[0180] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0181] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0182] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0184] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.< / eos> < / sos> < / sosp> < / laughter> < / sot> < / soi> < / eos> < / sos> < / sosp> < / laughter> < / sot> < / soi>
Claims
1. A voice interaction method, comprising: In response to the collected speech to be replied corresponding to the first object, obtaining multimodal information corresponding to the first object; Generate voice style prompt information corresponding to the first object based on the multimodal information using a preset style generation model; Based on the voice style prompt information, a reply voice corresponding to the voice to be replied is generated.
2. The voice interaction method according to claim 1, wherein the preset style generation model includes a label generation sub-model and a style prompt information generation sub-model, and wherein generating voice style prompt information corresponding to the first object based on the multimodal information using the preset style generation model comprises: Determining a user tag corresponding to the first object based on the multimodal information using the tag generation sub-model; The style prompt information is used to generate a sub-model, and based on the user tag, the voice style prompt information corresponding to the first object is generated.
3. The voice interaction method according to claim 2, after generating the voice style prompt information corresponding to the first object, the method further comprises: Establishing a correspondence between the identification information corresponding to the first object and the voice style prompt information, and storing the established correspondence in the preset database; The identification information includes the identity information of the first object.
4. The voice interaction method according to claim 2, wherein the label generation sub-model comprises: Feature extraction network and label generation network; The using the label generation sub-model to determine the user label corresponding to the first object based on the multimodal information includes: Using the feature extraction network, performing feature extraction and modality alignment on each modal information in the multimodal information to obtain corresponding multiple feature information; The tag generation network is used to generate the user tag based on the plurality of feature information.
5. The voice interaction method according to claim 4, wherein the step of performing feature extraction and modality alignment on each modal information in the multimodal information to obtain corresponding multiple feature information includes: Performing feature extraction on each modal information in the multimodal information to obtain corresponding multiple extracted feature information; Linearizing each of the plurality of extracted feature information respectively to obtain corresponding plurality of linearized feature information; Each piece of linearized feature information is normalized to obtain corresponding pieces of feature information.
6. The voice interaction method according to claim 4, wherein the multimodal information includes identification information and at least one of image information and voice information; the label generation network includes a feature fusion network layer and a label classification network layer; The step of generating the user tag based on the plurality of feature information by using the tag generation network includes: Using the feature fusion network layer, the feature information corresponding to the information other than the identification information in the plurality of feature information is fused to obtain fused feature information; The label classification network layer is used to perform matrix multiplication on the fused feature information and the feature information corresponding to the identification information in the plurality of feature information to obtain the user label.
7. The voice interaction method according to claim 1, further comprising: Performing identity recognition on the first object to determine identification information corresponding to the first object; In response to the voice style prompt information corresponding to the identification information being stored in the preset database, acquiring the voice style prompt information corresponding to the identification information from the preset database; Based on the voice style prompt information, a reply voice corresponding to the voice to be replied is generated.
8. The voice interaction method according to claim 7, after obtaining the voice style prompt information corresponding to the identification information from the preset database, the method further comprises: In response to a mismatch between the emotional state represented by the speech to be replied and the emotional state represented by the speech style prompt information, generating speech style prompt information corresponding to the first object based on the multimodal information, and generating a reply speech corresponding to the speech to be replied based on the speech style prompt information; In response to the emotional state represented by the speech to be replied matching the emotional state represented by the generated speech style prompt information, a reply speech corresponding to the speech to be replied is generated based on the speech style prompt information.
9. The voice interaction method according to any one of claims 1 to 8, further comprising: Get sample data; Generating style prompt information for the sample data using the style generation model to be trained to obtain style prompt information corresponding to the sample data; Calculating loss information between the style prompt information of the sample data and target style prompt information preset for the sample data; Based on the loss information, model parameters of the style generation model to be trained are adjusted to obtain the preset style generation model.
10. A voice interaction device, comprising: Sensors and processors; The sensor is configured to obtain multimodal information corresponding to the first object in response to the collected voice to be replied corresponding to the first object; The processor is configured to generate voice style prompt information corresponding to the first object based on the multimodal information using a preset style generation model; Based on the voice style prompt information, a reply voice corresponding to the voice to be replied is generated.
Citation Information
Cited By
Accent and style control method and device of man-machine interaction dialogue system
CN122347941A