Interaction method, apparatus, terminal, electronic device, and storage medium
By integrating the acoustic, semantic, and visual features of audio and video data, the system automatically identifies user emotions and intentions, generating images that match those emotions and intentions. This solves the problem of manually selecting images in existing video call technologies, enhancing user experience and engagement.
Patent Information
- Application Number
- CN202211399409.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2042-11-09
AI Technical Summary
Existing video call technology has simple functions, requires manual selection of images from the gallery, which affects the user experience. The gallery resources are limited and fixed, lacking fun and entertainment.
By fusing acoustic, semantic, and visual features of audio and video data, the system automatically identifies the user's emotions and intentions, and generates images that match the user's emotions and intentions for interaction.
It enables interactive image generation based on audio and video, enhancing the user experience, providing a rich variety of images, and increasing fun and entertainment.
Smart Images

Figure CN115578679B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to an interaction method and device, a terminal, an electronic device, and a storage medium. BACKGROUND
[0002] With the development of Internet technology and the popularity of mobile terminals, video call technology is increasingly favored by people.
[0003] Current video call technology has relatively simple functions and can only support microphone recording and camera capturing of the environment. Although some video chat software now embeds voice changing, background blurring, background changing, sending images outside the video, or changing the appearance of the person in the current video to enrich the fun of video calls. However, the functions provided by the current video call software all require manual operation, especially when sending additional interesting images outside the video, the user needs to manually select from the gallery and send each time. At the same time, the gallery resources are few and fixed, which affects the user experience.
[0004] Therefore, how to enrich the fun and entertainment of video calls and enhance the user experience is a problem that needs to be solved. SUMMARY
[0005] The present application provides an interaction method, device, terminal, electronic device and storage medium to solve the defects of simple video call function and manual selection from the gallery in the prior art, which affects the user experience.
[0006] The present application provides an interaction method, comprising:
[0007] obtaining audio and video data;
[0008] based on at least one of the acoustic features, semantic features and visual features of the audio and video data, performing intent and / or emotion recognition on the audio and video data to obtain intent and / or emotion recognition results of the audio and video data;
[0009] based on the intent and / or emotion recognition results of the audio and video data, determining an image matching the intent and / or emotion recognition results, and performing interaction based on the image.
[0010] According to the interaction method provided by the present application, the determination of the image matching the intent and / or emotion recognition results based on the intent and / or emotion recognition results of the audio and video data comprises:
[0011] The image prediction module is trained based on an intention and / or emotion label and a sample image matched with the intention and / or emotion label.
[0012] The pre-trained image decoding module is used for decoding the predicted image feature to obtain an image matched with the intention and / or emotion recognition result.
[0013] According to the interaction method provided in the application, the obtaining step of the image prediction module comprises:
[0014] An initial image prediction module is obtained.
[0015] The pre-trained image encoding module is used for encoding the sample image to obtain a sample image feature of the sample image.
[0016] The initial image prediction module is used for predicting an image feature of the intention and / or emotion label to obtain a predicted image feature corresponding to the intention and / or emotion label.
[0017] The initial image prediction module is iterated based on a similarity between the sample image feature and the predicted image feature to obtain an image prediction module.
[0018] According to the interaction method provided in the application, the intention and / or emotion recognition of the audio and video data based on at least one of the acoustic feature, the semantic feature and the visual feature of the audio and video data comprises:
[0019] The acoustic feature, the semantic feature and the visual feature of the audio and video data are fused to obtain a fused feature.
[0020] The intention and / or emotion recognition of the audio and video data is performed based on the fused feature to obtain an intention and / or emotion recognition result of the audio and video data.
[0021] According to the interaction method provided in the application, the intention and / or emotion recognition of the audio and video data based on the fused feature comprises:
[0022] The intention and emotion features of the audio and video data are extracted based on the fused feature to obtain a first intention feature and a first emotion feature.
[0023] interact the first intention feature and the first emotion feature based on a correlation between the first intention feature and the first emotion feature, to obtain a second intention feature and a second emotion feature;
[0024] perform intention recognition on the audio and video data based on the first intention feature and the second emotion feature, to obtain an intention recognition result of the audio and video data; and / or,
[0025] perform emotion recognition on the audio and video data based on the first emotion feature and the second intention feature, to obtain an emotion recognition result of the audio and video data.
[0026] According to the interaction method provided by the application, the determination of the acoustic feature and the semantic feature of the audio and video data comprises:
[0027] extract the acoustic feature of the audio data in the audio and video data based on an encoder in a pre-trained speech recognition model, to obtain the acoustic feature of the audio data;
[0028] extract the acoustic feature based on a decoder in the pre-trained speech recognition model, to obtain the hidden layer feature of the audio data;
[0029] extract the semantic feature of the hidden layer feature of the audio data based on a pre-trained language model, to obtain the semantic feature of the audio data.
[0030] According to the interaction method provided by the application, the obtaining of the audio and video data comprises:
[0031] obtaining the audio and video data in the video communication of the first client user;
[0032] The interaction based on the image comprises:
[0033] In the case that the first client opens the image interaction function, the image is sent to the second client, and the image is displayed on the second client.
[0034] The application further provides an interaction device, comprising:
[0035] a data acquisition unit configured to acquire audio and video data;
[0036] an identification unit configured to perform intention and / or emotion recognition on the audio and video data based on at least one of an acoustic feature, a semantic feature and a visual feature of the audio and video data, to obtain an intention and / or emotion recognition result of the audio and video data;
[0037] An interaction unit is configured to determine an image matching the intention and / or emotion recognition result based on the intention and / or emotion recognition result of the audio-video data, and perform interaction based on the image.
[0038] The application also provides a terminal comprising a camera, a microphone and a processor connected in sequence.
[0039] The camera is configured to acquire video data.
[0040] The microphone is configured to acquire audio data.
[0041] The processor is configured to perform intention and / or emotion recognition on the audio-video data based on at least one of acoustic features, semantic features and visual features of the audio-video data, to obtain an intention and / or emotion recognition result of the audio-video data, determine an image matching the intention and / or emotion recognition result based on the intention and / or emotion recognition result of the audio-video data, and perform interaction based on the image.
[0042] The application also provides an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the interaction method according to any one of the above when executing the program.
[0043] The application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the interaction method according to any one of the above.
[0044] The application also provides a computer program product comprising a computer program, wherein the computer program is executable on a processor to implement the interaction method according to any one of the above.
[0045] The interaction method, device, terminal, electronic device and storage medium provided by the application automatically extract emotion and / or intention information represented by audio-video based on at least one of acoustic features, semantic features and visual features of the audio-video, use the emotion and / or intention information to obtain an image matching the emotion and / or intention of a user, and then perform interaction based on the obtained image. Compared with the prior art in which a user manually selects an image to perform interaction, the application can automatically obtain an image based on audio-video and perform interaction based on the image, thereby enhancing the experience of the user. The obtained image matches the emotion and / or intention of the user, and more abundant and diversified images increase the interest and entertainment. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to make the technical solutions in the present application or prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those ordinary skilled in the art without creative effort based on these drawings.
[0047] Figure 1 is a flowchart of the interactive method provided by the present application;
[0048] Figure 2 is a flowchart of step 130 in the interactive method provided by the present application;
[0049] Figure 3 is a flowchart of the image prediction module acquisition method provided by the present application;
[0050] Figure 4 is a flowchart of the image prediction module acquisition method provided by the present application;
[0051] Figure 5 is a flowchart of step 120 in the interactive method provided by the present application;
[0052] Figure 6 is a flowchart of the feature fusion provided by the present application;
[0053] Figure 7 is a flowchart of step 120 in the interactive method provided by the present application;
[0054] Figure 8 is a structural diagram of the intention and emotion branch interaction module provided by the present application;
[0055] Figure 9 is a flowchart of the recognition method provided by the present application;
[0056] Figure 10 is a structural diagram of the interactive device provided by the present application;
[0057] Figure 11 is a structural diagram of a terminal provided by the present application;
[0058] Figure 12 is a structural diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0059] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the protection scope of the present application.
[0060] In current applications based on video call technology, whether it is video conference, live broadcast, video call and other software, in order to realize sending interesting images to the other party in the video call, it is necessary to manually find the image that the user wants from the gallery and click to send. The sending mode can be sent through the dialog box, or the image can be displayed on the user terminal by adding the image in the video. At the same time, due to the fixed and small number of gallery resources, the user experience is affected.
[0061] In view of the above problems, the embodiment of the present application provides an interaction method, and the technical concept of the method is to automatically extract the emotion and / or intention information represented by the audio and video data, and then use the emotion and / or intention information to obtain images that conform to the user's emotion and / or intention, and then interact based on the obtained images. Compared with the way of manually selecting images for interaction in the prior art, the present application can realize interaction based on automatically obtained images, which can enhance the user's experience. At the same time, the obtained images conform to the user's emotion and / or intention, and more rich and diverse images also increase the interest and entertainment.
[0062] It should be noted that the interaction method provided by the present application can be applied not only to various video call scenarios, but also to live broadcast, video conference, personalized recording and various scenarios.
[0063] Figure 1 is a flowchart of the interaction method provided by the present application, and the execution subject of each step in the method can be an interaction device, which can be realized by software and / or hardware, and the device can be integrated in an electronic device, which can be a smart phone, a tablet computer and other hardware devices with various operating systems, touch screens and / or display screens. As shown in Figure 1 The interaction method provided by the embodiment of the present application can include the following steps:
[0064] Step 110, obtaining audio and video data.
[0065] Specifically, the audio-video data can be pure audio data, pure video data, or audio and video data, and the embodiments of the present application do not make specific limitations thereon. The audio-video data can be collected by an audio-video device, for example, a camera can be used to collect a user's face or action video to obtain video data, and a sound collecting device can be used to collect a user's interactive voice to obtain audio data. The audio-video data can be obtained by local recording or through a network. In addition, the audio-video data can be a video segment formed after recording, or a video stream in a real-time recording process, for example, a video stream obtained in a video call or a live broadcast scenario.
[0066] In step 120, at least one of the acoustic features, semantic features, and visual features of the audio-video data is used to perform intent and / or emotion recognition on the audio-video data to obtain an intent and / or emotion recognition result of the audio-video data.
[0067] Specifically, after obtaining the audio-video data, intent and / or emotion recognition can be performed on the audio-video data to obtain an intent and / or emotion recognition result.
[0068] In most scenarios, there is a correlation between human intent and emotion. Here, emotion recognition can be performed alone, intent recognition can be performed alone, or emotion recognition and intent recognition can be performed simultaneously. The intent recognition result can be clapping with both hands, giving flowers, shopping, running, drinking water, etc., and the emotion recognition result can be any one of happy, angry, sad, surprised, disgusted, fearful, and neutral, and the embodiments of the present application do not make specific limitations thereon.
[0069] To obtain the intent and / or emotion recognition result, at least one of the acoustic features, semantic features, and visual features of the audio-video data can be used for recognition. It can be understood that any one of the acoustic features, semantic features, and visual features can be used for intent and / or emotion recognition, or any two or all of them can be used for intent and / or emotion recognition.
[0070] The acoustic features of the audio-video data can be obtained by performing acoustic feature extraction on the audio-video data, and can specifically include pitch, loudness, speech rate, etc. For example, when the pitch is high, the emotional state is more intense or angry, and when the pitch is low, the emotional state is more calm or rational. For another example, when the loudness is high, the emotional state is more intense or angry, and when the loudness is low, the emotional state is more comfortable or optimistic. For another example, when the speech rate is fast, the emotional state is more anxious or enthusiastic, and when the speech rate is slow, the emotional state is more calm or sad. Therefore, emotion recognition can be performed based on the acoustic features of the audio-video data. Because the emotional state is a long-term process, the acoustic features can be extracted in units of sentences or in units of paragraphs, which are not limited here.
[0071] The semantic feature of the audio-video data can be obtained by performing semantic feature extraction on the audio-video data. The semantic feature can represent the user's intention expressed by the audio-video data to some extent, such as "I want to send you flowers" or "I want to drink water". In addition, the keywords in the audio-video data can also reflect the user's emotional state, such as the Chinese keywords "huh", "ha", "hum", "alas", and the English keywords "nice", "opps", "yeah", which can be used as a basis for judging the user's emotional state.
[0072] The visual feature of the audio-video data can specifically include facial expression features and body movement features of the user in the video image. For example, when a person raises his eyebrows or squeezes them together, it shows a doubtful or worried intention; when the muscles around the mouth contract and rise, or when the mouth is wide open and laughing, it shows a positive or happy emotion. Some habitual movements, gestures, postures, and body postures can express the user's intention, such as nodding to indicate approval, shaking the head to indicate disapproval, and the like.
[0073] In step 130, based on the intention and / or emotion recognition result of the audio-video data, an image matching the intention and / or emotion recognition result is determined, and interaction is performed based on the image.
[0074] Specifically, after obtaining the intention and / or emotion recognition result of the audio-video data, an image matching the intention and / or emotion recognition result can be determined. For example, if the intention recognition result is "I want to drink coffee", the image matching it can be a cup of coffee; for example, if the emotion recognition result is happy, the image matching it can be a bouquet of flowers or a smiley image, and the like.
[0075] The image matching or image generation method can be used to obtain the image matching the intention and / or emotion recognition result. For example, based on the intention and / or emotion recognition result, a text-to-image generation technology can be used to generate an image matching the intention and / or emotion recognition result; or a pre-stored image library can be used, and an intention and / or emotion recognition label can be set for each image in the image library, and an image matching the intention and / or emotion recognition result can be selected from the image library based on the intention and / or emotion recognition result. The determination method of the image is not limited in the embodiments of the present application.
[0076] On this basis, interaction is performed based on the determined image matching the intention and / or emotion recognition result, for example, the image can be displayed on the display screen of the terminal, and the interaction method can be flexibly selected according to actual needs, for example, the image can be displayed in a gradually enlarged manner, or in a rotating manner, and the like.
[0077] The method provided by the embodiment of the present application can automatically extract the emotion and / or intention information represented by the audio and video through at least one of the acoustic features, semantic features and visual features of the audio and video, obtain an image that conforms to the emotion and / or intention of the user by using the emotion and / or intention information, and then interact based on the obtained image. Compared with the prior art in which the user manually selects an image for interaction, the present application can automatically obtain an image based on the audio and video and interact based on the image, thereby enhancing the experience of the user. The obtained image conforms to the emotion and / or intention of the user, and the more diverse and rich images increase the interest and entertainment.
[0078] Based on the above embodiment, Figure 2 is a flowchart of step 130 in the interaction method provided by the present application, as Figure 2 shown, in step 130, based on the intention and / or emotion recognition result of the audio and video data, an image that matches the intention and / or emotion recognition result is determined, specifically including:
[0079] In step 131, the image feature prediction module is used to predict the image features of the intention and / or emotion recognition result, to obtain the predicted image features corresponding to the intention and / or emotion recognition result. The image prediction module is trained based on the intention and / or emotion label and the sample image matching the intention and / or emotion label.
[0080] In step 132, the pre-trained image decoding module is used to decode the predicted image features, to obtain the image that matches the intention and / or emotion recognition result.
[0081] Specifically, in order to realize the determination of the image that matches the intention and / or emotion recognition result based on the intention and / or emotion recognition result of the audio and video data, a text-image generation model can be used. The text-image generation model can specifically include an image prediction module and an image decoding module. The image prediction module is used to predict the image features of the intention and / or emotion recognition result, to obtain the predicted image features corresponding to the intention and / or emotion recognition result. The image decoding module is used to decode the predicted image features, to obtain the image that matches the intention and / or emotion recognition result. The image decoding module can be pre-trained, for example, it can be a VQGAN decoder.
[0082] Before step 131 is performed, the image prediction module can be trained. First, a large amount of intention and / or emotion labeled text and sample images matching the intention and / or emotion label are collected. Then, the intention and / or emotion labeled text is input into the initial image prediction module, and the sample images are input into the image encoding module for training. During the training process, the initial image prediction module can amplify and learn the difference between the sample image features and the predicted image features, so that the trained image prediction module can better distinguish the difference between the sample image features and the predicted image features.
[0083] The intention and / or emotion label and the sample image matching the intention and / or emotion label can be obtained by crawling images commonly used to express intention and emotion in chats from the Internet as sample images, and labeling the sample images with and / or labels through manual labeling.
[0084] The method provided by the embodiment of the application is based on the image prediction module and the pre-trained image decoding module, and the intention and / or emotion recognition result is applied to obtain an image matching the intention and / or emotion recognition result, so that the image meeting the user's emotion and / or intention is automatically generated, the user experience is further enhanced compared with the manual selection from the gallery, and more rich and diverse images are obtained.
[0085] Based on any of the above embodiments, Figure 3 is one of the flowcharts of the image prediction module acquisition method provided by the application, as shown in Figure 3 The acquisition steps of the image prediction module include:
[0086] Step 310, acquiring an initial image prediction module;
[0087] Step 320, based on the initial image prediction module, predicting the image features of the intention and / or emotion label to obtain the predicted image features corresponding to the intention and / or emotion label;
[0088] Step 330, based on the pre-trained image encoding module, encoding the features of the sample images to obtain the sample image features of the sample images;
[0089] Step 340, based on the similarity between the sample image features and the predicted image features, iteratively adjusting the parameters of the initial image prediction module to obtain the image prediction module.
[0090] Specifically, to obtain the image prediction module, an initial image prediction module can be obtained first. The initial image prediction module can be a structure suitable for generating a task encoder-decoder. The intention and / or emotion label is input into the initial image prediction module, and the intention and / or emotion label is subjected to image feature prediction by the initial image prediction module to obtain a predicted image feature corresponding to the intention and / or emotion label.
[0091] It should be noted that the image prediction module can be combined with a pre-trained image encoding module during training. The sample image is input into the pre-trained image encoding module, and the sample image is subjected to feature encoding to obtain a sample image feature of the sample image.
[0092] On this basis, based on the similarity between the sample image feature and the predicted image feature, the initial image prediction module is subjected to parameter iteration to obtain the image prediction module. It can be understood that the greater the similarity between the sample image feature and the predicted image feature, the higher the matching degree between the obtained predicted image feature and the intention and / or emotion label; on the contrary, the smaller the similarity between the sample image feature and the predicted image feature, the lower the matching degree between the obtained predicted image feature and the intention and / or emotion label.
[0093] In one embodiment, Figure 4 is a flowchart of the image prediction module acquisition method provided by the present application, as shown in Figure 4 the initial image prediction module is an end-to-end structure, including a BART Encoder and a BART Decoder, and the intention and / or emotion label is input into the initial image prediction module to obtain a predicted image feature. The pre-trained image encoding module is a VQGAN-Encoder, and the sample image is input into the VQGAN-Encoder to obtain a sample image feature. Based on the similarity between the sample image feature and the predicted image feature, a cross-entropy loss function CELoss is determined. Based on the cross-entropy loss function, the initial image prediction module is subjected to parameter iteration to obtain the image prediction module.
[0094] Based on any one of the above embodiments, Figure 5 is a flowchart of step 120 in the interaction method provided by the present application, as shown in Figure 5 the step 120 specifically includes:
[0095] Step 121, the acoustic feature, semantic feature and visual feature of the audio and video data are subjected to feature fusion to obtain a fused feature;
[0096] Step 122, based on the fused feature, the intention and / or emotion of the audio and video data is identified to obtain an intention and / or emotion recognition result of the audio and video data.
[0097] Specifically, in the related art, most of the data of a single mode is applied for emotion recognition or intent recognition. For example, the data of a speech mode is applied for intent recognition, the data of a video mode is applied for emotion recognition, or the data of a text mode is applied for emotion recognition. Considering the recognition scheme based on the single-mode data, the recognition result accuracy is low, and in the face of complex actual scenes, the recognition scheme of the single-mode data cannot meet the needs of the actual scenes.
[0098] Meanwhile, considering that the data of each mode can complement each other, for example, visual features can not only capture the expression and action information of a speaker in a video, but also make full use of the lip information of the speaker to assist speech recognition when the mouth of the speaker is detected in the video.
[0099] Therefore, the embodiment of the present application fuses the acoustic features, semantic features and visual features of the audio-video data to obtain fused features. Here, the acoustic features, semantic features and visual features of the audio-video data can be spliced, or the acoustic features, semantic features and visual features can be weighted using an attention mechanism and then spliced, or the acoustic features, semantic features and visual features can be fused using a multilayer perceptron, which is not limited in the embodiment of the present application.
[0100] It can be understood that the fused features herein are features fused with the acoustic features, semantic features and visual features. Therefore, emotion recognition can be performed based on the fused features to obtain an emotion recognition result of the audio-video data, and / or intent recognition can be performed based on the fused features to obtain an intent recognition result of the audio-video data.
[0101] In one embodiment, due to the secondary complexity of the self-attention mechanism, the computational overhead is large, in order to be able to obtain the intent and emotion information of the speaker in real time, a multilayer perceptron (MLP, Multilayer Perceptron) can be used for feature fusion to obtain fused features. Figure 6 is a flowchart of feature fusion provided by the present application, as Figure 6 shown, the acoustic features, semantic features and visual features of the audio-video data are input into a multilayer perceptron to obtain fused features output by the multilayer perceptron. The multilayer perceptron includes a fusion layer, a linear layer and an activation layer. The fusion layer can use a Concatenate function, the linear layer can use a Linear function, and the activation layer can use a Relu activation function.
[0102] Based on any of the above embodiments, Figure 7 is a flowchart of step 120 in the interaction method provided by the present application, as Figure 7 shown, step 122 specifically includes:
[0103] Step 1221, based on the fusion feature, respectively, the audio and video data are subjected to intent and emotion feature extraction, and first intent feature and first emotion feature are obtained;
[0104] Step 1222, based on the correlation between the first intent feature and the first emotion feature, the first intent feature and the first emotion feature are interacted, and second intent feature and second emotion feature are obtained;
[0105] Step 1223, based on the first intent feature and the second emotion feature, the audio and video data are subjected to intent recognition, and the intent recognition result of the audio and video data is obtained; and / or,
[0106] Step 1224, based on the first emotion feature and the second intent feature, the audio and video data are subjected to emotion recognition, and the emotion recognition result of the audio and video data is obtained.
[0107] Specifically, based on the fusion feature for intent and / or emotion recognition, the fusion feature can be sent into two channels: intent branch and emotion branch. That is, the audio and video data are subjected to intent and emotion feature extraction, and first intent feature and first emotion feature are obtained.
[0108] Considering that in the training process, the positive correlation between intent and emotion can be effectively utilized through multi-task training. However, in some cases, the intent and emotion labels are not positively correlated, so in order to solve this problem, based on the correlation between the first intent feature and the first emotion feature, the first intent feature and the first emotion feature are interacted, and second intent feature and second emotion feature are obtained.
[0109] On this basis, based on the first intent feature and the second emotion feature, the audio and video data are subjected to intent recognition, and the intent recognition result of the audio and video data is obtained; and / or, based on the first emotion feature and the second intent feature, the audio and video data are subjected to emotion recognition, and the emotion recognition result of the audio and video data is obtained.
[0110] In one embodiment, the interaction of the first intent feature and the first emotion feature can be realized by an intent and emotion branch interaction module, Figure 8 is a structural schematic diagram of the intent and emotion branch interaction module provided by the application, as Figure 8 shown, the fusion feature enters the upper part intent module and the lower part emotion module respectively, after linear layer and activation function, the intent and emotion interaction module learns the influence of each other on both sides, and finally fuses its own feature and the other feature learned by the network to predict the respective labels.
[0111] The method provided by the embodiment of the application can improve the accuracy and reliability of recognition by interacting the intention feature and the emotion feature and identifying the features obtained after the interaction and the features of the self in the case that the intention and the emotion are not positively correlated.
[0112] Based on any of the above embodiments, the determining of the acoustic feature and the semantic feature of the audio-video data comprises:
[0113] Based on the encoder in the pre-trained speech recognition model, acoustic feature extraction is performed on the audio data in the audio-video data to obtain the acoustic feature of the audio data.
[0114] Based on the decoder in the pre-trained speech recognition model, feature extraction is performed on the acoustic feature to obtain the hidden layer feature of the audio data.
[0115] Based on the pre-trained language model, semantic feature extraction is performed on the hidden layer feature of the audio data to obtain the semantic feature of the audio data.
[0116] Specifically, the speech recognition model is a pre-trained model based on an end-to-end framework, and the encoder in the speech recognition model can perform acoustic feature extraction on the audio data in the audio-video data to obtain the acoustic feature of the audio data.
[0117] The semantic feature of the audio data can be obtained in combination with the pre-trained speech recognition model and the pre-trained language model, and the pre-trained language model can be a BERT model. The hidden layer feature of the audio data can be obtained by performing feature extraction on the acoustic feature through the decoder in the speech recognition model, and the semantic feature of the audio data can be obtained by performing semantic feature extraction on the hidden layer feature of the audio data based on the pre-trained language model.
[0118] In the fine-tuning process, the speech recognition and the BERT are jointly trained, and the embedding vector of the text input of the original BERT is replaced by the output of the hidden layer of the decoder of the speech recognition. In this way, the calculation overhead of obtaining the text by decoding the speech and then obtaining the semantic vector through the BERT in actual application can be reduced.
[0119] The method provided by the embodiment of the application can reduce the calculation overhead and improve the efficiency of feature extraction by replacing the input of the pre-trained language model with the output of the hidden layer of the decoder of the pre-trained speech recognition model.
[0120] Based on any of the above embodiments, the step 110 specifically comprises:
[0121] The audio-video data of the first client user in the video communication is obtained.
[0122] The interaction based on the image in the step 130 specifically comprises:
[0123] In a case where it is detected that the first client opens the image interaction function, the image is sent to the second client, and the image is displayed on the second client.
[0124] Specifically, the method provided by the embodiment of the present application can be applied in a video communication scenario to further enrich the video communication scenario. The audio and video data herein can be a video stream generated when the first client user and the second client user perform video communication.
[0125] In terms of function implementation, a selection switch can be added. When the user selects to use the image interaction function, that is, in a case where it is detected that the first client opens the image interaction function, the video communication software performs intention and / or emotion recognition on the audio and video data of the first client user in the communication process, obtains an intention and / or emotion recognition result of the audio and video data, then determines an image matched with the intention and / or emotion recognition result based on the intention and / or emotion recognition result of the audio and video data, and finally sends the generated image to be displayed above the video of the second client user.
[0126] Preferably, in order to enable the second client user to see the image sent by the first client user, a dynamic fade-in function is added to gradually display the image when the image is displayed.
[0127] It should be noted that, in the communication process, some audio and video segments cannot obtain the intention and emotion recognition result. Therefore, in order to reduce the calculation overhead caused by these useless video segments, it is determined whether the audio and video data is valid in actual application. If the obtained intention and emotion recognition result is NULL, step 130 is not executed, that is, the image does not need to be generated.
[0128] The method provided by the embodiment of the present application can increase the interestingness and entertainment of video communication.
[0129] Based on any of the above embodiments, an interaction method is provided, comprising:
[0130] S1, obtaining audio and video data of a first client user in video communication;
[0131] S2, in a case where it is detected that the first client opens the image interaction function, performing intention and / or emotion recognition on the audio and video data to obtain an intention and / or emotion recognition result of the audio and video data.
[0132] Figure 9 is a flowchart of the recognition method provided by the present application, as shown in Figure 9As shown, the audio data in the audio-video data is input into an encoder in a pre-trained speech recognition model to obtain acoustic features. Based on a decoder in the pre-trained speech recognition model, feature extraction is performed on the acoustic features to obtain hidden layer features of the audio data. Based on a pre-trained language model, semantic feature extraction is performed on the hidden layer features of the audio data to obtain semantic features of the audio data. The video front-end processing module is responsible for extracting the lip, expression and action information of a client user to obtain visual features of the audio data. The front-end processing module is composed of a 3DCNN and a RESNET18.
[0133] On this basis, the acoustic features, semantic features and visual features of the audio data are cross-modal fused to obtain fused features.
[0134] Based on the fused features, intent and emotion feature extraction is performed on the audio-video data to obtain first intent features and first emotion features. Based on the correlation between the first intent features and the first emotion features, interaction is performed on the first intent features and the first emotion features to obtain second intent features and second emotion features.
[0135] Based on the first intent features and the second emotion features, intent recognition is performed on the audio-video data to obtain an intent recognition result of the audio-video data; and / or,
[0136] Based on the first emotion features and the second intent features, emotion recognition is performed on the audio-video data to obtain an emotion recognition result of the audio-video data.
[0137] S3, based on the intent and / or emotion recognition result of the audio-video data, an image matched with the intent and / or emotion recognition result is generated, and interaction is performed based on the image. The image is sent to a second client, and the image is displayed on the second client.
[0138] The interactive device provided by the present application is described below. The interactive device described below can be referred to in correspondence with the interactive method described above.
[0139] Based on any of the above embodiments, Figure 10 The structure of the interactive device provided by the present application is shown in the figure. The interactive device includes a data acquisition unit 1010, an identification unit 1020 and an interaction unit 1030, wherein,
[0140] The data acquisition unit 1010 is configured to acquire audio-video data.
[0141] The identification unit 1020 is configured to perform intent and / or emotion recognition on the audio-video data based on at least one of the acoustic features, semantic features and visual features of the audio-video data to obtain an intent and / or emotion recognition result of the audio-video data.
[0142] The interaction unit 1030 is configured to determine an image matched with the intention and / or emotion recognition result based on the intention and / or emotion recognition result of the audio and video data, and perform interaction based on the image.
[0143] The interaction device provided by the embodiment of the present application automatically extracts the emotion and / or intention information represented by the audio and video through at least one of the acoustic features, semantic features and visual features of the audio and video, and then obtains an image that meets the user's emotion and / or intention using the emotion and / or intention information, and then performs interaction based on the obtained image. Compared with the prior art in which the user manually selects an image for interaction, the present application can automatically obtain an image based on the audio and video and perform interaction based on the image, which can enhance the user's experience. The obtained image meets the user's emotion and / or intention, and the more diverse and rich images also increase the interest and entertainment.
[0144] Based on any of the above embodiments, the interaction unit is specifically configured to:
[0145] The image prediction module is trained based on the intention and / or emotion label and a sample image matched with the intention and / or emotion label.
[0146] The pre-trained image decoding module is configured to perform feature decoding on the predicted image feature to obtain an image matched with the intention and / or emotion recognition result.
[0147] Based on any of the above embodiments, the device further comprises a module acquisition unit configured to:
[0148] Acquire an initial image prediction module.
[0149] The initial image prediction module is configured to perform image feature prediction on the intention and / or emotion label to obtain a predicted image feature corresponding to the intention and / or emotion label.
[0150] The pre-trained image encoding module is configured to perform feature encoding on the sample image to obtain a sample image feature of the sample image.
[0151] The initial image prediction module is configured to perform parameter iteration based on the similarity between the sample image feature and the predicted image feature to obtain an image prediction module.
[0152] Based on any of the above embodiments, the recognition unit is specifically configured to:
[0153] The recognition unit is configured to perform feature fusion on the acoustic features, semantic features and visual features of the audio and video data to obtain a fused feature.
[0154] perform intent and / or emotion recognition on the audio-video data based on the fusion feature to obtain an intent and / or emotion recognition result of the audio-video data.
[0155] According to any one of the above embodiments, the recognition unit is further specifically configured to:
[0156] perform intent and emotion feature extraction on the audio-video data respectively based on the fusion feature to obtain first intent features and first emotion features;
[0157] interact the first intent features and the first emotion features based on the correlation between the first intent features and the first emotion features to obtain second intent features and second emotion features;
[0158] perform intent recognition on the audio-video data based on the first intent features and the second emotion features to obtain an intent recognition result of the audio-video data; and / or,
[0159] perform emotion recognition on the audio-video data based on the first emotion features and the second intent features to obtain an emotion recognition result of the audio-video data.
[0160] According to any one of the above embodiments, the device further comprises a feature determination unit configured to:
[0161] perform acoustic feature extraction on audio data in the audio-video data based on an encoder in the pre-trained speech recognition model to obtain acoustic features of the audio data;
[0162] perform feature extraction on the acoustic features based on a decoder in the pre-trained speech recognition model to obtain hidden layer features of the audio data;
[0163] perform semantic feature extraction on the hidden layer features of the audio data based on a pre-trained language model to obtain semantic features of the audio data.
[0164] According to any one of the above embodiments, the data acquisition unit is specifically configured to:
[0165] acquire audio-video data of a first client user in video communication;
[0166] The interaction unit is specifically configured to:
[0167] In a case where it is detected that the first client opens an image interaction function, the image is sent to a second client, and the image is displayed on the second client.
[0168] According to any one of the above embodiments, Figure 11 is a structural schematic diagram of a terminal provided by the present application, such as Figure 11As shown, the terminal includes a camera 1110, a microphone 1120 and a processor 1130 connected in sequence:
[0169] The camera 1110 is configured to acquire video data.
[0170] The microphone 1120 is configured to acquire audio data.
[0171] The processor 1130 is configured to perform intent and / or emotion recognition on the audio and video data based on at least one of acoustic features, semantic features and visual features of the audio and video data, to obtain an intent and / or emotion recognition result of the audio and video data, determine an image matching the intent and / or emotion recognition result based on the intent and / or emotion recognition result of the audio and video data, and perform interaction based on the image.
[0172] The terminal provided by the embodiment of the present application automatically extracts emotion and / or intent information represented by audio and video based on at least one of acoustic features, semantic features and visual features of the audio and video, obtains an image matching the emotion and / or intent of a user based on the emotion and / or intent information, and then performs interaction based on the obtained image. Compared with the prior art in which a user manually selects an image for interaction, the present application can automatically obtain an image based on audio and video and perform interaction based on the image, thereby enhancing the experience of the user. The obtained image matches the emotion and / or intent of the user, and more diverse images increase the interest and entertainment.
[0173] Figure 12 An example of a schematic diagram of a physical structure of an electronic device is shown in FIG. 1. Figure 12 As shown, the electronic device can include a processor 1210, a communications interface 1220, a memory 1230 and a communications bus 1240, wherein the processor 1210, the communications interface 1220 and the memory 1230 complete communication with each other through the communications bus 1240. The processor 1210 can invoke a logical instruction in the memory 1230 to execute an interaction method, which includes acquiring audio and video data, performing intent and / or emotion recognition on the audio and video data based on at least one of acoustic features, semantic features and visual features of the audio and video data, obtaining an intent and / or emotion recognition result of the audio and video data, determining an image matching the intent and / or emotion recognition result based on the intent and / or emotion recognition result of the audio and video data, and performing interaction based on the image.
[0174] In addition, the logic instructions in the memory 1230 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0175] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor, so that the computer can execute the interaction method provided by the above-mentioned methods, and the method comprises the following steps: acquiring audio and video data; performing intent and / or emotion recognition on the audio and video data based on at least one of acoustic features, semantic features and visual features of the audio and video data, to obtain an intent and / or emotion recognition result of the audio and video data; determining an image matched with the intent and / or emotion recognition result based on the intent and / or emotion recognition result of the audio and video data, and interacting based on the image.
[0176] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor, so that the computer can execute the interaction method provided by the above-mentioned methods, and the method comprises the following steps: acquiring audio and video data; performing intent and / or emotion recognition on the audio and video data based on at least one of acoustic features, semantic features and visual features of the audio and video data, to obtain an intent and / or emotion recognition result of the audio and video data; determining an image matched with the intent and / or emotion recognition result based on the intent and / or emotion recognition result of the audio and video data, and interacting based on the image.
[0177] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0178] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0179] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An interaction method, characterized in that, The method comprises: acquiring audio and video data; performing intent and / or emotion recognition on the audio and video data based on at least one of acoustic features, semantic features and visual features of the audio and video data, to obtain an intent and / or emotion recognition result of the audio and video data; determining an image matching the intent and / or emotion recognition result based on the intent and / or emotion recognition result of the audio and video data, and performing interaction based on the image; the step of determining the image matching the intent and / or emotion recognition result based on the intent and / or emotion recognition result of the audio and video data comprises: performing image feature prediction on the intent and / or emotion recognition result based on an image prediction module, to obtain a predicted image feature corresponding to the intent and / or emotion recognition result, wherein the image prediction module is trained based on similarity between sample image features and predicted image features, the predicted image feature is determined based on an intent and / or emotion label, and the sample image feature is determined based on a sample image matching the intent and / or emotion label; performing feature decoding on the predicted image feature based on a pre-trained image decoding module, to obtain an image matching the intent and / or emotion recognition result.
2. The interaction method of claim 1, wherein, The acquisition step of the image prediction module comprises: acquiring an initial image prediction module; performing image feature prediction on the intent and / or emotion label based on the initial image prediction module, to obtain a predicted image feature corresponding to the intent and / or emotion label; performing feature encoding on the sample image based on a pre-trained image encoding module, to obtain a sample image feature of the sample image; performing parameter iteration on the initial image prediction module based on similarity between the sample image feature and the predicted image feature corresponding to the intent and / or emotion label, to obtain an image prediction module.
3. The interaction method of claim 1, wherein, The step of performing intent and / or emotion recognition on the audio and video data based on at least one of acoustic features, semantic features and visual features of the audio and video data, to obtain an intent and / or emotion recognition result of the audio and video data, comprises: performing feature fusion on acoustic features, semantic features and visual features of the audio and video data, to obtain fused features; performing intent and / or emotion recognition on the audio and video data based on the fused features, to obtain an intent and / or emotion recognition result of the audio and video data.
4. The interaction method of claim 3, wherein, The step of performing intent and / or emotion recognition on the audio and video data based on the fused features, to obtain an intent and / or emotion recognition result of the audio and video data, comprises: performing intent and emotion feature extraction on the audio and video data based on the fused features, respectively, to obtain first intent features and first emotion features; performing interaction on the first intent features and the first emotion features based on correlation between the first intent features and the first emotion features, to obtain second intent features and second emotion features; performing intent recognition on the audio and video data based on the first intent features and the second emotion features, to obtain an intent recognition result of the audio and video data; and / or, Based on the first emotional feature and the second intention feature, the audio and video data is subjected to emotional recognition, and an emotional recognition result of the audio and video data is obtained.
5. The interaction method of claim 1, wherein, The determination of the acoustic feature and the semantic feature of the audio and video data comprises: Based on an encoder in a pre-trained speech recognition model, acoustic feature extraction is performed on the audio data in the audio and video data, and an acoustic feature of the audio data is obtained. Based on a decoder in the pre-trained speech recognition model, feature extraction is performed on the acoustic feature, and a hidden layer feature of the audio data is obtained. Based on a pre-trained language model, semantic feature extraction is performed on the hidden layer feature of the audio data, and a semantic feature of the audio data is obtained.
6. The interaction method according to any one of claims 1-5, characterized in that, The audio and video data comprises: Audio and video data of a first client user during video communication is obtained. The interaction based on the image comprises: In a case where it is detected that the first client opens an image interaction function, the image is sent to a second client, and the image is displayed on the second client.
7. An interactive device, characterized by Comprise: A data acquisition unit is configured to acquire audio and video data. An identification unit is configured to perform intention and / or emotion identification on the audio and video data based on at least one of an acoustic feature, a semantic feature, and a visual feature of the audio and video data, and obtain an intention and / or emotion identification result of the audio and video data. An interaction unit is configured to determine an image matching the intention and / or emotion identification result based on the intention and / or emotion identification result of the audio and video data, and perform interaction based on the image. The determination of the image matching the intention and / or emotion identification result based on the intention and / or emotion identification result of the audio and video data comprises: An image feature prediction is performed on the intention and / or emotion identification result based on an image prediction module to obtain a predicted image feature corresponding to the intention and / or emotion identification result, the image prediction module is trained based on a similarity between a sample image feature and a predicted image feature, the predicted image feature is determined based on an intention and / or emotion label, and the sample image feature is determined based on a sample image matching the intention and / or emotion label. A feature decoding is performed on the predicted image feature based on a pre-trained image decoding module to obtain an image matching the intention and / or emotion identification result.
8. A terminal, characterized by comprising: Comprise a camera, a microphone and a processor connected in sequence: The camera is configured to acquire video data. The microphone is configured to acquire audio data. The processor is configured to perform intention and / or emotion identification on the audio and video data based on at least one of an acoustic feature, a semantic feature, and a visual feature of the audio and video data, obtain an intention and / or emotion identification result of the audio and video data, perform image feature prediction on the intention and / or emotion identification result based on an image prediction module, obtain a predicted image feature corresponding to the intention and / or emotion identification result, perform feature decoding on the predicted image feature based on a pre-trained image decoding module, obtain an image matching the intention and / or emotion identification result, and perform interaction based on the image. The image prediction module is trained based on similarity between a sample image feature and a predicted image feature, the predicted image feature is determined based on an intent and / or an emotion label, and the sample image feature is determined based on a sample image matching the intent and / or the emotion label.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the interaction method of any one of claims 1 to 6 when executing the computer program.
10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the interaction method of any one of claims 1 to 6.
Citation Information
Patent Citations
Man-machine interaction method and man-machine interaction device
CN110110169A
Learning apparatus and method for creating emotion expression video and apparatus and method for emotion expression video creation
US20210406554A1
Apparatus and method for emotional content services on telecommunication devices, apparatus and method for emotion recognition therefor, and apparatus and method for generating and matching the emotional content using same
WO2013027893A1