Method, device and storage medium for interaction with a virtual robot
By collecting multimodal data in real time and generating response text through virtual robots, the problem of insufficient interaction in live streaming rooms has been solved, enabling intelligent interaction and personalized services, thereby improving the activity level of live streaming rooms and the audience experience.
Patent Information
- Application Number
- CN202310736256.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-20
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-06-20
AI Technical Summary
Existing live streaming assistant robots cannot effectively interact with the streamer based on the content of the live stream, lack personalized services, resulting in insufficient activity in the live stream, and the threshold for streamers to join is relatively high.
The system collects multimodal data in real time through virtual robots, generates response text, and performs interactive operations, including chatting with the host, commenting, teasing, and giving rewards as thanks. Combined with AI technology, it automatically selects sound effects and commands, reducing the difficulty of operation for the host.
It increased the interactivity and enthusiasm of the live streamers, improved the audience experience, and reduced the operational difficulty and access barrier for the streamers.
Smart Images

Figure CN116756285B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a virtual robot interaction method, device and storage medium. BACKGROUND
[0002] With the rapid development of technology and the improvement of people's living standards, network live broadcast has gradually become one of the important ways for people to entertain their lives, and is deeply loved by young people. For live broadcast, cold field is a situation that greatly affects the mood of the host and the experience of the audience. Traditional live broadcast assistant robots can only execute some simple instructions, such as playing sound effects, reminding to follow, etc., and cannot effectively interact with the host according to the content of the live broadcast room, lack the ability of intelligent interaction and personalized service, and cannot really help the host improve the interaction and liveliness of the live broadcast room. In addition, the traditional robot also needs to manually configure various rules and instructions, and the access threshold of the host is high. SUMMARY
[0003] The embodiment of the present application provides a virtual robot interaction method, device and storage medium, to solve the problem of insufficient liveliness in the live broadcast room, and to interact very intelligently with the host through the virtual robot to improve the liveliness of the live broadcast room.
[0004] In a first aspect, the embodiment of the present application provides a virtual robot interaction method, which comprises:
[0005] In response to a virtual robot interaction request triggered by a user, real-time collection of multi-modal data of a live broadcast room of the user is performed;
[0006] Determination of text information corresponding to the multi-modal data is performed, and the text information is used to uniformly represent feature information of the multi-modal data;
[0007] Generation of response text according to the text information is performed;
[0008] Determination of interaction information based on the response text is performed, and control of the virtual robot to perform an interaction operation based on the interaction information is performed.
[0009] In a second aspect, the embodiment of the present application provides a virtual robot interaction device, which comprises:
[0010] A response module is configured to, in response to a virtual robot interaction request triggered by a user, real-time collection of multi-modal data of a live broadcast room of the user is performed;
[0011] A determination module is configured to determine text information corresponding to the multi-modal data, and the text information is used to uniformly represent feature information of the multi-modal data;
[0012] A generation module is configured to generate response text according to the text information;
[0013] The execution module is configured to determine interaction information based on the response text, and control the virtual robot to perform an interaction operation based on the interaction information.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a communication interface; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the interaction method of the virtual robot according to the first aspect.
[0015] In a fourth aspect, an embodiment of the present application provides a non-transitory machine readable storage medium, which stores executable code, and when the executable code is executed by a processor of an electronic device, the processor can at least implement the interaction method of the virtual robot according to the first aspect.
[0016] In the interaction scheme of the virtual robot provided by the embodiment of the present application, when processing the virtual robot interaction request triggered by the user, the different modal types of data of the live broadcast room of the user can be collected in real time in response to the virtual robot interaction request triggered by the user, such as live broadcast video data, live broadcast audio data, barrage information, and reward information. Then, the collected data is processed to convert the multi-modal data into text information in a unified form, and the text information is used to uniformly represent the feature information of the multi-modal data. Then, the response text is generated according to the text information. Based on the response text, the interaction information is determined, and the virtual robot is controlled to perform the interaction operation based on the interaction information.
[0017] In the above scheme, by analyzing and processing the multi-modal data of the live broadcast room, the information of the live broadcast room is fully mined, so that the virtual robot can automatically and intelligently interact with the user in combination with various live broadcast room information. Not only can the user operation difficulty be reduced, but also the live broadcast room activity can be improved, and the enthusiasm of the anchor and the audience experience are greatly improved. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0019] Figure 1 A flowchart of an interaction method of a virtual robot provided by an embodiment of the present application is provided.
[0020] Figure 2 A flowchart for determining interaction information based on response text provided by an embodiment of the present application is provided.
[0021] Figure 3 A flowchart for determining an instruction operation corresponding to the instruction information is provided for the embodiment of the present application.
[0022] Figure 4 A schematic diagram for determining a virtual robot interaction process in a cloud service mode is provided for the embodiment of the present application.
[0023] Figure 5 A structural schematic diagram of a virtual robot interaction device is provided for the embodiment of the present application.
[0024] Figure 6 A structural schematic diagram of an electronic device is provided for the embodiment. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application. In addition, the step sequence in each method embodiment below is only an example, not a strict limitation.
[0026] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or authorized by all parties, and the collection, use, and processing of related data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and provide corresponding operation portals for users to choose authorization or refusal.
[0027] With the rapid development of Internet technology, network live streaming as a new technical field enters the public view. Users can watch the wonderful performance of the anchor in the live streaming room on their own terminals and can interact with the anchor in real time. When the interaction in the anchor's live streaming room is less and the atmosphere is poor, the anchor often needs to mobilize the atmosphere or mobilize the atmosphere through the anchor's assistant robot.
[0028] The existing live streaming assistant robot can only execute some simple instructions, such as playing sound effects, reminding to follow, etc., and cannot effectively interact with the anchor according to the content of the live streaming room, lacks the ability of intelligent interaction and personalized service, and cannot really help the anchor improve the interaction and liveliness in the live streaming room. In addition, the traditional robot also needs to manually configure various rules and instructions, and the anchor's access threshold is high, thereby affecting the use of the anchor.
[0029] To solve the above technical problems, the embodiment of the present application provides a new interactive scheme of a virtual robot. In the interactive scheme, various data sources of a user live room are combined, live room information is fully mined, interactive information of the virtual robot is determined, and real-time interaction with the user is performed. For example, the virtual robot can chat with the anchor user in combination with live content, and can also comment on, make fun of, and boast about the live content according to user settings; at the same time, the virtual robot can also naturally thank and respond to the pop-up in the chat according to audience feedback, that is, the virtual robot can very intelligently interact with the anchor, which can improve the interaction and activity level of the live room.
[0030] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The embodiments described below and the features in the embodiments can be combined with each other without conflict.
[0031] Figure 1 A flowchart of an interactive method of a virtual robot according to an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the embodiment provides an interactive method of a virtual robot, and the execution subject of the method can be a server. It can be understood that the server can be implemented as software or a combination of software and hardware. Specifically, the method includes the following steps: Figure 1
[0032] 101. In response to a virtual robot interaction request triggered by a user, multi-modal data of a user live room is collected in real time.
[0033] 102. Text information corresponding to the multi-modal data is determined, and the text information is used to uniformly represent feature information of the multi-modal data.
[0034] 103. Response text is generated according to the text information.
[0035] 104. Interaction information is determined based on the response text, and a virtual robot performs an interactive operation based on the interaction information.
[0036] In the embodiment of the present application, when a user triggers a virtual robot interaction request, a server collects multi-modal data of the user live room in real time in response to the virtual robot interaction request triggered by the user. The virtual robot interaction request can include a user identifier, and the multi-modal data of the user live room is collected in real time according to the user identifier. Optionally, the multi-modal data can include live video data, live audio data, reward information, and pop-up information of different modal types. The specific types and quantities of data included can be set according to actual needs.
[0037] In order to facilitate subsequent analysis and processing of the collected multi-modal data, the data of various different modal types can be converted into a unified form first. Then, after obtaining the multi-modal data of the user live room, the multi-modal data can be processed to determine the text information corresponding to the multi-modal data. The text information is used to uniformly represent the feature information of the multi-modal data.
[0038] For example, the multi-modal data can include video data, audio data, reward information, and barrage information. The audio data can be converted into corresponding text information, where the text information can be used to describe the live content included in the live video data. The audio data is converted into corresponding text information, which can be used to describe the text content included in the audio data. The reward information is converted into corresponding text information, which can be used to describe the audience reward situation and audience feedback information. The barrage information is converted into corresponding text information, which can be used to describe the audience feedback information.
[0039] In an optional embodiment, if the collected multi-modal data includes user audio data, the text information corresponding to the user audio data can be determined by a speech recognition model. Specifically, the user audio data is input into the speech recognition model to obtain the text information corresponding to the user audio data. For example, the chat audio of the host is converted into chat text by using the speech recognition model to obtain the text of the host's speech.
[0040] The speech recognition model mainly includes an encoder and a decoder. The encoder is mainly used to convert the user's audio data into a vector representation for speech recognition. The decoder is mainly used to complete the speech-to-text recognition to recognize all the text spoken by the user in the audio data, and finally output the speech recognition result corresponding to the user to obtain the text information corresponding to the user audio data. Optionally, the encoder can include a plurality of cascaded encoders, and each encoder can include two sub-layers: an attention layer and a feedforward neural network layer. The decoder can include a plurality of cascaded decoders, and each decoder can include an attention layer and a feedforward neural network layer. The number of encoders included in the encoder can be set according to actual needs, and similarly, the number of encoders included in the decoder can be set according to actual needs, which is not limited here. In addition, the attention layer in the decoder here can include a self-attention layer and an attention layer.
[0041] In an optional embodiment, if the collected multi-modal data includes live video data, the text information corresponding to the live video data can be determined through an image recognition model. Specifically, the live video data is input into the image recognition model to obtain the text information corresponding to the live video data, and the text information is used to describe the live content included in the live video data. Alternatively, the collected live video data is processed by taking screenshots at a preset period, and the obtained multiple images are input into the image recognition model to obtain the text information corresponding to the live video data. Alternatively, the user live content is taken screenshots at a preset period to obtain multiple images, and the multiple images are input into the image recognition model to obtain the text information corresponding to the live video data. For example, the live of the anchor is taken screenshots regularly at a preset period, and a visual language pre-training model (BLIP-2 model) is used to generate a textual description corresponding to the screenshots to obtain a live video content description text.
[0042] The image recognition model mainly includes an image encoder, a text encoder, an image-text encoder, and an image-text decoder. The image encoder is mainly used to extract image feature information and convert the live video data of the user into a vector representation for image recognition. The text encoder is mainly used to extract text feature information and convert the vector representation of the image recognition into a vector representation for text recognition. The image-text encoder is used to predict whether the image-text pair is a positive match or a negative match. The image-text decoder is used to generate a text description of a given image. The pre-training tasks of BLIP mainly include: contrastive learning of output image features and text features, judging whether the image and text are consistent, and text generation. The three pre-training tasks are trained together, which can more fully utilize the collected multi-modal image-text data, and also enables the BLIP model to simultaneously satisfy the image-text understanding task and the image-text generation task.
[0043] In an optional embodiment, if the collected multi-modal data includes audience interaction data, the text information corresponding to the audience interaction data can be determined through a first language recognition model. Specifically, the audience interaction data is input into the first language recognition model to obtain the text information corresponding to the audience interaction data, and the text information is used to describe the audience feedback information. The audience interaction data can include audience reward information and bullet screen information. The first language recognition model can be a generative pre-training Transformer model (GTP model) that adopts unsupervised pre-training and supervised model fine-tuning. For example, using the GTP model, the audience reward gift information and the audience bullet screen information are converted into a text description, and the audience feedback information is summarized to obtain an audience feedback summary text.
[0044] From the above description, it can be known that the text information can include anchor speaking text, live video content description text, and audience feedback summary text, which can reflect the current live room information of the user. Therefore, the response text corresponding to the virtual robot can be determined according to the obtained text information.
[0045] Specifically, in the embodiment of the present application, after obtaining the text information corresponding to the multi-modal data, the text information is analyzed and processed to determine the response text corresponding to the virtual robot. The response text can include dialogue information and instruction information. That is, the current response mode and current response content corresponding to the virtual robot can be determined according to the current live situation. If it is determined that the virtual robot currently performs dialogue response according to the current obtained text information, then the specific dialogue information corresponding to the current dialogue response is determined. If it is determined that the virtual robot currently performs instruction response according to the current obtained text information, then the specific instruction information corresponding to the current instruction response is determined. In addition, in actual application, if it is determined that the virtual robot needs to perform dialogue response and instruction response at the same time according to the current obtained text information, then the specific dialogue information corresponding to the current dialogue response and the specific instruction information corresponding to the current instruction response can be determined.
[0046] The dialogue information can include real-time dialogue content with the anchor, comments on the live content, teasing and boasting, thanks to the audience and response to the scroll information, etc. For example, after processing the collected multi-modal data, the obtained text information is “How is the weather today?”. Then the dialogue information “It's not bad today” can be generated.
[0047] The instruction information can include sound effect playing instruction, volume adjustment instruction, background music switching instruction, etc. For example, after processing the collected multi-modal data, the obtained text information is “The sound of the music played is too loud”. Then the instruction information “Turn down the volume” can be generated.
[0048] In an optional embodiment, the response text corresponding to the text information can be determined by a second language recognition model. Specifically, the text information is input into the second language recognition model to obtain dialogue information and / or instruction information corresponding to the text information. The dialogue information is used to instruct the virtual robot to perform an operation corresponding to the dialogue information, and the instruction information is used to instruct the virtual robot to perform an operation corresponding to the instruction information. In addition, the second language recognition model can be a generative pre-trained transformer model (GPT model) that adopts unsupervised pre-training and supervised model fine-tuning. For example, the GPT model is used to analyze and process the text information to generate the response text. Specifically, when the GPT model analyzes the text information, the anchor speaking text is used as the main text, and the live video content description text and the audience feedback summary text are used as auxiliary texts to generate real-time dialogue information and / or instruction information. That is, in the language recognition model, the anchor speaking text has a high weight, and the live video content description text and the audience feedback summary text have a low weight.
[0049] After determining the response text corresponding to the current virtual robot, the interaction information is determined based on the response text, and the virtual robot is controlled to perform an interaction operation based on the interaction information. Different response texts correspond to different interaction information. For example, when the response text is dialogue information, the dialogue information is converted into speech by using a speech synthesis technology, so that the virtual robot performs speech playing. When the response text is instruction information, the instruction information is converted into an instruction operation, so that the virtual robot performs the instruction operation. In this way, the virtual robot can combine the live content and chat with the anchor in a personalized manner, can comment, joke and boast on the live content according to the anchor setting, can naturally thank and respond to the pop-up in the chat according to the audience feedback, and can improve user experience and user stickiness. At the same time, the virtual robot can also automatically generate an operation instruction according to the live content or the anchor dialogue, so that the virtual robot performs the corresponding operation, and reduces the operation difficulty of the anchor.
[0050] In addition, the voice style played by the virtual robot can be set, so that the virtual robot can chat with the anchor or interact with the audience in different voice styles, so as to increase the interest of the live broadcast and attract more audiences to watch the live broadcast of the user.
[0051] In the embodiments of the present application, the multi-modal data of the live room is analyzed and processed to fully mine the live room information, so that the virtual robot can automatically and intelligently interact with the user in combination with various live room information. Not only can the operation difficulty of the user be reduced, but also the live room activity can be improved, and the live broadcast enthusiasm of the anchor and the user experience can be greatly improved.
[0052] The execution subject of the virtual robot interaction scheme introduced in the above embodiments is the server. In actual application, the execution subject of the virtual robot interaction method can also be a virtual robot, which can be implemented as software or a combination of software and hardware, and the execution subject is not limited, and can be set according to actual needs. Then, if the execution subject of the method is a virtual robot, after the user virtual robot interaction request, the virtual robot responds to the virtual robot interaction request triggered by the user, and collects the multi-modal data of the live room of the user in real time. The multi-modal data is processed to determine the text information corresponding to the multi-modal data. The text information is used to uniformly represent the feature information of the multi-modal data. Then, the text information is analyzed and processed to determine the response text corresponding to the virtual robot. Then, based on the response text, the interaction information is determined, and the corresponding interaction operation is performed based on the interaction information.
[0053] In the embodiments of the present application, the specific implementation process involved can refer to the above embodiments, and will not be repeated here.
[0054] The above embodiments introduce the specific implementation process of determining the interaction operation performed by the virtual robot. In actual application, in order to increase the interactivity and interest of live broadcast, the user can set the person information of the robot, generate different voice information based on different person information, and combine the voice information with the response text to generate the interaction information. Figure 2 The processing process of determining the interaction information based on the response text and controlling the virtual robot to perform the interaction operation is exemplarily illustrated.
[0055] Figure 2 A flowchart for determining interaction information based on response text is provided in the embodiments of the present application; as shown in Figure 2 The response text includes dialogue information, and the embodiments of the present application provide a specific implementation method for determining interaction information based on response text, which includes the following steps:
[0056] 201, determine the person information of the virtual robot.
[0057] 202, determine the voice information corresponding to the dialogue information based on the person information, and control the virtual robot to play the voice information.
[0058] In the embodiments of the present application, when the generated response text is dialogue information, the interaction information can be generated in combination with the person information of the virtual robot. Specifically, first, the person information of the virtual robot is determined, the voice information corresponding to the dialogue information is determined based on the person information, and the virtual robot is controlled to play the voice information.
[0059] The user can pre-set the person information of the virtual robot according to the live content, or the server can set the person information of the virtual robot according to the current live content. The person information can include age, personality, speaking style, region, etc. Different styles of dialogue feedback will be generated according to the set person information, so that the virtual robot can play voice information in different style types.
[0060] In the embodiment of the application, the person information of the virtual robot is determined, and then the voice information corresponding to the dialogue information is determined based on the person information, and the virtual robot is controlled to play the voice information, thereby increasing the interactivity and interest of the live broadcast.
[0061] The above embodiment introduces an implementation manner of determining interactive information based on generated dialogue information. However, in actual application, instruction information can also be generated according to text information, and the interactive information is determined based on the instruction information, so that the virtual robot performs corresponding instruction operation. Specifically, the instruction operation corresponding to the instruction information is determined, and the virtual robot is controlled to perform the instruction operation. The operation information can include sound effect playing instruction, volume adjustment instruction, background music switching instruction, etc.
[0062] In the traditional live broadcast, the host usually needs to manually trigger sound effect playing instructions, volume adjustment instructions, background music switching instructions, etc. However, in some scenarios, such as dance scenarios, outdoor scenarios, game scenarios, etc., the host is not convenient to trigger the operation instruction, and cannot play the atmosphere sound effect in time, adjust the volume, etc. However, by using the interactive method of the virtual robot provided in the embodiment of the application, the instruction information can be automatically generated according to the live content, so that the virtual robot can automatically perform the corresponding instruction operation, and the user no longer needs to trigger, which can reduce the user operation and free the user's hands.
[0063] In addition, in actual application, the process of manually selecting sound effects has subjective factors, and the suitability of the sound effects cannot be guaranteed, and the best effect cannot be achieved. In order to solve the above problems, the AI technology is adopted in the embodiment of the application, and the matched sound effect playing is automatically selected in combination with multi-modal data, so that more efficient and intelligent live broadcast content atmosphere control is achieved. In order to facilitate the process of determining the matched preset sound effect, the accompanying drawings are combined Figure 3 The specific process of determining the preset sound effect in combination with multi-modal data to control the virtual robot to automatically play the preset sound effect is exemplarily described.
[0064] Figure 3 A flowchart for determining the instruction operation corresponding to the instruction information is provided for the embodiment of the application; as shown in Figure 3 The instruction information includes a sound effect playing instruction, and the embodiment of the application provides a specific implementation manner of determining the instruction operation corresponding to the instruction information, which includes the following steps:
[0065] 301、acquire live video data and live audio data in a current period.
[0066] 302、determine video features corresponding to the live video data and audio features corresponding to the live audio data.
[0067] 303、based on the audio features and the video features, determine a live scene, a live style, and a user emotion corresponding to a current live room.
[0068] 304、based on the live scene, the live style, and the user emotion, determine a preset sound effect matched with the current live from a sound effect feature library, and control the virtual robot to play the preset sound effect.
[0069] In actual application, after determining the corresponding instruction information according to the text information, the instruction operation corresponding to the instruction information is determined. It should be noted that if the specific instruction information determined is a sound effect playing instruction, the preset sound effect to be played by the virtual robot can be determined from the sound effect feature library before determining the instruction operation corresponding to the sound effect playing instruction.
[0070] Among them, the server can pre-construct a sound effect feature library. Specifically, first, a plurality of preset sound effects are acquired, and the live scene, user emotion, live style, live duration, and the like suitable for each preset sound effect are determined. An audio feature extraction model is used to extract the audio features of each preset sound effect to obtain the audio features corresponding to each sound effect. Then, the combination information of the audio features and the live scene, user emotion, live style, and live duration is constructed, and based on the corresponding relationship, the sound effect feature library is created.
[0071] Then, the preset sound effect matched with the current live content can be determined from the sound effect feature library in combination with the live content, and the virtual robot is controlled to automatically play the preset sound effect. Specifically, first, live video data and live audio data in a current period are acquired. Since the live content of the user may change every moment, in order to better determine the preset sound effect that meets the current live content of the user and makes the selected preset sound effect achieve the best effect, the live video data and live audio data of the user's live room in the current period are acquired.
[0072] Then, video features corresponding to the live video data and audio features corresponding to the live audio data are determined. In an optional embodiment, the collected live video data can be processed using a 3D convolutional neural network to extract time features and space features in the video data, and the extracted time features and space features are determined as the video features corresponding to the live video data. The live audio data can be converted into text information using a speech recognition model, and a Transformer model is used for text feature extraction to obtain the audio features corresponding to the live audio data.
[0073] After the audio features and the video features are determined, based on the audio features and the video features, a live scene, a live style, and a user emotion corresponding to the current live room are determined. Then, based on the live scene, the live style, and the user emotion, a preset sound effect that matches the current live is determined from a sound effect feature library. The determined preset sound effect is more matched with the current live content of the user to achieve the best effect, so that the live atmosphere is more active and lively, thereby improving the experience of the audience user. Finally, the virtual robot is controlled to play the preset sound effect.
[0074] In the embodiments of the present application, by extracting features from multi-modal data and analyzing feature information and emotions, comprehensive analysis of live content is realized, the atmosphere and emotions of the live content can be accurately grasped, appropriate preset sound effects can be selected for playing, and the quality and viewing experience of the live content are improved. Moreover, the preset sound effects available for the user are expanded from dozens to thousands, and the hands of the anchor are freed, so that more possibilities are provided for the live.
[0075] In the embodiments of the present application, the specific implementation process involved can refer to the content in the above embodiments, which will not be described here.
[0076] In addition, the interaction method of the virtual robot provided by the embodiments of the present application can be executed in the cloud. The cloud can be deployed with a plurality of computing nodes (cloud servers), each of which has computing, storage, and other processing resources. In the cloud, a plurality of computing nodes can be organized to provide a certain service. Of course, one computing node can also provide one or more services. The cloud can provide a service interface for the service, and the user can call the service interface to use the corresponding service.
[0077] For the scheme provided by the embodiments of the present application, the cloud can provide a service interface for the interaction service of the virtual robot, and the user can call the service interface through the terminal device to trigger a virtual robot interaction service request to the cloud. The request includes a user identifier, and the cloud determines a computing node responding to the request and uses the processing resources in the computing node to execute the following steps:
[0078] In response to a virtual robot interaction request triggered by a user, multi-modal data of a live room of the user is collected in real time;
[0079] Text information corresponding to the multi-modal data is determined, the text information being used to uniformly represent feature information of the multi-modal data;
[0080] A response text is generated according to the text information.
[0081] Interaction information is determined based on the response text, and an interaction operation is controlled to be performed by the virtual robot based on the interaction information.
[0082] The foregoing execution process can refer to the related descriptions in the other embodiments described above, and thus will not be described here.
[0083] For ease of understanding, exemplary descriptions are made in conjunction with Figure 4 A user can invoke an interaction service of a virtual robot through a terminal device E1 as shown in Figure 4 to collect multi-modal data of a live room of the user in real time, and to analyze and process the multi-modal data to obtain an interaction operation corresponding to the virtual robot. A service interface through which the user invokes the service includes a software development kit (SDK), an application programming interface (API), and the like. Figure 4 The case of an API interface is shown in Figure 4 In the cloud, as shown in the interaction service of the virtual robot is provided by a service cluster E2, and the service cluster E2 includes at least one computing node. After receiving the request, the service cluster E2 performs the steps in the foregoing embodiments to determine an interaction operation to be performed by the virtual robot, and feeds back to the terminal device E1.
[0084] The virtual robot interaction apparatus of one or more embodiments of the present application will be described in detail below. Those skilled in the art can understand that these apparatuses can all be configured by using commercially available hardware components through the steps taught by the present solution.
[0085] Figure 5 A structural schematic diagram of a virtual robot interaction apparatus provided by an embodiment of the present application is shown in Figure 5 The apparatus includes a response module 11, a determination module 12, a generation module 13, and an execution module 14.
[0086] The response module 11 is configured to collect multi-modal data of a live room of a user in real time in response to a virtual robot interaction request triggered by the user.
[0087] The determining module 12 is configured to determine text information corresponding to the multi-modal data, and the text information is used to uniformly represent feature information of the multi-modal data.
[0088] The generating module 13 is configured to generate response text according to the text information.
[0089] The executing module 14 is configured to determine interaction information based on the response text, and control the virtual robot to perform an interaction operation based on the interaction information.
[0090] Optionally, the multi-modal data includes user audio data, and the determining module 12 can be specifically configured to input the user audio data into a speech recognition model to obtain text information corresponding to the user audio data.
[0091] Optionally, the multi-modal data includes live video data, and the determining module 12 can be specifically configured to input the live video data into an image recognition model to obtain text information corresponding to the live video data, and the text information is used to describe live content included in the live video data.
[0092] Optionally, the multi-modal data includes audience interaction data, and the determining module 12 can be specifically configured to input the audience interaction data into a first language recognition model to obtain text information corresponding to the audience interaction data, and the text information is used to describe audience feedback information.
[0093] Optionally, the generating module 13 can be specifically configured to input the text information into a second language recognition model to obtain dialogue information and / or instruction information corresponding to the text information, the dialogue information is used to instruct the virtual robot to perform an operation corresponding to the dialogue information, and the instruction information is used to instruct the virtual robot to perform an operation corresponding to the instruction information.
[0094] Optionally, the executing module 14 can be specifically configured to determine person information of the virtual robot, determine voice information corresponding to the dialogue information based on the person information, and control the virtual robot to play the voice information.
[0095] Optionally, the executing module 14 can be specifically configured to determine an instruction operation corresponding to the instruction information, and control the virtual robot to perform the instruction operation.
[0096] Optionally, the instruction information includes sound effect playing instructions, and the execution module 14 can be specifically configured to: acquire live video data and live audio data in a current period; determine video features corresponding to the live video data and audio features corresponding to the live audio data; determine a live scene, a live style and user emotions corresponding to a current live room based on the audio features and the video features; and determine preset sound effects matched with the current live from a sound effect feature library based on the live scene, the live style and the user emotions, and control the virtual robot to play the preset sound effects.
[0097] Figure 5 The apparatus can perform the steps in the voice recognition method in the foregoing embodiments, and refer to the descriptions in the foregoing embodiments for detailed execution processes and technical effects, which will not be described here again.
[0098] The embodiment of the present application further provides an electronic device, such as Figure 6 As shown, the electronic device can include a processor 21, a memory 22 and a communication interface 23. The memory 22 stores executable codes, and when the executable codes are executed by the processor 21, the processor 21 implements the interactive method of the virtual robot as in the foregoing embodiments.
[0099] In addition, the embodiment of the present application provides a non-transitory machine readable storage medium, which stores executable codes, and when the executable codes are executed by the processor of the electronic device, the processor can at least implement the interactive method of the virtual robot as in the foregoing embodiments.
[0100] The apparatus embodiments described above are only schematic, and the units described as separate components can or can not be physically separate. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. Those skilled in the art can understand and implement without creative labor.
[0101] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of an appropriate general hardware platform, and of course can also be implemented by means of a combination of hardware and software. Based on such understanding, the above technical solutions can be embodied in the form of a computer program product, and the present application can be implemented in the form of a computer program product containing computer usable program codes in one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.).
[0102] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An interaction method of a virtual robot, characterized by, The method comprises: in response to a virtual robot interaction request triggered by a user, collecting multi-modal data of a live room of the user in real time; the multi-modal data comprises at least one of user audio data, live video data and audience interaction data; determining text information corresponding to the multi-modal data, the text information being used to uniformly represent feature information of the multi-modal data; generating response text according to the text information; determining interaction information based on the response text, and controlling the virtual robot to perform an interaction operation based on the interaction information; if the multi-modal data comprises live video data, the determining of the text information corresponding to the multi-modal data comprises: inputting the live video data into an image recognition model to obtain text information corresponding to the live video data, the text information being used to describe live content included in the live video data; the text information comprises anchor speaking text, live video content description text and audience feedback summary text; the generating of the response text according to the text information comprises: inputting the text information into a second language recognition model to obtain dialogue information and / or instruction information corresponding to the text information, the dialogue information being used to instruct the virtual robot to perform an operation corresponding to the dialogue information, and the instruction information being used to instruct the virtual robot to perform an operation corresponding to the instruction information; in the second language recognition model, the weight of the anchor speaking text is higher than the weight of the live video content description text, and is higher than the weight of the audience feedback summary text; in the case where the instruction information corresponding to the text information is obtained, the instruction information comprises a sound effect playing instruction; the determining of the interaction information based on the response text, and the controlling of the virtual robot to perform an interaction operation based on the interaction information, comprises: obtaining live video data and live audio data in a current period; determining video features corresponding to the live video data and audio features corresponding to the live audio data; based on the audio features and the video features, determining a live scene, a live style and a user emotion corresponding to a current live room; based on the live scene, the live style and the user emotion, determining a preset sound effect matched with the current live from a sound effect feature library, and controlling the virtual robot to play the preset sound effect.
2. The method of claim 1, wherein, if the multi-modal data comprises user audio data, the determining of the text information corresponding to the multi-modal data comprises: inputting the user audio data into a speech recognition model to obtain text information corresponding to the user audio data.
3. The method of claim 1, wherein, if the multi-modal data comprises audience interaction data, the determining of the text information corresponding to the multi-modal data comprises: inputting the audience interaction data into a first language recognition model to obtain text information corresponding to the audience interaction data, the text information being used to describe audience feedback information.
4. The method of claim 1, wherein, in the case where the dialogue information corresponding to the text information is obtained, the determining of the interaction information based on the response text, and the controlling of the virtual robot to perform an interaction operation based on the interaction information, comprises: determining persona information of the virtual robot; Based on the human setting information, determine voice information corresponding to the conversation information, and control the virtual robot to play the voice information.
5. An electronic device, comprising: The method comprises the steps of: The memory, the processor, the communication interface, wherein the memory has executable code stored thereon, and when the executable code is executed by the processor, the processor executes the interactive method of the virtual robot as claimed in any one of claims 1 to 4.
6. A non-transitory machine-readable storage medium, characterized in that, The non-transitory machine-readable storage medium has executable code stored thereon, and when the executable code is executed by the processor of the electronic device, the processor executes the interactive method of the virtual robot as claimed in any one of claims 1 to 4.
Citation Information
Patent Citations
Virtual robot multi-mode interaction method and system applied to video live-broadcasting platform
CN107423809A