Virtual image control method, device, apparatus, and medium

By using speech recognition and classification models to predict and control the behavior and voice output of virtual avatars, the problem of user interaction methods disrupting immersion in virtual scenes is solved, resulting in a more immersive user experience.

CN115345969BActive Publication Date: 2026-04-14BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2022-08-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, the way users interact in virtual scenes requires additional operations and disrupts the immersive experience, affecting users' sense of immersion and emotional expression.

Method used

By acquiring the user's voice audio, performing speech recognition and classification model prediction, determining the target behavior label, and controlling the virtual avatar to perform corresponding behaviors, the virtual avatar's actions and voice output are realized by combining the user's action and emotion categories.

Benefits of technology

It enhances users' perception and emotional expression in virtual scenarios, providing a more immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115345969B_ABST
    Figure CN115345969B_ABST
Patent Text Reader

Abstract

The present disclosure provides a virtual image control method, device, equipment and medium, the artificial intelligence technical field, specifically relates to image processing, deep learning and other technical fields, especially relates to 3D vision, virtual reality, augmented reality and metaverse and other scenes. The implementation scheme is: receiving the voice audio of the user;Based on the speech recognition result of the voice audio, the target behavior label is obtained by prediction;Based on the first action of the virtual image of the user and the first behavior corresponding to the target behavior label, the target behavior is determined;And control the virtual image of the user to make target behavior.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as image processing and deep learning, and particularly to scenarios such as 3D vision, virtual reality, augmented reality, and metaverse. Specifically, it relates to a method, device, electronic device, computer-readable storage medium, and computer program product for controlling a virtual image. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] The Metaverse is a new type of internet application and social form that integrates multiple new technologies, creating a blend of the virtual and the real. It provides immersive experiences based on extended reality technology, generates a mirror image of the real world based on digital twin technology, and builds an economic system based on blockchain technology. It closely integrates the virtual and real worlds in terms of economic, social, and identity systems, and allows each user to produce content and edit the world.

[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for controlling a virtual avatar.

[0006] According to one aspect of this disclosure, a method for controlling a virtual avatar is provided, comprising: receiving a user's voice audio; predicting a target behavior label based on the speech recognition result of the voice audio, wherein the target behavior label is used to obtain a first behavior corresponding to the target behavior label from a preset behavior set, the preset behavior set including at least one preset behavior and at least one behavior label corresponding to the at least one preset behavior respectively; determining a target behavior based on a first action of the user's virtual avatar and the first behavior corresponding to the target behavior label, wherein the first action indicates the current action of the user's virtual avatar in a virtual scene, the target behavior including at least one of a target action and the action of emitting target speech; and controlling the user's virtual avatar to perform the target behavior.

[0007] According to another aspect of this disclosure, a control device for a virtual avatar is provided, comprising: a receiving unit configured to receive a user's voice audio; a first prediction unit configured to predict a target behavior label based on a speech recognition result of the voice audio, wherein the target behavior label is used to obtain a first behavior corresponding to the target behavior label from a preset behavior set, the preset behavior set including at least one preset behavior and at least one behavior label corresponding to the at least one preset behavior; a determining unit configured to determine a target behavior based on a first action of the user's virtual avatar and the first behavior corresponding to the target behavior label, wherein the first action indicates the current action of the user's virtual avatar in a virtual scene, and the target behavior includes at least one of a target action and an action of emitting target speech; and a control unit configured to control the user's virtual avatar to perform the target behavior.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the aforementioned virtual avatar control method.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the control method of the virtual image described above.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program, when executed by a processor, implements the aforementioned method for controlling the virtual image.

[0011] According to one or more embodiments of this disclosure, it is possible to further enhance the user's perception and emotional expression in virtual scenes, thereby improving the user experience.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;

[0015] Figure 2 A flowchart illustrating a method for controlling a virtual avatar according to an embodiment of the present disclosure is shown;

[0016] Figure 3 A flowchart illustrating the determination of a target behavior based on a user's virtual avatar's first action and first behavior, according to an embodiment of the present disclosure, is shown.

[0017] Figure 4 A schematic diagram illustrating the overlay of special effects materials according to an exemplary embodiment of the present disclosure is shown;

[0018] Figure 5 A schematic diagram showing an enlarged display of special effects material according to an exemplary embodiment of the present disclosure is shown;

[0019] Figure 6 A structural block diagram of a control device for a virtual avatar according to an embodiment of the present disclosure is shown;

[0020] Figure 7 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0021] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0022] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0023] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0024] In related technologies, virtual scenarios typically involve operating virtual touch panels in the virtual world or using hardware devices such as controllers in the real world to input text or select and send emoticons. The sent text or emoticons are then displayed as bullet comments or special effects. This interaction method not only requires additional user input, but its presentation also disrupts the overall visual effect of the scene, thus interrupting the user's immersive experience.

[0025] In the embodiments of this disclosure, by acquiring the user's voice audio, performing speech recognition on it, and predicting the speech recognition results through a classification model to obtain corresponding behavior labels, and obtaining corresponding preset behaviors based on the behavior labels, the user's virtual avatar performs the preset behaviors. This can further enhance the user's perception and emotional expression in the virtual world, thereby bringing the user a more immersive experience.

[0026] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0027] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0028] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of the control methods for the virtual avatar described above.

[0029] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.

[0030] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0031] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to collect real-time motion and voice data. The client devices can provide interfaces that allow users to interact with them. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0032] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0033] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0034] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0035] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0036] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.

[0037] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0038] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0039] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0040] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0041] According to some embodiments, such as Figure 2 As shown, a method for controlling a virtual avatar is provided, including: step S201, receiving a user's voice audio; step S202, predicting a target behavior label based on the speech recognition result of the voice audio, wherein the target behavior label is used to obtain a first behavior corresponding to the target behavior label from a preset behavior set, the preset behavior set including at least one preset behavior and at least one behavior label corresponding to the at least one preset behavior; step S203, determining a target behavior based on the user's virtual avatar's first action and first behavior, wherein the first action indicates the current action of the user's virtual avatar in a virtual scene, and the target behavior includes at least one of a target action and the action of emitting target speech; and step S204, controlling the user's virtual avatar to perform the target behavior.

[0042] Therefore, by acquiring the user's voice audio, performing speech recognition, and using a classification model to predict the speech recognition results, corresponding behavior labels are obtained. Based on the behavior labels, corresponding preset behaviors are obtained, and the user's virtual avatar performs the preset behaviors. This can further enhance the user's perception and emotional expression in the virtual world, thereby bringing the user a more immersive experience.

[0043] In some embodiments, a virtual avatar may be a user's avatar in a virtual world, and the behavior of the virtual avatar in the virtual world is controlled by the user.

[0044] In some embodiments, hardware devices (such as sensors) can be used to collect information about the user during movement (e.g., relative skeletal positions) to obtain motion parameters, which are then used to control the corresponding virtual 3D model of the user to perform appropriate actions, thereby obtaining the movement of the virtual avatar. This process requires a large amount of equipment, needs to be operated by professional personnel, and involves complex information processing.

[0045] In some embodiments, the user's voice audio may be received via a hardware device (such as a radio device).

[0046] In some embodiments, the embodiments of this disclosure can be applied to specific virtual scenarios, such as virtual concert scenarios.

[0047] In a specific virtual scenario, in response to receiving a user's audio voice, speech recognition can be performed on the user's voice to obtain a speech recognition result. In some embodiments, the speech recognition result may include text information of the user's voice obtained through speech recognition.

[0048] In some embodiments, the text information of the user's speech (or the corresponding encoded vector) can be input into a pre-trained behavior prediction model, which can then perform predictive analysis to obtain the corresponding target behavior label. In some embodiments, the behavior prediction model can be a classification model built on a CNN network.

[0049] In some embodiments, a target behavior label is used to obtain a first behavior corresponding to the label from a preset behavior set. The preset behavior set may include at least one predefined preset behavior, and each preset behavior corresponds to a behavior label. Each preset behavior may include at least one of a preset action or the behavior of emitting a preset voice.

[0050] In some embodiments, those skilled in the art may set preset behaviors from the preset behavior set according to different specific virtual scenarios, without any restrictions.

[0051] In some exemplary embodiments, the preset set of behaviors for a virtual concert scene can include behaviors such as cheering, whistling, shouting, and waving. In one example, cheering behaviors may include, for example, the preset action of waving a glow stick and the behavior of uttering a preset cheering slogan (preset voice).

[0052] In some embodiments, predicting a target behavior label based on the speech recognition results of the audio recording may include: determining the user's emotion category based on the speech recognition results of the audio recording; and predicting the target behavior label based on the emotion category and the speech recognition results.

[0053] Therefore, emotion classification is performed based on speech recognition results, and behavior prediction is performed based on speech recognition results and corresponding emotion categories, thus making behavior prediction more accurate.

[0054] In some embodiments, the speech recognition result may include text information of the user's speech obtained through speech recognition.

[0055] In some embodiments, the text information of the user's speech (or the corresponding encoded vector) can first be input into a pre-trained emotion classification model, and the model can then perform predictive analysis to obtain the corresponding user emotion category. In some embodiments, the emotion classification model described above can be a classification model built based on CNN networks, etc., and there are no restrictions on this.

[0056] In some embodiments, the text information of the user's speech (or the corresponding encoded vector) and the user's emotion category can be further input into a pre-trained behavior prediction model. The model then performs prediction analysis to obtain the corresponding target behavior label. In some embodiments, the user emotion category input into the behavior prediction model can be an emotion label or a hidden feature obtained based on the aforementioned emotion classification model.

[0057] In some embodiments, the emotion category can be set according to actual needs, for example, it can include positive emotions, neutral emotions and negative emotions, without limitation.

[0058] In some embodiments, the speech recognition result includes text information corresponding to the speech audio and acoustic features. Determining the user's emotion category based on the speech recognition result of the speech audio may include: inputting the text information and acoustic features into an emotion classification model to obtain the emotion category output by the emotion classification model.

[0059] Therefore, by using multimodal data such as text information in user speech and acoustic features of user speech, emotion category prediction can be performed, thereby improving the accuracy of emotion classification.

[0060] In some embodiments, the speech recognition result may include text information of the user's speech obtained through speech recognition and the acoustic features of the speech audio.

[0061] In some embodiments, acoustic features may include basic acoustic features such as volume (amplitude of the audio) and pitch (frequency of the audio).

[0062] In some embodiments, acoustic features may also include, but are not limited to, Mel Frequency Cepstral Coefficient (MFCC) features, Constant Q-Cepstral Coefficients (CQCC) features, etc.

[0063] In some embodiments, the acquisition of the acoustic features described above may include: first, performing some preprocessing operations on the speech audio, such as dividing the speech audio into frames, wherein the frame division may divide the speech audio into multiple audio frames with a frame length of 25ms and a frame shift of 10ms; subsequently, the acoustic features described above may be extracted from each audio frame of the speech audio.

[0064] In some embodiments, the aforementioned text information (or the corresponding encoded vector of the text information) and one or more of the aforementioned acoustic features can be simultaneously input into a pre-trained emotion classification model. The model then performs predictive analysis to obtain the corresponding user emotion category. In some embodiments, the aforementioned emotion classification model can be a classification model built based on CNN networks, etc., and this is not limited thereto.

[0065] In some embodiments, determining a user's emotion category based on the speech recognition results of the speech audio may further include: detecting whether the text information includes preset text to obtain the detection results of the text information; and updating the emotion category based on the detection results and / or a first acoustic feature of the speech audio, wherein the first acoustic feature includes at least one of the amplitude and frequency of the speech audio.

[0066] Therefore, based on the obtained prediction results, further detection of whether the text information contains preset text and judgment of whether the volume (amplitude of audio) and pitch (frequency of audio) meet preset conditions can make the emotions more clearly defined, thereby enabling more accurate acquisition of the user's emotional state.

[0067] In some embodiments, the emotion categories obtained through the emotion classification model may include, for example, positive emotions, neutral emotions, and negative emotions.

[0068] In some embodiments, preset text can be set according to a specific virtual scene. For example, for a virtual concert scene, preset text may include the singer's name, the singer's cheering slogan, the song title, "It's so good!", "I love you", etc. In some embodiments, different preset texts may correspond to different emotion categories. For example, the singer's name, the singer's cheering slogan, the song title, "It's so good!", "I love you" may correspond to positive emotions; "It's so bad!", "The styling is bad" may correspond to negative emotions; "Performed normally" may correspond to neutral emotions.

[0069] In some embodiments, in response to the detection of preset text in the spoken text information, the emotion category can be updated based on the detection result. For example, if the model determines that the user's current emotion is neutral, but the detection result shows that the text information contains the preset text of a singer's name, the emotion category can be updated to positive emotion.

[0070] In some embodiments, the emotion category may be updated based on a first acoustic feature of the speech audio, namely at least one of the amplitude and frequency of the speech audio.

[0071] In some embodiments, different emotion categories can be determined based on the amplitude (corresponding to the volume) of the voice audio. For example, when the amplitude of the voice audio is greater than a first preset amplitude threshold, or the volume is greater than a first preset volume threshold (e.g., 60 decibels), the user's emotion can be judged as excited or agitated; when the amplitude of the voice audio is less than a second preset amplitude threshold, or the volume is less than a second preset volume threshold (e.g., 40 decibels), the user's emotion can be judged as depressed; otherwise, the user's emotion is judged as neutral or calm.

[0072] In some embodiments, different emotion categories can also be determined based on the frequency (corresponding to the pitch) of the speech audio. For example, when the frequency of the speech audio is greater than a first preset frequency threshold, the user's emotion can be judged as excited or agitated; when the frequency of the speech audio is less than a second preset frequency threshold, the user's emotion can be judged as depressed; otherwise, the user's emotion is judged as neutral or calm.

[0073] In some embodiments, the emotion category can be updated based on at least one of the amplitude and frequency of the speech audio. For example, if the model determines that the user's current emotion is negative and the volume of the speech audio is greater than 60 decibels, the emotion category can be updated to a highly negative emotion.

[0074] In some embodiments, the emotion category can be updated based on at least one of the detection results and the amplitude and frequency of the speech audio. For example, if the model determines that the user's current emotion is positive, and the detected text information includes preset text such as "It sounds so good" and the singer's name, and the volume of the speech audio is greater than a first preset volume threshold, then the emotion category can be updated to positive and excited.

[0075] In some embodiments, the aforementioned text information (or the encoding vector corresponding to the text information), one or more of the aforementioned acoustic features, and the user's emotion category can be further input into a pre-trained behavior prediction model, and the model can be used for prediction analysis to obtain the corresponding target behavior label.

[0076] In some embodiments, such as Figure 3 As shown, determining the target behavior based on the user's first action and first behavior of the virtual avatar may include: step S301, obtaining the first behavior that matches the target behavior tag in the preset behavior set based on the target behavior tag; step S302, obtaining the user's action parameters collected by the hardware device; step S303, determining the first action based on the action parameters; and step S304, fusing the first behavior and the first action to obtain the target behavior.

[0077] Therefore, by integrating preset behaviors with the user's current actions, the user's virtual avatar can perform the integrated behavior. This means the user's virtual avatar not only displays user actions collected by sensors and other hardware devices, but also further incorporates preset behaviors, thereby enhancing the user's perception and emotional expression in the virtual scene and providing a more immersive experience.

[0078] In some embodiments, the set of preset behaviors may include at least one predefined preset behavior, and each preset behavior corresponds to a behavior label. Each preset behavior may include at least one of a preset action or a preset voice action.

[0079] In some embodiments, those skilled in the art may set preset behaviors from the preset behavior set according to different specific virtual scenarios, without any restrictions.

[0080] In some exemplary embodiments, the preset set of behaviors for a virtual concert scene can include behaviors such as cheering, whistling, shouting, and waving, each corresponding to a different behavior label. In one example, cheering behaviors may include preset actions such as waving glow sticks and uttering preset cheering slogans (preset voices).

[0081] In some embodiments, the target behavior label can be obtained based on a behavior classification model corresponding to a specific virtual scene, and the output target behavior label is one of the behavior labels in a preset behavior set. Therefore, the first behavior corresponding to the target behavior label can be obtained by matching the target behavior label with the behavior labels in the preset behavior set.

[0082] In some embodiments, after a preset behavior is matched, the target behavior can be obtained by fusing the user's virtual avatar's current first action with the aforementioned first behavior.

[0083] In some embodiments, the user's virtual avatar's current first action can be determined by collecting the user's motion parameters through hardware devices (such as motion sensors) and determining the user's virtual avatar's first action based on the motion parameters.

[0084] In some embodiments, the first action may include a corresponding preset action, and the preset action and the first action in the first action can be fused. First, control parameters of the first action and control parameters of the preset action can be obtained (the control parameters may include the relative positions of multiple key points on the virtual avatar and the position of the user's virtual avatar in the virtual space). Based on the control parameters of the first action, the control parameters of the preset action are calibrated so that the multiple key points of the preset action correspond to the user's virtual avatar, thereby obtaining the fused target action and causing the user's virtual avatar to perform the target action (i.e., the target behavior).

[0085] In one example, if the first action of the user's virtual avatar is sitting in a chair and the preset action is waving, then the merged target action (i.e. the target behavior) could be the user's virtual avatar sitting in a chair and waving.

[0086] In some embodiments, the first action may include, for example, emitting a preset voice. Then, fusing the first action and the first movement may include maintaining the first movement while simultaneously emitting the preset voice (i.e., the target voice).

[0087] In some embodiments, the first behavior may include a corresponding preset action and the behavior of emitting a preset voice. The fusion of the first behavior and the first action may include fusing the preset action and the first action based on the above method, and controlling the user's virtual avatar to emit a preset voice (i.e., the target voice).

[0088] In some embodiments, the first behavior described above can also be used as the target behavior, and the user's virtual avatar can be controlled to display the target behavior.

[0089] In some embodiments, a display time can be set for the target behavior. After the preset display time after the user's virtual avatar performs the target behavior, the first behavior can be stopped from being merged with the user's virtual avatar behavior. That is, after the preset display time, the user's virtual avatar behavior can be restored to the virtual avatar behavior generated only by obtaining real-time motion parameters through hardware devices.

[0090] In some embodiments, the virtual avatar control method of this disclosure may further include: predicting a target special effect tag based on emotion category and speech recognition results; obtaining a target special effect material matching the target special effect tag from a preset special effect material set based on the target special effect tag, wherein the preset special effect material set includes at least one preset special effect material and at least one special effect tag corresponding to the at least one preset special effect material; and displaying the target special effect material in a virtual scene.

[0091] Therefore, based on the speech recognition results and emotion categories, predictions are made to obtain the predicted special effects labels and match the corresponding special effects materials. This enables the system to respond to the user's speech expression and display different special effects to the user in real time, further enhancing the user's perception and emotional expression in the virtual scene and bringing the user a more immersive experience.

[0092] In some embodiments, based on the emotion category and speech recognition results obtained by the above method (which may include at least one of the textual information and acoustic features of the speech audio), one or more of the above feature information can be input into a pre-trained effect prediction model, and the model can be used for prediction analysis to obtain the corresponding target effect label. In some embodiments, the above effect prediction model can be a classification model built based on CNN networks, etc., and there are no limitations on this.

[0093] In some embodiments, the preset effects materials in the preset effects material set may be set according to a specific virtual scene, and each preset effects material corresponds to an effects tag.

[0094] In some embodiments, the target effect tag can be obtained based on an effect prediction model corresponding to a specific virtual scene, and the output target effect tag is one of the effect tags in a preset set of effect materials. Therefore, the target effect material corresponding to that effect tag can be obtained by matching the target effect tag with the effect tags in the preset set of effect materials.

[0095] In some embodiments, the virtual avatar control method of this disclosure may further include: in response to multiple matching of target special effects material based on the user's voice audio within a preset time range, superimposing or expanding the display of the target special effects material in a virtual scene.

[0096] In some exemplary embodiments, a specific virtual scene may be, for example, a virtual concert scene, and the corresponding set of preset special effects materials may include special effects materials such as fireworks effects, light effects, and heart effects.

[0097] In some embodiments, when the same target special effect material is triggered multiple times within a preset time range using the above method, the target special effect material can be overlaid or expanded for display. Thus, by using diverse special effect display methods, the special effects can be enhanced, thereby further strengthening the user's visual experience and emotional expression in the virtual scene, and improving the user experience.

[0098] Figure 4 A schematic diagram showing the overlay of special effects materials according to an exemplary embodiment of the present disclosure is shown.

[0099] In one example, such as Figure 4As shown, if a user sends out multiple "I love you" audio messages within a preset time range (e.g., within 10 seconds) and triggers a heart-shaped effect each time, the effect can be stacked on top of each other for display.

[0100] In some embodiments, the display time of special effects can be set; when the display time of the special effects exceeds a preset display time, the special effects will automatically disappear. For example... Figure 4 As shown, the preset display time can be, for example, 10 seconds. Then, when the 10th second is reached, the first heart effect will automatically disappear.

[0101] Figure 5 A schematic diagram showing an enlarged display of special effects material according to an exemplary embodiment of the present disclosure is shown.

[0102] In one example, such as Figure 5 As shown, in response to a user uttering "I love you" multiple times within a preset time range (e.g., within 10 seconds) and triggering a heart-shaped effect each time, the effect can be gradually expanded and displayed.

[0103] In some embodiments, the display time of special effects can be set; when the display time of the special effects exceeds a preset display time, the special effects will automatically disappear. For example... Figure 5 As shown, the preset display time can be, for example, 10 seconds. When the 13th second is reached, that is, when the display time of the last of the multiple heart-shaped effects exceeds the preset display time, the heart-shaped effect will automatically disappear.

[0104] In some embodiments, such as Figure 6 As shown, a virtual avatar control device 600 is provided, comprising: a receiving unit 610 configured to receive a user's voice audio; a first prediction unit 620 configured to predict a target behavior label based on the voice recognition result of the voice audio, wherein the target behavior label is used to obtain a first behavior corresponding to the target behavior label from a preset behavior set, the preset behavior set including at least one preset behavior and at least one behavior label corresponding to the at least one preset behavior; a determining unit 630 configured to determine a user's target behavior based on a first action and a first behavior of the user's virtual avatar, wherein the first action indicates the current action of the user's virtual avatar in a virtual scene, and the target behavior includes at least a target action and / or the act of emitting a target voice; and a control unit 640 configured to control the user's virtual avatar to perform the target behavior.

[0105] The operation of units 610-640 of the virtual image control device 600 is similar to the operation of steps S201-S204 in the virtual image control method described above, and will not be described in detail here.

[0106] In some embodiments, the prediction unit may include: a first determining subunit configured to determine the user's emotion category based on the speech recognition result of the speech audio; and a prediction subunit configured to predict a target behavior label based on the emotion category and the speech recognition result.

[0107] In some embodiments, the speech recognition result includes text information corresponding to the speech audio and acoustic features. The determining subunit may include: an acquisition module configured to input the text information and acoustic features into an emotion classification model to obtain the emotion category output by the emotion classification model.

[0108] In some embodiments, the determining subunit may further include: a detection module configured to detect whether the text information includes preset text in order to obtain a detection result of the text information; and an updating module configured to update the emotion category based on the detection result and / or a first acoustic feature of the speech audio, wherein the first acoustic feature includes at least one of the amplitude and frequency of the speech audio.

[0109] In some embodiments, the determining unit may include: a first acquiring subunit configured to acquire a first behavior matching the target behavior tag from a preset behavior set based on the target behavior tag; a second acquiring subunit configured to acquire user action parameters collected by a hardware device; a second determining subunit configured to determine a first action based on the action parameters; and a fusing subunit configured to fuse the first behavior and the first action to obtain the target behavior.

[0110] In some embodiments, the control device for the virtual avatar disclosed herein may further include: a second prediction unit configured to predict a target effect tag based on an emotion category and a speech recognition result; an acquisition unit configured to acquire a target effect material matching the target effect tag from a preset set of effect materials based on the target effect tag, wherein the preset set of effect materials includes at least one preset effect material and at least one effect tag corresponding to the at least one preset effect material; and a first display unit configured to display the target effect material in a virtual scene.

[0111] In some embodiments, the control device for the virtual avatar disclosed herein may further include: a second display unit configured to overlay or expand the display of the target special effects material in a virtual scene in response to multiple matchings of the user's voice audio within a preset time range.

[0112] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0113] refer to Figure 7 The present invention describes a structural block diagram of an electronic device 700 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0114] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0115] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, storage unit 708, and communication unit 709. Input unit 706 can be any type of device capable of inputting information to electronic device 700. Input unit 706 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 707 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 708 may include, but is not limited to, hard disk and optical disk. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0116] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the virtual avatar control method described above. For example, in some embodiments, the virtual avatar control method described above can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the virtual avatar control method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the virtual avatar control method described above by any other suitable means (e.g., by means of firmware).

[0117] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0118] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0119] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0121] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0122] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0123] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0124] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A method for controlling a virtual avatar, the method comprising: Receive user's voice and audio; Based on the speech recognition results of the audio, the target behavior label is predicted; Based on the target behavior tag, the first behavior that matches the target behavior tag is obtained from a preset behavior set, wherein the preset behavior set includes at least one preset behavior and at least one behavior tag corresponding to the at least one preset behavior; Based on the user's virtual avatar's first action and first behavior, a target behavior is determined, wherein the virtual avatar is the user's avatar in a virtual scene, the first action indicates the current action of the user's virtual avatar in the virtual scene, which is controlled by the user in real time, and the target behavior includes a target action and the act of emitting a target voice. Determining the target behavior based on the user's virtual avatar's first action and first behavior includes: Obtain the user's real-time action parameters collected by the hardware device; Based on the real-time action parameters, the first action is determined; and The first behavior and the first action are fused to obtain the target behavior; and Control the user's virtual avatar to perform the target behavior.

2. The method of claim 1, wherein, The speech recognition results based on the speech audio predict the target behavior labels, including: Based on the speech recognition results of the audio, the user's emotion category is determined; and Based on the emotion category and the speech recognition result, the target behavior label is predicted.

3. The method according to claim 2, wherein, The speech recognition result includes the text information and acoustic features corresponding to the speech audio. Determining the user's emotion category based on the speech recognition result includes: The text information and the acoustic features are input into the emotion classification model to obtain the emotion category output by the emotion classification model.

4. The method according to claim 3, wherein, The process of determining the user's emotion category based on the speech recognition results of the audio audio further includes: Detect whether the text information includes preset text, and obtain the detection result of the text information; and The emotion category is updated based on the detection results and / or the first acoustic features of the speech audio, wherein the first acoustic features include at least one of the amplitude and frequency of the speech audio.

5. The method according to any one of claims 1-4, further comprising: Based on the emotion category and the speech recognition result, the target effect label is predicted; Based on the target effect tag, target effect materials matching the target effect tag are obtained from a preset effect material set, wherein the preset effect material set includes at least one preset effect material and at least one effect tag corresponding to the at least one preset effect material; as well as The target special effects material is displayed in the virtual scene.

6. The method according to claim 5, further comprising: In response to multiple matchings of the target special effects material based on the user's voice audio within a preset time range, the target special effects material is overlaid or expanded in the virtual scene.

7. A control device for a virtual avatar, the device comprising: The receiving unit is configured to receive the user's voice audio. The first prediction unit is configured to predict the target behavior label based on the speech recognition result of the speech audio. A unit for obtaining the first behavior that matches the target behavior tag from the preset behavior set based on the target behavior tag, wherein the preset behavior set includes at least one preset behavior and at least one behavior tag corresponding to the at least one preset behavior; A determining unit is configured to determine a target behavior based on a first action and a first behavior of the user's virtual avatar, wherein the virtual avatar is the user's avatar in a virtual scene, the first action indicates the current action of the user's virtual avatar in the virtual scene controlled by the user in real time, and the target behavior includes a target action and the act of emitting a target voice. The determining unit is configured to: Obtain the user's real-time action parameters collected by the hardware device; Based on the real-time action parameters, the first action is determined; and The first behavior and the first action are fused to obtain the target behavior; and The control unit is configured to control the user's virtual avatar to perform the target behavior.

8. The apparatus according to claim 7, wherein, The prediction unit includes: The first determining subunit is configured to determine the user's emotion category based on the speech recognition result of the speech audio; and The prediction subunit is configured to predict the target behavior label based on the emotion category and the speech recognition result.

9. The apparatus according to claim 8, wherein, The speech recognition result includes the text information corresponding to the speech audio and acoustic features, and the determining subunit includes: The acquisition module is configured to input the text information and the acoustic features into the emotion classification model to obtain the emotion category output by the emotion classification model.

10. The apparatus according to claim 9, wherein, The determining subunit further includes: The detection module is configured to detect whether the text information includes preset text, so as to obtain the detection result of the text information; and An update module is configured to update the emotion category based on the detection result and / or a first acoustic feature of the speech audio, wherein the first acoustic feature includes at least one of the amplitude and frequency of the speech audio.

11. The apparatus according to any one of claims 7-10, further comprising: The second prediction unit is configured to predict the target effect label based on the emotion category and the speech recognition result; The acquisition unit is configured to acquire target special effects material that matches the target special effects tag from a preset special effects material set based on the target special effects tag, wherein the preset special effects material set includes at least one preset special effects material and at least one special effects tag corresponding to the at least one preset special effects material respectively; as well as The first display unit is configured to display the target special effects material in the virtual scene.

12. The apparatus of claim 11, further comprising: The second display unit is configured to overlay or expand the target special effects material in the virtual scene in response to multiple matchings of the user's voice audio within a preset time range.

13. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

15. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Virtual anchor implementation method and device

    CN109637518A

  • Voice interaction method and apparatus, and electronic equipment

    CN111292743A

  • Video generation method and device, live broadcast processing method and device and readable medium

    CN113923462A