AVATAR VIDEO DEVICE AND METHOD

DE602014092786T2Active Publication Date: 2026-02-11TAHOE RES LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE602014092786
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2014-11-05
Publication Date
2026-02-11
Estimated Expiration
2034-11-05

AI Technical Summary

Technical Problem

Creating high-quality avatar videos without requiring large budgets or professional expertise is challenging due to the complexity of existing graphics editing software and animation techniques, especially in detecting dynamic facial features like tongue-out movements.

Method used

An avatar video generation system incorporating facial expression engines, an animation-rendering engine, and a video generator, which includes a tongue-out detector using a mouth region detector, extractor, and classifier, optimized for mobile devices, to efficiently analyze video, voice, and text inputs and generate realistic facial expressions and tongue-out animations.

Benefits of technology

Facilitates the creation of high-quality avatar videos with realistic tongue-out animations on mobile devices, reducing computational intensity and enabling amateur creators to produce engaging content with minimal resources.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing. More particularly, the present disclosure relates to the creation of avatar video, including tongue-out detection.Background

[0002] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art by inclusion in this section.

[0003] Micro-movies and cartoon videos made by amateurs have become increasingly popular, in particular, in social networks. An example is the 'Annoying Orange' American comedy series shared on YouTube ®< , where an anthropomorphic orange annoys other fruits, vegetables, and various other objects, and make jokes. Each of these videos usually consists of simple characters, but tells an interesting story. While these video typically do not require large budgets or major studio backings to produce them, it is still nonetheless not easy for an amateur to create it via today's graphics editing software and / or movie composition suites. Often, a small studio, with experienced artists having some years of accumulated art skills in, e.g., human action capturing and retargeting, character animation and rendering, is still required.

[0004] Hanser Eva: "SceneMaker: Intelligent Multimodal Visualisation of Natural Language Scripts",Ph.D. Dissertation, 1 January 2009, pages 1-45, discloses a system for performing plays or creating films and animations by automatically interpreting film and play scripts and automatically generating animated scenes from them. A genre-specific text-to-animation methodology is presented, wherein emotional expressivity inspired by the OCC (Ortony, Clore and Collins) emotion model enhances believability of virtual actors as well as scene presentation. The proposed system infers emotions from the story context, rather than relying on explicit emotion keywords, automatically detects genre from the script and applies appropriate cinematic direction.

[0005] JP2008241772A discloses a voice image processing device by which a processing for changing a mouth shape to be displayed on a screen by being synchronized with user's uttered voice is performed with a simple calculation. The voice image processing device includes: a memory section for storing a collation triangle string which approximates a waveform of a syllable, and a syllable image of the mouth shape which utters the syllable by relating them; an input section for receiving input of a voice signal; an approximation section for approximating a waveform of the received voice signal in the approximation triangle string; a determination section for determining whether or not, the approximation triangle string matches the collation triangle string; an output section for outputting the received voice signal; and a display section for displaying the syllable image corresponding to the collation triangle string, when it matches the approximation triangle corresponding to a period of the voice signal which is currently output.

[0006] TAEHOON CHO ET AL: "Vision-based animation of 3D facial avatars",2014 INTERNATIONAL CONFERENCE ON BIG DATA AND SMART COMPUTING (BIGCOMP), IEEE, 15 January 2014, pages 128-132, discloses a system that classifies a user's facial expressions captured by a webcam and applies notable characteristics of the classified expression to an animated 3D facial avatar. The facial avatar is composed of multiple 3D models where a face model is combined with accessorial models such as facial hair. Such accessorial models are used to decorate the avatar, and are automatically registered with the face model during animation.

[0007] WO96 / 17323A1 discloses a method and apparatus for synthesizing speech or facial movements to match selected speech sequences. A videotape of an arbitrary text sequence is obtained including a plurality of images of a user speaking various sequences. Video images corresponding to specific spoken phonemes are obtained. A video frame is digitized from that sequence which represents the extreme of mouth motion and shape. This is used to create a data base of images of different facial positions relative to spoken phonemes and diphthongs. An audio speech sequence is then used as the element to which a video sequence will be matched. The audio sequence is analyzed to determine spoken phoneme sequences and relative timings. The database is used to obtain images for each of these phonemes and these times, and morphing techniques are used to create transitions between the images. Different parts of the images can be processed in different ways to make a more realistic speech pattern.

[0008] US6272231B1 discloses an apparatus, and related method, for sensing a person's facial movements, features and characteristics and the like to generate and animate an avatar image based on facial sensing. The avatar apparatus uses an image processing technique based on model graphs and bunch graphs that efficiently represent image features as jets. The jets are composed of wavelet transforms processed at node or landmark locations on an image corresponding to readily identifiable features. The nodes are acquired and tracked to animate an avatar image in accordance with the person's facial movements. Also, the facial sensing may use jet similarity to determine the person's facial features and characteristic thus allows tracking of a person's natural characteristics without any unnatural elements that may interfere or inhibit the person's natural characteristics.

[0009] US2010 / 201693A1 discloses a system and method for capturing the voice and motion of a user and mapping the captured voice and motion to an avatar is disclosed. Other aspects include displaying the avatar in the virtual world of a movie or animation chosen by the user.

[0010] WO2013 / 149556A1 discloses a method and device for automatically playing an expression on a virtual image. The method includes the steps of : A, capturing a video image related to a game player in a game, and determining whether the captured video image contains a facial image of the game player, and if so, executing a step B; otherwise, returning to the step A; B: extracting facial features from the facial image, and obtaining a motion track of each of the facial features by motion directions of the extracted facial feature that are acquired for N times, where N is greater than or equal to 2; and C, determining a corresponding facial expression according to the motion track of the facial feature which is obtained in step B, and automatically playing the determined facial expression on a virtual image of the game player.

[0011] CN104112117A discloses an advanced local binary pattern feature tongue motion identification method. The method comprises the following steps: extracting a mouth area extraction image; detecting the mouth area in a facial image, carrying out graying and normalization on the mouth area extraction image, and setting the dimension as the 32*16 pixel; using the advanced local binary pattern algorithm and carrying out processing on a pixel difference value in the local binary pattern calculating zone to keep more vertical direction information; and carrying out tongue motion classification by using a classifier of a support vector machine.

[0012] US2007 / 168863A1 discloses a method and a system, wherein an avatar that represents an user in a communications session is animated, without user manipulation, based on the animation of another avatar that represents another user in the same instant messaging communication session. The avatars may be displayed in a single instant messaging window, and the displayed animations may create an appearance that the avatars are interacting with one another. An avatar animation may be based on the content communicated by a user and a category that is associated with a user.

[0013] JP2005148959A discloses a method, wherein the image of the face of a person is picked up, and the face image signal of the person is obtained, and a lip pattern identification result display information showing the identification result of whether or not a lip pattern shown by the face image signal is matched with or the closest to any of a plurality of lip patterns for reference is obtained by using the face image signal, and a conversation sentence corresponding to the lip pattern for reference with or to which the lip pattern shown by the face image signal is matched or the closest among the plurality of lip patterns for reference is displayed by using the lip pattern identification result display information.

[0014] US2014 / 112556A1 discloses a method, wherein features, including one or more acoustic features, visual features, linguistic features, and physical features may be extracted from signals obtained by one or more sensors with a processor. The acoustic, visual, linguistic, and physical features may be analyzed with one or more machine learning algorithms and an emotional state of a user may be extracted from analysis of the features.

[0015] KR20040014123A discloses a system for an emotion expression and an action implementation of a virtual personality and a method thereof are provided to implement a virtual personality which expresses an analyzed emotion as an operation by analyzing an emotion included in inputted operation data or natural language data.Brief Description of the Drawings

[0016] Embodiments will be readily understood by the following detailed description in conjunction with the accompanying drawings. To facilitate this description, like reference numerals designate like structural elements. Embodiments are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings. Figure 1 illustrates a block diagram of an avatar video generation system, according to the disclosed embodiments. Figure 2 illustrate a process for generating an avatar video, according to the disclosed embodiments. Figure 3 illustrates a block diagram of the tongue-out-detector of Figure 1 in further detail, according to the disclosed embodiments. Figure 4 illustrates sub-windows of an extracted mouth region, according to the disclosed embodiments. Figure 5 illustrates two image frames of a generated video, according to the disclosed embodiments. Figure 6 illustrates an example computer system suitable for use to practice various aspects of the present disclosure, according to the disclosed embodiments. Figure 7 illustrates a storage medium having instructions for practicing methods described with references to Figures 1-5, according to disclosed embodiments. Detailed Description

[0017] Apparatuses, methods and storage medium associated with creating an avatar video are disclosed herein. In embodiments, the apparatus may one or more facial expression engines, an animation-rendering engine, and a video generator coupled with each other. The one or more facial expression engines may be configured to receive video, voice and / or text inputs, and, in response, generate a plurality of animation messages having facial expression parameters that depict facial expressions for a plurality of avatars based at least in part on the video, voice and / or text inputs received. The animation-rendering engine may be coupled with the one or more facial expression engines, and configured to receive the one or more animation messages, and drive a plurality of avatar models, to animate and render the plurality of avatars with the facial expressions depicted. The video generator may be coupled with the animation-rendering engine, and configured to capture the animation and rendering of the plurality of avatars, to generate a video.

[0018] In embodiments, a video driven facial expression engine may include a tongue-out detector. The tongue-out detector may include a mouth region detector, a mouth region extractor, and a tongue classifier, coupled with each other. The mouth region detector may be configured to identify locations of a plurality of facial landmarks associated with identifying a mouth in the image frame. The mouth region extractor may be coupled with the mouth region detector, and configured to extract a mouth region from the image frame, based at least in part on the locations of the plurality of facial landmarks identified. The tongue classifier may be coupled with the mouth region extractor to analyze a plurality of sub-windows within the mouth region extracted to detect for tongue-out. In embodiments, the tongue-out detector may further include a temporal filter coupled with the tongue classifier and configured to receive a plurality of results of the tongue classifier for a plurality of image frames, and output a notification of tongue-out detection on successive receipt of a plurality of results from the tongue classifier indicating tongue-out detection for a plurality of successive image frames.

[0019] In the following detailed description, reference is made to the accompanying drawings which form a part hereof wherein like numerals designate like parts throughout, and in which is shown by way of illustration embodiments that may be practiced. It is to be understood that other embodiments may be utilized and structural or logical changes may be made without departing from the scope of the present disclosure. Therefore, the following detailed description is not to be taken in a limiting sense, and the scope of embodiments is defined by the appended claims and their equivalents.

[0020] Aspects of the disclosure are disclosed in the accompanying description. Alternate embodiments of the present disclosure and their equivalents may be devised without parting from the spirit or scope of the present disclosure. It should be noted that like elements disclosed below are indicated by like reference numbers in the drawings.

[0021] Various operations may be described as multiple discrete actions or operations in turn, in a manner that is most helpful in understanding the claimed subject matter. However, the order of description should not be construed as to imply that these operations are necessarily order dependent. In particular, these operations may not be performed in the order of presentation. Operations described may be performed in a different order than the described embodiment. Various additional operations may be performed and / or described operations may be omitted in additional embodiments.

[0022] For the purposes of the present disclosure, the phrase "A and / or B" means (A), (B), or (A and B). For the purposes of the present disclosure, the phrase "A, B, and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B and C).

[0023] The description may use the phrases "in an embodiment," or "in embodiments," which may each refer to one or more of the same or different embodiments. Furthermore, the terms "comprising," "including," "having," and the like, as used with respect to embodiments of the present disclosure, are synonymous.

[0024] As used herein, the term "module" may refer to, be part of, or include an Application Specific Integrated Circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or group) and / or memory (shared, dedicated, or group) that execute one or more software or firmware programs, a combinational logic circuit, and / or other suitable components that provide the described functionality.

[0025] Referring now to Figure 1, wherein an avatar video generation system, according to the disclosed embodiments, is shown. As illustrated, avatar video generation system 100 may include one or more facial expression engines 102, avatar animation-rendering engine 104, and video generator 106, coupled with each other as shown. As described earlier, and in more detail below, the one or more facial expression engines 102 may be configured to receive video, voice and / or text inputs, and, in response, generate a plurality of animation messages 108 having facial expression parameters that depict facial expressions for a plurality of avatars based at least in part on the video, voice and / or text inputs received. The facial expressions may include, but are not limited to, eye and / or mouth movements, head poses, such as, head rotation, movement, and / or coming closer or farther from the camera, and so forth. Animation-rendering engine 104 may be coupled with the one or more facial expression engines 102, and configured to receive the one or more animation messages 108, and drive a plurality of avatars model, to animate and render the plurality of avatars with the facial expressions depicted. Video generator 106 may be coupled with animation-rendering engine 104, and configured to capture the animation and rendering of the plurality of avatars, to generate a video.

[0026] In embodiments, facial expression engines 102 may include video driven facial expression engine (VDFEE) 112, voice recognition facial expression engine (VRFEE) 114, and text based facial expression engine (TBFEE) 116, coupled in parallel with avatar animation-rendering engine 104.

[0027] VDFEE 112 may be configured to receive video inputs having a plurality of image frames (e.g., from an image source, such as a camera (not shown)), and analyze the image frames for facial movements, such as, but not limited to eye and / or mouth movements, head poses, and so forth. Head poses may include head rotation, movement, and / or coming closer or farther from the camera. Additionally, VDFEE 112 may be configured to generate a number of animation messages 108 having facial expression parameters that depict facial expressions for a plurality of avatars. The generation of the animation messages 108 may be performed, based at least in part on results of the analysis of the image frames. For example, VDFEE 112 may be configured to analyze the image frames for landmarks of a face or head poses, and generate at least a subset of the plurality of animation messages 108 having facial expression parameters that depict facial expressions for the plurality of avatars. The facial expressions may include eye and mouth movements or head poses of the avatars, based at least in part on landmarks of a face or head poses of the image frames. In embodiments, VDFEE 112 may be configured with (or provided access to) data with respect to blend shapes (and optionally, corresponding weights) to be applied to a neutral version of an avatar to morph the avatar to have the various facial expressions. Accordingly, VDFEE 112 may generate animation messages 108 with identification of blend shapes (and optionally, corresponding weights) to be applied to the neutral version of an avatar to morph the avatar to have particular facial expressions.

[0028] Any number of known techniques may be employed to identify a face in each of a number of image frames, and track the face over the number of image frames to detect facial movements / expressions and / or head poses. In embodiments, VDFEE 112 may employ a facial mesh tracker to identify and track a face, and to detect the facial expressions. The facial mesh tracker may, e.g., be the facial mesh tracker disclosed in PCT Application PCT / CN2014 / 073695, entitled FACIAL EXPRESSION AND / OR INTERACTION DRIVEN AVATAR APPARATUS AND METHOD, filed on March 19, 2014.

[0029] In embodiments, the mouth movements may include the avatars sticking their tongues out. Any number of known techniques may be employed to detect a tongue-out condition. However, in embodiments, the face mesh tracker may include a tongue-out detector 122 of the present disclosure to more efficiently detect the tongue-out condition, to be described more fully below.

[0030] VRFEE 114 is configured to receive audio inputs, analyze the audio inputs, and generate a number of the plurality of animation messages 108 having facial expression parameters that depict facial expressions for the plurality of avatars. The generation of animation messages 108 is performed, based at least in part on results of the analysis of the audio inputs. VRFEE 114 is configured to analyze the audio inputs for at least volume or syllable, and generate the plurality of animation messages 108 having facial expression parameters that depict facial expressions for the plurality of avatars. The facial expressions include mouth movements of the plurality of avatars, and the mouth movements are selected, based at least in part on volume or syllable of the audio inputs. In embodiments, VRFEE 114 may be configured with (or provided access to) data with respect to correspondence of volume and / or syllables to facial expressions. Further, similar to VDFEE 112, VRFEE 114 may be configured with (or provided access to) data with respect to blend shapes (and optionally, corresponding weights) to be applied to a neutral version of an avatar to morph the avatar to have various facial expressions. Accordingly, VRFEE 114 may generate animation messages 108 with identification of blend shapes (and optionally, corresponding weights) to be applied to a neutral version of an avatar to morph the avatar to have particular facial expressions.

[0031] TBFEE 116 is configured to receive text inputs, analyze the text inputs, and generate a number of the plurality of animation messages 108 having facial expression parameters depicting facial expressions for the plurality of avatars. The generation of the animation messages 108 is performed, based at least in part on results of the analysis of the text inputs. TBFEE 116 is configured to analyze the text inputs for semantics, and generate at least a subset of the plurality of animation messages having facial expression parameters that depict facial expressions for the plurality of avatars. The facial expression include mouth movements of the plurality of avatars, and the mouth movements are selected, based at least in part on semantics of the text inputs. In embodiments, TBFEE 116 may be configured with (or provided access to) data with respect to correspondence of variance semantics to facial expressions. Further, similar to VDFEE 112 and VRFEE 114, TBFEE 116 may be configured with (or provided access to) data with respect to blend shapes (and optionally, corresponding weights) to be applied to a neutral version of an avatar to morph the avatar to have various facial expressions. Accordingly, TBFEE 116 may generate animation messages 108 with identification of blend shapes (and optionally, corresponding weights) to be applied to a neutral version of an avatar to morph the avatar to have particular facial expressions.

[0032] Continue to refer to Figure 1, avatar animation-rendering engine 104 is configured to receive animation messages 108, and drive one or more avatar models, in accordance with animation messages 108, to animate and render the avatars, replicating facial expressions and / or head movements depicted. In embodiments, avatar animation-rendering engine 104 may be configured with a number of avatar models to animate a number of avatars. Avatar animation-rendering engine 104 may also be configured with an interface for a user to select avatars to correspond to various characters of a story. Further, as alluded to earlier, avatar animation-rendering engine 104 may animate a facial expression, through blending of a plurality of pre-defined shapes, making avatar video generation system 100 suitable to be hosted by a wide range of mobile computing devices. In embodiments, a model with neutral expression and some typical expressions, such as mouth open, mouth smile, brow-up, and brow-down, blink, etc., may be first preconstructed, prior to facial tracking and animation. The blend shapes may be decided or selected for various facial expression engines 102 capabilities and target mobile device system requirements. During operation, facial expression engines 102 may output the blend shape weights (e.g., as part of animation messages 108) for avatar animation-rendering engine 104.

[0033] Upon receiving the blend shape weights (α i ) for the various blend shapes, avatar animation-rendering engine 104 may generate the expressed facial results with the formula: B * = B o + ∑ i α i ⋅ Δ B i where B* is the target expressed facial, B 0 is the base model with neutral expression, and ΔB i is i th< blend shape that stores the vertex position offset based on base model for specific expression.

[0034] Compared with other facial animation techniques, such as motion transferring and mesh deformation, using blend shape for facial animation may have several advantages: 1) Expressions customization: expressions may be customized according to the concept and characteristics of the avatar, when the avatar models are created. The avatar models may be made more funny and attractive to users. 2) Low computation cost: the computation may be configured to be proportional to the model size, and made more suitable for parallel processing. 3) Good scalability: addition of more expressions into the framework may be made easier.

[0035] Still referring to Figure 1, video generator 106 is configured to capture a plurality of image frames of the animation and rendering of the plurality of avatars, and generate a video based at least in part on the image frames of the animation and rendering captured. In embodiments, video generator 106 may capture an array of avatars animated by avatar animation-rendering engine 104. In other embodiments, video generator 106 may be coupled with multiple avatar animation-rendering engines 104. For these embodiments, a video scene may contain multiple avatars animated by the multiple animation-rendering engines 104 simultaneously.

[0036] Each of facial expression engines 102, avatar animation-rendering engine 104, and / or video generator 106 may be implemented in hardware, software or combination thereof. For example, each of facial expression engines 102, avatar animation-rendering engine 104, and / or video generator 106 may be implemented with Application Specific Integrated Circuits (ASIC), programmable circuits programmed with the implementation logic, software implemented in assembler languages or high level languages compilable into machine instructions supported by underlying general purpose and / or graphics processors.

[0037] Referring now to Figure 2, wherein a process for generating an avatar video, according to the disclosed embodiments, is shown. As illustrated, in embodiments, process 200 for generating an avatar video may include operations performed in blocks 202-216. The operations may be performed e.g., by facial expression engines 102, avatar animation-rendering engine 104, and / or video generator 106 of Figure 1.

[0038] Process 200 may start at block 202. At block 202, conversations between various characters of a story, for which a video is to be generated, may be received. As described earlier, the conversation may be received via video, voice and / or text inputs (e.g., by corresponding ones of the facial expression engines 102). At block 204, avatars to correspond to the various characters may be chosen (e.g., via a user interface of animation-rendering engine 104).

[0039] From block 204, process 200 may proceed to blocks 206, 208, and / or 210, where the video, voice and / or text inputs may correspondingly fed to e.g., respective ones of facial expression engines 102, to process. As described earlier, image frames of the video inputs may be analyzed to identify landmarks of a face and / or head poses in the image frames, and in turn, animation messages 108 with facial expression parameters depicting facial expressions, such as eye and / or mouth movements, or head poses, may be generated based at least in part on the identified landmarks of a face and / or head poses. Audio inputs are analyzed for volume and / or syllables, and in turn, animation messages 108 with facial expression parameters depicting facial expressions, such as mouth movements are generated based at least in part on the identified volume and / or syllables. Text is analyzed for semantics, and in turn, animation messages 108 with facial expression parameters depicting facial expressions, such as mouth movements, are generated based at least in part on the identified semantics.

[0040] From blocks 206, 208 and 210, process 200 proceeds to block 212. At block 212, the various avatars are animated and rendered with facial expressions, in accordance with the animation messages 108 received. Further, the animation and rendering is captured, e.g., in a number of image frames.

[0041] At block 214, a determination is made whether all conversations between the characters have been animated and captured. If more conversations between the characters are to be animated and captured, process 200 returns to block 204, and continue therefrom, as earlier described. On the other hand, if all conversations between the characters have been animated and captured, process 200 may proceed to block 216. At block 216, the captured image frames are combined / stitched together to form the video. Thereafter, process 200 may end.

[0042] Referring now briefly back to Figure 1, as described earlier, in embodiments, video driven facial expression engine 112 may be equipped with a tongue-out detector incorporated with teachings of the present disclosure to efficient support detection of a tongue-out condition. In general, tongue is a dynamic facial feature - it shows up only when opening mouth. The shape of the tongue varies across individuals, and its motion is very dynamic. Existing methods for tongue detection mainly fall into two categories - one uses deformable template or active contour model to track the shape of tongue; and the other calculates similarity score of mouth region with template images and then determining the tongue status. Both categories of methods are relatively computational intensive, and not particularly suitable for today's mobile client devices, such as smartphones, computing tablets, and so forth.

[0043] Referring now to Figure 3, wherein a block diagram of the tongue-out-detector of Figure 1, according to the disclosed embodiments, is shown with further details. As illustrated, tongue-out detector 122 may include mouth region detector 304, mouth region extractor 306, tongue-out classifier 308, and optionally, temporal filter 310, coupled with each other. In embodiments, mouth region detector 304 may be configured to receive an image frame with a face identified, e.g., image frame 302 with bounding box 303 identifying a region where a face is located. Further, mouth region detector 304 may be configured to analyze image frame 302, and identify a number of facial landmarks relevant to identifying the mouth region. In embodiments, mouth region detector 304 may be configured to analyze image frame 302, and identify a chin point, a location of the left corner of the mouth, and a location of the right corner of the mouth (depicted by the dots in Figure 3.

[0044] Mouth region extractor 306, on the other hand, may be configured to extract the mouth region from image frame 302, and provide the extracted mouth region to tongue-out classifier 308. In embodiments, mouth region extractor 306 may be configured to extract the mouth region from image frame 302, based at least in part on the relevant landmarks, e.g., the chin point, the location of the left corner of the mouth, and the location of the right corner of the mouth.

[0045] Tongue-out classifier 308 may be configured to receive the extracted mouth region for an image frame. In embodiments, tongue-out classifier 308 may be trained with hundreds or thousands of mouth regions without tongue-out or with tongue-out in a wide variety of manners (i.e., negative and positive samples of tongue-out) for a large number of tongues of different sizes and shapes. In embodiments, tongue-out classifier 308 is trained to recognize attributes of a number of sub-windows of an extracted mouth region with a tongue-out condition. In embodiments, tongue-out classifier 308 may employ any one of a number classifier methods including, but are not limited to, Adaboost, Neural network, Support vector machine, etc. See Figure 4 where a number of example potential relevant sub-windows using the Adaboost method, are illustrated. In embodiments, tongue-out classifier 308 may be configured to determine whether to classify an extracted mouth region as having a tongue-out condition by calculating and comparing attributes within the reference sub-windows 402 of the extracted mouth region being analyzed.

[0046] In embodiments, calculation and comparison may be performed for Haar-like features. A Haar-like feature analysis is an analysis that considers adjacent rectangular regions at a specific location in a detection window, sums up the pixel intensities in each region and calculates the difference between these sums. This difference is then used to categorize subsections of the image frame.

[0047] In other embodiments, calculation and comparison may be performed for histogram of oriented gradient (HOG), gradient, or summed gradient features. HOG features are feature descriptors used in computer vision and image processing for the purpose of object detection. The technique counts occurrences of gradient orientation in localized portions of the image frame. Summed Gradient features are feature descriptors that counts the sum of gradient-x and sum of gradient-y in a selected sub-window of the image frame.

[0048] Optional temporal filter 310 may be configured to avoid giving false indication of the detection of a tongue-out condition. In embodiments, optional temporal filter 310 may be configured to apply filtering to the output of tongue-out classifier 308. More specifically, optional temporal filter 310 may be configured to apply filtering to the output of tongue-out classifier 308 to provide affirmative notification of the tongue-out condition, only after N successive receipt of tongue-out classifier outputs indicating detection of tongue-out. N may be a configurable integer, empirically determined, depending on the accuracy desired. For example, a relatively higher N may be set, if it is desirable to avoid false positive, or a relatively lower N may be set, if it is desirable to avoid false negative. In embodiments, if false negative is not a concern, temporal filtering may be skipped.

[0049] Referring now to Figure 5, wherein two example image frames of an example generated video, according to the disclosed embodiments, are shown. As described earlier, video generator 106 may be configured to capture the animation and rendering of animation-rendering engine 104 into a number of image frames. Further, the captured image frames may be combined / stitched together to form a video. Illustrated in Figure 5 are two example image frames 502 and 504 of an example video 500. Example image frames 502 and 504 respectively capture animation of two avatars corresponding to two characters speaking their dialogues 506. While dialogues 506 are illustrated as captions in example images 502 and 504, in embodiments, dialogues 506 may additionally or alternatively captured as audio (with or without companion captions). As illustrated by example image frame 502, tongue-out detector 122 enables efficient detection and animation of a tongue-out condition for an avatar / character.

[0050] While tongue-out detector 122 has been described in the context of VDFEE 112 to facilitate efficient detection of tongue-out conditions for the animation and rendering of avatars in the generation of avatar video, the usage of tongue-out detector 122 is not so limited. It is anticipated tongue-out detector 122 may be used in a wide variety of computer-vision applications. For example, tongue-out detector 122 may be used in interactive applications, to trigger various control commands in video games in response to the detection of various tongue-out conditions.

[0051] Additionally, while avatar video generation system 100 is designed to be particularly suitable to be operated on a mobile device, such as a smartphone, a phablet, a computing tablet, a laptop computer, or an e-reader, the disclosure is not to be so limited. It is anticipated that avatar video generation system 100 may also be operated on computing devices with more computing power than the typical mobile devices, such as a desktop computer, a game console, a set-top box, or a computer server.

[0052] Figure 6 illustrates an example computer system that may be suitable for use to practice selected aspects of the present disclosure. As shown, computer 600 may include one or more processors or processor cores 602, and system memory 604. For the purpose of this application, including the claims, the terms "processor" and "processor cores" may be considered synonymous, unless the context clearly requires otherwise. Additionally, computer 600 may include mass storage devices 606 (such as diskette, hard drive, compact disc read only memory (CD-ROM) and so forth), input / output devices 608 (such as display, keyboard, cursor control and so forth) and communication interfaces 610 (such as network interface cards, modems and so forth). The elements may be coupled to each other via system bus 612, which may represent one or more buses. In the case of multiple buses, they may be bridged by one or more bus bridges (not shown).

[0053] Each of these elements may perform its conventional functions known in the art. In particular, system memory 604 and mass storage devices 606 may be employed to store a working copy and a permanent copy of the programming instructions implementing the operations associated with face and facial expression engines 102, avatar animation-rendering engine 104 and video generator 106, earlier described, collectively referred to as computational logic 622. The various elements may be implemented by assembler instructions supported by processor(s) 602 or high-level languages, such as, for example, C, that can be compiled into such instructions.

[0054] The number, capability and / or capacity of these elements 610 - 612 may vary, depending on whether computer 600 is used as a mobile device, a stationary device or a server. When use as mobile device, the capability and / or capacity of these elements 610 - 612 may vary, depending on whether the mobile device is a smartphone, a computing tablet, an ultrabook or a laptop. Otherwise, the constitutions of elements 610-612 are known, and accordingly will not be further described.

[0055] As will be appreciated by one skilled in the art, the present disclosure may be embodied as methods or computer program products. Accordingly, the present disclosure, in addition to being embodied in hardware as earlier described, may take the form of an entirely software embodiment (including firmware, resident software, micro-code, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to as a "circuit," "module" or "system." Furthermore, the present disclosure may take the form of a computer program product embodied in any tangible or non-transitory medium of expression having computer-usable program code embodied in the medium. Figure 7 illustrates an example computer-readable non-transitory storage medium that may be suitable for use to store instructions that cause an apparatus, in response to execution of the instructions by the apparatus, to practice selected aspects of the present disclosure. As shown, non-transitory computer-readable storage medium 702 may include a number of programming instructions 704. Programming instructions 704 may be configured to enable a device, e.g., computer 600, in response to execution of the programming instructions, to perform, e.g., various operations associated with facial expression engines 102, avatar animation-rendering engine 104 and video generator 106. In alternate embodiments, programming instructions 704 may be disposed on multiple computer-readable non-transitory storage media 702 instead. In alternate embodiments, programming instructions 704 may be disposed on computer-readable transitory storage media 702, such as, signals.

[0056] Any combination of one or more computer usable or computer readable medium(s) may be utilized. The computer-usable or computer-readable medium may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium. More specific examples (a non- exhaustive list) of the computer-readable medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a transmission media such as those supporting the Internet or an intranet, or a magnetic storage device. Note that the computer- usable or computer-readable medium could even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, via, for instance, optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory. In the context of this document, a computer-usable or computer-readable medium may be any medium that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer- usable medium may include a propagated data signal with the computer-usable program code embodied therewith, either in baseband or as part of a carrier wave. The computer usable program code may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc.

[0057] Computer program code for carrying out operations of the present disclosure may be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0058] The present disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0059] These computer program instructions may also be stored in a computer-readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instruction means which implement the function / act specified in the flowchart and / or block diagram block or blocks.

[0060] The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0061] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0062] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an" and "the" are intended to include plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specific the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operation, elements, components, and / or groups thereof.

[0063] Embodiments may be implemented as a computer process, a computing system or as an article of manufacture such as a computer program product of computer readable media. The computer program product may be a computer storage medium readable by a computer system and encoding a computer program instructions for executing a computer process.

[0064] The corresponding structures, material, acts, and equivalents of all means or steps plus function elements in the claims below are intended to include any structure, material or act for performing the function in combination with other claimed elements are specifically claimed. The description of the present disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the disclosure in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill without departing from the scope and spirit of the disclosure. The embodiment was chosen and described in order to best explain the principles of the disclosure and the practical application, and to enable others of ordinary skill in the art to understand the disclosure for embodiments with various modifications as are suited to the particular use contemplated.

[0065] Referring back to Figure 6, for one embodiment, at least one of processors 602 may be packaged together with memory having computational logic 622 (in lieu of storing on memory 604 and storage 606). For one embodiment, at least one of processors 602 may be packaged together with memory having computational logic 622 to form a System in Package (SiP). For one embodiment, at least one of processors 602 may be integrated on the same die with memory having computational logic 622. For one embodiment, at least one of processors 602 may be packaged together with memory having computational logic 622 to form a System on Chip (SoC). For at least one embodiment, the SoC may be utilized in, e.g., but not limited to, a smartphone or computing tablet.

[0066] It will be apparent to those skilled in the art that various modifications and variations can be made in the disclosed embodiments of the disclosed device and associated methods without departing from the scope of the claims.

[0067] Thus, it is intended that the present disclosure covers the modifications and variations of the embodiments disclosed above provided that the modifications and variations come within the scope of the claims.

Claims

1. An apparatus for animating and rendering a plurality of avatars, comprising: one or more facial expression engines (102) to generate a plurality of animation messages having facial expression parameters that depict facial expressions for the plurality of avatars, the one or more facial expression engines comprising a text based facial expression engine (116) to receive conversations between characters via text inputs and, for each conversation, analyze the text inputs for semantics, and generate at least a subset of the plurality of animation messages having facial expression parameters depicting facial expressions for the plurality of avatars, including mouth movements of the plurality of avatars, representing at least in part semantics of the text inputs, each of the plurality of avatars corresponding to a character of the conversations, wherein the one or more facial expression engines further comprise a voice recognition facial expression engine (114) to receive audio inputs, analyze the audio inputs for at least volume or syllable, and generate at least a subset of the plurality of animation messages having facial expression parameters that depict facial expressions for the plurality of avatars, including mouth movements of the plurality of avatars, representing at least in part volume or syllable of the audio inputs; an animation-rendering engine (104) coupled with the one or more facial expression engines (102) to, for each conversation, receive the one or more animation messages, and drive a plurality of avatar models, in accordance with the plurality of animation messages, to animate and render the plurality of avatars with the facial expressions depicted; and a video generator (106) coupled with the animation-rendering engine to capture a plurality of image frames of the animation and rendering of the plurality of avatars for each conversation, determine that all conversations have been captured, and stitch the plurality of image frames of each conversation to generate a video.

2. The apparatus of claim 1, wherein the one or more facial expression engines further (102) comprise a video driven facial expression engine (112) to receive video inputs having a plurality of image frames, analyze the image frames, the video driven facial expression engine (112) is to analyze the image frames for landmarks of a face or head poses, and generate at least a subset of the plurality of animation messages having facial expression parameters that depict facial expressions for the plurality of avatars, including eye and mouth movements or head poses of the avatars, representing at least in part landmarks of a face or head poses in the image frames.

3. The apparatus of claim 1, wherein the one or more facial expression engines (102) comprise a video driven facial expression engine (112) that includes a tongue-out detector (122) to detect a tongue-out condition in an image frame.

4. A method for animating and rendering a plurality of avatars, comprising: receiving, by a computing device, conversations between characters via text inputs, for each conversation, wherein the receiving comprises receiving audio inputs; : generating, by the computing device, a plurality of animation messages having facial expression parameters that depict facial expressions for the plurality of avatars, the generating comprising analyzing the text inputs for semantics and generating at least a subset of the plurality of animation messages having facial expression parameters depicting facial expressions for the plurality of avatars, including mouth movements of the plurality of avatars, representing at least in part semantics of the text inputs, each of the plurality of avatars corresponding to a character of the conversation, wherein the generating comprises analyzing the audio inputs for at least volume or syllable, and generating at least a subset of the plurality of animation messages having facial expression parameters that depict facial expressions for the plurality of avatars, including mouth movements of the plurality of avatars, representing at least in part volume or syllable of the audio inputs; driving, by the computing device, a plurality of avatar models, in accordance with the plurality of animation messages, to animate and render the plurality of avatars with the facial expressions depicted; and capturing, by the computing device, a plurality of image frames of the animation and rendering of the plurality of avatars; determining that all conversations have been captured; and stitching the plurality of image frames of each conversation to generate a video.

5. The method of claim 4, wherein receiving comprises receiving video inputs having a plurality of image frames; and generating comprises analyzing the image frames for landmarks of a face or head poses, and generating at least a subset of the plurality of animation messages having facial expression parameters that depict facial expressions for the plurality of avatars, including eye and mouth movements or head poses of the avatars, representing at least in part landmarks of a face or head poses in the image frames.

6. The method of claim 5, wherein analyzing the video image frames comprises detecting a tongue-out condition in an image frame.

7. At least one computer-readable medium having instructions to cause a computing device, in response to execution of the instruction by the computing device, to perform any one of the methods of claims 4-6.