Method and application for animating computer-generated images

The method and system capture and process facial and body movements to animate CGI characters on mobile devices, addressing the limitations of current animations by providing customizable and efficient digital asset creation.

JP2025523804APending Publication Date: 2025-07-25FD IP & LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025500915
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-13
Filing Date
2023-06-29
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Current digital asset animations lack customizability and uniqueness, particularly in showing complex facial expressions and full-body movements, often requiring external hardware and lacking flexibility in application.

Method used

A method and system for capturing facial expressions and body movements using mobile devices, processing them to generate digital assets, and mapping these movements onto a skeleton model to animate CGI characters, enabling real-time or pre-captured enhanced video content.

Benefits of technology

Enables customizable and unique digital asset animations directly on mobile devices, allowing for realistic and efficient creation of animated CGI characters without the need for external hardware, enhancing video content with interactive virtual objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523804000001_ABST
    Figure 2025523804000001_ABST
Patent Text Reader

Abstract

This disclosure relates to digital asset animation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Related Applications] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 368,260, filed July 13, 2022, entitled "METHOD AND APPLICATION FOR ANIMATING COMPUTER GENERATED IMAGES", which is hereby incorporated by reference in its entirety.

[0002] This disclosure relates to digital asset animation.

Background Art

[0003] Video artists currently have the option of supplementing their content with digital assets (e.g., computer generated imagery (CGI)). For example, some computer applications provide augmented reality filters that, when used, can apply visual effects to the person being filmed. For example, current filters can add dog ears or a crown to the user, or morph the user's face into a dragon's face. These filters lack customizability, and much of the associated animation lacks uniqueness and is limited in showing the movement of the user's entire body and complex facial expressions. Some custom animations of computer-generated graphics for video content are complex and often require hardware external to a mobile phone (e.g., a smartphone). Thus, it is desirable to provide a method for video artists to create content that includes custom animations of CGI characters.

Summary of the Invention

[0004] To provide a basic understanding, various details of the present disclosure are summarized below. This summary is not an extensive overview of the present disclosure and is not intended to identify any elements of the present disclosure or to delineate its scope. Rather, the primary purpose of this summary is to present some concepts of the present disclosure in a simplified form prior to the more detailed description presented below.

[0005] The method may include a processor generating face data characterizing captured face parts and / or expressions, the processor generating body data characterizing captured body movements, the processor mapping the captured face parts and / or expressions and the captured body movements onto a skeleton model, and the processor animating a digital asset based on the mapping.

[0006] In another example, the system may include a memory storing machine-readable instructions and data, and one or more processors accessing the memory and executing the machine-readable instructions. The machine-readable instructions may include a body capture module providing face data characterizing parts and / or expressions of a target face, a face capture module providing body data characterizing movements of the target body, a character module obtaining a digital asset from a digital asset database, and a skeleton module mapping movements of the target body and / or expressions of the target face onto the digital asset based on the body and / or face data to animate the digital asset, where extended video data is generated using the animated digital asset.

[0007] In yet another example, the method may include receiving, at a processor, first video signal data characterizing facial parts and / or expressions, receiving, at the processor, second video signal data characterizing body movements, the processor extracting the facial parts and / or expressions to provide facial data, the processor extracting the body movements to provide body data, the processor mapping the body movements and / or facial parts and / or expressions based on the facial and / or body data to a skeleton model to provide an animated skeleton model, and the processor animating a digital asset according to the actions of the animated skeleton model to provide an animated digital asset.

Brief Description of the Drawings

[0008] Embodiments will be described with reference to the accompanying drawings. In the drawings, like reference numerals may indicate identical or functionally similar elements. The drawing in which an element first appears is generally indicated by the leftmost digit of the corresponding reference numeral.

[0009]

Figure 1

[0010]

Figure 2

[0011]

Figure 3

[0012]

Figure 4

[0013]

Figure 5

[0014]

Figure 6

Mode for Carrying Out the Invention

[0015] Next, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Similar elements in the various figures may be denoted by similar reference numerals for consistency. Further, in the following detailed description of the embodiments of the present disclosure, numerous specific details are set forth in order to provide a more thorough understanding of the claimed subject matter. However, it will be apparent to those skilled in the art that the embodiments disclosed herein may be practiced without these specific details. In other instances, well-known features are not described in detail in order to avoid unnecessarily complicating the description. Further, it will be apparent to those skilled in the art that the scale of the elements presented in the accompanying drawings may vary without departing from the scope of the present disclosure.

[0016] This specification discloses examples of systems and methods for capturing the body movements and / or facial expressions of a subject (e.g., an actor), and generating digital assets based on the captured body movements and / or facial expressions. As used herein, the term “facial expression” or “facial movement” means the movement of the facial parts (eyes, mouth, nose, skin) on an individual's face. This includes, but is not limited to, eye movement, speaking, raising eyebrows, smiling, frowning, emotion, sticking out the tongue, moving ears, and puffing out the cheeks. Digital assets may be provided, and the captured body movements and / or facial expressions may be mapped onto the digital assets to animate the digital assets to reflect the actions of the digital assets. In some examples, the digital assets may be scaled and mapped according to the actions of the subject. In some examples, the mapped digital assets may be combined with a video signal to enable a user to view the digital assets on a display. In some examples, the body movement and facial expression data (calculated based on the captured video signal data of the subject) may be stored as collective body data, and the digital assets may be later mapped according to the collective body data.

[0017] This specification presents examples related to the creation of video clips using digital assets (or extended video data), and the systems and methods described herein can be used in various applications and / or industries such as the film and / or television industries. Thus, in some examples, the techniques / methods disclosed herein for capturing the movement of a subject's body and / or facial expressions, and for generating digital assets based on the captured body movement and / or facial expressions, can be incorporated into video generation systems related to film or industry. For example, a previsualization device and / or system may be configured to include the systems and / or methods disclosed herein for animating computer-generated images. An example of such a device / system is described in U.S. Patent Application No. 17 / 410,479, filed on August 24, 2021, entitled "Previsualization Devices and Systems for the Film Industry", which is hereby incorporated by reference in its entirety. Thus, the techniques / methods disclosed herein can be used in previsualization devices / systems to animate assets based on body movement and / or facial expressions during filming or production.

[0018] FIG. 1 is a block diagram of an animation system 10 that can be used to provide extended video content. Extended video content (or video data) can refer to video images that are enhanced or modified (either in real-time or pre-captured) using digital assets. Thus, extended video content can include animated video content (e.g., video images created using computer graphics rather than recorded using a camera). In the context of extended video, animated video data can be used to create virtual objects or characters that interact with the real-world environment captured by a camera. For example, an animated character can be placed within a real-world scene and made to interact with objects or people within the scene to create the illusion that the character exists in the real world.

[0019] In the example of FIG. 1, the animation system 10 includes a mobile device 100 that can be used to create extended video content. Although an example where the mobile device 100 is a smartphone is presented herein, in other examples, the mobile device 100 may be implemented as a tablet, a wearable device, a handheld game console, or as a portable device such as a laptop. In yet another example, the mobile device 100 may be implemented as a computer including a desktop computer or a portable computer.

[0020] Accordingly, in some examples, device 100 may include one or more image sensors (collectively shown as image sensor 102) for capturing the environment, a processor 106, and a memory medium / memory 108. Processor 106 may be variously embodied by, for example, a single-core processor, a dual-core processor (or multi-core processor), a digital processor and a cooperating numeric coprocessor, a digital controller, a graphics processing unit (GPU), a dedicated processor, an application-specific processor, an embedded processor, and the like. Processor 106 may include multiple processors, and each of the processors may include multiple cores. Processor 106 may be configured to control the operation and components of device 100 and may execute applications, apps, and instructions stored in memory 108 and / or accessible via communication device 110. Memory 108 may represent a non-transitory machine-readable memory (or other medium) such as random access memory (RAM), solid state drive, hard disk drive, and / or combinations thereof. In some embodiments, processor 106 and / or memory 108 may be implemented on a single die (e.g., chip) or on multiple chips.

[0021] The image sensor 102 can be, for example, a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) sensor configured to capture an environment and provide visual data characterizing the captured environment as, for example, one or more images. The image sensor 102 can be any type of device / system capable of capturing an environment and providing corresponding visual data. As another example, the image sensor 102 can be a camera that is analog, digital, and / or a combination thereof. The image sensor 102 is used to create an image or video corresponding to the visual data and detect and transmit the information stored in the memory 108.

[0022] In some examples, the image sensor 102 includes multiple cameras. As an example, the image sensor 102 includes a first camera 103 and a second camera 104, although in other examples a greater number of cameras may be used. In some examples, the first and second cameras 103 and 104 may be configured (or optimized) for a particular imaging / capturing purpose. Thus, each of the first and second cameras 103 and 104 may be configured for a certain use or associated with a specialty lens (e.g., a macro lens, an ultra-wide lens, a telephoto lens, a wide-angle lens, etc.). Thus, each of the first and second cameras 103 and 104 may be different from each other and may have unique features that make it suitable for various types of photography and videography.

[0023] In one example, the first camera 103 can be used to capture the parts and / or expressions of the face of a subject (e.g., an actor or a person), and provide first video signal data including the captured parts and / or expressions of the subject's face. The processor 206 can process the first video signal data to extract the captured parts and / or expressions of the face, and provide face data. The face data can characterize the captured parts and / or expressions of the subject's face. That is, in some embodiments, the first camera 103 can have a resolution and / or zoom for capturing details about the subject's face from a certain distance.

[0024] In some examples, the second camera 104 can be configured to capture the body of the subject (e.g., feet, head, and everything in between) and the movement of the subject, which can be provided as second video signal data. The second video signal data can be processed by the processor 106 to extract the movement of the subject and provide body data characterizing the movement of the subject. The subject can be a human, but in other examples, the subject can be a robot, an animal, or various types of objects that have a body and can experience movement. In some examples, the expression of the subject's face and the movement of the subject's body are captured by a single camera. In some examples, the face data is provided based on the parts and / or expressions of the face of a first subject, and the body data is provided based on the movement of the body of a second subject. In some examples, similar video signal data is used to generate the face and body data.

[0025] In some examples, the image capture and / or extraction of the facial expression can be implemented by the first camera 103, and the image capture and / or extraction of the movement of the subject's body can be implemented by the second camera 104. Thus, in some examples, separate cameras such as the first camera 103 and the second camera 104 can be utilized to extract the facial expression and the body movement of the subject.

[0026] In some examples, facial parts and / or expressions can be captured simultaneously with the movement of the subject's body. In other examples, the facial parts and / or expressions of the subject, and / or the movement of the body are captured independently. For example, first the movement of the subject's body (e.g., dance) may be captured, and then the facial expression of the subject (e.g., singing with emotion) may be captured, or vice versa.

[0027] In some examples, device 100 may include a communication device 110 that can be used to communicate with other devices 100n, server 160, and / or Internet 170 that may be configured similarly to device 100 using network 164. There may be "n" number of other devices. Network 164 may include any type of network, such as a local area network (LAN) (e.g., Wi-Fi (registered trademark) network), cellular network, Bluetooth (registered trademark) network, NFC network, VPN network.

[0028] Communication device 110 may include, for example, a wired communication interface, a wireless communication interface, a cellular interface, a short-range interface, a Bluetooth (registered trademark) interface, a Wi-Fi component, and / or other communication components / interfaces that provide communication via other means. Using device 100, data (e.g., face data, body data, other types of data) or an extended video can be uploaded to an online video platform via Internet 170. In some examples, device 100 can provide face and / or body data to other devices 100n, and other devices 100n can provide extended video content based on the face data and / or body data.

[0029] Continuing with the example of FIG. 1, device 100 may include a user interface 112 for receiving commands from a user of device 100. The user interface 112 may include, for example, a touch screen device, keyboard, mouse, motion sensor, button, knob, voice activation, headset, hand recognition, eye gaze recognition, and / or the like.

[0030] Processor 106 may access memory 108, storage 161 on or coupled to server 160, and / or cloud-based storage on Internet 170. In some embodiments, memory 108 or storage 161 may include a database of digital assets (e.g., CGI characters) that can map the movement of the body and / or facial expressions of the subject thereon. As disclosed herein, after capturing the movement of the body and / or facial expressions of the subject, digital assets (or a plurality of assets) can be selected from the database (e.g., using interface 112). The movement of the body and / or facial expressions of the subject can be mapped to digital assets based on body movement data and / or face data. In some examples, while an actor is being filmed or later, a user looking at display 114 of device 100 can view a CGI character with actions mapped to the actor (e.g., facial and / or body movements). As another example, an actor may be recorded by image sensor 102, and the movement of the actor's body and / or facial expressions may be extracted by image sensor 102 from the captured image and stored as data for the placement of digital assets with animation, based on subsequent digital asset mapping and / or mapping into a virtual or other recorded environment.

[0031] Figure 2 is a block diagram of an example of another animation system 200. The animation system 200 can include a processing system representing a computer system (or platform) 201, which can implement at least some of the functions / operations disclosed herein. The various components shown in system 200 of FIG. 2 are for the purpose of illustrating aspects of an exemplary embodiment, but it is possible to substitute it with other similar components implemented via hardware, software, or a combination thereof. The computer system 201 may be implemented, for example, as a desktop computer, a portable computer, a tablet, a mobile device (such as device 100 shown in FIG. 1), or as other known devices that can host a software platform and / or application. In some examples, since the computer system 201 can be implemented as device 100, reference may be made to the example of FIG. 1 in the example of FIG. 2. The computing platform 201 can be implemented within a computing cloud and on one or more servers including a server farm or a cluster of servers. In such a situation, the features of the computing platform 201 may represent a single instance of hardware, or multiple instances of hardware (e.g., distributed) of multiple instances of an application executed across hardware (such as a computer, a router, a memory, a processor, or a combination thereof). Alternatively, the computing platform 201 may be implemented on a single dedicated server or workstation.

[0032] The computer system 201 includes a processor 206 and a memory 208, which may be implemented similarly to the processor 106 and the memory 108 shown in FIG. 1. As disclosed herein, the computer system 201 can provide extended video data based on video data provided by an imaging sensor 202 of an environment with an embedded digital asset (e.g., a CGI character) having an animation based on body and / or face data. In some examples, as disclosed herein, the extended video data is provided based on environmental data. As shown in FIG. 2, the extended video data can be rendered on a display 210. The computer system 201 may include a user interface 207 that enables a user to provide the computer system 201 with commands for controlling various components of the animation system 200 and, in some examples, the generation of extended video data. In some examples, the computer system 201 may include an operating system, such as Windows®, Linux®, Apple®, Android®, or any other type of operating system. In some embodiments, the computer system 201 includes a graphics processing unit (GPU) that can process video data provided by the imaging sensor 202 to generate extended video data. As shown in FIG. 2, the imaging sensor 202 may include any number of cameras, such as cameras 203, 204.

[0033] In the example of FIG. 2, the components of computer system 201 may be connected to bus 220 to enable the exchange of data between components, for example, to provide extended video data. For example, processor 206 may communicate with data storage 261 using communication link 242. Cameras 203, 204 may provide their respective video data to processor 206 using communication link 242. In some examples, communication link 242 is a direct electrical wiring between components. In some examples, processor 206 may use communication link 243 to communicate with other components including server 160, cloud network 170. Each component is shown as being connected to computer system 201 via one of the two shown communication links 242, 243, but in other examples, multiple links 242, 243 may be used, and any component may use the corresponding link to transfer data to processor 206. Exemplary communication links may include, for example, proprietary communication networks, infrared, optical, or other suitable wired or wireless data communication links. The example of FIG. 2 shows computer system 201 communicating directly with server 160, but in other examples, computer system 201 may communicate with server 160 using network 170. In some examples, network 170 is implemented similar to network 164 shown in FIG. 1.

[0034] In some examples, the processor 206 may be configured to execute machine-readable instructions 211 stored within the memory 208 or at another memory location accessible by the processor 206. The machine-readable instructions 211 may include a body capture module 230 that can process video data from the image sensor 202, which may include video signals generated by the camera 203 and / or 204. In some examples, the video signal may be referred to as a raw video signal. The video signal may be characterized by a plurality of pixels supported in the horizontal direction (e.g., HD, 2K, and 1080P, also known as BT.709). In some examples, the camera 203 provides a first video signal 203a that includes one or more images representing the environment. For example, the first camera 203 may have a zoom and resolution defined or set to capture the movement of the body of the subject 240 (or actor) (e.g., feet, head, and everything in between). The first video signal 203a may be provided to the body capture module 230 for processing as the first video signal data.

[0035] In some examples, the machine-readable instructions 211 may also include a face capture module 232 that can receive first video data from the image sensor 202. In some embodiments, the face capture module 232 receives the first video signal data received by the body capture module 230. In some examples, the camera 204 provides a second video signal 203b that includes one or more images representing the environment. The second video signal 203b may be provided to the face capture module 232 (or in some examples to the body capture module 230) for processing as second video signal data. The second camera 204 may have a zoom and resolution set or defined to capture the facial expressions of the actor within the video frame. Thus, in some examples, the zoom and resolution of the second camera 204 may be such that the face (e.g., the head) is captured at a level of detail such that facial expressions are captured. For example, the level of detail captured by the second camera 204 may be such that the position (or movement) of the eyes of the subject 240 can be detected. Thus, the features of the subject 240 that can be captured may include where the subject 240 is looking, how the subject 240 is speaking, and the various emotions exhibited by the subject 240.

[0036] The body capture module 230 may provide body data characterizing the body movements of the subject 240 based on the first video signal data (or in other examples based on the second video signal data). The face capture module 232 may provide face data characterizing the captured facial parts and / or expressions of the subject 240 based on the second video signal (or in other examples based on the first video signal data). For example, the body capture module 230 can extract the body movements of the subject 240 from the first video signal data, and the face capture module 232 can extract the facial parts and / or expressions of the subject 240 from the second video signal data.

[0037] In some examples, each capture module 230, 232 may be configured to change the video coding format of the video signal from the image sensor 202. That is, the coding format of the video signal received from the image sensor 202 may be changed to a second format by the capture modules 230, 232. The change in format may facilitate additional processing of the video signal. Examples of video coding formats include, but are not limited to, H.262 (MPEG-2 Part 2), MPEG-4 Part 2, H.264 (MPEG-4 Part 10), HEVC (H.265), Theora, RealVideo RV40, VP9, and AV1. In other examples, the capture modules 230, 232 may be configured to change the video codec of the video signal.

[0038] In yet another embodiment, the capture modules 230, 232 may be configured to change the compression of the video signal from the image sensor 202. That is, the video signal may be compressed to have a smaller associated file size than its original format. In embodiments where the capture modules 230, 232 compress the video signal, the processing of the video signal by the processor 206 and other devices may be faster compared to an uncompressed video signal. The capture modules 230, 232 may also optimize the performance characteristics of the video signal, for example, but not limited to, frames per second (FPS), video quality, and resolution. The optimization may also enable the video signal to be processed in a more efficient and faster manner.

[0039] In some examples, as disclosed herein, the machine-readable instructions 211 can include a character module 234 that can obtain a digital asset (which can be referred to as a character in some examples) and map the selected digital asset to a skeleton model. The digital asset can be stored in a database accessible to the character module 234 on a storage device 261 or via cloud storage or remote storage 161 that communicates with the server 160. The digital assets in the database can be defined as having a shape in two or three dimensions. For example, a digital asset (e.g., a CGI character) can have a three-dimensional (3-D) body and a predefined dimensional scale (e.g., the CGI character can be configured to have a scaled size that is approximately 50 meters tall relative to the object 240 in the extended video data). Thus, when movement is mapped to the digital asset and inserted / incorporated into the video signal, the digital asset can be displayed much larger than the object 240.

[0040] In some embodiments, the digital asset can be uploaded directly or remotely (via the cloud 170 or a remote server 160) and be accessible by the character module 234. In some embodiments, the digital asset can include other assets such as, for example, an environment, a prop, a set extension, etc. in addition to the character and can also be stored in the database. In some embodiments, the character module 234 can be configured to change the level of detail (LOD) of the digital asset, e.g., the geometry detail and pixel complexity of the digital asset. By using LOD techniques, the efficiency of dimension rendering can be increased by reducing the load of graphic processing.

[0041] The machine-readable instruction 211 can include a skeleton module 236, which can animate a (virtual) skeleton or skeletal model of the object 240 based on the body movements and / or facial expressions captured by the cameras 203, 204, and provide an animated skeleton model. For example, the skeleton module 236 can receive face and / or body data and animate the skeleton or skeletal model to reflect the body movements and / or facial expressions captured by the cameras 203, 204. The skeleton module 236 can map the animated skeleton model to the digital assets (provided by the character module 234) to animate the digital assets according to the body movements and / or facial expressions of the object 240 captured by the cameras 203, 204. Therefore, the human actions (body movements and / or facial expressions) recorded by the image sensor 202 can be animated by the digital assets in one or more dimensions (e.g., 2 or 3 dimensions) based on the mapping of the animated skeleton model to the digital assets.

[0042] In some examples, a digital asset includes a digital asset frame (e.g., a CGI frame) and digital asset joints. A digital asset frame is a single still image from an animation or video sequence created using computer graphics. The frame may be considered a snapshot of a moment within the animation and, when combined with other frames, creates an illusion of motion. A joint is a connection point between two or more bones within a skeleton. In computer graphics, a joint is typically represented as a pivot point around which a bone can rotate, enabling the movement of the skeleton. In one example, to map the movement of a skeleton model's skeleton to a digital asset frame, motion data (e.g., body and face data) captured using cameras 203, 204 can be used to drive the movement of a virtual skeleton within a digital asset environment. The motion data may be processed by a skeleton module 236 and applied to the joints and bones of the virtual skeleton, thereby enabling the movement of the digital asset to be animated in a realistic and believable manner. The movement of the virtual skeleton is then used by the skeleton module 236 to animate the digital asset frame by positioning and rotating the joints and limbs of the digital asset to match the movement captured within the real-world environment.

[0043] In some examples, the body movements and facial expressions of subject 240 are captured ignoring the visual appearance of subject 240, and the combined movement data is mapped to a skeleton model or skeleton such that the digital asset performs the actions and facial expressions of subject 240. In the context of shooting, if the user desires to create a video of a dancing character, subject 240 can perform the action of dancing, and the dancing action is captured by image sensor 202 and transmitted to body capture module 230 and / or face capture module 232 for processing, and thus can provide face and / or body data. Skeleton module 236 receives body data from body capture module 230, extracts the dancing movements as the captured skeleton or series of joints of subject 240, and animates a digital asset having the corresponding skeleton or series of joints.

[0044] In some examples, skeleton module 236 receives the captured body movements and facial expressions from capture modules 230, 232, combines the extracted body data and face data, and generates synthesized target data that is a single data set for animating the movements and expressions of the digital asset. In some examples, skeleton module 236 uploads and / or stores the synthesized target data in local storage 261 or remote storage 161 and can later retrieve the synthesized actor data to generate extended video data. In some examples, capture modules 230, 232, character module 234, and skeleton module 236 can cooperate to generate an extended video signal, which is provided to display 210 and / or can be stored for later viewing / sharing on an online video platform. The extended video signal can include a video signal of the digital asset modeled on the environment and subject 240. Thus, the user of system 200 can view a video of the digital asset moving as subject 240 moves.

[0045] In some embodiments, a system, such as system 200, includes additional modules and hardware configured to virtually map an environment and determine the position of a camera relative to the environment. The additional captured and processed data enables a unique filming experience of an animated CGI character as described above. For example, a user of the system can walk around a physical location in the real world where a CGI character is virtually associated, and then film the CGI character mapped to the animation as if the CGI character actually exists within the real environment like a real actor.

[0046] In some examples, system 200 includes an environmental sensor 250 for capturing environmental data characterizing the environment. The environmental sensor 250 can scan the environment and determine local geometry and spatial configuration. For example, as described herein, the environmental sensor 250 can capture depth points used to construct a scaled virtual 3D model of the environment, which can be used to generate extended video data. The environmental sensor 250 can be implemented as, for example, an infrared system, a light detection and ranging (LIDAR) system, a thermal imaging system, an ultrasonic system, a stereoscopic system, an RGB camera, an optical system, and / or any device / sensor system capable of measuring the depth and distance of objects within the environment.

[0047] As another example, the environmental sensor 250 may include an infrared emitter and sensor. The infrared emitter projects a known pattern of infrared dots into the environment that are not included in the visible spectrum of the human eye. The projected dots may be captured by either an infrared sensor or an image camera 203, 204 for analysis. In some embodiments, the environmental sensor 250 may include a LIDAR-based system that can project a pulsed laser (signal) into the environment and use the amount of time it takes for the laser signal to return to generate a 3-D model of the environment. In some examples, any sensor system or combination of sensor systems may be utilized as the environmental sensor 250.

[0048] In some examples, the system 200 may include a motion sensor 238 for detecting the orientation of the device 100 and / or cameras 203, 204. In some examples, the cameras 203, 204 are integrated within a device such as the device 100. Thus, the examples herein related to measuring and determining environmental features and device orientation and position data may include detecting the orientation and position of the cameras 203, 204 within the environment. The motion sensor 238 may be a sensor or combination of sensors capable of detecting the motion, orientation, acceleration, and position of the device 100 and / or cameras 203, 204. The motion sensor 238 may be variously embodied as a gravity sensor, accelerometer, gyroscope, magnetometer, and the like, or a combination thereof. For example, the motion sensor 238 may be an inertial measurement unit (IMU) corresponding to a detection unit that may include an accelerometer, a gyroscope, and a magnetometer. In this way, the animation system 200 can calculate the position and orientation of the cameras 203, 204 relative to the generated 3-D model of the environment with reference to the origin.

[0049] In some examples, the machine-readable instructions 211 may include an environment module 252 that can receive environmental data from the environmental sensor 250. The environment module 252 processes the environmental data, locates objects and / or points, and determines distances between such objects and / or points in the environment. The environment module 252 uses the environmental data generated by the environmental sensor 250 to create a 3-D model / mesh of the environment. In some embodiments, the environment module 252 utilizes photogrammetric algorithms and uses one or a combination of the video signals generated by the camera and the environmental data generated by the environmental sensor 250 to generate a 3-D model of the environment. For example, multiple photographs from an image sensor or data sensor can be stitched together to construct a three-dimensional model of the environment. In other embodiments, point data generated from points observed in the real environment, such as point clouds (either object features or infrared point irradiations), is converted into a mesh (polygon or triangular mesh) model representing the real environment. In other embodiments, a neural radiance field (NeRF) is used to generate a 3-D model of the environment using one or a combination of the video signals generated by the camera and the environmental data generated by the environmental sensor.

[0050] In some examples, the machine-readable instructions 211 may include a tracking module 254 that determines the position of the device 100 and / or the cameras 203, 204 relative to the real-world environment and the 3-D model generated by the environment module 252. The tracking module 254 receives spatial tracking data from motion sensors 238 associated with the positioned device and / or cameras. Thus, as the operator of the camera moves to capture various angles of the environment or within the environment, the movement calculated by the tracking module can be taken into account in real time when displaying the digital asset. In other words, tracking of the device 100 and / or the cameras 203, 204 ensures that the digital asset remains in the selected position even as the perspective of the environment and the cameras changes, as if viewed from the changed camera perspective.

[0051] In some examples, the body capture module 230, the face capture module 232, the character module 234, the skeleton module 236, the environment module 252, and the tracking module 254 cooperate (or function together) to generate an extended video signal provided to the display 210 that can be recorded, stored, and shared as an animated production video. That is, the real-time animated video signal includes the video signal captured by the camera with digital assets that are displayed as being located or existing within the real-world environment.

[0052] In view of the foregoing structural and functional features described above, the exemplary method may be better understood with reference to FIGS. 3-4. For purposes of brevity, the exemplary method of FIGS. 3-4 is shown and described as being performed continuously, but in other examples, some actions may be performed multiple times and / or simultaneously in a different order than shown and described herein, so it should be understood and recognized that this example is not limited by the order shown. Further, it is not necessary that all of the described actions be performed to implement the method.

[0053] Figure 3 is a flowchart diagram of a method 300 for providing an animated video that can be implemented by a system such as system 10 shown in FIG. 1 or system 200 shown in FIG. 2. Thus, in the example of FIG. 3, reference may be made to the examples of FIGS. 1-2. Method 300 may begin at 302 by capturing (using, for example, image sensor 102 shown in FIG. 1 or image sensor 202 shown in FIG. 2) the body movements and facial expressions of a subject (such as subject 240 shown in FIG. 2). In some examples, separate cameras may be used to capture the body movements and facial expressions, or the same camera may be used. In some examples, the body movements and facial expressions of the actor may be captured simultaneously, or at different times from each other. The body movements and facial expressions of the subject may be captured and provided as video signal data, for example, as disclosed herein.

[0054] As shown in FIG. 3, the video signal data corresponding to the body movements and facial expressions of the actor may be represented as video signal data 301. At 304, the video signal can be processed to optimize the video signal data. Such processing includes optimizing the video signal data for additional processing including, but not limited to, changes in compression, format, codec, and the like. The optimization at 304 can provide optimized signal data 304, thereby potentially making subsequent processing faster. The processing may be performed, for example, by capture modules 230, 232. In some examples, step 304 may be omitted from the exemplary method 300.

[0055] At 306, body movements and facial expressions can be extracted from the video signal data 301 (or the optimized video signal data 303). For example, the skeleton module 236 shown in FIG. 2 receives the video signal data 301 or 302 from the capture module 232, extracts feature data as the skeleton or a series of joints captured from the subject 240, and provides the extracted feature data 305. The extracted feature data 305 representing body movements and / or facial expressions can be configured for additional processing. In some examples, as disclosed herein, the extracted feature data 305 corresponds to or includes body and / or face data.

[0056] At 308, feature data representing body movements and facial expressions can be combined to create a combined set of feature data 307. In some embodiments, the combined feature data 307 is further processed by a method and associated system, such as system 200. In some embodiments, the combined feature data 307 may be stored in the memory of a user device (such as device 100 shown in FIG. 1) at 310, or may be stored in the cloud or on a remote storage device (such as disclosed herein). At 312, an associated system (such as system 200) obtains digital assets from an accessible database that is stored either on the device or remotely; this may be referred to as digital asset data 309 (referred to as CGI character data 309 in the example of FIG. 3). The digital asset data 309 may include character skeleton and / or joint elements configured to be associated with the combined feature data 307 extracted from body movements and facial expressions.

[0057] In 314, the combined feature data 307 can be mapped to the digital asset data 309 to create a digital asset animation (or animated digital asset) 311. For example, the captured and extracted arm movements can be mapped to the digital asset, and as a result, the digital asset mimics the captured arm movements. As a result, the digital asset is then animated using the captured movements and facial parts of the subject without the need for a conventional motion capture system (e.g., involving the use of a motion capture suit and markers).

[0058] In some examples, in 316 of method 300, the animated digital asset 311 is inserted into the environment. In some examples, the environment is a CGI background such as a static image and / or an animated background. In other embodiments, the animated digital asset 311 is inserted into a recorded or ongoing (or real-time) video feed captured by an image sensor of the system, such as image sensor 202 of system 200 shown in FIG. 2. In some examples, system 200 may include other sensors, such as an environmental sensor, configured to map the captured environment and scale the animated digital asset 311 desirably within the captured environment. In this way, the animated digital asset 311 can be reused by inserting the character into various environments to create various videos.

[0059] In some examples, at 318, an animated production video 313 that includes background (environmental / virtual) and internally scaled and animated digital assets 311 can be made viewable as video content on a display, such as display 210 shown in FIG. 2. This process 300 may occur during filming, and it should be understood that the animated digital assets 311 can be overlaid on an actor within the viewfinder display of the capture device. Also, the animated production video 313 may be captured, recorded, and stored on a storage device such as the memory device mentioned above, or the video animation may be uploaded / shared to an online video platform.

[0060] FIG. 4 is a block diagram of a digital asset recommendation system 400 that can be used in one of the animation systems as shown in FIGS. 1 - 2. System 400 can be implemented on a computing device / platform such as device 100 shown in FIG. 1, or system 201 shown in FIG. 2, device 100n shown in FIGS. 1 - 2, etc., or on server 160 shown in FIGS. 1 - 2. Thus, in the example of FIG. 4, reference may be made to the examples of FIGS. 1 - 2. In some examples, system 400 can be implemented as a service for which payment is made for use within a cloud environment or over the Internet (such as Internet 170 shown in FIG. 1).

[0061] System 400 can be used to provide digital asset recommendations to a user (or a director when used in the context of a movie / TV application). The digital asset recommendations can include, for example, the type of digital asset, the attributes of the digital asset (such as clothing, weapons, size, color, shape, degree of brightness and darkness, etc.). System 400 can provide digital asset recommendations based on video signal data 402. The video signal data 402 can be provided by one or more cameras, such as cameras 203 and 204 shown in FIG. 2. The video signal data 402 can include a plurality of video frames generated by one or more cameras of the environment. In some examples, the environment can include an object (such as object 240 shown in FIG. 2).

[0062] System 400 can include an environment analysis unit 404. The environment analysis unit 404 can evaluate the video signal data 402 to detect or identify features from the environment and provide the detected / identified features as environment analysis data 406. For example, the environment analysis unit 404 can use object detection techniques to identify objects within the video signal data 402 and locate their positions. In some examples, the environment analysis unit 404 can use a convolutional neural network (CNN) model trained for image detection, and thus object detection, or various types of machine-learning (ML) models trained for image detection. The type of objects that can be detected using an object detection algorithm by the environment analysis unit 404 can depend on the training data used to train the ML model. The environment analysis unit 404 can detect objects and classify the objects into various categories based on their appearance and characteristics. The types of objects identified by the environment analysis unit can include, for example, people, vehicles, animals, furniture, electronic devices, buildings, landmarks, food, tools, sports equipment, and other types of objects.

[0063] In some examples, the environmental analysis unit 404 may identify objects according to the object rule 408. The object rule 408 may define or specify the type of (or specific) objects that the environmental analysis unit 404 should detect. The object rule may be provided by another system, software, or based on user input (in a user interface such as the input device, e.g., the user interface 112 shown in FIG. 1). Thus, the object rule 408 may adjust or modify the environmental analysis unit 404 to exclude some objects so as not to identify them. For example, when the environment is a beach, the object rule 408 can be used to exclude buildings from the considerations, and thus improve the efficiency and speed at which the environmental analysis unit 404 provides environmental analysis data 406. By using the object rule 408, the environmental analysis unit 404 is modified / configured to minimize or reduce false detections included in the environmental analysis data 406, and thus improve the overall accuracy at which the system 400 provides digital asset recommendations. In some examples, when an object is detected, its features may be extracted for further analysis and provided as part of the environmental analysis data 406. This can be done using techniques such as feature extraction that involve analyzing the appearance of the object and extracting relevant features such as color, texture, and shape.

[0064] The system 400 may include a digital asset recommendation engine 410. The digital asset recommendation engine 410 may process the environmental analysis data 406 to provide digital asset recommendation data 412. For example, the digital asset recommendation engine 410 may use the environmental analysis data 406 to evaluate the identified objects in the video signal data 402 and predict the digital assets most relevant to that environment.

[0065] In some examples, the engine 410 uses a machine learning (ML) algorithm (e.g., a trained ML model) trained to process identified objects / object characteristics from video data. For example, the ML algorithm of the engine 410 can be trained to analyze object features to make predictions about digital assets. For example, if the video data includes a ship, the ML algorithm can be trained to analyze the characteristics (e.g., properties and attributes) of the ship and recommend one or more possible digital assets (e.g., crew members, fish, other types of sea life) that can be used. The ML algorithm of the engine 410 can be implemented as a deep neural network model, a decision tree model, an ensemble model, or various types of ML models. Thus, the ML algorithm can be trained with respect to multiple objects for which the most relevant digital assets (e.g., CGI assets) are identified. Accordingly, the digital asset recommendation data 412 can identify multiple digital assets most relevant to the environment.

[0066] For example, during filming (production), the system 400 can identify the digital assets that may be needed for that scene, thereby eliminating the need to search for and identify the assets that can be used in that scene. Thus, once the creative direction of the scene is determined, the system 400 can be used to identify possible digital assets that match the creative direction of that scene.

[0067] System 400 may include a digital asset acquisition unit 414. The digital asset acquisition unit 414 may be used to query or search a digital asset database 416 based on digital asset recommendation data 412 to provide one or more digital assets 418. The digital asset database 416 may include a plurality of pre - constructed digital assets. In some examples, the digital asset acquisition unit 414 may evaluate technical feasibility during the identification of digital assets from the digital asset database 416. For example, the digital asset acquisition unit 414 may consider factors such as asset complexity, polygon count, level of detail, etc., and / or assess whether the digital asset (identified by data 412) can be efficiently rendered in the environment without affecting performance. In some examples, the digital asset acquisition unit 414 may optimize one or more selected digital assets. For example, the digital asset acquisition unit 414 can simplify the geometry of the asset, reduce the texture resolution, and / or adjust the level of detail to ensure that it is optimized for performance. In some examples, the digital asset acquisition unit 414 can render the digital asset recommendation data 412 on an output device (e.g., a display), and the user can select which assets will be provided as one or more digital assets 418. Once one or more digital assets 418 are provided, these assets may be supplied to a composition unit (e.g., a module, an application, dedicated hardware / device) to be combined with one or more images of the environment to provide extended video data, and this extended video data may be rendered on the same or a different output device.

[0068] In some examples, video signal data 402 may be received by a target action analysis unit 420, which can analyze the video signal data 402 to detect or identify a target action (e.g., body or face movement) in the environment. The target action analysis unit 420 may provide target action data 422 that identifies the target action. In some examples, the target action analysis unit 420 can use ML trained to process video of a moving human to detect the type of movement. The target action analysis unit 420 can use action recognition techniques to identify a specific action being performed by a target in the video signal data 402. In some examples, the ML algorithm of the target action analysis unit 420 is a trained CNN, although other types of ML models may also be used. For example, the ML algorithm can be trained on preprocessed video (with human movement and appearance of people) to learn patterns and features that distinguish different actions from each other.

[0069] In some examples, the ML algorithm of the target action analysis unit 420 can be trained to measure or determine the amount of hand movement and / or identify / interpret specific gestures made by a person. The target action analysis unit 420 can provide target action data 422 that characterizes the measured amount of hand movement and / or the type of gesture the target is making. The target action analysis unit 420 can use a computer vision algorithm to track the position and movement of the target's hand in 3D space over time (e.g., within an environment). In this case, the amount of hand movement can be calculated by the target action analysis unit 420 based on the distance the hand has moved, the speed of the movement, or other similar metrics. In some examples, the ML algorithm of the target action analysis unit 420 can be trained to recognize and classify specific hand gestures such as waving a hand, pointing, or making a fist. In this case, the amount of hand movement can be measured by the target action analysis unit 420 based on the number or frequency of gestures performed. In some examples, the target action analysis unit 420 can analyze the pose and shape of the target's hand. The target action analysis unit 420 can use a computer vision algorithm to analyze the pose and shape of the hand, including the position of the fingers and the curvature of the palm. In this case, the amount of hand movement can be measured by the target analysis unit 420 based on the degree of change in the hand's pose or shape over time.

[0070] In some examples, the target action analysis unit 420 may track the movement of one or more arms of a target within the video signal data 402 using a computer vision algorithm. For example, the target action analysis unit 420 can identify an arm and separate it (e.g., isolate it) from the background (e.g., the rest of the environment). This can be done using techniques such as, for example, background subtraction, edge detection, and / or color segmentation. The target action analysis unit 420 can track the movement of the arm by tracking the position and orientation of the arm (e.g., after separation) within each frame of the video signal data 402, which can be provided as tracking data. For example, the target action analysis unit 420 may use techniques such as optical flow, feature tracking, and / or template matching. The target action analysis unit can determine the movement of the arm by calculating arm movement parameters such as speed, acceleration, and / or trajectory based on the tracking data. For example, the target action analysis unit 420 may use one or more mathematical models and equations, or an ML algorithm trained to calculate the movement of the arm based on the tracking data. The target action analysis unit 420 may provide target action data 422 that characterizes the amount of movement of the target's arm.

[0071] In some examples, the target action analysis unit 420 may determine the degree of freedom (DOF) of the target arm. The target action analysis unit 420 may provide target action data 422 that characterizes the DOF of the target arm. For example, the target action analysis unit 420 can use a computer vision algorithm to track the position and orientation of each joint of the arm and provide tracking data. In one example, small markers can be attached to the joints of the arm, and the position of the small markers can be tracked using a computer vision algorithm. Then, the position and orientation of each joint can be calculated by the target action analysis unit 420 to provide tracking data based on the relative position of the markers. In some examples, a computer vision algorithm can be used to track features of the arm, such as the contour of the skin or the edge of the bone. The target action analysis unit 420 can calculate / determine the position and orientation of each joint and provide tracking data based on the location of these features. In other examples, a depth sensor or a stereo camera (in some examples, such as cameras 203 and / or 204 as disclosed herein) can be used to capture the 3D position (3D coordinates) of each joint of each arm. The target action analysis unit 420 can determine the position and orientation of each joint based on the 3D coordinates and provide tracking data. Once the position and orientation of each joint of the arm are tracked, the DOF can be calculated by the target action analysis unit 420 by analyzing the relative movement of the joints. For example, the target action analysis unit 420 can use mathematical models and equations or apply an ML algorithm to the tracking data.

[0072] In some examples, the digital asset recommendation engine 410 can receive the target action data 422 and provide digital asset recommendation data 412. For example, the digital asset recommendation engine 410 can include an ML algorithm trained to identify possible or most relevant digital assets that can be used based on the target action data 422. The ML algorithm can be the ML model disclosed herein, or various types of ML algorithms trained to identify relevant digital assets based on various types of target actions (e.g., facial and body movements), hand movements, hand gestures, arm movements, and / or arm DOFs. For example, the engine 410 can identify a CGI asset having an arm that enables the same DOF as the target based on the target action data 422. Thus, the engine 410 can identify the CGI asset most likely to mimic the DOF (and thus the movement capabilities) of the target's arm. Thus, a CGI asset can be identified that provides a more realistic and lifelike movement reflected in the movement of the target's arm.

[0073] Accordingly, the system 400 can identify the most relevant digital assets for digital synthesis (e.g., insertion into an environment). The system 400 can reduce the likelihood that an inappropriate or incorrect CGI asset is selected in a particular scene or environment. Further, the system 400 identifies the CGI asset most likely to mimic (as faithfully as possible) the target action / behavior in that scene or environment, or in a different use / usage.

[0074] FIG. 5 is another example of a digital asset recommendation method 500. The method 500 may be implemented by the system 400 shown in FIG. 4. Accordingly, reference may be made to FIGS. 1-4 in the example of FIG. 5. The method 500 may begin at 502 by receiving video signal data (e.g., video signal data 402 shown in FIG. 4) including a plurality of video frames generated by one or more cameras in an environment. At 504, the video signal data may be analyzed (e.g., by the environment analysis unit 404 shown in FIG. 4) to detect or identify features from the environment, which may be provided as environment analysis data (e.g., environment analysis data 406 shown in FIG. 4). The features may include one or more objects in the environment. In some examples, the analyzing step at 504 is implemented according to object rules (e.g., object rules 408 shown in FIG. 4). At 506, the environment analysis data is used to evaluate the identified features in the video signal data 402 (e.g., by the digital asset recommendation engine 410 shown in FIG. 4) to predict the digital asset (e.g., CGI character) most relevant to that environment. At 508, in some examples, the digital asset may be provided for digital composition / rendering.

[0075] FIG. 6 is another example of a method 600 for digital asset recommendation. The method 600 may be implemented by the system 400 shown in FIG. 4. Thus, reference may be made to FIGS. 1-4 in the example of FIG. 6. The method 600 may begin at 602 by receiving video signal data (e.g., video signal data 402 shown in FIG. 4) including a plurality of video frames generated by one or more cameras for a subject. At 604, the video signal data may be analyzed (e.g., by the subject action analysis unit 420 shown in FIG. 4) to identify one or more actions of the subject (e.g., body and / or face movements), which may be provided as subject action data (e.g., subject action data 422 shown in FIG. 4). In some examples, the one or more actions include movements of one or more arms and / or hands. At 606, the subject action data may be evaluated (e.g., by the digital asset recommendation engine 410 shown in FIG. 4) to predict a digital asset (e.g., a CGI character) that is most likely to mimic the actions of the subject (e.g., correctly, or within a degree or percentage of the subject). At 608, in some examples, the digital asset may be provided for digital composition / rendering.

[0076] The present disclosure also targets the following exemplary embodiments.

[0077] Embodiment 1: A method comprising: a processor generating face data characterizing captured face parts and / or expressions; the processor generating body data characterizing captured body movements; the processor mapping the captured face parts and / or expressions and the captured body movements onto a skeleton model; and the processor animating a digital asset based on the mapping.

[0078] Embodiment 2: The mapping step includes a step of providing an animated skeleton model based on the mapping step, and the causing step includes a step of mapping the animated skeleton model to the digital asset, the method according to Embodiment 1.

[0079] Embodiment 3: The method according to any one of Embodiments 1 to 2, further comprising a step of causing the processor to provide extended video data including the digital asset animated therein based on the mapping step.

[0080] Embodiment 4: The method according to Embodiment 3, wherein the extended video data is rendered on a display.

[0081] Embodiment 5: The method according to any one of Embodiments 1 to 4, further comprising: a step of the processor receiving first video signal data including one or more images of an environment including the object; and a step of the processor receiving second video signal data including one or more images of the environment including the object.

[0082] Embodiment 6: The method according to Embodiment 5, further comprising: a step of the processor extracting parts and / or expressions of the face of the object based on the first video signal data, the face data being generated based on the extracted parts and / or expressions of the face of the object; and a step of the processor extracting movements of the body of the object based on the second video signal data, the body data being generated based on the extracted movements of the body of the object.

[0083] Embodiment 7: The method according to any one of Embodiments 5 to 6, wherein the first and second video signal data are provided by different video cameras.

[0084] Embodiment 8: The method according to any one of Embodiments 1 to 7, further comprising: a step in which the processor receives video signal data including one or more images of an environment including a target; a step in which the processor extracts the facial parts and / or expression of the target based on the video signal data, the face data being generated based on the extracted facial parts and / or expression of the target; and a step in which the processor extracts the body movement of the target based on the video signal data, the body data being generated based on the extracted body movement of the target.

[0085] Embodiment 9: The method according to any one of Embodiments 1 to 8, further comprising: a step in which the processor queries a digital asset database to obtain the digital asset, the digital asset being defined by a digital asset joint, wherein the causing step includes a step of associating the digital asset joint with a joint of the skeleton model and mapping an action of the skeleton model onto the digital asset.

[0086] Embodiment 10: A system including: a memory storing machine-readable instructions and data; one or more processors accessing the memory and executing the machine-readable instructions, the machine-readable instructions including: a body capture module providing face data characterizing facial parts and / or expression of a target; a face capture module providing body data characterizing body movement of the target; a character module obtaining a digital asset from a digital asset database; and a skeleton module animating the digital asset by mapping the body movement and / or facial expression of the target based on the body and / or face data, wherein extended video data is generated using the animated digital asset.

[0087] Embodiment 11: The system according to Embodiment 10, further comprising an image sensor providing video signal data including one or more images of the target, the face and body data being generated based on the signal data.

[0088] Embodiment 12: The image sensor has a first camera and a second camera, each of the first and second cameras provides one of first and second video signal data, and the video signal data includes the first and second video signal data, the system according to Embodiment 11.

[0089] Embodiment 13: The first and second cameras are configured for different imaging purposes, the system according to Embodiment 12.

[0090] Embodiment 14: The first camera is used to capture the parts and / or expressions of the face of the subject to provide the first video signal data, and the second camera is used to capture the movement of the body of the subject to provide the second video signal data, the system according to Embodiments 12 to 13.

[0091] Embodiment 15: The first and second cameras are mounted on a single portable device, the system according to Embodiments 12 to 14.

[0092] Embodiment 16: Receiving, in a processor, first video signal data characterizing parts and / or expressions of a face; receiving, in the processor, second video signal data characterizing movement of a body; the processor extracting the parts and / or expressions of the face to provide face data; the processor extracting the movement of the body to provide body data; the processor mapping the movement of the body and / or parts and / or expressions of the face to a skeleton model based on the face and / or body data to provide an animated skeleton model; and the processor animating a digital asset according to an action of the animated skeleton model to provide an animated digital asset.

[0093] Embodiment 17: The method according to embodiment 16, further comprising a step of causing the processor to provide, on a display, extended video data including the animated digital asset.

[0094] Embodiment 18: The method according to any one of embodiments 16 to 17, wherein the first and second video signal data are provided by cameras of respective portable devices.

[0095] Embodiment 19: The method according to any one of embodiments 16 to 18, wherein the first video signal data characterizes parts and / or expressions of the face of a first object, and the second video signal data characterizes movements of the body of a second object.

[0096] Embodiment 20: The method according to any one of embodiments 16 to 19, wherein the first and second video signal data are generated non-simultaneously.

[0097] One or more exemplary embodiments incorporating aspects of the presently disclosed invention are presented herein. For clarity, not all features of a physical implementation are described or shown in this application. It is to be understood that in developing a physical implementation incorporating embodiments of the present invention, numerous implementation-specific decisions must be made to achieve the developer's goals, such as compliance with system-related, business-related, government-related, and other constraints, which vary by implementation and over time. While the developer's work can be time-consuming, such work would be routine for those of ordinary skill in the art who would benefit from this disclosure.

[0098] In view of the foregoing structural and functional descriptions, those skilled in the art will appreciate that some embodiments may be embodied as a method, a data processing system, or a computer program product. Accordingly, some of these embodiments of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Further, some embodiments of the present disclosure may be a computer program product on a computer-usable storage medium having computer-readable program code thereon. Any non-transitory tangible storage media processing structures may be utilized, including but not limited to static and dynamic storage devices, hard disks, optical storage devices, and magnetic storage devices, provided that any media that are not eligible for patent protection under 35 U.S.C. § 101 (e.g., propagating electrical or electromagnetic signals per se) are excluded.

[0099] By way of example, and not limitation, a computer-readable storage medium may include, as necessary, semiconductor-based circuits or devices or other integrated circuits (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disks, HDDs, hybrid hard drives (HHDs), optical disks, optical disc drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tapes, holographic storage media, solid-state drives (SSDs), RAM drives, secure digital cards, secure digital drives, or other suitable computer-readable storage media, or combinations of two or more thereof. The computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, as necessary.

[0100] Some embodiments are also described herein with reference to block diagrams of methods, systems, and computer program products. It will be understood that the blocks of the figures and combinations of blocks within the figures can be implemented by computer-executable instructions. These computer-executable instructions can be provided to one or more processors of a general purpose computer, a special purpose computer, or other programmable data processing apparatus (or combination of devices and circuits) such that the instructions executed via the processor create a machine that implements the functions specified within the block or blocks.

[0101] Also, these computer-executable instructions may be stored in a computer-readable memory, and the instructions can direct a computer or other programmable data processing apparatus to function in a particular manner, such that a manufactured article is created that includes instructions for implementing the functions specified within the block or blocks of a flowchart. Further, the computer program instructions can be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus, thereby creating a computer-implemented process. Accordingly, the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified within the block or blocks of a flowchart.

[0102] As used herein, the term "software" is intended to encompass any set or collection of instructions executable by a computer or other digital system to configure the computer or other digital system to perform the tasks for which the software is intended. As used herein, the term "software" is intended to encompass such instructions stored in a storage medium such as RAM, hard disk, optical disk, or the like, and is also intended to encompass so-called "firmware", which is software stored in ROM or the like. Such software may be organized in various ways and may include libraries, Internet-based programs stored on a remote server or the like, source code, interpreted code, object code, directly executable code, and software components organized as such. The software may be capable of calling system-level code or may be intended to perform some functions by calling other software existing on a server or elsewhere.

[0103] In interpreting the appended claims, unless the phrase "means for" or "step for" is expressly used in a particular claim, the applicant does not intend to invoke 35 U.S.C. § 112(f) with respect to any of the appended claims or claim elements, to assist the Patent Office and any reader of this application and any resulting patent.

[0104] Certain terms are used in the following description for clarity, but these terms are intended to refer only to the particular structure of the embodiments selected for illustration in the drawings and are not intended to define or limit the scope of the present disclosure. In the figures below and the following description, it should be understood that like numerals refer to like-functioning components.

[0105] Although several exemplary embodiments have been described in this disclosure, it will be understood by those skilled in the art that various changes may be made without departing from the spirit and scope of the invention, and equivalents may be substituted for its elements. Further, many modifications will be understood by those skilled in the art to adapt a particular apparatus, situation, or material to the embodiments of this disclosure without departing from its essential scope. Accordingly, the invention is not limited to the specific embodiments disclosed or to the best mode contemplated for carrying out the invention, but rather the invention is intended to cover all embodiments included within the scope of the appended claims. Further, references in the appended claims to an apparatus or system, or a component of an apparatus or system, that is adapted, arranged, enabled, configured, made possible, operable, or operative to perform a particular function include that apparatus, system, or component whether or not that particular function is activated, turned on, or unlocked, so long as that apparatus, system, or component is so adapted, arranged, enabled, configured, made possible, operable, or operative.

[0106] As used herein, the term "or" is intended to be inclusive rather than exclusive. Unless otherwise specified, "X uses A or B" is intended to mean any of a natural incisive permutation. That is, "X uses A or B" is satisfied if X uses A; X uses B; or X uses both A and B. Further, the articles "a" or "an" should generally be construed to mean "one or more" of each noun, unless otherwise specified. As used herein, the terms "example" and / or "exemplary" are used to indicate one or more features as an example, instance, or illustration. The subject matter described herein is not limited by such examples. Further, any aspect, feature, and / or design described herein as "example" or "exemplary" is not necessarily intended to be construed as preferred or advantageous. Similarly, any aspect, feature, and / or design described herein as "example" or "exemplary" is not meant to exclude equivalent embodiments (e.g., features, structures, and / or methodologies) known to those skilled in the art.

[0107] Understanding that it is not possible to describe every possible combination of the various features (e.g., components, products, and / or methods) described herein, one of ordinary skill in the art can recognize that many additional combinations and permutations of the various embodiments described herein are possible and contemplated. Further, as used herein, the terms "includes", "has", "possesses", and / or the like are intended to be inclusive in the same manner as the term "comprising" as construed when used as a transitional phrase in a claim.

[0108] What has been described above are examples. Of course, it is not possible to describe every conceivable combination of components or methods, but those skilled in the art will recognize that many additional combinations and permutations are possible. Accordingly, the present disclosure is intended to embrace all such modifications, corrections, and variations that fall within the scope of this application, including the appended claims. When the present disclosure or the claims enumerate "a", "an", "first", or "another" element or its equivalent, it should be construed to include one or more than one such element, and does not exclude or require two or more such elements. As used herein, the term "includes" means including but not limited to, and the term "including" means including but not limited to. The term "based on" means "based at least in part on".

[0109] As used herein, the terms "generally" and "substantially" are intended to encompass structural or numerical modifications that do not significantly affect the object of the element or number modified by such terms.

[0110] As used herein, the term "raw video signal" means a video signal directly obtained from an image sensor of a camera that captures, for example, the environment and actors within the field of view. The term "animated video signal" means a video signal that includes a combination of a raw video signal and CGI characters mapped to the captured movements of actors within the field of view.

Claims

1. The processor generates face data characterizing the captured face parts and / or expressions; The processor generates body data characterizing the captured body movements; The processor maps the captured face parts and / or expressions and the captured body movements onto a skeleton model; and The processor animates a digital asset based on the mapping. A method comprising the above steps.

2. The mapping step includes providing an animated skeleton model based on the mapping step, and the animating step includes mapping the animated skeleton model onto the digital asset. The method according to claim 1.

3. The processor further includes a step of providing extended video data including the internally animated digital asset based on the mapping step. The method according to claim 1.

4. The extended video data is rendered on a display. The method according to claim 3.

5. The processor receives first video signal data including one or more images of an environment including a subject; and The processor receives second video signal data including one or more images of the environment including the subject. The method according to claim 1, further comprising the above steps.

6. The processor extracts the face parts and / or expressions of the subject based on the first video signal data, and the face data is generated based on the extracted face parts and / or expressions of the subject; and The processor extracts the body movements of the subject based on the second video signal data, and the body data is generated based on the extracted body movements of the subject. The method according to claim 5, further comprising the above steps.

7. The first video signal data and the second video signal data are provided by different video cameras. The method according to claim 6.

8. The processor receives video signal data including one or more images of an environment including a subject. The processor extracts the facial parts and / or expressions of the subject based on the video signal data, the facial data being generated based on the extracted facial parts and / or expressions of the subject, and the processor extracts the body movements of the subject based on the video signal data, the body data being generated based on the extracted body movements of the subject, The method according to claim 1, further comprising.

9. The processor further comprises querying a digital asset database to obtain a digital asset, the digital asset being defined by a digital asset joint, The causing step includes associating the digital asset joint with the joint of the skeleton model and mapping the action of the skeleton model onto the digital asset. The method according to claim 1.

10. A memory storing machine-readable instructions and data, One or more processors accessing the memory and executing the machine-readable instructions Comprising, the machine-readable instructions being A body capture module providing facial data characterizing the facial parts and / or expressions of a subject, A face capture module providing body data characterizing the body movements of the subject, A character module for obtaining a digital asset from a digital asset database, and A skeleton module that maps the body movements and / or facial expressions of the subject based on the body and / or facial data onto the digital asset to animate the digital asset, wherein extended video data is generated using the animated digital asset. Including, System.

11. The system according to claim 10, further comprising an image sensor providing video signal data including one or more images of the subject, the facial data and the body data being generated based on the video signal data.

12. The image sensor has a first camera and a second camera, each of the first camera and the second camera providing one of first video signal data and second video signal data, the video signal data including the first video signal data and the second video signal data. The system according to claim 11.

13. The system according to claim 12, wherein the first camera and the second camera are configured for different imaging purposes.

14. The system according to claim 13, wherein the first camera is used to capture the parts and / or expression of the face of the subject and provide the first video signal data, and the second camera is used to capture the movement of the body of the subject and provide the second video signal data.

15. The system according to claim 14, wherein the first camera and the second camera are mounted on a single portable device.

16. Receiving, in a processor, first video signal data characterizing parts and / or expression of a face; Receiving, in the processor, second video signal data characterizing movement of a body; The processor extracting the parts and / or expression of the face and providing face data; The processor extracting the movement of the body and providing body data; The processor mapping the movement of the body and / or parts and / or expression of the face to a skeleton model based on the face and / or body data to provide an animated skeleton model; and The processor animating a digital asset according to an action of the animated skeleton model to provide an animated digital asset A method comprising.

17. The method according to claim 16, further comprising the processor causing extended video data including the animated digital asset to be provided on a display.

18. The method according to claim 16, wherein the first video signal data and the second video signal data are provided by cameras of respective portable devices.

19. The method according to claim 16, wherein the first video signal data characterizes parts and / or expression of a face of a first subject, and the second video signal data characterizes movement of a body of a second subject.

20. The method according to claim 19, wherein the first video signal data and the second video signal data are generated non-simultaneously.