Real-time rendering of video streams
By identifying and tracking facial features in the video stream, and generating and rendering the graphic representations input by users in real time, the problem of image modification in video communication in the prior art is solved, the effect of modifying the video stream in real time is realized, and the user interaction experience is improved.
Patent Information
- Application Number
- CN202111492976.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-07-18
- Filing Date
- 2017-07-18
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2037-07-18
AI Technical Summary
It is difficult for existing video communication devices to modify images in the video stream in real time during a communication session, especially by physical manipulation of the device such as changing orientation or manipulating input devices.
Modify the appearance of the video stream by identifying and tracking objects of interest, especially facial features, within the video stream, and generating and rendering user-input graphic representations, including colors, shapes, and lines.
The image of the video stream is modified in real time during video communication, enhanced user interaction experience, and able to track and maintain graphic representations in the video stream, such as drawing team marks or numbers on the athlete's cheeks.
Smart Images

Figure CN114143570B_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with the filing date of July 18, 2017, application number 201780056702.3, and invention title "Real-time Rendering of Video Streams".
[0002] Priority Claims
[0003] This application claims the priority of U.S. Patent Application Serial No. 15 / 213,186, filed on July 18, 2016, the benefit of the priority of each such application is hereby claimed, and each such application is incorporated herein by reference in its entirety. Technical Field
[0004] Embodiments of the present disclosure generally relate to the automatic processing of images. More specifically, but without limitation, the present disclosure presents systems and methods for generating a persistent graphical representation of user input within a video stream in real time. Background Art
[0005] Telecommunication applications and devices can use various media (such as text, images, sound recordings, and / or video recordings) to provide communication between multiple users. For example, video conferencing allows two or more individuals to communicate with each other using a combination of software applications, telecommunication devices, and telecommunication networks. Telecommunication devices can also record video streams for transmission as messages in a telecommunication network.
[0006] Although there are telecommunication applications and devices for providing two-way video communication between two devices, there may be problems with video streams, such as modifying an image within the video stream during an ongoing communication session. Telecommunication devices perform operations using physical manipulation of the device. For example, a device is typically operated by changing the orientation of the device or manipulating an input device (such as a touch screen). Thus, there is still a need in the art to improve video communication between devices and to modify a video stream in real time while the video stream is being captured. Brief Description of the Drawings
[0007] The various drawings in the figures merely illustrate example embodiments of the present disclosure and should not be considered as limiting its scope.
[0008] Figure 1 is a block diagram showing a networking system according to some example embodiments.
[0009] Figure 2 is a diagram showing a rendering system according to some example embodiments.
[0010] Figure 3 is a flowchart showing an example method for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device.
[0011] Figure 4 is a flowchart showing an example method for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device.
[0012] Figure 5 is a flowchart showing an example method for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device.
[0013] Figure 6 is a flowchart showing an example method for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device.
[0014] Figure 7 is a flowchart showing an example method for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device.
[0015] Figure 8 is a flowchart showing an example method for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device.
[0016] Figure 9 is a user interface diagram depicting an example mobile device and a mobile operating system interface according to some example embodiments.
[0017] Figure 10 is a block diagram showing an example of a software architecture that can be installed on a machine according to some example embodiments.
[0018] Figure 11 is a block diagram presenting a graphical representation of a machine in the form of a computer system within which a set of instructions can be executed to cause the machine to perform any method discussed herein.
[0019] The headings provided herein are for convenience only and do not necessarily affect the scope or meaning of the terms used. Detailed Description
[0020] The following description includes systems, methods, techniques, instruction sequences, and computer program products that illustrate embodiments of the present disclosure. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide an understanding of the various embodiments of the subject matter of the present invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of the present invention may be practiced without these specific details. Generally, well-known instruction examples, protocols, structures, and techniques need not be shown in detail.
[0021] A drawing system is described that identifies and tracks objects of interest across a video stream through a set of frames within the video stream. In various example embodiments, the drawing system identifies and tracks facial landmarks, user input, and the relative aspect ratio between landmarks or points on a face and user input. The drawing system enables a user to select colors, shapes, lines, and other graphical representations of user input to draw in real time within the frames of the video stream as the video stream is being captured by an image capture device. The drawing may be on a face or other object of interest. The drawing system is capable of modifying the video stream to incorporate the drawing to highlight or otherwise alter the appearance of the video stream. For example, the drawing system may reduce the color, shading, saturation, or brightness of the video stream and present the unmodified color in the graphical representation of the drawing, such that the drawn portion appears neon, flashing, or any other suitable color contrast. For example, a user of a device at a sporting event may use an interface to draw a team logo or player number on the cheek of a player who appears in a video stream generated by the device. The drawing system will then track the player and maintain the drawing on the player's cheek as the player moves within the video stream and as the video stream is transmitted to another device or other devices. In some cases, a user of a device at a sporting event may draw a team logo or player number on a portion of the screen near the player's face (such as in the air directly above the player's head). The drawing system will then track the player and maintain the drawing in the air above the player's head.
[0022] The above is a specific example. Various embodiments of the present disclosure relate to instructions for a device and one or more processors of the device to modify an image or video stream transmitted by the device to another device (e.g., modify the video stream in real time) when capturing the video stream. A drawing system is described that identifies and tracks objects of interest and regions within an image or video stream and through a set of images including the video stream. In various example embodiments, the drawing system identifies and tracks one or more facial features depicted within the video stream or image and performs image recognition, facial recognition, and facial processing functions with respect to the one or more facial features and the relationships between two or more facial features.
[0023] Figure 1FIG. 0 is a network diagram depicting a network system 100 according to one embodiment. The network system 100 has a client-server architecture configured to exchange data over a network. For example, the network system 100 can be a messaging system where clients transmit and exchange data within the network system 100. The data can relate to various functions (e.g., sending and receiving text and media communications, determining a geographical location, etc.) and aspects associated with the network system 100 and its users (e.g., transmitting communication data, receiving and sending indications of communication sessions, etc.). Although shown here as a client-server architecture, other embodiments can include other network architectures, such as peer-to-peer or distributed network environments.
[0024] As Figure 1 shown, the network system 100 includes a social messaging system 130. The social messaging system 130 generally is based on a three-tier architecture, including an interface layer 124, an application logic layer 126, and a data layer 128. As will be understood by those of ordinary skill in the relevant computer and Internet-related arts, Figure 1 each component or engine shown represents a set of executable software instructions and corresponding hardware for executing the instructions (e.g., memory and a processor), forming a hardware-implemented component or engine, and when the instructions are executed, serves as a dedicated machine configured to perform a specific set of functions. To avoid obscuring the subject matter of the present invention with unnecessary details, various functional components and engines not closely related to conveying an understanding of the subject matter of the present invention are omitted from Figure 1 FIG. 0. Of course, additional functional components and engines can be used with the social messaging system (such as the social messaging system shown in Figure 1 FIG. 0) to facilitate additional functions not specifically described herein. Additionally, Figure 1 the various functional components and engines depicted in FIG. 0 can reside on a single server computer or client device, or can be distributed across several server computers or client devices in various arrangements. Further, although Figure 1 the social messaging system 130 is depicted in FIG. 0 as a three-tier architecture, the subject matter of the present invention is in no way limited to such an architecture.
[0025] As Figure 1 shown, the interface layer 124 includes an interface component (e.g., a web server) 140 that receives requests from various client computing devices and servers (such as client device 110 executing a client application 112 and third-party server 120 executing a third-party application 122). In response to the received requests, the interface component 140 transmits an appropriate response to the requesting device via network 104. For example, the interface component 140 can receive requests such as Hypertext Transfer Protocol (HTTP) requests or other web-based application programming interface (API) requests.
[0026] The client device 110 may execute a traditional web browser application or an application (also referred to as an "app") developed for a specific platform to include various mobile computing devices and mobile-specific operating systems (e.g., IOS TM , ANDROID TM , PHONE) among any of them. Additionally, in some example embodiments, the client device 110 forms all or a part of the rendering system 160 such that components of the rendering system 160 configure the client device 110 to perform a set of specific functions regarding operations of the rendering system 160.
[0027] In an example, the client device 110 executes a client application 112. The client application 112 may provide functions of presenting information to the user 106 and communicating via the network 104 to exchange information with the social messaging system 130. Additionally, in some examples, the client device 110 executes functions of the rendering system 160 to segment images of a video stream during collection of the video stream and transmit the video stream (e.g., with image data modified based on the segmented images of the video stream).
[0028] Each of the client devices 110 may include a computing device that includes at least a display and communication capabilities with the network 104 to access the social messaging system 130, other client devices, and the third-party server 120. The client devices 110 include but are not limited to remote devices, workstations, computers, general-purpose computers, Internet appliances, handheld devices, wireless devices, portable devices, wearable computers, cellular or mobile phones, personal digital assistants (PDAs), smartphones, tablet computers, ultrabooks, netbooks, laptop computers, desktop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, gaming consoles, set-top boxes, network PCs, minicomputers, etc. The user 106 may be a person, a machine, or other components that interact with the client device 110. In some embodiments, the user 106 interacts with the social messaging system 130 via the client device 110. The user 106 may not be part of the networked environment but may be associated with the client device 110.
[0029] As Figure 1 shown, the data layer 128 has a database server 132 facilitating access to an information repository or database 134. The database 134 is a storage device storing data such as member profile data, social graph data (e.g., relationships among members of the social messaging system 130), image modification preference data, accessibility data, and other user data.
[0030] An individual can register with the social messaging system 130 to become a member of the social messaging system 130. After registration, the member can form social network relationships (e.g., friends, followers, or contacts) on the social messaging system 130 and interact with a wide range of applications provided by the social messaging system 130.
[0031] The application logic layer 126 includes various application logic components 150 that, in conjunction with the interface components 140, generate various user interfaces using data obtained from various data sources or data services in the data layer 128. Each application logic component 150 can be used to implement functions associated with various applications, services, and features of the social messaging system 130. For example, a social messaging application can be implemented with the application logic components 150. The social messaging application provides a messaging mechanism for users of the client device 110 to send and receive messages including text and media content such as pictures and videos. The client device 110 can access and view messages from the social messaging application for a specified period of time (e.g., limited or unlimited). In an example, a message recipient can access a particular message for a predetermined duration (e.g., specified by the message sender), which starts when the particular message is first accessed. After the predetermined duration has passed, the message will be deleted and the message recipient will no longer be able to access the message. Of course, other applications and services can be embodied in their respective application logic components 150.
[0032] As Figure 1 shown, the social messaging system 130 can include at least a portion of the rendering system 160 that is capable of identifying, tracking, and modifying video data during video data acquisition by the client device 110. Similarly, as described above, the client device 110 includes a portion of the rendering system 160. In other examples, the client device 110 can include the entire rendering system 160. In cases where the client device 110 includes a portion (or all) of the rendering system 160, the client device 110 can work alone or in cooperation with the social messaging system 130 to provide the functions of the rendering system 160 described herein.
[0033] In some embodiments, the social messaging system 130 can be an ephemeral messaging system that allows ephemeral communication, where content (e.g., video clips or images) is deleted after a deletion trigger event such as a viewing time or viewing completion. In such embodiments, the device uses the various components described herein in any case of generating, sending, receiving, or displaying ephemeral messages. For example, a device implementing the drawing system 160 can identify, track, and modify regions of interest, such as pixels representing skin on a face depicted in a video clip. The device can modify the regions of interest during video clip acquisition as part of the generation of the content of an ephemeral message, without performing image processing after video clip acquisition.
[0034] In Figure 2 various embodiments, the drawing system 160 can be implemented as a stand-alone system or in combination with the client device 110 and does not have to be included in the social messaging system 130. The drawing system 160 is shown to include an acquisition component 210, an input component 220, a location component 230, a linking component 240, a generation component 250, a rendering component 260, a scaling component 270, and a tracking component 280. All or some of the components 210-280 communicate with each other, for example, via a network, shared memory, etc. Each of the components 210-280 can be implemented as a single component, combined into other components, or further subdivided into multiple components. Other components that are not relevant to the example embodiments may also be included but are not shown.
[0035] The acquisition component 210 accesses or otherwise obtains frames acquired by an image acquisition device or otherwise received or stored in the client device 110 by the client device 110. In some cases, the acquisition component 210 can include part or all of an image acquisition component configured to cause the image acquisition device of the client device 110 to acquire frames of a video stream based on user interaction with a user interface presented on the display device of the client device 110. The acquisition component 210 can pass the frame or a portion of the frame to one or more other components of the drawing system 160.
[0036] The input component 220 identifies user input on an input device of the client device 110. The input component 220 can include part or all of an input device (such as a keyboard, mouse, cursor, touch screen, or any other suitable input device of the client device 110). In the case where the input device is a touch screen, in some example embodiments, the input component 220 identifies and differentiates the touch pressure applied to the touch screen to identify or differentiate relevant features selected based on the touch pressure.
[0037] The location component 230 identifies one or more locations within a frame of the video stream. The location component 230 can identify locations on portions of a face depicted within a frame of the video stream. In some embodiments, the location component 230 can identify or define coordinates, facial landmarks, or points on portions of the face. In some cases, the location component 230 can determine or generate a facial mesh or mesh elements on portions of the face to identify facial landmarks and to determine distances between two or more facial landmarks.
[0038] The linking component 240 links the user input received by the input component 220 and the one or more locations identified by the location component 230. The linking component 240 can directly link locations, points, or landmarks on portions of the face with one or more points of the user input received by the input component 220. In some cases, the linking component 240 links the locations, points, or landmarks of portions of the face with the user input by linking one or more pixels associated with the user input and one or more pixels associated with portions of the face.
[0039] The generating component 250 generates a graphical representation of the user input. The graphical representation can be generated as lines, portions of lines, shapes, colors, patterns, or any other representation. In some cases, the generating component 250 can generate at least in part from a user interface selection of a location and a graphical type for the user input. In some cases, the generating component generates or assigns edge or dimension features to the graphical representation. The edge and dimension features modify the depiction of the graphical representation and can be selected by user input of an input device of the client device 110. In some cases, one or more of the edge and dimension features can be determined by the generating component 250 or other components of the drawing system 160 based on features of portions of the face or relative distances of portions of the face.
[0040] The rendering component 260 renders the graphical representation within a frame of the video stream. In some cases, the rendering component 260 renders the graphical representation on portions of the face or within and outside of the frame and portions of the face. The rendering component 260 can modify the frame of the video stream to include one or more graphical representations. In some cases, the rendering component 260 stores the graphical representation within a processor-readable storage device. After storage, the rendering component 260 or the fetching component 210 can access the graphical representation and render the previously stored graphical representation in real time within subsequent video streams.
[0041] The scaling component 270 scales the graphical representation based on the movement of the object depicted within the frame of the video stream. In some example embodiments, the scaling component 270 scales the dimensional characteristics of the graphical representation based on the movement of a portion of the face relative to the client device 110 or the image capture device of the client device 110. In some embodiments, the scaling component 270 may maintain the line width or aspect ratio between the graphical representation and one or more facial landmarks, positions, or points on the portion of the face or one or more coordinates outside the set of coordinates associated with the portion of the face but within the frame while scaling the graphical representation as the portion of the face moves within the frame and relative to the image capture device.
[0042] The tracking component 280 tracks the movement of an object within the frame of the video stream. In some cases, the tracking component 280 tracks the movement of a portion of the face, one or more facial landmarks of the portion of the face, or a designated point on the object of interest. The tracking component 280 may match the movement of one or more facial landmarks, facial portions, or other objects with the movement of the graphical representation overlaid on the graphical plane.
[0043] Figure 3 A flowchart depicting an example method 300 for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device (e.g., a computing device) is shown. The operations of method 300 may be performed by components of the drawing system 160 and are described below for illustrative purposes.
[0044] In operation 310, the acquisition component 210 receives or otherwise accesses one or more frames of the video stream. At least a portion of the one or more frames depicts at least a portion of a face. In some embodiments, the acquisition component 210 receives one or more frames of the video stream as captured by an image capture device associated with the client device 110 and presented on the user interface of the face drawing application. The acquisition component 210 may include the image capture device as part of the hardware that includes the acquisition component 210. In these embodiments, the acquisition component 210 directly receives one or more frames or the video stream captured by the image capture device. In some cases, the acquisition component 210 passes all or a portion of one or more images or video streams (e.g., a set of images including the video stream) to one or more components of the drawing system 160, as described in more detail below.
[0045] In operation 320, input component 220 identifies user input on an input device of client device 110 (e.g., a computing device or a mobile computing device). Input component 220 can identify user input based on receiving input or an indication of input from an input device of client device 110. For example, input component 220 can identify user input received from a mouse, cursor, touchscreen device, button, keypad, or any other suitable input device that is part of, connected to, or communicates with client device 110.
[0046] In some embodiments, user input is received via a touchscreen device. The touchscreen device can be configured to be pressure sensitive to touches that provide user input. In these embodiments, input component 220 determines the touch pressure of the user input. Based on the touch pressure of the user input, input component 220 identifies an edge feature associated with the touch pressure. The touch pressure can be binary. In cases where the touch pressure is binary, a pressure above a specified threshold can be associated with a first aspect of the user input, while a pressure below the specified threshold can be associated with a second aspect of the user input. In these cases, input component 220 determines that the touch pressure exceeds the pressure threshold and identifies the first aspect as being associated with the user input. In some embodiments, the first aspect and the second aspect can be first and second edge features. For example, the first edge feature can be a diffused edge of an edge of a blended or blurred line, and the second edge feature can be a sharp edge with a well-defined boundary. In some embodiments, the touch pressure can be non-binary. In these cases, the edge feature can be associated with two or more pressures or specified thresholds between two terminal values associated with the pressure of the user input. For example, the first terminal value can represent the lightest touch pressure perceivable by the touchscreen, while the second terminal value can represent the heaviest pressure perceivable by the touchscreen. In these cases, two or more thresholds can be located at values between the first terminal value and the second terminal value. Each of the two or more thresholds can be associated with an edge feature or other user input or graphical representation feature such that a pressure exceeding a threshold of the two or more thresholds can indicate a single feature or edge feature. The feature can be an edge feature (e.g., sharpness of an edge), line width or thickness, line color or color intensity, or any other suitable feature.
[0047] Once the acquisition component 210 begins receiving frames of the video stream, user input can include a selection of a color or a set of colors. The drawing system 160 can cause a set of user input elements presented on the video stream to be presented to enable selection of colors, graphic effects (e.g., flash, neon glow, variable or timed glow, event-based appearance), line widths, designs, and other options that affect the style, color, and timing of the drawn graphical representation. For example, the user input can include a selection of a timing mechanism or a location mechanism for presenting the drawn graphical representation. In these cases, the drawing system 160 can cause the drawn graphical representation to be presented in accordance with an event (such as a face rotation or an eye blink) or based on a timing element (such as a time interval or a set time period). In the case where the user input includes a selection of a style, subsequent user input can be formatted according to one or more sets of style characteristics of the selected style. In some embodiments, selectable styles include drawing styles (e.g., cubism, impressionism, Shijo school, landscape), palette styles (e.g., neon, pastel), or other styles that can be created by a set of predefined style characteristics. For example, in the case of selecting an impressionist style, the drawn graphical representation of subsequent user input can be configured within a soft palette with diffused edges and simulated brush strokes. As another example, in the case of selecting a neon palette or style, the drawn graphical representation generated from subsequent user input can be generated to have a glowing effect. The glowing effect can be generated by the brightness values of the colors within the drawn graphical representation and by adjusting the brightness values of at least a portion of the frames of the video stream.
[0048] In operation 330, the location component 230 identifies one or more locations on a portion of the face corresponding to the user input. The location component 230 can identify the one or more locations based on the pixel locations of one or more locations on the portion of the face, facial landmarks on the portion of the face, one or more coordinates mapped to the portion of the face, or any other suitable location on the portion of the face. For example, the face can depict a set of known facial landmarks. The location component 230 can identify the one or more locations relative to the facial landmarks.
[0049] In some example embodiments of performing operation 330, the position component 230 determines a facial mesh on a portion of the face. The facial mesh identifies one or more facial landmarks depicted within the portion of the face. Determining the facial mesh may include projecting the mesh onto a regular grid to divide the mesh into 100×100 cells by the regular grid. Although described as a mesh, it should be understood that the mesh may be formed by intersecting or touching polygons. For example, the mesh may be formed by a set of connected triangles, where the points of one or more of the triangles represent facial landmarks, landmark points, or points determined based on the relative distances of the facial landmarks. Once the mesh projection is formed, the position component 230 may determine, for each cell, the mesh elements corresponding to the grid cell. Then, the position component 230 may determine the pixels corresponding to each determined mesh element. In some cases, a breadth-first search may be used to perform the determination of the pixels corresponding to each mesh element.
[0050] The position component 230 may identify one or more coordinates input by the user. The one or more coordinates may be aligned with or identified relative to one or more cells, points, or intersections of the mesh. In some cases, the one or more coordinates may be identified as one or more pixels depicted within the mesh. In some embodiments, identifying one or more coordinates input by the user may include determining one or more pixels selected within the user input (e.g., one or more pixels drawn by the user input intersecting a line). Once the one or more pixels are determined, the position component 230 may determine the portion, point, or coordinates within the mesh corresponding to the one or more determined pixels.
[0051] The position component 230 may map the one or more coordinates to at least a portion of one or more facial landmarks. In some cases, the relative distances of the one or more coordinates are determined relative to the facial landmarks, mesh elements, or points on the mesh. In some cases, the position component 230 maps the one or more coordinates to one or more pixels within the frame that are mapped or otherwise associated with the mesh. The mapping may establish a reference for tracking the user input and a graphical representation of the user input across multiple frames of the video stream. In these cases, tracking the user input or the graphical representation may enable the generation, tracking, and presentation of a persistent graphical representation within the video stream in real time, as described further below and in the embodiments.
[0052] In operation 340, the linking component 240 links the user input to one or more locations on a portion of the face. In some embodiments, linking the user input to one or more locations on a portion of the face generates a pixel-independent association between the user input or points within the user input and one or more locations on the face. The pixel-independent association enables the user input to move with the portion of the face across multiple frames of the video stream. The link can be generated as a database, table, or data structure that associates points of the user input and one or more locations on the face. The database, table, or data structure can be referenced when it is determined that the location occupied by the portion of the face within a frame is different from the location of the portion of the face occupied in the previous frame of the video stream.
[0053] In operation 350, the generating component 250 generates a graphical representation of the user input. As in operation 340, the graphical representation of the user input is linked to one or more locations on a portion of the face. The graphical representation can include lines, portions of lines, shapes, colors, patterns, or any other representation of the user input. For example, the graphical representation can be a line that occupies the area generated by the user input within the frame. In some embodiments, the graphical representation can be a drawn line, shape, or design to be presented in real time within a frame of the video stream. The graphical representation can be generated to appear as a painting, makeup, tattoo, or other representation located on or within the skin or hair depicted on the portion of the face. In some cases, the graphical representation can be generated as a physical object or structure, such as antlers, a mask, hair, a hat, glasses, or other object or structure positioned in, on, or near the portion of the face.
[0054] In embodiments where the user input includes touch pressure, one or more of the input component 220 and the generating component 250 assign edge features identified as being associated with the touch pressure to the graphical representation generated from the user input. For example, in the case where the edge feature is a diffuse edge, the graphical representation generated as a drawn line can be positioned on the face and include an edge that spreads out evenly or unevenly outward from a central point within the line. As another example, the diffuse edge can blend the edge of the graphical representation into one or more additional colors of the portion of the face.
[0055] In operation 360, the rendering component 260 renders the graphical representation on a portion of the face within one or more subsequent frames of the video stream. The graphical representation can be presented on a portion of the face at one or more locations. The rendering component 260 causes the graphical representation to be presented within a frame of the video stream. Once rendered by the rendering component 260, the graphical representation can be continuously rendered on frames of the video stream where one or more locations of the portion of the face corresponding to the graphical representation appear within the frame. In frames where one or more locations of the portion of the face do not exist within one or more frames, the rendering component 260 can skip or otherwise not render the graphical representation within one or more frames.
[0056] In some embodiments, once rendered, the rendering component 260 may store the graphical representation in a processor-readable storage device. Once stored, the graphical representation may be invoked based on one or more events. In some embodiments, both may correspond to opening the face painting application after storing the graphical representation or initializing an instance of the face painting application after storage. In some cases, in response to launching the face painting application, the acquisition component 210 detects portions of a face in a new video stream. The new video stream may be a video stream different from the video stream used to generate and render the graphical representation. In response to detecting portions of a face in the new video stream, the rendering component 260 renders the graphical representation stored in the processor-readable storage device within one or more frames of the new video stream.
[0057] Figure 4 A flowchart depicting an example method 400 for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device (e.g., a computing device) is shown. Operations of method 400 may be performed by components of the painting system 160. In some cases, certain operations of method 400 may be performed using one or more operations of method 300 or as sub-operations of one or more operations of method 300, as will be explained in more detail below. For example, as Figure 4 shown, method 400 may be performed as part of operation 320.
[0058] In operation 410, the position component 230 determines a first relative distance of a portion of the face from an image capture device of the computing device. The first relative distance may be determined in a frame of the video stream. The position component 230 may determine the relative distance of the portion of the face based on the amount of the frame occupied by the portion of the face. In cases where the first relative distance is determined based on the occupied portion of the frame, the position component 230 may determine the distance between face landmarks and temporarily store the distance for comparison to determine changes in the relative position of the portion of the face with respect to the image capture device.
[0059] In operation 420, the input component 220 determines the dimensional characteristics of the graphical representation based on a first relative distance of a portion of the face. In some example embodiments, the input component 220 determines the dimensional characteristics based on the first relative distance and a constant dimension input by the user. In these embodiments, the user input is received from a touchscreen device based on contact from a finger or a stylus. The finger or stylus may have a predetermined width or a width within a predetermined width range. The input component 220 can know the width or the predetermined width range such that the input received from the touchscreen has a fixed input width. The input component 220 can receive user input having a fixed input width. The input component 220, in cooperation with the position component 230 and the linking component 240, can link one or more points on the fixed input width of the user input to one or more points, one or more positions, or one or more facial landmarks depicted on a portion of the face.
[0060] In some cases, the input component 220 can determine the dimensional characteristics at least in part based on a selection of a line width option in the user interface. For example, prior to receiving user input corresponding to the graphical representation, the rendering component 260 can present a set of selectable line width elements. The input component 220 can receive user input from the set of selectable line width elements, the user input including a selection of a line width, thickness, shape, or any other suitable user interface option for the dimensional characteristics. The input component 220 combines the selection of the line width element to modify a default line width associated with the user input for the graphical representation.
[0061] In operation 430, the scaling component 270 scales the dimensional characteristics of the graphical representation based on the movement of a portion of the face from a first relative distance to a second relative distance in subsequent frames of the video stream. For example, as depicted within one or more frames of the video stream, as the face moves closer to the image capture device, the line width of the graphical representation can increase such that the lines of the graphical representation become thicker as the face approaches the image capture device. Similarly, as a portion of the face moves away from the image capture device, the line width of the graphical representation decreases such that the lines become thinner. In some embodiments, the scaling component 270 maintains the line width ratio of the graphical representation relative to a point, position, or facial landmark of a portion of the face. For example, although the line width becomes thicker as the portion of the face moves closer to the image capture device during scaling, the aspect ratio of the line width remains constant. In these cases, although the pixel values of the graphical representation change as the dimensional characteristics scale, the ratio is maintained such that the graphical representation occupies the same apparent space on the portion of the face regardless of the relative distance of the portion of the face from the image capture device.
[0062] In some embodiments, the scaling component 270 may include or determine a minimum face ratio for scaling the graphical representation. The scaling component 270 may use a minimum face ratio preset within the rendering system 160 for the size of the facial portion relative to the line width or other dimensional features of the graphical representation. In some cases, where the scaling component 270 includes a preset minimum face ratio, the scaling component 270 may generate and cause a rendering of an instruction (e.g., via a user interface element) to be shown on the video stream to indicate to the user to move a portion of the face or the image capture device until the portion of the face is positioned within the video frame at a size corresponding to the minimum face ratio. In some embodiments, the scaling component 270 determines the minimum face ratio by assuming that the initially presented portion of the face within the frame of the video stream is the minimum size of the face to be presented within the video stream.
[0063] In operation 432, when scaling the dimensional feature, the scaling component 270 identifies a first line width for the graphical representation at a first relative distance. The scaling component 270 may identify the first line width based on one or more of user input selected for the line width, user input corresponding to the graphical representation, and the first relative distance. In some embodiments, as described above, the scaling component 270 may identify the first line width as the selected line width, a fixed input width, or other specified value. In some cases, the scaling component 270 identifies the first line width as determined using the fixed input width and the first relative distance of the portion of the face.
[0064] In operation 434, the scaling component 270 determines a first relative position of at least one point on the graphical representation and two or more facial landmarks depicted on the portion of the face. The scaling component 270 may determine the first relative position of the at least one point at one or more pixels within one or more frames of the video stream. In some cases, the scaling component 270 determines the first relative position as one or more coordinates within one or more frames of the video stream, where the video stream is segmented into a set of coordinates that partition the frames for object tracking.
[0065] In operation 436, the scaling component 270 determines a change in the distance between two or more facial landmarks at a second relative distance. In some embodiments, the scaling component 270 determines the change in distance by identifying a second relative position of two or more facial landmarks depicted on the portion of the face. The scaling component 270 determines the change in distance by comparing the first relative position of the two or more facial landmarks and the second relative position of the two or more facial landmarks at the second relative distance. The scaling component 270 may compare the pixel or coordinate positions of the first relative position and the second relative position to determine whether the pixel or coordinate positions match. In the case where the pixel or coordinate positions do not match, the scaling component 270 may identify the change in distance.
[0066] In operation 438, the scaling component 270 modifies a first line width of the graphical representation based on a change in the distance between two or more facial landmarks to generate a second line width of the graphical representation. In some embodiments, in response to identifying a change in the distance between the first relative position and the second relative position, the scaling component 270 modifies the width of the graphical representation by changing the first line width to the second line width. The second line width maintains the aspect ratio that exists between the first line width and two or more facial features at the first relative distance.
[0067] Figure 5 A flowchart depicting an example method 500 for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device (e.g., a computing device) is shown. Operations of method 500 may be performed by components of the drawing system 160. In some cases, in one or more of the embodiments described, certain operations of method 500 may be performed using one or more operations of method 300 or 400, or as sub-operations of one or more operations of method 300 or 400, as will be explained in more detail below.
[0068] In operation 510, the generation component 250 generates a graphical plane that is located in front of a portion of the face. The graphical plane may be transparent and extend outward from the portion of the face. The graphical plane may include a set of coordinates distributed with respect to the graphical plane. The graphical plane may be generated as an image layer within a frame of the video stream. In some cases, the image layer is generated as a transparent overlay that is located on top of an image depicted within a frame of the video stream.
[0069] In some embodiments, the graphic plane is contoured such that graphic representations can be linked to coordinates on the graphic plane that include two or more three-dimensional positions of portions of the face. In some embodiments, the contoured graphic plane is generated to span two or more image layers. The two or more image layers can overlap one another and be positioned as a transparent overlay over an image depicted within a frame of the video stream. As the image capture device rotates about a portion of the face, the two or more image layers serve as a contoured plane that covers the portion of the face from two or more perspectives within the video stream. For example, in a case where a user has drawn a pair of antlers that extend upward from a portion of the face input by the user and are rendered as a graphic representation, portions of the antlers can span two or more images that serve as the contoured graphic plane. As the image capture device rotates about the portion of the face, the antlers are represented as two-dimensional or three-dimensional antlers that extend upward from the portion of the face. As another example, in a case where user input for a graphic representation is received on respective portions of a portion of the face within the video stream, the generating component 250 can generate and can at least temporarily store graphic representations for the portion of the face that are discontinuous in view. In this way, the drawing system 160 can receive input and generate a graphic representation of a drawing that covers the entire face of the user. In a case where the user faces the image capture device, a graphic representation of a drawing on the front of the face can be invoked or generated for rendering and presentation within the video stream. In a case where the user rotates or the image capture device is rotated to show a side of the face, a graphic representation for the side of the face can be invoked or generated for rendering. Additionally, as the face or the image capture device rotates, a graphic representation of a drawing for each newly included portion of the face (e.g., an angle from a front-facing to a side-facing view) can be invoked or generated such that a smooth rotation can be rendered and displayed to include graphic representations associated with each region, location, point, or facial landmark as the regions, locations, points, or landmarks rotate into view.
[0070] In operation 520, the linking component 240 links one or more points on the graphic plane to one or more facial landmarks depicted on a portion of the face. The linking component 240 can link one or more points on the graphic plane to one or more facial landmarks in a manner similar or identical to the linking performed in operation 340.
[0071] In operation 530, the tracking component 280 tracks the movement of one or more facial landmarks. The tracking component 280 matches the movement of the one or more facial landmarks to the movement of the graphic plane. The tracking component 280 may use one or more facial tracking algorithms to track the movement of one or more facial landmarks across the frames of a video stream. For example, the tracking component 280 may use an active appearance model, principal component analysis, eigen tracking, deformable surface model, or any other suitable tracking method to track the facial landmarks. The movement of the facial landmarks may be tracked with respect to forward movement, backward movement, rotation, translation, and other movements between the frames of the video stream.
[0072] Figure 6 A flowchart illustrating an example method 600 for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device (e.g., a computing device) of a client device is shown. The operations of method 600 may be performed by components of the rendering system 160. In some cases, certain operations of method 600 may be performed using one or more operations of method 300, 400, or 500, or as sub-operations of one or more operations of method 300, 400, or 500, as will be explained in more detail below. Method 600 may enable tracking of one or more positions on a portion of a face to render the graphical representation as a persistent graphical representation.
[0073] In operation 610, the position component 230 identifies one or more positions on a portion of the face in a first subsequent frame at a first location. In some embodiments, the position component 230 may identify one or more positions on the facial portion in a manner similar or identical to operation 330 or 410.
[0074] In operation 620, the rendering component 260 renders a graphical representation at one or more positions on the portion of the face at the first location. The rendering component 260 may render the graphical representation in a manner similar or identical to the rendering performed in operation 360 above. The graphical representation may be rendered in real time within the video stream such that after the graphical representation is generated, the rendering component 260 modifies the video stream by depicting the graphical representation within one or more frames of the video stream. The graphical representation may be depicted on one or more frames of the video by being included in an image layer (e.g., a graphic plane) that overlays one or more frames.
[0075] In operation 630, the position component 230 identifies one or more positions on a portion of the face in a second subsequent frame at a second location. In some cases, the position component 230 may identify the one or more positions in the second subsequent frame in a manner similar or identical to the identification performed in operation 330, 410, or 436 above.
[0076] In operation 640, the rendering component 260 renders a graphical representation at one or more locations on a portion of the face at a second location. The rendering component 260 may render the graphical representation within a frame of the video stream in a manner similar to or the same as operation 360 or 620.
[0077] Figure 7 A flowchart illustrating an example method 700 for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device (e.g., a computing device) is shown. Operations of method 700 may be performed by components of the drawing system 160. In some cases, certain operations of method 700 may be performed using one or more operations of method 300, 400, 500, or 600, or as sub-operations of one or more operations of method 300, 400, 500, or 600, as will be explained in more detail below.
[0078] In operation 710, the input component 220 determines that a user input extends outward from a portion of the face within the video stream. In some embodiments, the input component 220 receives an input that traverses a portion of the frame of the video stream that does not include a portion of the face portion before the user input terminates. The input component 220 may determine that the user input extends outward from the portion of the face based on the position component 230 identifying the boundary of the portion of the face within the frame of the video stream.
[0079] In operation 720, the position component 230 determines one or more coordinates of a set of coordinates corresponding to the position of the user input. In some embodiments, the one or more coordinates are coordinates, points, or facial landmarks located on the portion of the face. In these embodiments, operation 720 may be performed in a manner similar to or the same as operation 330 described above. In some cases, the position component 230 identifies a set of frame coordinates that are independent of the coordinates of the portion of the face depicted within the frame. The frame coordinates divide the frame into frame regions. At least a portion of the frame coordinates may be occupied by the portion of the face. At least a portion of the frame coordinates may be occupied by the background rather than the portion of the face. In these cases, the position component 230 determines one or more coordinates of a set of coordinates for the user input corresponding to one or more frame coordinates that are not related to the portion of the face.
[0080] In operation 730, the linking component 240 links the user input to one or more coordinates on the graphical plane. In some embodiments, the linking component 240 links the user input to one or more coordinates on the graphical plane in a similar or identical manner as in operation 340 or 520 described above. The user input may be linked to one or more coordinates on the graphical plane corresponding to one or more frame coordinates. In some cases, at least a portion of the one or more coordinates linked to the user input are coordinates corresponding to one or more frame coordinates, and a portion of the one or more coordinates linked to the user input are coordinates corresponding to portions of the face.
[0081] In operation 740, the rendering component 260 renders the graphical representation. The graphical representation may be rendered to extend outward from a portion of the face on the graphical plane within one or more subsequent frames of the video stream. In some example embodiments, the rendering component 260 renders the graphical representation in a similar or identical manner as in operation 360, 620, or 640.
[0082] Figure 8 A flowchart illustrating an example method 800 for generating a graphical representation on a face within one or more frames of a video stream while the video stream is being captured and presented on a display device of a client device (e.g., a computing device) is shown. The operations of method 800 may be performed by components of the drawing system 160. In some cases, certain operations of method 800 may be performed using one or more operations of method 300, 400, 500, 600, or 700, or as sub-operations of one or more operations of method 300, 400, 500, 600, or 700, as will be explained in more detail below.
[0083] In operation 810, the input component 220 determines a first user input. The first user input extends outward from a portion of the face within the video stream. In some embodiments, the input component 220 determines the first user input in a similar or identical manner as one or more of the operations in operation 320, 420, or 710.
[0084] In operation 812, the position component 230 determines one or more coordinates of a set of coordinates corresponding to the position of the first user input. In some example embodiments, the position component 230 determines the one or more coordinates in a similar or identical manner as the operations described in operation 330, 410, or 720 above.
[0085] In operation 814, the linking component 240 links the first user input to one or more coordinates on the first graphical plane. The first graphical plane may be positioned at a first three-dimensional location with respect to a portion of the face. In some example embodiments, the linking component 240 performs operation 814 in a manner similar or identical to the operations described with respect to operation 340 or 520 above. In some cases, the linking component 240 creates a data structure that stores, at least temporarily, on a processor-readable storage device, the link between the first user input and the one or more coordinates on the first graphical plane.
[0086] In operation 816, the rendering component 260 renders a first graphical representation. The first graphical representation may be rendered to extend outward from a portion of the face on the first graphical plane within one or more subsequent frames of the video stream. In some example embodiments, the rendering component 260 performs operation 816 in a manner similar or identical to one or more of the operations described with respect to operation 360, 620, 640, or 740 above.
[0087] In operation 818, the generating component 250 generates a second graphical plane with respect to a portion of the face at a second three-dimensional location. In some example embodiments, the generating component 250 generates the second graphical plane in a manner similar or identical to one or more of operation 350 or 510. The second graphical plane may be generated as an overlay that overlaps at least a portion of the first graphical plane and at least a portion of the portion of the face. In some cases, the second graphical plane may be positioned such that as the image capture device rotates around the portion of the face, the second graphical plane appears to be located at a three-dimensional location spaced apart from the position of the first graphical plane or the portion of the face depicted within the frame of the video stream.
[0088] In operation 820, the input component 220 determines a second user input located within the frame of the video stream. In some example embodiments, the input component 220 determines the second user input location in a manner similar or identical to operation 320 or 710.
[0089] In operation 822, the linking component 240 links the second user input to one or more coordinates on the second graphical plane. In some example embodiments, the linking component 240 performs operation 822 in a manner similar or identical to operation 340, 520, or 730.
[0090] In operation 824, the generating component 250 generates a second graphical representation of the second user input at one or more coordinates on the second graphical plane. In some example embodiments, the generating component 250 generates the second graphical representation of the second user input in a manner similar or identical to operation 350, 510, or 818.
[0091] In operation 826, the rendering component 260 renders a first graphical representation on a first graphical plane and a second graphical representation on a second graphical plane. The rendering component 260 may render the first graphical representation and the second graphical representation within subsequent frames of the video stream. In some example embodiments, the rendering component 260 renders each of the first graphical representation and the second graphical representation in a manner similar or identical to operation 360, 620, 640, or other operations described above. The first graphical representation and the second graphical representation may be rendered on frames of the video on the first and second graphical planes, along with a depiction of portions of a face within the frame. The rendering component 260 may render the first graphical representation and the second graphical representation in one or more frames of the video stream in real time once generated.
[0092] Example
[0093] To better illustrate the apparatus and methods disclosed herein, a non-limiting list of examples is provided here:
[0094] 1. A method, comprising: receiving, by one or more processors, one or more frames of a video stream, at least a portion of the one or more frames depicting at least a portion of a face; identifying a user input on an input device of a computing device; identifying one or more locations on a portion of the face corresponding to the user input; linking the user input to the one or more locations on the portion of the face; generating a graphical representation of the user input, the graphical representation of the user input being linked to the one or more locations on the portion of the face; and rendering the graphical representation on the portion of the face within one or more subsequent frames of the video stream, the graphical representation being presented at the one or more locations on the portion of the face.
[0095] 2. The method according to example 1, wherein identifying the user input further comprises: determining a first relative distance of the portion of the face from an image capture device of the computing device in a frame of the video stream; determining a size characteristic of the graphical representation based on the first relative distance of the portion of the face; and scaling the size characteristic of the graphical representation based on movement of the portion of the face from the first relative distance to a second relative distance in subsequent frames of the video stream.
[0096] 3. The method according to example 1 or 2, wherein scaling the size characteristic of the graphical representation further comprises: identifying a first line width of the graphical representation at the first relative distance; determining a first relative position of at least one point on the graphical representation and two or more facial landmarks depicted on the portion of the face; determining a change in distance between the two or more facial landmarks at the second relative distance; and modifying the first line width for the graphical representation based on the change in distance between the two or more facial landmarks to generate a second line width for the graphical representation.
[0097] 4. The method according to any one or more of Examples 1-3, wherein identifying one or more positions on a portion of the face corresponding to the user input further comprises: determining a face mesh on the portion of the face, the face mesh identifying one or more facial landmarks depicted within the portion of the face; identifying one or more coordinates of the user input; and mapping the one or more coordinates to at least a portion of the one or more facial landmarks.
[0098] 5. The method according to any one or more of Examples 1-4, further comprising: generating a graphic plane in front of the portion of the face; linking one or more points on the graphic plane to one or more facial landmarks depicted on the facial portion; and tracking the movement of the one or more facial landmarks and matching the movement of the one or more facial landmarks with the movement of the graphic plane.
[0099] 6. The method according to any one or more of Examples 1-5, wherein the graphic plane is contoured such that the graphic representation linked to the coordinates on the graphic plane includes two or more three-dimensional positions relative to the portion of the face.
[0100] 7. The method according to any one or more of Examples 1-6, wherein the graphic plane is transparent, extends outward from the face, and has a set of coordinates distributed with respect to the graphic plane, and further comprises: determining that the user input extends outward from the facial portion; determining one or more coordinates of the set of coordinates corresponding to the position of the user input; linking the user input to one or more coordinates on the graphic plane; and rendering, within one or more subsequent frames of the video stream, a graphic representation extending outward from the face on the graphic plane.
[0101] 8. The method according to any one or more of Examples 1-7, wherein the graphic plane is a first graphic plane located at a first three-dimensional position relative to the facial portion, and the graphic representation is a first graphic representation, and further comprises: generating a second graphic plane at a second three-dimensional position relative to the portion of the face; determining a second user input located within the frame of the video stream; linking the second user input to one or more coordinates on the second graphic plane; generating a second graphic representation of the second user input at the one or more coordinates on the second graphic plane; and rendering the first graphic representation on the first graphic plane and rendering the second graphic representation on the second graphic plane.
[0102] 9. The method according to any one or more of Examples 1-8, wherein the method further comprises tracking one or more positions on the facial portion to render the graphical representation as a persistent graphical representation by: identifying one or more positions on the portion of the face in a first subsequent frame at a first location; rendering the graphical representation at the one or more positions on the facial portion at the first location; identifying one or more positions on the portion of the face in a second subsequent frame at a second location; and rendering the graphical representation at the one or more positions on the facial portion at the second location.
[0103] 10. The method according to any one or more of Examples 1-9, further comprising: storing the graphical representation in a processor-readable storage device; detecting a portion of a face within a new video stream in response to launching an application; and rendering the graphical representation within one or more frames of the new video stream in response to detecting the portion of the face within the new video stream.
[0104] 11. The method according to any one or more of Examples 1-10, wherein the input device is a touchscreen device configured to be pressure-sensitive to touches providing user input, and wherein identifying the user input further comprises: determining the touch pressure of the user input; identifying edge features associated with the touch pressure; and assigning the edge features to the graphical representation generated from the user input.
[0105] 12. A system comprising: one or more processors; and a non-transitory processor-readable storage medium coupled to the one or more processors and storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including: receiving, by the one or more processors, one or more frames of a video stream, at least a portion of the one or more frames depicting at least a portion of a face; identifying a user input on an input device of a computing device; identifying one or more positions on the portion of the face corresponding to the user input; linking the user input to the one or more positions on the portion of the face; generating a graphical representation of the user input, the graphical representation of the user input being linked to the one or more positions on the portion of the face; and rendering the graphical representation on the portion of the face within one or more subsequent frames of the video stream, the graphical representation being presented at the one or more positions on the portion of the face.
[0106] 13. The system according to Example 12, wherein identifying the user input further comprises: determining a first relative distance of the facial portion from an image capture device of the computing device in a frame of the video stream; determining a size feature of the graphical representation based on the first relative distance of the portion of the face; and scaling the size feature of the graphical representation based on movement of the portion of the face from the first relative distance to a second relative distance in subsequent frames of the video stream.
[0107] 14. The system according to example 12 or 13, wherein the dimensional feature of the scaled graphical representation further comprises: identifying a first line width of the graphical representation at a first relative distance; determining a first relative position of at least one point on the graphical representation and two or more facial landmarks depicted on a portion of the face; determining a change in the distance between two or more facial landmarks at a second relative distance; and modifying the first line width of the graphical representation based on the change in the distance between the two or more facial landmarks to generate a second line width of the graphical representation.
[0108] 15. The system according to any one or more of examples 10-14, wherein identifying one or more positions on a portion of the face corresponding to a user input further comprises: determining a facial mesh on the portion of the face, the facial mesh identifying one or more facial landmarks depicted within the portion of the face; identifying one or more coordinates of the user input; and mapping the one or more coordinates to at least a portion of the one or more facial landmarks.
[0109] 16. The system according to any one or more of examples 10-15, wherein the operation further comprises: generating a graphical plane in front of the portion of the face; linking one or more points on the graphical plane to one or more facial landmarks depicted on the portion of the face; and tracking the movement of the one or more facial landmarks and matching the movement of the one or more facial landmarks with the movement of the graphical plane.
[0110] 17. A processor-readable storage medium storing processor-executable instructions that, when executed by a processor of a machine, cause the machine to perform operations including: receiving, by one or more processors, one or more frames of a video stream, at least a portion of the one or more frames depicting at least a portion of a face; identifying a user input on an input device of the computing device; identifying one or more positions on a portion of the face corresponding to the user input; linking the user input to the one or more positions on the portion of the face; generating a graphical representation of the user input, the graphical representation of the user input being linked to the one or more positions on the portion of the face; and rendering the graphical representation on the portion of the face within one or more subsequent frames of the video stream, the graphical representation being presented at the one or more positions on the portion of the face.
[0111] 18. The processor-readable storage medium according to example 17, wherein identifying the user input further comprises: determining a first relative distance of the portion of the face to an image capture device of the computing device in a frame of the video stream; determining a dimensional feature of the graphical representation based on the first relative distance of the portion of the face; and scaling the dimensional feature of the graphical representation based on the movement of the portion of the face from the first relative distance to a second relative distance in subsequent frames of the video stream.
[0112] 19. The processor-readable storage medium according to Example 17 or 18, wherein the size feature of the scaled graphical representation further includes: identifying a first line width of the graphical representation at a first relative distance; determining a first relative position of at least one point on the graphical representation and two or more facial landmarks depicted on a portion of the face; determining a change in the distance between two or more facial landmarks at a second relative distance; and modifying the first line width of the graphical representation based on the change in the distance between two or more facial landmarks to generate a second line width of the graphical representation.
[0113] 20. The processor-readable storage medium according to any one or more of Examples 17 - 19, wherein identifying one or more positions on a portion of the face corresponding to a user input further includes: determining a facial mesh on the portion of the face, the facial mesh identifying one or more facial landmarks depicted within the portion of the face; identifying one or more coordinates of the user input; and mapping the one or more coordinates to at least a portion of the one or more facial landmarks.
[0114] 21. A machine-readable medium carrying processor-executable instructions that, when executed by one or more processors of a machine, cause the machine to perform the method according to any one of Examples 1 to 11.
[0115] These and other examples and features of the apparatus and method are set forth in part above in the detailed description. The summary and examples are intended to provide non-limiting embodiments of the subject matter. It is not intended to provide an exclusive or exhaustive explanation. The detailed description is included to provide further information about the subject matter.
[0116] Modules, Components, and Logic
[0117] Certain embodiments are described herein as including logic or multiple components, modules, or mechanisms. A component can constitute a hardware component. A "hardware component" is a tangible unit capable of performing certain operations and can be configured or arranged in a certain physical manner. In various example embodiments, a computer system (e.g., a stand-alone computer system, a client computer system, or a server computer system) or a hardware component of a computer system (e.g., at least one hardware processor, a processor, or a group of processors) is configured by software (e.g., an application or a portion of an application) as a hardware component for performing certain operations as described herein.
[0118] In some embodiments, the hardware components are implemented mechanically, electronically, or any suitable combination thereof. For example, the hardware components can include dedicated circuits or logic that are permanently configured to perform certain operations. For example, the hardware component can be a dedicated processor, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). The hardware components can also include programmable logic or circuits that are temporarily configured by software to perform certain operations. For example, the hardware components can include software contained within a general purpose processor or other programmable processor. It should be understood that the decision to implement the hardware components mechanically in dedicated and permanently configured circuits or in temporarily configured circuits (e.g., configured by software) can be driven by cost and time considerations.
[0119] Accordingly, the phrase "hardware component" should be understood to include a tangible entity, an entity that is a physical construction, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) so as to operate in a certain manner or to perform certain operations as described herein. As used herein, "hardware-implemented component" refers to a hardware component. Considering embodiments in which a hardware component is temporarily configured (e.g., programmed), it is not necessary to configure or instantiate every hardware component in the hardware components at any one time. For example, in the case where a hardware component includes a general purpose processor that is configured by software to be a dedicated processor, the general purpose processor can be configured to be different dedicated processors (e.g., including different hardware components) at different times. Thus, software can configure a particular one or more processors, e.g., to constitute a particular hardware component at one moment and a different hardware component at a different moment.
[0120] The hardware components can provide information to and receive information from other hardware components. Accordingly, the described hardware components can be regarded as communicatively coupled. In cases where there are multiple hardware components present simultaneously, communication can be achieved via signal transmission (e.g., via appropriate circuits and buses) between or among two or more hardware components. In embodiments in which multiple hardware components are configured or instantiated at different times, communication between these hardware components can be achieved, for example, by storing and retrieving information in a memory structure accessible by the multiple hardware components. For example, one hardware component performs an operation and stores the output of that operation in a memory device communicatively coupled thereto. Then, another hardware component can later access the memory device to retrieve and process the stored output. The hardware components can also initiate communication with input or output devices and can operate on resources (e.g., collections of information).
[0121] The various operations of the example methods described herein can be performed, at least in part, by a processor temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such a processor constitutes a processor-implemented component that operates to perform the operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using a processor.
[0122] Similarly, the methods described herein can be implemented, at least in part, by a processor, where a particular processor or processors are examples of hardware. For example, at least some of the operations of the method can be performed by a processor or a processor-implemented component. Additionally, the processor can also operate to support the performance of the relevant operations in a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some of the operations can be performed by a group of computers (as an example of a machine including a processor), and these operations can be accessed via a network (e.g., the Internet) and via an appropriate interface (e.g., an application programming interface (API)).
[0123] The performance of certain operations can be distributed among processors, not only residing within a single machine but also deployed across multiple machines. In some example embodiments, the processor or processor-implemented component is located in a single geographical location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, the processor or processor-implemented component is distributed across multiple geographical locations.
[0124] Application
[0125] Figure 9 An example mobile device 900 is shown that executes a mobile operating system (e.g., IOS TM , ANDROID TM , Phone or other mobile operating system) consistent with some embodiments. In one embodiment, the mobile device 900 includes a touch screen operable to receive touch data from a user 902. For example, the user 902 can physically touch 904 the mobile device 900, and in response to the touch 904, the mobile device 900 can determine touch data such as the touch location, touch force, or gesture movement. In various example embodiments, the mobile device 900 displays a home screen 906 (e.g., IOS TMon the Springboard), which is operable to launch applications or otherwise manage various aspects of the mobile device 900. In some example embodiments, the home screen 906 provides status information such as battery life, connectivity, or other hardware status. The user 902 can activate a user interface element by touching the area occupied by the corresponding user interface element. In this way, the user 902 interacts with the applications of the mobile device 900. For example, touching the area occupied by a specific icon included in the home screen 906 causes the launch of the application corresponding to the specific icon.
[0126] As Figure 9 shown, the mobile device 900 may include an imaging device 908. The imaging device 908 may be a camera or any other device coupled to the mobile device 900 capable of capturing a video stream or one or more consecutive images. The imaging device 908 may be triggered by the rendering system 160 or a selectable user interface element to initiate the capture of a video stream or continuum of frames and pass the video stream or continuum of images to the rendering system 160 for processing according to one or more methods described in this disclosure.
[0127] Many types of applications (also referred to as "application software") may be executed on the mobile device 900, such as native applications (e.g., applications programmed in Objective-C, Swift, or another suitable language running on IOS TM or applications programmed in Java running on ANDROID TM ), mobile web applications (e.g., applications written in Hypertext Markup Language - 5 (HTML5)), or hybrid applications (e.g., native shell applications that launch an HTML5 session). For example, the mobile device 900 includes messaging application software, audio recording application software, camera application software, book reader application software, media application software, fitness application software, file management application software, location application software, browser application software, settings application software, contacts application software, phone call application software, or other application software (e.g., game application software, social network application software, biometric monitoring application software). In another example, the mobile device 900 includes such as A social messaging application software 910, which is consistent with some embodiments, allows users to exchange ephemeral messages including media content. In this example, the social messaging application software 910 may incorporate aspects of the embodiments described herein. For example, in some embodiments, the social messaging application includes an ephemeral media gallery created by the user's social messaging application. These galleries may consist of videos or pictures posted by the user and viewable by the user's contacts (e.g., "friends"). Alternatively, a public gallery may be created by an administrator of the social messaging application, which consists of media from any user of the application (and accessible by all users). In yet another embodiment, the social messaging application may include a "magazine" feature, which consists of articles and other content generated by publishers on the platform of the social messaging application and accessible by any user. Any of these environments or platforms can be used to implement the concepts of the present invention.
[0128] In some embodiments, the ephemeral messaging system may include a message having an ephemeral video clip or image that is deleted after a deletion trigger event such as a viewing time or viewing completion. In such an embodiment, the apparatus implementing the rendering system 160 may identify, track, extract, and generate a facial representation within the ephemeral video clip as the ephemeral video clip is acquired by the apparatus, and send the ephemeral video clip to another apparatus using the ephemeral messaging system.
[0129] Software architecture
[0130] Figure 10 is a block diagram 1000 showing the architecture of software 1002 that can be installed on the above-described apparatus. Figure 10 This is merely a non-limiting example of a software architecture, and it will be understood that many other architectures can be implemented to facilitate the functions described herein. In various embodiments, the software 1002 is implemented by hardware of a machine 1100 such as Figure 11 The machine 1100 includes a processor 1110, a memory 1130, and I / O components 1150. In this example architecture, the software 1002 can be conceptualized as a stack of layers, where each layer may provide a specific function. For example, the software 1002 includes layers such as an operating system 1004, libraries 1006, frameworks 1008, and applications 1010. Operationally, consistent with some embodiments, the application 1010 makes application programming interface (API) calls 1012 through the software stack and receives messages 1014 in response to the API calls 1012.
[0131] In various embodiments, the operating system 1004 manages hardware resources and provides common services. The operating system 1004 includes, for example, a kernel 1020, services 1022, and drivers 1024. Consistent with some embodiments, the kernel 1020 acts as an abstraction layer between the hardware and other software layers. For example, the kernel 1020 provides functions such as memory management, processor management (e.g., scheduling), component management, network connectivity, and security settings. The services 1022 may provide other common services to other software layers. According to some embodiments, the drivers 1024 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 1024 may include a display driver, a camera driver, a driver, a flash drive, a serial communication driver (e.g., a Universal Serial Bus (USB) driver), a driver, an audio driver, a power management driver, etc.
[0132] In some embodiments, the library 1006 provides a low-level common infrastructure utilized by the applications 1010. The library 1006 may include system libraries 1030 (e.g., the C standard library), which may provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, the library 1006 may include API libraries 1032, such as media libraries (e.g., libraries that support the presentation and manipulation of various media formats such as Moving Picture Experts Group-4 (MPEG4), High Efficiency Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., the OpenGL framework for presenting two-dimensional (2D) and three-dimensional (3D) graphics content on a display), database libraries (e.g., SQLite that provides various relational database functions), web libraries (e.g., WebKit that provides web browsing functionality), etc. The library 1006 may also include a variety of other libraries 1034 to provide many other APIs to the applications 1010.
[0133] According to some embodiments, the framework 1008 provides a high-level common architecture that can be utilized by the applications 1010. For example, the framework 1008 provides various Graphical User Interface (GUI) functions, high-level resource management, high-level location nodes, etc. The framework 1008 may provide a wide range of other APIs that can be utilized by the applications 1010, some of which may be specific to a particular operating system or platform.
[0134] In an example embodiment, the application 1010 includes a home page application 1050, a contacts application 1052, a browser application 1054, a book reader application 1056, a location application 1058, a media application 1060, a messaging application 1062, a game application 1064, and other widely classified applications such as third-party applications 1066. According to some embodiments, the application 1010 is a program that executes functions defined in the program. The application 1010 can be created using various programming languages and structured in various ways, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a specific example, the third-party application 1066 (e.g., an application developed using an ANDROID TM or IOS TM software development kit (SDK) by an entity other than the vendor of a specific platform) can be mobile software that runs on a mobile operating system such as IOS TM 、ANDROID TM 、 PHONE or other mobile operating systems). In this example, the third-party application 1066 can call the API call 1012 provided by the operating system 1004 to facilitate the execution of the functions described herein.
[0135] Example Machine Architecture and Machine-Readable Medium
[0136] Figure 11 is a block diagram showing components of a machine 1100 that can read instructions (e.g., processor-executable instructions) from a machine-readable medium (e.g., a non-transitory processor-readable storage medium or a processor-readable storage device) and execute any of the methods discussed herein. Specifically, Figure 11FIG. 0 shows a schematic diagram of a machine 1100 in the example form of a computer system within which instructions 1116 (e.g., software, program, application, applet, application program, or other executable code) can be executed to cause the machine 1100 to perform any of the methods discussed herein. In alternative embodiments, the machine 1100 operates as a stand-alone device or can be coupled (e.g., network-connected) to other machines. In a networked deployment, the machine 1100 can operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1100 can include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a network device, a network router, a network switch, a network bridge, or any machine capable of executing the instructions 1116 that specifies actions to be taken by the machine 1100, either continuously or otherwise. Further, although only a single machine 1100 is shown, the term "machine" can also be regarded as including a collection of machines 1100 that individually or jointly execute the instructions 1116 to perform any of the methods discussed herein.
[0137] In various embodiments, the machine 1100 includes a processor 1110, a memory 1130, and I / O components 1150 that can be configured to communicate with each other via a bus 1102. In an example embodiment, the processor 1110 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) includes, for example, a processor 1112 and a processor 1114 that can execute the instructions 1116. The term "processor" is intended to include multi-core processors that can include more than two independent processors (also referred to as "cores") that can execute the instructions 1116 simultaneously. Although Figure 11 multiple processors 1110 are shown, the machine 1100 can include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0138] According to some embodiments, the memory 1130 includes a main memory 1132, a static memory 1134, and a storage unit 1136 accessible to the processor 1110 via the bus 1102. The storage unit 1136 may include a machine-readable medium 1138 on which instructions 1116 embodying any of the methods or functions described herein are stored. The instructions 1116 may also reside, completely or at least partially, within at least one of the main memory 1132, the static memory 1134, the processor 1110 (e.g., within a cache of the processor) or any suitable combination during execution of the instructions by the machine 1100. Thus, in various embodiments, the main memory 1132, the static memory 1134, and the processor 1110 are considered to be machine-readable media 1138.
[0139] As used herein, the term "memory" refers to a machine-readable medium 1138 capable of storing data temporarily or permanently and may be considered to include, but not be limited to, random access memory (RAM), read-only memory (ROM), cache, flash memory, and buffer memory. Although the machine-readable medium 1138 is shown as a single medium in the exemplary embodiments, the term "machine-readable medium" should be considered to include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) capable of storing the instructions 1116. The term "machine-readable medium" may also be regarded as including any medium or combination of media capable of storing instructions (e.g., the instructions 1116) for execution by a machine (e.g., the machine 1100) such that the instructions, when executed by a processor (e.g., the processor 1110) of the machine 1100, cause the machine 1100 to perform any of the methods described herein. Thus, the term "machine-readable medium" refers to a single storage device or apparatus, as well as a "cloud-based" storage system or storage network including multiple storage devices or apparatuses. Thus, the term "machine-readable medium" may be considered to include, but not be limited to, data repositories in the form of solid-state memory (e.g., flash memory), optical media, magnetic media, other non-volatile memory (e.g., erasable programmable read-only memory (EPROM)), or any suitable combination thereof. The term "machine-readable medium" may include the signal itself.
[0140] The I / O components 1150 include a variety of components for receiving input, providing output, generating output, sending information, exchanging information, taking measurements, and the like. In general, it is understood that the I / O components 1150 may include Figure 11Many other components not shown. The I / O components 1150 are grouped according to function only for the sake of simplicity in the following discussion, and the grouping is in no way restrictive. In various example embodiments, the I / O components 1150 include an output component 1152 and an input component 1154. The output component 1152 includes visual components (e.g., a display, such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), auditory components (e.g., speakers), tactile components (e.g., a vibration motor), other signal generators, etc. The input component 1154 includes alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instrument), tactile input components (e.g., physical buttons, a touch screen that provides the location and force of a touch or touch gesture, or other tactile input components), audio input components (e.g., a microphone), etc.
[0141] In some additional example embodiments, the I / O components 1150 include a biometric component 1156, a motion component 1158, an environmental component 1160, or a location component 1162 among various other components. For example, the biometric component 1156 includes components that detect expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or mouth postures), measure biometric signals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identify people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion component 1158 includes an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. The environmental component 1160 includes, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., a thermometer that detects the ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., a microphone that detects background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor component (e.g., a machine olfactory detection sensor, a gas detection sensor for detecting the concentration of a hazardous gas for safety or measuring pollutants in the atmosphere), or other components that may provide an indication, measurement, or signal corresponding to the surrounding physical environment. The location component 1162 includes a positioning sensor component (e.g., a global positioning system (GPS) receiver component), an altitude sensor component (e.g., an altimeter or a barometer that can detect the air pressure from which the altitude can be derived), an orientation sensor component (e.g., a magnetometer), etc.
[0142] Communication can be implemented using a variety of technologies. The I / O component 1150 can include a communication component 1164 that is operable to couple the machine 1100 to the network 1180 or the device 1170 via the coupler 1182 and the coupler 1172, respectively. For example, the communication component 1164 includes a network interface component or another suitable device that interfaces with the network 1180. In another example, the communication component 1164 includes a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, components (e.g., low power ), components and other communication components that provide communication via other modes. The device 1170 can be any one of another machine or a variety of peripheral devices (e.g., a peripheral device coupled via a universal serial bus (USB)).
[0143] In addition, in some embodiments, the communication component 1164 detects an identifier or includes a component operable to detect an identifier. For example, the communication component 1164 includes a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as universal product code (UPC) barcodes, and multi-dimensional barcodes such as quick response (QR) codes, Aztec codes, data matrix, digital graphics, maxi codes, PDF417, hypercodes, uniform commercial code reduced space symbol (UCC RSS)-2D barcodes, and other optical codes), an acoustic detection component (e.g., a microphone for identifying an audio signal of a tag), or any suitable combination thereof. In addition, various information such as a location via Internet protocol (IP) geolocation, a location via signal triangulation, a location via detecting or an NFC beacon signal, etc. can be derived via the communication component 1164 that can indicate a specific location.
[0144] Transmission medium
[0145] In various example embodiments, portions of the network 1180 can be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, A network, another type of network, or a combination of two or more such networks. For example, network 1180 or a portion of network 1180 may include a wireless or cellular network, and coupling 1182 may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or another type of cellular or wireless coupling. In this example, coupling 1182 may implement any one of a variety of types of data transfer technologies, such as Single-Carrier Radio Transmission Technology (1xRTT), Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, GSM Enhanced Data Rates for GSM Evolution (EDGE) technology, 3rd Generation Partnership Project (3GPP) including 3G, Fourth Generation Wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Worldwide Interoperability for Microwave Access (WiMAX), Long Term Evolution (LTE) standards, other standards defined by various standards organizations, other remote protocols, or other data transfer technologies.
[0146] In an example embodiment, instructions 1116 are sent or received via network 1180 using a transmission medium (e.g., a network interface component included in communication component 1164) and utilizing any one of a plurality of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, in other example embodiments, instructions 1116 are sent or received to or from device 1170 via coupling 1172 (e.g., a peer-to-peer coupling) using a transmission medium. The term "transmission medium" may be regarded as including any non-transitory medium that is capable of storing, encoding, or carrying instructions 1116 executable by machine 1100 and includes digital or analog communication signals or other non-transitory media for facilitating such communication implementation of software.
[0147] Furthermore, since machine-readable medium 1138 does not embody a propagated signal, machine-readable medium 1138 is non-transitory (in other words, does not have any transient signals). However, labeling machine-readable medium 1338 as "non-transitory" should not be construed to mean that the medium cannot be moved. The medium should be considered to be transportable from one physical location to another. Additionally, since machine-readable medium 1138 is tangible, the medium can be considered to be a machine-readable device.
[0148] Language
[0149] Throughout the specification, multiple instances may implement components, operations, or structures described as a single instance. Although the separate operations of a method are shown and described as separate operations, the separate operations may be performed concurrently, and the operations need not be performed in the order shown. Structures and functions presented as separate components in an example configuration may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate multiple components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0150] Although the subject matter of the present invention has been described in terms of specific example embodiments, various modifications and changes may be made to these embodiments without departing from the broader scope of the disclosed embodiments. Such embodiments of the subject matter of the present invention may herein be referred to individually or collectively by the term "invention," solely for convenience, and are not intended to limit the scope of the present application to any single disclosure or inventive concept if more than one is in fact disclosed.
[0151] The embodiments shown herein are described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments may be used and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of the disclosure. Accordingly, the detailed description should not be considered limiting, and the scope of the various embodiments is defined only by the appended claims and the full scope of equivalents given by those claims.
[0152] As used herein, the term "or" may be interpreted in an inclusive or exclusive manner. In addition, multiple instances may be provided for resources, operations, or structures described herein as a single instance. Further, the boundaries between various resources, operations, components, engines, and data stores are to some extent arbitrary, and particular operations are shown in the context of a particular illustrative configuration. Other allocations of functionality are contemplated and may fall within the scope of the various embodiments of the disclosure. In general, structures and functions presented as separate resources in an example configuration may be implemented as a combined structure or resource. Similarly, structures and functions presented as a single resource may be implemented as separate resources. These and other variations, modifications, additions, and improvements fall within the scope of the embodiments of the disclosure as expressed by the appended claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
Claims
1. A method for real-time rendering of a video stream, comprising: Receiving, by one or more processors, a video stream depicting an object; Receiving, from a user, an input selecting a timing mechanism that defines when to render a graphical representation; Determining that the input extends outward from the object; And Rendering the graphical representation relative to the object within one or more frames of the video stream based on the timing mechanism and in response to determining that the input extends outward from the object, the graphical representation extending outward from the object on a transparent graphical plane within one or more frames of the video stream.
2. The method according to claim 1, wherein, The input is received during sequential presentation of multiple frames of the video stream, wherein the timing mechanism includes a set time period for presenting the graphical representation, and wherein the object includes at least a portion of a face, and the method further comprises: Identifying one or more positions on the portion of the face corresponding to the input; Linking the input to the one or more positions on the portion of the face; Generating the graphical representation of the input, the graphical representation of the input being linked to the one or more positions on the portion of the face; and Rendering the graphical representation on the portion of the face within one or more subsequent frames of the video stream in response to detecting an event for triggering presentation of the graphical representation and based on the timing mechanism, the graphical representation being presented at the one or more positions on the portion of the face.
3. The method according to claim 1, wherein Receiving the input includes: Determining a first relative distance of the object in a frame of the video stream to an image acquisition device; Determining a size characteristic of the graphical representation based on the first relative distance of the object; and Scaling the size characteristic of the graphical representation based on the object moving from the first relative distance to a second relative distance in subsequent frames of the video stream.
4. The method according to claim 3, wherein Scaling the size characteristic of the graphical representation includes: Identifying a first line width of the graphical representation at the first relative distance; Determining a first relative position of at least one point on the graphical representation and two or more landmarks depicted on the object; Based on the first relative position, determining a change in distance between the two or more landmarks at the second relative distance; and Modifying the first line width of the graphical representation based on the change in distance between the two or more landmarks to generate a second line width of the graphical representation.
5. The method according to claim 1, further comprising: Generating a graphical plane located in front of the object; Connecting one or more points on the graphical plane to one or more landmarks depicted on the object; And Tracking movement of the one or more landmarks and matching the movement of the one or more landmarks with movement of the graphical plane.
6. The method according to claim 5, wherein, The graphical plane is contoured such that graphical representations linked to coordinates on the graphical plane include two or more three-dimensional positions relative to the object.
7. The method according to claim 5, wherein The graphical plane has a set of coordinates distributed with respect to the graphical plane, and the method further comprises: Determine one or more coordinates in the set of coordinates corresponding to the position of the input; and Link the input to the one or more coordinates on the transparent graphic plane.
8. The method according to claim 7, wherein The transparent graphic plane is a first graphic plane located at a first three-dimensional position relative to the object, and the graphic representation is a first graphic representation, and the method further includes: Generate a second graphic plane at a second three-dimensional position relative to the object; Determine a second input located within a frame of the video stream; Link the second input to one or more coordinates on the second graphic plane; Generate a second graphic representation of the second input at the one or more coordinates on the second graphic plane; and [[ID= 11. The method according to claim 1, wherein, 14. The method according to claim 2, wherein 15. The method according to claim 1, wherein, 16. A system for real-time rendering of a video stream, comprising: One or more processors; And A non-transitory processor-readable storage medium coupled to the one or more processors and storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: Receiving a video stream depicting an object; Receiving from a user an input selecting a timing mechanism that defines when to render a graphical representation; Determining that the input extends outward from the object; and Based on the timing mechanism and in response to determining that the input extends outward from the object, rendering the graphical representation within one or more frames of the video stream relative to the object, the graphical representation extending outward from the object on a transparent graphical plane within one or more frames of the video stream.
17. The system according to claim 16, wherein The input is received during sequential presentation of multiple frames of the video stream, wherein the timing mechanism includes a set time period for presenting the graphical representation, and wherein the object includes at least a portion of a face, and the operations further include: Identifying one or more locations on the portion of the face corresponding to the input; Linking the input to the one or more locations on the portion of the face; Generating the graphical representation of the input, the graphical representation of the input being linked to the one or more locations on the portion of the face; and In response to detecting an event for triggering presentation of the graphical representation and based on the timing mechanism, rendering the graphical representation on the portion of the face within one or more subsequent frames of the video stream, the graphical representation being presented at the one or more locations on the portion of the face.
18. The system according to claim 16, wherein, The operation for receiving the input includes: Determining a first relative distance from the object in a frame of the video stream to an image acquisition device; Determining a size characteristic of the graphical representation based on the first relative distance of the object; and Scaling the size characteristic of the graphical representation based on the object moving from the first relative distance to a second relative distance in subsequent frames of the video stream.
19. A non-transitory processor-readable storage medium storing processor-executable instructions that, when executed by a processor of a machine, cause the machine to perform operations, the operations including: Receiving a video stream depicting an object; Receiving from a user an input selecting a timing mechanism that defines when to render a graphical representation; Determining that the input extends outward from the object; And Based on the timing mechanism and in response to determining that the input extends outward from the object, rendering the graphical representation within one or more frames of the video stream relative to the object, the graphical representation extending outward from the object on a transparent graphical plane within one or more frames of the video stream.
20. The non-transitory processor-readable storage medium according to claim 19, wherein, The input is received during sequential presentation of multiple frames of the video stream, wherein the timing mechanism includes a set time period for presenting the graphical representation, and wherein the object includes at least a portion of a face, and the operations further include: Identifying one or more locations on the portion of the face corresponding to the input; Linking the input to the one or more locations on the face; Generating a graphical representation of the input, the graphical representation of the input being linked to the one or more locations on the portion of the face; and In response to detecting an event for triggering presentation of the graphical representation and based on the timing mechanism, rendering the graphical representation on the portion of the face in one or more subsequent frames of the video stream, the graphical representation being presented at the one or more locations on the portion of the face.
21. A method for real-time rendering of a video stream, comprising: Receiving, by one or more processors, a video stream depicting an object; Receiving an input during sequential presentation of multiple frames of the video stream, the input representing a selection of a timing mechanism that includes a set time period for presenting a graphical representation; Determining that the input extends outward from the object; And In response to determining that the input extends outward from the object and in response to detecting an event for triggering presentation of the graphical representation and based on the timing mechanism, rendering the graphical representation relative to the object in one or more subsequent frames of the video stream, the graphical representation extending outward from the object on a transparent graphical plane in one or more subsequent frames of the video stream.
22. A system for real-time rendering of a video stream, comprising: One or more processors; And A non-transitory processor-readable storage medium coupled to the one or more processors and storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations that include: Receiving a video stream depicting an object; Receiving an input during sequential presentation of multiple frames of the video stream, the input representing a selection of a timing mechanism that includes a set time period for presenting a graphical representation; Determining that the input extends outward from the object; and In response to determining that the input extends outward from the object and in response to detecting an event for triggering presentation of the graphical representation and based on the timing mechanism, rendering the graphical representation relative to the object in one or more subsequent frames of the video stream, the graphical representation extending outward from the object on a transparent graphical plane in one or more frames of the video stream.
23. A non-transitory processor-readable storage medium storing processor-executable instructions that, when executed by a processor of a machine, cause the machine to perform operations that include: Receiving a video stream depicting an object; Receiving an input during sequential presentation of multiple frames of the video stream, the input representing a selection of a timing mechanism that includes a set time period for presenting a graphical representation; Determine that the input extends outward from the object; and In response to determining that the input extends outward from the object and in response to detecting an event for triggering the rendering of the graphical representation and based on the timing mechanism, render the graphical representation relative to the object within one or more subsequent frames of the video stream, the graphical representation extending outward from the object on a transparent graphical plane within one or more frames of the video stream.
24. A method for real-time rendering of a video stream, comprising: Receiving, by one or more processors, a video stream depicting an object; Receiving an input from a user; Determining that the input extends outward from the object; and In response to determining that the input extends outward from the object, rendering a graphical representation relative to the object within one or more frames of the video stream, the graphical representation extending outward from the object on a transparent graphical plane within one or more frames of the video stream.
25. An apparatus for real-time rendering of a video stream, comprising: Means for receiving, by one or more processors, a video stream depicting an object; Means for receiving an input from a user; Means for determining that the input extends outward from the object; and Means for rendering a graphical representation relative to the object within one or more frames of the video stream in response to determining that the input extends outward from the object, the graphical representation extending outward from the object on a transparent graphical plane within one or more frames of the video stream.
Citation Information
Patent Citations
Real-time animations of emoticons using facial recognition during a video chat
US20120069028A1
Modifying Video Call Data
US20160127681A1