Object modeling and replacement in video streams
By identifying and tracking objects of interest in a video stream, generating graphical interface elements and enabling interaction with them, the limitations of existing technologies in video stream modification and game interaction are overcome, achieving real-time modeling of video streams and a rich interactive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SNAP INC
- Filing Date
- 2017-06-30
- Publication Date
- 2026-04-10
AI Technical Summary
Existing telecommunications devices and video game applications have limitations in modifying and interacting with video streams, making it impossible to effectively model and modify them during video stream acquisition, resulting in limited communication and gaming experiences.
By identifying and tracking objects of interest, such as hands and fingers, in a video stream, graphical interface elements can be generated and interacted with, enabling real-time modification of the video stream and creation of the game environment.
It enables real-time object modeling and modification during video stream acquisition, enhancing the flexibility and richness of video communication and game interaction.
Smart Images

Figure CN116139466B_ABST
Abstract
Description
[0001] This application is a continuation of patent application with application number 201780053008.6, filed on June 30, 2017, having the title "Object Modeling and Replacement in Video Streams", and priority claim.
[0002] CLAIM OF PRIORITY
[0003] This application claims priority to U.S. Patent Application Serial No. 15 / 199,482, filed on June 30, 2016, the priority of each of which is hereby claimed, and each of which is incorporated by reference herein in its entirety. TECHNICAL FIELD
[0004] Embodiments of the present disclosure generally relate to automatic image segmentation of video streams. More specifically, but not by way of limitation, the present disclosure presents systems and methods for segmenting video streams to generate models of objects and replace depictions of objects within the video streams. BACKGROUND
[0005] Telecommunications applications and devices can use various media, such as text, images, sound recordings, and / or video recordings, to provide communication between multiple users. For example, video conferencing allows two or more individuals to communicate with each other using a combination of software applications, telecommunications devices, and telecommunications networks. Telecommunications devices can also record video streams for transmission as messages between telecommunications networks.
[0006] Video games generally enable a user to control an interactive element depicted on a display device to interact with predetermined objects within a programmed interaction. Users playing video games often control a predetermined character in a game environment determined prior to the start of a game session or generally proceed through a predetermined game play that includes a set of preprogrammed events or challenges. BRIEF DESCRIPTION OF DRAWINGS
[0007] The various drawings of the attached figures only show example embodiments of the present disclosure and should not be considered limiting the scope thereof.
[0008] Figure 1 is a block diagram illustrating a networking system in accordance with some example embodiments.
[0009] Figure 2 is a diagram illustrating a video modification system in accordance with some example embodiments.
[0010] Figure 3 is a flowchart illustrating an example method for segmenting portions of a video stream and modifying the portions of the video stream based on the segmentation in accordance with some example embodiments.
[0011] Figure 4is a flow diagram illustrating an example method for segmenting portions of a video stream and modifying the portions of the video stream based on the segmentation, in accordance with some example embodiments.
[0012] Figure 5 is a flow diagram illustrating an example method for segmenting portions of a video stream and modifying the portions of the video stream based on the segmentation, in accordance with some example embodiments.
[0013] Figure 6 is a flow diagram illustrating an example method for segmenting portions of a video stream and modifying the portions of the video stream based on the segmentation, in accordance with some example embodiments.
[0014] Figure 7 is a flow diagram illustrating an example method for segmenting portions of a video stream and modifying the portions of the video stream based on the segmentation, in accordance with some example embodiments.
[0015] Figure 8 is a user interface diagram depicting an example mobile device and mobile operating system interface, in accordance with some example embodiments.
[0016] Figure 9 is a block diagram illustrating an example of a software architecture that can be installed on a machine, in accordance with some example embodiments.
[0017] Figure 10 is a block diagram that presents a graphical representation of a machine in the form of a computer system within which a set of instructions, for causing the machine to perform any of the methodologies discussed herein, can be executed, in accordance with example embodiments.
[0018] The headings provided herein are merely for convenience and are not necessarily indicative of the extent or content of the subject matter covered by each section. DETAILED DESCRIPTION
[0019] The description provided herein includes systems, methods, techniques, instruction sequences, and computing machine program products illustrative of embodiments of the disclosure. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide an understanding of various embodiments of the inventive subject matter. It will be evident, however, to those skilled in the art, that embodiments of the inventive subject matter can be practiced without some or all of these specific details. Generally, well-known instruction instances, protocols, structures, and techniques have not been shown in detail.
[0020] While there are telecommunication applications and devices to provide video communication between two devices or video recording operations within a single device, telecommunication devices are generally limited in the modification of video streams. For example, telecommunication applications are generally limited in operations to capture, measure, model, and modify aspects of a video stream while displaying a modified version of the video stream. Methods generally accepted for editing or modifying video do not model or modify the video or video communication while the video is being captured or in the midst of a video communication. Similarly, there are video game applications and devices to provide preprogrammed game environments and interactions within a set of possible simulated environments and predetermined possible interactions. Video game applications and devices are generally limited to preprogrammed interactions and environments and cannot modify a video stream during capture of the video stream and cannot interact within a game environment that includes an unmodified or modified video stream. Thus, there remains a need in the art for improvements in devices, video capture and modification operations, and video communication between devices and video game applications and devices.
[0021] In one embodiment, an application operating on a device includes a component to generate a video game environment and interactions based on a video stream while the video stream is being captured by the device. The application identifies a hand and fingers connected to the hand within a field of view of the video stream. The application detects a direction of the fingers based on the locations identified for the fingers. The hand and fingers within the field of view can be a hand and fingers of a user of the device. The application renders the video stream (e.g., a video stream during a video game session) at a client device and generates a graphical interface element that aligns with the direction of the fingers and replaces at least a portion of the hand and fingers within the video stream. In one embodiment, the graphical interface element aligns with the direction of the fingers such that a target point of the graphical interface element represents the direction of the fingers as determined by the application. The application changes the direction of the fingers and the corresponding target point of the graphical interface element through movement of one or more of the fingers and the hand to enable interaction with the video stream and the video game environment.
[0022] The above is one specific example. Various embodiments of the present disclosure relate to devices and instructions that can be executed by one or more processors of a device to identify an object of interest within a field of view of at least a portion of frames of a video stream and model and modify a depiction of the object of interest depicted within a video stream captured by the device during capture of the video stream. In some embodiments, the video stream is transmitted to another device while the video stream is being captured. In these embodiments, the receiving device can be a video game environment server that transmits data representing game elements for rendering within the video stream and for interaction using graphical interface elements of the modified video stream. In some embodiments, the device capturing the video stream generates and renders game elements within the video stream to create a game environment within the video stream during capture of the video stream.
[0023] A video modification system is described that identifies and tracks objects of interest present in at least a portion of a video stream and through a set of images comprising the video stream. In various example embodiments, the video modification system identifies and tracks a hand and one or more environmental elements depicted in the video stream. Based on user input of finger movement and corresponding target points of graphical interface elements, the video modification system can render game elements that interact with graphical interface elements and one or more environmental elements. The video modification system renders graphical interface elements, game elements, and their interactions, and renders the elements and interactions within the video stream during video stream acquisition. Although described in relation to identifying finger direction and target points of graphical interface elements, it should be understood that the video modification system discussed below can track any object of interest.
[0024] Figure 1 This is a network diagram depicting a network system 100 according to one embodiment, having a client-server architecture configured to exchange data over a network. For example, network system 100 may be a messaging system in which clients transmit and exchange data within network system 100. The data may relate to various functions (e.g., sending and receiving text and media communications, determining geographic location, etc.) and aspects associated with network system 100 and its users (e.g., transmitting communication data, receiving and sending indications of communication sessions, etc.). While network system 100 is shown herein as having a client-server architecture, other embodiments may include other network architectures, such as peer-to-peer or distributed network environments.
[0025] like Figure 1 As shown, network system 100 includes social messaging system 130. Social messaging system 130 is typically based on a three-tier architecture, consisting of an interface layer 124, an application logic layer 126, and a data layer 128. As understood by those skilled in the art of computers and the Internet, Figure 1 Each module, component, or engine shown represents a set of executable software instructions and corresponding hardware (e.g., memory and processor) for executing those instructions, forming a hardware-implemented module, component, or engine, which, when executing instructions, serves as a dedicated machine configured to perform a specific set of functions. To avoid obscuring the subject matter of this invention with unnecessary detail, from... Figure 1 Various functional modules, components, and engines that are not closely related to conveying an understanding of the subject matter of this invention have been omitted. Of course, additional functional modules, components, and engines can be integrated with social messaging systems (such as...) Figure 1 The social messaging system shown can be used together to enable additional functionalities not specifically described herein. Furthermore, Figure 1The various functional modules, components, and engines depicted in FIG. 1 can reside on a single server computer or client device, or can be distributed among several server computers or client devices in various arrangements. Moreover, although Figure 1 The social messaging system 130 is depicted in FIG. 1 as a three-tiered architecture, but the subject matter of the present application is in no way limited to such an architecture.
[0026] As Figure 1 As shown in FIG. 1, the interface layer 124 includes an interface component (e.g., a web server) 140 that receives requests from various client computing devices and servers, such as the client device 110 executing the client application 112, and the third-party server 120 executing the third-party application 122. In response to the received requests, the interface component 140 transmits appropriate responses to the requesting devices via the network 104. For example, the interface component 140 can receive requests, such as hypertext transfer protocol (HTTP) requests or other web-based application programming interface (API) requests.
[0027] The client device 110 can execute a conventional web browser application or an application (also referred to as an "app") that has been developed for a particular platform to include various mobile computing devices and mobile-specific operating systems (e.g., IOS TM , ANDROID TM , PHONE). Moreover, in some example embodiments, the client device 110 forms all or part of the video modification system 160, such that components of the video modification system 160 configure the client device 110 to perform a particular set of functions related to the operation of the video modification system 160.
[0028] In an example, the client device 110 executes the client application 112. The client application 112 can provide functionality to present information to the user 106 and to communicate via the network 104 to exchange information with the social messaging system 130. Moreover, in some examples, the client device 110 performs the functionality of the video modification system 160 to segment images of a video stream during capture of the video stream, modify objects depicted in the video stream, and transmit the video stream in real-time (e.g., based on image data modified based on the segmented images of the video stream).
[0029] Each of the client devices 110 may include a computing device, which includes at least a display and the ability to communicate with the network 104 to access the social messaging system 130, other client devices, and the third-party server 120. Client devices 110 include, but are not limited to, remote devices, workstations, computers, general-purpose computers, internet devices, handheld devices, wireless devices, portable devices, wearable computers, cellular or mobile phones, personal digital assistants (PDAs), smartphones, tablet computers, ultrabooks, netbooks, laptop computers, desktop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, set-top boxes, network PCs, minicomputers, etc. User 106 may be a person, a machine, or other component that interacts with client device 110. In some embodiments, user 106 interacts with the social messaging system 130 via client device 110. User 106 may not be part of network system 100, but may be associated with client device 110.
[0030] like Figure 1 As shown, data layer 128 has a database server 132 that facilitates access to an information repository or database 134. Database 134 is a storage device that stores data such as member profile data, social graph data (e.g., relationships between members of the social messaging system 130), image modification preference data, accessibility data, and other user data.
[0031] Individuals can register using the social messaging system 130 to become members of the social messaging system 130. After registration, members can form social network relationships (e.g., friends, followers, or contacts) on the social messaging system 130 and interact with a wide range of applications offered by the social messaging system 130.
[0032] Application logic layer 126 includes various application logic components 150, which, in conjunction with interface component 140, use data obtained from various data sources or data services in data layer 128 to generate various user interfaces. Each application logic component 150 can be used to implement functionality associated with various applications, services, and features of social messaging system 130. For example, a social messaging application may be implemented using one or more application logic components 150. The social messaging application provides users of client device 110 with a messaging mechanism for sending and receiving messages, including text and media content such as pictures and videos. Client device 110 can access and view messages from the social messaging application for a specified time period (e.g., limited or unlimited). In the example, a message recipient can access a specific message for a predefined duration (e.g., specified by the message sender), the predefined duration starting upon the first access to the specific message. After the predefined duration has elapsed, the message is deleted, and the message recipient can no longer access the message. Of course, other applications and services can be embodied in their own application logic components 150.
[0033] like Figure 1 As shown, the social messaging system 130 may include at least a portion of the video editing system 160, which is capable of identifying, tracking, modeling, and modifying objects within video data during video data acquisition by the client device 110. Similarly, as described above, the client device 110 includes a portion of the video editing system 160. In other examples, the client device 110 may include the entire video editing system 160. Where the client device 110 includes a portion (or all) of the video editing system 160, the client device 110 may operate independently or collaboratively with the social messaging system 130 to provide the functionality of the video editing system 160 described herein.
[0034] In some embodiments, the social messaging system 130 may be a short-lived messaging system capable of enabling brief communication, where content (e.g., video clips or images) is deleted after a deletion trigger event, such as viewing time or viewing completion. In this embodiment, the apparatus uses various components described herein in any aspect of generating, sending, receiving, or displaying short-lived messages. For example, an apparatus implementing the video modification system 160 can identify, track, model, and modify objects of interest, such as a hand depicted in a video stream. The apparatus can modify the object of interest as part of the content generation of a short-lived message during the acquisition of the video stream without requiring image processing after the video stream is acquired.
[0035] In some embodiments, the video modification system 160 can be part of a video game system. The video modification system 160 can identify, track, and modify a focus object within a video stream. In addition to the modified focus object, other portions of the video game system can render interactive objects within the video stream. In some cases, portions of the video game system can implement interactions with the interactive objects using one or more of the following: a pose or movement of the focus object; a movement of the mobile computing device; a selection of a physical control of the mobile computing device; a selection of a virtual or graphical control presented on a display device of the mobile computing device; or a combination thereof to provide a display of the modified video stream, the modified focus object, and the interactive objects in a game environment.
[0036] Figure 2 FIG. 1 is a diagram illustrating a video modification system 160, in accordance with some example embodiments. In various embodiments, the video modification system 160 can be implemented in conjunction with the client device 110, or as a standalone system, and need not be included in the social messaging system 130. The video modification system 160 is shown to include an acquisition component 210, an identification component 220, a direction component 230, a modification component 240, a threshold component 250, and a binarization component 260. All or some of the components 210-260 are in communication with each other, e.g., via a network coupling, shared memory, etc. Each of the components 210-260 can be implemented as a single component, combined into other components, or further subdivided into multiple components. Other components unrelated to the example embodiments can also be included, but are not shown.
[0037] The acquisition component 210 receives or accesses a set of images (e.g., frames) in a video stream. In some embodiments, the acquisition component 210 receives the set of images directly from an image capture device of the client device 110. In some cases, an application or component of the client device 110 passes the set of images to the acquisition component 210 for one or more methods described herein.
[0038] The identification component 220 determines or otherwise identifies features related to a focus object within the set of images accessed by the acquisition component 210 and graphical interface elements to be inserted into the set of images. In some embodiments, the identification component 220 identifies individual pixels associated with the focus object. The identification component 220 can identify the individual pixels based on color values associated with the pixels. In some embodiments, the identification component 220 identifies the focus object or portions of the focus object using one or more image or object recognition processes or by identifying points for generating a convex polygon. In cases where the identification component 220 identifies a convex polygon, the identification component 220 can identify separate portions of the focus object by identifying defects within the convex polygon. The defects can represent points on the focus object that are located a distance from an outline of the convex polygon.
[0039] The direction component 230 determines a direction of a portion of a focus object depicted within one or more frames of a video stream or set of images accessed by the acquisition component 210. For example, where the focus object is a hand, the direction component 230 can determine a direction in which a finger of the hand is pointing. The direction determined by the direction component 230 can include analog or actual three-dimensional values along x, y, and z axes. For example, the direction can represent a target point of the portion of the focus object. The target point can indicate a pixel or region within an image of a set of images and can also indicate a depth of representation within the image. In some embodiments, the direction component 230 can determine the direction as a direction line or vector. The direction line or vector can extend between two points identified along the portion of the focus object for which the direction value is determined. In some cases, the direction component 230 can use three or more points along the portion of the focus object to determine the direction. The three or more points can be connected together to form a chevron for determining the direction of the portion of the focus object.
[0040] The modification component 240 performs modification operations on images within a video stream accessed by the acquisition component 210. In some embodiments, the modification component 240 modifies a depiction of a focus object within an image or frame of the video stream. The modification component 240 can modify the depiction of the focus object by replacing at least a portion of the focus object with a graphical interface element. In some cases, the modification component 240 replaces the focus object entirely or positions the graphical interface element to obscure a portion of the focus object without covering the entirety. For example, where the focus object is a hand and the graphical interface element is a weapon such as a laser blast, the modification component 240 can position the laser blast (e.g., the graphical interface element) such that the hand appears to be holding the laser blast. In some other examples, the graphical interface element can include a representation of a portion of the focus object. For example, the graphical interface element can depict a portion of a hand holding a paint can and a can holder in a position to engage a nozzle. In this example, the graphical interface element can be aligned with at least a portion of the focus object, such as a wrist connected to the hand, and replace or otherwise cover the focus object.
[0041] The threshold component 250 dynamically modifies thresholds set within the video modification system 160 to improve access, interpretation, and modification of frames of a video stream. In some embodiments, the threshold component 250 dynamically modifies histogram thresholds for images within a video stream. The threshold component 250 can be used in quality assurance operations to remove undesirable or unexpected movements, sizes, scales, or other performance characteristics that can inhibit or otherwise adversely affect display, gameplay, or other presentation of modified frames of the video stream.
[0042] The binarization component 260 generates a binary image based on a set of images accessed by the acquisition component 210. In some cases, the binarization component 260 generates one or more binarization matrices for one or more images. The binarization matrices can represent a binary version of one or more images. In some cases, the binarization component 260 performs one or more binarization operations to generate the binarization matrices. Although described below with specific examples of binarization operations, it should be understood that the binarization component 260 can perform any suitable binarization operation to generate the binarization matrices.
[0043] Figure 3 A flowchart depicting an example method 300 for segmenting portions of a video stream and modifying portions of the video stream based on the segmentation (e.g., a representation or depiction of a focus object) is illustrated. The operations of the method 300 can be performed by components of the video modification system 160 and are described below for purposes of illustration.
[0044] In operation 310, the acquisition component 210 receives or otherwise accesses a set of images within a video stream. The set of images can be represented by one or more images depicted within a field of view of an image capture device. In some cases, the acquisition component 210 accesses a video stream captured by an image capture device associated with the client device 110 and presented on the client device 110 as part of hardware that includes the acquisition component 210. In these embodiments, the acquisition component 210 directly receives the video stream captured by the image capture device. In some cases, as described in greater detail below, the acquisition component 210 passes all or a portion of the video stream (e.g., a set of images comprising the video stream) to one or more components of the video modification system 160. The set of images can depict at least a portion of a focus object. In some cases, as will be explained in greater detail below, the focus object can be at least a portion of a hand.
[0045] In some cases, operation 310 occurs in response to the launch of a video game application or a communication application. After launching the application, the application can receive one or more selections of user interface elements presented at a display device of the client device 110. For example, the application can present and cause to be presented on the display device a “start” button. Operation 310 can be launched in response to the application receiving a selection of the “start” button. In some cases, selecting the “start” button launches an image capture device operably connected to the client device 110, such as a camera of a smartphone or tablet computer. After launching the image capture device, the image capture device begins capturing images or frames of a video stream for access by the acquisition component 210.
[0046] In operation 320, the identification component 220 determines pixels within one or more images of a set of images (e.g., a video stream) corresponding to a focus object. In some embodiments, the focus object is a portion of a hand. The identification component 220 can determine the pixels corresponding to the focus object in a predetermined location or region of the field of view of the image capture device. In some example embodiments, operation 320 includes all or part of the process for detecting and generating a binary image of the focus object. The binary image can be generated by the binarization component 260. In the case that the focus object is a hand, the binary image can be a binary skin image that detects skin within the field of view, as represented by the pixels corresponding to the hand, and disregards other elements present within the field of view of the video stream.
[0047] The binary image can be generated within the video game application without being presented. In these embodiments, the video modification system 160 processes the video stream to identify the focus object in real-time as the video stream is received, including identifying the pixels corresponding to the focus object and generating the binary image from the frames within the video stream.
[0048] In operation 330, the identification component 220 identifies a portion of the focus object. The identified portion of the focus object can be a portion of the focus object that extends from another portion of the focus object. In some embodiments, the portion of the focus object is a finger portion of a hand. The identification component 220 can identify the portion of the focus object using edge detection or matching, such as Canny edge detection; feature-based object recognition, such as scale-invariant feature transform (SIFT) or speeded up robust features (SURF); or other suitable object recognition processes, operations, or algorithms. The identification component 220 can identify the portion of the object using one or more object recognition processes on one or more frames (e.g., images) of the video stream, the pixels corresponding to the focus object, and the binary image generated from the frames by the binarization component 260 using the color profile and the pixels corresponding to the focus object.
[0049] In the case that the portion of the focus object is a finger, the identification component 220 identifies a finger position of the portion of the hand. The finger position can be identified based on the pixels corresponding to the portion of the hand or the binary image generated from the frames using the pixels corresponding to the portion of the hand. To identify the finger position, the identification component 220 can form a convex polygon that encloses at least a portion of the portion of the hand.
[0050] In some cases, the identification component 220 generates the convex polygon by generating a contour line that encloses a portion of the hand and at least a portion of the fingers. The identification component 220 identifies a set of vertices within the contour line. In some embodiments, the vertices can be indicated by changes in the direction of the contour line. A change in the direction of the contour line can be identified as the contour line changing direction or the contour line forming a point or a set of points at an angle that exceeds a predetermined angle threshold. The angle threshold can be represented by a vertex along the contour line having an angle greater than 1 degree to 10 degrees, indicating a change from 180 degrees to between 179-170 degrees or less. In some cases, the angle threshold is a point or a set of points representing a vertex in the contour line having an angle greater than 90 degrees.
[0051] In some embodiments, the identification component 220 identifies vertices at edge pixels near or along the contour of the hand indicated by intersections of white and black pixels within the binary image. A first portion of the set of vertices can be identified near or along the contour line. In some cases, one or more vertices (e.g., a second portion of the set of vertices) are identified a distance from the contour line.
[0052] After forming the convex polygon and identifying the set of vertices, the identification component 220 identifies one or more defects within the convex polygon. In some cases, the defects indicate spaces between two fingers on a portion of the hand. The defects can be identified as a second portion of the set of vertices that are a distance from the contour line of the convex polygon. The identification component 220 can determine that a vertex is a defect within the convex polygon by determining an angle between two vertices positioned along the contour line and the vertex being tested as a defect. The angle between the vertex and the defect vertex can be measured as an angle formed between a first line extending between the first vertex positioned on the contour line and the defect vertex and a second line extending between a second vertex positioned on the contour line and the defect vertex.
[0053] The identification component 220 can identify defect vertices where the angle is greater than a defect angle threshold. For example, in some cases, the defect angle threshold can be set to fifty degrees, eighty degrees, ninety degrees, or a value therebetween. In cases where the angle exceeds the defect angle threshold, the identification component 220 can ignore the vertex as a defect. In some cases, in cases where the angle is less than the defect angle threshold, the identification component 220 can determine that the vertex positioned a distance from the contour line of the convex polygon is a point between two fingers on a portion of the hand.
[0054] In some cases, after the identification component 220 identifies one or more defects, the identification component 220 identifies a vertex connection extending between a vertex along the contour line and a vertex located a distance from the contour line (e.g., a vertex that is not identified as a defect). The vertex connection can be represented by a line extending between the two vertices. In some cases, the identification component 220 identifies two vertex connections, a first vertex connection between a vertex along the contour line and a first vertex located a distance from the contour line, and a second vertex connection between a vertex along the contour line and a second vertex located a distance from the contour line and a distance from the first vertex. In these cases, the first vertex connection and the second vertex connection join at the vertex located along the contour line. The combination of the first vertex connection and the second vertex connection can generate a V-shape. The V-shape can extend between a point at or near the tip of the finger and a connection point between adjacent fingers. In cases where the identification component 220 generates a V-shape, the V-shape can be proximate to the location of the finger.
[0055] In embodiments where the video modification system 160 is part of a video game application, the operation 330 can identify a finger extending from a portion of a hand within one or more frames of the video stream. The hand and the finger can be identified and tracked for use as input or control for the video game application. In some cases, the hand is positioned within a predetermined portion of the one or more frames such that positioning the hand in the predetermined portion enables identification of the hand and the finger.
[0056] In operation 340, the direction component 230 determines a direction of the portion of the object of interest. In cases where the portion of the object of interest is a finger, the direction component 230 determines a direction of the finger based on the finger position. The direction component 230 can determine the direction of the portion of the object of interest based at least in part on the position of the object of interest determined in operation 330. In some example embodiments, the direction component 230 identifies a point along or within the portion of the object of interest that is related to the position and orientation of the portion of the object of interest to identify the direction of the portion of the object of interest. The direction component 230 can determine the direction of the portion of the object of interest based at least in part on pixels identified as corresponding to the object of interest. In some cases, the direction of the portion of the object of interest is determined using two or more pixels or points selected from the pixels corresponding to the object of interest.
[0057] In some example embodiments, the direction component 230 determines a direction of a portion of the object of interest by identifying a first point and a second point on the portion of the object of interest. For example, the direction component 230 can identify a tip of a finger that is delineated on a portion of a hand. The tip can represent a first point of the finger. The direction component 230 then identifies a second point on the finger. The second point can be spaced apart from the first point by a distance along the finger that is delineated within the field of view. In some cases, the second point can be selected from pixels that correspond to the finger. The second point can be selected from the pixels of the finger based on a position of the finger that is determined in operation 330. For example, the second point can be selected as a pixel that is located at a middle position (relative to a width of the finger as represented by pixels within the field of view) of the finger. In this way, the first point (located at the tip of the finger) and the second point (located at the middle of the finger relative to the width of the finger) can be used to determine a direction of the finger. In some cases where a first vertex connection and a second vertex connection are determined, the second point can be selected as a pixel that is located between the first vertex connection and the second vertex connection. A distance between the first vertex connection and the second vertex connection can be determined so as to align the second point to a middle position relative to a width of the finger as represented by pixels within the field of view.
[0058] In some cases, the direction component 230 generates a direction line that extends between the first point and the second point. The direction line can be represented as a vector that extends between the two points in the field of view and that indicates a target point within the field of view that is spaced apart from the first point in the direction line (e.g., a vector). The first point and the second point can be associated with relative positions along the finger such that changes in position, direction (e.g., along x and y axes of the display device), and three-dimensional direction (e.g., along a z axis that extends inward and outward from the display device) adjust the target point along the finger in the field of view. For example, in a case where the position of the first point is moved right along the x axis of the display device by a number of pixels and the second point remains relatively fixed, the target point can be adjusted right with little change in position of the target point along the y axis. As another example, in a case where the position of the first point is moved down along the y axis and the position of the second point is moved up along the y axis, the target point can be adjusted along the z axis to indicate a point that is delineated within a background of the field of view that is presented on the display device.
[0059] In cases where the direction component 230 generates a direction for a finger of a hand extending distally from within a video game application, the direction component 230 can identify a direction relative to an x-axis and a y-axis of a display device or frame of a client device. The z-axis corresponds to a depth for the direction determination. In these cases, the direction component 230 can identify a target point for the finger within a frame in a simulated three-dimensional matrix. By determining a direction in the simulated three-dimensional matrix, the direction component 230 enables the finger to move to provide user control over a video game environment depicted in frames of a video stream. For example, movement of the finger between frames of the video stream can cause the direction component 230 to recalculate a direction for the finger. Movement of the finger and modification of the target point can additionally cause the video game application to generate or render interactive elements within the frame, such as targets, monsters, aliens, and other objects with which the user can interact using the target point generated from the finger.
[0060] In some embodiments, the direction component 230 determines a direction of a portion of the object of interest using the first vertex connection and the second vertex connection. In these cases, the direction component 230 uses the locations and relative positions of the first vertex and the second vertex spaced a distance apart from the contour line and the vertices located on the contour line to form a V-shape (e.g., a shape formed by the angled connection between the first vertex connection and the second vertex connection). The V-shape and the associated vertices of the points forming the V-shape can be used to calculate a direction of the finger and a target point. As described above, relative movement of the vertices along the contour line and the first and second vertices spaced a distance apart from the contour line along the x, y, and z-axes can result in adjustments to the target point along the x, y, and z-axes as presented on the field of view by the display device.
[0061] The direction component 230 can perform one or more operations to ensure smooth tracking, translation, and modification of the portion of the object of interest. In some embodiments, the direction component 230 determines a direction of a portion of the object of interest (e.g., a finger of a portion of a hand within a field of view) for each frame of a video stream. Using the determined direction for each frame, the direction component 230 can determine a direction of the portion of the object of interest for each frame of a set of previous frames of one or more images (e.g., frames of the video stream) for a given or current frame within the video stream. In some cases, the direction component 230 can maintain a direction buffer including direction information (e.g., vectors) determined in each frame as new frames of the video stream are received by the acquisition component 210.
[0062] After the direction component 230 has populated the direction buffer with the direction (e.g., vector) information for a predetermined number of frames of a set of previous frames, the direction component 230 combines the direction information for portions of the focus object of at least a portion of the set of previous frames to identify an aggregate direction of the portions of the focus object. For example, where the portions of the focus object are fingers of a portion of a hand depicted in the field of view and the direction buffer includes vector information for the fingers determined for a set of previous frames, the direction component 230 can identify an aggregate direction of the fingers to average direction variations determined due to movement of the fingers. By averaging the direction variations, the direction component 230 can eliminate spurious, jittery, or otherwise erroneous finger directions identified in individual frames of the video stream.
[0063] The predetermined number of frames can be initially set, for example, to three to seven frames within the video stream. In some cases, the number of frames used in the smoothing and tracking operations can be dynamically changed in response to determinations by the direction component 230. For example, as a finger moves, the direction component 230 can determine that the distance of movement of the finger between frames of the video stream is sufficient to cause jitter, spuriousness, or other undesirable signals. The direction component 230 can increase the number of frames used in the smoothing and tracking operations. In cases where the direction component 230 determines that the number of frames is greater than a minimum number of frames used to smooth and track movement of the finger or portion of the focus object and can impact memory or other resource consumption of the client device 110, the direction component 230 can decrease the number of frames over which movement is averaged to reduce resource consumption in the client device 110 for the smoothing and tracking operations.
[0064] In operation 350, the modification component 240 replaces at least a portion of the focus object with a graphical interface element aligned with the direction of the portion of the focus object. The alignment of the graphical interface element with the direction (e.g., finger direction) of the portion of the focus object can be achieved by the modification component 240 positioning a distal end of the graphical interface element proximate to a distal end of the portion of the focus object (e.g., finger). For example, a muzzle of a shockwave (e.g., graphical interface element) can be positioned at or near a first point of a direction line or at or near a vertex on a contour line and aligned such that a target point of the simulation of the shockwave is positioned at or near a target point of the portion of the focus object.
[0065] In cases where the portion of the object of interest is a hand, the modification component 240 replaces at least a portion of the hand or the hand with a graphical interface element that is aligned with the direction of the finger. For example, the graphical interface element can be a representation of a paint can. The modification component 240 can modify the portion of the hand by rendering the paint can in the hand with the nozzle of the paint can aligned with the direction of the finger such that the target point of the paint can is proximate to the target point of the finger. In some embodiments, the modification component 240 can modify the portion of the hand by replacing the hand depicted within the field of view with a graphical representation of a hand holding the paint can.
[0066] In some embodiments, the graphical interface element can be a weapon (e.g., a gun, a laser gun, a shockwave, an energy weapon, a golf club, a spear, a knife). The modification component 240 modifies the portion of the hand by rendering the weapon in the hand with the offensive end of the weapon (e.g., the muzzle of the gun, the point of the blade, the head of the golf club) aligned with the direction of the finger. In some cases, the modification component 240 modifies the portion of the hand by replacing the hand depicted within the field of view with a graphical interface element of the weapon, a portion of the graphical interface element, or a portion of the representation of the hand that depicts at least a portion of the hand.
[0067] In some embodiments, the binarization component 260 generates one or more binary images by isolating pixels corresponding to the object of interest prior to the modification component 240 replacing the portion of the object of interest. The binarization component 260 can isolate the pixels by converting the pixels corresponding to the portion of the object of interest to a first value and converting the remaining pixels within the field of view to a second value.
[0068] Figure 4 A flowchart illustrating an example method 400 for segmenting a portion of a video stream and modifying the portion of the video stream (e.g., a representation of an object of interest or a depiction) based on the segmentation is shown. The operations of method 400 can be performed by components of the video modification system 160. In some cases, certain operations of the method 400 can be performed using one or more operations of the method 300, or as sub-operations of one or more operations of the method 300, as will be explained in more detail below. For example, as shown in FIG. 3, the operations of the method 400 can represent a set of sub-operations of the operation 320. Figure 4 As shown in FIG. 3, the operations of the method 400 can represent a set of sub-operations of the operation 320.
[0069] In operation 410, the identification component 220 samples one or more color values from one or more pixels within the portion of the field of view of the image capture device. In some embodiments, the identification component 220 determines the pixels corresponding to the object of interest (e.g., a portion of a hand) by sampling one or more color values from one or more pixels within the portion of the field of view of the image capture device. The identification component 220 can sample the one or more color values by identifying a subset of pixels located within a predetermined portion of the field of view of the frame of the video stream.
[0070] The identification component 220 can select a predetermined number of pixels for sampling color values. In some embodiments, the identification component 220 selects pixels for sampling color values until an average change in color values between pixels that have been sampled and newly sampled pixels is below a predetermined color change threshold. The identification component 220 can perform one or more operations or sub-operations to sample one or more color values from one or more pixels within the portion of the field of view, as described in more detail below.
[0071] In some embodiments, operation 412 is performed by the identification component 220 selecting a first pixel within the portion of the field of view in operation 410. The identification component 220 can randomly select the first pixel within the portion of the field of view. In some cases, the identification component 220 selects the first pixel at a predetermined location within the portion of the field of view. Although a specified method of selecting the first pixel is presented here, it should be understood that the identification component 220 can use any suitable method to select the first pixel.
[0072] In operation 414, the identification component 220 determines that the first pixel includes a first color value within a predetermined color value range. The desired color range can be selected based on a desired object of interest. For example, where the desired object of interest is a user’s hand, the desired color range includes color values associated with human skin color. In some cases, after selecting the first pixel and determining that the color value is within the desired color range for the object of interest, the identification component 220 refines the color range from a first color range to a second color range. The second color range can be a portion of the first color range. For example, where the first color range includes color values associated with human skin, the second color range can include a subset of color values associated with human skin. The second color range can narrow the expected color range to a range of color values associated with a subset of human skin that is more closely related to the color value of the first pixel.
[0073] In operation 416, the identification component 220 selects a second pixel within the portion of the field of view. The second pixel can be spaced apart from the first pixel by a distance and remain within the portion of the field of view. The second pixel can be selected similarly or identically to the first pixel in operation 412. In some cases, the identification component 220 selects the second pixel based on the location of the first pixel. In these embodiments, the second pixel can be selected a predetermined distance from the first pixel within the portion of the field of view. For example, the second pixel can be selected between one and one hundred pixels from the first pixel, so long as the distance between the first pixel and the second pixel does not place the second pixel outside of the portion of the field of view.
[0074] In operation 418, the identification component 220 determines that the second pixel includes a second color value within a predetermined range of color values. Determining that the second pixel includes a second color value within a range of color values can be performed similarly or identically to determining that the first pixel includes a first color value within a range of color values in operation 414 described above.
[0075] In operation 420, the identification component 220 compares the first color value and the second color value to determine that the second color value is within a color threshold of the first color value. In some embodiments, the color threshold is a value that places both the first color value and the second color value within a predetermined range of color values that the identification component 220 determines the first color value is included within. In these cases, the color threshold can be dynamically determined based on the first color value and the predetermined range of color values such that the second color value is acceptable if it falls within the predetermined range of color values (e.g., a predetermined threshold). In some example embodiments, the color threshold can be a predetermined color threshold. In these cases, the second color value can be discarded and a second color value reselected in the event that the second color value falls outside of the predetermined color threshold, despite still being within the predetermined color range.
[0076] In operation 430, the identification component 220 determines a color profile for the hand (e.g., the object of interest) based on the one or more color values (e.g., the first color value and the second color value) sampled from the one or more pixels (e.g., the first pixel and the second pixel). The identification component 220 includes the first color value and the second color value in the color profile for the hand based on the second color value being within a predetermined color threshold. The color profile represents an intermediate color value for the hand. The intermediate color value can be a single color value, such as a midpoint of the color values sampled by the identification component 220, or a range of color values, such as a range of color values that includes the color values sampled by the identification component 220.
[0077] In some embodiments, after the identification component 220 samples one or more colors and determines a color profile, the identification component 220 can identify pixels within the predetermined portion of the field of view that have color values associated with the color profile. The identification component 220 can determine that a pixel has a color value associated with the color profile by identifying that the color value of the pixel is within a predetermined range of the intermediate color value (e.g., a single intermediate color value) or within a range of color values of the color profile. The identification component 220 can identify pixels associated with the color profile in the predetermined portion of the field of view for each frame of the video stream.
[0078] In some cases, in response to identifying the pixels associated with the color profile, the binarization component 260 extracts the object of interest by computing a binary image of the object of interest. In the binary image computed by the binarization component 260, the pixels associated with the color profile can be assigned a first value and the remaining pixels within the field of view can be assigned a second value. For example, the first value can be 1, indicating a white pixel, and the second value can be 0, indicating a black pixel. The binarization component 260 can also employ a nonlinear median blur filter to filter the binary image of each frame to remove artifacts, inclusions, or other false pixel conversions.
[0079] In some embodiments, the identification component 220 determines the color profile by generating a histogram from one or more color values sampled within the portion of the field of view. The histogram can represent a distribution of pixels having a specified color value among the pixels sampled by the identification component 220. The histogram can be a color histogram generated in a three-dimensional color space such as red, green, blue (RGB); hue, lightness, saturation (HLS); and hue, saturation, value (HSV), among others. The histogram can be generated using any suitable frequency identification operation or algorithm capable of identifying a frequency of a color value appearing in the pixels from which the color is sampled and selected.
[0080] In some cases, the histogram can be a two-dimensional histogram. The two-dimensional histogram can identify a combination of intensity or value and a number of pixels having the identified combination of intensity or value. In some example embodiments, the histogram generated as a two-dimensional histogram identifies a combination of hue values and color saturation values of the pixels having the specified color values identified during color sampling. The hue and color saturation values can be extracted from an HSV color space.
[0081] In response to generating the histogram, the identification component 220 removes one or more bins of the histogram below a predetermined pixel threshold. A bin of the histogram can be understood as a set of partitions within the histogram. Each bin can represent a color value among the color values sampled from the selected pixels or the pixels within the portion of the field of view. The bin can indicate a value associated with or depicting a number of pixels having the specified color value for the bin. In embodiments in which the histogram is a two-dimensional histogram, a bin of the histogram indicates hue and color saturation values and a number of pixels having the hue and color saturation values. In these embodiments, the bin having the greatest number of pixels associated with the hue and color saturation values of the color profile is associated with the object of interest (e.g., the hand).
[0082] The predetermined pixel threshold is a threshold used to estimate pixels associated with or not associated with the object of interest. The predetermined pixel threshold can be applied to bins of the histogram such that bins having a number of pixels above the predetermined pixel threshold are estimated or determined to be associated with the object of interest. Bins having a number of pixels above the predetermined pixel threshold can be associated with color values within a desired color range for the object of interest identification. The predetermined pixel threshold can be a percentage value, such as a percentage of a total number of pixels within the field of view portion included in the specified bin. For example, the predetermined pixel threshold can be five percent, ten percent, or fifteen percent of the total pixels within the field of view portion. In this example, bins containing less than five percent, ten percent, or fifteen percent of the total pixels, respectively, can be determined to contain pixels not associated with the object of interest. In some cases, the predetermined pixel threshold is a numerical value of pixels that occurs within a bin. For example, the predetermined pixel threshold can be between 1,000 and 200,000 pixels. In this example, bins having less than 1,000 pixels or 200,000 pixels, respectively, can be determined to contain pixels not associated with the object of interest.
[0083] The identification component 220 includes color values in the color profile associated with bins having a number or percentage of pixels above the predetermined pixel threshold. In some cases, the binarization component 260 generates a binary image, the binarization component 260 converts color values of pixels in bins exceeding the predetermined pixel threshold to a first value, indicating white pixels representing a portion of the object of interest. The binarization component 260 can convert color values of pixels in bins below the predetermined pixel threshold to a second value, indicating black pixels not associated with the object of interest.
[0084] Figure 5 A flowchart illustrating an example method 500 for segmenting portions of a video stream and modifying portions of the video stream based on the segmentation (e.g., representations or depictions of the object of interest) is shown. Operations of the method 500 can be performed by components of the video modification system 160. In some cases, certain operations of the method 500 can be performed using one or more operations of the methods 300 or 400, or as sub-operations of one or more operations of the methods 300 or 400, as will be explained in more detail below. For example, as shown in FIG. 3, operations of the method 400 can represent a set of sub-operations of the operation 320. Figure 5 As shown in FIG. 3, operations of the method 400 can represent a set of sub-operations of the operation 320.
[0085] In operation 510, the identification component 220 determines that a set of artifacts within the first set of pixels exceeds an artifact threshold. In some embodiments, the first set of pixels is determined by the identification component 220 in operation 320 as described above. The first set of pixels can correspond to a portion of the object of interest (e.g., a finger on a portion of a hand). In some cases, the first set of pixels can correspond to the entire object of interest identified by the identification component 220.
[0086] In some cases, the artifact threshold is a number of pixels or pixel regions within a convex polygon that have a different value than the pixels or pixel regions surrounding the artifact. For example, in cases where a convex polygon is identified and a binary image is generated, the artifact can be a pixel or pixel region within the convex polygon that has a value of zero (e.g., a black pixel) and is depicted on the object of interest.
[0087] Although described as an artifact within the first set of pixels, the artifact can be identified outside of the first set of pixels. In some cases, the artifact can be a pixel or pixel region outside of the convex polygon that has a value substantially different from the values of the surrounding pixels or pixel regions. For example, the artifact outside of the convex polygon can be a pixel or pixel region having a value of one (e.g., a white pixel) outside of one or more of the convex polygon or the object of interest.
[0088] To determine that a set of artifacts exceeds the artifact threshold, the identification component 220 can identify a total number of pixels included in the set of artifacts and determine that the total pixel count exceeds the total pixel count of the artifact threshold. In some cases, the identification component 220 identifies a number of artifacts (e.g., discrete groupings of pixels having similar values surrounded by pixels having substantially different values) and determines that the number of identified artifacts exceeds a total number of artifacts as calculated by the artifact threshold.
[0089] In operation 520, the threshold component 250 dynamically modifies the histogram threshold to identify pixels corresponding to the portion of the object of interest. In some embodiments, the threshold component 250 dynamically modifies the histogram threshold based on an orientation of the portion of the object of interest (e.g., a finger). In some embodiments, the histogram threshold is a predetermined color threshold. Adjustment of the histogram threshold can increase or decrease the predetermined color threshold to include a greater or lesser number of color values within the color threshold and the histogram threshold. In some cases, the histogram threshold is associated with a predetermined pixel threshold. Modification of the histogram threshold can increase or decrease the predetermined pixel threshold to include color values associated with a greater or lesser number of pixels as part of the object of interest or the color profile.
[0090] In response to determining that the artifact exceeds the artifact threshold, the threshold component 250 or the identification component 220 can determine one or more color values of pixels adjacent or proximate to the one or more artifacts. In some embodiments, the threshold component 250 or the identification component 220 determines a location of the one or more color values within the histogram. The threshold component 250 or the identification component 220 can also determine a location of the one or more color values relative to a predetermined color range. Based on the location of the one or more color values, the threshold component 250 modifies the histogram threshold to include additional color values associated with the one or more colors of the adjacent or proximate pixels. For example, in instances where the identification component 220 or the threshold component 250 identifies one or more color values of one or more pixels adjacent or proximate to an artifact located at a low end of the color threshold or the pixel threshold in the histogram, the threshold component 250 can increase the low end of the color threshold or the pixel threshold to include one or more colors or one or more pixels previously excluded from the low end of the color threshold or the pixel threshold.
[0091] In operation 530, the identification component 220 determines a second set of pixels within the one or more images corresponding to the object of interest (e.g., a hand or a portion of a hand) depicted within the field of view. The second set of pixels is determined based on the modified histogram threshold. In some embodiments, the second set of pixels includes at least a portion of the first set of pixels. Operation 530 can be performed similarly or identically to operation 320 described above with reference to FIG. 3. Figure 3 Operation 530 can be performed similarly or identically to operation 320 described above with reference to FIG. 3.
[0092] Figure 6 A flowchart illustrating an example method 600 for segmenting portions of a video stream and modifying portions of the video stream based on the segmentation (e.g., a representation or depiction of an object of interest) is shown. The operations of method 600 can be performed by components of the video modification system 160. In some instances, certain operations of the method 600 can be performed using one or more operations of the methods 300, 400, or 500, or as sub-operations of one or more operations of the methods 300, 400, or 500, as will be explained in more detail below. For example, the operations of the method 600 can represent a set of operations performed in response to performance of operation 340.
[0093] In operation 610, the direction component 230 determines a current direction of a portion of the object of interest (e.g., a finger depicted on a portion of a hand) from a designated corner of the field of view. The designated corner can be a predetermined corner, edge, set of pixels, pixel, coordinate, or other portion of the field of view. For example, in some embodiments, the designated corner can be presented as a lower left corner of the field of view as presented on a display device of the client device 110. In some embodiments, operation 610 is performed similarly or identically to operation 340.
[0094] In operation 620, the direction component 230 identifies a combined direction of the portion of the attention object for one or more images (e.g., a set of previous frames of a video stream) to indicate a previous combined direction. In some embodiments, the direction component 230 identifies a combined direction of the portion of the attention object for three or more previous frames of the video stream. The direction component 230 identifies one or more directions of the portion of the attention object determined for each of the one or more images. The direction component 230 can average the one or more directions of the one or more images. In some cases, the direction component 230 averages the one or more directions by generating an average of one or more vectors representing the one or more directions. The direction component 230 can generate a weighted moving average of the one or more directions, where two or more directions or vectors are close to each other. In some embodiments, the weighted moving average places greater weight on directions determined for frames immediately preceding a current frame and the currently identified direction.
[0095] In operation 630, the direction component 230 determines that the change in position between the current direction and the previous combined direction exceeds a position threshold. The direction component 230 can determine that the change in position exceeds the position threshold by comparing the combined direction and the current direction to determine the change in position. The change in position can be measured in pixels, degrees, radians, or any other suitable measure extending between two directions, vectors, or points along a direction line. In response to measuring the change in position, the direction component 230 can compare the measurement to the position threshold. The position threshold can be a value in radians, degrees, pixels, or other threshold. In cases where the change in position is below the position threshold, the direction component 230 can ignore the combined direction of the previous frames.
[0096] In operation 640, the direction component 230 selects from a set of first direction identification operations and a set of second direction identification operations based on the change in position exceeding the position threshold. The set of first direction identification operations includes one or more threshold modification operations. The set of second direction identification operations includes determining a direction from the current frame and two or more of the previous frames.
[0097] In some embodiments, the first direction identification operations modify one or more of a histogram threshold, a color profile, or a pixel threshold. The direction component 230, alone or in combination with one or more of the threshold component 250 and the identification component 220, can modify the one or more of the histogram threshold, the color profile, or the pixel threshold in the same manner as described above with respect to the direction component 230 of FIG. 2. Figures 3-5The histogram threshold, color profile, or pixel threshold are modified in a similar manner as described. In some cases, in response to determining that the change in position exceeds the position threshold, the direction component 230 can identify color values of pixels associated with points on a direction line in the current frame and in one or more previous frames. The direction component 230 can also identify color values of pixels or clusters of pixels adjacent or nearest to the points on the direction line. In cases where the direction component 230 determines that the color values of the pixels adjacent or nearest to the points on the direction line and the color profile or color values within the histogram associated with a number of pixels above the pixel threshold are not correlated, the direction component 230 can disregard the direction of the current frame. In these cases, the direction component 230 can revert to the direction in a previous frame. The direction component 230 can also recalculate the direction of the direction line by identifying points on the portion of the object of interest and generating a new direction line or vector. In some embodiments, the direction component 230, the threshold component 250, or the identification component 220 can modify one or more of the histogram threshold, the color profile, or the pixel threshold to remove false positive or false positive pixel identifications outside of one or more of the object of interest and the convex polygon. For example, the modification of one or more of the histogram threshold, the color profile, or the pixel threshold can cause the resulting binary image to exhibit fewer false positives (e.g., values of 1 indicating white pixels) with the first pixel value, where the frame is captured in a dark room or using a suboptimal International Organization for Standardization (ISO) speed or aperture size (e.g., focal length ratio, f-ratio, or f-stop).
[0098] In some cases, a set of second direction identification operations cause the direction component 230, alone or in combination with the identification component 220 or the threshold component 250, to modify a previous frame threshold. The previous frame threshold represents a number of previous frames of the video stream used to calculate an average direction or a weighted moving average of the direction of the portion of the object of interest. In some embodiments, the direction component 230 initially calculates the average direction using directions calculated for three frames immediately preceding the current frame of the video stream. In cases where the change in position exceeds the position threshold, the direction component 230 modifies the previous frame threshold to include one or more additional frames and recalculates the average direction. In cases where the modification of the previous frame threshold causes the change in position to fall within the position threshold, the direction component 230 can continue to identify changes in direction in the additional frames using the modified previous frame threshold.
[0099] In some embodiments, the direction component 230 identifies a change in direction between a set of previous frames of the video stream that indicates motion of the object of interest or the portion of the object of interest above a predetermined motion threshold. In these cases, the direction component 230 terminates or discontinues the calculation of the average direction and the use of the previous frame threshold. In response to discontinuing the use of the previous frame threshold or determining that the motion exceeds the motion threshold, the direction component 230 determines a change in the angle of the direction line or vector and a change in the location of the direction line or vector, or one or more points along the direction line or vector. In response to the change in the location and the change in the angle of the direction line or vector exceeding a modified location threshold or an angle threshold, the direction component 230 can modify or discard the change in the location or the change in the angle. In cases where the direction component 230 modifies the change in the location or the change in the angle, the direction component 230 can average the change in the location or the change in the angle with the location or the angle in one or more previous frames and use the average location or average angle as the location or angle of the current frame. In cases where the direction component 230 discards the change in the location or the change in the angle, the direction component 230 can replace the location or the angle of the previous frame with the location or the angle of the current frame.
[0100] Figure 7 A flow diagram illustrating an example method 700 for segmenting portions of a video stream and modifying portions of the video stream based on the segmentation (e.g., a representation or depiction of an object of interest) is shown. The operations of the method 700 can be performed by the components of the video modification system 160. In some cases, certain operations of the method 700 can be performed using one or more operations of the methods 300, 400, 500, or 600, or as sub-operations of one or more operations of the methods 300, 400, 500, or 600, as will be explained in more detail below. For example, the operations of the method 700 can represent a set of operations performed in response to the performance of one or more of the operations 320, 330, or 340.
[0101] In operation 710, the direction component 230 determines a direction of the portion of the object of interest. In cases where the portion of the object of interest is a finger, the direction component 230 determines the direction of the finger based on the finger location. The direction of the finger can be represented as a vector extending between two pixels corresponding to a portion of the hand (e.g., along the finger location). Operation 710 can be performed similarly or the same as operation 340 described above with respect to FIG. 4. Figure 3 Operation 710 can be performed similarly or the same as operation 340 described above with respect to FIG. 4.
[0102] In operation 720, the threshold component 250 dynamically modifies the histogram threshold. The threshold component 250 can modify the histogram threshold in response to determining a direction of the portion of the object of interest (e.g., a finger) and determining an error, artifact, or unexpected motion of the object of interest. In some embodiments, operation 720 is performed similarly or the same as operation 520.
[0103] In operation 730, the threshold component 250 identifies a first performance characteristic of the vector. The first performance characteristic can include a presence of a jitter at the object of interest or the direction line (e.g., the vector), a presence of an artifact on the object of interest, an alignment of the graphical interface element to the object of interest, a coverage of the graphical interface element to the object of interest, a location of the convex polygon, a contour size of the convex polygon, a contour proportion of the convex polygon, and other suitable performance characteristics. The artifact can be similar to the artifacts described above and appear within the depiction of the object of interest.
[0104] The presence of the jitter can be represented as a movement from a first position of one or more points or edges of the object of interest to a second position between frames and a subsequent movement from the second position to the first position in a subsequent frame that is temporally proximate to the initial movement. For example, when the video stream is depicted as a binary image with the object of interest represented using white pixels, the jitter in the rapid, unstable, or irregular movement within the video stream can be seen, often repeating one or more times before the subsequent movement.
[0105] The alignment of the graphical interface element to the object of interest can be determined based on a direction line or vector for the finger (e.g., a portion of the object of interest) and a direction line or vector generated for the graphical interface element. As described above, the direction line or vector for the graphical interface element can be generated similarly or identically to the direction line or vector for the finger. For example, the direction line or vector for the graphical interface element can be generated by selecting two or more points or pixels depicted on the graphical interface element, such as a point proximate to a distal end of the graphical interface element (e.g., a muzzle of a shockwave) and a point spaced a distance apart from the distal end and toward a proximal end of the graphical interface element. The angles of the direction lines of the finger and the graphical interface element relative to a reference point can be compared to generate a value of the performance characteristic. For example, the direction component 230 can determine a percentage difference between the angles of the direction line of the graphical interface element and the direction line of the finger and use the percentage difference as the performance characteristic value.
[0106] The coverage of the graphical interface element to the object of interest can be determined by identifying two or more boundaries of the graphical interface element. The identification component 220 can identify the boundaries of the graphical interface element as one or more edges of the graphical interface element or one or more edges of an image depicting the graphical interface element. The identification component 220 can determine an amount of the object of interest that is not covered by the graphical interface element based on portions of the object of interest extending outward from the two or more boundaries of the graphical interface element. The coverage performance characteristic can be assigned a value such as a percentage of the object of interest that is covered by the graphical interface element.
[0107] The position of the convex polygon can be identified by the identifying component 220 determining a boundary of a predetermined portion of a frame. In some cases, the boundary is a rectangular portion of each frame of the video stream. The predetermined portion of the frame can be located in a lower left corner or side of the frame. The identifying component 220 can identify a percentage, amount, or number of points on the contour line, or a portion of the contour line that overlaps or extends outward from the predetermined portion of the frame. The polygon position performance feature can include a value that indicates whether or how much the convex polygon is in the predetermined portion of the frame.
[0108] The contour size of the convex polygon can be identified with respect to the predetermined portion of the frame. In some embodiments, the identifying component 220 determines a size of the convex polygon. The size of the convex polygon can include an area (e.g., a pixel area) occupied by the convex polygon, a shape or set of edges of the convex polygon, or any other suitable size of the convex polygon. The area of the convex polygon can be determined by a number of pixels enclosed within the convex polygon. In some embodiments, the contour size of the convex polygon is determined as a percentage of the predetermined portion of the frame. The performance feature value associated with the contour size can be an area value, a percentage of the predetermined portion of the frame, or any other suitable value that describes the contour size.
[0109] The contour ratio of the convex polygon can be identified as an expected proportional ratio of a portion of the convex polygon with respect to other portions of the convex polygon. For example, in a case where the object of interest defined by the convex polygon is a hand and the portion of the object of interest used to determine a direction of the object of interest is a finger, the contour ratio can fall in an expected proportional ratio of the finger to the hand. The identifying component 220 can identify the portion of the object of interest and generate a dividing line that divides the portion of the object of interest from the remaining portion of the object of interest. The identifying component 220 can then compare an area or other size measurement of the portion of the object of interest to the remaining portion of the object of interest to generate the contour ratio.
[0110] In operation 740, the threshold component 250 determines that the performance feature exceeds the feature threshold. The feature threshold (which the threshold component 250 compares to the performance feature value) can be specific to the performance feature. For example, in a case where the performance feature value is a value of a contour size of a convex polygon, the feature threshold can be a maximum expected area value or a maximum percentage value of an area of the predetermined portion of the frame that is occupied by the convex polygon. In some cases, the threshold component 250 determines that the performance feature exceeds the feature threshold by identifying which value, performance feature, or feature threshold is greater. For example, in a case where the performance feature is a graphical interface element covering the object of interest, the threshold component 250 can determine that eighty-five percent of the object of interest is covered by the graphical interface element, leaving fifteen percent uncovered. In a case where the feature threshold is an uncovered area of five percent of the object of interest, the fifteen percent of uncovered area exceeds the feature threshold.
[0111] In operation 750, the threshold component 250 modifies the histogram threshold based on the performance characteristic. The histogram threshold can be modified in a similar or identical manner as described in operation 520. The modification of the histogram threshold can include or exclude pixels contained in the object of interest or the convex polygon. For example, in cases where the convex polygon has a contour size that occupies a larger portion of the predetermined portion of the frame than the desired characteristic threshold, the histogram threshold can be increased to reduce the identified size or dimension of the convex polygon.
[0112] In operation 760, the threshold component 250 identifies that the second performance characteristic of the vector is within the characteristic threshold. The second performance characteristic can be the same performance characteristic as identified in operation 730, where the performance characteristic value is modified based on the modification of the histogram threshold performed in operation 750. After the threshold component 250 determines the performance characteristic is within the characteristic threshold, the video modification system 160 can proceed to operation 350. In some embodiments, after the video modification system 160 completes operation 760 or operation 350, the video modification system 160 can proceed to modify the next frame within the video stream.
[0113] Examples
[0114] To better illustrate the devices and methods disclosed herein, a non-limiting list of examples is provided herein:
[0115] 1. A method comprising: receiving, by one or more processors, one or more images depicting at least a portion of a hand; determining, in a predetermined portion of a field of view of an image capture device, pixels within the one or more images corresponding to the portion of the hand, the portion of the hand having a finger; identifying, based on the pixels corresponding to the portion of the hand, a finger position of the finger; determining, based on the finger position, a direction of the finger; dynamically modifying, based on the direction of the finger, a histogram threshold to identify the pixels corresponding to the portion of the hand; and replacing the portion of the hand and the finger with a graphical interface element aligned with the direction of the finger.
[0116] 2. The method of example 1, wherein identifying the finger position further comprises: forming a convex polygon that encloses at least a portion of the portion of the hand; and identifying one or more defects within the convex polygon, the defects indicating spaces between two fingers positioned on the portion of the hand.
[0117] 3. The method of example 1 or 2, wherein determining the direction of the finger further comprises: identifying a tip of the finger, the tip representing a first point of the finger; identifying a second point on the finger, the second point spaced apart from the first point along the finger; generating a directional line extending between the first point and the second point; and determining a direction of the directional line within the one or more images.
[0118] 4. The method of any one or more of examples 1-3, wherein determining the pixels corresponding to the portion of the hand further comprises: sampling one or more color values from one or more pixels within the portion of the field of view of the image capture device; and determining a color profile for the hand based on the one or more color values sampled from the one or more pixels, the color profile representing an intermediate color value of the hand.
[0119] 5. The method of any one or more of examples 1-4, wherein sampling the one or more color values further comprises: selecting a first pixel within the portion of the field of view; determining that the first pixel includes a first color value within a predetermined range of color values; selecting a second pixel within the portion of the field of view, the second pixel being spaced apart from the first pixel and located within the portion of the field of view; determining that the second pixel includes a second color value within the predetermined range of color values; comparing the first color value and the second color value to determine that the second color value is within a predetermined threshold of the first color value; and including the first color value and the second color value in the color profile for the hand.
[0120] 6. The method of any one or more of examples 1-5, wherein determining the color profile further comprises: generating a histogram from the one or more color values sampled from the portion of the field of view; removing one or more bins of the histogram associated with a number of pixels below a predetermined pixel threshold; and including one or more bins of the histogram associated with a number of pixels above the predetermined pixel threshold within the color profile.
[0121] 7. The method of any one or more of examples 1-6, wherein the method further comprises: generating one or more binary images by isolating the pixels corresponding to the portion of the hand by converting the pixels corresponding to the portion of the hand to a first value and converting remaining pixels within the field of view to a second value.
[0122] 8. The method of any one or more of examples 1-7, wherein determining the direction of the digit further comprises: determining a direction of the digit for each frame of a set of previous frames for the one or more images; and combining the directions of the digit for the set of previous frames to identify an aggregated direction of the digit.
[0123] 9. The method of any one or more of examples 1-8, wherein determining the direction of the digit further comprises: determining a current direction of the digit from a specified angle of the field of view; identifying a combined direction of the digit for the one or more images to indicate a previous combined direction; determining that a change in position between the current direction and the previous combined direction exceeds a position threshold; and selecting from a set of first direction identification operations and a set of second direction identification operations based on the change in position exceeding the position threshold.
[0124] 10. The method of any one or more of examples 1-9, wherein the direction of the finger is represented by a vector extending between two pixels corresponding to the portion of the hand, and dynamically modifying the histogram threshold further comprises: identifying a first performance characteristic of the vector; determining that the first performance characteristic exceeds a characteristic threshold; modifying the histogram threshold based on the first performance characteristic; and identifying a second performance characteristic of the vector within the characteristic threshold based on the modification of the histogram threshold.
[0125] 11. The method of any one or more of examples 1-10, wherein determining the pixels within the one or more images corresponding to the portion of the hand identify a first set of pixels, and further comprising: determining that a set of artifacts within the first set of pixels exceeds an artifact threshold; and determining a second set of pixels within the one or more images corresponding to the portion of the hand based on modifying the histogram threshold, the second set of pixels comprising at least a portion of the first set of pixels.
[0126] 12. A system comprising: one or more processors; and processor-readable storage storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: receiving, by the one or more processors, one or more images depicting at least a portion of a hand; determining, within a predetermined portion of a field of view of an image capture device, pixels within the one or more images corresponding to the portion of the hand, the portion of the hand having a finger; identifying, based on the pixels corresponding to the portion of the hand, a finger position of the finger; determining, based on the finger position, a direction of the finger; dynamically modifying, based on the direction of the finger, a histogram threshold to identify the pixels corresponding to the portion of the hand; and replacing the portion of the hand and the finger with a graphical interface element aligned with the direction of the finger.
[0127] 13. The system of example 12, wherein determining the direction of the finger further comprises: identifying a tip of the finger, the tip representing a first point of the finger; identifying a second point on the finger, the second point spaced apart from the first point along the finger; generating a directional line extending between the first point and the second point; and determining a direction of the directional line within the one or more images.
[0128] 14. The system of example 12 or 13, wherein determining the pixels corresponding to the portion of the hand further comprises: sampling one or more color values from the one or more pixels within the portion of the field of view of the image capture device; and determining, based on the one or more color values sampled from the one or more pixels, a color profile for the hand, the color profile representing an intermediate color value of the hand.
[0129] 15. The system of any one or more of examples 12-14, wherein determining the color profile further comprises: generating a histogram from the one or more color values sampled from within the portion of the field of view; removing one or more bins of the histogram associated with a number of pixels below a predetermined pixel threshold; and including one or more bins of the histogram associated with a number of pixels above the predetermined pixel threshold within the color profile.
[0130] 16. The system of any one or more of examples 12-15, wherein determining the direction of the finger further comprises: determining a current direction of the finger from a specified angle of the field of view; identifying a combined direction of the finger for the one or more images to indicate a previous combined direction; determining a change in position between the current direction and the previous combined direction exceeds a position threshold; and selecting from a first set of direction identification operations and a second set of direction identification operations based on the change in position exceeding the position threshold.
[0131] 17. The system of any one or more of examples 12-16, wherein the direction of the finger is represented by a vector extending between two pixels corresponding to the portion of the hand, and dynamically modifying the histogram threshold further comprises: identifying a first performance characteristic of the vector; determining the first performance characteristic exceeds a characteristic threshold; modifying the histogram threshold based on the first performance characteristic; and identifying a second performance characteristic of the vector is within the characteristic threshold based on the modification of the histogram threshold.
[0132] 18. A processor-readable storage device storing processor-executable instructions that, when executed by a processor of a machine, cause the machine to perform operations comprising: receiving one or more images depicting at least a portion of a hand; determining, within a predetermined portion of a field of view of an image capture device, pixels within the one or more images corresponding to a portion of the hand, the portion of the hand having a finger; identifying a finger position of the finger based on the pixels corresponding to the portion of the hand; determining a direction of the finger based on the finger position; dynamically modifying a histogram threshold to identify the pixels corresponding to the portion of the hand based on the direction of the finger; and replacing the portion of the hand and the finger with a graphical interface element aligned with the direction of the finger.
[0133] 19. The processor-readable storage device of example 18, wherein determining the direction of the finger further comprises: identifying a tip of the finger, the tip representing a first point for the finger; identifying a second point on the finger, the second point spaced apart from the first point along the finger; generating a directional line extending between the first point and the second point; and determining a direction of the directional line within the one or more images.
[0134] 20. The processor-readable storage device of example 18 or 19, wherein determining the direction of the finger further comprises: determining a current direction of the finger from a specified angle of the field of view; identifying a combined direction of the finger for the one or more images to indicate a previous combined direction; determining a change in position between the current direction and the previous combined direction exceeds a position threshold; and selecting from a first set of direction identification operations and a second set of direction identification operations based on the change in position exceeding the position threshold.
[0135] 21. A machine-readable medium carrying processor-executable instructions, which when executed by one or more processors of a machine, cause the machine to perform the method of any of examples 1-11.
[0136] These and other examples and features of the present devices and methods are set forth in part in the DETAILED DESCRIPTION. The SUMMARY and examples are intended to provide non-limiting illustrations of the subject matter. They are not intended to provide an exhaustive explanation of the subject matter. The DETAILED DESCRIPTION is included to provide further information about the subject matter.
[0137] Modules, Components, and Logic
[0138] Certain examples are described herein as including logic or a number of components, modules, or mechanisms. Modules and components can constitute either software modules (e.g., code, instructions, instruction sets, or other software) or hardware modules or components (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs) or both, or one or more hardware processors configured with software modules) configured to perform the various processes described herein. In some embodiments, software modules are stored in memory. Various software modules can perform one or more tasks in the example embodiments. In some embodiments, software modules are written in C, C++, Java, Python, Ruby, Perl, JavaScript, HTML5, or other suitable software languages or programming or scripting languages. In some embodiments, software modules are written in multiple languages.
[0139] In some embodiments, a hardware module or hardware component is implemented in mechanical, electronic, or any suitable combination thereof. For example, a hardware module or hardware component can include dedicated circuitry or logic that is permanently configured to perform certain operations. For example, a hardware module or hardware component can be a special-purpose processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). A hardware module or hardware component can also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. For example, a hardware module or hardware component can include software encompassed within a general-purpose processor or other programmable processor. It will be appreciated that the decision to implement a hardware module or hardware component mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) can be driven by cost and time considerations.
[0140] Accordingly, the phrase "hardware module" or "hardware component" should be understood to encompass a tangible entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner and / or to perform certain operations described herein. As used herein, "hardware-implemented module" or "hardware-implemented component" refers to a hardware module or hardware component, respectively. Considering embodiments in which hardware modules or hardware components are temporarily configured (e.g., programmed), each of the hardware modules or hardware components need not be configured or instantiated at any one instance in time. For example, where a hardware module or hardware component comprises a general-purpose processor configured by software to become a special-purpose processor, the general-purpose processor can be configured as
[0141] Hardware modules or hardware components can provide information to, and receive information from, other hardware modules or hardware components. Accordingly, the described hardware modules or hardware components can be regarded as being communicatively coupled. Where multiple hardware modules or hardware components exist contemporaneously, communications can be achieved through signal transmission (e.g., over appropriate circuits and buses) between or among two or more of the hardware modules or hardware components. In embodiments in which multiple hardware modules or hardware components are configured or instantiated at different times, communications between such hardware modules or hardware components can be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware modules or hardware components have access. For example, one hardware module or hardware component performs an operation and stores the output of that operation in a memory device to which it is communicatively coupled. A further hardware module or hardware component can then access the memory device to retrieve and process the stored output. Hardware modules or hardware components can also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information).
[0142] The various operations of example methods described herein can be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors constitute processor-implemented modules or components that operate to perform one or more operations or functions described herein. As used herein, "processor-implemented module" or "processor-implemented component" refers to a hardware module or hardware component, using a processor.
[0143] Similarly, the methods described herein can be at least partially processor-implemented, with a particular processor or processors being an example of hardware. For example, at least some of the operations of a method can be performed by one or more processors or processor-implemented modules. Moreover, the processor(s) can also operate to support performance of the relevant operations in a "cloud computing" environment or as a "software as a service" (SaaS). For example, at least some of the operations can be performed by a group of computers (as examples of machines including processors), with these operations being accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., an API).
[0144] The performance of certain of the operations can be distributed among the processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the processors or processor-implemented modules can be located in a single geographic location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, the processors or processor-implemented modules can be distributed across a number of geographic locations.
[0145] Application
[0146] Figure 8 An example mobile device 800 executing a mobile operating system (e.g., IOS TM , ANDROID TM , Phone, or other mobile operating system) consistent with some embodiments is shown. In one embodiment, the mobile device 800 includes a touch screen operable to receive tactile data from a user 802. For example, the user 802 can physically touch 804 the mobile device 800, and in response to the touch 804, the mobile device 800 can determine tactile data such as a touch location, a touch force, or a gesture action. In various example embodiments, the mobile device 800 displays a home screen 806 (e.g., the Springboard on IOS TM In some example embodiments, the home screen 806 provides status information such as battery life, connectivity, or other hardware status. The user 802 can activate user interface elements by touching the area occupied by the respective user interface element. In this manner, the user 802 interacts with the applications of the mobile device 800. For example, touching the area occupied by a particular icon included in the home screen 806 causes the launching of the application corresponding to the particular icon.
[0147] As Figure 8As shown, the mobile device 800 may include an imaging device 808. The imaging device 808 may be a camera or any other device coupled to the mobile device 800 capable of acquiring a video stream or one or more consecutive images. The imaging device 808 may be triggered by the video modification system 160 or optional user interface elements to initiate the acquisition of a video stream or continuum of images and to pass the video stream or continuum of images to the video modification system 160 for processing according to one or more methods described in this disclosure.
[0148] Many kinds of applications (also known as "application software") can be executed on the mobile device 800, such as native applications (e.g., in iOS). TM Applications running in Objective-C, Swift, or another suitable language, or in Android TM Applications running on this device include Java-programmed applications, mobile web applications (e.g., applications written in Hypertext Markup Language-5 (HTML5)), or hybrid applications (e.g., native shell applications that launch HTML5 sessions). For example, mobile device 800 includes messaging applications, audio recording applications, camera applications, book reader applications, media applications, fitness applications, file management applications, location applications, browser applications, settings applications, contact applications, phone calling applications, or other applications (e.g., game applications, social networking applications, biometric monitoring applications). In another example, mobile device 800 includes applications such as... A social messaging application 810, consistent with some embodiments, allows users to exchange ephemeral messages including media content. In this example, the social messaging application 810 may incorporate aspects of the embodiments described herein. For example, in some embodiments, the social messaging application 810 includes ephemeral media galleries created by users of the social messaging application 810. These galleries may consist of videos or pictures posted by users and viewable by the users' contacts (e.g., "friends"). Alternatively, public galleries may be created by the administrator of the social messaging application 810, which consists of media from any user of the application (and accessible to all users). In yet another embodiment, the social messaging application 810 may include a "magazine" feature, consisting of articles and other content generated by publishers on the platform of the social messaging application 810 and accessible to any user. Any of these environments or platforms can be used to implement the concepts of this disclosure.
[0149] In some embodiments, the ephemeral messaging system can include messages with ephemeral video clips or images that are deleted after a deletion trigger event such as a viewing time or a viewing completion. In such embodiments, a device implementing the video modification system 160 can identify, track, and modify an object of interest within an ephemeral video clip as the ephemeral video clip is being captured by the device and sent to another device using the ephemeral messaging system.
[0150] Software Architecture
[0151] Figure 9 is a block diagram 900 illustrating an architecture of software 902 that can be installed on the above-described devices. Figure 9 The software architecture 900 is merely an example of a non-limiting example of a software architecture and it would be understood that many other architectures can be implemented to facilitate the functionality described herein. In various embodiments, the software 902 is implemented by the hardware of the machine 1000 such as that shown in FIG. 10 that includes processors 1010, memory 1030, and I / O components 1050. In this example architecture, the software 902 can be conceptualized as a stack of layers, where each layer can provide a particular functionality. For example, the software 902 includes layers such as an operating system 904, libraries 906, frameworks 908, and applications 910. Operationally, the applications 910 invoke application programming interface (API) calls 912 through the software stack and receive messages 914 in response to the API calls 912, consistent with some embodiments. Figure 10
[0152] In various implementations, the operating system 904 manages hardware resources and provides common services. The operating system 904 includes, for example, a kernel 920, services 922, and drivers 924. Consistent with some embodiments, the kernel 920 acts as an abstraction layer between the hardware and the other software layers. For example, the kernel 920 provides memory management, processor management (e.g., scheduling), component management, networking, and security settings, among other functionality. The services 922 can provide other common services for the other software layers. According to some embodiments, the drivers 924 are responsible for controlling or interfacing with the underlying hardware. For instance, the drivers 924 can include display drivers, camera drivers, Bluetooth® drivers, flash drivers, flash drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), WiFi® drivers, audio drivers, power management drivers, and so forth. drivers, flash drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), WiFi® drivers, audio drivers, power management drivers, and so forth.
[0153] In some embodiments, libraries 906 provide a low-level common infrastructure executed by the applications 910. Libraries 906 can include system libraries 930 (e.g., C standard library) that can provide functions such as storage allocation functions, string manipulation functions, mathematical functions, and the like. In addition, libraries 906 can include API libraries 932 such as media libraries (e.g., libraries to support presentation and manipulation of various media formats such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., an OpenGL framework for rendering two-dimensional (2D) and three-dimensional (3D) graphics on a display), database libraries (e.g., SQLite to provide various relational database functions), web libraries (e.g., WebKit to provide web browsing functionality), and the like. Libraries 906 also include a wide variety of other libraries 934 to provide many other APIs to the applications 910.
[0154] According to some embodiments, frameworks 908 provide a high-level common infrastructure that can be utilized by the applications 910. For example, frameworks 908 provide various graphical user interface (GUI) functions, high-level resource management, high-level location node functions, and so forth. The frameworks 908 can provide a broad spectrum of other APIs that can be utilized by the applications 910, some of which are specific to a particular operating system or platform.
[0155] In an example embodiment, the applications 910 include a home application 950, a contacts application 952, a browser application 954, a book reader application 956, a location application 958, a media application 960, a messaging application 962, a game application 964, and a broad assortment of other applications such as a third party application 966. According to some embodiments, the applications 910 are programs that execute functions defined in the programs. Various programming languages can be employed to create the applications 910, structured in a variety of manners, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a specific example, the third party application 966 (e.g., an application developed by a vendor, other than the vendor of the particular platform) can be an Android or IOS software development kit (SDK) application developed using an Android or IOS SDK respectively. TM or IOS TM software development kit (SDK). TM , ANDROID TM , mobile software running on a mobile operating system (e.g., iOS®by Apple Inc., Android®by Google, Inc., Windows®by the Microsoft Corporation, or the like, or any combination thereof, or other mobile operating systems such as phone or other mobile operating systems). In this example, the third-party application 966 can invoke the API calls 912 provided by the operating system 904 to facilitate performance of the functionality described herein.
[0156] Example Machine Architectures and Machine-Readable Media
[0157] Figure 10 is a block diagram illustrating components of a machine 1000, according to some embodiments, able to read instructions (e.g., processor-executable instructions) from a machine-readable medium (e.g., a non-transitory processing- readable storage medium or a processor-readable storage device) and perform any one or more of the methodologies discussed herein. Specifically, the machine 1000 can be a special-purpose Figure 10 A schematic illustration of a machine 1000 in the example form of a computer system is shown, within which instructions 1016 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 1000 to perform any one or more of the methodologies discussed herein can be executed. In alternative embodiments, the machine 1000 operates as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1000 can operate in the capacity of a server machine or a client machine in server-client network environments, or as a peer machine in peer-to-peer (or distributed) network environments. The machine 1000 can comprise, but not be limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular telephone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing the instructions 1016, sequentially or otherwise, that specify actions to be taken by machine 1000. Further, while only a single machine 1000 is illustrated, the term "machine" shall also be taken to include a collection of machines 1000 that individually or jointly execute the instructions 1016 to perform any one or more of the methodologies discussed herein.
[0158] In various embodiments, the machine 1000 includes processors 1010, memory 1030, and I / O components 1050, which can be configured to communicate with one another via a bus 1002. In example embodiments, the processors 1010 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio-frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) include, for example, a processor 1012 and a processor 1014, which can execute the instructions 1016. The term “processor” is intended to include multi-core processors that can include two or more independent processors (also referred to as “cores”) that can execute instructions 1016 Figure 10 Multiple processors 1010 are shown, but the machine 1000 can include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0159] According to some embodiments, the memory 1030 includes a main memory 1032, a static memory 1034, and a storage unit 1036 accessible to the processors 1010 via a bus 1002. The storage unit 1036 can include a machine-readable medium 1038 on which are stored instructions 1016 embodying any one or more of the methodologies or functions described herein. The instructions 1016 can also reside, completely or at least partially, within the main memory 1032, within static memory 1034, within at least one of the processors 1010 (e.g., within the processor’s cache memory), or any suitable combination thereof, during their execution by the machine 1000. Thus, in various embodiments, the main memory 1032, the static memory 1034, and the processors 1010 are considered machine-readable media 1038.
[0160] As used herein, the term "memory" refers to a machine-readable medium used for storage of information that is temporary or permanent in nature, and can be considered to include, but is not limited to, random access memory (RAM), read only memory (ROM), buffers, flash memory, and cache memory. While a machine-readable medium 1038 is shown in the example embodiment to be a single medium, the term "machine-readable medium" should be taken to include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) that is capable of storing instructions 1016. The term "machine-readable medium" can also be taken to include any medium or combination of media that is capable of storing instructions (e.g., instructions 1016) for execution by a machine (e.g., machine 1000), such that the instructions, when executed by one or more processors of the machine (e.g., processors 1010), cause the machine to perform any one or more of the methods described herein. Accordingly, a "machine-readable medium" refers to a single storage apparatus or device, as well as "cloud-based" storage systems or storage networks that include multiple storage apparatus or devices. Thus, the term "machine-readable medium" can be taken to include, but not be limited to, one or more data repositories in the form of a solid-state memory (e.g., flash memory), optical media, magnetic media, other non-volatile memory (e.g., erasable programmable read-only memory (EPROM)), or any suitable combination thereof. The term "machine-readable medium" expressly excludes a transitory signal per se.
[0161] I / O components 1050 include a wide variety of components to receive input, provide output, produce output, transmit information, exchange information, capture measurements, and so on. In general, it will be understood that the I / O components 1050 can include many other components that are not shown in FIG. 10. The I / O components 1050 are grouped as shown primarily to simplify the following discussion and are not meant to limit the Figure 10 In various example embodiments, I / O components 1050 include output components 1052 and input components 1054. Output components 1052 include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibratory motor), other signal generators, and so on. Input components 1054 include alphanumeric input components (e.g., a keyboard, a touchscreen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), tactile input components (e.g., a physical button, a touch screen that provides location and force of touches or touch gestures, or other tactile input components), audio input components (e.g., a microphone), and the like.
[0162] In some additional example embodiments, the I / O components 1050 include biometric components 1056, motion components 1058, environmental components 1060, or position components 1062, among a wide array of other components. For example, biometric components 1056 include components to detect expressions (e.g., hand expressions, facial expressions, vocal expressions, body gestures, or mouth gestures), measure biosignals (e.g., blood pressure, heart rate, body temperature, perspiration, or brain waves), identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or electroencephalogram-based identification), and the like. The motion components 1058 include acceleration sensor components (e.g., accelerometer), gravitation sensor components, rotation sensor components (e.g., gyroscope), and so forth. The environmental components 1060 include, for example, illumination sensor components (e.g., photometer), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., barometer), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors components (e.g., machine olfaction detection sensors, gas detection sensors to detect concentrations of hazardous gases for safety or to measure pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to a surrounding physical environment. The position components 1062 include location sensor components (e.g., a Global Position System (GPS) receiver component), altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), and the like.
[0163] Communication can be enabled via a wide variety of technologies. The I / O components 1050 can include communication components 1064 operable to couple the machine 1000 to networks 1080 or to other devices 1070 via coupling 1082 and coupling 1072, respectively. For example, communication components 1064 can include a network interface component or another suitable device to interface with a network 1080. In further examples, communication components 1064 can include wired communication components, wireless communication components, cellular communication components, Near Field Communication (NFC) components, Bluetooth®components (e.g., Bluetooth®low energy), Wi-Fi®components, and other communication components to provide communication via other modalities. The other devices 1070 can be another machine or any of a wide array of peripheral devices (e.g., a peripheral device coupled via a USB). The devices 1070 can be another machine or any of a wide array of peripheral devices (e.g., a peripheral device coupled via a USB).
[0164] Also, in some embodiments, the communication components 1064 detect identifiers or include components operable to detect identifiers (e.g., an RFID tag reader component, a NFC smart tag detection component, an optical reader component (e.g., an optical sensor to detect one- dimensional bar codes such as Universal Product Code (UPC) bar codes, multidimensional bar codes such as Quick Response (QR) code, Aztec code, Data Matrix, Datagiyph, MaxiCode, PDF417, Ultra Code, Uniform Commercial Code Reduced Space Symbol (UCC RSS)-2D bar codes, and other optical codes), an acoustic detection component (e.g., a microphone to identify tagged audio signals), or any suitable combination thereof). Additionally, a variety of information can be derived via the communication components 1064, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, location via Bluetooth® signal detection, or NFC beacon signal detection, etc.
[0165] Transmission Medium
[0166] In various example embodiments, one or more portions of the network 1080 can be a self- organizing network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WW AN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, a network, another type of network, or a combination of two or more such networks. For example, the network 1080 or a portion of the network 1080 can include a wireless or cellular network and the coupling 1082 can be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or another type of cellular or wireless coupling. In this example, the coupling 1082 can implement any of a variety of types of data transfer technology, such as Single Carrier Radio Transmission Technology (lxRTT), Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data Rates for GSM Evolution (EDGE) technology, third Generation Partnership Project (3GPP) including 3G, fourth generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Worldwide Interoperability for Microwave Access (WiMAX), Long Term Evolution (LTE) standard, others defined by various standards setting organizations, other long range protocols, or other data transfer technology.
[0167] In an example embodiment, instructions 1016 are sent or received over network 1080 via a transmission medium using a network interface device (e.g., a network interface component included in communication component 1064), and utilizing any of a plurality of known transport protocols (e.g., HTTP). Similarly, in other example embodiments, instructions 1016 are sent or received to device 1070 via a transmission medium using a coupling 1072 (e.g., peer-to-peer coupling). The term “transmission medium” can be considered to include any intangible medium capable of storing, encoding, or carrying instructions 1016 executed by machine 1000, and includes digital or analog communication signals or other intangible media to facilitate the communication implementation of such software.
[0168] Furthermore, because the machine-readable medium 1038 does not embody a propagating signal, it is non-transient (in other words, it does not possess any transient signals). However, labeling the machine-readable medium 1038 as "non-transient" should not be interpreted as meaning that the medium cannot be moved. The medium should be considered as transferable from one physical location to another. Additionally, since the machine-readable medium 1038 is tangible, it can be considered a machine-readable device.
[0169] language
[0170] Throughout this specification, multiple instances can implement components, operations, or structures described as single instances. While individual operations of one or more methods are shown and described as separate operations, one or more individual operations can be performed simultaneously, and they do not need to be performed in the order shown. Structures and functionalities presented as individual components in example configurations can be implemented as combined structures or components. Similarly, structures and functionalities presented as single components can be implemented as multiple separate components. These and other variations, modifications, additions, and improvements fall within the scope of this document's subject matter.
[0171] While an overview of the subject matter of the invention has been described with reference to specific exemplary embodiments, various modifications and changes may be made to these embodiments without departing from the broader scope of the embodiments of this disclosure. Such embodiments of the subject matter of the invention may be referred to herein individually or collectively by the term "invention," which is merely for convenience and is not intended to limit the scope of this application to any single disclosure or inventive concept if more than one is disclosed in fact.
[0172] The embodiments shown herein are described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments may be used and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Therefore, the specific implementation should not be considered limiting, and the scope of the various embodiments is defined only by the appended claims and the full scope of their equivalents.
[0173] As used herein, the term "or" can be construed in either an inclusive or exclusive sense. Furthermore, instances of the resources, operations, or structures described herein can be provided as a single instance. Moreover, the boundaries between various resources, operations, modules, components, engines, and data stores are somewhat arbitrary, and the same particular operation can be implemented in a different context in which it is combined with other operations. Other allocations of functionality are envisioned and can fall within the scope of the various embodiments of the present disclosure. In general, the boundaries of the various resources, operations, and structures are not as significant as the functionality to be performed by the various resources. Also, structure and functionality presented as discrete components in example configurations can be implemented as a combined structure or resource. Similarly, structures and functionality presented as a single resource can be implemented as discrete resources. These and other variations, modifications, additions, and improvements fall within the scope of the embodiments of the present disclosure as represented by the appended claims. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.
Claims
1. A method for object modeling and replacement in a video stream, comprising: receiving, by one or more processors, a video stream captured by a user device; identifying, in at least one frame of the video stream, a focus object by determining pixels within the video stream corresponding to the focus object; identifying a portion of the focus object; determining, for each frame in a set of previous frames of the video stream, a direction associated with the identified focus object by determining a direction of the identified portion of the focus object; combining the directions of the identified focus object for the set of previous frames to identify an aggregate direction of the identified focus object; and modifying the video stream to include a graphical element by positioning the graphical element within the at least one frame of the video stream based on the aggregate direction of the identified focus object; wherein the focus object is a user’s hand and the portion of the focus object is a finger of the user’s hand; wherein the method further comprises: dynamically increasing a number of frames in the set of previous frames when it is determined that a distance of movement of the finger is sufficient to cause jitter, artifacts, or other undesirable signals; and / or dynamically decreasing the number of frames in the set of previous frames when it is determined that the number of frames is greater than a minimum number. modifying the video stream includes replacing the portion of the identified focus object with the graphical element aligned with the direction of the identified focus object.
2. The method of claim 1, wherein, modifying the video stream includes adding the graphical element to the portion of the identified focus object.
3. The method of claim 1, wherein, the graphical element comprises a weapon in a video game and wherein the weapon is pointed in a direction associated with the identified focus object.
4. The method of claim 1, wherein, 5. The method of claim 1, further comprising: detecting movement of the identified focus object over a plurality of frames of the video stream; and wherein modifying the video stream to include the graphical element comprises repositioning the graphical element within the plurality of frames based on the detected movement. dynamically modifying a histogram threshold used to identify pixels as corresponding to the identified focus object based on the direction of the identified focus object.
7. A system for object modeling and replacement in a video stream, comprising:
6. The method of claim 1, further comprising: one or more processors; and processor-readable storage storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: receiving a video stream captured by a user device; identifying, in at least one frame of the video stream, a focus object by determining pixels within the video stream corresponding to the focus object; identifying a portion of the focus object; determining, for each frame in a set of previous frames of the video stream, a direction associated with the identified focus object by determining a direction of the identified portion of the focus object; combining the directions of the identified focus object for the set of previous frames to identify an aggregate direction of the identified focus object; and modify the video stream to include the graphical element by positioning the graphical element within the at least one frame of the video stream based on the aggregate direction of the identified focus object; wherein the focus object is a hand of a user and the portion of the focus object is a finger of the user; wherein the operations further comprise: dynamically increasing a number of frames of the set of previous frames when it is determined that a distance of movement of the finger is sufficient to cause jitter, artifacts, or other undesirable signals; and / or dynamically decreasing the number of frames when it is determined that the number of frames is greater than a minimum number.
8. The system of claim 7, wherein, modifying the video stream includes replacing a portion of the identified focus object with the graphical element aligned with the direction of the identified focus object.
9. The system of claim 7, wherein, modifying the video stream includes adding the graphical element to a portion of the identified focus object.
10. The system of claim 7, wherein, the graphical element comprises a weapon in a video game and wherein the weapon is pointed in a direction associated with the identified focus object.
11. The system of claim 7, wherein, the operations further comprise: detecting movement of the identified focus object over a plurality of frames of the video stream; and wherein modifying the video stream to include the graphical element comprises repositioning the graphical element within the plurality of frames based on the detected movement.
12. The system of claim 7, wherein, the operations further comprise dynamically modifying a histogram threshold used to identify pixels as corresponding to the identified focus object based on the direction of the identified focus object.
13. A non-transitory processor-readable storage device storing processor-executable instructions that, when executed by a processor of a machine, cause the machine to perform operations comprising: receiving a video stream captured by a user device; identifying a focus object in at least one frame of the video stream by determining pixels within the video stream that correspond to the focus object; identifying a portion of the focus object; determining a direction associated with the identified focus object for each frame of a set of previous frames of the video stream by determining a direction of the identified portion of the focus object; combining the directions of the identified focus object for the set of previous frames to identify an aggregate direction of the identified focus object; and modify the video stream to include the graphical element by positioning the graphical element within the at least one frame of the video stream based on the aggregate direction of the identified focus object; wherein the focus object is a hand of a user and the portion of the focus object is a finger of the user; wherein the operations further comprise: dynamically increasing a number of frames of the set of previous frames when it is determined that a distance of movement of the finger is sufficient to cause jitter, artifacts, or other undesirable signals; and / or dynamically decreasing the number of frames when it is determined that the number of frames is greater than a minimum number.
14. A method for object modeling and replacement in a video stream, comprising: accessing, by one or more processors, a video stream; generating a display of the video stream that depicts a focus object in at least one frame of the video stream; identifying the object of interest in the at least one frame of the video stream by determining pixels within the video stream that correspond to the object of interest; identifying a portion of the object of interest; determining a direction associated with the object of interest by determining a direction of a portion of the object of interest for a set of previous frames of the video stream; combining the directions of the object of interest for the set of previous frames to identify a target direction of the object of interest; modifying the video stream such that a graphical element is positioned and depicted within the at least one frame of the video stream that depicts the object of interest based on the target direction of the object of interest; and repositioning the graphical element within the video stream based on detected movement of the object of interest; wherein the object of interest is a hand of a user and the portion of the object of interest is a finger of the user; wherein the method further comprises: dynamically increasing a number of frames of the set of previous frames when it is determined that a distance of movement of the finger is sufficient to cause jitter, artifacts, or other undesirable signals; and / or dynamically decreasing a number of frames of the set of previous frames when it is determined that the number of frames is greater than a minimum number. replacing a portion of the object of interest with the graphical element that is aligned with the direction of the object of interest.
15. The method of claim 14, further comprising: adding the graphical element to a portion of the object of interest.
16. The method of claim 14, further comprising: the graphical element comprises a weapon in a video game and wherein the weapon is pointed in a direction associated with the object of interest.
17. The method of claim 14, wherein, detecting movement of the object of interest over a plurality of frames of the video stream.
18. The method of claim 14, further comprising: dynamically modifying a histogram threshold used to identify pixels as corresponding to the object of interest based on the direction of the object of interest.
19. The method of claim 14, further comprising:
20. A system for object modeling and replacement in a video stream, comprising: one or more processors; and a processor-readable storage device storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: accessing a video stream; generating a display of the video stream that depicts an object of interest in at least one frame of the video stream; identifying the object of interest in the at least one frame of the video stream by determining pixels within the video stream that correspond to the object of interest; identifying a portion of the object of interest; determining a direction associated with the object of interest by determining a direction of a portion of the object of interest for a set of previous frames of the video stream; combining the directions of the object of interest for the set of previous frames to identify a target direction of the object of interest; modifying the video stream such that a graphical element is positioned and depicted within the at least one frame of the video stream that depicts the object of interest based on the target direction of the object of interest; and repositioning the graphical element within the video stream based on detected movement of the object of interest; wherein the object of interest is a hand of a user and the portion of the object of interest is a finger of the user; wherein the operations further comprise: dynamically increasing a number of frames of the set of previous frames when it is determined that a distance of movement of the finger is sufficient to cause jitter, artifacts, or other undesirable signals; and / or dynamically decreasing a number of frames of the set of previous frames when it is determined that the number of frames is greater than a minimum number. dynamically increasing a number of frames of the set of previous frames when it is determined that a distance of movement of the finger is sufficient to cause jitter, artifacts, or other undesirable signals; and / or dynamically decreasing the number of frames of the set of previous frames when it is determined that the number of frames is greater than a minimum number.
21. The system of claim 20, wherein, the operations further comprise replacing a portion of the object of interest with the graphical element aligned with the direction of the object of interest.
22. The system of claim 20, wherein, the operations further comprise adding the graphical element to a portion of the object of interest.
23. The system of claim 20, wherein, the graphical element comprises a weapon in a video game, and wherein the weapon is pointed in a direction associated with the object of interest.
24. The system of claim 20, wherein, the operations further comprise detecting movement of the object of interest over a plurality of frames of the video stream.
25. The system of claim 20, wherein, the operations further comprise dynamically modifying a histogram threshold used to identify pixels as corresponding to the object of interest based on the direction of the object of interest.
26. A non-transitory processor-readable storage device storing processor-executable instructions that, when executed by a processor of a machine, cause the machine to perform operations comprising: accessing a video stream; generating a display of the video stream depicting an object of interest in at least one frame of the video stream; identifying the object of interest in the at least one frame of the video stream by determining pixels within the video stream corresponding to the object of interest; identifying a portion of the object of interest; determining a direction associated with the object of interest by determining a direction of a portion of the object of interest for a set of previous frames of the video stream; combining the directions of the object of interest for the set of previous frames to identify a target direction of the object of interest; modifying the video stream such that a graphical element is positioned and depicted within the at least one frame of the video stream depicting the object of interest based on the target direction of the object of interest; and repositioning the graphical element within the video stream based on detected movement of the object of interest; wherein the object of interest is a hand of a user and the portion of the object of interest is a finger of the user; wherein the operations further comprise: dynamically increasing a number of frames of the set of previous frames when it is determined that a distance of movement of the finger is sufficient to cause jitter, artifacts, or other undesirable signals; and / or dynamically decreasing the number of frames of the set of previous frames when it is determined that the number of frames is greater than a minimum number.
27. The non-transitory processor-readable storage device of claim 26, wherein, the operations comprise replacing a portion of the object of interest with the graphical element aligned with the direction of the object of interest.
28. The non-transitory processor-readable storage device of claim 26, wherein, the graphical element comprises a weapon in a video game, and wherein the weapon is pointed in a direction associated with the object of interest.
Citation Information
Patent Citations
Method and device for implementing dynamic image processing
CN101247482A
Camera-based user input for compact devices
CN101689244A