Image segmentation and modification of a video stream
By identifying and tracking facial features in video streams through an image segmentation system, and cropping and modifying regions of interest, the problem of image modification in video communication is solved, thus improving the quality of video communication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2016-11-29
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies make it difficult to modify images within a video stream in real time during video communication, especially for operations such as image segmentation of facial features and teeth whitening, resulting in poor video communication performance.
The system identifies and tracks objects of interest, such as facial features, in the video stream using an image segmentation system, crops image data outside the region of interest, and modifies the color values of the objects of interest through image processing operations to generate a modified video stream.
It enables real-time image segmentation of facial features and teeth whitening in video communication, improving the effect of video communication and user experience.
Smart Images

Figure CN116342622B_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application filed on November 29, 2016, with application number 201680080497.X and invention title "Image Segmentation and Modification of Video Stream".
[0002] Priority requirements
[0003] This application claims the benefit of priority to U.S. Patent Application Serial No. 14 / 953,726, filed November 30, 2015, entitled “Image Segmentation and Modification of a Video Stream,” and each of those applications is hereby claimed, and each of those applications is incorporated herein by reference in its entirety. Technical Field
[0004] Embodiments of this disclosure generally relate to automatic image segmentation of video streams. More specifically, but not limitingly, this disclosure relates to systems and methods for image segmentation of identified regions of interest within a face depicted in a video stream. Background Technology
[0005] Telecommunications applications and devices can use various media such as text, images, audio recordings, and / or video recordings to provide communication between multiple users. For example, video conferencing allows two or more individuals to communicate with each other using a combination of software applications, telecommunications devices, and telecommunications networks. Telecommunications devices can also record video streams to send messages across telecommunications networks. Attached Figure Description
[0006] The various figures in the accompanying drawings illustrate only exemplary embodiments of this disclosure and should not be construed as limiting its scope.
[0007] Figure 1 This is a block diagram illustrating a networking system according to some example embodiments.
[0008] Figure 2 This is a diagram illustrating an image segmentation system according to some example embodiments.
[0009] Figure 3 This is a flowchart illustrating an example method for segmenting an image within a video stream and modifying portions of the video stream based on the segmentation, according to some example embodiments.
[0010] Figure 4 The regions of interest within one or more images of a video stream according to some example embodiments are shown.
[0011] Figure 5 The image shows a binarized image of the region of interest according to some example embodiments.
[0012] Figure 6 This is a flowchart illustrating an example method for segmenting an image within a video stream and modifying portions of the video stream based on the segmentation, according to some example embodiments.
[0013] Figure 7 A binarized image of a region of interest with noisy pixels is shown according to some example embodiments.
[0014] Figure 8 This is a flowchart illustrating an example method for tracking and modifying objects of interest in a video stream, according to some example embodiments.
[0015] Figure 9 A set of marked pixels is shown within a region of interest according to some example embodiments.
[0016] Figure 10 These are user interface diagrams depicting example mobile devices and mobile operating system interfaces according to some example embodiments.
[0017] Figure 11 This is a block diagram illustrating an example of a software architecture that can be installed on a machine according to some example embodiments.
[0018] Figure 12 It is a block diagram of a graphical representation of a machine in the form of a computer system according to an example embodiment, within which a set of instructions can be executed to cause the machine to perform any of the methods discussed herein.
[0019] The headings provided here are for convenience only and are not intended to affect the scope or meaning of the terms used. Detailed Implementation
[0020] The following description includes systems, methods, techniques, instruction sequences, and computer program products illustrating embodiments of the present disclosure. In this description, numerous specific details are set forth for purposes of explanation in order to provide an understanding of various embodiments of the subject matter of the invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of the invention can be practiced without these specific details. Generally, well-known examples of instructions, protocols, structures, and techniques need not be shown in detail.
[0021] Although telecommunications applications and devices exist to provide bidirectional video communication between two devices, problems with video streaming, such as modifying images within the video stream during a communication session, can exist. Methods commonly accepted for editing or modifying video alter the video or video communication itself while it is being captured or during video communication. Therefore, there remains a need in the art to improve video communication between devices.
[0022] Embodiments of this disclosure generally relate to automatic image segmentation of video streams. Some embodiments relate to image segmentation of regions of interest identified within a face depicted in a video stream. For example, in one embodiment, an application running on a device receives video captured by the device. The video captured by the device is a video stream such as a video conference or video chat between mobile devices. The application identifies the mouth and exposed teeth within the mouth in the video stream. The application tracks the exposed teeth across the video stream. When receiving a video stream from the device's camera and sending the video stream to another mobile device, the application modifies the exposed teeth in the video stream when teeth are visible in the video stream. The application modifies the teeth by whitening them. A mobile device receiving a streaming video conference from a mobile device running the application displays the teeth of a person in the video chat whitened. The application can also change the color of the teeth to any desired color, rotate between colors, and display multiple colors.
[0023] The above is a specific example. Various embodiments of this disclosure relate to apparatus and instructions executed by one or more processors of the apparatus to modify a video stream transmitted from one apparatus to another while simultaneously acquiring (e.g., modifying the video stream in real time). An image segmentation system is described that identifies and tracks objects of interest across a video stream and through a set of images including the video stream. In various example embodiments, the image segmentation system identifies and tracks one or more facial features depicted in the video stream. Although facial features have been described, it should be understood that, as discussed below, the image segmentation system can track any object of interest.
[0024] An image segmentation system receives a video stream from an imaging device and identifies the approximate location of an object of interest within the images of the video stream. It identifies a region of interest surrounding the object of interest. In some embodiments, the image containing the object of interest in a portion of the video stream is cropped to remove image data outside the region of interest. The image segmentation system may perform one or more image processing operations on the region of interest to increase contrast and manipulate pixel values to identify the object of interest within the region of interest and isolate the object of interest from other objects, shapes, textures, or other features within the region of interest. Once a specific pixel for the object of interest is identified, the image segmentation system may track the pixels of the object of interest across other portions of the video stream. In some embodiments, the image segmentation system modifies the values of the object of interest within the video stream by identifying the relative positions of the pixels of the object of interest by referencing other points within the image and tracking the pixels corresponding to the positions of the objects of interest. The image segmentation system can modify the appearance of the object of interest within the video stream by modifying the color values of the pixels representing the object of interest. In some cases, the image segmentation system generates an image layer that is overlaid on the images within the video stream to modify the appearance of the object of interest.
[0025] Figure 1This is a network diagram depicting a network system 100 with a client-server architecture configured for exchanging data over a network, according to one embodiment. For example, network system 100 may be a messaging system in which clients communicate and exchange data within network system 100. This data may relate to various functions (e.g., sending and receiving text and media communications, determining geographic location, etc.) and aspects (e.g., transmitting communication data, receiving and sending indications of communication sessions, etc.) associated with network system 100 and its users. Although shown herein as a client-server architecture, other embodiments may include other network architectures, such as peer-to-peer or distributed network environments.
[0026] like Figure 1 As shown, network system 100 includes social messaging system 130. Social messaging system 130 is typically based on a three-tier architecture, consisting of an interface layer 124, an application logic layer 126, and a data layer 128. As understood by those skilled in the art of computer and internet technology, Figure 1 Each module or engine shown represents a set of executable software instructions and corresponding hardware (e.g., memory and processor) for executing those instructions, forming a hardware-implemented module or engine, and acting as a dedicated machine configured to perform a specific set of functions when executing the instructions. To avoid obscuring the subject matter of the invention with unnecessary detail, from... Figure 1 Various functional modules and engines irrelevant to conveying the subject matter of this invention have been omitted. Of course, additional functional modules and engines can be used with social messaging systems to implement additional functions not specifically described herein, such as... Figure 1 As shown in the image. Furthermore... Figure 1 The various functional modules and engines described can reside on a single server computer or client device, or they can be distributed across multiple server computers or client devices in various arrangements. Furthermore, although... Figure 1 The social messaging system 130 is described as having a three-layer architecture, but the subject matter of this invention is not limited to this architecture.
[0027] like Figure 1 As shown, interface layer 124 comprises interface module 140 (e.g., web server) that receives requests from various client computing devices and servers, such as client device 110 executing client application 112 and third-party server 120 executing third-party application 122. In response to a received request, interface module 140 transmits an appropriate response to the requesting device via network 104. For example, interface module 140 may receive requests such as Hypertext Transfer Protocol (HTTP) requests or other web-based application programming interface (API) requests.
[0028] Client device 110 can execute conventional web browser applications or applications (also referred to as "application software") developed for specific platforms to include any of a variety of mobile computing devices and mobile-specific operating systems (such as iOS™, Android™, Windows® Phone). Furthermore, in some example embodiments, client device 110 forms all or part of image segmentation system 160, such that modules of image segmentation system 160 configure client device 110 to perform a specific set of functions relating to the operation of image segmentation system 160.
[0029] In the example, client device 110 is executing client application 112. Client application 112 can provide the functionality to present information to user 106 and communicate via network 104 to exchange information with social messaging system 130. Furthermore, in some examples, client device 110 performs the functions of image segmentation system 160 to segment images of a video stream during video stream acquisition and transmit the video stream (e.g., using image data modified from segmented images based on the video stream).
[0030] Each of the client devices 110 may include a computing device that includes at least the ability to display and communicate with the social messaging system 130, other client devices, and third-party server 120 via network 104. Client devices 110 include, but are not limited to, remote devices, workstations, computers, general-purpose computers, internet infrastructure, handheld devices, wireless devices, portable devices, wearable computers, cellular or mobile phones, personal digital assistants (PDAs), smartphones, tablet computers, ultrabooks, netbooks, laptop computers, desktop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, set-top boxes, network PCs, minicomputers, etc. User 106 may be a person, a machine, or other tool that interacts with client device 110. In some embodiments, user 106 interacts with social messaging system 130 via client device 110. User 106 may not be part of a networked environment but may be associated with client device 110.
[0031] like Figure 1 As shown, data layer 128 has a database server 132 that facilitates access to an information repository or database 134. Database 134 is a storage device for storing data such as member profile data, social graph data (e.g., relationships between members of the social messaging system 130), image modification preference data, accessibility data, and other user data.
[0032] Individuals can register using the social messaging system 130 to become members of the social messaging system 130. Once registered, members can form social network relationships (e.g., friends, followers, or contacts) on the social messaging system 130 and interact with a wide range of applications provided by the social messaging system 130.
[0033] Application logic layer 126 includes various application logic modules 150, which, together with interface module 140, generate various user interfaces with data obtained from various data sources or data services in data layer 128. Each application logic module 150 can be used to implement functionality associated with various applications, services, and features of social messaging system 130. For example, a social messaging application can be implemented by application logic module 150. The social messaging application provides a messaging mechanism for users of client device 110 to send and receive messages including text and media content (such as pictures and videos). Client device 110 can access and view messages from the social messaging application for a specific time period (e.g., limited or unlimited). In the example, a message recipient can access a specific message for a predefined duration (e.g., specified by the message sender), which begins upon the first access to the specific message. After the predefined duration has elapsed, the message is deleted, and the message recipient can no longer access the message. Of course, other applications and services can be implemented separately in their own application logic modules 150.
[0034] like Figure 1 As shown, the social messaging system 130 may include at least a portion of an image segmentation system 160 capable of identifying, tracking, and modifying video data during video data acquisition by the client device 110. Similarly, as described above, the client device 110 includes a portion of the image segmentation system 160. In other examples, the client device 110 may include the entire image segmentation system 160. In instances where the client device 110 includes a portion (or all) of the image segmentation system 160, the client device 110 may operate independently or in conjunction with the social messaging system 130 to provide the functionality of the image segmentation system 160 described herein.
[0035] In some embodiments, the social messaging system 130 may be a short-time messaging system that enables short-term communication in which content (e.g., video clips or images) is deleted after a deletion trigger event, such as viewing time or viewing completion. In this embodiment, the apparatus uses various modules described herein within any context of generating, sending, receiving, or displaying short-time messages. For example, an apparatus implementing the image segmentation system 160 can identify, track, and modify objects of interest, such as a set of exposed teeth within a mouth depicted in a video clip. The apparatus can modify the object of interest during the acquisition of the video clip without performing image processing after the video clip has been acquired as part of content generation for short-time messaging.
[0036] exist Figure 2 In various embodiments, the image segmentation system 160 may be implemented as a standalone system or in conjunction with the client device 110, and is not necessarily included in the social messaging system 130. The image segmentation system 160 is shown as including a location module 210, a video processing module 220, a recognition module 230, a modification module 240, a tracking module 250, and a communication module 260. For example, all or some of the modules 210-260 communicate with each other, for example via network coupling, shared memory, etc. Each module 210-260 may be implemented as a single module, combined with other modules, or further subdivided into multiple modules. Although not shown, other modules unrelated to the example embodiment may also be included.
[0037] The location module 210 performs a localization operation within the image segmentation system 160. In various example embodiments, the location module 210 identifies and provides the location of an object of interest depicted by images (e.g., one or more frames of the video stream) from a video stream. In some embodiments, the location module 210 may be part of a face tracking module or system. In some cases, where the object of interest is part of a face, the location module 210 identifies the location of the face depicted in one or more images within the video stream and one or more facial features depicted on the face. For example, when the location module 210 is configured to locate exposed teeth within a mouth, the location module 210 may identify the face depicted in the image, identify the mouth on the face, and identify exposed teeth within a portion of an image of the video stream that includes the mouth.
[0038] In at least some embodiments, the location module 210 locates a region of interest within one or more images containing the object of interest. For example, the region of interest identified by the location module 210 may be a portion of an image within a video stream, such as a rectangle in which the object of interest appears. Although referred to as a rectangle, the region of interest may be any suitable shape or combination of shapes, as described below. For example, the region of interest may be represented as a circle, a polygon, or similar in shape and size to the object of interest (e.g., a mouth, a wall, the outline of a vehicle) and include the outline of the object of interest.
[0039] In some embodiments, the location module 210 performs a cropping function. For example, after determining the region of interest within an image, the location module 210 crops the image, removing the region of interest outside the region of interest. In some cases, after cropping, the region of interest is processed by one or more other modules of the image segmentation system 160. In the case of processing by other modules, the location module 210, either alone or in cooperation with the communication module 260, transmits the cropped region of interest to one or more other modules (e.g., the video processing module 220).
[0040] Video processing module 220 performs one or more video processing functions on one or more regions of interest identified by location module 210 on one or more images within a video stream. In various embodiments, video processing module 220 converts the regions of interest of one or more images of the video stream into binarized regions of interest, where each pixel has one of two values. As described below, the binarized image generated by video processing module 220 contains two possible values, 0 or 1, for any pixel within the binarized region of the image. Video processing module 220 can convert the regions of interest or one or more images of the video to represent the depicted object of interest as a contrasting color image. For example, pixels in the binarized image can be converted to represent only black (e.g., value 1) and white (e.g., value 0). Although described in this disclosure as a contrasting image composed of black and white pixels, video processing module 220 can convert the regions of interest of one or more images of the video stream into any two contrasting colors (e.g., red and blue). After binarization, video processing module 220 can send or otherwise pass the binarized regions of interest to one or more additional modules of image segmentation system 160. For example, the video processing module 220, alone or in collaboration with the communication module 260, can pass the binarized region of interest to the recognition module 230.
[0041] In some embodiments, in order to increase the contrast between the first group of pixels and the second group of pixels while maintaining the shape, size and pixel region of the object of interest represented by the first group of pixels or the second group of pixels, the video processing module 220 processes the region of interest to increase the contrast between the first group of pixels and the second group of pixels without binarizing the region of interest or generating a binarized region of interest through an iterative process.
[0042] The recognition module 230 identifies a first group of pixels and a second group of pixels within the area of interest modified by the video processing module 220. The recognition module 230 can identify the group of pixels based on the different color values between the first group of pixels and the second group of pixels, or based on a combination of color differences between the first group of pixels, the second group of pixels, and a threshold.
[0043] In some embodiments, the recognition module 230 marks a first group of pixels within a set of images in the video stream, generating a set of marked pixels. In some cases, the recognition module 230 also identifies the location of the first group of pixels against a set of reference landmarks within the video stream or the region of interest. The recognition module 230 identifies the color values of the first group of pixels within one or more of the first and second sets of images in the video stream. With the first group of pixels already marked, the recognition module 230 may cooperate with the tracking module 250 to identify pixels in the second set of images in the video stream corresponding to the marked pixels and to identify the color values of the corresponding pixels.
[0044] Modification module 240 can perform one or more modifications to an object of interest or a portion of an object of interest within the second set of images of the video stream based on the tracking module 250. For example, tracking module 250 can identify the location of a portion of an object of interest within the second set of images of the video stream and transmit that location to modification module 240. Modification module 240 can then perform one or more modifications to the portion of the object of interest to generate a modified second set of images of the video stream. Modification module 240 can then transmit the modified second set of images of the video stream to communication module 260 for transmission to another client device, social messaging system 130, or storage device of client device 110 while the video stream is being acquired. In some embodiments, modification module 240 can perform modifications in real time to transmit the modified second set of images of the video stream in full-duplex communication between two or more client devices. For example, when tracking module 250 tracks pixels or landmarks identified by recognition module 230, modification module 240 can modify the color of an object of interest within the second set of images of the video stream (e.g., whitening exposed teeth in the mouth) to generate a modified second set of images of the video stream. Modification module 240, in cooperation with communication module 260, can send a modified second set of images from the video stream of client device 110 to one or more other client devices.
[0045] The tracking module 250 tracks at least one of an object of interest and a portion thereof, based at least in part on data generated by the identification module 230 (e.g., identified pixels, marker pixels, and landmark reference pixels). In some embodiments, the tracking module 250 identifies and tracks corresponding pixels within one or more images of a second set of images in the video stream. Corresponding pixels represent a first set of pixels (e.g., the object of interest) within the second set of images in the video stream. In some cases, the tracking module 250 tracks marker pixels (e.g., corresponding pixels) based on identified landmarks or a set of landmarks. For example, marker pixels may be included in a binary mask as additional landmark points for tracking locations identified with respect to other landmarks in the binary mask. The tracking module 250 may also track occlusions of the object of interest so that the modification module 240 modifies or avoids modifying pixels based on whether the object of interest is displayed within images of the second set of images in the video stream.
[0046] Communication module 260 provides various communication functions. For example, communication module 260 receives communication data indicating data received from input of client device 110. The communication data may indicate a modified video stream created by a user on client device 110 for storage or transmission to another user's client device. Communication module 260 may transmit communication data between client devices via a communication network. Communication module 260 may exchange network communications with database server 132, client device 110, and third-party server 120. Information obtained by communication module 260 includes user-related data (e.g., member profile data from online accounts or social networking service data) or other data to facilitate the implementation of the functions described herein. In some embodiments, communication module 260 enables communication between one or more of location module 210, video processing module 220, identification module 230, modification module 240, and tracking module 250.
[0047] Figure 3 A flowchart illustrating an example method 300 for segmenting a portion of a video stream and modifying that portion based on the segmentation is provided. The operation of method 300 can be performed by components of the image segmentation system 160 and is described below for illustrative purposes.
[0048] In operation 310, the location module 210 determines the approximate location of an object of interest within the video stream. The video stream comprises a set of images. In some embodiments, the set of images in the video stream is divided into a first set of images and a second set of images. In these cases, the first set of images represents a portion of the video stream, wherein the image segmentation system 160 processes one or more images and identifies the object of interest and its relationship to one or more reference objects within the video stream. The second set of images represents a portion of the video stream in which the image segmentation system 160 tracks the object of interest. The image segmentation system 160 may modify one or more aspects of the tracked object of interest within the second set of images.
[0049] In various example embodiments, the location module 210 is configured to identify and locate predetermined objects or object types. For example, the location module 210 may be configured to identify and locate walls, vehicles, facial features, or any other objects appearing in the video stream. In some cases, the location module 210 is configured to identify objects selected from multiple object types. In the case of selecting objects from multiple object types, the image segmentation system 160 may receive user input selecting objects from a list, table, or other organized set of objects. Objects may also be automatically selected by the image segmentation system 160.
[0050] When the image segmentation system 160 selects objects from multiple object types, it identifies one or more objects within the video stream as members of the multiple objects. The image segmentation system 160 may determine the object of interest from the one or more objects based on location, salience, size, or any other suitable method. In some cases, the image segmentation system 160 identifies the object of interest from one or more objects by selecting an object located near the center of an image. In some embodiments, the image segmentation system 160 identifies the object of interest based on its location within multiple images of a first set of images. For example, the image segmentation system 160 selects an object if it is salientally placed (e.g., near the center of an image) and if the object is detected as a percentage or number of images in the first set of images (e.g., where the percentage or number exceeds a predetermined threshold).
[0051] In some embodiments, such as Figure 4 As shown, the position module 210 determines the approximate position of the mouth 410 within the video stream (e.g., within the first set of images in the video stream). When the position module 210 determines the mouth, it can utilize a set of face tracking operations to determine one or more landmarks within the face depicted in the image or set of images, and identify landmarks representing the mouth.
[0052] In operation 320, the location module 210 identifies a region of interest within one or more images of the first set of images. A region of interest is a portion of one or more images containing the approximate location of the object of interest. In some cases, where the object of interest is the mouth, the region of interest is a portion of an image or a set of images extending across the width of the face to include the mouth (e.g., the width of the mouth extending between the junctures at each corner of the mouth) and portions of the face surrounding the mouth. The region of interest may also extend across the height of the face to include the mouth (e.g., the height of the mouth extending between the uppermost vermilion border of the upper lip and the lowermost vermilion border of the lower lip) and portions of the face surrounding the mouth.
[0053] In some embodiments where the focus is on the mouth, the region of interest can be a bounded region, or represented by a bounded region 400 surrounding the portion of the mouth 410 and the face 420, such as... Figure 4 As shown in the diagram. The bounded region can be rectangular, circular, elliptical, or any other suitable shape of the boundary region. In cases where facial landmarks are used to identify the mouth (object of interest) via one or more face tracking operations, the region of interest can be a portion of an image or a set of images, including a predetermined portion of the image or extending a predetermined distance from any landmark for mouth identification. For example, the region of interest can occupy five percent, fifteen percent, or twenty percent of an image region or a set of images. As a further example, the region of interest can extend in any given direction between 10 pixels and 100 pixels. While embodiments of this disclosure present measurements and percentages of images, it should be understood that measurements and percentages can be higher or lower based on one or more aspects of the image (e.g., resolution, size, pixel count) and one or more aspects of the display device depicting the image (e.g., display size, display aspect ratio, resolution). For example, the region of interest can be any size, ranging from one pixel in an image in the first set of images of a video stream to an entire region.
[0054] In various situations, the area of interest can be a single area of interest or multiple areas of interest. For example, when the location module 210 is identifying areas of interest for a pair of walls on opposite sides of a room, the location module 210 identifies a first area of interest for the first wall on the first side of the room and a second area of interest for the second wall on the second side of the room.
[0055] In operation 330, video processing module 220 generates a modified region of interest. Video processing module 220 generates the modified region of interest by performing one or more image processing functions on the region of interest. In some embodiments, operation 330 includes a set of sub-operations for generating the modified region of interest.
[0056] In operation 332, video processing module 220 crops one or more portions of one or more images from the first group of images outside the region of interest. For example, the region of interest can be cropped into... Figure 4 The bounded region 400 of the region of interest is depicted in the figure. In some embodiments, when cropping one or more images, the video processing module 220 isolates the region of interest by removing portions of one or more images that occur outside the region of interest. For example, if the region of interest is located near the center of the region of interest and includes 15 percent of the images in the one or more images, the video processing module 220 removes 85 percent of the images that are not bounded within the region of interest.
[0057] In various example embodiments, after cropping one or more images to remove image data outside the bounded region of the region of interest, the video processing module 220 performs one or more operations to generate a modified region of interest, such that the recognition module 230 can distinguish objects of interest within the modified region of interest from irrelevant shapes, features, textures, and other aspects of images or a set of images also located within the region of interest. In some cases, in operation 334, the video processing module 220 generates the modified region of interest through binarization. In these embodiments, the video processing module 220 generates a binary image version of the region of interest 500, such as... Figure 5 As shown in the diagram. The video processing module 220 can use any suitable binarization method to generate a binary version of the region of interest. After binarization, the modified region of interest can depict the object of interest as a first set of pixels, and other features of the region of interest as a second set of pixels. The following is about... Figure 6 Describe other operations used to generate the modified region of interest.
[0058] In operation 340, the identification module 230 identifies a first group of pixels and a second group of pixels within the modified region of interest. In various embodiments, the identification module 230 identifies the first group of pixels as distinct from the second group of pixels based on a comparison of the values assigned to the first group and the second group of pixels. The identification module 230 may identify the first group of pixels as having a first value and the second group of pixels as having a second value. In the case where the modified region of interest is binarized (containing pixel color values of 1 or 0), the identification module 230 identifies the first group of pixels as pixels with a color value of 1 (e.g., the first value) and the second group of pixels as pixels with a color value of 0 (e.g., the second value).
[0059] When the modified region of interest contains multiple pixel values (such as grayscale), the recognition module 230 performs a value comparison for each pixel based on a predetermined threshold. In these cases, the recognition module 230 identifies pixels included in a first group of pixels whose pixel values are higher than the predetermined threshold. Pixels included in a second group of pixels are identified as having values lower than the predetermined threshold. For example, if the grayscale (e.g., intensity) values of the pixels in the modified region of interest range from 0 to 256, the recognition module 230 may identify pixels having values greater than 50 percent of that range as included in the first group of pixels. In some cases, the recognition module 230 identifies pixels as belonging to a first group of pixels whose pixel values are within 20 percent of the grayscale (e.g., greater than 240). While specific values for grayscale intensity values have been provided, it should be understood that these values can be modified based on multiple factors related to the image and the display device depicting the image.
[0060] In operation 350, modification module 240 modifies the color values of a first group of pixels within a second group of images in the video stream. In various embodiments, the color values of the first group of pixels are first color values within the second group of images. In these cases, modification module 240 identifies the first color values of the first group of pixels within image data of one or more images comprising the second group of images. In response to identifying the first color values of the first group of pixels, modification module 240 replaces the first color values with second color values that are different from the first color values. After modification module 240 replaces the second color values with the first color values in one or more images, when the second group of images within the video stream is displayed, the object of interest is represented by the color of the second color value.
[0061] In some cases, the modification module 240 modifies the color values of the first group of pixels within the second group of images by generating a modified image layer for the first group of pixels. In these embodiments, the modified image layer is generated using the first pixels with modified color values. The modification module 240 applies the modified image layer to at least a portion of the second group of images (e.g., one or more images in the second group of images, or a portion of images in the second group of images). The modification module 240 may cooperate with the tracking module 250 to align the positions of the first group of pixels within the second group of images with the modified image layer.
[0062] Figure 6 A flowchart illustrating an example method 600 for segmenting a video stream and modifying one or more segmented portions of the video stream is shown. The operations of method 600 can be performed by components of the image segmentation system 160. In some cases, certain operations of method 600 can be performed using one or more operations of method 300, or can be performed as sub-operations of one or more operations of method 300, as will be explained in more detail below.
[0063] In various example embodiments, method 600 is initially performed by image segmentation system 160, performing operations 310 and 320, as described above regarding... Figure 3 In these embodiments, the image segmentation module determines the approximate location of the object of interest and identifies regions of interest within one or more images of the first set of images in the video stream.
[0064] In operation 610, video processing module 220 converts pixels within a region of interest to grayscale values to generate a grayscale region of interest. The grayscale region of interest is the result of grayscale conversion, removing color information from pixels. The remaining value of any given pixel is an intensity value, where higher intensities are displayed as brighter and weaker intensities as darker. The range of intensity values displayed within the grayscale region of interest can vary based on factors associated with one or more images. In some cases, the intensity value range can be from 0 (e.g., represented as black) to 65536 (e.g., represented as white). The intensity range may be limited to less than or greater than 65536. For example, in some cases, the intensity range extends from zero to 256. Although reference conversion to grayscale has been discussed, in some embodiments, pixels within a region of interest may be converted to a single channel other than grayscale. In some cases, video processing module 220 identifies intensity values representing grayscale without converting the region of interest to a grayscale region of interest.
[0065] In some example embodiments, the video processing module 220 uses Equation 1 to convert the color value of a pixel into a grayscale value:
[0066]
[0067] As shown in Equation 1 above, "g" is the green value of the pixel, "r" is the red value of the pixel, "b" is the blue value of the pixel, and "v" is the resulting grayscale value. In embodiments using Equation 1, each pixel within the region of interest has a set of color values. This set of color values can be expressed as a triplet. Each value within this set of color values can represent the saturation value of a specified color. In these embodiments, the video processing module 220 performs Equation 1 on each pixel within the region of interest to generate a grayscale value for each pixel, and thereby modifies the region of interest into a grayscale region of interest.
[0068] In various embodiments, the video processing module 220 identifies pixels in a grayscale region of interest and generates a histogram of values within the grayscale range. The histogram of values can indicate the distribution of intensity values associated with pixels in the grayscale region of interest.
[0069] In operation 620, video processing module 220 equalizes the histogram values within the grayscale region of interest to generate a balanced region of interest. Equalizing the histogram values results in increased contrast within the grayscale region of interest. For example, when the region of interest includes the mouth and exposed teeth, histogram equalization brightens pixels representing teeth (e.g., increases grayscale values) and darkens pixels representing lips, gums, and facial areas (e.g., decreases grayscale values). Video processing module 220 may utilize palette-changing or image-changing histogram equalization techniques to adjust the intensity values of pixels within the region of interest, increasing the overall contrast of the grayscale region of interest and generating a balanced region of interest.
[0070] In operation 630, video processing module 220 thresholds the balanced region of interest to generate a binarized region of interest. A first group of pixels within the binarized region of interest has a first value. A second group of pixels within the binarized region of interest has a second value that differs from the first value. For example, in some embodiments, the first value may be 1, and the second value may be 0. In some embodiments, video processing module 220 uses Otsu thresholding (e.g., the Otsu method) to threshold the balanced region of interest. In these embodiments, video processing module 220 performs clustering-based image thresholding on the balanced region of interest. The Otsu thresholding process causes video processing module 220 to generate a binarized region of interest that separates the first group of pixels from the second group of pixels. For example, after thresholding, the region of interest has a first group of pixels and a second group of pixels, the first group of pixels having a first value (e.g., a value of 1 indicating white pixels) and the second group of pixels having a second value (e.g., a value of 0 indicating black pixels). In performing Otsu thresholding, video processing module 220 may calculate a threshold to separate the first group of pixels and the second group of pixels. The video processing module 220 then applies a threshold to the intensity value (e.g., grayscale value) of the pixels to distinguish between the first group of pixels and the second group of pixels. If a pixel has an intensity value greater than the threshold, the video processing module 220 converts the pixel value to 1, generating a white pixel. If a pixel has an intensity value less than the threshold, the video processing module 220 converts the pixel value to 0, generating a black pixel. The resulting first group of pixels and second group of pixels form a binary region of interest.
[0071] In some embodiments, operations 610, 620, and 630 can be replaced by a collected binarization operation. The binarization operation computes a binarization matrix based on the pixel values within the region of interest (e.g., the binary value of each pixel within the region of interest or image). For example, where the pixel's color value consists of red, green, and blue values, the binarization operation computes the binarization matrix from the RGB values of the region of interest or image. In some cases, the binarization matrix is generated by marking the pixel's binary value as either 1 or 0. If... The binarization operation determines that binarization[i][j] = 1, otherwise binarization[i][j] = 0. In some cases, the video processing module 220 generates the binarization matrix based on the fact that the red value is greater than the sum of the green and blue values divided by a constant. Although specific examples of the equations for generating the binary pixel values of the binarization matrix have been described, it should be understood that different equations or standards can be used to generate the binarization matrix.
[0072] In operation 640, video processing module 220 identifies one or more pixels that have a first value and are located between two or more segments of a plurality of segments having a second value. In some cases, the second group of pixels includes a plurality of segments interrupted by one or more pixels 700 having the first value, such as... Figure 7 As shown. In some cases, an interruption in one or more pixels having a first value indicates noise, lighting defects, or other unexpected intersections between the first group of pixels and the second group of pixels. The video processing module 220 can identify one or more noise pixels based on the position of the noise pixel relative to other pixels with the same value and the number of pixels with the same value as the noise pixel.
[0073] For example, if the video processing module 220 identifies a group of potential noise pixels with a first value within a group of pixels with a second value, the video processing module 220 can determine the number of noise pixels. If the number of noise pixels is below a predetermined threshold (e.g., value, percentage of the region of interest, percentage of the original image), the video processing module 220 can determine whether the noise pixels are close to (e.g., adjacent to or connected to) a larger set of pixels in the first group.
[0074] In operation 650, the video processing module 220 replaces the first value of one or more identified pixels with a second value. If the video processing module 220 determines that a noise pixel is not connected to a larger set of the first group of pixels and is below a predetermined threshold, the video processing module 220 converts the value of the noise pixel to a second value that matches the second group of pixels.
[0075] In some embodiments, when performing operation 650, the video processing module 220 identifies a group of pixels within a first group of pixels having a number or proportion less than a predetermined error threshold. For example, if the pixels remaining in the first group of pixels after noise pixel conversion represent exposed teeth in a mouth, the video processing module 220 may identify a subset of the first group of pixels. The subset of the first group of pixels may be a group of pixels with a first value, separated from another subset by a portion of a second group of pixels. For example, if the binarized region of interest includes exposed teeth in a mouth as the object of interest, the subset of pixels is a group of pixels representing teeth (e.g., white pixels), separated from other teeth by lines or gaps (e.g., black pixels) in the second group of pixels. If the subset of the first group of pixels is below the error threshold, the subset is converted to a second value; otherwise, it is considered for removal by the image segmentation system 160. In some embodiments, the error threshold is a percentage of the original image region. In some cases, the error threshold is between 0% and 2% of the original image region. For example, if the region of a subset of the first set of pixels is less than 1% of the total image region, the video processing module 220 can convert the subset to a second value; otherwise, it will consider removing the subset.
[0076] Figure 8 A flowchart illustrating an example method 800 for tracking and modifying objects of interest in a video stream using an image segmentation system 160 is depicted. Operations of method 800 can be performed by components of the image segmentation system 160. In some cases, in one or more of the described embodiments, certain operations of method 800 can be performed using one or more operations of method 300 or 600, or as sub-operations of one or more operations of method 300 or 600, as explained in more detail below.
[0077] In various example embodiments, method 800 is initially performed by operations 310, 320, and 330. In some cases, operation 330 includes one or more operations from method 600.
[0078] In operation 810, the identification module 230 determines that the first group of pixels has a value equal to or greater than a predetermined threshold (e.g., a focus threshold). In some example embodiments, operation 810 may be performed as a sub-operation of operation 340. The identification module 230 may determine that the first group of pixels in the binarized region of interest generated in operation 330 or 630 has a value greater than the focus threshold. In some embodiments, the focus threshold is 0.5 or 1.
[0079] In operation 820, the recognition module 230 marks the first group of pixels within the second group of images in the video stream as generating a set of marker pixels 900, such as... Figure 9As shown in the diagram. This set of marked pixels are the pixels to be tracked within the second set of images in the video stream. Marked pixels can be identified and included in a set of points of interest to establish a relationship between the marked pixels and one or more orientation points. In some cases, marking the first set of pixels involves one or more sub-operations.
[0080] In operation 822, the recognition module 230 identifies the position of a first set of pixels relative to one or more landmarks representing the object of interest within the first set of images. In embodiments where the first set of pixels represents exposed teeth within a mouth depicted within a region of interest (e.g., a binary region of interest), the recognition module 230 identifies the position of the first set of pixels relative to facial recognition landmarks representing the mouth or the face in its vicinity. For example, the first set of pixels may be added as a set of landmark points on a binary mask of a face including a set of facial landmarks.
[0081] In operation 824, recognition module 230 identifies the color values of a first set of pixels within a first set of images. In embodiments where the first set of images has been marked for tracking (e.g., a set of landmark points within a set of reference landmarks identified as objects depicted in the first set of images), in a second set of images of the video stream, one or more of recognition module 230 and tracking module 250 identify corresponding pixels within one or more images of the second set of images that correspond to the first set of pixels. In some embodiments, recognition module 230 and tracking module 250 use one or more image tracking operations (such as a suitable set of operations for face tracking) to establish a correspondence. After establishing the corresponding pixels, recognition module 230 identifies one or more color values associated with the corresponding pixels. For example, recognition module 230 may identify a set of values for the red, green, and blue channels (e.g., red / green / blue (RGB) values) of the corresponding pixels.
[0082] In operation 830, tracking module 250 tracks a set of marker pixels across a second set of images in the video stream. In some embodiments, to track marker pixels, tracking module 250 tracks landmark points representing marker pixels or corresponding pixels in the second set of images. Tracking module 250 can use any suitable tracking operation to identify changes in the depiction (e.g., whether landmarks are occluded), position, and orientation of marker pixels in the images across the second set of images in the video stream. In embodiments where marker pixels represent exposed teeth, tracking module 250 detects changes in the position and orientation of teeth across the second set of images in the video stream. Tracking module 250 can also detect and track occlusion of teeth (e.g., marker pixels), where lips, hands, hair, tongue, or other occluders blur teeth and prevent teeth from being exposed in one or more images of the second set of images.
[0083] In operation 840, when a color value is present near a marker pixel, modification module 240 modifies the color value of the first group of pixels. In various embodiments, modification module 240 receives an indication from tracking module 250 of a corresponding pixel (e.g., a pixel within the image that is close to or at the location of a marker pixel or marker). As described above, regarding operation 350, modification module 240 may modify the color value of a marker pixel within an image in a second group of images, or modify the color by applying an image layer to an image within the second group of images.
[0084] If the tracking module 250 determines that a marker pixel is occluded, the modification module 240 does not modify the color of pixels within the second set of images or apply an image layer. For example, if the marker pixel represents exposed teeth inside the mouth, the tracking module 250 sends a color change instruction to the modification module 240 to trigger the modification module 240 to modify the color value of the corresponding pixel. If the tracking module 250 detects an occlusion of a marker pixel within the image, the tracking module 250 does not send an image color change instruction, or sends an instruction to the modification module 240 to wait or stop modifying the image, until the tracking module 250 detects a marker pixel in a subsequent image (e.g., detects that the occlusion no longer exists).
[0085] Example
[0086] To better illustrate the apparatus and methods disclosed herein, a non-limiting table of examples is provided herein:
[0087] 1. A computer-implemented method for manipulating a portion of a video stream, comprising: using one or more processors of a client device to determine a general location of a mouth within a video stream comprising a first set of images and a second set of images; identifying a region of interest within one or more images of the first set of images, the region of interest being a portion of the general location of the mouth surrounded by the one or more images; generating a modified region of interest by the client device; identifying a first set of pixels and a second set of pixels within the modified region of interest by the client device; and modifying color values of the first set of pixels within the second set of images of the video stream by the client device.
[0088] 2. According to the method of Example 1, wherein the method further includes cropping one or more portions of one or more images of a first set of images outside the region of interest.
[0089] 3. According to the method of Example 1 or 2, wherein generating the modified region of interest further includes converting the pixels within the region of interest into grayscale values to generate a grayscale region of interest.
[0090] 4. Based on any one or more of the methods in Examples 1-3, wherein the pixels within the grayscale region of interest include a histogram of values within the grayscale, and generating the modified region of interest further includes balancing the histogram values within the grayscale region of interest to generate a balanced region of interest.
[0091] 5. According to any one or more of the methods in Examples 1-4, wherein generating the modified region of interest further includes thresholding the balanced region of interest to generate a binarized region of interest, wherein a first group of pixels within the binarized region of interest has a first value, and a second group of pixels within the binarized region of interest has a second value different from the first value.
[0092] 6. The method according to any one or more of Examples 1-5, wherein the second group of pixels comprises a plurality of segments interrupted by one or more pixels having a first value, and generating the binarized region of interest further comprises: identifying one or more pixels having a first value located between two or more segments of a plurality of segments having a second value; and replacing the first value of the identified one or more pixels with the second value.
[0093] 7. According to any one or more of the methods in Examples 1-6, wherein identifying the first group of pixels and the second group of pixels within the modified region of interest further includes: determining that the first group of pixels has a value greater than a predetermined threshold; and marking the first group of pixels within the second group of images of the video stream to generate a set of marked pixels.
[0094] 8. According to any one or more of the methods in Examples 1-7, wherein marking the first group of pixels within the second group of images further includes: identifying the position of the first group of pixels relative to one or more landmarks of the face depicted within the first group of images; and identifying the color value of the first group of pixels within the first group of images.
[0095] 9. The method according to any one or more of Examples 1-8, wherein modifying the color value of the first group of pixels within the second group of images further includes: tracking a set of marker pixels across the second group of images in the video stream; and modifying the color value of the first group of pixels when the color value is rendered close to the marker pixels.
[0096] 10. A system for manipulating a portion of a video stream, comprising: one or more processors; and a machine-readable storage medium carrying processor-executable instructions that, when executed by the processor of the machine, cause the machine to perform operations including: determining a general location of a mouth within a video stream comprising a first set of images and a second set of images; identifying a region of interest within one or more images of the first set of images, the region of interest being a portion of the general location of the mouth enclosing the one or more images; generating a modified region of interest; identifying a first set of pixels and a second set of pixels within the modified region of interest; and modifying the color values of the first set of pixels within the second set of images of the video stream.
[0097] 11. According to the system of Example 10, generating the modified region of interest causes the machine to perform the following operations: converting the pixels within the region of interest into grayscale values to generate a grayscale region of interest.
[0098] 12. A system according to Example 10 or 11, wherein the pixels in the grayscale region of interest include a histogram of values within the grayscale, and generating a modified region of interest causes the machine to perform the following operation: equalizing the histogram values in the grayscale region of interest to generate a balanced region of interest.
[0099] 13. A system according to any one or more of Examples 10-12, wherein generating a modified region of interest causes the machine to perform the following operations: thresholding the balanced region of interest to generate a binarized region of interest, wherein a first group of pixels within the binarized region of interest has a first value, and a second group of pixels within the binarized region of interest has a second value different from the first value.
[0100] 14. A system according to any one or more of Examples 10-13, wherein the second group of pixels comprises a plurality of segments interrupted by one or more pixels having a first value, and a binarized region of interest is generated such that the machine performs the following operations: identifying one or more pixels having a first value located between two or more segments of a plurality of segments having a second value; and replacing the first value of the identified one or more pixels with the second value.
[0101] 15. A system according to any one or more of Examples 10-14, wherein identifying a first group of pixels and a second group of pixels within a modified region of interest causes the machine to perform operations including: determining that the first group of pixels has a value greater than a predetermined threshold; and marking the first group of pixels within a second group of images of the video stream to generate a set of marked pixels.
[0102] 16. A system according to any one or more of Examples 10-15, wherein marking a first group of pixels within a second group of images causes the machine to perform operations including: identifying the position of the first group of pixels relative to one or more landmarks of a face depicted within the first group of images; and identifying the color value of the first group of pixels within the first group of images.
[0103] 17. A system according to any one or more of Examples 10-16, wherein modifying the color values of the first set of pixels within the second set of images causes the machine to perform operations including: tracking a set of marked pixels across the second set of images in the video stream; and modifying the color values of the first set of pixels when the color values are rendered close to the marked pixels.
[0104] 18. A machine-readable storage medium carrying processor-executable instructions that, when executed by a processor of a machine, cause the machine to perform operations including: determining the approximate location of a mouth within a video stream comprising a first set of images and a second set of images; identifying a region of interest within one or more images of the first set of images, the region of interest being a portion of the approximate location of the mouth enclosing the one or more images; generating a modified region of interest; identifying a first set of pixels and a second set of pixels within the modified region of interest; and modifying the color values of the first set of pixels within the second set of images of the video stream.
[0105] 19. A machine-readable storage medium according to Example 18, wherein identifying a first group of pixels and a second group of pixels within a modified region of interest causes the machine to perform operations including: determining that the first group of pixels has a value within the modified region of interest that is greater than a predetermined threshold; and marking the first group of pixels within a second group of images of a video stream to generate a set of marked pixels.
[0106] 20. A machine-readable storage medium according to Example 18 or 19, wherein modifying the color values of a first set of pixels within a second set of images causes the machine to perform operations including: identifying the position of the first set of pixels relative to one or more landmarks of a face depicted within the first set of images; identifying the color values of the first set of pixels within the first set of images; tracking a set of marker pixels across a second set of images in a video stream; and modifying the color values of the first set of pixels when the color values are presented close to the marker pixels.
[0107] These and other examples and features of the apparatus and method have been partially set forth in the foregoing detailed description. The summary and embodiments are intended to provide non-limiting examples of the subject matter. They are not intended to provide an exclusive or exhaustive explanation. Detailed descriptions are included to provide further information regarding the subject matter.
[0108] Modules, components and logic
[0109] Some embodiments are described herein as including logic or multiple components, modules, or mechanisms. Modules can constitute software modules (e.g., code embodied in a machine-readable medium) or hardware modules. A “hardware module” is a tangible unit capable of performing certain operations and which can be configured or arranged in some physical manner. In various exemplary embodiments, a computer system (e.g., a standalone computer system, a client computer system, or a server computer system) or a hardware module of a computer system (e.g., at least one hardware processor, processor, or a group of processors) can be configured by software (e.g., an application or a portion of an application) to operate as a hardware module performing certain operations as described herein.
[0110] In some embodiments, the hardware module may be implemented mechanically, electronically, or any suitable combination thereof. For example, the hardware module may include dedicated circuitry or logic permanently configured to perform certain operations. For example, the hardware module may be a dedicated processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). The hardware module may also include programmable logic or circuitry temporarily configured by software to perform certain operations. For example, the hardware module may include software included within a general-purpose processor or other programmable processor. It should be understood that the decision to implement the hardware module mechanically in dedicated and permanently configured circuitry or in temporarily configured circuitry (e.g., configured by software) may be driven by cost and time considerations.
[0111] Therefore, the phrase "hardware module" should be understood to include tangible entities, i.e., physical structures, permanent configurations (e.g., hardwired) or temporary configurations (e.g., programming), entities that operate or perform certain operations described herein in some way. As used herein, "hardware-implemented module" refers to a hardware module. Considering embodiments in which hardware modules are temporarily configured (e.g., programmed), each hardware module in the hardware module does not need to be configured or instantiated at any one time. For example, in the case where the hardware module includes a general-purpose processor configured by software as a dedicated processor, the general-purpose processor can be configured as correspondingly different dedicated processors (e.g., including different hardware modules) at different times. Accordingly, the software can configure a particular processor, for example, constituting a particular hardware module at one point in time and different hardware modules at different points in time.
[0112] Hardware modules can provide information to and receive information from other hardware modules. Therefore, the described hardware modules can be considered communication-coupled. When multiple hardware modules exist simultaneously, communication can be achieved through signal transmission (e.g., via appropriate circuitry and buses) between or among two or more hardware modules. In embodiments where multiple hardware modules are configured or instantiated at different times, such communication between hardware modules can be achieved, for example, through the storage and retrieval of information in a memory structure accessed by the multiple hardware modules. For example, one hardware module can perform an operation and store the output of that operation in a memory device to which it is communication-coupled. Another hardware module can then access the memory device at a later time to retrieve and process the stored output. Hardware modules can also initiate communication with input or output devices and can operate on resources (e.g., collections of information).
[0113] The various operations of the example methods described herein are performed at least in part by a processor configured, either temporarily (e.g., by software) or permanently, to perform the relevant operations. Whether temporarily or permanently configured, such a processor can constitute a processor-implemented module that operates to perform the operations or functions described herein. As used herein, "processor-implemented module" means a hardware module implemented using a processor.
[0114] Similarly, the methods described herein may be at least partially implemented by a processor, where a particular processor or processors are examples of hardware. For example, at least some operations of the method may be performed by a processor or a module implemented by a processor. Furthermore, the processor may also be operable to support the performance of related operations in a “cloud computing” environment or as “Software as a Service” (SaaS). For example, at least some operations may be performed by a set of computers (as an example of a machine including a processor), which may be accessible via a network (e.g., the Internet) and via a suitable interface (e.g., an application programming interface, API).
[0115] The performance of certain operations can be distributed across processors, residing not only within a single machine but also deployed across multiple machines. In some example embodiments, the processor or processor-implemented module may reside in a single geographic location (e.g., in a home environment, office environment, or server farm). In other example embodiments, the processor or processor-implemented module may be distributed across multiple geographic locations.
[0116] application
[0117] Figure 10An example mobile device 1000 executing a mobile operating system (e.g., iOS™, Android™, Windows® Phone, or other mobile operating systems) is shown, consistent with some embodiments. In one embodiment, the mobile device 1000 includes a touchscreen operable to receive tactile data from a user 1002. For example, the user 1002 may physically touch 1004 of the mobile device 1000, and in response to the touch 1004, the mobile device 1000 may determine tactile data such as touch location, touch force, or gesture. In various example embodiments, the mobile device 1000 displays a home screen 1006 (e.g., Springboard on iOS™), operable to launch applications or otherwise manage various aspects of the mobile device 1000. In some example embodiments, the home screen 1006 provides status information such as battery life, connectivity, or other hardware status. The user 1002 can activate user interface elements by touching areas occupied by corresponding user interface elements. In this way, the user 1002 interacts with applications on the mobile device 1000. For example, touching the area occupied by a specific icon included in the home screen 1006 causes the application corresponding to that specific icon to be launched.
[0118] like Figure 10 As shown, the mobile device 1000 may include an imaging device 1008. The imaging device may be a camera or any other device coupled to the mobile device 1000 capable of acquiring a video stream or one or more consecutive images. The imaging device 1008 may be triggered by the image segmentation system 160 or optional user interface elements to initiate the acquisition of the video stream or consecutive images and to pass the video stream or consecutive images to the image segmentation system for processing according to one or more methods described in this disclosure.
[0119] Many kinds of applications (also referred to as "application software") can be executed on mobile device 1000, such as native applications (e.g., applications running on iOS™ programmed in Objective-C, Swift, or another suitable language, or applications running on Android™ programmed in Java), mobile web applications (e.g., applications written in Hypertext Markup Language-5 (HTML5)), or hybrid applications (e.g., native shell applications that launch HTML5 sessions). For example, mobile device 1000 includes messaging applications, audio recording applications, camera applications, book reader applications, media applications, fitness applications, file management applications, location applications, browser applications, settings applications, contact applications, phone calling applications, or other applications (e.g., game applications, social networking applications, biometric monitoring applications). In another example, mobile device 1000 includes social messaging application software 1010 such as SNAPCHAT®, which, consistent with some embodiments, allows users to exchange short messages including media content. In this example, social messaging application software 1010 may incorporate aspects of the embodiments described herein. For example, in some embodiments, the social messaging application includes short-lived media galleries created by the user's social messaging application. These galleries may consist of videos or pictures posted by the user and viewable by the user's contacts (e.g., "friends"). Alternatively, public galleries may be created by an administrator of the social messaging application, which consists of media from any user of the application (and accessible to all users). In yet another embodiment, the social messaging application may include a "magazine" feature, comprising articles and other content generated by publishers on the social messaging application's platform and accessible to any user. Any of these environments or platforms can be used to implement the concepts of the present invention.
[0120] In some embodiments, the short message delivery system may include a message containing a short video clip or image that is deleted after a deletion trigger event, such as viewing time or viewing completion. In this embodiment, while a short video clip is being captured by a device and sent to another device using the short message delivery system, the device implementing the image segmentation system 160 can identify, track, and modify objects of interest within the short video clip.
[0121] Software Architecture
[0122] Figure 11 This is a block diagram 1100 showing a software architecture 1102 that can be installed on the device described above. Figure 11This is merely a non-limiting example of a software architecture, and it should be understood that many other architectures can be implemented to achieve the functionality described herein. In various embodiments, software 1102 can be composed of, for example... Figure 12 The hardware execution of machine 1200 includes a processor 1210, memory 1230, and I / O components 1250. In this example architecture, software 1102 can be conceptualized as a stack of layers, where each layer provides a specific function. For example, software 1102 includes layers such as an operating system 1104, libraries 1106, frameworks 1108, and applications 1110. Operationally, consistent with some embodiments, application 1110 invokes application programming interface (API) calls 1112 through the software stack and receives messages 1114 in response to API call 1112.
[0123] In various implementations, operating system 1104 manages hardware resources and provides public services. Operating system 1104 includes, for example, a kernel 1120, services 1122, and drivers 1124. Consistent with some embodiments, kernel 1120 serves as an abstraction layer between hardware and other software layers. For example, kernel 1120 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. Services 1122 can provide other public services to other software layers. According to some embodiments, driver 1124 is responsible for controlling or interfaced with the underlying hardware. For example, driver 1124 may include a display driver, camera driver, Bluetooth® driver, flash memory driver, serial communication driver (e.g., Universal Serial Bus (USB) driver), Wi-Fi® driver, audio driver, power management driver, and so on.
[0124] In some embodiments, library 1106 provides low-level public infrastructure that can be utilized by application 1110. Library 1106 may include system libraries 1130 (e.g., the C standard library) that provide functions such as memory allocation, string manipulation, and mathematical functions. Furthermore, library 1106 may include media libraries (e.g., libraries supporting the rendering and manipulation of various media formats such as Moving Picture Experts Group 4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer 3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) Audio Coding, Joint Picture Experts Group (JPEG or JPG), or Portable Web Graphics (PNG),) graphics libraries (e.g., OpenGL frameworks for displaying two-dimensional (2D) and three-dimensional (3D) graphics content on a display), database libraries (e.g., SQLite providing various relational database functions), web libraries (e.g., WebKit providing web browsing capabilities), and API libraries 1132. Library 1106 may also include a variety of other libraries 1134 to provide numerous other APIs to application 1110.
[0125] According to some embodiments, framework 1108 provides advanced public infrastructure that can be utilized by application 1110. For example, framework 1108 provides various graphical user interface (GUI) functions, advanced resource management, advanced location services, and so on. Framework 1108 can provide a wide range of other APIs that can be utilized by application 1110, some of which may be targeted to a specific operating system or platform.
[0126] In an example embodiment, application 1110 includes a home application 1150, a contacts application 1152, a browser application 1154, a book reader application 1156, a location application 1158, a media application 1160, a messaging application 1162, a game application 1164, and other applications in a broad category (such as third-party applications 1166). According to some embodiments, application 1110 is a program that performs functions defined in the program. Application 1110 can be created using various programming languages in various ways, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a particular example, third-party application 1166 (e.g., an application developed by an entity other than a platform-specific vendor using the Android™ or iOS™ Software Development Kit (SDK)) can be mobile software running on a mobile operating system (such as iOS™, Android™, Windows® Phone, or other mobile operating systems). In this example, third-party application 1166 may invoke API calls 1112 provided by operating system 1104 to facilitate the implementation of the functions described herein.
[0127] Example machine architecture and machine-readable media
[0128] Figure 12 This is a block diagram illustrating components of a machine 1200, according to some example embodiments, capable of reading instructions (e.g., processor-executable instructions) from a machine-readable medium (e.g., a non-transitory machine-readable storage medium) and performing any one or more methods discussed herein. Specifically, Figure 12A schematic diagram of machine 1200 in the form of an example computer system is shown, wherein instructions 1216 (e.g., software, programs, applications, applets, application software, or other executable code) for causing machine 1200 to perform any one or more methods discussed herein can be executed. In alternative embodiments, machine 1200 operates as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 1200 can operate as a server machine or client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1200 can be, but is not limited to, server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart home appliances), other smart devices, network devices, network routers, network switches, network bridges, or any machine capable of executing instructions 1216, which sequentially or otherwise specify the actions to be taken by machine 1200. Furthermore, although only a single machine 1200 is shown, the term "machine" should also be understood to include a collection of machines 1200 that individually or in combination execute instructions 1216 to perform any of the methods discussed herein.
[0129] In various embodiments, machine 1200 includes processor 1210, memory 1230, and I / O components 1250, which may be configured to communicate with each other via bus 1202. In example embodiments, processor 1210 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) includes, for example, processors 1212 and 1214 capable of executing instructions 1216. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Figure 12 Multiple processors are shown, but machine 1200 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0130] According to some embodiments, memory 1230 includes main memory 1232, static memory 1234, and storage cell 1236, which are accessible via bus 1202 to processor 1210. Storage cell 1236 may include machine-readable medium 1238 on which instructions 1216 embodying any methods or functions described herein are stored. Instructions 1216 may also reside wholly or at least partially within main memory 1232, static memory 1234, at least one of processor 1210 (e.g., within the processor's cache memory), or any suitable combination thereof during execution by machine 1200. Therefore, in various embodiments, main memory 1232, static memory 1234, and processor 1210 are considered to be machine-readable medium 1238.
[0131] As used herein, the term "memory" refers to machine-readable medium 1238 capable of temporarily or permanently storing data and may be considered to include, but is not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, and cache memory. While machine-readable medium 1238 is shown as a single medium in the example embodiment, the term "machine-readable medium" should be understood to include a single medium or multiple media (e.g., a centralized or distributed database or associated cache and server). The term "machine-readable medium" should also be understood to include any medium or combination of media (e.g., machine 1200) capable of storing instructions (e.g., instruction 1216) that, when executed by a processor (e.g., processor 1210) of machine 1200, cause machine 1200 to perform any of the methods described herein. Accordingly, "machine-readable medium" refers to a single storage device or apparatus, as well as a "cloud-based" storage system or storage network that includes multiple storage devices or apparatuses. Accordingly, the term "machine-readable medium" should be understood to include, but is not limited to, solid-state memory (e.g., flash memory), optical media, magnetic media, other non-volatile memory (e.g., erasable programmable read-only memory (EPROM)) or any suitable combination thereof as a data repository. The term "machine-readable medium" itself specifically excludes non-legal signals themselves.
[0132] I / O component 1250 may include various components for receiving input, providing output, generating output, sending information, exchanging information, acquiring measurements, etc. Generally, it should be understood that I / O component 1250 may include components not in... Figure 12Many other components are shown. The I / O components 1250 are grouped by function only to simplify the following discussion, and the grouping is not limiting. In various example embodiments, the I / O components 1250 may include output components 1252 and input components 1254. Output components 1252 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tube (CRT) displays), acoustic components (e.g., speakers), haptic components (e.g., vibration motors), other signal generators, etc. Input components 1254 include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other indicating instruments), haptic input components (e.g., physical buttons, touchscreens providing the location and force of a touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc.
[0133] In some other example embodiments, I / O component 1250 includes a wide range of other components, such as biometric component 1256, motion component 1258, environmental component 1260, or position component 1262. For example, biometric component 1256 may include components for detecting expressions (e.g., hand gestures, facial expressions, vocal expressions, body posture, or mouth gestures), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 1258 includes accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), and so on. Environmental component 1260 includes, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., a thermometer for detecting ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., a microphone for detecting background noise), a proximity sensor component (e.g., an infrared sensor for detecting nearby objects), a gas sensor component (e.g., a machine olfactory detection sensor, a gas detection sensor for detecting the concentration of harmful gases for safety purposes or for measuring pollutants in the atmosphere), or other components that may provide indications, measurements, or signals corresponding to the surrounding physical environment. Position component 1262 includes a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects from which altitude air pressure can be derived), an orientation sensor component (e.g., a magnetometer), etc.
[0134] Various technologies can be used to implement communication. I / O component 1250 may include communication component 1264, operable to couple machine 1200 to network 1280 or device 1270 via couplings 1282 and 1272, respectively. For example, communication component 1264 includes a network interface component or other suitable device that interfaces with network 1280. In another example, communication component 1264 includes wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components that provide communication via other modes. Device 1270 may be another machine or any variety of peripheral devices (e.g., peripheral devices coupled via USB).
[0135] Furthermore, the communication component 1264 can detect identifiers or include components operable to detect identifiers. For example, the communication component 1264 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) codes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, data matrices, data icons, maximum codes, PDF417, supercodes, Uniform Commercial Code Shrinkable Space Symbol System (UCC RSS-2D barcodes and other optical codes)), an acoustic detection component (a microphone for identifying audio signals from the tag), or any suitable combination thereof. Additionally, various information can be derived via the communication component 1264, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, location via detection of BLUETOOTH® or NFC beacon signals that can indicate a specific location, etc.
[0136] transmission medium
[0137] In various example embodiments, a portion of network 1280 may be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), World Wide Web (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a common old-style telephone service (POTS) network, a cellular telephone network, a wireless network, a Wi-Fi® network, another type of network, or a combination of two or more such networks. For example, network 1280 or a portion of network 1280 may include a wireless or cellular network, and coupling 1282 may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or other types of cellular or wireless coupling. In this example, the Coupler 1282 can implement any of various types of data transmission technologies, such as Single Carrier Radio Transmission (1xRTT), Evolved Data Optimization (EVDO), General Packet Radio Service (GPRS), Evolution Enhanced Data Rate (EDGE) for GSM, the 3rd Generation Partnership Project (3GPP) including 3G, 4G networks, Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Global Microwave Access Interoperability (WiMAX), Long Term Evolution (LTE) standards, other standards defined by various standards-setting organizations, other remote protocols, or other data transmission technologies.
[0138] In an example embodiment, instructions 1216 may be sent or received over network 1280 via a transmission medium using a network interface device (e.g., a network interface component included in communication component 1264) and utilizing any of a plurality of known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, in other example embodiments, instructions 1216 may be sent or received over a transmission medium using a coupling 1272 (e.g., peer-to-peer coupling). The term “transmission medium” should be considered to include any intangible medium capable of storing, encoding, or carrying instructions 1216 for execution by machine 1200, and includes digital or analog communication signals or other intangible media that facilitate such software communication.
[0139] Furthermore, the machine-readable medium 1238 is non-transient (in other words, it does not have any transient signals) because it does not contain propagating signals. However, labeling the machine-readable medium 1238 as "non-transient" should not be interpreted as meaning that the medium cannot be moved; the medium should be considered as being movable from one physical location to another. Additionally, since the machine-readable medium 1238 is tangible, it can be considered a machine-readable device.
[0140] language
[0141] Throughout this specification, multiple instances can implement components, operations, or structures described as single instances. While individual operations of methods are shown and described as separate operations, these individual operations can be performed simultaneously and do not need to be performed in the order shown. Structures and functionalities presented as individual components in the example configuration can be implemented as composite structures or components. Similarly, structures and functionalities presented as single components can be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of this document's subject matter.
[0142] While an overview of the subject matter of the invention has been described with reference to specific exemplary embodiments, various modifications and changes may be made to these embodiments without departing from the broader scope of the embodiments disclosed herein. Such embodiments of the subject matter of the invention may be referred to individually or collectively by the term "invention" and are intended merely for convenience and not to limit the scope of this application to any single disclosure or inventive concept, if more than one is disclosed in fact.
[0143] The embodiments shown herein are described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments may be used and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Therefore, the specific implementation should not be considered limiting, and the scope of the various embodiments is defined only by the appended claims and the full scope of their equivalents.
[0144] As used herein, the term "or" may be interpreted in an inclusive or exclusive manner. Furthermore, multiple instances of the resources, operations, or structures described herein may be provided as a single instance. Moreover, the boundaries between various resources, operations, modules, engines, and data stores are somewhat arbitrary, and specific operations are shown in the context of a particular illustrative configuration. Other allocations of functionality are contemplated and may fall within the scope of various embodiments of this disclosure. Typically, structures and functions presented as individual resources in example configurations may be implemented as combined structures or resources. Similarly, structures and functions presented as single resources may be implemented as separate resources. These and other variations, modifications, additions, and improvements fall within the scope of embodiments of this disclosure as represented by the appended claims. Therefore, the specification and drawings are to be considered illustrative rather than restrictive.
Claims
1. A computer-implemented method for manipulating a portion of a video stream, comprising: The approximate location of the mouth within the video stream is determined by the client device; the video stream includes a face and comprises a first set of images and a second set of images. The client device identifies a region of interest, which includes multiple pixels within one or more images in the first set of images, and the region of interest is the portion of the one or more images that surrounds the approximate location of the mouth. The client device generates a binarized matrix by performing the following operations: For each of the plurality of pixels within the region of interest: Obtain a set of color values associated with the pixel; The binary value of the pixel is determined by comparing the first value of the first part of the set of color values with the second value of the second part of the set of color values. The binary value of the pixel is stored in the binarization matrix; The binarization matrix is used to modify each of the plurality of pixels within the region of interest to create a binarized region of interest. The client device identifies a set of teeth visible within the mouth in the binarized region of interest; The client device identifies a first group of pixels and a second group of pixels within the binarized region of interest, and adds at least a portion of the first group of pixels as a set of landmarks within the binary mask of the face, wherein the first group of pixels corresponds to the set of teeth within the mouth. as well as When the group of teeth is visible within the second group of images, the client device modifies the color value of the first group of pixels within the second group of images of the video stream.
2. The method according to claim 1, further comprising: The plurality of pixels within the region of interest are converted to grayscale values to generate a grayscale region of interest. The plurality of pixels are converted to grayscale by the following method: for each pixel, a grayscale value is generated by multiplying a fixed intensity value by the quotient of a triplet of the pixel’s color saturation value.
3. The method according to claim 2, wherein, The pixels within the grayscale region of interest include histogram values within the grayscale, and the method further includes: The histogram values within the grayscale region of interest are balanced to generate a balanced region of interest.
4. The method of claim 3, further comprising: The balanced region of interest is thresholded to generate a binarized region of interest, wherein the first group of pixels within the binarized region of interest has a first value, and the second group of pixels within the binarized region of interest has a second value that is different from the first value.
5. The method according to claim 4, wherein, The second group of pixels includes multiple segments interrupted by one or more pixels having the first value, and generating the binarized region of interest further includes: Identify one or more pixels having the first value, located between two or more segments of the plurality of segments having the second value; and The first value of one or more identified pixels is replaced with the second value.
6. The method of claim 4, further comprising: Determine that the first group of pixels has a value greater than a predetermined threshold; as well as The first group of pixels is labeled within the second group of images in the video stream to generate a set of labeled pixels.
7. The method according to claim 6, wherein, Marking the first group of pixels within the second group of images further includes: Identify the position of the first group of pixels relative to one or more landmarks of the face depicted within the first group of images; and Identify the color values of the first group of pixels within the first group of images.
8. The method according to claim 6, wherein, Modifying the color values of the first group of pixels within the second group of images further includes: Tracking the set of marked pixels of the second set of images across the video stream; and When the color value is presented close to the marked pixel, the color value of the first group of pixels is modified.
9. A system for manipulating a portion of a video stream, comprising: One or more processors; as well as A non-transitory machine-readable storage medium storing processor-executable instructions that, when executed by a machine's processor, cause the machine to perform operations including: The approximate location of the mouth within the video stream is determined by the client device; the video stream includes a face and comprises a first set of images and a second set of images. The client device identifies a region of interest, which includes multiple pixels within one or more images in the first set of images, and the region of interest is the portion of the one or more images that surrounds the approximate location of the mouth. The client device generates a binarized matrix by performing the following operations: For each of the plurality of pixels within the region of interest: Obtain a set of color values associated with the pixel; The binary value of the pixel is determined by comparing the first value of the first part of the set of color values with the second value of the second part of the set of color values. The associated binary value of the pixel is stored in the binarization matrix; The binarization matrix is used to modify each of the plurality of pixels within the region of interest to create a binarized region of interest. The client device identifies a set of teeth visible within the mouth in the binarized region of interest; The client device identifies a first group of pixels and a second group of pixels within the binarized region of interest, and adds at least a portion of the first group of pixels as a set of landmarks within the binary mask of the face, wherein the first group of pixels corresponds to the set of teeth within the mouth. as well as When the group of teeth is visible within the second group of images, the client device modifies the color value of the first group of pixels within the second group of images of the video stream.
10. The system according to claim 9, further comprising: The plurality of pixels within the region of interest are converted to grayscale values to generate a grayscale region of interest. The plurality of pixels are converted to grayscale by the following method: for each pixel, a grayscale value is generated by multiplying a fixed intensity value by the quotient of a triplet of the pixel’s color saturation value.
11. The system according to claim 10, wherein, The pixels within the grayscale interest region include the histogram values within the grayscale, and further include: The histogram values within the grayscale region of interest are balanced to generate a balanced region of interest.
12. The system of claim 11, further comprising: The balanced region of interest is thresholded to generate a binarized region of interest, wherein the first group of pixels within the binarized region of interest has a first value, and the second group of pixels within the binarized region of interest has a second value that is different from the first value.
13. The system according to claim 12, wherein, The second group of pixels includes multiple segments interrupted by one or more pixels having the first value, and generating the binarized region of interest further includes: Identify one or more pixels having the first value, located between two or more segments of the plurality of segments having the second value; and The first value of one or more identified pixels is replaced with the second value.
14. The system of claim 12, further comprising: Determine that the first group of pixels has a value greater than a predetermined threshold; as well as The first group of pixels is labeled within the second group of images in the video stream to generate a set of labeled pixels.
15. The system according to claim 14, wherein, Marking the first group of pixels within the second group of images further includes: Identify the position of the first group of pixels relative to one or more landmarks of the face depicted within the first group of images; and Identify the color values of the first group of pixels within the first group of images.
16. The system according to claim 14, wherein, Modifying the color values of the first group of pixels within the second group of images further includes: Tracking the set of marked pixels of the second set of images across the video stream; and When the color value is presented close to the marked pixel, the color value of the first group of pixels is modified.
17. A non-transitory machine-readable storage medium storing processor-executable instructions, which, when executed by a processor of a machine, cause the machine to perform operations including: The approximate location of the mouth within the video stream is determined by the client device; the video stream includes a face and comprises a first set of images and a second set of images. The client device identifies a region of interest, which includes multiple pixels within one or more images in the first set of images, and the region of interest is the portion of the one or more images that surrounds the approximate location of the mouth. The client device generates a binarized matrix by performing the following operations: For each of the plurality of pixels within the region of interest: Obtain a set of color values associated with the pixel; The binary value of the pixel is determined by comparing the first value of the first part of the set of color values with the second value of the second part of the set of color values. The associated binary value of the pixel is stored in the binarization matrix; The binarization matrix is used to modify each of the plurality of pixels within the region of interest to create a binarized region of interest. The client device identifies a set of teeth visible within the mouth in the binarized region of interest; The client device identifies a first group of pixels and a second group of pixels within the binarized region of interest, and adds at least a portion of the first group of pixels as a set of landmarks within the binary mask of the face, wherein the first group of pixels corresponds to the set of teeth within the mouth. as well as When the group of teeth is visible within the second group of images, the client device modifies the color value of the first group of pixels within the second group of images of the video stream.
18. The non-transitory machine-readable storage medium of claim 17, further comprising: The plurality of pixels within the region of interest are converted to grayscale values to generate a grayscale region of interest. The plurality of pixels are converted to grayscale by the following method: for each pixel, a grayscale value is generated by multiplying a fixed intensity value by the quotient of a triplet of the pixel’s color saturation value.
19. The non-transitory machine-readable storage medium according to claim 18, wherein, The pixels within the grayscale interest region include the histogram values within the grayscale, and further include: The histogram values within the grayscale region of interest are balanced to generate a balanced region of interest.
20. The non-transitory machine-readable storage medium of claim 19, further comprising: The balanced region of interest is thresholded to generate a binarized region of interest, wherein the first group of pixels within the binarized region of interest has a first value, and the second group of pixels within the binarized region of interest has a second value that is different from the first value.
Citation Information
Patent Citations
Image processing apparatus and method thereof
CN101206761A
Locating and augmenting object features in images
WO2015015173A1