Repair based on real-time machine learning
Patent Information
- Application Number
- CN202380069936.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-03
- Filing Date
- 2023-09-27
- Publication Date
- 2025-05-13
Smart Images

Figure CN119998828A_ABST
Abstract
Description
[0001] Priority declaration
[0002] This application claims the benefit of U.S. patent application serial number 17937,734, filed on October 3, 2022, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to providing an augmented reality (AR) experience using a messaging application. Background Art
[0004] AR is a modification of a virtual environment. For example, in virtual reality (VR), the user is completely immersed in a virtual world, while in AR, the user is immersed in a world that combines or superimposes virtual objects on the real world. AR systems are designed to generate and present virtual objects that realistically interact with the real world environment and with each other. Examples of AR applications can include single-player or multi-player video games, instant messaging systems, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] In the drawings, which are not necessarily drawn to scale, like reference numerals may describe similar components in different views. To easily identify the discussion of any particular element or action, the most significant digit or digits in the reference numeral refer to the figure number in which the element is first introduced. Some non-limiting examples are shown in the figures of the accompanying drawings, in which:
[0006] Figure 1 is a diagrammatic representation of a networking environment in which the present disclosure may be deployed, according to some examples.
[0007] Figure 2 is a diagrammatic representation of a messaging client application according to some examples.
[0008] Figure 3 is a diagrammatic representation of data structures maintained in a database according to some examples.
[0009] Figure 4 is a graphical representation of messages according to some examples.
[0010] Figure 5 is a block diagram illustrating an example AR in-painting system according to some examples.
[0011] Figure 6 is a block diagram illustrating a more detailed example of an AR repair system according to some examples.
[0012] Figure 7 and Figure 8 is a graphical representation of the output of an AR inpainting system according to some examples.
[0013] Fig. 9 is a flow chart illustrating example operation of an AR repair system according to an example.
[0014] Fig.10 is a diagrammatic representation of a machine in the form of a computer system according to some examples within which a set of instructions may be executed for causing the machine to perform any one or more of the methodologies discussed herein.
[0015] Fig.11 is a block diagram illustrating a software architecture in which examples may be implemented. DETAILED DESCRIPTION
[0016] The following description includes systems, methods, techniques, instruction sequences, and computer program products that embody illustrative examples of the present disclosure. In the following description, for the purpose of illustration, many specific details are set forth to provide an understanding of various examples. However, it will be apparent to those skilled in the art that examples can be practiced without these specific details. Typically, well-known instruction instances, protocols, structures, and techniques are not necessarily shown in detail.
[0017] Typically, VR systems and AR systems allow users to add AR elements in places or locations that may interfere with real-world objects depicted in the video. For example, a user may desire to place an AR coffee table in a video that already includes a coffee table. However, the user is limited to placing the AR coffee table in an area with free space; otherwise, the AR coffee table will overlap with real-world objects such as the real-world coffee table. The overlap of the AR coffee table with real-world objects destroys the illusion that the AR coffee table looks as if it is part of the real-world environment. In addition, if the AR coffee table is placed on top of a real-world coffee table, it may be difficult to fully understand what it would look like if the AR coffee table replaced the real-world coffee table. This severely limits the functionality of typical AR systems and reduces the overall interest in using these systems.
[0018] Some systems perform inpainting tasks to remove portions or areas of an image or video and reconstruct the missing areas in the image or video. This allows the user to place an AR object in a portion of the video that has now been modified to exclude certain real-world objects. In such systems, pixels in the image are synthesized based on a hole mask. These systems cannot be used for real-time video to modify the video as it is captured. In addition, when applied to previously captured video, these systems may remove objects and present blurred areas in the portion of the video from which the object has been removed. This destroys the illusion that the real-world object has been removed and reduces the overall use and enjoyment of the system.
[0019] The disclosed technology improves the efficiency of using electronic devices that implement or otherwise access AR / VR systems by intelligently and automatically modifying real-time video to remove one or more real-world objects in a realistic manner. Specifically, the disclosed technology applies a machine learning model such as a generative adversarial network to the video. The machine learning model can generate a new frame based on a previously captured video frame and a current frame of the video, in which certain real-world objects are removed in real time as the video is captured. The machine learning model is trained to learn the texture of the video and can complete different areas in a convincing manner. Then, one or more AR objects can be placed in the area from which the real-world objects have been removed. This improves the overall user experience and enhances the illusion that the AR elements are part of the real-world environment depicted in the video.
[0020] In some cases, the disclosed technology receives a video including a depiction of a real-world object in a real-world environment. The disclosed technology accesses a segmentation associated with the real-world object and removes the depiction of the real-world object from a region of a first frame of the video. The disclosed technology processes the first frame of the video and one or more previous frames before the first frame through a machine learning model to generate a new frame in which a portion of the first frame has been blended into the region from which the depiction of the real-world object has been removed.
[0021] In this way, the disclosed technology can automatically display one or more AR objects in the current image or video when capturing an image or video without further input from the user. This improves the user's overall experience when using an electronic device and reduces the total amount of system resources required to complete the task.
[0022] Networked computing environment
[0023] Figure 1 1 is a block diagram illustrating an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network. The messaging system 100 includes multiple instances of a client device 102, each of the multiple instances hosting a number of applications including a messaging client 104 and other external applications 109 (e.g., third-party applications). Each messaging client 104 is communicatively coupled to other instances of the messaging client 104 (e.g., hosted on corresponding other client devices 102), a messaging server system 108, and an external app server 110 via a network 112 (e.g., the Internet). The messaging client 104 may also communicate with locally hosted third-party applications (also referred to as "external applications" and "external apps") 109 using an application programming interface (API).
[0024] The client device 102 can operate as a standalone device, or can be coupled (e.g., networked) to other machines. In a networked deployment, the client device 102 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The client device 102 may include, but is not limited to: a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, web appliances, a network router, a network switch, a network bridge, or any machine capable of performing the disclosed operations. In addition, although only a single client device 102 is shown, the term "client device" should also be considered to include a collection of machines that perform the disclosed operations individually or jointly.
[0025] In some examples, client device 102 may include AR glasses or AR headsets, where virtual content is displayed within the lenses of the glasses while the user views the real world environment through the lenses. For example, an image may be presented on a transparent display that allows the user to simultaneously view content presented on the display and real-world objects.
[0026] The messaging clients 104 are able to communicate and exchange data with other messaging clients 104 and messaging server systems 108 via the network 112. The data exchanged between the messaging clients 104 and between the messaging clients 104 and messaging server systems 108 include functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).
[0027] The messaging server system 108 provides server-side functionality to specific messaging clients 104 via the network 112. Although certain functions of the messaging system 100 are described herein as being performed by the messaging client 104 or by the messaging server system 108, it may be a design choice for the certain functions to be located in the messaging client 104 or in the messaging server system 108. For example, it may be technically preferred to initially deploy certain technologies and functions within the messaging server system 108, but to later migrate the technologies and functions to the messaging client 104 where the client device 102 has sufficient processing power.
[0028] The messaging server system 108 supports various services and operations provided to the messaging clients 104. Such operations include sending data to the messaging clients 104, receiving data from the messaging clients 104, and processing data generated by the messaging clients 104. As examples, the data may include message content, client device information, geo-location information, media enhancements and overlays, message content permanence conditions, social networking information, and live event information. Data exchanges within the messaging system 100 are activated and controlled by functions available via the user interface (UI) of the messaging clients 104.
[0029] Turning now specifically to the messaging server system 108, an application program interface (API) server 116 is coupled to the application server 114 and provides a programming interface to the application server 114. The application server 114 is communicatively coupled to a database server 120, which facilitates access to a database 126 that stores data associated with messages processed by the application server 114. Similarly, a web server 128 is coupled to the application server 114 and provides a web-based interface to the application server 114. To this end, the web server 128 handles incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.
[0030] The API server 116 receives and sends message data (e.g., commands and message payloads) between the client device 102 and the application server 114. Specifically, the API server 116 provides a collection of interfaces (e.g., routines and protocols) that can be called or queried by the messaging client 104 to activate the functionality of the application server 114. The API server 116 exposes various functions supported by the application server 114, including: account registration, login functionality, sending messages from a particular messaging client 104 to another messaging client 104 via the application server 114, sending media files (e.g., images or videos) from a messaging client 104 to the messaging server 118, and possible access by another messaging client 104, setting of collections (e.g., stories) of media data, retrieval of a friend list of a user of a client device 102, retrieval of such collections, retrieval of messages and content, addition and removal of entities (e.g., friends) of an entity graph (e.g., a social graph), position of friends within the social graph, and opening of application events (e.g., related to a messaging client 104).
[0031] The application server 114 hosts several server applications and subsystems, including, for example, a messaging server 118, an image processing server 122, and a social network server 124. The messaging server 118 implements several message processing technologies and functions, particularly those related to the aggregation and other processing of content (e.g., text and multimedia content) included in messages received from multiple instances of the messaging client 104. As will be described in more detail, text and media content from multiple sources can be aggregated into collections of content (e.g., referred to as stories or galleries). These collections are then made available to the messaging client 104. In view of the hardware requirements of other processor and memory intensive data processing, the processing can also be performed on the server side by the messaging server 118.
[0032] The application servers 114 also include an image processing server 122 dedicated to performing various image processing operations, typically on images or videos within the payload of messages sent from or received at the messaging server 118 .
[0033] The image processing server 122 is used to implement ( Figure 2 2 ). The scanning function includes activating and providing one or more augmented reality experiences on the client device 102 when an image is captured by the client device 102. Specifically, the messaging client 104 on the client device 102 can be used to activate the camera. The camera displays one or more real-time images or videos to the user along with one or more icons or identifiers of one or more augmented reality experiences. The user can select a given identifier in the identifier to launch the corresponding augmented reality experience or perform a desired image modification.
[0034] The social network server 124 supports various social networking functions and services and makes them available to the messaging server 118. To this end, the social network server 124 maintains and accesses the database 126 ( Figure 3 ) entity map 308. Examples of functions and services supported by social network server 124 include identifying other users with whom a particular user of messaging system 100 has relationships or "follows," and also identifying interests and other entities of a particular user.
[0035] Returning to the messaging client 104, the features and functions of the external resource (e.g., an external application 109 or an applet) are made available to the user via the interface of the messaging client 104. The messaging client 104 receives a user selection of an option for launching or accessing features of an external resource (e.g., a third-party resource) such as an external application 109. The external resource can be a third-party application 106 (e.g., a "native application") installed on the client device 102 or a small-scale version of a third-party application (e.g., a "applet") hosted on the client device 102 or at a remote end of the client device 102 (e.g., on an external resource or application server 110). The small-scale version of the third-party application includes a subset of the features and functions of the third-party application (e.g., a full-scale, native version of a third-party standalone application) and is implemented using a markup language document. In one example, the small-scale version of the third-party application (e.g., a "applet") is a web-based markup language version of the third-party application and is embedded in the messaging client 104. In addition to using markup language documents (eg, .*ml files), applets may also include scripting languages (eg, .*js files or .json files) and style sheets (eg, .*ss files).
[0036] In response to receiving a user selection of an option for launching or accessing a feature of an external resource (e.g., an external application 109), the messaging client 104 determines whether the selected external resource is a web-based external resource or a locally installed external application. In some cases, an external application 109 locally installed on the client device 102 can be launched independently of the messaging client 104 and separately from the messaging client 104, for example, by selecting an icon corresponding to the external application 109 on a home screen of the client device 102. Small-scale versions of such external applications can be launched or accessed via the messaging client 104, and in some examples, no part of the small-scale external application or only a limited part can be accessed outside the messaging client 104. The small-scale external application can be launched by the messaging client 104 receiving a markup language document associated with the small-scale external application from the external application server 110 and processing the document.
[0037] In response to determining that the external resource is a locally installed external application 109, the messaging client 104 instructs the client device 102 to launch the external application 109 by executing the locally stored code corresponding to the external application 109. In response to determining that the external resource is a web-based resource, the messaging client 104 communicates with the external application server 110 to obtain a markup language document corresponding to the selected resource. The messaging client 104 then processes the obtained markup language document to present the web-based external resource within the user interface of the messaging client 104.
[0038] The messaging client 104 may notify the user of the client device 102 or other users (e.g., "friends") associated with such users of activities occurring in one or more external resources. For example, the messaging client 104 may provide a participant in a conversation (e.g., a chat session) in the messaging client 104 with a notification related to the current or recent use of an external resource by one or more members of a user group. One or more users may be invited to join an active external resource or to start an external resource (in the friend group) that has been recently used but is currently inactive. The external resource may provide the participants in the conversation who each use the corresponding messaging client 104 with the ability to share items, conditions, states, or locations in the external resource with one or more members of the user group entering the chat session. The shared items may be interactive chat cards that members of the chat may interact with, such as starting a corresponding external resource, viewing specific information within an external resource, or bringing members of the chat to a specific location or state within an external resource. Within a given external resource, a response message may be sent to the user on the messaging client 104. The external resource may selectively include different media items in the response based on the current context of the external resource.
[0039] The messaging client 104 may present a list of available external resources (e.g., third-party or external applications 109 or applets) to the user to launch or access a given external resource. The list may be presented in a context-sensitive menu. For example, icons representing different external applications (or applets) of the external application 109 (or applets) may vary based on how the user launches the menu (e.g., from a conversational interface or from a non-conversational interface).
[0040] System Architecture
[0041] Figure 21 is a block diagram showing additional details about the messaging system 100 according to some examples. Specifically, the messaging system 100 is shown to include a messaging client 104 and an application server 114. The messaging system 100 contains several subsystems that are supported on the client side by the messaging client 104 and on the server side by the application server 114. These subsystems include, for example, a transient timer system 202, a collection management system 204, an enhancement system 208, a map system 210, a game system 212, and an external resource system 220.
[0042] The transient timer system 202 is responsible for implementing temporary or time-limited access to content by the messaging client 104 and the messaging server 118. The transient timer system 202 includes several timers that selectively enable access to messages and associated content (e.g., for presentation and display) via the messaging client 104 based on duration and display parameters associated with a message or a collection of messages (e.g., a story). Additional details regarding the operation of the transient timer system 202 are provided below.
[0043] The collection management system 204 is responsible for managing collections or collections of media (e.g., collections of text, images, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into "event galleries" or "event stories." The collection can be made available for a specified time period, such as the duration of an event related to the content. For example, content related to a concert can be made available as a "story" for the duration of the concert. The collection management system 204 can also be responsible for publishing an icon to the user interface of the messaging client 104 that provides notification of the existence of a particular collection.
[0044] The collection management system 204 also includes a curation interface 206 that allows collection managers to manage and curate specific content collections. For example, the curation interface 206 enables an event organizer to curate a collection of content related to a specific event (e.g., to delete inappropriate content or redundant messages). In addition, the collection management system 204 uses machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, compensation can be paid to users for including user-generated content in a collection. In such cases, the collection management system 204 operates to automatically pay such users for using their content.
[0045] The enhancement system 208 provides various functions that enable users to enhance (e.g., annotate or otherwise modify or edit) media content associated with a message. For example, the enhancement system 208 provides functions related to the generation and publication of media overlays for messages processed by the messaging system 100. The enhancement system 208 is operable to provide media overlays or enhancements (e.g., image filters) to the messaging client 104 based on the geographic location of the client device 102. In another example, the enhancement system 208 is operable to provide media overlays to the messaging client 104 based on other information such as social network information of the user of the client device 102. Media overlays can include audio and visual content and visual effects. Examples of audio and visual content include pictures, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects can be applied to media content items (e.g., photos) at the client device 102. For example, media overlays can include text, graphical elements, or images that can be superimposed on photos taken by the client device 102. In another example, the media overlay includes a location identification overlay (e.g., Venice Beach), the name of a live event, or a business name overlay (e.g., Beach Cafe). In another example, the augmentation system 208 uses the geographic location of the client device 102 to identify a media overlay that includes the name of a business at the geographic location of the client device 102. The media overlay may include other tags associated with the business. The media overlay may be stored in the database 126 and accessed by the database server 120.
[0046] In some examples, the enhancement system 208 provides a user-based publishing platform that enables a user to select a geographic location on a map and upload content associated with the selected geographic location. The user can also specify situations in which a particular media overlay should be provided to other users. The enhancement system 208 generates a media overlay that includes the uploaded content and associates the uploaded content with the selected geographic location.
[0047] In other examples, the augmentation system 208 provides a merchant-based publishing platform that enables merchants to select specific media overlays associated with a geographic location via a bidding process. For example, the augmentation system 208 associates the media overlays of the merchant with the highest bid with the corresponding geographic location for a predetermined amount of time. The augmentation system 208 communicates with the image processing server 122 to obtain an augmented reality experience and presents an identifier of such experience in one or more user interfaces (e.g., as an icon on top of a real-time image or video, or as a thumbnail or icon in the interface dedicated to the identifier of the augmented reality experience presented). Once the augmented reality experience is selected, one or more images, videos, or augmented reality graphical elements are obtained and presented as an overlay on the image or video captured by the client device 102. In some cases, the camera is switched to face-up (e.g., the front camera of the client device 102 is activated in response to the activation of a specific augmented reality experience), and an image from the front camera of the client device 102 instead of the rear camera of the client device 102 begins to be displayed on the client device 102. One or more images, videos, or augmented reality graphical elements are captured and presented as an overlay on top of the image captured and displayed by the front-facing camera of the client device 102 .
[0048] In other examples, the augmentation system 208 can communicate and exchange data with a server and with another augmentation system 208 on another client device 102 via the network 112. The exchanged data may include: a session identifier identifying the shared AR session; a transformation between a first client device 102 and a second client device 102 (e.g., the plurality of client devices 102 includes the first device and the second device) for aligning the shared AR session with a common origin; a common coordinate system; functionality (e.g., commands for activating functions), and other payload data (e.g., text, audio, video, or other multimedia data).
[0049] The augmentation system 208 sends the transformation to the second client device 102 so that the second client device 102 can adjust the AR coordinate system based on the transformation. In this way, the first client device 102 and the second client device 102 synchronize their coordinate systems and frames to display content in the AR session. Specifically, the augmentation system 208 calculates the origin of the second client device 102 in the coordinate system of the first client device 102. The augmentation system 208 can then determine the offset of the coordinate system of the second client device 102 based on the position of the origin in the coordinate system of the second client device 102 from the perspective of the second client device 102. The offset is used to generate a transformation so that the second client device 102 generates AR content according to a common coordinate system or frame with the first client device 102.
[0050] The augmentation system 208 can communicate with the client device 102 to establish a separate or shared AR session. The augmentation system 208 can also be coupled to the messaging server 118 to establish an electronic group communication session (e.g., group chat, instant messaging) for the client device 102 in the shared AR session. The electronic group communication session can be associated with a session identifier provided by the client device 102 to obtain access rights to the electronic group communication session and the shared AR session. In one example, the client device 102 first obtains access rights to the electronic group communication session and then obtains a session identifier in the electronic group communication session that allows the client device 102 to access the shared AR session. In some examples, the client device 102 is able to access the shared AR session without the assistance of the augmentation system 208 in the application server 114 or communication with the augmentation system 208 in the application server 114.
[0051] The mapping system 210 provides various geolocation functions and supports the presentation of map-based media content and messages by the messaging client 104. For example, the mapping system 210 enables the display of user icons or avatars (e.g., stored in the profile data 316) on a map to indicate the current or past locations of the user's "friends" and media content (e.g., a collection of messages including photos and videos) generated by such friends in the context of the map. For example, on the map interface of the messaging client 104, messages posted by the user to the messaging system 100 from a particular geographic location can be displayed to the "friends" of a particular user in the context of a map of the particular location. The user can also share his or her location and status information with other users of the messaging system 100 (e.g., using appropriate status avatars) via the messaging client 104, where the location and status information is similarly displayed to selected users within the context of the map interface of the messaging client 104.
[0052] The gaming system 212 provides various gaming functions within the context of the messaging client 104. The messaging client 104 provides a gaming interface that provides a list of available games (e.g., web-based games or web-based applications) that can be launched by a user within the context of the messaging client 104 and played with other users of the messaging system 100. The messaging system 100 also enables a particular user to invite other users to participate in the play of a particular game by sending invitations to such other users from the messaging client 104. The messaging client 104 also supports both voice and text messaging (e.g., chat) within the context of game play, provides game leaderboards, and supports the provision of in-game rewards (e.g., coins and power-ups).
[0053] The external resource system 220 provides an interface for the messaging client 104 to communicate with the external application server 110 to start or access external resources. Each external resource (application) server 110 hosts, for example, an application based on a markup language (e.g., HTML5) or a small-scale version of an external application (e.g., a game, utility, payment, or shared travel application outside the messaging client 104). The messaging client 104 can start the web-based resource (e.g., application) by accessing an HTML5 file from an external resource (application) server 110 associated with a web-based resource. In some examples, the application hosted by the external resource server 110 is programmed in JavaScript using a software development kit (SDK) provided by the messaging server 118. The SDK includes an application programming interface (API) having functions that can be called or activated by a web-based application. In some examples, the messaging server 118 includes a JavaScript library that provides a given third-party resource access to certain user data of the messaging client 104. HTML5 is used as an example technology for programming games, but applications and resources programmed based on other technologies can be used.
[0054] In order to integrate the functionality of the SDK into a web-based resource, the SDK is downloaded from the messaging server 118 by the external resource (application) server 110, or received in other ways by the external resource (application) server 110. Once the SDK is downloaded or received, the SDK is included as part of the application code of the web-based external resource. The code of the web-based resource can then call or activate certain functions of the SDK to integrate the features of the messaging client 104 into the web-based resource.
[0055] The SDK stored on the messaging server 118 effectively provides a bridge between external resources (e.g., a third party or external application 109 or applet) and the messaging client 104. This provides users with a seamless experience of communicating with other users on the messaging client 104 while also preserving the look and feel of the messaging client 104. In order to bridge the communication between the external resources and the messaging client 104, in some examples, the SDK facilitates the communication between the external resource server 110 and the messaging client 104. In some examples, the WebViewJavaScriptBridge running on the client device 102 establishes two one-way communication channels between the external resources and the messaging client 104. Messages are sent asynchronously between the external resources and the messaging client 104 via these communication channels. Each SDK function activation is sent as a message and a callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message with the callback identifier.
[0056] By using the SDK, not all information from the messaging client 104 is shared with the external resource server 110. The SDK limits which information is shared based on the needs of the external resource. In some examples, each external resource server 110 provides an HTML5 file corresponding to a web-based external resource to the messaging server 118. The messaging server 118 can add a visual representation (e.g., box art or other graphics) of the web-based external resource in the messaging client 104. Once the user selects the visual representation or instructs the messaging client 104 to access the features of the web-based external resource through the GUI of the messaging client 104, the messaging client 104 obtains the HTML5 file and instantiates the resources required to access the features of the web-based external resource.
[0057] The messaging client 104 presents a graphical user interface (e.g., a login page or title screen) for an external resource. During, before, or after presenting the login page or title screen, the messaging client 104 determines whether the launched external resource has been previously authorized to access the user data of the messaging client 104. In response to determining that the launched external resource has been previously authorized to access the user data of the messaging client 104, the messaging client 104 presents another graphical user interface of the external resource including the functions and features of the external resource. In response to determining that the launched external resource has not been previously authorized to access the user data of the messaging client 104, after a threshold time period (e.g., 3 seconds) of displaying the login page or title screen of the external resource, the messaging client 104 slides up a menu (e.g., animating the menu to emerge from the bottom of the screen to the middle or other portion of the screen) for authorizing the external resource to access the user data. The menu identifies the type of user data that the external resource will be authorized to use. In response to receiving the user selection of the accept option, the messaging client 104 adds the external resource to the list of authorized external resources and enables the external resource to access the user data from the messaging client 104. In some examples, the messaging client 104 authorizes the external resource to access the user data according to the OAuth 2 framework.
[0058] The messaging client 104 controls the type of user data shared with the external resource based on the type of external resource that is authorized. For example, an external resource including a full-scale external application (e.g., a third party or external application 109) is provided with access to a first type of user data (e.g., only a two-dimensional avatar of the user with or without different avatar characteristics). As another example, an external resource including a small-scale version of an external application (e.g., a web-based version of a third-party application) is provided with access to a second type of user data (e.g., payment information, a two-dimensional avatar of the user, a three-dimensional avatar of the user, and an avatar with various avatar characteristics). Avatar characteristics include different ways to customize the look and feel of an avatar (e.g., different poses, facial features, clothing, etc.).
[0059] The AR repair system 224 receives an image or video depicting a real-world environment (e.g., a room in a home) including one or more real-world objects (chairs, sofas, televisions, tables, people, etc.) from the client device 102. The AR repair system 224 receives or accesses a segmentation for the real-world object. The AR repair system 224 applies the segmentation to the image or video to remove the area of the image or video corresponding to the segmentation (e.g., remove the depiction of the real-world object). The AR repair system 224 can then apply a previously trained machine learning model to mix pixels from other parts of the image or video from which the real-world object has not been removed into the area of the image from which the real-world object has been removed. Specifically, the AR repair system 224 can use features of a previously captured image or frame and features of a currently captured image or frame to generate a new image in which the portion of the image or video that has been removed is filled with new pixel values.
[0060] In some examples, the AR restoration system 224 modifies the pixels of a given real-world object to blend the given real-world object with the background. In this way, the given real-world object can be removed or blended out from the image or video. The AR restoration system 224 can then place the AR item on the area that has been blended out (e.g., the AR item can be placed on top of where the real-world object was without overlapping or interfering with the real-world object). This allows the user to see how the AR item would look in the real-world environment as a replacement for the real-world object. Figure 5 An illustrative implementation of an AR repair system 224 is shown and described.
[0061] The AR repair system 224 is a component that can be accessed by an AR / VR application implemented on the client device 102. The AR / VR application uses an RGB camera to capture images or videos of the real-world environment. When capturing images or videos, the AR repair system 224 continuously or periodically modifies the images or videos to remove one or more areas depicting real-world objects and blend other parts of the image or video (e.g., background) into the removed one or more areas.
[0062] Data Architecture
[0063] Figure 3 is a schematic diagram illustrating a data structure 300 that may be stored in a database 126 of a messaging server system 108 according to some examples. Although the contents of the database 126 are illustrated as including several tables, it will be appreciated that data may be stored in other types of data structures (eg, an object-oriented database).
[0064] Database 126 includes message data stored in message table 302. For any particular message, the message data includes at least message sender data, message recipient (or receiver) data, and payload. Figure 4 Additional details regarding information that may be included in a message and included within the message data stored in message table 302 are described.
[0065] The entity table 306 stores entity data and is linked (e.g., by reference) to the entity graph 308 and profile data 316. The entities whose records are maintained within the entity table 306 may include individuals, corporate entities, organizations, objects, places, events, etc. Regardless of the entity type, any entity for which the messaging server system 108 stores data may be an identified entity. Each entity is provided with a unique identifier as well as an entity type identifier (not shown).
[0066] The entity graph 308 stores information about relationships and associations between entities. Such relationships may be social, professional (e.g., working in a common company or organization), interest-based, or activity-based, by way of example only.
[0067] The profile data 316 stores multiple types of profile data about a particular entity. Based on the privacy settings specified by the particular entity, the profile data 316 can be selectively used and presented to other users of the messaging system 100. In the case where the entity is a person, the profile data 316 includes, for example, a user name, phone number, address, settings (e.g., notification and privacy settings), and an avatar representation (or a collection of such avatar representations) selected by the user. The particular user can then selectively include one or more of these avatar representations in the content of messages transmitted via the messaging system 100 and on a map interface displayed to other users by the messaging client 104. The collection of avatar representations can include a "situation avatar" that presents a graphical representation of a situation or activity that the user can choose to transmit at a particular time.
[0068] Where the entity is a group, the group profile data 316 may similarly include one or more avatar representations associated with the group, in addition to the group name, members, and various settings (eg, notifications) of the relevant group.
[0069] Database 126 also stores enhancement data, such as overlays or filters, in enhancement table 310. The enhancement data is associated with and applied to videos (whose data is stored in video table 304) and images (whose data is stored in image table 312).
[0070] The database 126 may also store data related to individual and shared AR sessions. The data may include data transmitted between an AR session client controller of a first client device 102 and another AR session client controller of a second client device 102, and data transmitted between an AR session client controller and an augmented system 208. The data may include data for establishing a common coordinate system for a shared AR scene, transformations between devices, session identifiers, images depicting the body, skeletal joint positions, wrist joint positions, feet, and the like.
[0071] In one example, a filter is an overlay that is displayed as an overlay on an image or video during presentation to a recipient user. Filters can be of various types, including filters that a user selects from a collection of filters presented to a sending user by messaging client 104 when the sending user is composing a message. Other types of filters include geolocation filters (also known as geofilters), which can be presented to a sending user based on a geographic location. For example, a geolocation filter specific to a neighboring or special location can be presented by messaging client 104 within a user interface based on geographic location information determined by a global positioning system (GPS) unit of client device 102.
[0072] Another type of filter is a data filter, which can be selectively presented to the sending user by the messaging client 104 based on other input or information collected by the client device 102 during the message creation process. Examples of data filters include the current temperature at a particular location, the current speed at which the sending user is traveling, the battery life of the client device 102, or the current time.
[0073] Other augmented data that may be stored in the image table 312 include augmented reality content items (eg, corresponding to an applied augmented reality experience). Augmented reality content items or augmented reality items may be real-time special effects and sounds that may be added to an image or video.
[0074] As described above, augmented data includes augmented reality content items, overlays, image transformations, AR images, and similar items involving modifications that can be applied to image data (e.g., videos or images). This includes real-time modifications that modify images when they are captured using a device sensor (e.g., one or more cameras) of the client device 102 and then display the modified images on the screen of the client device 102. This also includes modifications to stored content such as video clips that can be modified in a gallery. For example, in a client device 102 with access rights to multiple augmented reality content items, a user can use a single video clip with multiple augmented reality content items to see how different augmented reality content items will modify the stored clips. For example, by selecting different augmented reality content items for the content, multiple augmented reality content items that apply different pseudo-random movement models can be applied to the same content. Similarly, real-time video capture can be used with the modifications shown to show how the video image currently being captured by the sensor of the client device 102 will modify the captured data. Such data can be displayed only on the screen without being stored in the memory, or the content captured by the device sensor can be recorded and stored in the memory with or without modification (or both). In some systems, the preview feature can simultaneously display what different augmented reality content items will look like in different windows in the display. For example, this can enable viewing multiple windows with different pseudo-random animations on the display at the same time.
[0075] Thus, data and various systems using augmented reality content items or other such transformation systems using the data to modify content may involve detection of objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.) in video frames, tracking of such objects as they leave the field of view, enter the field of view, and move around the field of view, and modification or transformation of such objects as they are tracked. In various examples, different methods for implementing such transformations may be used. Some examples may involve generating a three-dimensional mesh model of one or more objects, and using transformations and animated textures of the models within the video to implement the transformations. In other examples, tracking of points on an object may be used to place an image or texture (which may be two-dimensional or three-dimensional) at the tracked location. In yet another example, neural network analysis of a video frame may be used to place an image, model, or texture in content (e.g., an image or video frame). Thus, an augmented reality content item refers to both images, models, and textures used to create transformations in content, and additional modeling and analysis information required to implement the transformation using object detection, tracking, and placement.
[0076] Real-time video processing can be performed with any kind of video data (e.g., video streams, video files, etc.) stored in the memory of any kind of computerized system. For example, a user can load video files and store them in the memory of the device, or a sensor of the device can be used to generate a video stream. In addition, any object, such as human faces and body parts, animals, or non-living things such as chairs, cars, or other objects, can be processed using computer animation models.
[0077] In some examples, when a specific modification is selected together with the content to be transformed, the element to be transformed is identified by the computing device, and then the element is detected and tracked if it exists in the frame of the video. The elements of the object are modified according to the request for modification, so the frame of the video stream is transformed. For different types of transformations, the transformation of the frame of the video stream can be performed by different methods. For example, for the transformation of the frame that mainly involves changing the form of the elements of the object, the characteristic points of each element of the object are calculated (for example, using an active shape model (ASM) or other known methods). Then, for each element of at least one element of the object, a grid based on the characteristic points is generated. The grid is used for the subsequent stage of tracking the elements of the object in the video stream. During the tracking process, the grid mentioned for each element is aligned with the position of each element. Then, additional points are generated on the grid. A collection of first points is generated for each element based on the request for modification, and a collection of second points is generated for each element based on the collection of first points and the request for modification. Then, the frame of the video stream can be transformed by modifying the elements of the object based on the collection of first points and the collection of second points and the grid. In this type of method, the background of the modified object can also be changed or distorted by tracking and modifying the background.
[0078] In some examples, a transformation that changes some areas of an object using elements of an object can be performed by calculating characteristic points for each element of the object and generating a grid based on the calculated characteristic points. Points are generated on the grid, and then various areas based on these points are generated. Then, the elements of the object are tracked by aligning the area for each element with the position for each element of at least one element, and the properties of the area can be modified based on a request for modification, thereby transforming the frame of the video stream. Depending on the specific request for modification, the properties of the mentioned area can be transformed in different ways. Such modifications may involve: changing the color of the area; removing at least some parts of the area from the frame of the video stream; including one or more new objects in the area based on the request for modification; and modifying or distorting the elements of the area or object. In various examples, any combination of such modifications or other similar modifications may be used. For certain models to be animated, some characteristic points may be selected as control points for determining the entire state space of options for model animation.
[0079] In some examples of computer animation models that use face detection to transform image data, faces are detected on an image using a specific face detection algorithm (e.g., Viola-Jones). An active shape model (ASM) algorithm is then applied to the face region of the image to detect facial feature reference points.
[0080] Other methods and algorithms suitable for face detection can be used. For example, in some examples, landmarks are used to locate features, and landmarks represent distinguishable points that are present in most of the images considered. For example, for facial landmarks, the location of the left pupil can be used. If the initial landmarks are not identifiable (for example, if the person has an eye mask), secondary landmarks can be used. This type of landmark identification process can be used for any such object. In some examples, a collection of landmarks forms a shape. The coordinates of the points in the shape can be used to represent the shape as a vector. A shape is aligned with another shape using a similarity transformation (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between the shape points. The mean shape (meanshape) is the average of the aligned training shapes.
[0081] In some examples, the landmark search begins with an average shape that aligns with the location and size of the face determined by a global face detector. This type of search then repeats the following steps until convergence occurs: tentative shapes are proposed by adjusting the positions of shape points through template matching of the image texture around each point, and then the tentative shapes are conformed to the global shape model. In some systems, individual template matches are unreliable, and the shape model aggregates the results of weak template matches to form a stronger overall classifier. The entire search is repeated at each level of the image pyramid from coarse resolution to fine resolution.
[0082] The transformation system can capture an image or video stream on a client device (e.g., client device 102) and perform complex image manipulations locally on the client device 102 while maintaining an appropriate user experience, computation time, and power consumption. Complex image manipulations can include size and shape changes, emotion transitions (e.g., changing a face from a frown to a smile), state transitions (e.g., aging a subject, reducing apparent age, changing gender), style transitions, application of graphical elements, and any other suitable image or video manipulations enabled by a convolutional neural network that has been configured to execute efficiently on the client device 102.
[0083] In some examples, a computer animation model for transforming image data can be used by a system in which a user can capture an image or video stream of the user (e.g., a selfie) using a client device 102 having a neural network, where the neural network is running as part of a messaging client 104 running on the client device 102. A transformation system running within the messaging client 104 determines the presence of a face within the image or video stream and provides a modification icon associated with the computer animation model for transforming the data image, or the computer animation model can exist in association with an interface described herein. The modification icon includes the following changes, which can be the basis for modifying the user's face in the image or video stream as part of the modification operation. Once the modification icon is selected, the transformation system initiates a process of transforming the user's image to reflect the selected modification icon (e.g., generating a smiling face for the user). Once the image or video stream is captured and the specified modification is selected, the modified image or video stream can be presented in a graphical user interface displayed on the client device 102. The transformation system can implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification. That is, a user can capture an image or video stream, and once a modification icon is selected, the modification result can be presented in real time or near real time. In addition, while capturing the video stream, the modification can be continuous and the selected modification icon remains toggled. Machine learning neural networks can be used to implement such modifications.
[0084] The graphical user interface presenting the modifications performed by the transformation system can provide additional interactive options to the user. Such options can be based on the interface for initiating the selection of a particular computer animation model and content capture (e.g., initiated from a content creator user interface). In various examples, the modification can be persistent after the initial selection of the modification icon. The user can switch the modification on or off by tapping or otherwise selecting the face being modified by the transformation system, and store it for later viewing or browsing to other areas of the imaging application. In the case of multiple faces being modified by the transformation system, the user can globally switch the modification on or off by tapping or selecting a single face displayed and modified within the graphical user interface. In some examples, each face in a group of multiple faces can be modified separately, or such modifications can be switched separately by tapping or selecting an individual face or a series of individual faces displayed within the graphical user interface.
[0085] The story table 314 stores data related to a collection of messages and associated images, video or audio data compiled into a collection (e.g., a story or gallery). The creation of a particular collection can be initiated by a particular user (e.g., each user whose record is maintained in the entity table 306). A user can create a "personal story" in the form of a collection of content that has been created and sent / broadcasted by the user. To this end, the user interface of the messaging client 104 may include a user-selectable icon to enable the sending user to add specific content to his or her personal story.
[0086] A collection may also constitute a "live story" that is a collection of content from multiple users, created manually, automatically, or using a combination of manual and automatic techniques. For example, a "live story" may constitute a curated stream of user-submitted content from various locations and events. Users whose client devices have location services enabled and who are at a common location event at a particular time may be presented with an option, for example via a user interface of the messaging client 104, to contribute content to a particular live story. Live stories may be identified to the user by the messaging client 104 based on his or her location. The end result is a "live story" told from a community perspective.
[0087] Another type of content collection is called a "location story," which enables users whose client devices 102 are located in a particular geographic location (e.g., on a college or university campus) to contribute to a particular collection. In some examples, contributions to location stories may require a second level of authentication to verify that the end user belongs to a particular organization or other entity (e.g., is a student on a university campus).
[0088] As mentioned above, the video table 304 stores video data, which in one example is associated with a message whose record is maintained within the message table 302. Similarly, the image table 312 stores image data associated with a message for which message data is stored in the entity table 306. The entity table 306 may associate various enhancements from the enhancement table 310 with the various images and videos stored in the image table 312 and the video table 304.
[0089] Data structure 300 can also store training data for training one or more machine learning techniques (models) to classify real-world environments. The training data can include multiple images and videos and their corresponding true value real-world environment classifications. The images and videos can include a mixture of all categories of real-world objects that may appear in different real-world environments (e.g., rooms in a home or residence). One or more machine learning techniques can be trained to extract features of received input images or videos, and establish a relationship between the extracted features and the real-world environment classification. Once trained, the machine learning technology can receive new images or videos, and can calculate the real-world environment classification for the newly received images or videos.
[0090] The data structure 300 may also store training data for training one or more machine learning techniques (models) to determine the mixing mode of an object. The training data may include multiple images and videos, real-world object labels of real-world objects appearing in the images and videos, and their corresponding true value mixing modes. The images and videos may include a mixture of all kinds of real-world objects that may appear in different real-world environments (e.g., rooms in a home or residence). One or more machine learning techniques may be trained to extract features of received input images or videos, and establish a relationship between the extracted features, the objects detected in the images or videos, and the mixing mode of each object. Once trained, the machine learning technique may receive a new image or video, a given real-world object, and may calculate a mixing mode for the real-world objects appearing in the newly received image or video.
[0091] The data structure 300 may also store a plurality of different expected objects or lists for different real-world environment classifications. For example, the data structure 300 may store a first list of expected objects for a first real-world environment classification. That is, a real-world environment classified as a kitchen may be associated with a list of expected objects such as electrical appliances and / or furniture items, including: a tea maker, an oven, a kettle, a mixer, a refrigerator, a blender, a storage cabinet, a cabinet, a range hood, a range hood, a microwave, a dishwashing liquid, a kitchen countertop, a dining table, a kitchen scale, a pedal trash can, a grill, and a drawer. As another example, a real-world environment classified as a living room may be associated with a list of expected objects, including: a wing chair, a TV stand, a sofa, a cushion, a phone, a TV, a speaker, a side table, a tea set, a fireplace, a remote control, a fan, a floor lamp, a carpet, a table, blinds, curtains, a picture, a vase, and a grandfather clock.
[0092] As another example, a real-world environment classified as a bedroom may be associated with a list of expected objects such as furniture items, including: headboards, footboards and mattress frames, mattresses and box springs, mattress pads, sheets and pillowcases, blankets, quilts, comforters, bedspreads, duvets, bed skirts, sleeping pillows, specialty pillows, decorative pillows, pillow covers and pillowcases, throws (blankets), decorative fabrics, rods, brackets, draperies, curtains, roller blinds, blinds, nightstands, casual tables; lighting fixtures: floor lamps, table lamps, pendant lamps; wall lamps , alarm clocks, radios, plants and plant containers, vases, flowers, candles, candlesticks, artwork, posters, prints, photos, photo frames, photo albums, decorations and knickknacks, dressing tables and clothing, wardrobes, closets, TV cabinets, chairs, love chairs, benches, footstools, bookshelves, decorative shelves, books, magazines, bookends, luggage, benches, writing desks, dressing tables, mirrors, rugs, jewelry boxes and jewelry, storage boxes, baskets, trays, telephones; televisions, cable boxes, satellite receivers, DVD players and videos, tablets, night lights.
[0093] Each of the expected objects may be associated with a tag identifying a type of the expected object and a size range associated with the expected object. The tag may then be used to selectively modify pixels of the expected object (if the expected object appears in the real-world environment) in response to an AR item being placed over the expected object in the real-world environment (e.g., applying a blending mode to remove the expected object from an image or video).
[0094] Data communication architecture
[0095] Figure 4 is a schematic diagram illustrating the structure of a message 400 according to some examples, which message 400 is generated by a messaging client 104 for transmission to another messaging client 104 or a messaging server 118. The content of a particular message 400 is used to populate a message table 302 stored in a database 126 accessible by a messaging server 118. Similarly, the content of the message 400 is stored in memory as "in-flight" or "in-flight" data of a client device 102 or application server 114.
[0096] Message 400 is shown to include the following example components:
[0097] Message identifier 402 : a unique identifier that identifies the message 400 .
[0098] Message text payload 404 : text to be generated by a user via the user interface of the client device 102 and included in the message 400 .
[0099] Message image payload 406 : Image data captured by a camera component of the client device 102 or obtained from a memory component of the client device 102 and included in the message 400 . The image data for a sent or received message 400 may be stored in the image table 312 .
[0100] Message video payload 408 : Video data captured by the camera component or obtained from the memory component of the client device 102 and included in the message 400 . The video data for the message 400 sent or received may be stored in the video table 304 .
[0101] Message audio payload 410 : audio data captured by a microphone or obtained from a memory component of the client device 102 and included in the message 400 .
[0102] Message enhancement data 412 : Enhancement data (eg, filters, stickers, or other annotations or enhancements) representing enhancements to be applied to the message image payload 406 , message video payload 408 , or message audio payload 410 of the message 400 . The enhancement data 412 for a sent or received message 400 may be stored in the enhancement table 310 .
[0103] Message duration parameter 414: A parameter value indicating the amount of time in seconds that the content of a message (eg, message image payload 406, message video payload 408, message audio payload 410) is to be made accessible or presented to a user via messaging client 104.
[0104] Message geolocation parameters 416: Geolocation data (e.g., latitude and longitude coordinates) associated with the content payload of the message. Multiple message geolocation parameter 416 values may be included in the payload, each of which is associated with a content item included in the content (e.g., a specific image within the message image payload 406 or a specific video in the message video payload 408).
[0105] Message story identifier 418: An identifier value that identifies one or more content collections (e.g., "Stories" identified in stories table 314) with which a particular content item in message image payload 406 of message 400 is associated. For example, multiple images within message image payload 406 may each be associated with multiple content collections using an identifier value.
[0106] Message tags 420: Each message 400 may be tagged with a plurality of tags, each of which indicates the subject of the content included in the message payload. For example, where a particular image included in the message image payload 406 depicts an animal (e.g., a lion), a tag value indicating the relevant animal may be included within the message tags 420. The tag values may be manually generated based on user input, or may be automatically generated using, for example, image recognition.
[0107] Message sender identifier 422: An identifier (eg, a messaging system identifier, an email address, or a device identifier) that indicates the user of the client device 102 on which the message 400 was generated and from which the message 400 was sent.
[0108] Message recipient identifier 424: An identifier (eg, a messaging system identifier, an email address, or a device identifier) that indicates the user of the client device 102 to which the message 400 is addressed.
[0109] The content (e.g., value) of each component of message 400 may be a pointer to a location in a table in which the content data value is stored. For example, the image value in message image payload 406 may be a pointer to a location in image table 312 (or the address of a location in image table). Similarly, the value in message video payload 408 may point to data stored in video table 304, the value stored in message enhancement data 412 may point to data stored in enhancement table 310, the value stored in message story identifier 418 may point to data stored in story table 314, and the values stored in message sender identifier 422 and message recipient identifier 424 may point to user records stored in entity table 306.
[0110] AR Repair System
[0111] Figure 5 is a block diagram illustrating an example AR repair system 224 according to an illustrative example. The AR repair system 224 includes a set of components 510 that operate on a set of input data, such as a monocular image or video 501 depicting a real-world object in a real-world environment. The AR repair system 224 includes a segmentation module 512, an object removal module 514, a video mixing module 516, an image modification module 518, an AR item selection module 519, and an image display module 520. All or some of the components of the AR repair system 224 can be implemented by a server, in which case the monocular image or video 501 is provided to the server by the client device 102. In some cases, some or all of the components of the AR repair system 224 can be implemented by the client device 102.
[0112] In some examples, the AR restoration system 224 receives a video including a depiction of a real-world object in a real-world environment. The AR restoration system 224 accesses a segmentation associated with the real-world object. The AR restoration system 224 removes the depiction of the real-world object from an area of a first frame of the video. The AR restoration system 224 uses a machine learning model to process the first frame of the video and one or more previous frames before the first frame to generate a new frame in which a portion of the first frame has been blended into the area from which the depiction of the real-world object has been removed.
[0113] In some examples, the AR repair system 224 generates a new frame of the video for display that does not include a depiction of the real-world object and in which a portion of the first frame has been blended into the area from which the depiction of the real-world object has been removed. In some aspects, the AR repair system 224 receives input defining a segmentation. In some aspects, the AR repair system 224 displays multiple AR experiences and receives input selecting a given AR experience from the multiple AR experiences. The AR repair system 224 obtains a segmentation associated with the selected given AR experience.
[0114] In some examples, the machine learning model includes a generative adversarial network (GAN). In some cases, processing the first frame and one or more previous frames by the machine learning model includes generating a modified first frame in response to removing a depiction of a real-world object from a region of the first frame. The modified previous frame is generated in response to applying segmentation to a given previous frame of the one or more previous frames to remove the depiction of the real-world object from the given previous frame. An encoder is applied to the modified first frame and the modified previous frame to generate a first plurality of features associated with the modified first frame and a second plurality of features associated with the modified previous frame.
[0115] In some examples, the AR repair system 224 selects a subset of the first plurality of features and selects a subset of the second plurality of features. The AR repair system 224 generates a combined feature subset by combining the subset of the first plurality of features with the subset of the second plurality of features based on a similarity metric associated with the first frame and the previous frame. The AR repair system 224 applies a decoder to the combined feature subset to generate a new frame. In some examples, the combined feature subset is generated based on a weighted average of the subset of the first plurality of features and the subset of the second plurality of features. In some aspects, the AR repair system 224 determines an optical flow frame representing a difference or distance between the first frame and the second frame, and calculates a similarity metric based on the optical flow frame.
[0116] In some examples, the AR repair system 224 determines a camera movement difference between the first frame and the previous frame, and calculates a similarity metric based on the camera movement difference. In some examples, the AR repair system 224 multiplies the similarity metric with a subset of the first plurality of features to generate a first set of adjusted features. The AR repair system 224 calculates the difference between the value "1" and the similarity metric, and multiplies the difference with a subset of the second plurality of features to generate a second set of adjusted features. The AR repair system 224 adds the first set of adjusted features and the second set of adjusted features to generate a combined feature subset.
[0117] In some examples, the AR repair system 224 applies the decoder by identifying one or more features in the first plurality of features that are not included in the subset of the first plurality of features, and adding the identified one or more features to the combined feature subset to generate a complete feature set. The decoder can be applied to the complete feature set to generate a new frame.
[0118] In some examples, the AR repair system 224 trains a machine learning model by receiving training data, the training data including a plurality of training videos depicting one or more training real-world objects and an associated ground-truth video in which one or more training real-world objects have been removed. Each of the plurality of training videos depicts a different real-world environment. The AR repair system 224 accesses a first training video associated with a first ground-truth video in the ground-truth video in which one or more training real-world objects have been removed, and modifies a first frame and an adjacent frame of the first training video to remove a given training real-world object in one or more training real-world objects depicted in the first frame and an adjacent second frame. The AR repair system 224 applies the machine learning model to the first frame and the adjacent frame of the first training video to estimate a new training frame in which portions of the first frame and the adjacent frame have been mixed into an area from which a depiction of a given training real-world object in one or more training real-world objects has been removed. The AR repair system 224 calculates a deviation between the estimated new training frame and a ground-truth frame in the first ground-truth video corresponding to a playback position of the first frame and the adjacent frame of the first training video, and updates one or more parameters of the machine learning model based on the calculated deviation. This process is repeated for all or part of the training video until a stopping criterion is reached.
[0119] In some examples, the AR restoration system 224 applies a full body segmentation machine learning model to the video to estimate the segmentation associated with the real-world object. In such a case, the depiction of the removed real-world object corresponds to the full body of the person.
[0120] In some aspects, the AR repair system 224 obtains one or more AR objects associated with the AR experience and overlays the one or more AR objects over a portion of the new frame corresponding to the area from which the depiction of the real-world object has been removed. In some aspects, the one or more AR objects include fashion items. In some aspects, the one or more AR objects include furniture items. In some aspects, the depiction of the removed real-world object corresponds to a watermark in the video.
[0121] Reference Figure 5 , the segmentation module 512 receives a monocular image or video 501. The image or video 501 may be received as a new image captured by a (front or rear) camera of the client device 102, a previously captured video stream, or a portion of a real-time video stream. The segmentation module 512 receives or accesses segmentations associated with real-world objects. For example, the segmentation module 512 may receive input from a user drawing a shape on the image or video 501. The segmentation module 512 uses the shape drawn by the user to generate a segmentation corresponding to the shape to remove areas of the image or video 501 within the generated segmentation.
[0122] In some examples, segmentation module 512 receives input from a user selecting a particular type of AR experience. For example, a user may select a virtual try-on experience. In response, segmentation module 512 accesses the AR experience to identify the type of clothing or fashion item corresponding to the virtual try-on experience. Segmentation module 512 determines that the type of AR experience is a full-body virtual try-on experience. In such a case, segmentation module 512 accesses a full-body segmentation system to generate a full-body segmentation of one or more people depicted in image or video 501. The full-body segmentation system can implement one or more trained machine learning models configured to generate segmentations corresponding to a person's full body.
[0123] In some cases, segmentation module 512 determines that the type of AR experience is a shirt or upper body virtual try-on experience. In such a case, segmentation module 512 accesses an upper body segmentation system to generate upper body segmentations of one or more persons depicted in image or video 501. The full body segmentation system implements one or more trained machine learning models configured to generate segmentations corresponding to the upper body of a person.
[0124] Similarly, segmentation module 512 may determine that the type of AR experience is a lower body virtual try-on experience. In such a case, segmentation module 512 accesses the lower body segmentation system to generate a lower body segmentation of one or more people depicted in image or video 501. The lower body segmentation system may implement one or more trained machine learning models configured to generate segmentations corresponding to the lower body of a person. This allows the user to virtually try on upper body fashion items (e.g., shirts), lower body fashion items (e.g., pants), and / or full body fashion items (e.g., full body suits or dresses).
[0125] The segmentation module 512 outputs the segmentation that has been accessed or generated to the object removal module 514. The object removal module 514 is configured to apply the segmentation to the image or video 501 as a mask to remove all pixel values within the segmentation or set them to black. This results in an image or frame in which the real-world objects corresponding to the segmentation depicted in the image or video 501 have been removed. Other parts of the image or video 501 outside the segmentation are retained, such as the background. In order to improve the illusion or maintain the illusion that the image or video 501 with the removed real-world objects is real, the image or video from which the real-world objects have been removed is provided to the video mixing module 516. The video mixing module 516 generates a new image or video in which other parts of the image or video 501 that have not been removed are mixed into the area of the image or video 501 that has been removed. In some cases, the video mixing module 516 applies one or more machine learning models (e.g., GAN) to perform the mixing operation. The video mixing module 516 can access or obtain one or more previously captured frames (e.g., an immediately previous frame or a frame at a certain distance from the current frame) of the image or video 501. Using both the current frame and the previously captured frames, the video mixing module 516 generates a new image or video in which the area from which the real-world object has been removed is filled with new pixel values based on other parts of the image or video 501.
[0126] Figure 66 is a block diagram 600 illustrating an example AR repair system according to an exemplary example. Specifically, the block diagram 600 is an example implementation of the video mixing module 516. The video mixing module 516 may include an encoder 610 and a decoder 620. The video mixing module 516 receives a current frame 602 of an image or video from the object removal module 514 in which the area corresponding to the real-world object has been removed based on the segmentation generated by the segmentation module 512. Together with the current frame 602, the video mixing module 516 receives a previous frame 601 of the image or video from the object removal module 514 in which the area corresponding to the real-world object has been removed based on the segmentation generated by the segmentation module 512. The previous frame 601 may be an immediately preceding frame before the current frame 602, or may be a frame that is a certain distance in time before the current frame 602.
[0127] The encoder 610 generates a first set of features 612 corresponding to the previous frame 601 and a second set of features 614 corresponding to the current frame 602. The video mixing module 516 selects a particular subset 634 of the first set of features 612 (e.g., features output by layers 3 to 5 of the encoder 610). The video mixing module 516 also selects a corresponding subset 632 or counterpart of the second set of features 614 (e.g., features output by layers 3 to 5 of the encoder 610).
[0128] The video mixing module 516 generates a combined feature subset 630 by combining a specific subset 634 and a corresponding subset 632 according to a mixing parameter or similarity measure 636 between the previous frame 601 and the current frame 602. In some examples, the video mixing module 516 calculates the difference between the previous frame 601 and the current frame 602 to generate the similarity measure 636 (e.g., an optical flow frame). As the difference between the previous frame 601 and the current frame 602 increases, the video mixing module 516 selects or weights the features of the current frame 602 more heavily than the features of the previous frame 601. Alternatively, as the difference between the previous frame 601 and the current frame 602 increases, the video mixing module 516 selects or weights the features of the previous frame 601 more heavily than the features of the current frame 602. In some cases, the video mixing module 516 accesses camera movement information (e.g., gyroscope measurements) associated with the camera used to capture the previous frame 601 and the current frame 602. Alternatively or in addition to the optical flow frames, the video mixing module 516 can adjust or calculate the similarity metric 636 based on the camera movement information.
[0129] In some examples, the video mixing module 516 calculates the combined feature subset 630 according to a function of the similarity metric 636. For example, the video mixing module 516 applies the function to calculate a weighted average of the features of the specific subset 634 and the corresponding subset 632. The function can be calculated according to the following formula: Faveraged = α * Fcurrent + (1-α) * Fprevious, where α represents the similarity metric 636, Fcurrent represents the corresponding subset 632, and Fprevious represents the specific subset 634.
[0130] The video mixing module 516 provides the combined feature subset 630 to the decoder 620. The decoder 620 also receives one or more other features 616 (e.g., features from layer 2 of the encoder 610) from the features 614 of the current frame 602 that were not used to generate the combined feature subset 630. The decoder 620 processes the other features 616 and the combined feature subset 630 to generate a new image or video in which portions of the current frame 602 have been mixed into the region from which the depiction of the real-world object has been removed.
[0131] In some examples, the video mixing module 516 (e.g., encoder 610 and / or decoder 620) is trained using training data. The training data can be initially generated by capturing a first video depicting a real-world object in a real-world environment. The real-world object can be removed, and then a second video of the real-world environment is captured in which the real-world object is not depicted. The second video can represent a true value image or frame corresponding to the first video. Multiple videos can be captured in a similar manner with different types of real-world objects and in different types of real-world environments.
[0132] During training, the machine learning model of the video mixing module 516 receives a given training video depicting a real-world object from the training data stored in the data structure 300 from a plurality of training videos. The video mixing module 516 applies segmentation to the training video to remove one or more portions of the training video corresponding to the real-world object. The video mixing module 516 obtains a previous frame (to which segmentation has been applied to remove the real-world object) relative to the current frame from the given training video and extracts one or more features from the previous frame and the current frame, and combines the extracted features according to a similarity measure between the previous frame and the current frame. Then, the video mixing module 516 generates or estimates a new frame that mixes pixels from other portions of the current frame into one or more portions of the training video that have been removed.
[0133] The video mixing module 516 obtains a known or predetermined true value video that does not depict real-world objects corresponding to a given training video. The video mixing module 516 determines the playback position corresponding to the current frame of the given training video, and obtains an image or frame corresponding to the determined playback position from the true value video. The acquired frame is a true value frame corresponding to the training video in which real-world objects are not depicted. The video mixing module 516 compares the estimated new frame with the true value image or frame. Based on the difference threshold of the comparison, the video mixing module 516 updates one or more coefficients or parameters of the machine learning model (e.g., encoder 610 and / or decoder 620), and obtains one or more additional training videos of the real-world environment. After a specified number of cycles or batches of training videos have been processed and / or when the difference threshold reaches a specified value, the video mixing module 516 completes the training, and the parameters and coefficients of the video mixing module 516 are stored as a trained machine learning model.
[0134] The AR item selection module 519 selects a given AR object. For example, the AR item selection module 519 may receive a user selection of an AR fashion item or an AR furniture item. The AR item selection module 519 provides the identification of the AR object to the image modification module 518. The image modification module 518 receives a new image or video generated by the video mixing module 516, in which one or more parts corresponding to the segmentation have been removed and mixed with other parts that have not been removed. The image modification module 518 receives placement information from user input to position the AR object at a specific location within the new image or video generated by the video mixing module 516. In some cases, the image modification module 518 automatically places the selected AR object on an area of the image or video 501 corresponding to the segmentation generated by the segmentation module 512. The image modification module 518 can access movement information from the camera used to capture the image or video 501, and can dynamically update the position of the AR object based on the tracking of the movement of the camera. Similarly, the video mixing module 516 can continuously track the movement of the camera to dynamically adjust the position of the segmentation and dynamically adjust which parts of the image or video 501 are removed. As the position of the segmentation is adjusted in a similar manner (e.g., based on the previous frame), the video mixing module 516 continuously mixes in pixels from other areas of the image or video 501 that have not been removed.
[0135] The image modification module 518 modifies the pixels of the area of the new image or video on which the AR object is placed to blend the AR object with the background of the real-world environment depicted in the new image or video 501. As a result, the real-world object is removed from the image or video 501 and replaced with the AR object. The image modification module 518 then displays the AR object on the new image or video to make it appear as if the AR object has been placed in the real-world environment instead of the real-world object corresponding to the segmentation. The image modification module 518 provides the modified image or video to the image display module 520 for presentation to the user on an output device.
[0136] Figure 7 and Figure 8 700 is a diagrammatic representation of the output of the AR repair system 224 according to some examples. Specifically, as shown in the output 700 of the AR repair system 224, an image or video 710 can be captured in real time by a camera device. The image or video 710 includes a depiction of a real-world object 712 (e.g., a person). The AR repair system 224 can receive a segmentation (e.g., a full-body segmentation) corresponding to the real-world object 712. The AR repair system 224 can use the segmentation to remove the real-world object 712 from the video 710. The AR repair system 224 can generate a new frame 720 based on a previously captured frame of the video 710, in which the real-world object 712 is removed and the area 722 corresponding to the segmentation is filled with mixed pixels from other parts of the video 710 (current frame and / or previous frame) that have not been removed.
[0137] As shown in output 800 of the AR repair system 224, an image or video 810 is captured in real time by a camera. The image or video 810 includes a depiction of a real-world object 812 (e.g., a person) and a watermark 814. The AR repair system 224 receives a segmentation (e.g., an upper body segmentation) corresponding to the real-world object 812. The AR repair system 224 uses the segmentation to remove the upper body portion of the real-world object 812 from the video 810. The AR repair system 224 generates a new frame based on a previously captured frame of the video 810, in which the upper body of the real-world object 812 is removed and an area 820 corresponding to the upper body segmentation is filled with pixels mixed in from other parts of the video 810 (current frame and / or previous frame) that have not been removed. Similarly, the AR repair system 224 generates a new frame based on a previously captured frame of the video 810, in which the watermark 814 is removed and the area 824 corresponding to the watermark is filled with mixed pixels from other parts of the video 810 (the current frame and / or the previous frame) that have not been removed.
[0138] After generating the new frame or video, the AR repair system 224 receives input for selecting an option to virtually try on one or more AR fashion items. In response to receiving input for selecting a given fashion item (e.g., an upper body garment such as a shirt), the AR repair system 224 modifies the new image or video to display the selected AR object 830 above the area 820 corresponding to the upper body segmentation, which has been filled with mixed pixels from other parts of the video 810 (current frame and / or previous frame) that have not been removed. As the real-world object 812 continues to move around in the image or video 810, the AR repair system 224 continuously tracks the real-world object 812 to continuously update the position of the segmentation for removing the portion of the real-world object 812. As those parts are removed, the AR repair system 224 continuously mixes in pixels from other parts that have not been removed based on the previous video frame, and maintains the placement of the AR object that has been added to the area that has been segmented out and removed.
[0139] In other examples, the real-world objects 812 include items of real-world furniture in a store or home. In such cases, the AR restoration system 224 uses furniture item segmentation to remove the real-world furniture items from the video 810. The AR restoration system 224 generates a new frame based on a previously captured frame of the video 810 in which the real-world furniture is removed and the area 820 corresponding to the real-world furniture is filled with pixels mixed in from other portions of the video 810 (the current frame and / or the previous frame) that were not removed.
[0140] After generating the new frame or video, the AR repair system 224 receives input selecting an option to virtually try out one or more AR furniture items. In response to receiving input selecting a given furniture item, the AR repair system 224 modifies the new image or video to display the selected AR furniture item over an area 820 corresponding to the furniture item segmentation, which has been filled with mixed-in pixels from other portions of the video 810 (current frame and / or previous frames) that have not been removed. As the camera moves around, the AR repair system 224 continuously tracks the real-world furniture items to continuously update the locations of the segmentations for removing the furniture items. As those portions are removed, the AR repair system 224 continuously mixes in pixels from other portions that have not been removed based on previous video frames, and maintains the placement of AR furniture items that have been added over the areas that have been segmented out and removed.
[0141] Fig. 9900 is a flow chart of a process according to some examples. Although the flow chart may describe the operations as a sequential process, many of these operations may be performed in parallel or simultaneously. In addition, the order of the operations may be rearranged. When the operations of the process are completed, the process is terminated. The process may correspond to a method, a process, etc. The steps of a method may be performed in whole or in part, may be performed in combination with some or all of the steps in other methods, and may be performed by any number of different systems or any part thereof (e.g., a processor included in any system).
[0142] At operation 901 , as discussed above, the client device 102 receives a video including a depiction of a real-world object in a real-world environment.
[0143] At operation 902 , as discussed above, the client device 102 accesses a segmentation associated with a real-world object.
[0144] At operation 903 , as discussed above, the client device 102 removes the depiction of the real-world object from the area of the first frame of the video.
[0145] At operation 904, as discussed above, the client device 102 processes a first frame of the video and one or more previous frames preceding the first frame through a machine learning model to generate a new frame in which portions of the first frame have been blended into areas from which depictions of real-world objects have been removed.
[0146] Machine Architecture
[0147] Fig.101000 within which instructions 1008 (e.g., software, programs, applications, applet, app, or other executable code) may be executed for causing the machine 1000 to perform any one or more of the methods discussed herein. For example, the instructions 1008 may cause the machine 1000 to perform any one or more of the methods described herein. The instructions 1008 transform a general purpose, unprogrammed machine 1000 into a specific machine 1000 that is programmed to perform the functions described and illustrated in the manner described. The machine 1000 may operate as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1000 may operate as a server machine or a client machine in a server-client network environment or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1000 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, web appliances, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing instructions 1008 specifying actions to be taken by the machine 1000. In addition, although only a single machine 1000 is shown, the term "machine" should also be considered to include a collection of machines that individually or jointly execute instructions 1008 to perform any one or more of the methods discussed herein. For example, the machine 1000 may include any of the client devices 102 or a number of server devices that form part of the messaging server system 108. In some examples, the machine 1000 may also include both a client system and a server system, wherein certain operations of a particular method or algorithm are performed on the server side and certain operations of the particular method or algorithm are performed on the client side.
[0148] The machine 1000 may include a processor 1002, a memory 1004, and an input / output (I / O) component 1038 that may be configured to communicate with each other via a bus 1038. In an example, the processor 1002 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 1006 and a processor 1010 that execute instructions 1008. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that may execute instructions simultaneously. Although Fig.10 Multiple processors 1002 are shown, but the machine 1000 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0149] The memory 1004 includes a main memory 1012, a static memory 1014, and a storage unit 1016, which are all accessible by the processor 1002 via the bus 1040. The main memory 1004, the static memory 1014, and the storage unit 1016 store instructions 1008 that embody any one or more of the methodologies or functions described herein. The instructions 1008 may also reside, completely or partially, within the main memory 1012, within the static memory 1014, within a machine-readable medium within the storage unit 1016, within at least one of the processors 1002 (e.g., within a cache memory of a processor), or within any suitable combination thereof during execution thereof by the machine 1000.
[0150] I / O components 1038 may include a variety of components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurements, etc. The specific I / O components 1038 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine will most likely not include such a touch input device. It will be appreciated that I / O components 1038 may include Fig.101024 and 1026. In various examples, the I / O components 1038 may include a user output component 1024 and a user input component 1026. The user output component 1024 may include a visual component (e.g., a display such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), an acoustic component (e.g., a speaker), a tactile component (e.g., a vibration motor, a resistance mechanism), other signal generators, etc. The user input component 1026 may include an alphanumeric input component (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input component), a point-based input component (e.g., a mouse, a touch pad, a trackball, a joystick, a motion sensor, or other pointing instrument), a tactile input component (e.g., a physical button, a touch screen or other tactile input component that provides the location and force of a touch or touch gesture), an audio input component (e.g., a microphone), etc.
[0151] In other examples, I / O component 1038 may include biometric component 1028, motion component 1030, environment component 1032, or positioning component 1034, as well as a variety of other components. For example, biometric component 1028 includes components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition), etc. Motion component 1030 includes acceleration sensor components (e.g., accelerometers), gravity sensor components, and rotation sensor components (e.g., gyroscopes).
[0152] Environmental components 1032 include, for example, one or more cameras (with still image / photo and video capabilities), illumination sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors (e.g., gas detection sensors that detect concentrations of hazardous gases for safety or measure pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment.
[0153] With respect to cameras, client device 102 may have a camera system including, for example, a front-facing camera on a front surface of client device 102 and a rear-facing camera on a rear surface of client device 102. The front-facing camera may, for example, be used to capture still images and videos (e.g., “selfies”) of a user of client device 102, which may then be enhanced with the enhancement data (e.g., filters) described above. The rear-facing camera may, for example, be used to capture still images and videos in a more traditional camera mode, which are similarly enhanced with the enhancement data. In addition to the front-facing camera and the rear-facing camera, client device 102 may also include a 360° camera for capturing 360° photos and videos.
[0154] Additionally, the camera system of the client device 102 may include dual rear cameras (e.g., a main camera and a depth sensing camera), or even triple, quad, or quintuple rear camera configurations on the front and back of the client device 102. For example, these multi-camera systems may include a wide-angle camera, an ultra-wide-angle camera, a telephoto camera, a macro camera, and a depth sensor.
[0155] The positioning component 1034 includes a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., a barometer or altimeter that detects air pressure from which the altitude can be derived), an orientation sensor component (e.g., a magnetometer), and the like.
[0156] A variety of technologies can be used to achieve communication. The I / O components 1038 also include a communication component 1036 that is operable to couple the machine 1000 to the network 1020 or the device 1022 via a corresponding coupling or connection. For example, the communication component 1036 may include a network interface component or other suitable device that interfaces with the network 1020. In other examples, the communication component 1036 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, Parts (e.g. Low energy consumption), Device 1022 may be another machine or any of a variety of peripheral devices (eg, a peripheral device coupled via USB).
[0157] In addition, the communication component 1036 can detect an identifier or include a component that can be operated to detect an identifier. For example, the communication component 1038 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional bar codes such as universal product codes (UPC) bar codes, multi-dimensional bar codes such as quick response (QR) codes, Aztec codes, data matrix, data symbols (Dataglyph), MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar codes, and other optical codes) or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tagged product). In addition, various information can be obtained via the communication component 1036, such as location via Internet Protocol (IP) geolocation, ... Signal triangulation to obtain location, location obtained via detection of NFC beacon signals that can indicate a specific location, etc.
[0158] Various memories (e.g., main memory 1012, static memory 1014, and memory of processor 1002) and storage unit 1016 may store one or more collections of data structures (e.g., software) and instructions used by or embodying any one or more of the methods or functions described herein. When executed by processor 1002, these instructions (e.g., instructions 1008) cause various operations to implement the disclosed examples.
[0159] The instructions 1008 may be sent or received over the network 1020 using a transmission medium via a network interface device (e.g., a network interface component included in the communication component 1036) and using any of a number of well-known transmission protocols (e.g., the Hypertext Transfer Protocol (HTTP)). Similarly, the instructions 1008 may be sent or received via a coupling (e.g., a peer-to-peer coupling) with the device 1022 using a transmission medium.
[0160] Software Architecture
[0161] Fig.111100 is a block diagram illustrating a software architecture 1104 that can be installed on any one or more of the devices described herein. The software architecture 1104 is supported by hardware such as a machine 1102 including a processor 1120, a memory 1126, and an I / O component 1138. In this example, the software architecture 1104 can be conceptualized as a stack of layers, where each layer provides specific functionality. The software architecture 1104 includes layers such as an operating system 1112, a library 1110, a framework 1108, and an application 1106. In operation, the application 1106 activates an API call 1150 through the software stack and receives a message 1152 in response to the API call 1150.
[0162] The operating system 1112 manages hardware resources and provides public services. The operating system 1112 includes, for example, a kernel 1114, services 1116, and drivers 1122. The kernel 1114 serves as an abstraction layer between the hardware and other software layers. For example, the kernel 1114 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. Services 1116 can provide other public services to other software layers. Drivers 1122 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 1122 may include display drivers, camera drivers, or Low-power drivers, Flash drivers, Serial communication drivers (e.g., USB drivers), Drivers, audio drivers, power management drivers, etc.
[0163] The library 1110 provides a common low-level infrastructure used by the application 1106. The library 1110 may include a system library 1118 (e.g., a C standard library) that provides functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, the library 1110 may include an API library 1124, such as a media library (e.g., a library for supporting the presentation and manipulation of various media formats such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), a graphics library (e.g., an OpenGL framework for rendering in two dimensions (2D) and three dimensions (3D) in graphical content on a display), a database library (e.g., SQLite for providing various relational database functions), a web library (e.g., WebKit for providing web browsing functions), etc. The library 1110 may also include a variety of other libraries 1128 to provide many other APIs to the application 1106 .
[0164] The framework 1108 provides a common high-level infrastructure used by the applications 1106. For example, the framework 1108 provides various graphical user interface (GUI) functions, advanced resource management, and advanced location services. The framework 1108 can provide a wide range of other APIs that can be used by the applications 1106, some of which may be specific to a particular operating system or platform.
[0165] In an example, applications 1106 may include a home application 1136, a contacts application 1130, a browser application 1132, a book reader application 1134, a location application 1142, a media application 1144, a messaging application 1146, a game application 1148, and a variety of other applications such as external applications 1140. Applications 1106 are programs that execute functions defined in the program. Various programming languages may be used to create one or more of the applications 1106 structured in various ways, such as an object-oriented programming language (e.g., Objective-C, Java, or C++) or a procedural programming language (e.g., C or assembly language). In a specific example, external applications 1140 (e.g., applications created by an entity other than the vendor of a particular platform using ANDROID TM or IOS TM Software Development Kit (SDK) applications can be developed for example IOS TM ANDROID TM , Mobile software running on a mobile operating system such as a mobile phone or another mobile operating system. In this example, external application 1140 can activate API calls 1150 provided by operating system 1112 to facilitate the functions described herein.
[0166] Glossary
[0167] "Carrier signal" means any intangible medium capable of storing, encoding or carrying instructions for execution by a machine and includes a digital or analog communications signal or other intangible medium to facilitate transmission of such instructions. Instructions may be sent or received over a network using a transmission medium via a network interface device.
[0168] "Client Device" refers to any machine that interfaces with a communications network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, desktop computer, laptop computer, portable digital assistant (PDA), smart phone, tablet computer, ultrabook, netbook, notebook computer, multiprocessor system, microprocessor-based or programmable consumer electronics, game console, set-top box, or any other communications device that a user may use to access a network.
[0169] "Communications Network" means one or more parts of a network, which may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a Plain Old Telephone Service (POTS) network, a cellular telephone network, a wireless network, A network, other type of network, or a combination of two or more such networks. For example, the network or a portion of the network may include a wireless network or a cellular network, and the coupling may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, the coupling may implement any of various types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rate GSM evolution (EDGE) technology, the third generation partnership project (3GPP) including 3G, fourth generation wireless (4G) network, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), world wide interoperability for microwave access (WiMAX), long term evolution (LTE) standards, other data transmission technologies defined by various standard setting organizations, other long distance protocols, or other data transmission technologies.
[0170] "Component" means a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other techniques of segmentation or modularization that provide specific processing or control functionality. Components may be combined via their interfaces with other components to implement machine processing. A component may be a packaged functional hardware unit designed for use with other components and is part of a program that typically performs a specific one of the related functions.
[0171] Components may constitute software components (e.g., code embodied on a machine-readable medium) or hardware components. A "hardware component" is a tangible unit that is capable of performing certain operations and may be configured or arranged in some physical manner. In various examples, one or more computer systems (e.g., stand-alone computer systems, client computer systems, or server computer systems) or one or more hardware components of a computer system (e.g., a processor or group of processors) may be configured by software (e.g., an application or application portion) as hardware components that operate to perform certain operations described herein.
[0172] Hardware components may also be implemented mechanically, electrically, or in any appropriate combination thereof. For example, a hardware component may include a dedicated circuit system or logic that is permanently configured to perform certain operations. A hardware component may be a dedicated processor, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). A hardware component may also include a programmable logic or circuit system that is temporarily configured to perform certain operations by software. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific component of a machine) that is uniquely customized to perform the configured function, and is no longer a general-purpose processor. It will be recognized that the decision to implement a hardware component mechanically, in a dedicated and permanently configured circuit, or in a circuit system that is temporarily configured (e.g., configured by software) may be driven for cost and time considerations. Therefore, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, whether it is a physically constructed, permanently configured (e.g., hardwired) entity or a temporarily configured (e.g., programmed) entity that operates in some way or performs certain operations described herein.
[0173] Consider an example where hardware components are temporarily configured (e.g., programmed), without each of the hardware components being configured or instantiated at any one time. For example, where the hardware components include a general-purpose processor that is configured by software to become a special-purpose processor, the general-purpose processor can be configured into different special-purpose processors (e.g., including different hardware components) at different times. The software configures one or more specific processors accordingly, such as to constitute a specific hardware component at one time and to constitute different hardware components at different times.
[0174] Hardware components can provide information to other hardware components and receive information from other hardware components. Accordingly, the described hardware components can be considered to be communicatively coupled. In the case of multiple hardware components simultaneously, communication can be realized by (for example, by appropriate circuits and buses) signal transmission between or among two or more hardware components in the hardware components. In the example that multiple hardware components are configured or instantiated at different times, the communication between such hardware components can be realized, for example, by the storage and acquisition of information in the memory structure to which multiple hardware components have access rights. For example, a hardware component can perform an operation and store the output of the operation in a memory device communicatively coupled thereto. Then, another hardware component can access the memory device at a subsequent time to obtain the stored output and process it. The hardware component can also initiate communication with an input device or an output device, and can operate on resources (for example, a collection of information).
[0175] The various operations of the example methods described herein may be performed at least in part by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform related operations. Whether temporarily configured or permanently configured, such processors may constitute a processor-implemented component that operates to perform one or more operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the method described herein may be implemented at least in part by a processor using one or more specific processors as hardware examples. For example, at least some of the operations of the method may be performed by one or more processors 1002 or processor-implemented components. In addition, one or more processors may also operate to support the execution of related operations in a "cloud computing" environment or operate as "software as a service" (SaaS). For example, at least some of the operations may be performed by a computer group (as an example of a machine including a processor), where these operations may be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of certain operations may be distributed between processors, not residing only in a single machine, but deployed across several machines. In some examples, the processor or processor-implemented components may be located in a single geographic location (e.g., within a home environment, an office environment, or a server farm). In other examples, the processor or processor-implemented components may be distributed across several geographic locations.
[0176] "Computer-readable storage media" refers to both machine storage media and transmission media. Thus, the term includes both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" mean the same thing and may be used interchangeably in this disclosure.
[0177] "Ephemeral messages" are messages that are accessible for a limited duration. Ephemeral messages can be text, images, videos, etc. The access time for ephemeral messages can be set by the sender of the message. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is transient.
[0178] "Machine storage media" refers to a single or multiple storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Thus, the term should be considered to include, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memory including, by way of example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGA, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage media", "device storage media", "computer storage media" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage media", "device storage media", "computer storage media" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are encompassed by the term "signal media".
[0179] “Non-transitory computer-readable storage medium” refers to a tangible medium capable of storing, encoding, or carrying instructions for machine execution.
[0180] “Signal medium” refers to any intangible medium that is capable of storing, encoding or carrying instructions for machine execution and includes digital or analog communication signals, or other intangible media that facilitates the communication of software or data. The term “signal medium” should be deemed to include any form of modulated data signal, carrier wave, etc. The term “modulated data signal” means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. The terms “transmission medium” and “signal medium” mean the same thing and may be used interchangeably in this disclosure.
[0181] Changes and modifications may be made to the disclosed examples without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the appended claims.
Claims
1. A method comprising: receiving, by one or more processors, a video including a depiction of a real-world object in a real-world environment; accessing a segmentation associated with the real-world object; removing a depiction of the real-world object from an area of a first frame of the video; as well as The first frame of the video and one or more previous frames preceding the first frame are processed by a machine learning model to generate a new frame in which a portion of the first frame has been blended into the area from which the depiction of the real-world object has been removed.
2. The method of claim 1 further comprising generating the new frame of the video for display, the new frame not including the depiction of the real-world object, and in the new frame, the portion of the first frame has been blended into the area from which the depiction of the real-world object has been removed.
3. The method according to any one of claims 1 to 2, further comprising: An input defining the segmentation is received.
4. The method according to any one of claims 1 to 3, further comprising: Display multiple augmented reality (AR) experiences; Receiving input selecting a given AR experience from the plurality of AR experiences; as well as The segmentation associated with the selected given AR experience is obtained.
5. The method according to any one of claims 1 to 4, wherein: The machine learning model includes a generative adversarial network (GAN).
6. The method according to any one of claims 1 to 5, wherein: Processing the first frame and the one or more previous frames by the machine learning model includes: generating a modified first frame in response to removing the depiction of the real-world object from the region of the first frame; generating a modified previous frame in response to applying the segmentation to a previous frame of the one or more previous frames to remove the depiction of the real-world object from the previous frame; and An encoder is applied to the modified first frame and the modified previous frame to generate a first plurality of features associated with the modified first frame and a second plurality of features associated with the modified previous frame.
7. The method according to claim 6, further comprising: selecting a subset of the first plurality of features; selecting a subset of the second plurality of features; generating a combined subset of features by combining the subset of the first plurality of features with the subset of the second plurality of features based on a similarity metric associated with the first frame and the previous frame; as well as A decoder is applied to the combined subset of features to generate the new frame.
8. The method according to claim 7, wherein: The combined subset of features is generated based on a weighted average of the subset of the first plurality of features and the subset of the second plurality of features.
9. The method according to claim 8, further comprising: determining an optical flow frame representing a difference or distance between the first frame and the previous frame; as well as The similarity measure is calculated based on the optical flow frame.
10. The method according to claim 8, further comprising: determining a camera movement difference between the first frame and the previous frame; as well as The similarity measure is calculated based on the camera movement difference.
11. The method according to claim 8, further comprising: multiplying the similarity measure with the subset of the first plurality of features to generate a first adjusted set of features; Calculating the difference between value 1 and the similarity measure; multiplying the difference with the subset of the second plurality of features to generate a second adjusted set of features; as well as The first set of adjusted features and the second set of adjusted features are added to generate the combined subset of features.
12. The method according to claim 7, wherein: Applying the decoder comprises: identifying one or more features of the first plurality of features not included in the subset of the first plurality of features; and The identified one or more features are added to the combined subset of features to generate a complete set of features, wherein the decoder is applied to the complete set of features to generate the new frame.
13. The method according to any one of claims 1 to 12, further comprising training the machine learning model by: receiving training data comprising a plurality of training videos depicting one or more training real-world objects and associated ground-truth videos from which the one or more training real-world objects have been removed, each of the plurality of training videos depicting a different real-world environment; accessing a first training video of the plurality of training videos that is associated with a first ground-truth video of the ground-truth videos from which the one or more training real-world objects have been removed; modifying a first frame of the first training video and a second frame adjacent to the first frame to remove a given training real-world object from among the one or more training real-world objects depicted in the first frame and the second frame; applying the machine learning model to the first frame and the second frame of the first training video to estimate a new training frame in which portions of the first frame and the second frame have been blended into regions from which a depiction of the given one of the one or more training real-world objects has been removed; Calculating a deviation between the estimated new training frame and a true frame in the first true video corresponding to a playback position of the first frame and the second frame of the first training video; as well as One or more parameters of the machine learning model are updated based on the calculated deviations.
14. The method according to any one of claims 1 to 13, further comprising: A full body segmentation machine learning model is applied to the video to estimate the segmentation associated with the real-world object, wherein the depiction of the real-world object that is removed corresponds to the full body of a person.
15. The method according to any one of claims 1 to 14, further comprising: Obtaining one or more augmented reality (AR) objects associated with an AR experience; as well as The one or more AR objects are superimposed over a portion of the new frame corresponding to the area from which the depiction of the real-world object has been removed.
16. The method according to claim 15, wherein: The one or more AR objects include a fashion item.
17. The method according to claim 15, wherein: The one or more AR objects include items of furniture.
18. The method according to claim 15, wherein: The depiction of the real-world object being removed corresponds to a watermark in the video.
19. A system comprising: A processor configured to perform operations comprising: receiving a video including a depiction of a real-world object in a real-world environment; accessing a segmentation associated with the real-world object; removing the depiction of the real-world object from an area of the first frame of the video; and The first frame of the video and one or more previous frames preceding the first frame are processed by a machine learning model to generate a new frame in which a portion of the first frame has been blended into the area from which the depiction of the real-world object has been removed.
20. A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising: receiving a video including a depiction of a real-world object in a real-world environment; accessing a segmentation associated with the real-world object; removing a depiction of the real-world object from an area of a first frame of the video; as well as The first frame of the video and one or more previous frames preceding the first frame are processed by a machine learning model to generate a new frame in which a portion of the first frame has been blended into the area from which the depiction of the real-world object has been removed.