Synthetic view for try-on experience
The synthetic view system solves the problem of users having difficulty viewing small fashion items by detecting and generating AR items in the view of the second camera device, improving the efficiency and user experience of the virtual try-on experience.
Patent Information
- Application Number
- CN202480009065.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-25
- Filing Date
- 2024-01-23
- Publication Date
- 2025-09-05
AI Technical Summary
In the prior art, when users use electronic devices for a virtual try-on experience, it is difficult to effectively view the details of small fashion items in the real world, and the operation is cumbersome, which affects the experience.
Through the synthetic view system, a fashion item being worn by a person in a first camera view is detected, and a second camera view is generated, an AR item visually similar to the fashion item is synthesized, and the second image is modified to present a synthetic view of the person wearing the fashion item.
The efficiency and experience of users viewing fashion items on electronic devices are improved, without the need for awkward operation of the camera device, and the time for trying on and the consumption of system resources are reduced.
Smart Images

Figure CN120604270A_ABST
Abstract
Description
[0001] Priority claim
[0002] This application claims the benefit of priority to U.S. patent application serial number 18 / 159,554, filed on January 25, 2023, which is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure generally relates to using a messaging application to generate images and provide a virtual try-on experience. Background Art
[0004] Augmented reality (AR) is the modification of the real environment with virtual objects, visual effects, or additional information. For example, in virtual reality (VR), the user is completely immersed in the virtual world, while in AR, the user is immersed in a world where virtual objects are combined with the real world or superimposed on the real world. AR systems are designed to generate and present virtual objects that interact realistically with the real-world environment and with each other. Examples of AR applications can include single-player or multi-player video games, instant messaging systems, etc.
[0005] Brief description of the several views of the accompanying drawings
[0006] In the accompanying drawings (which are not necessarily drawn to scale), similar reference numerals may describe similar components in different views. To easily identify the discussion of any particular element or action, the most significant digit or digits in a reference numeral refer to the figure in which the element is first introduced. Some non-limiting examples are shown in the figures of the accompanying drawings, in which:
[0007] Figure 1 is a diagrammatic representation of a networked environment in which the present disclosure may be deployed, according to some examples.
[0008] Figure 2 is a diagrammatic representation of a messaging client application according to some examples.
[0009] Figure 3 is a diagrammatic representation of data structures maintained in a database according to some examples.
[0010] Figure 4 is a diagrammatic representation of messages according to some examples.
[0011] Figure 5 is a block diagram illustrating an example synthetic view system according to some examples.
[0012] Figures 6 to 9 is a diagrammatic representation of inputs and outputs of a synthetic view system according to some examples.
[0013] Figure 10is a flow diagram illustrating example operation of a synthetic view system according to some examples.
[0014] Figure 11 is a diagrammatic representation of a machine in the form of a computer system according to some examples within which a set of instructions may be executed, causing the machine to perform any one or more of the methodologies discussed herein.
[0015] Figure 12 is a block diagram illustrating a software architecture in which examples may be implemented. DETAILED DESCRIPTION
[0016] The following description includes systems, methods, techniques, instruction sequences, and computer program products that embody illustrative examples of the present disclosure. In the following description, for illustrative purposes, many specific details are set forth to provide an understanding of the various examples. However, it will be apparent to those skilled in the art that the examples can be practiced without these specific details. Typically, well-known instruction instances, protocols, structures, and techniques are not necessarily shown in detail.
[0017] Typically, messaging applications and other social networking platforms allow users to view a live or real-time camera feed depicting the user wearing different items. For example, a user can activate a virtual try-on experience in which one or more AR items are placed on the user depicted in the image in real time. This provides the user with the ability to visualize how different products will look on the user before the user purchases the product. Sometimes, the item being tried on and depicted as being worn by the user is small and difficult to see in the image captured by the camera of the device used to present the try-on experience. To get a better idea of how the object looks on the user, the user needs to move the camera closer to the object of interest or turn / move their body or head in awkward ways to get a different perspective.
[0018] For example, a user may be wearing earrings and may be depicted as wearing earrings in a camera feed captured by the device's front-facing camera. The camera feed may depict the user's entire face and the user's ear on which the earring is worn. However, since the earring appears relatively small in the overall image, it is difficult to see the earring. The user must manipulate the phone in an awkward manner to bring the phone closer to the ear of interest on which the earring is worn or to obtain a better view, such as capturing the image from a side perspective. This results in a distorted image of the ear and earring being presented, and the user's face may no longer be visible. This makes it difficult for the user to accurately understand and visualize how the earring looks relative to the user's entire face. Furthermore, manipulating the camera and the device with the camera in an awkward manner to obtain a correct photo of the earring takes time and is very disruptive to the overall experience. Sometimes, when the camera is too close to the AR earring (in the case of AR earrings) and no longer captures an image of the user's face, the facial tracking information used to generate the display of the earring is completely lost. When this happens, the AR earrings begin to behave erratically or may be removed from the displayed image entirely, which undermines the overall illusion that the earrings are part of the real-world image captured by the camera device.
[0019] The disclosed technology seeks to improve the efficiency of using electronic devices by intelligently and automatically generating images of fashion items detected in a real-time camera feed from different composite views in a simple, seamless, and intuitive manner. The disclosed technology detects a fashion item being worn by a person depicted in a camera feed captured and accessed in real time. The camera feed may depict a first camera view of the person, such as a front view (showing only the frontal profile of the face). The disclosed technology can use the first camera view image to synthesize a second camera view of the person, such as a side view or side profile of the face. The disclosed technology also modifies the fashion item to be depicted from an angle corresponding to the first camera view and synthesizes a new image of the second camera view depicting the person wearing the fashion item. This allows a person to capture a single image representing one view of a person wearing the fashion item and seamlessly (without awkward camera manipulation) access or view one or more additional images depicting the same person wearing the fashion item from different views or perspectives. The additional views can be presented together with or separately from the single captured image.
[0020] This can reduce the overall time and expense a user spends trying on different fashion items (whether real-world fashion items or AR fashion items) such as shoes, shirts, earrings, watches, or other fashion items. As used herein, "clothing," "fashion items," and "apparel" are used interchangeably and should be understood to have the same meaning. Clothing, clothing, or fashion items may include shirts, pants, skirts, dresses, shoes, wallets, furniture items, household items, goggles, lenses, AR logos, AR badges, trousers, shorts, skirts, jackets, T-shirts, blouses, glasses, jewelry, earrings, bunny ears, hats, earmuffs, or any other suitable items or objects.
[0021] Specifically, the disclosed technology accesses a first image depicting a first portion of a person's body from a first camera view. The disclosed technology detects a fashion item being worn by the person in the first image and, based on the first image, generates a second image depicting a second portion of the person's body from a second camera view. The disclosed technology generates an AR item from the second camera view that visually resembles the fashion item. The disclosed technology modifies the second image with the AR item to present a composite view of the second portion of the person's body wearing the AR item corresponding to the fashion item. This improves the user's overall experience when using the electronic device and reduces the total amount of system resources required to complete tasks.
[0022] Networked computing environment
[0023] Figure 1 1 is a block diagram illustrating an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network. The messaging system 100 includes multiple instances of a client device 102, each of which hosts several applications, including a messaging client 104 and other external applications 109 (e.g., third-party applications). Each messaging client 104 is communicatively coupled to other instances of the messaging client 104 (e.g., hosted on corresponding other client devices 102), a messaging server system 108, and an external app server 110 via a network 112 (e.g., the Internet). The messaging clients 104 can also communicate with locally hosted third-party applications (also referred to as "external applications" and "external apps") 109 using an application programming interface (API).
[0024] The client device 102 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the client device 102 can operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The client device 102 can include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web device, a network router, a network switch, a network bridge, or any machine capable of performing the disclosed operations. In addition, although only a single client device 102 is shown, the term "client device" should also be considered to include a collection of machines that perform the disclosed operations individually or in combination.
[0025] In some examples, client device 102 may include AR glasses or an AR head-mounted device in which virtual content is displayed within the lenses of the glasses while the user views the real-world environment through the lenses. For example, an image may be presented on a transparent display, allowing the user to simultaneously view the content presented on the display and real-world objects.
[0026] The messaging clients 104 are able to communicate and exchange data with other messaging clients 104 and messaging server systems 108 via the network 112. The data exchanged between the messaging clients 104 and between the messaging clients 104 and messaging server systems 108 includes functions (e.g., commands to activate functions) and payload data (e.g., text, audio, video, or other multimedia data).
[0027] The messaging server system 108 provides server-side functionality to certain messaging clients 104 via the network 112. Although certain functionality of the messaging system 100 is described herein as being performed by either the messaging client 104 or the messaging server system 108, the location of certain functionality within the messaging client 104 or within the messaging server system 108 may be a matter of design choice. For example, it may be technically preferable to initially deploy certain technologies and functionality within the messaging server system 108 but later migrate the technologies and functionality to the messaging client 104 if the client device 102 has sufficient processing power.
[0028] The messaging server system 108 supports various services and operations provided to the messaging clients 104. Such operations include sending data to the messaging clients 104, receiving data from the messaging clients 104, and processing data generated by the messaging clients 104. By way of example, this data may include message content, client device information, geographic location information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. The exchange of data within the messaging system 100 is invoked and controlled by functions available through the user interface of the messaging clients 104.
[0029] Turning now specifically to the messaging server system 108, an API server 116 is coupled to the application server 114 and provides a programming interface to the application server 114. The application server 114 is communicatively coupled to a database server 120, which facilitates access to a database 126 that stores data associated with messages processed by the application server 114. Similarly, a web server 128 is coupled to the application server 114 and provides a web-based interface to the application server 114. To this end, the web server 128 handles incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.
[0030] The API server 116 receives and sends message data (e.g., commands and message payloads) between the client device 102 and the application server 114. Specifically, the API server 116 provides a set of interfaces (e.g., routines and protocols) that can be called or queried by the messaging client 104 to activate functionality of the application server 114. The API server 116 exposes various functionality supported by the application server 114, including: account registration; login functionality; sending messages from a particular messaging client 104 to another messaging client 104 via the application server 114; sending media files (e.g., images or videos) from a messaging client 104 to the messaging server 118 for possible access by another messaging client 104; setting up collections of media data (e.g., stories); retrieving a friend list of the user of the client device 102; retrieving such collections; retrieving messages and content; adding and removing entities (e.g., friends) from an entity graph (e.g., a social graph); locating friends in a social graph; and opening application events (e.g., related to the messaging client 104).
[0031] The application server 114 hosts several server applications and subsystems, including, for example, a messaging server 118, an image processing server 122, and a social network server 124. The messaging server 118 implements several message processing technologies and functions, particularly those related to the aggregation and other processing of content (e.g., text and multimedia content) included in messages received from multiple instances of the messaging client 104. As will be described in further detail, text and media content from multiple sources can be aggregated into collections of content (e.g., called stories or libraries). These collections are then made available to the messaging client 104. Given the hardware requirements for other processor- and memory-intensive data processing, such processing of data can also be performed on the server side by the messaging server 118.
[0032] The application server 114 also includes an image processing server 122 that is dedicated to performing various image processing operations, typically with respect to images or videos within the payload of messages sent from or received at the messaging server 118 .
[0033] The image processing server 122 is used to implement the enhancement system 208 ( Figure 2 The scanning functionality includes activating and providing one or more AR experiences on the client device 102 when an image is captured by the client device 102. Specifically, the messaging client 104 on the client device 102 can be used to activate the camera. The camera displays one or more real-time images or videos and one or more icons or identifiers of one or more AR experiences to the user. The user can select a given identifier from the identifiers to launch the corresponding AR experience or perform the desired image modification.
[0034] The social network server 124 supports various social networking functions and services and makes these functions and services available to the messaging server 118. To this end, the social network server 124 maintains and accesses an entity graph 308 (e.g., Figure 3 ). Examples of functions and services supported by the social network server 124 include identifying other users of the messaging system 100 who have relationships with a particular user or who the particular user is "following," and also identifying interests and other entities of a particular user.
[0035] Returning to the messaging client 104, the features and functionality of the external resource (e.g., a third-party application 109 or mini-program) are made available to the user via the interface of the messaging client 104. The messaging client 104 receives a user selection of an option to launch or access features of an external resource (e.g., a third-party resource) such as an external app 109. The external resource can be a third-party application (external app 109) installed on the client device 102 (e.g., a "native application"), or a small-scale version of a third-party application (e.g., a "mini-program") hosted on the client device 102 or remote from the client device 102 (e.g., on an external resource or app server 110). The small-scale version of the third-party application includes a subset of the features and functionality of the third-party application (e.g., a full-scale native version of a third-party standalone application) and is implemented using a markup language document. In one example, the small-scale version of the third-party application (e.g., a "mini-program") is a web-based markup language version of the third-party application and is embedded in the messaging client 104. In addition to using markup language documents (eg, *ml files), applets may include scripting languages (eg, .*js files or .json files) and style sheets (eg, *ss files).
[0036] In response to receiving a user selection of an option for launching or accessing a feature of an external resource (external app 109), the messaging client 104 determines whether the selected external resource is a web-based external resource or a locally installed external application. In some cases, an external application 109 installed locally on the client device 102 can be launched independently of and separately from the messaging client 104, for example, by selecting an icon corresponding to the external application 109 on the home screen of the client device 102. A small-scale version of such an external application can be launched or accessed via the messaging client 104, and in some examples, no part of the small-scale external application can be (or a limited part can be) accessed outside of the messaging client 104. The small-scale external application can be launched by the messaging client 104 receiving a markup language document associated with the small-scale external application from the external app server 110 and processing such a document.
[0037] In response to determining that the external resource is a locally installed external application 109, the messaging client 104 instructs the client device 102 to launch the external application 109 by executing locally stored code corresponding to the external application 109. In response to determining that the external resource is a web-based resource, the messaging client 104 communicates with the external app server 110 to obtain a markup language document corresponding to the selected resource. The messaging client 104 then processes the obtained markup language document to present the web-based external resource within the user interface of the messaging client 104.
[0038] The messaging client 104 can notify the user of the client device 102 or other users associated with such user (e.g., "friends") of activities occurring in one or more external resources. For example, the messaging client 104 can provide participants in a conversation (e.g., a chat session) within the messaging client 104 with notifications regarding current or recent use of an external resource by one or more members of the user group. One or more users can be invited to join an active external resource or a recently used but currently inactive external resource (in the friend group) can be launched. The external resource can provide participants in the conversation (each using a corresponding messaging client 104) with the ability to share items, conditions, states, or locations within the external resource with one or more members of the user group entering the chat session. The shared items can be interactive chat cards that members of the chat can interact with, for example, to launch a corresponding external resource, view specific information within the external resource, or take members of the chat to a specific location or state within the external resource. Within a given external resource, a response message can be sent to the user on the messaging client 104. The external resource may selectively include different media items in the response based on the current context of the external resource.
[0039] The messaging client 104 can present a list of available external resources (e.g., third-party or external applications 109 or applets) to the user to launch or access a given external resource. The list can be presented in a context-sensitive menu. For example, the icons representing different external applications in the external applications 109 (or applets) can change based on how the user launches the menu (e.g., from a conversational interface or from a non-conversational interface).
[0040] System Architecture
[0041] Figure 21 is a block diagram illustrating further details regarding the messaging system 100 according to some examples. Specifically, the messaging system 100 is shown as including a messaging client 104 and an application server 114. The messaging system 100 includes several subsystems that are supported on the client side by the messaging client 104 and on the server side by the application server 114. These subsystems include, for example, a transient timer system 202, a collection management system 204, an enhancement system 208, a mapping system 210, a gaming system 212, and an external resource system 220.
[0042] The transient timer system 202 is responsible for implementing temporary or time-limited access to content by the messaging client 104 and the messaging server 118. The transient timer system 202 includes several timers that selectively enable access to messages and associated content (e.g., for presentation and display) via the messaging client 104 based on duration and display parameters associated with the message or collection of messages (e.g., a story). Additional details regarding the operation of the transient timer system 202 are provided below.
[0043] The collection management system 204 is responsible for managing collections or collections of media (e.g., collections of text, images, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into "event libraries" or "event stories." Such collections can be made available for a specified time period (e.g., the duration of the event to which the content relates). For example, content related to a concert can be made available as a "story" for the duration of the concert. The collection management system 204 can also be responsible for publishing an icon to the user interface of the messaging client 104 that provides notification of the existence of a particular collection.
[0044] The collection management system 204 also includes a curation interface 206 that allows collection managers to manage and curate specific content collections. For example, the curation interface 206 enables event organizers to curate content collections related to a specific event (e.g., removing inappropriate content or redundant messages). In addition, the collection management system 204 uses machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, users can be paid compensation for including user-generated content in a collection. In such cases, the collection management system 204 operates to automatically pay such users for use of their content.
[0045] The enhancement system 208 provides various functions that enable users to enhance (e.g., annotate or otherwise modify or edit) media content associated with a message. For example, the enhancement system 208 provides functions related to generating and publishing media overlays for messages processed by the messaging system 100. The enhancement system 208 is operable to provide media overlays or enhancements (e.g., image filters) to the messaging client 104 based on the geographic location of the client device 102. In another example, the enhancement system 208 is operable to provide media overlays to the messaging client 104 based on other information such as social network information of the user of the client device 102. Media overlays can include audio and visual content and visual effects. Examples of audio and visual content include pictures, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects can be applied to media content items (e.g., photos) at the client device 102. For example, a media overlay can include text, graphic elements, or images that can be superimposed on a photo taken by the client device 102. In another example, the media overlay includes a location identification overlay (e.g., Venice Beach), the name of a live event, or a business name overlay (e.g., Beach Cafe). In another example, the augmentation system 208 uses the geographic location of the client device 102 to identify a media overlay that includes the name of a business at the geographic location of the client device 102. The media overlay may include other logos associated with the business. The media overlay may be stored in the database 126 and accessed through the database server 120.
[0046] In some examples, the enhancement system 208 provides a user-based publishing platform that enables a user to select a geographic location on a map and upload content associated with the selected geographic location. The user can also specify situations in which a particular media overlay should be provided to other users. The enhancement system 208 generates a media overlay that includes the uploaded content and associates the uploaded content with the selected geographic location.
[0047] In other examples, the augmentation system 208 provides a merchant-based publishing platform that enables merchants to select specific media overlays associated with a geographic location via a bidding process. For example, the augmentation system 208 associates the media overlay of the highest-bidding merchant with the corresponding geographic location for a predefined amount of time. The augmentation system 208 communicates with the image processing server 122 to obtain an AR experience and presents an identifier for such an experience in one or more user interfaces (e.g., as an icon on a live image or video, or as a thumbnail or icon in an interface dedicated to the identifier of the presented AR experience). Once the AR experience is selected, one or more images, videos, or AR graphical elements are retrieved and presented as an overlay on top of the image or video captured by the client device 102. In some cases, the camera is switched to a front view (e.g., the front-facing camera of the client device 102 is activated in response to activation of a particular AR experience), and an image from the front-facing camera of the client device 102, rather than the rear-facing camera of the client device 102, begins to be displayed on the client device 102. One or more images, videos, or AR graphical elements are retrieved and rendered as an overlay on top of the image captured and displayed by the front-facing camera of the client device 102 .
[0048] In other examples, the augmentation system 208 can communicate and exchange data with another augmentation system 208 on another client device 102 and with a server via the network 112. The exchanged data may include: a session identifier identifying the shared AR session; a transformation between the first client device 102 and the second client device 102 (e.g., a plurality of client devices 102 including the first device and the second device), the transformation used to align the shared AR session to a common origin; a common coordinate system; functionality (e.g., a command to activate the functionality); and other payload data (e.g., text, audio, video, or other multimedia data).
[0049] The augmentation system 208 sends the transformation to the second client device 102 so that the second client device 102 can adjust the AR coordinate system based on the transformation. In this way, the first client device 102 and the second client device 102 synchronize their coordinate systems to display content in the AR session. Specifically, the augmentation system 208 calculates the origin of the second client device 102 in the coordinate system of the first client device 102. The augmentation system 208 can then determine an offset in the coordinate system of the second client device 102 based on the position of the origin in the coordinate system of the second client device 102 from the perspective of the second client device 102. The offset is used to generate a transformation so that the second client device 102 generates AR content according to a common coordinate system with the first client device 102.
[0050] The augmentation system 208 can communicate with the client device 102 to establish an individual or shared AR session. The augmentation system 208 can also be coupled to the messaging server 118 to establish an electronic group communication session (e.g., group chat, instant messaging) for the client device 102 in the shared AR session. The electronic group communication session can be associated with a session identifier provided by the client device 102 to obtain access to the electronic group communication session and the shared AR session. In one example, the client device 102 first obtains access to the electronic group communication session and then obtains a session identifier in the electronic group communication session that allows the client device 102 to access the shared AR session. In some examples, the client device 102 is able to access the shared AR session without the assistance of the augmentation system 208 in the application server 114 or without communicating with the augmentation system 208 in the application server 114.
[0051] The mapping system 210 provides various geolocation capabilities and supports the presentation of map-based media content and messages by the messaging client 104. For example, the mapping system 210 enables the display of user icons or avatars (e.g., stored in the profile data 316) on a map to indicate the current or past locations of the user's "friends," as well as media content (e.g., a collection of messages including photos and videos) generated by such friends within the context of the map. For example, a message posted by a user to the messaging system 100 from a particular geolocation can be displayed to the particular user's "friends" on the map interface of the messaging client 104 within the context of that particular location on the map. A user can also share his or her location and status information with other users of the messaging system 100 via the messaging client 104 (e.g., using an appropriate status avatar), where the location and status information is similarly displayed to selected users within the context of the mapping interface of the messaging client 104.
[0052] The gaming system 212 provides various gaming functions within the context of the messaging client 104. The messaging client 104 provides a gaming interface that provides a list of available games (e.g., web-based games or web-based applications) that can be launched by a user within the context of the messaging client 104 and played with other users of the messaging system 100. The messaging system 100 also enables a particular user to invite other users to play a particular game by sending an invitation to such other users from the messaging client 104. The messaging client 104 also supports both voice and text messaging (e.g., chat) within the context of gaming, provides leaderboards for gaming, and also supports the provision of in-game rewards (e.g., game coins and items).
[0053] The external resource system 220 provides the messaging client 104 with an interface for communicating with the external app server 110 to launch or access external resources. Each external resource (app) server 110 hosts, for example, an application based on a markup language (e.g., HTML5) or a small-scale version of an external application (e.g., a game, utility, payment, or ride-sharing application external to the messaging client 104). The messaging client 104 can launch a web-based resource (e.g., an application) by accessing an HTML5 file from the external resource (app) server 110 associated with the web-based resource. In some examples, the applications hosted by the external resource server 110 are programmed in JavaScript using a software development kit (SDK) provided by the messaging server 118. The SDK includes an API with functions that can be called or activated by the web-based application. In some examples, the messaging server 118 includes a JavaScript library that provides a given third-party resource with access to certain user data of the messaging client 104. HTML5 is used as an example technology for programming games, but applications and resources programmed based on other technologies can also be used.
[0054] To integrate the SDK's functionality into a web-based resource, the external resource (app) server 110 downloads the SDK from the messaging server 118 or receives the SDK in other ways. Once downloaded or received, the SDK is included as part of the application code of the web-based external resource. The code of the web-based resource can then call or activate certain functions of the SDK to integrate the features of the messaging client 104 into the web-based resource.
[0055] The SDK stored on the messaging server 118 effectively provides a bridge between external resources (e.g., third-party or external applications 109 or applets) and the messaging client 104. This provides users with a seamless experience of communicating with other users on the messaging client 104 while also preserving the look and feel of the messaging client 104. In order to bridge the communication between the external resources and the messaging client 104, in some examples, the SDK facilitates communication between the external resource server 110 and the messaging client 104. In some examples, the WebViewJavaScriptBridge running on the client device 102 establishes two one-way communication channels between the external resources and the messaging client 104. Messages are sent asynchronously between the external resources and the messaging client 104 via these communication channels. Each SDK function call is sent as a message and a callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message with the callback identifier.
[0056] By using the SDK, not all information from the messaging client 104 is shared with the external resource server 110. The SDK limits which information is shared based on the requirements of the external resource. In some examples, each external resource server 110 provides an HTML5 file corresponding to a web-based external resource to the messaging server 118. The messaging server 118 can add a visual representation (e.g., box art or other graphics) of the web-based external resource in the messaging client 104. Once the user selects the visual representation through the GUI of the messaging client 104 or instructs the messaging client 104 to access a feature of the web-based external resource, the messaging client 104 obtains the HTML5 file and instantiates the resources required to access the feature of the web-based external resource.
[0057] The messaging client 104 presents a graphical user interface (GUI) (e.g., a login page or title screen) for the external resource. During, before, or after presenting the login page or title screen, the messaging client 104 determines whether the launched external resource has been previously authorized to access the user data of the messaging client 104. In response to determining that the launched external resource has been previously authorized to access the user data of the messaging client 104, the messaging client 104 presents another GUI of the external resource, which includes the functions and features of the external resource. In response to determining that the launched external resource has not been previously authorized to access the user data of the messaging client 104, after displaying the login page or title screen of the external resource for a threshold time period (e.g., 3 seconds), the messaging client 104 slides a menu for authorizing the external resource to access the user data (e.g., animating the menu to emerge from the bottom of the screen to the middle or other part of the screen). The menu identifies the type of user data that the external resource will be authorized to use. In response to receiving the user selection of the accept option, the messaging client 104 adds the external resource to the list of authorized external resources and enables the external resource to access the user data from the messaging client 104. In some examples, the messaging client 104 authorizes the external resource to access the user data according to the OAuth 2 framework.
[0058] The messaging client 104 controls the type of user data shared with external resources based on the type of authorized external resources. For example, external resources including full-scale external applications (e.g., third-party or external application 109) are provided with access to a first type of user data (e.g., only a two-dimensional (2D) avatar of the user with or without different avatar characteristics). As another example, external resources including small-scale versions of external applications (e.g., web-based versions of third-party applications) are provided with access to a second type of user data (e.g., payment information, a 2D avatar of the user, a 3D avatar of the user, and avatars with various avatar characteristics). Avatar characteristics include different ways to customize the look and feel of an avatar (e.g., different poses, facial features, clothing, etc.).
[0059] Synthetic view system 224 accesses a first image of a first camera view that depicts a first portion of a person's body. Synthetic view system 224 detects a fashion item being worn by the person in the first image and, based on the first image, generates a second image of a second camera view that depicts a second portion of the person's body. Synthetic view system 224 generates an augmented reality (AR) item for the second camera view that visually resembles the fashion item and modifies the second image with the AR item to present a synthetic view of the second portion of the person's body wearing the AR item corresponding to the fashion item. The AR item visually resembles the fashion item by having the same or similar appearance as the fashion item. That is, the AR item has the same visual features and attributes as the fashion item.
[0060] Combined with the following Figure 5 An exemplary implementation of the synthetic view system 224 is shown and described.
[0061] Data Architecture
[0062] Figure 3 is a diagram illustrating a data structure 300 that may be stored in the database 126 of the messaging server system 108 according to certain examples. Although the contents of the database 126 are shown as including several tables, it should be understood that data may be stored in other types of data structures (e.g., an object-oriented database).
[0063] The database 126 includes message data stored in the message table 302. For any particular message, the message data includes at least message sender data, message recipient (or receiver) data, and payload. Figure 4 Additional details regarding information that may be included in a message and included within the message data stored in message table 302 are described.
[0064] The entity table 306 stores entity data and is linked (e.g., by reference) to the entity graph 308 and profile data 316. Entities whose records are maintained within the entity table 306 may include individuals, corporate entities, organizations, objects, places, events, and the like. Regardless of the entity type, any entity for which the messaging server system 108 stores data may be an identified entity. Each entity is provided with a unique identifier and an entity type identifier (not shown).
[0065] The entity graph 308 stores information about relationships and associations between entities. By way of example only, such relationships may be social, professional (e.g., working in a common company or organization), interest-based, or activity-based.
[0066] The profile data 316 stores various types of profile data about a particular entity. Based on the privacy settings specified by the particular entity, the profile data 316 can be selectively used and presented to other users of the messaging system 100. In the case where the entity is a person, the profile data 316 includes, for example, the user's name, phone number, address, settings (e.g., notification and privacy settings), and an avatar representation (or a collection of such avatar representations) selected by the user. The particular user can then selectively include one or more of these avatar representations within the content of messages transmitted via the messaging system 100 and on a map interface displayed to other users by the messaging client 104. The collection of avatar representations can include a "status avatar," which presents a graphical representation of a status or activity that a user can select to transmit at a particular time.
[0067] Where the entity is a group, profile data 316 for the group may similarly include one or more avatar representations associated with the group, in addition to the group name, members, and various settings for the relevant group (eg, notifications).
[0068] The database 126 also stores enhancement data, such as overlays or filters, in an enhancement table 310. The enhancement data is associated with and applied to videos (data for videos is stored in the video table 304) and images (data for images is stored in the image table 312).
[0069] The database 126 may also store data related to individual and shared AR sessions. This data may include data transferred between an AR session client controller of a first client device 102 and another AR session client controller of a second client device 102, as well as data transferred between the AR session client controller and the augmented system 208. The data may include data used to establish a common coordinate system for a shared AR scene, transformations between devices, session identifiers, images depicting the body, skeletal joint positions, wrist joint positions, feet, and the like.
[0070] In one example, a filter is an overlay that is displayed as an overlay on an image or video during presentation to a recipient user. Filters can be of various types, including filters that a user selects from a set of filters presented to a sending user by messaging client 104 when the sending user is composing a message. Other types of filters include geolocation filters (also known as geofilters), which can be presented to a sending user based on geographic location. For example, a geolocation filter specific to a nearby or specific location can be presented by messaging client 104 within a user interface based on geographic location information determined by a global positioning system (GPS) unit of client device 102.
[0071] Another type of filter is a data filter, which can be selectively presented to the sending user by the messaging client 104 based on other input or information collected during the message creation process by the client device 102. Examples of data filters include the current temperature at a particular location, the current speed the sending user is traveling, the battery life of the client device 102, or the current time.
[0072] Other augmentation data that may be stored in the image table 312 include AR content items (eg, corresponding to an applied AR experience). AR content items or AR items may be real-time special effects and sounds that may be added to an image or video.
[0073] As described above, augmented data includes AR content items, overlays, image transformations, AR images, and similar items involving modifications that can be applied to image data (e.g., video or images). This includes real-time modifications, where the image is modified as it is captured using the device sensors of client device 102 (e.g., one or more cameras) and the modified image is then displayed on the screen of client device 102. This also includes modifications to stored content (e.g., video clips in a library that can be modified). For example, on a client device 102 with access to multiple AR content items, a user can use a single video clip with multiple AR content items to see how different AR content items will modify the stored clip. For example, by selecting different AR content items for the content, multiple AR content items applying different pseudo-random movement models can be applied to the same content. Similarly, real-time video capture can be used with displayed modifications to illustrate how the video image currently captured by the client device 102's sensors will modify the captured data. Such data can be simply displayed on the screen without being stored in memory, or the content captured by the device sensors can be recorded with or without modifications (or both) and stored in memory. In some systems, a preview feature can simultaneously show how different AR content items will look in different windows in the display. This can, for example, enable viewing multiple windows with different pseudo-random animations on the display at the same time.
[0074] Thus, using data from an AR content item and various systems or other such transformation systems that use this data to modify the content can involve: detecting objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.) in a video frame; tracking such objects as they leave, enter, and move around the field of view; and modifying or transforming such objects while tracking them. In various examples, different methods for implementing such transformations can be used. Some examples can involve: generating a three-dimensional mesh model of one or more objects and using transformations and animated textures of the models within a video to implement the transformations. In other examples, tracking of points on an object can be used to place an image or texture (which can be two-dimensional or three-dimensional) at the tracked location. In yet another example, neural network analysis of a video frame can be used to place an image, model, or texture within the content (e.g., an image or video frame). Thus, an AR content item refers to both the images, models, and textures used to create the transformations within the content, as well as the additional modeling and analysis information required to implement such transformations using object detection, tracking, and placement.
[0075] Real-time video processing can be performed using any type of video data (e.g., video streams, video files, etc.) stored in the memory of any type of computerized system. For example, a user can load a video file and store it in the device's memory, or a device's sensors can be used to generate a video stream. In addition, computer animation models can be used to process any object, such as a human face and body parts, an animal, or an inanimate object (e.g., a chair, a car, or other objects).
[0076] In some examples, when a specific modification is selected along with the content to be transformed, a computing device identifies the elements to be transformed and then detects and tracks them if present in a frame of a video. Elements of an object are modified based on the modification request, thereby transforming the frame of the video stream. For different types of transformations, the transformation of the frame of the video stream can be performed using different methods. For example, for frame transformations that primarily involve changing the form of an element of an object, characteristic points are calculated for each element of the object (e.g., using an active shape model (ASM) or other known methods). A mesh based on the characteristic points is then generated for each of at least one element of the object. This mesh is used in a subsequent stage of tracking the elements of the object in the video stream. During the tracking process, the mesh for each element is aligned with the position of each element. Additional points are then generated on the mesh. A set of first points is generated for each element based on the modification request, and a set of second points is generated for each element based on the set of first points and the modification request. The frame of the video stream can then be transformed by modifying the elements of the object based on the set of first and second points and the mesh. In such an approach, the background of the modified object may also be changed or distorted by tracking and modifying the background of the modified object.
[0077] In some examples, a transformation that uses the elements of an object to alter some regions of the object can be performed by calculating characteristic points for each element of the object and generating a grid based on the calculated characteristic points. Points are generated on the grid, and various regions are then generated based on these points. The elements of the object are then tracked by aligning the region for each element with the position of each of at least one of the elements, and the properties of the region can be modified based on a request for modification, thereby transforming the frame of the video stream. Depending on the specific modification request, the properties of the referenced region can be transformed in different ways. Such modifications can involve: changing the color of the region; removing at least some portion of the region from the frame of the video stream; including one or more new objects in the region based on the modification request; and modifying or distorting elements of the region or object. In various examples, any combination of such modifications or other similar modifications can be used. For certain models to be animated, some characteristic points can be selected as control points to determine the entire state space of options for the model animation.
[0078] In some examples of computer animation models that use face detection to transform image data, a face is detected on an image using a specific face detection algorithm (e.g., Viola-Jones). An ASM algorithm is then applied to the face region of the image to detect facial feature reference points.
[0079] Other methods and algorithms suitable for face detection can be used. For example, in some examples, features are located using landmarks, which represent distinguishable points that are present in most of the images considered. For example, for facial landmarks, the location of the left eye pupil can be used. If the initial landmarks are not identifiable (for example, in the case of a person wearing an eye patch), secondary landmarks can be used. Such a landmark identification process can be used for any such object. In some examples, a set of landmarks forms a shape. The coordinates of the points in the shape can be used to represent the shape as a vector. One shape is aligned with another shape using a similarity transformation (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between the shape points. The average shape is the average of the aligned training shapes.
[0080] In some examples, the landmark search begins with an average shape aligned with the position and size of a face determined by a global face detector. This search then repeats until convergence: the positions of shape points are adjusted by template matching the image texture around each point to propose tentative shapes, which are then fit to a global shape model. In some systems, individual template matches are unreliable, and the shape model pools the results of weak template matches to form a stronger overall classifier. This entire search is repeated at each level of the image pyramid, from coarse to fine resolution.
[0081] The transformation system can capture an image or video stream on a client device (e.g., client device 102) and perform complex image manipulations locally on the client device 102 while maintaining an appropriate user experience, computational time, and power consumption. Complex image manipulations can include size and shape changes, emotion transitions (e.g., changing a face from a frown to a smile), state transitions (e.g., aging a subject, reducing apparent age, changing gender), style transitions, application of graphical elements, and any other suitable image or video manipulations enabled by a convolutional neural network that has been configured to execute efficiently on the client device 102.
[0082] In some examples, a computer-animated model for transforming image data can be used by a system in which a user can capture an image or video stream of the user (e.g., a selfie) using a client device 102 having a neural network operating as part of a messaging client 104 operating on the client device 102. A transformation system operating within the messaging client 104 determines the presence of a face within the image or video stream and provides a modification icon associated with the computer-animated model to transform the image data, or the computer-animated model can be presented in association with an interface described herein. The modification icon includes a change that can be used to modify the basis for modifying the user's face within the image or video stream as part of the modification operation. Once the modification icon is selected, the transformation system initiates a process that transforms the user's image to reflect the selected modification icon (e.g., generating a smiley face on the user). Once the image or video stream is captured and the specified modification is selected, the modified image or video stream can be presented in a GUI displayed on the client device 102. The transformation system can implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification. That is, a user can capture an image or video stream, and once a modification icon has been selected, the modified result can be presented in real time or near real time. Furthermore, the modification can be persistent while the video stream is being captured and the selected modification icon remains toggled. A machine-learned neural network can be used to implement such modification.
[0083] The GUI presenting the modifications performed by the transformation system can provide the user with additional interactive options. Such options can be based on the interface used to initiate content capture and selection of a particular computer animation model (e.g., initiated from a content creator user interface). In various examples, the modifications can be persistent after the initial selection of the modification icon. The user can toggle the modifications on or off by tapping or otherwise selecting the face being modified by the transformation system, and store them for later viewing or browsing other areas of the imaging application. In the event that the transformation system modifies multiple faces, the user can globally toggle the modifications on or off by tapping or selecting a single face that is being modified and displayed within the graphical user interface. In some examples, each face in a plurality of face groups can be modified individually, or such modifications can be individually toggled by tapping or selecting a single face or series of faces displayed within the GUI.
[0084] The story table 314 stores data about a collection of messages and associated image, video, or audio data that are compiled into a collection (e.g., a story or library). The creation of a particular collection can be initiated by a particular user (e.g., each user whose record is maintained in the entity table 306). A user can create a "personal story" in the form of a collection of content that has been created and sent / broadcasted by the user. To this end, the user interface of the messaging client 104 can include a user-selectable icon that enables the sending user to add specific content to his or her personal story.
[0085] A collection may also constitute a "live story," which is a collection of content from multiple users that is created manually, automatically, or using a combination of manual and automatic techniques. For example, a "live story" may constitute a curated stream of user-submitted content from various locations and events. Users whose client devices have location services enabled and who are at a co-located event at a particular time may be presented with the option to contribute content to a particular live story, for example, via the user interface of the messaging client 104. Live stories may be identified to a user by the messaging client 104 based on his or her location. The end result is a "live story" told from the perspective of the community.
[0086] Another type of content collection is called a "location story," which enables users whose client devices 102 are located within a specific geographic location (e.g., on a college or university campus) to contribute to a particular collection. In some examples, contributions to location stories may require secondary authentication to verify that the end user belongs to a specific organization or other entity (e.g., is a student on a university campus).
[0087] As mentioned above, video table 304 stores video data that, in one example, is associated with messages whose records are maintained within message table 302. Similarly, image table 312 stores image data that is associated with messages whose message data is stored in entity table 306. Entity table 306 may associate various enhancements from enhancement table 310 with the various images and videos stored in image table 312 and video table 304.
[0088] The data structure 300 may also store training data for training one or more machine learning techniques (models) to generate 2D bounding boxes. The training data may include multiple training videos and corresponding ground truth bounding boxes. Images and videos may include a mixture of various real-world objects that may appear in different real-world environments, such as different rooms in a home or family. One or more machine learning techniques or models may be trained to extract features of a received input image or video and establish a relationship between the extracted features and the 2D bounding boxes of the real-world objects depicted in the image or video. Once trained, the machine learning technique may receive a new image or video and estimate a 2D bounding box for the newly received image or video.
[0089] Data communication architecture
[0090] Figure 4 is a schematic diagram illustrating the structure of a message 400, according to some examples, generated by a messaging client 104 for transmission to another messaging client 104 or a messaging server 118. The content of a particular message 400 is used to populate a message table 302, which is stored within a database 126 and accessible by a messaging server 118. Similarly, the content of the message 400 is stored in memory as "in-flight" or "in-transit" data of a client device 102 or an application server 114. The message 400 is shown as including the following example components:
[0091] Message identifier 402 : A unique identifier that identifies the message 400 .
[0092] Message text payload 404 : Text to be generated by the user via the user interface of the client device 102 and included in the message 400 .
[0093] Message image payload 406: Image data captured by the camera component of the client device 102 or retrieved from the memory of the client device 102 and included in the message 400. Image data for a message 400 sent or received may be stored in the image table 312. Message video payload 408: Video data captured by the camera component or retrieved from the memory component of the client device 102 and included in the message 400. Video data for a message 400 sent or received may be stored in the video table 304.
[0094] • Message audio payload 410 : audio data captured by a microphone or retrieved from a memory component of the client device 102 and included in the message 400 .
[0095] Message enhancement data 412: Enhancement data (e.g., filters, stickers, or other annotations or enhancements) representing enhancements to be applied to the message image payload 406, message video payload 408, or message audio payload 410 of the message 400. The enhancement data 412 for a sent or received message 400 may be stored in the enhancement table 310.
[0096] • Message duration parameter 414: A parameter value indicating the amount of time in seconds that the content of a message (eg, message image payload 406, message video payload 408, message audio payload 410) will be presented to or made accessible to a user via messaging client 104.
[0097] Message geolocation parameter 416: Geolocation data (e.g., latitude and longitude coordinates) associated with the content payload of the message. Multiple message geolocation parameter 416 values may be included in the payload, each of which is associated with a content item included in the content (e.g., a specific image within the message image payload 406, or a specific video within the message video payload 408).
[0098] Message story identifier 418: An identifier value that identifies one or more content collections (e.g., a "story" identified in stories table 314) associated with a particular content item in message image payload 406 of message 400. For example, the identifier value may be used to associate multiple images within message image payload 406 with multiple content collections.
[0099] Message tags 420: Each message 400 can be tagged with a plurality of tags, each of which indicates the subject matter of the content included in the message payload. For example, if a particular image included in the message image payload 406 depicts an animal (e.g., a lion), a tag value indicating the relevant animal can be included within the message tags 420. The tag values can be manually generated based on user input, or can be automatically generated using, for example, image recognition.
[0100] • Message sender identifier 422: An identifier (eg, a messaging system identifier, an email address, or a device identifier) that indicates the user of the client device 102 on which the message 400 was generated and from which the message 400 was sent.
[0101] • Message recipient identifier 424: An identifier (eg, a messaging system identifier, an email address, or a device identifier) that indicates the user of the client device 102 to which the message 400 is addressed.
[0102] The content (e.g., value) of each component of message 400 may be a pointer to a location in a table in which content data values are stored. For example, the image value in message image payload 406 may be a pointer to a location in image table 312 (or the address of a location in image table 312). Similarly, the value in message video payload 408 may point to data stored in video table 304, the value stored in message enhancement data 412 may point to data stored in enhancement table 310, the value stored in message story identifier 418 may point to data stored in story table 314, and the values stored in message sender identifier 422 and message recipient identifier 424 may point to user records stored in entity table 306.
[0103] Synthetic View System
[0104] Figure 5 is a block diagram illustrating an example synthetic view system 224 according to an example. The synthetic view system 224 includes a set of components 510 that operate on a set of input data, such as an image 501 of a person wearing a real-world object or an AR object (e.g., a fashion item) in a first camera view (e.g., a front view of the person). The synthetic view system 224 includes an image access module 512, a fashion item detection module 514, an AR item generation module 517, a new view image generation module 516, an image modification module 518, and an image display module 520. All or some of the components of the synthetic view system 224 may be implemented by a server, in which case the image 501 is provided to the server by a client device 102. In some cases, some or all of the components of the synthetic view system 224 may be implemented by a client device 102 or may be distributed across a group of client devices 102.
[0105] In some examples, synthetic view system 224 accesses a first image of a first camera view that depicts a first portion of a person's body. Synthetic view system 224 detects in the first image a fashion item being worn by the person depicted in the image, the fashion item being depicted from the first camera view. Synthetic view system 224 generates a second image of a second camera view that depicts a second portion of the person's body based on the first image. Synthetic view system 224 generates an AR item for the second camera view that visually resembles the fashion item based on the fashion item detected in the first image. Synthetic view system 224 modifies the second image with the AR item to present a synthetic view of the second portion of the person's body wearing the AR item corresponding to the fashion item.
[0106] In some examples, the fashion item in the first image comprises an AR fashion item. In some examples, the first camera view of the fashion item comprises a front view of the AR fashion item. In some examples, the second camera view of the fashion item comprises a side view of the AR fashion item. In some examples, the first portion of the person's body comprises a front view of the person's face, and the second portion of the person's body comprises a side view of the person's face.
[0107] In some examples, synthetic view system 224 applies a pre-trained generative machine learning model to the first image to generate the second image. In some examples, synthetic view system 224 applies a 3D rendering process to the first image to generate the second image.
[0108] In some examples, synthetic view system 224 activates the front-facing camera of the device to capture a depiction of a first portion of a person's body for a first camera view. Synthetic view system 224 synthetically generates a second image that approximates the depiction of a second portion of the body for a second camera view based on the depiction of the first portion of the person's body for the first camera view.
[0109] In some examples, synthetic view system 224 presents the second image along with the first image. In some examples, the second image is superimposed on the first image. In some examples, the second image is presented adjacent to the first image.
[0110] In some examples, synthetic view system 224 synthetically rotates the person's body to generate a second image of the second camera view depicting a second portion of the person's body. In some examples, the fashion item includes earrings, the first portion of the body includes a front view of the person's ear wearing the earrings, and the synthetic view includes a side view of the person's ear wearing the earrings.
[0111] In some examples, the fashion item includes a shoe, the first portion of the body includes a front view of a leg and foot of a person wearing the shoe, and the synthetic view includes a side view of the leg and foot of the person wearing the shoe. In some examples, synthetic view system 224 overlays the AR item on top of the second image to generate the synthetic view.
[0112] In some examples, synthetic view system 224 receives input selecting a second camera view. A second image and a synthetic image can be generated in response to receiving the input. In some examples, the second camera view includes a view of the person's body from a closer distance to the camera than the distance of the person's body to the camera corresponding to the first camera view.
[0113] For example, the image access module 512 can access a live camera feed from the front or rear camera of the client device 102. In some other examples, the image access module 512 can receive a previously recorded or stored camera feed, for example, from a friend or another user associated with another client device 102. The camera feed can depict a real-world person or a body part of a real-world person from a first camera view (e.g., a front view of a body part of the person). The image access module 512 provides the camera feed, including one or more frames depicting the person or the body part of the person from the first camera view, to the fashion item detection module 514.
[0114] For example, Figure 6 As shown, the image access module 512 receives a video 600 including a frame 601 depicting the face of a person 610 in a front view (e.g., a first camera view). The face of the person 610 includes a front view of an ear 620, which is depicted as wearing a fashion item 630, such as an earring. The earring can be a real-world earring, or it can be an AR earring corresponding to a virtual try-on experience. For example, the client device 102 can receive an input selecting or activating the virtual try-on experience by selecting an icon (not shown) corresponding to the virtual try-on experience. In response, the image access module 512 generates a frame 601 in which the virtual fashion item corresponding to the virtual try-on experience is presented as the fashion item 630. In some cases, the frame 601 includes a second fashion item (e.g., another earring depicted as being worn on the other ear of the person 610).
[0115] In some examples, image access module 512 may receive input selecting a new view icon (not shown). The new view icon may be presented above or adjacent to frame 601. The new view icon may allow a user to specify a new view (e.g., a specific camera angle or a side view), or may present a menu listing various types of camera views (e.g., a top view, a side view, a bottom view, etc.). In response to receiving input selecting the new view icon identifying a new camera view, image access module 512 provides video 600 to fashion item detection module 514 so that a composite image or image portion corresponding to fashion item 630 is automatically generated and presented. The composite image may depict a person or a body portion of a person wearing fashion item 630 from the selected new camera view (e.g., a side view).
[0116] Return to reference Figure 5, the fashion item detection module 514 implements one or more machine learning models (e.g., one or more neural networks) that are trained to detect and track one or more fashion items in one or more video frames. For example, during training, the machine learning model of the fashion item detection module 514 receives a given training image (or video) from a plurality of training images (or videos) depicting one or more people wearing fashion items. The plurality of training images are associated with corresponding ground truth labels or indications of the type and location of the fashion items depicted in the training images and can be received or accessed from the training image data stored in the data structure 300. The fashion item detection module 514 applies the one or more machine learning models to the given training images. The fashion item detection module 514 generates estimated tracking information for the fashion item depicted in the given training images.
[0117] The fashion item detection module 514 obtains known or predetermined ground truth tracking information corresponding to a given training image. The new view image generation module 516 compares the estimated tracking information with the ground truth tracking information (calculating the deviation between the two). Based on a difference threshold in the comparison (or deviation), the fashion item detection module 514 updates one or more coefficients or parameters and obtains one or more additional training images from the training data.
[0118] After processing a specified number of epochs or batches of training images and / or when a difference threshold (or deviation) (calculated based on the difference or deviation between the estimated tracking information and the true tracking information) reaches a specified value, the fashion item detection module 514 completes training, and the parameters and coefficients of the fashion item detection module 514 are stored as a trained machine learning technique.
[0119] In some examples, fashion item detection module 514 may be used to enhance a received image or video with fashion items corresponding to an AR experience by overlaying one or more AR fashion items on the received image or video. Fashion item detection module 514 may use one or more previously trained machine learning models to track a person's real-world object of interest (e.g., an ear) and may place an AR object (e.g., an AR earring) relative to the tracked real-world object of interest.
[0120] The fashion item detection module 514 can continuously track the fashion item depicted in the image or video. The fashion item detection module 514 provides tracking information of the fashion item to the new view image generation module 516 and the AR item generation module 517. The new view image generation module 516 can selectively, dynamically, and automatically synthesize a second image that depicts the person in the image accessed by the image access module 512 from a different camera angle or selected camera view.
[0121] Specifically, the new view image generation module 516 may receive an image or video from the image access module 512. The new view image generation module 516 may also receive input indicating a new camera view for synthesizing a new image. For example, the input may request a view of a person synthesized from a second camera view (e.g., a side view or a perspective view).
[0122] In some examples, the new-view image generation module 516 can apply one or more pre-trained generative machine learning models corresponding to the selected new camera view to the image received from the image access module 512. The one or more pre-trained generative machine learning models generate a composite image and output the composite image as a second image, which depicts a body portion of a person in the original image from the selected new camera view. In some examples, each of the one or more pre-trained generative machine learning models is trained based on corresponding training data. Specifically, during training, the machine learning model of the new-view image generation module 516 receives a given training image (or video) from a plurality of training images (or videos) that depict one or more people wearing a fashion article from a first camera view (e.g., a front view). The plurality of training images are associated with corresponding ground-truth images of the same person wearing the fashion article from a second camera view (e.g., a side view) and can be received or accessed from the training image data stored in the data structure 300. The new-view image generation module 516 applies one or more machine learning models to a given training image depicting a person wearing or not wearing a fashion article from a first camera view perspective (e.g., a front view of a body part of the person, such as the front of the face or the front of the legs). In some cases, the new-view image generation module 516 applies the one or more trained machine learning models to determine the type of view of the body part in the image (e.g., to determine whether the current view is a front view or a side view). The new-view image generation module 516 generates an estimated image of the person from a second camera view perspective (e.g., a side view of a body part of the person, such as the side of the face or the side of the legs).
[0123] The new view image generation module 516 obtains a known or predetermined ground truth image corresponding to a given training image (which has been captured and depicts the same person from the second camera view). The new view image generation module 516 compares the estimated image with the ground truth image (calculates the deviation between the two). Based on a difference threshold in the comparison (or deviation), the new view image generation module 516 updates one or more coefficients or parameters and obtains one or more additional training images from the training data.
[0124] After processing a specified number of epochs or batches of training images and / or when a difference threshold (or deviation) (calculated based on the difference or deviation between the estimated image and the true tracking image) reaches a specified value, the new view image generation module 516 completes training, and the parameters and coefficients of the new view image generation module 516 are stored as a trained machine learning technique.
[0125] In some examples, new-view image generation module 516 can apply a 3D rendering process corresponding to the selected new camera view to the image received from image access module 512. The output of the 3D rendering process is a composite image of the selected new camera view that depicts the person depicted in the image received from image access module 512. In this manner (using a generative machine learning model and / or using a 3D rendering process), new-view image generation module 516 synthetically rotates a body portion of a person depicted in an image received from image access module 512 (e.g., captured by a camera) to synthetically generate a new image of the body portion from a different camera view.
[0126] For example, Figure 7 6 , a user interface 700 is presented in which a composite image 701 depicts a person 710 and an ear 720 of person 710 from a camera view that is different from the camera view depicted and used to capture frame 601. Specifically, composite image 701 corresponds to a side view of the face of person 610 depicted in frame 601 and is a composite version or image 701 that depicts the side of the face of person 610 depicted in frame 601. Composite image 701 resembles a view of person 710 that is a side perspective relative to the view captured by the camera of client device 102 represented by frame 601.
[0127] The AR item generation module 517 receives tracking information for a fashion item depicted in the image accessed by the image access module 512 and the new camera view. In some examples, the fashion item detection module 514 may segment the detected fashion item and provide only the portion of the image received from the image access module 512 that corresponds to the fashion item segmentation. The AR item generation module 517 may generate a 3D model of the fashion item depicted in the image received from the image access module 512. In some cases, the AR item generation module 517 receives the AR fashion item depicted in the image received from the image access module 512. The AR item generation module 517 applies a 3D object model process to rotate, turn, or modify the view of the 3D model of the fashion item to represent the view of the fashion item in the new camera view. Based on the 3D object model, the AR item generation module 517 generates an AR item that resembles the fashion item in the new camera view. The AR item generation module 517 provides the AR item representing the fashion item in the new camera view to the new view image generation module 516 and / or the image modification module 518.
[0128] Image modification module 518 may receive a composite image of a body portion of a person in the new camera view from new view image generation module 516 and an AR item representing a fashion item in the new camera view from AR item generation module 517. Image modification module 518 combines (e.g., projects or overlays) the AR item with the composite image to generate a new image or a new composite image of the body portion of the person wearing the fashion item in the new camera view. In some cases, image modification module 518 may overlay one or more AR objects on the scene generated by new view image generation module 516. Image modification module 518 may receive input from the user to move (reposition) the AR object within the scene and may update the position, orientation, and / or size of the AR object based on the input.
[0129] In some examples, such as Figure 7As shown, image modification module 518 can generate image 701 in which a composite view of person 710 (e.g., a composite or generated side view of a person's body part, such as a left ear) is combined with an AR item 730 representing a fashion item (e.g., a left earring). Composite image 701 depicts the composite body part of person 710 wearing the AR item from a different camera view (e.g., a side camera view) than frame 601, in which the same body part is depicted wearing the fashion item (e.g., a front camera view). Image display module 520 receives the image from image modification module 518 and displays the image on a screen of client device 102. The image provided by image display module 520 can be shared with other users on messaging system 100. For example, the image provided by image display module 520 can be included as part of an advertisement or promotion associated with the real-world object depicted in the image.
[0130] In some examples, image display module 520 presents composite image 701 over frame 601, for example, in a picture-in-picture arrangement. In such cases, composite image 701 is presented as a larger image relative to frame 601, and frame 601 is presented in a window over composite image 701. In some examples, composite image 701 is presented as a smaller image relative to frame 601, and composite image 701 is presented in a window over frame 601. In some examples, composite image 701 is presented next to frame 601, or above or below frame 601.
[0131] In some examples, the synthetic view system 224 presents Figure 8 8. User interface 800 is shown. User interface 800 includes an image 801 captured by a front-facing camera or a rear-facing camera of client device 102. Image 801 depicts a side view of a person's leg 810 wearing a fashion item 820, such as a shoe. Synthetic view system 224 may receive input requesting that leg 810 and fashion item 820 be viewed from a different camera view (e.g., a front camera view). In response, synthetic view system 224 presents Figure 9 User interface 900 is shown. In user interface 900, synthetic view system 224 generates and presents synthetic image 901 in which leg 810 of the person depicted in image 801 is depicted from a front camera view. Synthetic view system 224 also generates a front camera view of fashion item 820 and overlays the front camera view of fashion item 820 on synthetic image 901. This results in synthetic image 901 in which leg 810, which was depicted from a side camera view as wearing fashion item 820, is depicted from a front camera view as wearing fashion item 820.
[0132] Figure 101 is a flowchart of a process 1000 according to some examples. Although the flowchart describes the operations as a sequential process, many of these operations can be performed in parallel or simultaneously. In addition, the order of the operations can be rearranged. The process terminates when its operations are completed. A process can correspond to a method, a procedure, etc. The steps of a method can be performed in whole or in part, can be combined with some or all of the steps in other methods, and can be performed by any number of different systems or any part thereof (e.g., a processor included in any system).
[0133] At operation 1001 , as described above, the synthetic view system 224 (eg, a server or client device 102 ) accesses a first image of a first camera view depicting a first portion of a person's body.
[0134] At operation 1002, as described above, the synthetic view system 224 detects a fashion item being worn by a person depicted in the image in a first image, the fashion item being depicted from a first camera view. The synthetic view system 224 may also detect a body part of the person depicted in the first image.
[0135] At operation 1003, as described above, the synthetic view system 224 generates a second image depicting a second portion of a person's body from a second camera view based on the first image. For example, the second portion of the body may correspond to a different perspective, such as a side view versus a front view, of the same body portion (e.g., the first portion of the body) depicted in the first image.
[0136] At operation 1004 , as described above, the synthetic view system 224 generates an AR item of the second camera view that visually resembles the fashion item based on the fashion item detected in the first image.
[0137] At operation 1005 , as described above, the synthetic view system 224 modifies the second image with the AR item to present a synthetic view of a second portion of the body of a person wearing the AR item corresponding to the fashion item.
[0138] Machine Architecture
[0139] Figure 111 is a diagrammatic representation of a machine 1100 within which instructions 1108 (e.g., software, programs, applications, applet, apps, or other executable code) may be executed for causing the machine 1100 to perform any one or more of the methodologies discussed herein. For example, the instructions 1108 may cause the machine 1100 to perform any one or more of the methodologies described herein. The instructions 1108 transform a general-purpose, unprogrammed machine 1100 into a specialized machine 1100 that is programmed to perform the functions described and illustrated in the manner described. The machine 1100 may operate as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1100 may operate as a server or a client machine in server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1100 may include, but is not limited to, a server computer, a client computer, a PC, a tablet computer, a laptop computer, a netbook, an STB, a PDA, an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smartwatch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing instructions 1108 that specify actions to be taken by the machine 1100, either sequentially or otherwise. Furthermore, while only a single machine 1100 is shown, the term "machine" should also be considered to include a collection of machines that individually or jointly execute instructions 1108 to perform any one or more of the methodologies discussed herein. For example, the machine 1100 may include the client device 102 or any of several server devices that form part of the messaging server system 108. In some examples, the machine 1100 may also include both a client system and a server system, with certain operations of a particular method or algorithm being performed on the server side and certain operations of a particular method or algorithm being performed on the client side.
[0140] The machine 1100 may include a processor 1102, a memory 1104, and input / output (I / O) components 1138 that may be configured to communicate with each other via a bus 1140. In an example, the processor 1102 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 1106 that executes instructions 1108 and a processor 1110. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that may execute instructions concurrently. Although Figure 11 Multiple processors 1102 are shown, but the machine 1100 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0141] The memory 1104 includes a main memory 1112, a static memory 1114, and a storage unit 1116, all of which are accessible by the processor 1102 via the bus 1140. The main memory 1104, the static memory 1114, and the storage unit 1116 store instructions 1108 that embody any one or more of the methodologies or functions described herein. The instructions 1108 may also reside, completely or partially, within the main memory 1112, within the static memory 1114, within a machine-readable medium within the storage unit 1116, within at least one of the processors 1102 (e.g., within a cache memory of the processor), or within any suitable combination thereof during execution thereof by the machine 1100.
[0142] The I / O components 1138 may include various components that receive input, provide output, generate output, send information, exchange information, capture measurements, etc. The specific I / O components 1138 included in a particular machine will depend on the type of machine. For example, a portable machine (such as a mobile phone) may include a touch input device or other such input mechanism, while a headless server machine will likely not include such a touch input device. It should be understood that the I / O components 1138 may include Figure 11Many other components are not shown in the drawings. In various examples, the I / O components 1138 may include user output components 1124 and user input components 1126. The user output components 1124 may include visual components (e.g., displays such as plasma display panels (PDPs), light emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), acoustic components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. The user input components 1126 may include alphanumeric input components (e.g., keyboards, touch screens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touch pads, trackballs, joysticks, motion sensors, or other pointing instruments), tactile input components (e.g., physical buttons, touch screens or other tactile input components that provide location and force of touch or touch gestures), audio input components (e.g., microphones), etc.
[0143] In other examples, the I / O component 1138 may include a biometric component 1128, a motion component 1130, an environmental component 1132, or a position component 1134, as well as various other components. For example, the biometric component 1128 includes components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion component 1130 includes an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, and a rotation sensor component (e.g., a gyroscope).
[0144] Environmental components 1132 include, for example, one or more cameras (with still image / photo and video capabilities), lighting sensor components (e.g., a photometer), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., a barometer), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., an infrared sensor that detects nearby objects), gas sensors (e.g., a gas detection sensor that detects concentrations of hazardous gases for safety or measures pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment.
[0145] With respect to cameras, client device 102 can have a camera system that includes, for example, a front-facing camera on the front surface of client device 102 and a rear-facing camera on the rear surface of client device 102. The front-facing camera can be used, for example, to capture still images and videos of a user of client device 102 (e.g., “selfies”), which can then be enhanced using the enhancement data (e.g., filters) described above. The rear-facing camera can be used, for example, to capture still images and videos in a more traditional camera mode, where the images are similarly enhanced using the enhancement data. In addition to the front-facing and rear-facing cameras, client device 102 can also include a 360° camera for capturing 360° photos and videos.
[0146] Additionally, the camera system of the client device 102 may include dual rear cameras (e.g., a main camera and a depth sensing camera), or even triple, quad, or quintuple rear camera configurations on the front and back sides of the client device 102. For example, these multi-camera systems may include a wide-angle camera, an ultra-wide-angle camera, a telephoto camera, a macro camera, and a depth sensor.
[0147] The location component 1134 includes a positioning sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure, from which altitude can be obtained), an orientation sensor component (e.g., a magnetometer), and the like.
[0148] A variety of technologies can be used to implement communications. The I / O components 1138 also include a communications component 1136 that is operable to couple the machine 1100 to the network 1120 or device 1122 via corresponding couplings or connections. For example, the communications component 1136 may include a network interface component or other suitable device that interfaces with the network 1120. In other examples, the communications component 1136 may include a wired communications component, a wireless communications component, a cellular communications component, a near field communications (NFC) component, a Components (e.g. Low energy consumption), Components and other communication components for providing communication via other modalities. Device 1122 can be another machine or any of a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0149] In addition, the communication component 1136 can detect an identifier or include a component operable to detect an identifier. For example, the communication component 1136 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional bar codes such as Universal Product Code (UPC) bar codes, multi-dimensional bar codes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar codes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tag). In addition, various information can be derived via the communication component 1136, such as location via Internet Protocol (IP), location via Internet Protocol (IP), location information ... Positioning can be derived through signal triangulation, positioning can be derived through detection of NFC beacon signals that can indicate a specific position, etc.
[0150] Various memories (e.g., main memory 1112, static memory 1114, and memory of processor 1102) and storage unit 1116 may store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. When executed by processor 1102, these instructions (e.g., instructions 1108) cause various operations to implement the disclosed examples.
[0151] The instructions 1108 may be sent or received over the network 1120 using a transmission medium via a network interface device (e.g., a network interface component included in the communications component 1136) and using any of several well-known transmission protocols (e.g., HTTP). Similarly, the instructions 1108 may be sent or received over a coupling (e.g., a peer-to-peer coupling) with the device 1122 using a transmission medium.
[0152] Software Architecture
[0153] Figure 1212 is a block diagram 1200 illustrating a software architecture 1204 that can be installed on any one or more of the devices described herein. The software architecture 1204 is supported by hardware, such as a machine 1202 including a processor 1220, memory 1226, and I / O components 1238. In this example, the software architecture 1204 can be conceptualized as a stack of layers, each providing specific functionality. The software architecture 1204 includes layers such as an operating system 1212, libraries 1210, frameworks 1208, and applications 1206. Operationally, the applications 1206 invoke API calls 1250 through the software stack and receive messages 1252 in response to the API calls 1250.
[0154] The operating system 1212 manages hardware resources and provides public services. The operating system 1212 includes, for example, a kernel 1214, services 1216, and drivers 1222. The kernel 1214 serves as an abstraction layer between the hardware and other software layers. For example, the kernel 1214 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. Services 1216 can provide other public services to other software layers. Drivers 1222 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 1222 may include display drivers, camera drivers, or Low-power drivers, Flash drivers, serial communication drivers (e.g., USB drivers), drivers, audio drivers, power management drivers, etc.
[0155] The libraries 1210 provide a common low-level infrastructure used by the applications 1206. The libraries 1210 may include system libraries 1218 (e.g., C standard libraries) that provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, the libraries 1210 may include API libraries 1224, such as media libraries (e.g., libraries for supporting the presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., the OpenGL framework for rendering graphical content on a display in 2D and 3D), database libraries (e.g., SQLite providing various relational database functions), web libraries (e.g., WebKit providing web browsing functions), etc. The library 1210 may also include various other libraries 1228 to provide many other APIs to the application 1206 .
[0156] The framework 1208 provides a common high-level infrastructure used by the applications 1206. For example, the framework 1208 provides various GUI functions, advanced resource management, and advanced positioning services. The framework 1208 can provide a wide range of other APIs that can be used by the applications 1206, some of which may be specific to a particular operating system or platform.
[0157] In an example, applications 1206 may include a home application 1236, a contacts application 1230, a browser application 1232, a book reader application 1234, a location application 1242, a media application 1244, a messaging application 1246, a game application 1248, and a variety of other applications such as external applications 1240. Applications 1206 are programs that perform functions defined in the program. Various programming languages may be used to create one or more of the applications 1206 structured in various ways, such as an object-oriented programming language (e.g., Objective-C, Java, or C++) or a procedural programming language (e.g., C or assembly language). In a specific example, external applications 1240 (e.g., written by an entity other than the vendor of a particular platform using ANDROID) may be used to create a program that is not a part of the application 1206. TM or IOS TM SDK developed applications) can be mobile operating systems such as IOS TM ANDROID TM 、 Phone, or another mobile operating system running mobile software. In this example, the external application 1240 can activate the API call 1250 provided by the operating system 1212 to facilitate the functions described herein.
[0158] Glossary
[0159] "Carrier signal" refers to any intangible medium that can store, encode, or carry instructions for execution by a machine and includes digital or analog communication signals or other intangible media that facilitates communication of such instructions. Instructions can be sent or received over a network using a transmission medium via a network interface device.
[0160] "Client Device" refers to any machine that interfaces with a communications network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, desktop computer, laptop computer, PDA, smartphone, tablet computer, ultrabook, netbook, laptop computer, multiprocessor system, microprocessor-based or programmable consumer electronics, game console, STB, or any other communications device that a user may use to access a network.
[0161] "Communications network" means one or more parts of a network, which may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a Plain Old Telephone Service (POTS) network, a cellular telephone network, a wireless network, The network, other types of networks, or a combination of two or more such networks. For example, the network or a portion of the network may include a wireless network or a cellular network, and the coupling may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other types of cellular or wireless coupling. In this example, the coupling may implement any of various types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rates for GSM evolution (EDGE) technology, the third generation partnership project (3GPP) including 3G, fourth generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), world wide interoperability for microwave access (WiMAX), long term evolution (LTE) standards, other data transmission technologies defined by various standards setting organizations, other long distance protocols, or other data transmission technologies.
[0162] "Component" means a logical, device, or physical entity having boundaries defined by function or subroutine calls, branch points, APIs, or other techniques that provide for partitioning or modularization of specific processing or control functions. Components can be combined with other components via their interfaces to perform machine processing. A component can be a packaged functional hardware unit designed for use with other components and can be part of a program that generally performs a specific one of the related functions.
[0163] Components may constitute software components (e.g., code implemented on a machine-readable medium) or hardware components. A "hardware component" is a tangible unit that is capable of performing certain operations and that can be configured or arranged in some physical manner. In various examples, one or more computer systems (e.g., stand-alone computer systems, client computer systems, or server computer systems) or one or more hardware components of a computer system (e.g., a processor or group of processors) may be configured by software (e.g., an application or application portion) to be hardware components that operate to perform certain operations described herein.
[0164] Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component may include a dedicated circuit system or logic that is permanently configured to perform certain operations. A hardware component may be a dedicated processor, such as a field programmable gate array (FPGA) or an ASIC. A hardware component may also include a programmable logic or circuit system that is temporarily configured to perform certain operations through software. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured through such software, the hardware component becomes a specific machine (or specific component of a machine) that is uniquely customized to perform the configured function, rather than a general-purpose processor. It should be recognized that it can be decided for cost and time considerations whether to implement the hardware component mechanically in a dedicated and permanently configured circuit system or in a temporarily configured circuit system (e.g., configured by software). Therefore, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, that is, an entity that is physically constructed, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in some manner or perform certain operations described herein.
[0165] Consider an example where hardware components are temporarily configured (e.g., programmed), without having to configure or instantiate every hardware component in the hardware components at any one time. For example, where the hardware components include a general-purpose processor that is configured by software to become a special-purpose processor, the general-purpose processor can be configured at different times to become different special-purpose processors (e.g., including different hardware components). The software accordingly configures one or more specific processors to, for example, constitute a specific hardware component at one time and to constitute different hardware components at different times.
[0166] Hardware components can provide information to other hardware components and receive information from other hardware components. Therefore, described hardware components can be considered to be coupled in communication. In the case of having multiple hardware components simultaneously, communication can be realized by the signal transmission between two or more hardware components in the hardware components (for example, by suitable circuit and bus). In the example where multiple hardware components are configured or instantiated at different times, communication between such hardware components can be realized, for example, by storing information in a memory structure that multiple hardware components can access and retrieving information in this memory structure. For example, a hardware component can perform an operation and store the output of this operation in the memory device to which it is coupled in communication. Then, another hardware component can access this memory device at a subsequent time to retrieve stored output and process stored output. Hardware components can also initiate communication with input device or output device, and can operate on resources (for example, the set of information).
[0167] The various operations of the example methods described herein may be performed at least in part by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily configured or permanently configured, such a processor may constitute a processor-implemented component that operates to perform one or more operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein may be at least partially implemented by a processor, wherein one or more specific processors are examples of hardware. For example, at least some of the operations of the method may be performed by one or more processors 1102 or a processor-implemented component. In addition, the one or more processors may also operate to support the execution of the relevant operations in a "cloud computing" environment or as a "software as a service" (SaaS) operation. For example, at least some of the operations may be performed by a group of computers (as an example of a machine including a processor), wherein these operations may be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of certain operations in the operation may be distributed between processors, not residing only within a single machine, but deployed across several machines. In some examples, the processor or processor-implemented components may be located in a single geographic location (e.g., in a home environment, an office environment, or a server farm). In other examples, the processor or processor-implemented components may be distributed across several geographic locations.
[0168] "Computer-readable storage media" refers to both machine storage media and transmission media. Thus, the term encompasses both storage devices / medium and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" mean the same thing and may be used interchangeably in this disclosure.
[0169] A "transient message" is a message that is accessible for a limited duration. Transient messages can be text, images, videos, and more. The access time for a transient message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is transient.
[0170] “Machine storage media” refers to a single or multiple storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Thus, the term should be taken to include, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGAs, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms “machine storage media,” “device storage media,” and “computer storage media” mean the same thing and are used interchangeably in this disclosure. The terms “machine storage media,” “computer storage media,” and “device storage media” expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are encompassed by the term “signal media.”
[0171] “Non-transitory computer-readable storage medium” refers to a tangible medium capable of storing, encoding, or carrying instructions for execution by a machine.
[0172] "Signal medium" refers to any intangible medium capable of storing, encoding, or carrying instructions to be executed by a machine and includes digital or analog communication signals or other intangible media that facilitate the communication of software or data. The term "signal medium" should be construed to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" refers to a signal that has one or more of its characteristics set or changed to encode information therein. The terms "transmission medium" and "signal medium" refer to the same thing and may be used interchangeably in this disclosure.
[0173] Changes and modifications may be made to the disclosed examples without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the following claims.
Claims
1. A method comprising: accessing, by one or more processors of the device, a first image of a first camera view depicting a first portion of a person's body; detecting a fashion item being worn by the person in the first image in the first image, the fashion item being depicted in the first image of the first camera view; generating a second image of a second camera view depicting a second portion of the person's body based on the first image; generating an augmented reality (AR) item of the second camera view that visually resembles the fashion item based on the fashion item detected in the first image; as well as The second image is modified with the AR article to present a composite view of a second portion of the person's body wearing the AR article corresponding to the fashion article.
2. The method according to claim 1, wherein The fashion item in the first image includes an AR fashion item.
3. The method according to any one of claims 1 to 2, wherein The first camera view of the fashion item includes a front view of the AR fashion item.
4. The method according to claim 3, wherein: The second camera view of the fashion item includes a side view of the AR fashion item.
5. The method according to any one of claims 1 to 4, wherein The first portion of the person's body includes the front of the person's face, and wherein the second portion of the person's body includes the side of the person's face.
6. The method according to any one of claims 1 to 5, further comprising: A generative machine learning model is applied to the first image to generate the second image.
7. The method according to any one of claims 1 to 6, further comprising: A three-dimensional (3D) rendering process is applied to the first image to generate the second image.
8. The method according to any one of claims 1 to 7, further comprising: activating a front-facing camera of the device to capture a depiction of a first portion of the person's body from the first camera view; as well as Based on the depiction of the first portion of the person's body by the first camera view, a second image is synthetically generated that approximates the depiction of the second portion of the body by the second camera view.
9. The method according to any one of claims 1 to 8, further comprising: The second image is presented together with the first image.
10. The method according to claim 9, wherein: The second image is superimposed on the first image.
11. The method according to any one of claims 1 to 10, wherein The second image is presented adjacent to the first image.
12. The method according to any one of claims 1 to 11, further comprising: The person's body is synthetically rotated to generate the second image of the second camera view depicting a second portion of the person's body.
13. The method according to any one of claims 1 to 12, wherein The fashion article comprises an earring, wherein the first body portion comprises a front view of an ear of the person wearing the earring, and wherein the composite view comprises a side view of the ear of the person wearing the earring.
14. The method according to any one of claims 1 to 13, wherein The fashion article comprises a shoe, wherein the first body portion comprises a front view of a leg and foot of the person wearing the shoe, and wherein the composite view comprises a side view of the leg and foot of the person wearing the shoe.
15. The method according to any one of claims 1 to 14, further comprising: The AR item is superimposed on the second image to generate the composite view.
16. The method according to any one of claims 1 to 15, further comprising: Input is received selecting the second camera view, wherein the second image and the composite image are generated in response to receiving the input.
17. The method according to any one of claims 1 to 16, wherein The second camera view comprises a view of the person's body that is from a closer distance to the camera relative to the distance of the person's body to the camera corresponding to the first camera view.
18. A system comprising: A processor of a device, the processor being configured to perform operations comprising: accessing a first image of a first camera view depicting a first portion of a person's body; detecting a fashion item being worn by the person in the first image in the first image, the fashion item being depicted in the first image of the first camera view; generating a second image of a second camera view depicting a second portion of the person's body based on the first image; generating an augmented reality (AR) item of the second camera view that visually resembles the fashion item based on the fashion item detected in the first image; and The second image is modified with the AR article to present a composite view of a second portion of the person's body wearing the AR article corresponding to the fashion article.
19. The system according to claim 18, wherein: The fashion item in the first image includes an AR fashion item.
20. A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a device, cause the device to perform operations comprising: accessing a first image of a first camera view depicting a first portion of a person's body; detecting a fashion item being worn by the person in the first image in the first image, the fashion item being depicted in the first image of the first camera view; generating a second image of a second camera view depicting a second portion of the person's body based on the first image; generating an augmented reality (AR) item of the second camera view that visually resembles the fashion item based on the fashion item detected in the first image; as well as The second image is modified with the AR article to present a composite view of a second portion of the person's body wearing the AR article corresponding to the fashion article.