Representing visual enhancement
By training machine learning models through self-supervised learning and automated rendering services, embeddings of visual enhancements are generated, solving the problem that existing models cannot effectively represent visual enhancements. This achieves more efficient visual enhancement recognition and retrieval, and optimizes the utilization of computing resources.
Patent Information
- Application Number
- CN202480026589.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-20
- Filing Date
- 2024-04-17
- Publication Date
- 2025-11-18
AI Technical Summary
Existing machine learning models fail to effectively capture the features of visual enhancements when generating representations of visual enhancements, resulting in limited use of visual enhancements in interactive systems for identification, comparison, and retrieval, and they cannot provide a suitable content pipeline for training.
A machine learning model is trained using a self-supervised learning method to generate embeddings for specific visual enhancements. Unsupervised learning and an automated lens rendering service are used to map the visual enhancements to a lower-dimensional space, and the embeddings of multiple sample content items are aggregated to generate enhancement identifiers.
It improves the efficiency of visually enhanced recognition, comparison, and retrieval, reduces duplication in databases and interactive systems, promotes enhanced sorting and retrieval, and optimizes the utilization of computing resources.
Smart Images

Figure CN120981801A_ABST
Abstract
Description
[0001] Priority Statement
[0002] This application claims the benefit of priority to U.S. Patent Application Serial No. 18 / 304,078, filed April 20, 2023, which is incorporated herein by reference in its entirety. Technical Field
[0003] The topics disclosed in this paper generally relate to machine learning. More specifically, but not exclusively, the topics disclosed in this paper relate to representation learning, which includes example machine learning models trained to generate vector representations of visual content. Background Technology
[0004] Machine learning models are applications that provide computer systems with the ability to perform tasks by reasoning based on patterns discovered in the analysis of data, without explicit programming. Machine learning explores the study and construction of algorithms (also referred to as models in this paper) that can learn from existing data and make predictions about new data.
[0005] Representation learning is a subfield of machine learning that focuses on learning useful features or representations of data. Machine learning models can automatically learn a set of features that capture the essential characteristics of input data. For example, a machine learning model can find a transformation or mapping from input data to a new representation that is more compact than the input data while containing important information about the input data, thus making the new representation useful in downstream tasks.
[0006] In the context of visual representation learning (e.g., learning from input images), machine learning models can be trained to map high-dimensional visual inputs to a lower-dimensional space. As used in this disclosure, the term "embedding" refers to such a representation of visual content in a lower-dimensional space that is designed to capture or preserve important visual features of the content. Attached Figure Description
[0007] In the accompanying drawings (which are not necessarily drawn to scale), similar reference numerals can describe similar parts in different views. To facilitate identification of any discussion of a particular element or action, one or more of the highest-order digits in the reference numerals indicate the drawing number in which the element was first introduced. Some non-limiting examples are shown in the accompanying drawings:
[0008] Figure 1 It is a diagrammatic representation of a networked environment in which the content of this disclosure can be deployed, based on some examples.
[0009] Figure 2 It is a graphical representation of an interactive system with both client-side and server-side functionalities, based on some examples.
[0010] Figure 3 It is a block diagram illustrating the interaction between certain components of an interactive system, including artificial intelligence and machine learning systems, as well as embedded and mapping systems, based on some examples.
[0011] Figure 4 This is a diagram illustrating the training and use of machine learning programs based on some examples.
[0012] Figure 5 This is a graph illustrating how machine learning models are trained based on some examples to generate embeddings representing visual enhancements within input video items.
[0013] Figure 6 It is a conceptual illustration of generating embeddings to capture visually enhanced visual features or effects.
[0014] Figure 7 Example frames from each of a pair of visually enhanced videos are shown, based on some examples.
[0015] Figure 8 Example frames from each of a pair of visually enhanced videos are shown, based on some examples.
[0016] Figure 9 Example frames from each of a pair of visually enhanced videos are shown, based on some examples.
[0017] Figure 10 This is a flowchart illustrating a method for generating embeddings using a machine learning model based on some examples and mapping target visual enhancements to enhancement identifiers.
[0018] Figure 11 It is a conceptual illustration of generated embeddings based on some examples to capture visually enhanced visual effects or features.
[0019] Figure 12 It is a conceptual illustration of generated embeddings based on some examples to capture visually enhanced visual effects or features.
[0020] Figure 13 This is a conceptual diagram illustrating the use of enhanced identifiers in the automated nearest neighbor process, based on some examples.
[0021] Figure 14 It is a graphical representation based on examples such as data structures maintained in a database.
[0022] Figure 15 It is a graphical representation based on some example messages.
[0023] Figure 16 The system includes a head-mounted wearable device, as shown in some examples.
[0024] Figure 17 It is a graphical representation of a machine in the form of a computer system, based on some examples, within which a set of instructions can be executed to cause the machine to perform any or more of the methods discussed herein.
[0025] Figure 18 It is a block diagram showing the software architecture in which examples can be implemented. Detailed Implementation
[0026] As used in this disclosure, the term "visual enhancement" refers to an effect, modification, removal, or addition that alters an image (or sequence of images) compared to the image (or sequence of images) that would be observed if captured and rendered without any such effect, modification, removal, or addition. Examples of visual enhancements may include two-dimensional or three-dimensional augmented reality effects, filters, lenses, media overlays (e.g., text, color, or image overlays), augmented reality experiences, virtual reality experiences, extended reality (XR) experiences, or combinations thereof.
[0027] Visual enhancements introduce special, interesting, entertaining, or useful effects onto content items such as images or videos. For example, a user of an interactive application can select which visual enhancements to apply to video content captured using the user's device's camera. The visual enhancements can then be applied in real time (e.g., before or during content capture, to objects presented in the camera's feed interface) or after content capture (e.g., to a video file retrieved from the user's device's memory).
[0028] Examples of this disclosure provide computerized mechanisms for generating representations of visually enhanced effects. Visual effects may include, for example, changes in visual appearance caused by the visual enhancement (e.g., applying an overlay to a person's face, or changing the colors of an image). In some cases, visually enhanced effects may include changes in the format or style of content items (e.g., applying a "green screen" effect to a video captured by a user). In some cases, visual effects may cause both aesthetic and formatting changes.
[0029] While conventional systems can enable the generation of vector representations of images or videos, for example, using machine learning models, a drawback of such systems is that the generated representations do not adequately represent the visual enhancements applied within the image or video. For example, a user might capture a first "selfie" (self-portrait) video at a beach and apply visual enhancements to the captured facial features, and then capture a second "selfie" video in a jungle and apply the same visual enhancements to the second video. For instance, due to significant differences in the background scenes between the two videos, the representations generated from these videos may be substantially different. Therefore, the output of a conventional system may have limited use in the context of visual enhancements (e.g., for recognizing, comparing, retrieving, or ranking visual enhancements within an interactive system). Furthermore, conventional systems may not provide a suitable pipeline for training machine learning models to generate outputs that represent visual enhancements in a useful or appropriate manner.
[0030] In some examples of this disclosure, technical shortcomings of conventional systems can be addressed or mitigated by generating embeddings for specific visual enhancements to capture the effects of visual enhancements, such as those in numerical representation. Specific visual enhancements can be mapped to embeddings (e.g., a list of floating-point numbers). Embeddings for specific visual enhancements can be determined by aggregating multiple embeddings generated using sample content items to which specific visual enhancements have been applied.
[0031] Furthermore, embeddings can capture the visual effect of a particular visual augmentation relative to other visual augmentations, for example, making similar visual augmentations have similar embeddings, thus providing useful output when compared to the output of a regular system. The mapping of augmented visual effects to corresponding embeddings provides a mechanism for understanding and comparing different augmentations. Therefore, in some examples, augmented visual effects are mapped to features (e.g., vectors) that can be consumed by downstream machine learning models or other tools. Embeddings generated in this way can be used for various downstream tasks, such as those consumed by machine learning models trained for augmented labeling, augmented ranking, or augmented retrieval within interactive systems.
[0032] Before deploying a machine learning model that can generate embeddings representing important features of visual enhancements, the model can be trained (e.g., using unsupervised learning methods). In some examples, self-supervised learning is employed to train the model to focus on the appearance or effect of visual enhancements within a video, rather than on other content within the video (e.g., background content not altered by the visual enhancements). The machine learning model can be trained to minimize the loss between the training video representations generated in each of multiple training sets. Each training set can include different training video items, each including the same predefined visual enhancements.
[0033] Within the context of video content, machine learning models can be trained on video pairs. A pair can include two different plate videos or background videos, to which the same visual enhancements are applied (e.g., using an automated lens rendering service). An automated lens rendering service can provide a pipeline of suitable content to train machine learning models to generate outputs that represent visual enhancements in a useful or appropriate manner.
[0034] In some examples, for each of multiple visual enhancements, two videos rendered on different backgrounds can be sampled. A representation is generated for each video, and one of the representations is transformed (e.g., via a predictor network) to generate a transformed representation. The loss function can take into account the similarity between the transformed representation and the original, untransformed representation. The process can be repeated, for example, a predefined number of times for each visual enhancement, or repeated with a predefined number of training pairs to ensure sufficient training. In this way, a machine learning model is trained to generate similar embeddings for video pairs including the same (or similar) visual enhancements, and dissimilar embeddings for video pairs including dissimilar visual enhancements.
[0035] In some examples, the training phase of a machine learning model can include self-supervised positive-only contrastive learning. In positive-only learning, the machine learning model can be trained specifically on positive pairs (e.g., video pairs including the same visual enhancements). In other examples, the training phase of a machine learning model can include contrastive learning in which the model is also trained on negative pairs, such as video pairs with different visual enhancements applied to them (e.g., videos with the same background but different visual enhancements between the two videos). In some cases, this can facilitate the model learning to map similar inputs to nearby points in the feature space, while pushing dissimilar inputs further away. Specifically, the model can learn more easily or better to distinguish between items with the same or similar enhancements and items with different (or dissimilar) enhancements.
[0036] Once trained, inference can be run on input video items that include visual enhancements to generate embeddings that capture the visual effects of the visual enhancements applied to the input video items. In some examples, multiple embeddings are generated for a specific visual enhancement by running inference on multiple different videos, all of which have the specific visual enhancement applied. Aggregated embeddings can be determined (e.g., based on the mean of the embeddings), and the aggregated embeddings can be designated as enhancement identifiers for a specific visual enhancement. In this way, in the case of an interactive system employing a large number of visual enhancements, enhancement identifiers can be obtained for each visual enhancement and stored in a database; for example, a one-to-one mapping between each visual enhancement and its corresponding enhancement identifier (in vector form) can be stored for further use.
[0037] Therefore, examples of this disclosure relate to training machine learning models to generate vector representations of visual enhancements applied to content items, deploying such machine learning models, and mapping the vector representations to their corresponding visual enhancements. The systems and methods described herein can be useful in solving the example technical problem of creating vector representations of visual enhancements from a video (or a representation of a video) containing visual enhancements and underlying content to which the visual enhancements are applied.
[0038] The examples described herein can address or mitigate technical problems associated with representation learning for visual augmentation. Specifically, but not exclusively, this disclosure describes systems and methods that facilitate the application of self-supervised learning methods to videos to map augmentation-specific features from the video to a lower-dimensional space. Examples of this disclosure consider the temporal dimension associated with the augmented video to ensure that the effects of the augmentation along the temporal dimension are captured in the embedding.
[0039] The examples described herein can provide efficient or scalable systems for generating high-quality embeddings. Furthermore, the examples in this disclosure can solve or mitigate technical problems related to generating useful mappings for visual augmentation understanding or comparison purposes.
[0040] Embedsions generated using the techniques described herein can be used for a variety of purposes, such as reducing or eliminating duplication of visual enhancements in databases or interactive systems, facilitating the retrieval of enhancements from storage locations, generating insights into the similarity or differences between enhancements, predicting enhancement tags within content items, ranking enhancements, or clustering meaningful sets of visual enhancements for presentation to users.
[0041] When taken into account the effects of this disclosure, one or more of the methods described herein can eliminate the need for certain work or resources that would otherwise be involved in data extraction. Computational resources used by one or more machines, databases, or networks can be utilized more efficiently or even reduced, for example, due to more efficient and accurate embedding generation, reduced duplication, or improved augmented sorting or retrieval. Examples of such computational resources may include processor cycles, network traffic, memory usage, graphics processing unit (GPU) resources, data storage capacity, power consumption, and cooling capacity.
[0042] Networked computing environment
[0043] Figure 1 This is a block diagram illustrating an example interactive system 100 for facilitating interactions over a network, such as exchanging text messages, making text, audio, and video calls, or playing games. Interactive system 100 includes multiple user systems 102, each of which hosts multiple applications including an interactive client 104 and other applications 106. Each interactive client 104 is communicatively coupled to other instances of the interactive client 104 (e.g., hosted on corresponding other user systems 102), an interactive server system 110, and a third-party server 112 via one or more communication networks including a network 108 (e.g., the Internet). The interactive client 104 may also communicate with the locally hosted applications 106 using an application programming interface (API).
[0044] Each user system 102 may include multiple user devices, such as mobile devices 114, head-mounted wearable devices 116, and computer client devices 118, that are communicatively connected to exchange data and messages.
[0045] Interactive client 104 interacts with other interactive clients 104 and with interactive server system 110 via network 108. The data exchanged between interactive clients 104 (e.g., interaction 120) and between interactive client 104 and interactive server system 110 includes functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).
[0046] Interactive server system 110 provides server-side functionality to interactive client 104 via network 108. While some functions of interactive system 100 are described herein as being performed by interactive client 104 or interactive server system 110, the location of certain functions within interactive client 104 or interactive server system 110 may be a design choice. For example, it may be technically preferred that certain technologies and functions are initially deployed within interactive server system 110, but later migrated to interactive client 104 where user system 102 has sufficient processing power.
[0047] The interactive server system 110 supports various services and operations provided to the interactive client 104. Such operations include sending data to and receiving data from the interactive client 104, and processing data generated by the interactive client 104. This data may include message content, client device information, geolocation information, content enhancements (e.g., filters and overlays), message content persistence conditions, entity relationship information, and live event information. Data exchange within the interactive system 100 is activated and controlled via functions available through the user interface (UI) of the interactive client 104.
[0048] Specifically, turning to interactive server system 110, API server 122 is coupled to interactive server 124 and provides it with a programming interface, making the functionality of interactive server 124 accessible to interactive client 104, other applications 106, and third-party server 112. Interactive server 124 is communicatively coupled to database server 126, thereby facilitating access to database 128, which stores data associated with the interactions processed by interactive server 124. Similarly, web server 130 is coupled to interactive server 124 and provides a web-based interface to interactive server 124. To this end, web server 130 handles incoming network requests via Hypertext Transfer Protocol (HTTP) and several other related protocols.
[0049] API server 122 receives and sends interactive data (e.g., command and message payloads) between interactive server 124 and user system 102 (as well as interactive client 104 and other applications 106) and third-party server 112. Specifically, API server 122 provides a set of interfaces (e.g., routines and protocols) that can be invoked or queried by interactive client 104 and other applications 106 to activate the functionality of interactive server 124. API server 122 exposes various functions supported by interactive server 124, including account registration; login functionality; sending interactive data from one interactive client 104 to another interactive client 104 via interactive server 124; transferring media files (e.g., images or videos) from interactive client 104 to interactive server 124; setting media data sets (e.g., stories); retrieving the friend list of users in user system 102; retrieving messages and content; adding and deleting entities (e.g., friends) against an entity relationship graph (e.g., entity graph 1410); locating friends in the entity relationship graph; and opening (e.g., application events associated with interactive client 104). Interactive server 124 hosts multiple systems and subsystems, as shown below. Figure 2 Describe it.
[0050] System Architecture
[0051] Figure 2 This is a block diagram illustrating further details of an interactive system 100 according to some examples. Specifically, the interactive system 100 is shown as including an interactive client 104 and an interactive server 124. The interactive system 100 includes multiple subsystems supported on the client side by the interactive client 104 and on the server side by the interactive server 124. In some examples, these subsystems are implemented as microservices. A microservice subsystem (e.g., a microservice application) may have components that enable the microservice subsystem to operate independently and communicate with other services. Example components of a microservice subsystem may include:
[0052] Functional logic: Functional logic implements the functions of the microservice subsystem and represents the specific capabilities or functions provided by the microservice.
[0053] API Interface: Microservices can communicate with each other through well-defined APIs or interfaces using lightweight protocols such as REST or messaging. API interfaces define the inputs and outputs of a microservice subsystem and how it interacts with other microservice subsystems within the interactive system 100.
[0054] Data storage: The microservice subsystem can be responsible for its own data storage, which can take the form of a database, cache, or other storage mechanisms (e.g., using database server 126 and database 128). This allows the microservice subsystem to operate independently of other microservices in the interactive system 100.
[0055] Service discovery: Microservice subsystems can find and communicate with other microservice subsystems in the interactive system 100. The service discovery mechanism enables microservice subsystems to locate and communicate with other microservice subsystems in a scalable and efficient manner.
[0056] Monitoring and logging: Microservice subsystems may need to be monitored and logged to ensure availability and performance. Monitoring and logging mechanisms enable the tracking of the health and performance of microservice subsystems.
[0057] In some examples, the interactive system 100 may employ a monolithic architecture, a service-oriented architecture (SOA), a function-as-a-service (FaaS) architecture, or a modular architecture. Example subsystems are discussed below.
[0058] Image processing system 202 provides various functions that enable users to capture and enhance (e.g., annotate or otherwise modify or edit) media content associated with a message. Camera device system 204 includes (e.g., in a camera device application) control software that (e.g., directly or via an operating system) interacts with and controls the hardware camera device hardware of user system 102 to modify and enhance real-time images captured and displayed via interactive client 104.
[0059] Enhancement system 206 provides functionality related to the generation and distribution of enhancements (e.g., filters or media overlays) for images captured in real-time by the camera device of user system 102 or retrieved from the memory of user system 102. For example, enhancement system 206 is operable to select, present, and display media overlays (e.g., image filters or image lenses) to interactive client 104 for enhancing real-time images received via camera device system 204 or stored images retrieved from memory 1602 of user system 102. These enhancements are selected by enhancement system 206 and presented to the user of interactive client 104 based on several inputs and data, such as:
[0060] The geographic location of user system 102; and
[0061] User entity relationship information of users in user system 102.
[0062] Enhancements can include audio and visual content as well as visual effects. As explained above, "visual enhancement" refers to visual effects applied to or capable of being applied to content. Examples of audio and visual content include images, text, logos, animations, and sound effects. Audio and visual content or visual effects can be applied to media content items (e.g., photos or videos) at user system 102 for transmission in messages, or to video content such as video content streams or feeds sent from interactive client 104. Therefore, image processing system 202 can interact with and support various subsystems of communication system 208, such as messaging system 210 and video communication system 212.
[0063] Media overlays may include text or image data that can be superimposed on photographs taken by user system 102 or video streams produced by user system 102. In some examples, media overlays may be location overlays (e.g., Venice Beach), names of live events, or names of businesses (e.g., beach cafes). In other examples, image processing system 202 uses the geolocation of user system 102 to identify media overlays that include the names of businesses located at the geolocation of user system 102. Media overlays may include additional identifiers associated with the businesses. Media overlays may be stored in database 128 and accessed through database server 126.
[0064] Image processing system 202 provides a user-based publishing platform that allows users to select a geographic location on a map and upload content associated with that location. Users can also specify which media overlays should be provided to other users. Image processing system 202 generates a media overlay that includes the uploaded content and associates it with the selected geographic location.
[0065] The augmented reality creation system 214 supports augmented reality developer platforms and includes applications for content creators (e.g., artists and developers) to create and publish augmentations (e.g., augmented reality experiences) to interactive clients 104. The augmented reality creation system 214 provides content creators with a library of built-in features and tools, including, for example, custom shaders, tracking technologies, and templates. In some examples, the augmented reality creation system 214 provides a merchant-based publishing platform that allows merchants to select specific augmentations associated with geolocation via a bidding process. For example, the augmented reality creation system 214 associates the media overlay of the highest-bidding merchant with a corresponding geolocation for a predefined amount of time.
[0066] Communication system 208 is responsible for enabling and processing various forms of communication and interaction within interactive system 100, and includes messaging system 210, audio communication system 216, and video communication system 212. Messaging system 210 is responsible for enabling temporary or time-limited access to content by interactive client 104. Messaging system 210 includes (e.g., in a short-lived timer system) multiple timers that selectively enable access (e.g., for presentation and display) of messages and associated content via interactive client 104 based on duration and display parameters associated with a message or set of messages (e.g., a story). Audio communication system 216 enables and supports audio communication (e.g., real-time audio chat) between multiple interactive clients 104. Similarly, video communication system 212 enables and supports video communication (e.g., real-time video chat) between multiple interactive clients 104.
[0067] User management system 218 is operationally responsible for managing user data and profiles, and maintaining entity information about users of interactive system 100 and the relationships between users (e.g., stored in entity table 1408, entity diagram 1410, and profile data 1402).
[0068] The collection management system 220 is operationally responsible for managing collections or sets of media (e.g., collections of text, images, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into "event galleries" or "event stories." Such collections can be made available for a specified time period (e.g., the duration of the event to which the content relates). For example, content related to a concert can be available as a "story" for the duration of the concert. The collection management system 220 can also be responsible for publishing icons that notify the user interface of the interactive client 104 of the availability of specific collections. The collection management system 220 includes curation functions that enable collection managers to manage and curate specific content collections. For example, a curation interface enables event organizers to curate collections of content related to a specific event (e.g., removing inappropriate content or redundant messages). Additionally, the collection management system 220 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, users may be compensated for including user-generated content in a collection. In such cases, the collection management system 220 operates to automatically pay such users for using their content.
[0069] Map system 222 provides various geolocation functions and supports the presentation of map-based media content and messages by interactive client 104. For example, map system 222 enables the display (e.g., stored in profile data 1402) of user icons or avatars on a map to indicate the current or past locations of the user's "friends" within the context of the map, as well as media content generated by these friends (e.g., a collection of messages including photos and videos). For example, on the map interface of interactive client 104, messages posted by a user from a specific geolocation to interactive system 100 can be displayed to the specific user's "friends" within the context of that specific location on the map. Users can also share their location and status information with other users of interactive system 100 via interactive client 104 (e.g., using an appropriate status avatar), where the location and status information is similarly displayed to selected users within the context of the map interface of interactive client 104.
[0070] Game system 224 provides various game functions within the context of interactive client 104. Interactive client 104 provides a game interface that offers a list of available games that can be initiated by a user within the context of interactive client 104 and played with other users of interactive system 100. Interactive system 100 also enables specific users to invite other users to participate in specific games by sending invitations from interactive client 104. Interactive client 104 also supports sending and receiving voice, video, and text messages (e.g., chat) within the context of playing the game, provides leaderboards for the game, and also supports providing in-game rewards (e.g., game currency and items).
[0071] External resource system 226 provides interactive client 104 with an interface to communicate with remote servers (e.g., third-party server 112) to launch or access external resources (i.e., applications or applets). Each third-party server 112 hosts applications or smaller versions of applications (e.g., game applications, utility applications, payment applications, or ride-sharing applications) based on markup languages (e.g., HTML5). Interactive client 104 can launch web-based resources (e.g., applications) by accessing HTML5 files from the third-party server 112 associated with the web-based resource. The application hosted by the third-party server 112 is programmed in JavaScript using a software development kit (SDK) provided by interactive server 124. The SDK includes application programming interfaces (APIs) with functionality that can be called or activated by the web-based application. Interactive server 124 hosts a JavaScript library that provides access to a given external resource for specific user data of interactive client 104. HTML5 is an example of a technology used for programming games, but applications and resources programmed using other technologies can be used.
[0072] To integrate the SDK's functionality into the web-based resource, the third-party server 112 downloads the SDK from the interactive server 124, or the third-party server 112 otherwise receives the SDK. Once downloaded or received, the SDK is included as part of the application code of the web-based external resource. The code of the web-based resource can then call or activate certain functions of the SDK to integrate the features of the interactive client 104 into the web-based resource.
[0073] The SDK stored on the interactive server system 110 effectively bridges the gap between external resources (e.g., application 106 or applet) and the interactive client 104. This provides users with a seamless experience communicating with other users on the interactive client 104 while preserving the look and feel of the interactive client 104. To bridge communication between the external resources and the interactive client 104, the SDK facilitates communication between a third-party server 112 and the interactive client 104. A bridging script running on the user system 102 establishes two unidirectional communication channels between the external resources and the interactive client 104. Messages are sent asynchronously between the external resources and the interactive client 104 via these communication channels. Each SDK function activation is sent as a message and callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message with that callback identifier.
[0074] By using the SDK, not all information from the interactive client 104 is shared with the third-party server 112. The SDK limits which information is shared based on the needs of the external resources. Each third-party server 112 provides the interactive server 124 with an HTML5 file corresponding to the web-based external resource. The interactive server 124 can add a visual representation (e.g., box design or other graphics) of the web-based external resource to the interactive client 104. Once the user selects the visual representation or instructs the interactive client 104 to access the features of the web-based external resource through the interactive client 104's GUI, the interactive client 104 obtains the HTML5 file and instantiates the resource for accessing the features of the web-based external resource.
[0075] Interactive client 104 presents a graphical user interface (GUI) for an external resource (e.g., a login page or title screen). During, before, or after presenting the login page or title screen, interactive client 104 determines whether the initiated external resource has previously been authorized to access user data of interactive client 104. In response to determining that the initiated external resource has previously been authorized to access user data of interactive client 104, interactive client 104 presents another GUI of the external resource, including its functionality and characteristics. In response to determining that the initiated external resource has not previously been authorized to access user data of interactive client 104, after displaying the login page or title screen of the external resource for a threshold time period (e.g., 3 seconds), interactive client 104 slides up a menu (e.g., animates the menu to appear from the bottom of the screen to the middle or other parts of the screen) to authorize the external resource to access user data. This menu identifies the type of user data that the external resource will be authorized to use. In response to receiving a user selection of the accept option, interactive client 104 adds the external resource to the list of authorized external resources and allows the external resource to access user data from interactive client 104. External resources are authorized by the interactive client 104 to access user data under the OAuth 2 framework.
[0076] Interactive client 104 controls the type of user data shared with external resources based on the type of authorized external resource. For example, it provides access to a first type of user data (e.g., two-dimensional avatars of users with or without different avatar characteristics) to external resources including full-scale applications (e.g., application 106). As another example, it provides access to a second type of user data (e.g., payment information, two-dimensional avatars of users, three-dimensional avatars of users, and avatars with various avatar characteristics) to external resources including smaller versions of applications (e.g., web-based versions of applications). Avatar characteristics include different ways of customizing the appearance and feel of an avatar (e.g., different poses, facial features, clothing, etc.).
[0077] The advertising system 228 enables third parties to purchase advertisements to be presented to end users via the interactive client 104, and also handles the delivery and presentation of these advertisements.
[0078] Artificial intelligence and machine learning system 230 provides various services to different subsystems within interaction system 100. For example, AI and machine learning system 230 operates in conjunction with image processing system 202 and camera device system 204 to analyze images and extract information, such as objects, text, or faces. This information can then be used by image processing system 202 to enhance, filter, or manipulate (e.g., apply visual enhancements) the image. AI and machine learning system 230 can be used by enhancement system 206 to generate enhanced content and augmented reality experiences, such as adding virtual objects or animations to real-world images. Communication system 208 and messaging system 210 can use AI and machine learning system 230 to analyze communication patterns and provide insights into how users interact with each other, as well as to provide intelligent message classification and tagging, such as classifying messages based on sentiment or topic.
[0079] The artificial intelligence and machine learning system 230 can also provide chatbot functionality for message interactions 120 between user systems 102 and between user systems 102 and the interaction server system 110. The artificial intelligence and machine learning system 230 can also provide generative capabilities, such as allowing users to generate text, images, or video content based on prompts. The artificial intelligence and machine learning system 230 can work with the audio communication system 216 to provide speech recognition and natural language processing capabilities, enabling users to interact with the interaction system 100 using voice commands.
[0080] Embedding and mapping system 232 is responsible for generating embeddings that capture visual enhancements within the interaction system 100 that are available to the user or are considered to be available to the user. Embedding and mapping system 232 may work with artificial intelligence and machine learning system 230 to provide training for one or more machine learning models for this purpose and to run inference on input content items (e.g., including enhanced videos) to extract embeddings from the input content items. Embedding and mapping system 232 is also responsible for storing embeddings, for example, storing aggregated embeddings in the form of enhancement identifiers, thereby creating a mapping between enhancements and their corresponding embeddings within the interaction system 100. Embedding and mapping system 232 may work with artificial intelligence and machine learning system 230 to activate machine learning models that consume these embeddings. For example, embedding and mapping system 232 may provide embeddings to an enhancement ranking machine learning model implemented by artificial intelligence and machine learning system 230, which uses the embeddings as input data to generate rankings for a set of visual enhancements.
[0081] Figure 3 Figure 300 illustrates the interaction between certain components of an interactive system 100, including an embedding and mapping system 232 and an artificial intelligence and machine learning system 230, according to some examples.
[0082] In some examples, pipelines are created to serve production traffic for creating embeddings for large-scale visual enhancements. The enhancement metadata service 302 of the embedding and mapping system 232 provides data related to the enhancements used within the interactive system 100, such as identification information, event information, usage statistics, or other metadata. This data may be stored in a database, for example, in the enhancement data storage component 304 that forms part of the embedding and mapping system 232.
[0083] Enhancement rendering service 306 is configured to render specific visual enhancements on base videos (also referred to as background videos). In other words, enhancement rendering service 306 provides an automated service capable of supplying input video items from which embeddings can be generated. Enhancement rendering service 306 can retrieve base videos from a library in database 128 and render visual enhancements onto each base video. For example, for a specific visual enhancement, enhancement rendering service 306 can render visual enhancements onto 50 base videos, 100 base videos, or 1000 base videos for use downstream in training or inference, as further described below. The rendered videos can be converted into a binary format (e.g., TFRecord format) for downstream processing by processing unit 308 of embedding and mapping system 232.
[0084] Enhanced rendering service 306 can also be used during the training phase. For example, and as further described below, enhanced rendering components (such as enhanced rendering service 306) can be used to apply predefined visual enhancements to a first base video and to a second base video to define positive training pairs for training a machine learning model.
[0085] Still refer to Figure 3 The processing unit 308 can perform various processing functions of the embedding and mapping system 232, as well as processing functions associated with the artificial intelligence and machine learning system 230. The processing unit 308 communicates with the enhanced metadata service 302, the enhanced rendering service 306, and the enhanced data storage unit 304 to obtain input video items and the metadata required to process those input video items. The processing unit 308 can also communicate with the media delivery platform 310, for example, to extract the frame rate of the corresponding input video. The processing unit 308 can instruct or invoke the artificial intelligence and machine learning system 230 to run inference on one or more input video items. For example, the processing unit 308 can communicate with the artificial intelligence and machine learning system 230 to utilize a trained machine learning model of the artificial intelligence and machine learning system 230, which generates embeddings that capture visual enhancements applied to a given input video. The training and use of such a machine learning model are further described below with reference to examples.
[0086] In use, the processing unit 308 can be handled by another component of the interactive system 100 (in... Figure 3 The model activation component 312 (shown as model activation component 312) is triggered or indicated. For example, in the case where embedding is used for the purpose of enhancing sorting (e.g., to create a list of similar enhancements to be displayed to the user on interactive client 104), model activation component 312 may be a component of a sorting system that sends a request for embedding data related to one or more specific visual enhancements.
[0087] Embedsions can be stored in the embedding data storage unit 314 by the processing unit 308. In some examples, embeddings generated for the same visual enhancement can be aggregated, for example, by generating embeddings of mean, average, concatenated, or other combinations thereof. Aggregated embeddings can be stored in the embedding data storage unit 314 and used, for example, as identifiers or representations of the corresponding visual enhancements for downstream tasks (e.g., sorting).
[0088] In some examples, the processing component 308 may periodically update the embedded data storage component 314, for example, to add embeddings for new visual enhancements within the interactive system 100.
[0089] Figure 4 This is a block diagram illustrating, in general, a machine learning program 400 based on some examples. The machine learning program (also referred to as a machine learning algorithm or tool) is used as part of the techniques and systems described herein to perform operations associated with the generation of embeddings.
[0090] This paper explores and constructs machine learning algorithms (also referred to as tools in this paper) that can learn from or be trained using existing data and make predictions about or based on new data. Such machine learning tools operate by building models based on training data 408 to make data-driven predictions or decisions that are expressed as outputs or evaluations (e.g., evaluations 416). Although the examples are presented relative to several machine learning tools, the principles presented in this paper can be applied to other machine learning tools.
[0091] In some examples, different machine learning tools can be used. For example, logistic regression (LR), Naive Bayes, random forest (RF), neural networks (NN), matrix factorization, and support vector machines (SVM) can be used. Two common types of problems in machine learning are classification problems and regression problems. Classification problems (also known as categorization problems) aim to classify an item into one of several category values (e.g., is the object an apple or an orange?). Regression algorithms aim to quantify some items (e.g., by providing values as real numbers).
[0092] The machine learning program 400 supports two types of phases: a training phase 402 and a prediction phase 404. In the training phase 402, supervised learning, unsupervised learning, or reinforcement learning can be used. For example, the machine learning program 400 (1) receives features 406 (e.g., structured or labeled data in supervised learning) and / or (2) identifies features 406 in the training data 408 (e.g., unstructured or unlabeled data for unsupervised learning). In the prediction phase 404, the machine learning program 400 uses features 406 to analyze query data 412 to generate results or predictions, as an example of evaluation 416 (this phase is also referred to as inference).
[0093] In training phase 402, feature engineering can be used to identify features 406 and may include identifying distinctive and independent features that provide useful information for efficient operation of machine learning program 400 in pattern recognition, classification, and regression. In some examples, training data 408 includes labeled data, which is known data about the pre-identified features 406 and one or more outcomes. Each of features 406 can be a variable or attribute, such as various measurable properties of a process, item, system, or phenomenon represented by a dataset (e.g., training data 408). By way of example only, features 406 may also have different types such as numerical features, strings, and graphs, and may include one or more of content 418, concepts 420, attributes 422, historical data 424, and / or user data 426.
[0094] In this context, the concept of features is related to the concept of explanatory variables used in statistical techniques such as linear regression. Selecting distinctive and independent features that provide useful information is crucial for the efficient operation of machine learning programs in pattern recognition, classification, and regression. Features can be of different types, such as numeric features, strings, and graphs.
[0095] In training phase 402, machine learning program 400 uses training data 408 to find correlations between features 406 that influence prediction results or evaluations 416. Using training data 408 and the identified features 406, machine learning program 400 is trained during training phase 402 at machine learning program training 410. Machine learning program 400 evaluates the value of feature 406 when it is correlated with training data 408. The result of training is a trained machine learning program 1514 (e.g., a trained or learned model).
[0096] Furthermore, training phase 402 may involve machine learning, in which training data 408 is structured (e.g., labeled during preprocessing operations), and the trained machine learning program 414 implements a relatively simple neural network 428 capable of performing, for example, classification and clustering operations. In other examples, training phase 402 may involve deep learning, in which training data 408 is unstructured, and the trained machine learning program 414 implements a deep neural network 428 capable of performing both feature extraction and classification / clustering operations.
[0097] The neural network 428, generated during training phase 402 and implemented within the trained machine learning program 414, may include a hierarchical (e.g., layered) organization of neurons. For example, neurons (or nodes) may be arranged hierarchically into several layers, including an input layer, an output layer, and multiple hidden layers. Each layer within the neural network 428 may have one or more neurons, and each of these neurons is operationally computed to a small function (e.g., an activation function). For example, if the activation function produces a result exceeding a certain threshold, the output can be transmitted from that neuron (e.g., a sending neuron) to connected neurons (e.g., receiving neurons) in successive layers. The connections between neurons also have associated weights that define the influence of the input from the sending neuron to the receiving neuron.
[0098] In some examples, and by way of example only, neural network 428 may also be one of several different types of neural networks, including single-layer feedforward networks, artificial neural networks (ANNs), recurrent neural networks (RNNs), symmetrically connected neural networks, unsupervised pre-trained networks, convolutional neural networks (CNNs), or recurrent neural networks (RNNs).
[0099] During the prediction phase 404 or inference, an evaluation is performed using a trained machine learning program 414. Query data 412 is provided as input to the trained machine learning program 414, and in response to receiving query data 412, the trained machine learning program 414 generates an evaluation 416 as output.
[0100] As mentioned above, Figure 4 This provides a general overview of some aspects of training and using machine learning programs. Now turn to Figure 5 Based on some examples, a more detailed process 500 is shown to illustrate the training of a machine learning model. Specifically, process 500 involves training a machine learning model to generate embeddings representing visual enhancements appearing within input video items.
[0101] Training is performed using self-supervised contrastive learning (specific forms of "contrastive learning" are discussed below). Figure 5 Machine learning models. Self-supervised learning is a type of unsupervised learning that involves training a neural network to make predictions without using explicit labels. In self-supervised learning, the model creates its own supervision by defining a task that requires the network to learn useful things about the data. Specifically, in Figure 5 In this context, the training process involves a machine learning model learning to generate useful representations of the input data from representations of different views that include the same visual enhancements. Figure 5 The operations in the process can be performed, for example, by an artificial intelligence and machine learning system 230 and / or an embedding and mapping system 232.
[0102] Figure 5 The goal of the training phase is to bring views with similar enhancements closer together in the feature space, while pushing views with dissimilar enhancements apart as training progresses. This is achieved by allowing, for example, the generation of embeddings that can be used to compare enhancements, identify enhancements within other content, or deduplicate enhancements during the retrieval phase.
[0103] Unsupervised training phase Figure 5 The diagram is shown as including a term used as input 502, and a view layer 504, a representation layer 506, a projection layer 508, and a prediction layer 510. Each layer processes a pair of training video terms. At higher levels, Figure 5 The architecture shown can be similar to certain aspects of the BYOL (Bootstrap Your Own Latent) architecture used to train deep neural networks, for example, by adding the predictor component to one "branch" (associated with one of the input items in the input pair), but not to another "branch" (associated with the other item in the input pair). Figure 5 The method is similar to BYOL's, which lies in Figure 5 The method utilizes positive training pairs. In the traditional sense, contrastive learning usually involves learning to distinguish between positive and negative pairs, where negative pairs are drawn from a distribution different from the positive pairs.
[0104] Referring to input 502, and as mentioned above, while contrastive learning typically utilizes both positive and negative pairs of input samples, the example in this disclosure utilizes only positive pairs, similar to the BYOL method. However, unlike the BYOL method, which requires a single sample as input and then transforms that single sample into two augmented views, Figure 5The architecture leverages a direct comparison of two different videos. For example, two input video items can be generated by applying the same visual enhancement 512 to two different videos (e.g., a first base video 514 and a second base video 516 with different backgrounds, objects, content, etc.). This differs from the BYOL method, where two enhanced views are typically generated by applying random modifications (e.g., rotations, cropping, etc., applied to the same original image) to the same input samples.
[0105] The enhanced rendering service 306 can generate the input 502 required for each training set during the training phase. It should be understood that a large training set (e.g., video pairs) covering multiple different visual enhancements can be utilized, and Figure 5 For illustrative purposes, only one set is shown. Therefore, each training set may include a pair of training video items. Within each pair, the first training video item may include a first video content with a first predefined visual enhancement applied, and the second training video item may include a second video content (different from the first video content) with the same predefined visual enhancement applied.
[0106] In some examples, the enhancement rendering service 306 uses the same large set of background videos for each visual enhancement. In other words, each visual enhancement is rendered onto all background videos in the set utilized by the enhancement rendering service 306. Therefore, a machine learning model can be trained on a pair of these training sets to improve its ability to extract important visual features of the enhancements from larger videos.
[0107] Now return to Figure 5 The first base video 514, which applies visual enhancement 512, defines the first training video item 518, and the second base video 516, which applies visual enhancement 512, defines the second training video item 520 within the view layer 504. The view layer 504 acquires the "raw" video content items and preprocesses them into a form suitable for downstream analysis. Preprocessing may involve resizing video frames, cropping video frames, or normalizing video frames.
[0108] Starting from the view layer 504, the two input samples are transformed into latent space vectors in the representation layer 506 to obtain representations of the first training video item 522 and the second training video item 524. For example, the first training video item 518 and the second training video item 520 can be input to the video encoder in binary format to obtain representations. In some examples, a video encoder such as MoViNet (Mobile Video Network) can be applied in the representation layer 506.
[0109] The term "MoViNet" refers to a lightweight family of CNNs designed for video understanding tasks. MoViNet models have an encoder-decoder architecture and can be trained to process video frames using 3D convolutional layers to capture both spatial and temporal information in the input video. In conventional 2D CNNs, only spatial information in the image is captured. However, in models such as MoViNet, filters can slide along both the spatial and temporal dimensions of the video, capturing information about how objects move and change over time. This can facilitate the capture of visual effects in the video, including how those effects change over time. A MoViNet model may include a backbone network followed by one or more fully connected layers. The backbone network uses convolutional layers to extract features from the input video frames, and the one or more fully connected layers generate a vector representation of the input video (e.g., a representation of a first training video term 522 or a representation of a second training video term 524).
[0110] Therefore, for a specific sample, representation layer 506 acquires preprocessed video frames and transforms them into compact representations that capture the key features of the sample. For example, the representations of the first training video item 522 and the second training video item 524 generated by the MoViNet model can be viewed as compact representations of key features and patterns in the frames of the first training video item 518 and the second training video item 520, respectively. In this context, the representation can be a sequence of feature vectors or a single vector summarizing the entire video. It should be understood that MoViNet is merely an example of the type of model that can be used to generate representations in some examples. Other types of models, such as other deep neural network designs, can be employed in the examples of this disclosure.
[0111] In some examples, and such as Figure 5 As shown, the representations of the first training video item 522 and the second training video item 524 can be projected, for example, by corresponding multilayer perceptrons (MLPs) to a smaller space to obtain a first projection 526 and a second projection 528. In some examples, the input video items can thus undergo a nonlinear transformation to be mapped to a high-dimensional representation in representation layer 506 and then compressed into a low-dimensional space in projection layer 508 while preserving the most important information. Projection layer 508 can facilitate downstream comparison of the vector representations of the original video items.
[0112] The representation of the first training video item 518 can then be transformed to produce a transformed representation 530. For example, a predictor component (e.g., a predictor neural network) can transform the first projection 526 into the transformed representation 530. In some examples, the predictor component of the prediction layer 510 includes a neural network that takes the first projection 526 as input and applies a series of linear and nonlinear transformations to the first projection. The predictor network may include several fully connected layers followed by a nonlinear activation function (e.g., ReLU (rectifier linear unit)). The output of the predictor network is presented as... Figure 5 The transformed representation of 530 is depicted as a low-dimensional vector in an example form.
[0113] Note that in some examples, and such as Figure 5 As in the case described above, assuming the predictor component is applied to only one "branch" during training, only one of the terms in the training pair is transformed to create the final prediction target. Therefore, the first projection 526 can be viewed as representing the "online" branch of the model ( Figure 5 The left side), while the second projection 528 can be regarded as representing the "target" branch of the model ( Figure 5 (to the right). The transformed representation 530 (final predicted target) is used for comparison with the representation of the "target" (e.g., the second projection 528).
[0114] Specifically, such as Figure 5 As shown in the figure, loss function 532 analyzes the transformed representation q(z) 530 and the projected representation of the second input video term. The similarity between two representations. For example, loss function 532 can utilize the cosine similarity between these two representations. Cosine similarity provides a measure of the similarity between two non-zero vectors in space. Cosine similarity measures the cosine of the angle between two vectors in multidimensional space. A suitable formula for cosine similarity is:
[0115] ,
[0116] Here, AB represents the dot product of vectors A and B, and ||A|| and ||B|| represent the Euclidean norms of A and B, respectively.
[0117] As defined above, cosine similarity ranges from -1 to 1, where 1 indicates that two vectors are identical, and -1 indicates that two vectors are completely dissimilar. To minimize the loss, in the context of cosine loss, it is desirable to make the cosine similarity as close to 1 as possible (e.g., ...). ).
[0118] The use of loss functions (such as those described above) encourages machine learning models (e.g., example predictor networks described above) to learn representations similar to the “target”, for example, to make the two vectors as similar as possible.
[0119] Assuming that both the predicted representation and the "target" are based on different videos including the same visual enhancements, this encourages the machine learning model to extract a useful representation of the visual enhancements from the representation of the entire input video item. As training progresses, the machine learning model can learn a mapping between a high-dimensional representation of the input video item and a target embedding that captures the most relevant features of the visual enhancements included in the input video item, regardless of (or largely regardless of) other elements present in the video (e.g., unenhanced background objects).
[0120] Update the parameters of the machine learning model to minimize the loss between training video representations within subsequent training video item pairs based on the loss function. For example, update the weights of the predictor network using backpropagation (e.g., using stochastic gradient descent (SGD) or other optimization algorithms) to minimize the loss function.
[0121] In some examples, a sufficiently large number of training samples can be used, such as 20 million, 25 million, 30 million, 35 million, or 40 million (for example only). See reference. Figure 5 One or more components of the online branch (e.g., at 522, 526, 530) can be updated via gradient descent, while gradients are blocked for the target branch (e.g., at 524, 528). In the latter case, the weights can be updated by periodically incorporating the weights of the online branch into the weights of the target branch. Thus, the target branch can provide a stable and slowly shifting objective from which the online branch learns, for example, by using an exponential moving average to update the parameters of the target branch.
[0122] Therefore, in some examples, for each training pair, a transformed representation is generated based on the embedding of the first training video item, and the transformed representation is compared with the embedding of the second training video item. Minimizing the loss between these training video representations involves using an appropriate loss function to measure the similarity between the transformed representation and the embedding of the second training video item. During training, it is assumed that the visual enhancements are common to the items in each training pair, and minimizing the loss function causes the machine learning model to generate output embeddings that capture the visual enhancements applied to the input video items.
[0123] Once the machine learning model has been trained, for example, as referenced Figure 5 As described, this machine learning model can be used to map specific visual enhancements available within the interaction system 100 to enhancement identifiers. Enhancement identifiers can be, for example, Figure 6The embedding shown is a list of floating-point numbers that captures the visual effect of the augmentation identifier. The model's architecture and training ensure that embeddings generated from content items that include dissimilar content (e.g., different background videos) but include the same visual augmentation are similar vectors, while embeddings generated from content items that include dissimilar visual augmentations (even if other content is similar) are dissimilar vectors.
[0124] Figure 6 Figure 600 conceptually illustrates the generation of the embedding 602 that captures the visual features or effects of enhancement 604. It should be noted that... Figure 6 as well as Figures 11 to 13 The example embeddings shown are simplified examples intended to represent, but not depict, the actual embeddings.
[0125] Use the above reference Figure 5 The training method outlined herein enables a trained machine learning model to output useful embeddings within the context of an interactive system 100 that allows users to apply visual enhancements to content. Specifically, in some examples, a trained “online branch” can be used to run inference on new inputs to generate such a useful embedding. In some cases, the “target branch” used during training may be discarded for inference purposes.
[0126] For example only, Figure 7 , Figure 8 and Figure 9 Frames from corresponding enhanced video pairs are shown, based on some examples. Figure 7 In the first video item 702, a first enhancement 706 is applied to the frame sequence. The first enhancement 706 presents flames or a "burning appearance" on the facial region of a person depicted in the video. A second video item 704 has a second enhancement 708 applied to the frame sequence. The second enhancement 708 presents flames above the head of a person depicted in the video. The first enhancement 706 and the second enhancement 708 are not identical, but they have similar visual effects and context. Therefore, running inference on the first video item 702 and the second video item 704 using a machine learning model trained according to the example in this paper produces similar outputs (vectors that are close to each other in the relevant space). In other words, because the machine learning model is trained to focus on visual enhancements and is substantially invariant to background elements and other parts of the video unrelated to the enhancements, the machine learning model can extract comparable and useful outputs from different videos with the same or similar enhancements applied. Therefore, the model learns to encode the input video item in a way that captures desired information about any visual enhancements (or multiple enhancements) that can be applied to the content of the input video item.
[0127] exist Figure 8In the first video item 802, a first enhancement 806 is applied to the frame sequence. The first enhancement 806 renders a "green screen" effect, where the user captures a video being shown in front of a selected "green screen" style background. The second video item 804 has a second enhancement 808 applied to the frame sequence. The second enhancement 808 also has a "green screen" effect, but the user has selected a different background than that used in the first video item 802. Therefore, the first enhancement 806 and the second enhancement 808 may not be visually or aesthetically similar, but they share a similar visual effect in that they render a "green screen." Therefore, performing inference on these two videos can also produce similar embeddings.
[0128] exist Figure 9 In the first video item 902, there is an enhancement 906 applied to the frame sequence. Enhancement 906 is for dancing Supermario. TM The character is rendered as a 3D augmented reality effect on frames captured by the user. The second video item 904 has the same enhancement 906 applied to its frames. Similarly, assuming that two video items 902 and 904 have visual enhancements applied to them that cause substantially the same visual effect, running inference on these two videos can produce similar embeddings, even if other content in video items 902 and 904 may be visually different.
[0129] The above description and Figures 7 to 9 The enhancements shown are just examples, and other visual enhancements such as various augmented reality effects or video filters can be employed.
[0130] As described above, the trained machine learning model generates embeddings from the input video items, which are vector representations of visual effects created by enhancements within the input video items. Figure 10 This is a flowchart illustrating a method 1000 for generating embeddings of target visual enhancements using a machine learning model based on some examples and mapping the target visual enhancements to enhancement identifiers.
[0131] Machine learning models can be based on references Figure 5 The unsupervised training process described describes the model trained. (Refer to...) Figure 10 The described operations can be performed by components of the interactive system 100, including, for example, an artificial intelligence and machine learning system 230 and an embedding and mapping system 232.
[0132] Method 1000 begins at the start loop element 1002 and proceeds to operation 1004, where the embedding and mapping system 232 accesses a first input video item including a target visual enhancement at operation 1004. The "target" visual enhancement is a specific visual enhancement used within the interactive system 100 to which it seeks embedding. In other words, a target visual enhancement can be considered as an enhancement of interest. For example, it can be targeted at... Figure 7 One of the visual enhancements depicted in the image, the "burning appearance," seeks or requests embedding.
[0133] The first input video item includes the video content and the target visual enhancement. In some examples, the input video item can be generated in post-production using Enhanced Rendering Service 306 or another Enhanced Rendering component. Alternatively, user-captured video can be retrieved from database 128, in which the target visual enhancement is applied to the video stream.
[0134] At operation 1006, a first input video item is fed as query data to a machine learning model, and the machine learning model (e.g., using a trained “online branch” as described above) generates a first embedding. The first embedding is a first vector representation capturing the visual effect of the target visual enhancement within the first input video item. This can be obtained using a suitable video model (e.g., a MoViNet-type model as described above). As mentioned, the embedding and mapping system 232 can leverage the capabilities of the artificial intelligence and machine learning system 230 to run inference using the machine learning model. The input video item can be consumed by the machine learning model in an appropriate format (e.g., binary format), transformed into a representation, and then transformed into a lower-dimensional representation by the predictor network to output the final model output (evaluation).
[0135] At operation 1008, the embedding and mapping system 232 accesses the second input video item. The second input video item also includes a targeted visual enhancement, but this targeted visual enhancement is not applied to the same video content as in the first input video item. For example, the first input video item can be rendered by applying the targeted visual enhancement to the first substrate video, and the second input video item can be rendered by applying the targeted visual enhancement to the second substrate video.
[0136] At operation 1010, a second input video item is fed as query data into a machine learning model, and the machine learning model generates a second embedding. The second embedding is a second vector representation that captures the visual effect of the target visual enhancement within the first input video item. Given that the machine learning model is trained to minimize the loss between pairs, and assuming that the same target visual enhancement is applied to both input video items, the machine learning model produces two similar embeddings. In the context of this disclosure, when used relative to vector representations such as embeddings, the term "similarity" means that the two representations are relatively close to each other in the vector space, as opposed to representations generated for different visual enhancements that are relatively far apart in the vector space. This "similarity" can be measured using appropriate metrics such as cosine similarity or Euclidean distance.
[0137] At operation 1012, the embedding and mapping system 232 generates an aggregated embedding based on the first and second embeddings. For example, processing unit 308 can obtain these two embeddings and determine the mean of the first and second embeddings (e.g., a new vector based on the average of the corresponding values of the two initial vectors). At operation 1014, the embedding and mapping system 232 then designates the aggregated embedding as an enhancement identifier for the target visual enhancement, and this relationship can be recorded in the embedding data storage unit 314 or another database 128. In this way, a one-to-one mapping between the target visual enhancement and the enhancement identifier (in this case, the aggregated embedding) is determined (operation 1016) and can be used to facilitate various tasks. The operations of method 1000 can be repeated for multiple other target visual enhancements to map these additional target visual enhancements to their corresponding embeddings. Method 1000 ends at the closed-loop element 1018.
[0138] exist Figure 10 In the example, two embeddings are generated for a specific visual enhancement, and these two embeddings are aggregated to obtain a final embedding used to identify the visual enhancement and compare it with other visual enhancements. However, it should be understood that this approach is merely an example. In some cases, only one embedding may be generated for each visual enhancement without any aggregation. In other cases, more than two embeddings may be generated for each target visual enhancement. For example, 50, 100, or 1000 embeddings may be generated based on all the different videos that include a specific target visual enhancement (e.g., all background videos in the dataset used by the enhancement rendering service 306) to produce a reliable final or aggregated enhancement identifier.
[0139] Figure 11Figure 1100 conceptually illustrates an embedding 1102 generated to capture the visual effect of a target visual enhancement in the example form of a “big eye” enhancement 1104. Enhancement 1104 is a visual enhancement rendered, for example, by the image processing system 202 of the interactive system 100 by modifying the user’s eye region captured in the video stream.
[0140] Figure 12 Figure 1200 conceptually illustrates an embedding 1202 generated to capture the visual effect of another targeted visual enhancement in the example form of a “funny face” enhancement 1204. Enhancement 1204 is a visual enhancement rendered, for example, by the image processing system 202 of the interactive system 100 by adapting the facial features of the user captured in the video stream to have a different appearance.
[0141] As mentioned above, the embeddings generated based on the examples in this disclosure can be used for a variety of purposes. Figure 13 Figure 1300 is a conceptual illustration of the enhancement identifier generated relative to the visual enhancement of interest during the nearest neighbor process.
[0142] like Figure 13 As shown, query embedding 1304 is generated for query enhancement 1302. For example, the embedding and mapping system 232 of the interactive system 100 can receive a video or still image including query enhancement 1302, obtain its embedding, and perform an automated query process 1306 to retrieve similar enhancements or “nearest neighbor” enhancements from the set of embeddings stored in the embedding data storage component 314.
[0143] The embedding and mapping system 232 uses a vector representation of the visual effect of query enhancement 1302 to capture query enhancement 1302, and retrieves four similar embedding sets, for example, the vector representation of the visual enhancement that is closest to query embedding 1304 in the vector space. Figure 13 As shown, the retrieved embedding or nearest neighbor 1308 is similar to but not the same as the query embedding 1304, and their corresponding visual enhancements therefore present similar visual effects when compared with the query enhancement 1302.
[0144] Retrieving such embeddings can be useful, for example, for presenting a possible set of visual enhancements to the user of interactive client 104 for selection. A similar approach can also be followed to identify visual enhancements that have been copied or may be copied, for example, in database 128 of interactive system 100, thereby ensuring that the user of interactive client 104 is not presented with copied or irrelevant content.
[0145] The embedding and mapping system 232 can also compare a target enhancement identifier (e.g., query embedding 1304) with identifiers of one or more reference visual enhancements to determine the similarity between the target visual enhancement and the reference visual enhancements. In some examples, instead of identifying nearest neighbors 1308 (or other than identifying nearest neighbors 1308), the query process 1306 can identify matches. For example, if query embedding 1304 is the same as or has a similarity threshold greater than that of an embedding stored in the embedding data storage component 314, the embedding and mapping system 232 can determine that query enhancement 1302 matches the stored enhancement. Matching can be useful, for example, for accurately identifying visual enhancements in content uploaded by a user to the interaction system 100. In such a case, the interaction system 100 can query the database 128 to identify the target visual enhancement based on the enhancement identifier (the embedding generated for the uploaded content).
[0146] These and other downstream activities can leverage machine learning models. Examples include augmentation ranking models (e.g., for suggesting the best augmentations for a particular scene or for a particular scene), augmentation retrieval models (which may be referred to as lens-to-lens retrieval models), augmentation deduplication models, augmentation labeling models, or matching models designed to estimate the fit between augmentations and scenes (e.g., images or videos).
[0147] Data Architecture
[0148] Figure 14 This is a schematic diagram illustrating a data structure 1400 that may be stored in a database 1404 (e.g., database 128 or another database) of an interactive server system 110, according to certain examples. Although the contents of database 1404 are shown as including multiple tables, it should be understood that data may be stored in other types of data structures (e.g., object-oriented databases).
[0149] Database 1404 includes message data stored in message table 1406. For any given message, this message data includes at least message sender data, message receiver (or recipient) data, and payload. See below for reference. Figure 15 Further details are provided regarding information that can be included in the message and is contained within the message data stored in message table 1406.
[0150] Entity table 1408 stores entity data and (for example, links to) entity diagram 1410 and profile data 1402. Entities for which records are maintained in entity table 1408 can include individuals, company entities, organizations, objects, locations, events, etc. Regardless of the entity type, any entity for which the interactive server system 110 stores data can be an identifiable entity. Each entity is assigned a unique identifier and an entity type identifier (not shown).
[0151] Entity Graph 1410 stores information about relationships and associations between entities. As an example only, such relationships can be social, professional (e.g., working in a common company or organization), interest-based, or activity-based. Some relationships between entities can be one-way, such as an individual user subscribing to digital content (e.g., a newspaper or other digital media channel or brand) for a business or publishing user. Other relationships can be two-way, such as the "friend" relationship between the various users of Interactive System 100.
[0152] Certain licenses and relationships can be attached to each relationship, and also to each direction of the relationship. For example, a two-way relationship (e.g., a friend relationship between individual users) may include authorization for the public disclosure of digital content items between the individual users, but certain restrictions or filters may be imposed on the public disclosure of these digital content items (e.g., based on content characteristics, location data, or time of day data). Similarly, a subscription relationship between an individual user and a business user may impose varying degrees of restrictions on the public disclosure of digital content from the business user to the individual user, and may significantly limit or prevent the public disclosure of digital content from the individual user to the business user. As an example of an entity, a specific user may (e.g., through privacy settings) record certain restrictions in the records for that entity within entity table 1408. Such privacy settings may be applied to all types of relationships in the context of the interaction system 100, or selectively applied to certain types of relationships.
[0153] Profile data 1402 stores various types of profile data about a specific entity. Based on privacy settings specified by the specific entity, profile data 1402 can be selectively used and presented to other users of the interaction system 100. In the case of an individual, profile data 1402 includes, for example, a username, phone number, address, settings (e.g., notification and privacy settings), and an avatar representation (or a set of such avatar representations) selected by the user. The specific user can then selectively include one or more of these avatar representations within the content of messages transmitted via the interaction system 100 and on a map interface displayed to other users by the interaction client 104. The set of avatar representations may include “status avatars,” which present a graphical representation of a status or activity that the user can choose to transmit at a specific time.
[0154] In the case that the entity is a group, in addition to the group name, members and various settings for the relevant group (e.g., notifications), the profile data 1402 for the group may similarly include one or more avatars associated with the group.
[0155] Database 1404 also stores enhancement data, such as overlays or filters, in enhancement table 1412. Enhancement data is associated with and applied to videos (video data is stored in video table 1414) and images (image data is stored in image table 1416). Enhancements can be visual enhancements or other types of enhancements, such as audio enhancements or combinations thereof.
[0156] In some examples, filters are displayed as overlays on images or videos during presentation to the recipient user. Filters can be of various types, including user-selected filters from a set of filters presented to the sending user by the interactive client 104 when the sending user is composing a message. Other types of filters include geolocation filters (also known as geographic filters), which can be presented to the sending user based on geographic location. For example, geolocation filters specific to nearby or particular locations can be presented by the interactive client 104 within the user interface based on geolocation information determined by the Global Positioning System (GPS) unit of the user system 102.
[0157] Another type of filter is a data filter, which can be selectively presented to the sending user by the interactive client 104 based on other inputs or information collected by the user system 102 during the message creation process. Examples of data filters include the current temperature at a specific location, the sending user's current speed, the battery life of the user system 102, or the current time.
[0158] Other augmented data that can be stored in image table 1416 includes augmented reality content items (e.g., corresponding to the application of a "lens" or augmented reality experience). Augmented reality content items can be real-time special effects and sounds that can be added to images or videos.
[0159] Collection table 1418 stores data about collections of messages and associated image, video, or audio data, compiled into collections (e.g., stories or galleries). The creation of a specific collection can be initiated by a specific user (e.g., each user for whom records are maintained in entity table 1408). A user can create a "personal story" in the form of a collection of content that has been created and sent / broadcast by that user. For this purpose, the user interface of interactive client 104 may include user-selectable icons that allow the sending user to add specific content to his or her personal story.
[0160] The collection can also constitute a "live story," which is a collection of content from multiple users created manually, automatically, or using a combination of manual and automatic technologies. For example, a "live story" can constitute a curated stream of user-submitted content from various locations and events. Users whose client devices have location services enabled and who are at a co-located event at a specific time can be presented with the option to contribute content to a specific live story, for example, via the user interface of interactive client 104. Live stories can be identified to a user by interactive client 104 based on their location. The end result is a "live story" told from a community perspective.
[0161] Another type of content collection is called a "location story," which allows users of user system 102 located in a specific geographic location (e.g., on a college or university campus) to contribute to a specific collection. In some examples, contributions to a location story may employ secondary authentication to verify that the end user belongs to a specific organization or other entity (e.g., is a student on a university campus).
[0162] As mentioned above, Video Table 1414 stores video data, which in some examples is associated with messages for which records are maintained within Message Table 1406. Video Table 1414 may also store video items used as input to machine learning models to generate embeddings as described herein, for example, as query data or training data for the inference phase. Similarly, Image Table 1416 stores image data associated with messages whose message data is stored in Entity Table 1408. Entity Table 1408 can associate various enhancements from Enhancement Table 1412 with various images and videos stored in Image Table 1416 and Video Table 1414.
[0163] Database 1404 also includes an embedding table 1420, which stores data relating to embeddings generated according to the examples described herein. Embedding table 1420 may include, for example, a mapping of visual enhancements to corresponding embeddings, aggregated embeddings, or details of individual embeddings used to determine aggregated embeddings.
[0164] Data communication architecture
[0165] Figure 15This is a schematic diagram illustrating the structure of message 1500 according to some examples, generated by interactive client 104 for transmission to another interactive client 104 via interactive server 124. The content of a particular message 1500 is used to populate message table 1406 stored in database 1404 accessible by interactive server 124. Similarly, the content of message 1500 is stored in memory as "in-transit" or "in-flight" data for user system 102 or interactive server 124. Message 1500 is shown to include the following example components:
[0166] Message Identifier 1502: A unique identifier that identifies message 1500.
[0167] Message text payload 1504: The text to be generated by the user via the user interface of user system 102 and included in message 1500.
[0168] Message image payload 1506: Image data captured by the camera device component of user system 102 or retrieved from the memory component of user system 102 and included in message 1500. The image data for the sent or received message 1500 can be stored in image table 1416.
[0169] Message video payload 1508: Video data captured by the camera device component or retrieved from the memory component of the user system 102 and included in message 1500. The video data for the sent or received message 1500 can be stored in image table 1416.
[0170] Message audio payload 1510: Audio data captured by the microphone or retrieved from the memory component of the user system 102 and included in message 1500.
[0171] Message enhancement data 1512: This represents enhancement data (e.g., filters, stickers, or other annotations or enhancements) to be applied to the message image payload 1506, message video payload 1508, or message audio payload 1510 of message 1500. Enhancement data for the sent or received message 1500 can be stored in enhancement table 1412.
[0172] Message duration parameter 1514: A parameter value, in seconds, indicating the amount of time that the content of the message (e.g., message image payload 1506, message video payload 1508, message audio payload 1510) will be presented to the user via the interactive client 104 or made accessible to the user.
[0173] Message geolocation parameter 1516: Geolocation data (e.g., latitude and longitude coordinates) associated with the message's content payload. Multiple message geolocation parameter 1516 values may be included in the payload, each of which is associated with a content item included in the content (e.g., a specific image within the message image payload 1506 or a specific video within the message video payload 1508).
[0174] Message Story Identifier 1518: An identifier value that identifies one or more sets of content (e.g., "Stories" identified in set table 1418) associated with a specific content item in the message image payload 1506 of message 1500. For example, the identifier value can be used to associate multiple images within the message image payload 1506 with multiple sets of content, respectively.
[0175] Message Tag 1520: Each message 1500 can be labeled with multiple tags, each of which indicates the subject of the content included in the message payload. For example, in the case where a specific image depicts an animal (e.g., a lion) is included in the message image payload 1506, a tag value can be included within the message tag 1520 indicating the relevant animal. Tag values can be manually generated based on user input, or can be automatically generated using, for example, image recognition.
[0176] Message sender identifier 1522: An identifier (e.g., message sending system identifier, email address, or device identifier) indicating the user of the user system 102 on which message 1500 is generated and from which message 1500 is sent.
[0177] Message receiver identifier 1524: An identifier (e.g., message sending and receiving system identifier, email address, or device identifier) indicating the user of the user system 102 to which message 1500 is addressed.
[0178] The content (e.g., values) of each component of message 1500 can be pointers to locations in tables where content data values are stored. For example, an image value in message image payload 1506 can be a pointer to a location within image table 1416 (or the address of a location within image table 1416). Similarly, a value in message video payload 1508 can point to data stored in image table 1416, a value in message enhancement data 1512 can point to data stored in enhancement table 1412, a value in message story identifier 1518 can point to data stored in set table 1418, and values in message sender identifier 1522 and message receiver identifier 1524 can point to user records stored in entity table 1408.
[0179] Systems with head-worn devices
[0180] Figure 16 A system 1600 including a head-worn wearable device 116 with a selector input device is shown according to some examples. Figure 16 This is a high-level functional block diagram of an example head-mounted wearable device 116 that is communicatively coupled to mobile devices 114 and various server systems 1604 (e.g., interactive server system 110) via various networks 108.
[0181] The head-mounted wearable device 116 includes one or more camera devices, each of which may be, for example, a visible light camera 1606, an infrared emitter 1608, and an infrared camera 1610.
[0182] Mobile device 114 connects to head-mounted wearable device 116 using both low-power wireless connection 1612 and high-speed wireless connection 1614. Mobile device 114 also connects to server system 1604 and network 1616.
[0183] The head-mounted wearable device 116 also includes two image displays in the image display 1618 of the optical components. The two image displays 1618 of the optical components include an image display associated with the left lateral side of the head-mounted wearable device 116 and an image display associated with the right lateral side of the head-mounted wearable device 116. The head-mounted wearable device 116 also includes an image display driver 1620, an image processor 1622, a low-power circuitry system 1624, and a high-speed circuitry system 1626. The image displays 1618 of the optical components are used to present images and videos to the user of the head-mounted wearable device 116, including images that may include a graphical user interface.
[0184] The image display driver 1620 commands and controls the image display 1618 of the optical components. The image display driver 1620 can directly deliver image data to the image display 1618 of the optical components for presentation, or it can convert image data into a signal or data format suitable for delivery to the image display device. For example, the image data can be video data formatted according to compression formats such as H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, etc., while still image data can be formatted according to compression formats such as Portable Network Group (PNG), Joint Photographic Experts Group (JPEG), Tagged Image File Format (TIFF), or Exchangeable Image File Format (Exif).
[0185] The head-wearable device 116 includes a frame and a stem (or temple) extending laterally from the frame. The head-wearable device 116 also includes a user input device 1628 (e.g., a touch sensor or a press button), comprising an input surface on the head-wearable device 116. The user input device 1628 (e.g., a touch sensor or a press button) is used to receive input selections from a user for manipulating a graphical user interface of the presented image.
[0186] Figure 16 Components of the illustrated head-wearable device 116 are located on one or more circuit boards (e.g., PCBs or flexible PCBs) in the frame or temples. Alternatively or additionally, the depicted components may be located in chunks, frames, hinges, or nose bridges of the head-wearable device 116. The left and right visible light camera devices 1606 may include digital camera elements, such as complementary metal-oxide-semiconductor (CMOS) image sensors, charge-coupled devices, camera lenses, or any other corresponding visible light or light-capturing elements that can be used to capture data, including images of scenes with unknown objects.
[0187] The head-mounted wearable device 116 includes a memory 1602 that stores instructions for performing a subset or all of the functions described herein. The memory 1602 may also include a storage device.
[0188] like Figure 16As shown, the high-speed circuit system 1626 includes a high-speed processor 1630, a memory 1602, and a high-speed wireless circuit system 1632. In some examples, an image display driver 1620 is coupled to the high-speed circuit system 1626 and operated by the high-speed processor 1630 to drive the left and right image displays in the image display 1618 of the optical components. The high-speed processor 1630 can be any processor capable of managing the high-speed communication and operation of any general-purpose computing system required by the head-worn device 116. The high-speed processor 1630 includes the processing resources required to manage high-speed data transmission to a wireless local area network (WLAN) over the high-speed wireless connection 1614 using the high-speed wireless circuit system 1632. In some examples, the high-speed processor 1630 executes the operating system of the head-worn device 116 (e.g., a LINUX operating system) or other such operating system, and this operating system is stored in the memory 1602 for execution. Among other duties, the high-speed processor 1630, which executes the software architecture of the head-worn device 116, manages data transmission with the high-speed wireless circuit system 1632. In some examples, the high-speed wireless circuit system 1632 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, also referred to herein as Wi-Fi®. In some examples, the high-speed wireless circuit system 1632 may implement other high-speed communication standards.
[0189] The low-power wireless circuitry system 1634 and high-speed wireless circuitry system 1632 of the head-mounted wearable device 116 may include a short-range transceiver (Bluetooth™) and a wireless wide-area network transceiver, a wireless local area network transceiver, or a wide-area network transceiver (e.g., cellular or Wi-Fi®). The mobile device 114, including transceivers communicating via the low-power wireless connection 1612 and the high-speed wireless connection 1614, can be implemented using the architectural details of the head-mounted wearable device 116, as can other components of the network 1616.
[0190] Memory 1602 includes any storage device capable of storing various data and applications, including camera data generated by the left and right visible light cameras 1606, the infrared camera 1610, and the image processor 1622, as well as the generated images to be displayed on an image display in an image display 1618 of the optical components via an image display driver 1620. While memory 1602 is shown as integrated with the high-speed circuitry 1626, in some examples, memory 1602 may be a separate, independent component of the head-mounted wearable device 116. In some such examples, electrical wiring may provide a connection from the image processor 1622 or the low-power processor 1636 to memory 1602 via a chip including a high-speed processor 1630. In some examples, the high-speed processor 1630 may manage addressing of memory 1602, such that the low-power processor 1636 will activate the high-speed processor 1630 whenever a read or write operation involving memory 1602 is required.
[0191] like Figure 16 As shown, the low-power processor 1636 or high-speed processor 1630 of the head-mounted wearable device 116 may be coupled to a camera device (visible light camera 1606, infrared emitter 1608 or infrared camera 1610), an image display driver 1620, a user input device 1628 (e.g., a touch sensor or a press button), and a memory 1602.
[0192] The head-mounted wearable device 116 is connected to a host computer. For example, the head-mounted wearable device 116 is paired with the mobile device 114 via a high-speed wireless connection 1614 or connected to the server system 1604 via a network 1616. The server system 1604 may be one or more computing devices as part of a service or network computing system, for example, it includes a processor, memory, and network communication interfaces to communicate with the mobile device 114 and the head-mounted wearable device 116 via the network 1616.
[0193] Mobile device 114 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via network 1616, low-power wireless connection 1612, or high-speed wireless connection 1614.
[0194] The output components of the head-worn wearable device 116 include visual components, such as displays (e.g., liquid crystal displays (LCDs), plasma display panels (PDPs), light-emitting diode (LED) displays, projectors, or waveguides). The image display of the optical components is driven by an image display driver 1620. The output components of the head-worn wearable device 116 also include acoustic components (e.g., speakers), haptic components (e.g., vibration motors), other signal generators, etc. The input components (e.g., user input devices 1628) of the head-worn wearable device 116, mobile device 114, and server system 1604 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens that provide the position and force of touch or touch gestures, or other haptic input components), audio input components (e.g., microphones), etc.
[0195] The head-mounted wearable device 116 may also include additional peripheral device elements. Such peripheral device elements may include biometric sensors, additional sensors, or display elements integrated with the head-mounted wearable device 116. For example, peripheral device elements may include any I / O components, including output components, motion components, position components, or any other such components described herein.
[0196] For example, biometric components include those for detecting expressions (e.g., hand gestures, facial expressions, vocal expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Biometric components may include brain-computer interface (BMI) systems that allow communication between the brain and external devices or machines. This can be achieved by recording brain activity data, converting that data into a format that can be understood by a computer, and then using the resulting signals to control the device or machine.
[0197] Examples of BMI technology types include:
[0198] BMI based on electroencephalography (EEG) uses electrodes placed on the scalp to record electrical activity in the brain.
[0199] Invasive BMI uses electrodes surgically implanted in the brain.
[0200] Optogenetics BMI uses light to control the activity of specific nerve cells in the brain.
[0201] Any biometric data collected by the biometric component is captured and stored only with user approval and is deleted upon user request. Furthermore, such biometric data may be used for very limited purposes (e.g., identification and verification). To ensure the restricted and authorized use of biometric information and other personally identifiable information (PII), access to this data is limited to authorized personnel (if access to the data occurs). Any use of biometric data may be strictly limited to identification and verification purposes, and such biometric data is not shared or sold to any third party without the user's explicit consent. In addition, appropriate technical and organizational measures are implemented to ensure the security and confidentiality of this sensitive information.
[0202] Motion components include accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Position components include positioning sensor components (e.g., GPS receiver components) for generating position coordinates, Wi-Fi or Bluetooth™ transceivers for generating positioning system coordinates, altitude sensor components (e.g., altimeters or barometers that detect air pressure, from which altitude can be obtained), orientation sensor components (e.g., magnetometers), etc. Such positioning system coordinates can also be received from mobile device 114 via low-power wireless circuit system 1634 or high-speed wireless circuit system 1632 through low-power wireless connection 1612 and high-speed wireless connection 1614.
[0203] Machine architecture
[0204] Figure 17This is a schematic representation of machine 1700, within which instructions 1702 (e.g., software, program, application, app, or other executable code) can be executed to cause machine 1700 to perform any or more of the methods discussed herein. For example, instructions 1702 can cause machine 1700 to perform any or more of the methods described herein. Instructions 1702 transform the general, unprogrammed machine 1700 into a specific machine 1700 programmed to perform the described and illustrated functions in the described manner. Machine 1700 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 1700 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1700 may include, but is not limited to, server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing instructions 1702 specifying actions to be taken by machine 1700. Furthermore, although only a single machine 1700 is shown, the term "machine" should also be considered as a collection of machines that individually or jointly execute instructions 1702 to perform any or more of the methods discussed herein. For example, machine 1700 may include user system 102 or any of a plurality of server devices forming part of interactive server system 110. In some examples, machine 1700 may also include both client and server systems, wherein certain operations of a particular method or algorithm are performed on the server side and certain operations of said particular method or algorithm are performed on the client side.
[0205] Machine 1700 may include a processor 1704, a memory 1706, and an input / output (I / O) unit 1708 that can be configured to communicate with each other via a bus 1710. In the example, processor 1704 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processors 1712 and 1714 that execute instruction 1702. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Figure 17 Multiple processors 1704 are shown, but machine 1700 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0206] Memory 1706 includes main memory 1716, static memory 1718, and memory cell 1720, all of which are accessible by processor 1704 via bus 1710. Main memory 1706, static memory 1718, and memory cell 1720 store instructions 1702 that implement any one or more of the methods or functions described herein. Instructions 1702 may also reside wholly or partially in main memory 1716, static memory 1718, machine-readable medium 1722 within memory cell 1720, at least one processor of processor 1704 (e.g., within the processor's cache memory), or any suitable combination thereof during execution by machine 1700.
[0207] I / O component 1708 may include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurement results, etc. The specific I / O component 1708 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine is unlikely to include such a touch input device. It should be recognized that I / O component 1708 may include... Figure 17Many other components are not shown. In various examples, I / O component 1708 may include user output component 1724 and user input component 1726. User output component 1724 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tube (CRT) displays), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. User input component 1726 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens or other haptic input components that provide the position and force of a touch or touch gesture), audio input components (e.g., microphones), etc.
[0208] In other examples, I / O component 1708 may include biometric component 1728, motion component 1730, environmental component 1732, or position component 1734, as well as a wide range of other components. For example, biometric component 1728 includes components for detecting expressions (e.g., hand expressions, facial expressions, vocal expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and recognizing a person (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Biometric components may include brain-computer interface (BMI) systems that allow communication between the brain and external devices or machines. This can be achieved by recording brain activity data, converting that data into a format that can be understood by a computer, and then using the resulting signals to control devices or machines.
[0209] Any biometric data collected by the biometric component is captured and stored only with user approval and is deleted upon user request. Furthermore, such biometric data may be used for very limited purposes (e.g., identification and verification). To ensure the restricted and authorized use of biometric information and other personally identifiable information (PII), access to this data is limited to authorized personnel (if access to the data occurs). Any use of biometric data may be strictly limited to identification and verification purposes, and such biometric data is not shared or sold to any third party without the user's explicit consent. In addition, appropriate technical and organizational measures are implemented to ensure the security and confidentiality of this sensitive information.
[0210] The moving part 1730 includes an acceleration sensor part (e.g., an accelerometer), a gravity sensor part, and a rotation sensor part (e.g., a gyroscope).
[0211] The environmental component 1732 includes, for example, one or more camera devices (with still image / photograph and video capabilities), lighting sensor components (e.g., a photometer), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., a barometer), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., an infrared sensor for detecting nearby objects), gas sensors (e.g., gas detection sensors for detecting the concentration of hazardous gases for safety purposes or for measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment.
[0212] Regarding the camera device, user system 102 may have a camera device system including, for example, a front-facing camera on the front surface of user system 102 and a rear-facing camera on the rear surface of user system 102. The front-facing camera may be used, for example, to capture still images and videos (e.g., "selfies") of the user of user system 102, which can then be enhanced with the aforementioned enhancement data (e.g., filters). For example, the rear-facing camera may be used to capture still images and videos in a more conventional camera device mode, which are similarly enhanced with the enhancement data. In addition to the front-facing and rear-facing cameras, user system 102 may also include a 360° camera for capturing 360° photos and videos.
[0213] Furthermore, the camera system of user system 102 may include dual rear cameras (e.g., a main camera and a depth-sensing camera), or even triple, quadruple, or quintuple rear camera configurations on the front and rear sides of user system 102. For example, these multi-camera systems may include wide-angle cameras, ultra-wide-angle cameras, telephoto cameras, macro cameras, and depth sensors.
[0214] The position component 1734 includes a positioning sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure and from which altitude can be obtained), an orientation sensor component (e.g., a magnetometer), and the like.
[0215] Various technologies can be used to implement communication. I / O component 1708 also includes a communication component 1736 operable to couple machine 1700 to network 1738 or device 1740 via a suitable coupling or connection. For example, communication component 1736 may include a network interface component or another suitable device that interfaces with network 1738. In further examples, communication component 1736 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components for providing communication via other modalities. Device 1740 may be another machine or any peripheral device from a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0216] Furthermore, the communication component 1736 can detect identifiers, or include components operable to detect identifiers. For example, the communication component 1736 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph™, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying audio signals from the tag). Additionally, various information can be obtained via the communication component 1736, such as location obtained via Internet Protocol (IP) geolocation, location obtained via Wi-Fi® signal triangulation, location obtained by detecting NFC beacon signals that can indicate a specific location, etc.
[0217] Various memories (e.g., main memory 1716, static memory 1718, and the memory of processor 1704) and storage unit 1720 may store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 1702) cause various operations to implement the disclosed examples when executed by processor 1704.
[0218] Instructions 1702 can be sent or received over network 1738 via a transmission medium using a network interface device (e.g., a network interface component included in communication component 1736) and using any of several known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 1702 can be sent or received via a transmission medium coupled to device 1740 (e.g., peer-to-peer coupling).
[0219] Software Architecture
[0220] Figure 18 This is a block diagram 1800 illustrating a software architecture 1802 that can be installed on any one or more of the devices described herein. The software architecture 1802 is supported by hardware such as a machine 1804 including a processor 1806, memory 1808, and I / O components 1810. In this example, the software architecture 1802 can be conceptualized as a stack of layers, where each layer provides a specific function. The software architecture 1802 includes layers such as an operating system 1812, libraries 1814, frameworks 1816, and applications 1818. Operationally, application 1818 activates API calls 1820 via the software stack and receives messages 1822 in response to API calls 1820.
[0221] Operating system 1812 manages hardware resources and provides public services. Operating system 1812 includes, for example, a kernel 1824, services 1826, and drivers 1828. Kernel 1824 acts as an abstraction layer between the hardware layer and other software layers. For example, kernel 1824 provides memory management, processor management (e.g., scheduling), component management, networking and security settings, and other functions. Services 1826 can provide other public services to other software layers. Drivers 1828 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 1828 may include display drivers, camera drivers, BLUETOOTH® or BLUETOOTH® low-power drivers, flash memory drivers, serial communication drivers (e.g., USB drivers), Wi-Fi® drivers, audio drivers, power management drivers, etc.
[0222] Library 1814 provides common low-level infrastructure used by application 1818. Library 1814 may include system library 1830 (e.g., the C standard library), which provides functions such as memory allocation, string manipulation, and mathematical functions. Additionally, library 1814 may include API library 1832, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Picture Experts Group (JPEG or JPG), or Portable Web Graphics (PNG)), graphics libraries (e.g., the OpenGL framework for rendering graphic content on a display in two-dimensional (2D) and three-dimensional (3D) formats), database libraries (e.g., SQLite, which provides various relational database functions), web libraries (e.g., WebKit, which provides web browsing capabilities), and so on. Library 1814 may also include various other libraries 1834 to provide many other APIs to application 1818.
[0223] Framework 1816 provides common high-level infrastructure for use by application 1818. For example, Framework 1816 provides various graphical user interface (GUI) functions, advanced resource management, and advanced location services. Framework 1816 can provide a wide range of other APIs that can be used by application 1818, some of which may be specific to a particular operating system or platform.
[0224] In the example, application 1818 may include home application 1836, contact application 1838, browser application 1840, book reader application 1842, location application 1844, media application 1846, messaging application 1848, game application 1850, and a wide variety of other applications such as third-party application 1852. Application 1818 is a program that performs the functions defined in the program. One or more applications 1818 can be created using various programming languages, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a particular example, third-party application 1852 (e.g., an application developed by an entity other than a platform vendor using the Android™ or iOS™ Software Development Kit (SDK)) may be mobile software that runs on a mobile operating system such as iOS™, Android™, Windows® phones, or other mobile operating systems. In this example, a third-party application 1852 can activate API call 1820 provided by the operating system 1812 to facilitate the functionality described herein.
[0225] Example
[0226] In view of the above-mentioned implementation methods, this application discloses the following list of examples, wherein a feature of a single example or more than one feature of an example is combined together, and optionally, is combined with one or more features of one or more other examples, which are also further examples falling within the disclosure of this application.
[0227] Example 1 is a system comprising: at least one processor; at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations including: accessing an input video item comprising a target visual enhancement; generating an embedding of the input video item by a machine learning model trained in an unsupervised training phase to minimize a loss between training video representations generated in each of a plurality of training sets, each training set comprising a plurality of different training video items, each training video item comprising the same predefined visual enhancement; and mapping the target visual enhancement to an enhancement identifier based on the generation of the embedding of the input video item.
[0228] In Example 2, the subject of Example 1 includes, wherein the embedding of the input video item includes a vector representation of the visual effect of the target visual enhancement within the input video item.
[0229] In Example 3, the subject of any one of Examples 1 to 2 includes, wherein the unsupervised training phase includes self-supervised contrastive learning of only positive samples.
[0230] In Example 4, the subject of any one of Examples 1 to 3 includes, wherein each training set includes a pair of training video items, and wherein the self-supervised training phase includes: accessing a first training video item in a first pair of training video items, the first training video item including first video content with a first predefined visual enhancement applied; accessing a second training video item in the first pair of training video items, the second training video item including second video content with the first predefined visual enhancement applied; generating vector representations of the first training video item and vector representations of the second training video item; and generating a transformed representation based on the vector representation of the first training video item, wherein minimizing the loss between the training video representations includes: for the first pair of training video items, using a loss function to measure the similarity between the transformed representation and the vector representation of the second training video item.
[0231] In Example 5, the subject of Example 4 includes, wherein the self-supervised training phase further includes: automatically updating the parameters of the machine learning model to minimize the loss between training video representations within subsequent training video item pairs based on the loss function.
[0232] In Example 6, the subject of any one of Examples 4 to 5 includes, wherein the first video content is a first base video and the second video content is a second base video, and the first base video is different from the second base video.
[0233] In Example 7, the subject of Example 6 includes the operation further comprising: applying the first predefined visual enhancement to the first base video by the enhancement rendering component; and applying the first predefined visual enhancement to the second base video by the enhancement rendering component.
[0234] In Example 8, the subject of any one of Examples 4 to 7 includes, wherein generating the transformed representation includes: transforming the vector representation of the first training video item into a transformed representation by the predictor component of the machine learning model to create a final prediction target for the machine learning model.
[0235] In Example 9, the subject of any one of Examples 1 to 8 includes, wherein the input video item is a first input video item and the embedding of the first input video item is a first embedding, and the operation further includes: accessing a second input video item including the target visual enhancement; generating a second embedding of the second input video item by the machine learning model; and generating an aggregated embedding based on the first embedding and the second embedding, the aggregated embedding representing the visual effect of the target visual enhancement.
[0236] In Example 10, the subject of Example 9 includes, wherein generating the aggregated embedding includes determining the mean of the first embedding and the second embedding.
[0237] In Example 11, the subject of any of Examples 9 to 10 includes, wherein the enhanced identifier is the aggregated embedding.
[0238] In Example 12, the subject of any one of Examples 1 to 11 includes the operation further comprising: automatically comparing the embedding of the enhancement identifier with that of a reference visual enhancement to determine the similarity between the target visual enhancement and the reference visual enhancement.
[0239] In Example 13, the subject of any one of Examples 1 to 12 includes, wherein the machine learning model is a first machine learning model, and the operation further includes: sending the augmented identifier to a second machine learning model, the second machine learning model including at least one of the following: an augmented ranking model, an augmented retrieval model, an augmented deduplication model, an augmented labeling model, or an augmented clustering model.
[0240] In Example 14, the subject of any of Examples 1 to 13 includes, and the operation further includes: querying a database to identify the target visual enhancement based on the enhancement identifier.
[0241] In Example 15, the subject of any one of Examples 1 to 14 includes the operation further comprising: receiving the input video item from a user device, wherein the target visual enhancement has been applied to the input video item by a user of the user device.
[0242] In Example 16, the subject of any of Examples 1 to 15 includes, wherein the target visual enhancement is an augmented reality effect.
[0243] In Example 17, the subject of any of the items in Example 16 includes, wherein the augmented reality effect is a video filter.
[0244] In Example 18, the subject of any of Examples 1 to 17 includes, wherein the input video item is in binary file format.
[0245] Example 19 is a method comprising: accessing an input video item including a target visual enhancement; generating an embedding of the input video item by a machine learning model trained in an unsupervised training phase to minimize a loss between training video representations generated in each of a plurality of training sets, each training set including a plurality of different training video items, each training video item including the same predefined visual enhancement; and mapping the target visual enhancement to an enhancement identifier based on the generation of the embedding of the input video item.
[0246] Example 20 is a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations including: accessing an input video item comprising a target visual enhancement; generating an embedding of the input video item by a machine learning model trained in an unsupervised training phase to minimize a loss between training video representations generated in each of a plurality of training sets, each training set comprising a plurality of distinct training video items, each training video item comprising the same predefined visual enhancement; and mapping the target visual enhancement to an enhancement identifier based on the generation of the embedding of the input video item.
[0247] Example 21 is at least one machine-readable medium including instructions that, when executed by a processing circuitry system, cause the processing circuitry system to perform operations for implementing any one of Examples 1 to 20.
[0248] Example 22 is an apparatus that includes means for implementing any one of Examples 1 to 20.
[0249] Example 23 is a system for implementing any one of Examples 1 through 20.
[0250] Example 24 is a method for implementing any one of Examples 1 through 20.
[0251] in conclusion
[0252] Examples of this disclosure provide techniques for training machine learning models to generate vector representations of visual enhancements applied to content items, deploying such machine learning models, and mapping vector representations to their corresponding visual enhancements.
[0253] Examples of this disclosure can allow for enhanced understanding of how enhancements or user-applied enhancements are implemented, thereby improving the ability to personalize content or enhancement options presented to users of interactive systems. In this way, the quality or diversity of content can be improved, recommendations can be strengthened, and ultimately, users can be empowered to express themselves in more diverse and creative ways using technological tools.
[0254] While the examples in this disclosure describe embeddings generated based on video items, it should be understood that the techniques and systems described herein can also be applied to other types of content items, such as images (e.g., still images or single frames). Furthermore, while the examples in this disclosure describe embeddings that capture or represent enhanced visual effects, it should be understood that the techniques and systems described herein can also be applied to other types of embeddings, such as embeddings designed to capture or represent audio-related features of content items.
[0255] As used in this disclosure, the term "machine learning model" (or simply "model") can refer to a single, independent model or a combination of models. The term can also refer to a system, component, or module that includes a machine learning model and one or more supporting or supplementary components that do not necessarily perform machine learning tasks.
[0256] As used in this disclosure, phrases of the form "at least one of A, B, or C", "at least one of A, B, or C", "at least one of A, B, and C", etc., should be interpreted as selecting at least one from the group including "A, B, and C". In this disclosure, unless explicitly stated otherwise in conjunction with specific examples, this phrasing does not imply "at least one of A, at least one of B, and at least one of C". As used in this disclosure, the example "at least one of A, B, or C" will cover any of the following selections: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, and {A, B, C}.
[0257] Unless the context explicitly requires it, throughout the specification and claims, the words “comprising,” “including,” etc., shall be interpreted in an inclusive sense rather than an exclusive or exhaustive sense; that is, meaning “including but not limited to.” As used herein, the terms “connection,” “coupled,” or any variation thereof mean any direct or indirect connection or coupling between two or more elements; the coupling or connection between elements may be physical, logical, or a combination thereof. Furthermore, when used in this application, the words “in this document,” “above,” “below,” and words with similar meanings refer to the entire application and not any particular part of the application. Where the context permits, the use of singular or plural terms may also include the plural or singular, respectively. When referring to a list of two or more items, the word “or” covers all the following interpretations of the word: any item in the list, all items in the list, and any combination of items in the list. Similarly, with respect to a list of two or more items, the word “and / or” covers all the following interpretations of the word: any item in the list, all items in the list, and any combination of items in the list.
[0258] While some examples (such as those depicted in the accompanying figures) include a specific order of operations, that order may be changed without departing from the scope of this disclosure. For example, some operations in the depicted operations may be performed in parallel or in a different order that does not substantially affect the functionality described in the examples. In other examples, different components of an example device or system implementing the example methods may perform their functions substantially simultaneously or in a specific order.
[0259] Glossary
[0260] For example, "carrier signal" refers to any intangible medium or other intangible medium capable of storing, encoding, or carrying instructions executed by a machine and including digital or analog communication signals. Instructions can be sent or received over a network using a transmission medium via a network interface device.
[0261] For example, "client device" refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. Client devices can be, but are not limited to, mobile phones, desktop computers, laptop computers, portable digital assistants (PDAs), smartphones, tablet computers, ultrabooks, netbooks, laptop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, set-top boxes, or any other communication device that a user can use to access the network.
[0262] For example, "communication network" refers to one or more parts of a network, which can be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a Common Old-Style Telephone Service (POTS) network, a cellular telephone network, a wireless network, a Wi-Fi® network, other types of networks, or a combination of two or more such networks. For example, a network or part of a network may include a wireless network or a cellular network, and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or other types of cellular or wireless coupling. In this example, coupling can enable any data transmission technology of various types, such as single-carrier radio transmission technology (1xRTT), evolved data optimization (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rate GSM evolution (EDGE) technology, the 3rd Generation Partnership Project (3GPP) including 3G, fourth-generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Global Microwave Access Interoperability (WiMAX), Long Term Evolution (LTE) standards, other data transmission technologies defined by various standards setting organizations, other long-distance protocols, or other data transmission technologies.
[0263] For example, a “component” refers to a logical or physical entity having boundaries defined by functional or subroutine calls, branch points, APIs, or other technologies that provide partitioning or modularity for a particular processing or control function. A component can be combined with other components via its interface to perform machine processing. A component can be an encapsulated functional hardware unit designed for use with other components, and part of a program that typically performs a related function. A component can constitute a software component (e.g., code implemented on a machine-readable medium) or a hardware component. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various examples, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components (e.g., processors or processor groups) of a computer system can be configured by software (e.g., an application or application portion) to operate to perform certain operations as described herein. Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component can include a dedicated circuit system or logic permanently configured to perform certain operations. Hardware components can be dedicated processors, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). Hardware components can also include programmable logic or circuit systems temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a particular machine (or a specific part of a machine), which is uniquely tailored to perform the configured function and is no longer a general-purpose processor. It will be appreciated that a decision can be made, for cost and time considerations, whether to implement a hardware component mechanically in a dedicated and permanently configured circuit system or in a temporarily configured (e.g., configured by software) circuit system. Therefore, the phrase “hardware component” (or “hardware-implemented component”) should be understood to include tangible entities, i.e., entities physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain way or perform certain operations described herein. Consider the example of a hardware component being temporarily configured (e.g., programmed), without needing to configure or instantiate each hardware component at any given time. For example, in cases where the hardware components include a general-purpose processor that is configured as a dedicated processor via software, this general-purpose processor can be configured as different dedicated processors (e.g., including different hardware components) at different times. The software accordingly configures one or more specific processors to constitute a specific hardware component at one time and different hardware components at different times. The hardware components can provide information to and receive information from other hardware components.Therefore, the described hardware components can be considered communicatively coupled. In the presence of multiple hardware components, communication can be achieved through signal transmission between or among two or more hardware components (e.g., via appropriate circuitry and buses). In examples where multiple hardware components are configured or instantiated at different times, such communication between hardware components can be achieved, for example, by storing information in a memory structure accessible to the multiple hardware components and retrieving information from the memory structure. For example, a hardware component can perform an operation and store the output of that operation in a memory device communicatively coupled to it. Another hardware component can then access the memory device at a subsequent time to retrieve and process the stored output. Hardware components can also initiate communication with input or output devices and can operate on resources (e.g., collections of information). The various operations of the example methods described herein can be performed, at least in part, by one or more processors configured, either temporarily (e.g., by software) or permanently, to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented components that operate to perform one or more operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be implemented at least in part by processors, where a particular processor or one or more processors are examples of hardware. For example, at least some operations of the methods can be performed by one or more processors or processor-implemented components. Furthermore, one or more processors can also operate to support the execution of related operations in a “cloud computing” environment or as a “Software as a Service” (SaaS) operation. For example, at least some operations can be performed by a group of computers (as an example of a machine including processors), where these operations are accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of some operations can be distributed among processors, not residing within a single machine, but deployed across multiple machines. In some examples, the processor or processor-implemented component may reside in a single geographic location (e.g., within a home environment, office environment, or server cluster). In other examples, the processor or processor-implemented component may be distributed across multiple geographic locations.
[0264] For example, "computer-readable storage medium" refers to both machine-readable storage media and transmission media. Therefore, these terms include both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" refer to the same thing and can be used interchangeably in this disclosure.
[0265] For example, "machine storage medium" refers to one or more storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Therefore, this term should be considered to include, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGAs, and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium," "device storage medium," and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage medium," "computer storage medium," and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium."
[0266] For example, a "non-transitory computer-readable storage medium" refers to a tangible medium capable of storing, encoding, or carrying instructions that can be executed by a machine.
[0267] For example, "signal medium" refers to any intangible medium capable of storing, encoding, or carrying instructions executable by a machine, and includes digital or analog communication signals or other intangible media that facilitate the communication of software or data. The term "signal medium" should be considered to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal whose characteristics are set or altered in a manner that encodes information in the signal. The terms "transmission medium" and "signal medium" refer to the same thing and may be used interchangeably in this disclosure.
[0268] "User equipment" means, for example, a device that is accessed, controlled, or owned by a user and that the user interacts with to perform actions or interact with other users or computer systems.
Claims
1. A system comprising: At least one processor; as well as At least one memory component storing instructions, which, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including: Access includes input video items that include target visual enhancement; The embeddings of the input video items are generated by a machine learning model trained in an unsupervised training phase to minimize the loss between training video representations generated in each of multiple training sets, each training set comprising multiple distinct training video items, and each training video item comprising the same predefined visual augmentations; and Based on the generation of the embedding of the input video item, the target visual enhancement is mapped to an enhancement identifier.
2. The system according to claim 1, wherein, The embedding of the input video item includes a vector representation of the visual effect of the target visual enhancement within the input video item.
3. The system according to claim 1, wherein, The unsupervised training phase includes self-supervised contrastive learning with only positive samples.
4. The system according to claim 1, wherein, Each training set includes a pair of training video items, and wherein the unsupervised training phase includes: Access the first training video item in the first pair of training video items, the first training video item including first video content with the first predefined visual enhancement applied; Access the second training video item in the first pair of training video items, the second training video item including second video content with the first predefined visual enhancement applied; Generate vector representations of the first training video item and vector representations of the second training video item; and Generating a transformed representation based on the vector representation of the first training video item, wherein minimizing the loss between the training video representations includes: for the first pair of training video items, using a loss function to measure the similarity between the transformed representation and the vector representation of the second training video item.
5. The system according to claim 4, wherein, The unsupervised training phase also includes: The parameters of the machine learning model are automatically updated to minimize the loss between training video representations within subsequent training video item pairs based on the loss function.
6. The system according to claim 4, wherein, The first video content is a first base video, and the second video content is a second base video, wherein the first base video is different from the second base video.
7. The system according to claim 6, wherein the operation further comprises: The enhanced rendering component applies the first predefined visual enhancement to the first base video; as well as The enhanced rendering component applies the first predefined visual enhancement to the second base video.
8. The system according to claim 4, wherein, Generating the transformed representation includes: transforming the vector representation of the first training video item into the transformed representation by the predictor component of the machine learning model to create a final prediction target for the machine learning model.
9. The system according to claim 1, wherein, The input video item is a first input video item, and the embedding of the first input video item is a first embedding. The operation further includes: Access the second input video item that includes the visual enhancement of the target; The machine learning model generates a second embedding of the second input video item; and An aggregated embedding is generated based on the first embedding and the second embedding, the aggregated embedding representing the visual effect of the target visual enhancement.
10. The system according to claim 9, wherein, Generating the aggregated embeddings includes determining the mean of the first embedding and the second embedding.
11. The system according to claim 9, wherein, The enhanced identifier is the aggregated embedding.
12. The system according to claim 1, wherein the operation further comprises: The embedding of the enhancement identifier is automatically compared with that of a reference visual enhancement to determine the similarity between the target visual enhancement and the reference visual enhancement.
13. The system according to claim 1, wherein, The machine learning model is a first machine learning model, and the operation further includes: The enhanced identifier is sent to a second machine learning model, which includes at least one of the following: an enhanced ranking model, an enhanced retrieval model, an enhanced deduplication model, an enhanced labeling model, or an enhanced clustering model.
14. The system according to claim 1, wherein the operation further comprises: The database is queried to identify the target visual enhancement based on the enhancement identifier.
15. The system according to claim 1, wherein the operation further comprises: The user receives the input video item from the user equipment, and the target visual enhancement has been applied to the input video item by the user of the user equipment.
16. The system according to claim 1, wherein, The target visual enhancement mentioned above is an augmented reality effect.
17. The system according to claim 16, wherein, The augmented reality effect is a video filter.
18. The system according to claim 1, wherein, The input video item is in binary file format.
19. A method comprising: Access includes input video items that include target visual enhancement; The embeddings of the input video items are generated by a machine learning model, which is trained in an unsupervised training phase to minimize the loss between the training video representations generated in each of multiple training sets, each training set including multiple different training video items, each training video item including the same predefined visual augmentations; as well as Based on the generation of the embedding of the input video item, the target visual enhancement is mapped to an enhancement identifier.
20. A non-transitory computer-readable storage medium storing instructions, said instructions, when executed by at least one processor, causing said at least one processor to perform an operation, said operation comprising: Access includes input video items that include target visual enhancement; The embeddings of the input video items are generated by a machine learning model, which is trained in an unsupervised training phase to minimize the loss between the training video representations generated in each of multiple training sets, each training set including multiple different training video items, each training video item including the same predefined visual augmentations; as well as Based on the generation of the embedding of the input video item, the target visual enhancement is mapped to an enhancement identifier.