Visual style transformation using machine learning
The method uses Gram matrix pairs and convolutional neural networks to automatically extract and apply distinctive style features, addressing the limitations of existing 3D stylization methods by efficiently transforming 3D objects to match target styles, enhancing virtual experience design.
Patent Information
- Application Number
- JP2025074600
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-29
- Filing Date
- 2025-04-28
- Publication Date
- 2025-11-11
AI Technical Summary
Existing methods for 3D stylization fail to effectively modify the style of a source object to mimic a target scene or object, as they rely on handcrafted shape descriptors that lack generalization and fail to capture important style features due to complex view selection in similarity metrics.
A method using a loss function based on Gram matrix pairs to identify distinctive stylistic features, applying convolutional neural networks to modify visual assets to match a target style by optimizing convolution filter weights, enabling automatic extraction of discriminative style characteristics.
This approach simplifies the process of creating visually appealing virtual experiences by automatically identifying and manipulating 3D objects to match a desired aesthetic, reducing time and effort in designing aesthetically coherent scenes.
Smart Images

Figure 2025168668000001_ABST
Abstract
Description
[Technical Field]
[0001] Embodiments relate generally to online virtual experience platforms, and more particularly to methods, systems, and computer-readable media for identifying distinctive style characteristics of visual assets and performing visual style transformations to match the visual assets to a target style. [Background technology]
[0002] Online platforms, such as virtual experience platforms and online gaming platforms, allow users to create virtual experiences or games in a particular style, which may be applied to visual assets within the virtual experience or game.
[0003] Prior art techniques for 3D stylization are designed to address specific design styles or to use exemplary shapes where the visual style of a target can be encoded by their surface normals. However, neither of these address the general stylization problem where a user attempts to modify the source of a 3D object to mimic the style of a target scene or object. To solve this problem, we use a method that can extract a set of discriminative features that distinguish the visual styles across source and target objects.
[0004] Existing methods for identifying similarity metrics used to evaluate style similarity between various shapes often rely on handcrafted sets of shape descriptors that do not guarantee generalization across various design styles. Finding a weighted balance among such descriptors is complex, and some methods use deep learning models to jointly infer a set of shape descriptors and their balancing weights. This method calculates style similarity in image space. Unfortunately, preselecting a set of rendering views is complex and directly impacts the final metric, as important style features may not be captured from the selected views. As a result, existing methods for estimating style similarity fail to provide techniques for extracting a set of style-discriminatory features.
[0005] The background description provided herein is intended to provide a context for the present disclosure. The work of the inventors named herein, to the extent that it is described in this Background section, and aspects of the description that may not otherwise be considered prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure. Summary of the Invention [Problem to be solved by the invention]
[0006] Aspects of the present disclosure are directed to methods, systems, and computer-readable media for identifying distinctive style characteristics of an input shape for application to a virtual experience or asset. [Means for solving the problem]
[0007] In one aspect of the present disclosure, a computer-implemented method is provided. The method may include receiving, by a processor, a visual asset. The method may include receiving, by the processor, an indicator of a visual style. The method may include identifying, by the processor, a set of distinctive stylistic features for the visual style using a loss function calculated based on a first sum of differences between a first set of Gram matrix pairs and a second sum of differences between a second set of Gram matrix pairs. The first set of Gram matrix pairs may be associated with a set of positive input samples of the same visual style, and the second set of Gram matrix pairs may be associated with a set of negative input samples of a different visual style. The method may include modifying, by the processor, a shape of the visual asset to match the identified set of stylistic features distinctive to the visual style. After modification, the method may include rendering, by the processor, the visual asset to match the set of distinctive stylistic features.
[0008] In some implementations, the method may include receiving, by a processor, a set of positive input samples of the same visual style and a set of negative input samples of a different visual style. In some implementations, the method may include generating, by the processor, a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples. In some implementations, the method may include obtaining, by the processor, a first set of Gram matrix pairs associated with the set of positive input samples based on the first set of feature maps and a second set of Gram matrix pairs associated with the set of negative input samples based on the second set of feature maps. In some implementations, the method may include calculating, by the processor, a loss function based on a first summation of differences of the first set of Gram matrix pairs and a second summation of differences of the second set of Gram matrix pairs.
[0009] In some implementations, a set of positive input samples associated with the same visual style may include a first set of shape pairs of the same style, and in some implementations, a set of negative input samples associated with a different visual style may include a second set of shape pairs of a different style.
[0010] In some implementations, one shape in the second set of shape pairs may have the same style as the style of the first set of shape pairs.
[0011] In some implementations, generating a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples may include applying, by the processor, a set of convolution filters to the set of positive input samples to generate the first set of feature maps. In some implementations, generating a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples may include applying, by the processor, a set of convolution filters to the set of negative input samples to generate the second set of feature maps.
[0012] In some implementations, the set of convolution filters may include kernels of various sizes.
[0013] In some implementations, obtaining a first set of Gram matrix pairs associated with a set of positive input samples based on a first set of feature maps and a second set of Gram matrix pairs associated with a set of negative input samples based on a second set of feature maps may include determining, by a processor, a first correlation between the first set of feature maps associated with the set of positive input samples for obtaining the first set of Gram matrix pairs. In some implementations, obtaining a first set of Gram matrix pairs associated with a set of positive input samples based on a first set of feature maps and a second set of Gram matrix pairs associated with a set of negative input samples based on a second set of feature maps may include determining, by a processor, a second correlation between the second set of feature maps associated with the set of negative input samples for obtaining the second set of Gram matrix pairs.
[0014] According to another aspect of the present disclosure, a computing device is provided. The computing device may include a processor and a memory coupled to the processor, and instructions stored therein, when executed by the processor, cause the processor to perform operations. The operations may include receiving visual assets by the processor. The operations may include receiving an indicator of the visual assets by the processor. The operations may include identifying a set of distinctive stylistic features for the visual style by the processor using a loss function calculated based on a first sum of differences between a first set of Gram matrix pairs and a second sum of differences between a second set of Gram matrix pairs. The first set of Gram matrix pairs may be associated with a set of positive input samples of the same visual style, and the second set of Gram matrix pairs may be associated with a set of negative input samples of a different visual style. The operations may include modifying the visual assets to match the identified set of distinctive stylistic features for the visual style by the processor. After modification, the operations may include rendering the visual assets to match the set of distinctive stylistic features by the processor.
[0015] In some implementations, the operations may include receiving, by the processor, a set of positive input samples of the same visual style and a set of negative input samples of a different visual style. In some implementations, the operations may include generating, by the processor, a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples. In some implementations, the operations may include obtaining, by the processor, a first set of Gram matrix pairs associated with the set of positive input samples based on the first set of feature maps and a second set of Gram matrix pairs associated with the set of negative input samples based on the second set of feature maps. In some implementations, the operations may include calculating, by the processor, a loss function based on a first summation of differences of the first set of Gram matrix pairs and a second summation of differences of the second set of Gram matrix pairs.
[0016] In some implementations, a set of positive input samples associated with the same visual style may include a first set of shape pairs of the same style, and in some implementations, a set of negative input samples associated with a different visual style may include a second set of shape pairs of a different style.
[0017] In some implementations, one shape in the second set of shape pairs may have the same style as the style of the first set of shape pairs.
[0018] In some implementations, generating a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples may include applying, by a processor, a set of convolutional filters to the set of positive input samples to generate the first set of feature maps. In some implementations, generating a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples may include applying, by a processor, a set of convolutional filters to the set of negative input samples to generate the second set of feature maps.
[0019] In some implementations, the set of convolution filters may include kernels of various sizes.
[0020] In some implementations, obtaining a first set of Gram matrix pairs associated with a set of positive input samples based on a first set of feature maps and a second set of Gram matrix pairs associated with a set of negative input samples based on a second set of feature maps may include determining a first correlation between the first set of feature maps associated with the set of positive input samples for obtaining the first set of Gram matrix pairs by the processor. In some implementations, obtaining a first set of Gram matrix pairs associated with a set of positive input samples based on the first set of feature maps and a second set of Gram matrix pairs associated with a set of negative input samples based on the second set of feature maps may include determining a second correlation between the second set of feature maps associated with the set of negative input samples for obtaining the second set of Gram matrix pairs by the processor.
[0021] According to a further aspect of the present disclosure, a non-transitory computer-readable medium having instructions stored thereon is provided. When executed by a processor, the instructions cause the processor to perform operations. The operations may include receiving a visual asset by the processor. The operations may include receiving an indicator of the visual asset by the processor. The operations may include identifying a set of distinctive stylistic features for the visual style by the processor using a loss function calculated based on a first sum of differences between a first set of Gram matrix pairs and a second sum of differences between a second set of Gram matrix pairs. The first set of Gram matrix pairs may be associated with a set of positive input samples of the same visual style, and the second set of Gram matrix pairs may be associated with a set of negative input samples of a different visual style. The operations may include modifying the visual asset to match the identified set of distinctive stylistic features for the visual style by the processor. After modification, the operations may include rendering the visual asset to match the set of distinctive stylistic features by the processor.
[0022] In some implementations, the operations may include receiving, by the processor, a set of positive input samples of the same visual style and a set of negative input samples of a different visual style. In some implementations, the operations may include generating, by the processor, a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples. In some implementations, the operations may include obtaining, by the processor, a first set of Gram matrix pairs associated with the set of positive input samples based on the first set of feature maps and a second set of Gram matrix pairs associated with the set of negative input samples based on the second set of feature maps. In some implementations, the operations may include calculating, by the processor, a loss function based on a first summation of differences of the first set of Gram matrix pairs and a second summation of differences of the second set of Gram matrix pairs.
[0023] In some implementations, a set of positive input samples associated with the same visual style may include a first set of shape pairs of the same style, and in some implementations, a set of negative input samples associated with a different visual style may include a second set of shape pairs of a different style.
[0024] In some implementations, one shape in the second set of shape pairs may have the same style as the style of the first set of shape pairs.
[0025] In some implementations, generating a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples may include applying, by a processor, a set of convolutional filters to the set of positive input samples to generate the first set of feature maps. In some implementations, generating a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples may include applying, by a processor, a set of convolutional filters to the set of negative input samples to generate the second set of feature maps.
[0026] In some implementations, the set of convolution filters may include kernels of various sizes.
[0027] In some implementations, obtaining a first set of Gram matrix pairs associated with a set of positive input samples based on a first set of feature maps and a second set of Gram matrix pairs associated with a set of negative input samples based on a second set of feature maps may include determining a first correlation between the first set of feature maps associated with the set of positive input samples for obtaining the first set of Gram matrix pairs by the processor. In some implementations, obtaining a first set of Gram matrix pairs associated with a set of positive input samples based on the first set of feature maps and a second set of Gram matrix pairs associated with a set of negative input samples based on the second set of feature maps may include determining a second correlation between the second set of feature maps associated with the set of negative input samples for obtaining the second set of Gram matrix pairs by the processor.
[0028] According to yet other aspects, portions, features, and implementation details of the systems, methods, and non-transitory computer-readable media may be combined to form additional aspects, including aspects that exclude and / or modify some or portions of the individual components or features, include additional components or features, and / or other modifications, and all such modifications are within the scope of the present disclosure. [Brief explanation of the drawings]
[0029] [Figure 1] FIG. 1 illustrates an exemplary network environment, according to some implementations. [Figure 2] 1 is a diagram of a set of positive input samples and a set of negative input samples according to some implementations. [Figure 3A] 1 shows an example diagram of a feature map, according to some implementations. [Figure 3B] 3B shows a diagram of an exemplary Gram matrix calculation based on the feature map of FIG. 3A according to some implementations. [Figure 4] 3 is a diagram of a set of distinctive style features identified based on the set of positive and negative input samples of FIG. 2, according to some implementations. [Figure 5] 1A-1C are diagrams of input shapes and corresponding rendered assets according to some implementations. [Figure 6A] 1 is a flowchart of an example method for training and using a convolutional neural network (CNN) to identify distinctive style features, according to some implementations. [Figure 6B] 1 is a flowchart of an exemplary method for training and using a convolutional neural network (CNN) to identify distinctive style features, according to some implementations. [Figure 7] FIG. 1 is a block diagram illustrating an exemplary computing device, according to some implementations. DETAILED DESCRIPTION OF THE INVENTION
[0030] In the following detailed description, reference is made to the accompanying drawings, which form a part of the detailed description. Like symbols in the drawings typically identify like components unless context dictates otherwise. The illustrative implementations described in the detailed description, drawings, and claims are not intended to be limiting. Other implementations may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. Aspects of the present disclosure, as generally described herein and shown in the drawings, can be arranged, substituted, combined, separated, and designed into a wide variety of different configurations, all of which are contemplated herein.
[0031] References herein to "some implementations," "implementations," "example implementations," and the like indicate that the described implementations may include a particular feature, structure, or characteristic, but each implementation may not necessarily include that particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same implementation. Furthermore, when a particular feature, structure, or characteristic is described in connection with an implementation, such feature, structure, or characteristic may also be achieved in connection with other implementations, whether or not explicitly described.
[0032] Recent advances in gaming, metaverses, and user-generated content (UGC) platforms have lowered the expertise barrier when designing visual assets (e.g., 2D or 3D avatars, objects, and scenes). However, designing a virtual environment remains a complex undertaking, requiring significant user time to create new 3D objects, modify 3D objects, or select existing 3D objects to compose a desired scene. While designing 3D objects from scratch is complex and time-consuming, the rapid growth of commercial and platform-specific 3D asset markets allows users to browse and find existing objects that match a particular aesthetic. However, beyond finding objects that match a particular aesthetic, composing scenes with style-compatible objects and assets is important for designing aesthetically appealing virtual experiences. Toward this end, users may spend significant time and effort finding objects that match a particular style. As a result, users may reuse existing ones to avoid the complex task of creating new 3D objects from scratch, thereby compromising the specific style.
[0033] Providing easy-to-use tools for geometric style analysis and manipulation has the potential to improve how both novice and expert users compose 3D environments and objects. For example, guiding users to select style-compatible assets can greatly simplify their navigation through a 3D asset marketplace. This can help non-expert users create more visually appealing and professional scenes. Furthermore, providing editing tools that automatically change the visual style of existing objects allows users to reuse any existing 3D asset and coherently compose virtual experiences of any style. Finally, combining geometric style analysis and control capabilities with a generative model can effectively remove the expertise barrier involved in designing an overall virtual experience. For example, a generative model could be tasked with creating visual assets that match style embeddings and other style specifications (e.g., "generate a rabbit in a cubic style" when a cubic style is implied from the style embedding).
[0034] This disclosure provides techniques for automatically extracting a set of discriminative style features for each style using a loss function generated by comparing different image style forms. For example, this disclosure uses an artificial intelligence (AI)-based technique that automatically learns a set of distinctive style features specific to each input visual style. Distinct style features are features that are found consistently within the visual assets of the input visual style while simultaneously being common to the input visual style and distinct from other visual styles.
[0035] For example, corner shape, ornamentation, thickness, or curvature may be distinctive stylistic features that distinguish serif fonts from sans serif fonts. By training machine learning models to identify distinguishing features of one style from another, implementations described herein can estimate stylistic similarities between shapes and manipulate the shape of an asset to match the shape of an input for a virtual environment.
[0036] To that end, the method optimizes a set of convolution filter weights that activate geometric features that characterize the distinctive style characteristics of an input style. Using the optimized set of convolution filter weights, a convolutional neural network (CNN) learns the distinctive style characteristics based on a loss function generated using the optimized set of convolution filter weights. Once trained, the CNN can identify the distinctive style characteristics of any 2D or 3D shape that a user may use as input. In this way, the present disclosure can manipulate the shapes of objects and assets for a virtual environment to match the distinctive style characteristics of a particular aesthetic, thereby reducing the time and effort spent creating a visually appealing experience.
[0037] Figure 1: System architecture
[0038] FIG. 1 illustrates an exemplary network environment 100 according to some implementations of the present disclosure. FIG. 1 and other figures use similar reference numbers to identify similar elements. A letter following a reference number, such as "110a," indicates that the text specifically refers to the element with that particular reference number. A reference number in text without a letter following it, such as "110," refers to any or all elements in the figure with that reference number (e.g., "110" in the text refers to reference numbers "110a," "110b," and / or "110n" in the figures).
[0039] The network environment 100 (also referred to herein as a “platform”) includes an online virtual experience server 102, a data store 108, a client device 110 (or multiple client devices), and a third-party server 118, all connected through a network 122.
[0040] The online virtual experience server 102 may include, among other things, a virtual experience engine 104, one or more virtual experiences 105, and a machine learning component 130. The machine learning component 130 may include a model that is trained to identify distinctive stylistic features. Once trained, the model may modify or generate visual assets of a particular style that are associated with the distinctive stylistic features of input visual assets. In some implementations, the online virtual experience server 102 may be configured to provide the virtual experiences 105 to one or more client devices 110, provide automatic generation of a loss function through the machine learning component 130, and provide automatic generation of virtual experience assets (e.g., objects, avatars, etc.) that match the stylistic features of input images using a CNN trained using the loss function.
[0041] The data store 108 is shown coupled to the online virtual experience server 102. However, in some implementations, the data store 108 may be provided as part of the online virtual experience server 102. In some implementations, the data store 108 may be configured to store advertising data, user data, engagement data, and / or other contextual data associated with the machine learning component 130.
[0042] Each client device 110 (e.g., 110a, 110b, 110n) may include a virtual experience application 112 (e.g., 112a, 112b, 112n) and an I / O interface 114 (e.g., 114a, 114b, 114n) for interacting with the online virtual experience server 102 and for viewing graphical user interfaces (GUI), for example, through a computer monitor or display (not shown). In some implementations, client devices 110 may be configured to run and display virtual experiences, which may include the virtual user engagement portal described herein.
[0043] Network environment 100 is provided for illustrative purposes. In some implementations, network environment 100 may include the same, fewer, more, or different elements, configured in the same or different manner as shown in FIG.
[0044] In some implementations, the network 122 may include a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., a 702.11 network, a Wi-Fi network, or a wireless LAN (WLAN)), a cellular network (e.g., a Long Term Evolution (LTE) network, a Fifth Generation (5G) New Radio (NR) network), a router, a hub, a switch, a server computer, or a combination thereof.
[0045] In some implementations, the data store 108 may be non-transitory computer-readable memory (e.g., random access memory), a cache, a drive (e.g., a hard drive), a flash drive, a database system, or another type of component or device capable of storing data. The data store 108 may also include multiple storage components (e.g., multiple drives or multiple databases), which may span multiple computing devices (e.g., multiple server computers).
[0046] In some implementations, the online virtual experience server 102 may include a server having one or more computing devices (e.g., a cloud computing system, a rack-mounted server, a server computer, a cluster of physical servers, a virtual server, etc.). In some implementations, a server may be included in the online virtual experience server 102, may be a separate system, or may be part of another system or platform. In some implementations, the online virtual experience server 102 may be a single server, or any combination of multiple servers, load balancers, network devices, and other components. The online virtual experience server 102 may also be implemented on a physical server, although some implementations may utilize virtualization technology. Other variations of the online virtual experience server 102 also apply.
[0047] In some implementations, the online virtual experience server 102 may include one or more computing devices (e.g., rack-mounted servers, router computers, server computers, personal computers, mainframe computers, laptop computers, tablet computers, desktop computers, etc.), data stores (e.g., hard disks, memory, databases), networks, software components, and / or hardware components that may be used to perform operations on the online virtual experience server 102 and to provide users (through client devices 110) with access to the online virtual experience server 102.
[0048] The online virtual experience server 102 may also include a website (e.g., one or more web pages) or application backend software that may be used to provide users with access to content provided by the online virtual experience server 102. For example, users (or developers) may access the online virtual experience server 102 using a virtual experience application 112 on each client device 110.
[0049] In some implementations, the online virtual experience server 102 may contain the digital assets and requirements for creating digital virtual experiences. For example, the platform may provide an administrator interface that allows design, modification, personalization, and other modification capabilities. In some implementations, the virtual experience may include, for example, a two-dimensional (2D) game, a three-dimensional (3D) game, a virtual reality (VR) game, or an augmented reality (AR) game. In some implementations, virtual experience creators and / or developers may search for virtual experiences, combine parts of virtual experiences, or tailor virtual experiences for specific activities (e.g., group virtual experiences) and other features offered through the virtual experience server 102.
[0050] In some implementations, the online virtual experience server 102 or the client device 110 may include a virtual experience engine 104 or a virtual experience application 112. In some implementations, the virtual experience engine 104 may be used for developing or executing the virtual experience 105. For example, the virtual experience engine 104 may include, among other features, a rendering engine ("renderer") for 2D, 3D, VR, or AR graphics, a physics engine, a collision detection engine (and collision response), a sound engine, scripting capabilities, a haptic engine, an artificial intelligence engine, networking capabilities, streaming capabilities, memory management capabilities, threading capabilities, scene graph capabilities, or video support for filmmaking techniques. Components of the virtual experience engine 104 may generate instructions (e.g., rendering instructions, collision instructions, physics instructions, etc.) that assist in computing and rendering the virtual experience.
[0051] The online virtual experience server 102 may perform some or all virtual experience engine functions (e.g., generating physics instructions, rendering instructions, etc.) using the virtual experience engine 104, or may release some or all virtual experience engine functions to a virtual experience engine (not shown) at the client device 110. In some implementations, each virtual experience 105 may have a different ratio between virtual experience engine functions implemented on the online virtual experience server 102 and virtual experience engine functions implemented on the client device 110.
[0052] In some implementations, virtual experience instructions may refer to instructions that enable client device 110 to render gameplay, graphics, and other features of the virtual experience. The instructions may include one or more user inputs (e.g., positioning of physics objects), character position and velocity information, or instructions (e.g., physics instructions, rendering instructions, collision instructions, etc.).
[0053] In some implementations, each client device 110 may include a computing device such as a personal computer (PC), a mobile device (e.g., a laptop, a mobile phone, a smartphone, a tablet computer, or a netbook computer), a network-connected television, a game console, etc. In some implementations, client devices 110 may also be referred to as "user devices." In some implementations, one or more client devices 110 may connect to the online virtual experience server 102 at any one time. It should be noted that the number of client devices 110 is not limiting and is provided by way of example. In some implementations, any number of client devices 110 may be used.
[0054] In some implementations, each client device 110 may include an instance of a virtual experience application 112. The virtual experience application 112 may be rendered for interaction on the client device 110. During user interaction within the virtual experience or other GUI of the network environment 100, the user may create an asset (e.g., an object, an avatar, etc.) that includes distinctive stylistic features. For example, the user may input image data to be associated with an image style into the machine learning component 130. The machine learning component 130 may identify a set of distinctive stylistic features associated with the visual style using a CNN with kernel weights trained using a loss function to identify the distinctive stylistic features. The machine learning component 130 may adjust various features (e.g., shape, curvature, angle, etc.) of the asset, object, and / or scene to match the set of distinctive stylistic features of the visual style indicator (e.g., the input image data). Once the asset's features have been adjusted to match the set of distinctive stylistic features, the asset may be rendered by the machine learning component 130.
[0055] The machine learning component 130 may use convolutional filters and their corresponding feature maps to represent the identified set of distinctive style features. However, because the set of distinctive style features is initially unknown, the kernel weights are also unknown before training and cannot be handcrafted a priori. To identify the kernel weights, the machine learning component 130 may iteratively optimize the convolutional filters of the CNN so that the extracted features are maximally aligned to identify the distinctive style features. To that end, the learned features must satisfy the following conditions: 1) the features must be frequent within a style, and 2) the features must be distinct from features common to two or more styles simultaneously.
[0056] The machine learning component 130 softly models these requirements as individual conditions within a contrastive framework. Contrastive learning aims to encourage the extracted style features of objects belonging to the same style to be identical and different for objects with different styles. However, these goals are position-dependent and cannot be directly modeled based on learned feature maps. This means that convolving the same set of learned filters on inputs of the same style but with slightly different positions will lead to different feature maps. As a result, feature maps do not provide a reliable metric for comparing styles.
[0057] Below, a more detailed discussion of the machine learning component 130 and its operation is presented with reference to FIGS. 2-5.
[0058] Figure 2-5: Learning and using model components of unique features
[0059] Figure 2 shows a diagram 200 of a set of positive input samples and a set of negative input samples, according to some implementations. Figure 3A shows a diagram 300 of generated feature maps, according to some implementations. Figure 3B shows a diagram 301 of an example Gram matrix calculation based on the feature maps of Figure 3A, according to some implementations. Figure 4 shows a diagram 400 of a set of distinctive style features identified based on the set of positive input samples and the set of negative input samples of Figure 2, according to some implementations. Figure 5 shows a diagram 500 of input shapes and corresponding rendered assets, according to some implementations. Figures 2-5 are discussed simultaneously.
[0060] Instead of using the feature maps directly, the machine learning component 130 identifies distinctive style features based on correlations between feature maps using Gram matrices, as described below with reference to FIGS. 2, 3A, and 3B.
[0061] 2 and 3A, the feature map 302 may be generated based on the set of positive input samples 202 and the set of negative input samples 204. For example, the set of positive input samples 202 may include images of shapes of the same style (e.g., characters in a serif font), while the set of negative input samples 204 may include images of shapes of different styles (e.g., characters in a serif font and a sans-serif font). The machine learning component 130 may measure the similarity between two shapes (e.g., letter or number fonts, 2D objects, 3D objects, etc.) based on the overall difference in the correlation of their filter responses. This representation has the advantage of removing position information and making the representation position-independent.
[0062] In the non-limiting example of Figure 2, positive input samples 202 include images of the letters D, K, S, and W, all in a sans-serif font. That is, set of positive input samples 202 includes shapes of the same style. Negative input samples 204 include shape pairs of one or more different styles. For example, negative input samples 204 include a pair of Ds, a pair of Ks, a pair of Ss, and a pair of Ws. Each shape pair in set of negative input samples 204 includes a first shape in a first style (e.g., a sans-serif font) and a second shape in a second style (e.g., a serif font). While Figure 2 illustrates styles defined in a human-understandable manner, more complex styles may be used for visual assets on a virtual environment platform and may be distinguished based on visual appearance without specific style names being associated with each style.
[0063] The machine learning component 130 may generate a first set of feature maps associated with the set of positive input samples 202 and a second set of feature maps associated with the set of negative input samples 204. The first set of feature maps may be generated by applying a set of convolution filters to the set of positive input samples 202, and the second set of feature maps may be generated by applying the set of convolution filters to the set of negative input samples 204. By way of example, the set of convolution filters may include kernels of size 5×5, 7×7, and 9×9, or any other suitable size. Each filter in the set of convolution filters operates independently of the other filters in the set. This means that the machine learning component 130 may input the same input shape to each filter separately. The output of the set of convolution filters may be a set of feature maps 302 (hereinafter referred to as “feature maps 302”), as depicted in FIG. 3A .
[0064] 3A, the machine learning component 130 may flatten the feature map 302 to obtain a flat horizontal feature map 304a and a flat vertical feature map 304b.
number
number
[0065] Referring to Figure 3B,
number
number
number
number
number
number
number
[0066] 3B, the Gram matrix 306 captures the feature correlation between each matrix pair in the flat horizontal feature map 304a and the flat vertical feature map 304b. A high feature correlation value for a matrix pair indicates that the image has the feature captured by the first feature map and the feature captured by the second feature map. Conversely, a relatively low feature correlation value indicates that the image does not have the feature captured by the first feature map from the flat horizontal feature map 304a and the feature captured by the second feature map from the flat vertical feature map 304b.
[0067] Using the feature correlation differences, the machine learning component 130 can calculate a loss function that compares the Gram matrices and solves the function in equation (1) for the set of filter weights that minimizes the correlation of the corresponding feature maps.
number
[0068] In [Number 10], I + represents the set of positive input samples 202, and I - represents the set of negative input samples 204, i represents the first shape, and j represents the second shape. In Equation 2, I + The summation for the stylistic shapes (
number
number
[0069] An example of distinctive stylistic features captured by the loss function shown in Equation 2 is depicted in Figure 4. Using the loss function, the machine learning component 130 identifies a set of distinctive stylistic features for serif fonts. For example, the set of distinctive stylistic features may capture serif ornamentation 402 and thickness change 404 as distinctive stylistic features that distinguish between sans serif font styles (left column) and serif font styles (right column).
[0070] Referring to FIG. 5 , once the kernel weights for the CNN filter are trained using the loss function, the machine learning component 130 can identify distinctive style features for any input shape 502. The machine learning component 130 can render the asset 504 to include the distinctive style features of the input shape 502. The visual style of an input asset (not shown) (e.g., an existing cow shape in the asset repository or a new cow shape) can be manipulated or transformed to match the visual style of the input shape 502. For example, when designing a virtual experience, a user may input image data associated with a circle or sphere into the machine learning component 130. Using the trained CNN filter, distinctive style features, including circular features, can be identified for the circle or sphere. The machine learning component 130 can apply circular features to the cow's features such that the asset's body shape, head shape, leg shape, horn shape, ear shape, etc. are styled with the circular features of the shape in the user's input image.
[0071] Below, a more detailed discussion of the methods for training and using the CNNs of the present disclosure is presented with reference to FIGS. 6A and 6B.
[0072] Figures 6A and 6B: Exemplary methods for training and using CNNs
[0073] 6A and 6B are flowcharts of a method 600 for training and using a CNN to render a scene based on distinctive stylistic features of an input shape, according to some implementations.
[0074] In some implementations, method 600 may be implemented on, for example, server 102 described with reference to FIG. 1 . In some implementations, some or all of method 600 may be implemented on one or more client devices 110, one or more developer devices (not shown), one or more server devices 102, and / or a combination of developer devices, server devices, and client devices, as shown in FIG. 1 . In the described embodiments, the implemented system includes one or more digital processors or processing circuits (“processors”) and one or more storage devices (e.g., data store 108 or other storage). In some implementations, different components of one or more servers and / or clients may perform different blocks or other portions of method 600. In some embodiments, a first device is described as performing blocks of method 600. Some implementations may have one or more blocks of method 600 performed by one or more devices (e.g., other client devices or server devices) that can send results or data to the first device.
[0075] In some implementations, method 600, or portions thereof, may be initiated automatically by a system. In some implementations, the implementing system is a first device. For example, the method (or portions thereof) may be performed periodically, or may be performed based on one or more specific events or conditions, such as upon a user request, upon receipt of an input form, and / or one or more occurring conditions, which may be specified in a setting read by the method. With reference to FIG. 6A , method 600 may begin at block 602. At block 602, a set of positive input samples associated with the same visual style and a set of negative input samples associated with different visual styles may be received. For example, with reference to FIGS. 1 and 2 , the machine learning component 130 may receive a set of positive input samples 202 including images of forms in the same style (e.g., characters in a serif font) and a set of negative input samples 204 including images of forms in different styles (e.g., characters in a serif font and a sans serif font).
[0076] Block 604 may follow block 602. In block 604, a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples may be generated. For example, with reference to FIGS. 1 and 3A , the machine learning component 130 may generate feature maps 302 based on the set of positive input samples 202 (e.g., the first set of feature maps) and the set of negative input samples 204 (e.g., the second set of feature maps). For example, the set of positive input samples 202 may include images of features in the same style (e.g., characters in a serif font), while the set of negative input samples 204 may include images of features in different styles (e.g., characters in a serif font and a sans serif font).
[0077] Block 606 may follow block 604. In block 606, a first set of Gram matrix pairs associated with a set of positive input samples based on a first set of feature maps and a second set of Gram matrix pairs associated with a set of negative input samples based on a second set of feature maps may be obtained. For example, with reference to FIGS. 1 and 3B,
number
number
[0078] Block 608 may follow block 606. In block 608, a loss function may be calculated that is based on a first sum of differences between a first set of Gram matrix pairs and a second sum of differences between a second set of Gram matrix pairs. For example, with reference to FIGS. 1 and 4, using feature correlation differences, the machine learning component 130 may calculate the loss function. The loss function compares Gram matrices and solves for a set of filter weights that minimizes the energy of equation (2).
[0079] Block 610 may follow block 608. At block 610, a visual asset may be received. For example, with reference to FIG. 1 , the machine learning component 130 may receive image data associated with a visual asset (e.g., an image, an object, or a scene). The image data may be associated with a single visual asset or may be associated with all visual assets in a scene to be rendered.
[0080] Block 612 may follow block 610. In block 612, a visual style indicator may be received. For example, with reference to FIGS. 1 and 5 , the machine learning component 130 may receive an input form 502 (visual style indicator) of an image, object, or scene. The input form 502 may be selected and entered, for example, by a user. The input form 502 may be associated with a particular visual style.
[0081] Referring to FIG. 6B , block 614 may follow block 612. In block 614, a set of distinctive stylistic features for the visual style may be identified using a loss function calculated based on a first sum of differences between a first set of Gram matrix pairs and a second sum of differences between a second set of Gram matrix pairs. For example, referring to FIGS. 1 and 4 , an example of distinctive stylistic features obtained by the loss function shown in equation (2) is depicted. Using the loss function, the machine learning component 130 may identify a set of distinctive stylistic features for a serif font. For example, the set of distinctive stylistic features may obtain serif ornamentation 402 and thickness change 404 as distinctive stylistic features that distinguish between a sans serif font style (left column) and a serif font style (right column).
[0082] Block 616 may follow block 614. In block 616, the shape of the visual asset may be modified to match the identified set of distinctive stylistic features for the visual style. For example, with reference to FIGS. 1 and 5 , once the kernel weights for the CNN filter are trained using the loss function, the machine learning component 130 may identify distinctive stylistic features for any input shape 502. The machine learning component 130 may adjust the shape of the asset 504 to include the distinctive stylistic features of the input shape 502. For example, when designing a virtual experience, a user may input image data associated with a circle or a sphere into the machine learning component 130. Using the trained CNN filter, distinctive stylistic features, including circular features, may be identified for the circle or sphere. The machine learning component 130 may apply circular features to the features of a cow, such that the asset's body shape, head shape, leg shape, horn shape, ear shape, etc. are styled with the circular features of the shape in the user's input image. In some implementations, the machine learning component 130 may reposition one or more vertices of the mesh of the visual asset to align with distinctive style features identified for the visual style and / or texture of the visual asset.
[0083] Block 618 may follow block 616. At block 618, the modified visual asset may be rendered based on the set of unique style characteristics. Block 618 may end the operation. For example,
[0084] Figure 7: Computing Device
[0085] Below, with reference to FIG. 7, a more detailed description of various computing devices that may be used to implement the different devices and / or components shown in FIG. 1 is provided.
[0086] FIG. 7 is a block diagram of an exemplary computing device 700 that may be used to implement one or more features described herein, according to some implementations. In one example, device 700 may be used to implement a computer device (e.g., 102, 110 of FIG. 1 ) and perform suitable operations as described herein. Computing device 700 may be any suitable computer system, server, or other electronic or hardware device. For example, computing device 700 may be a mainframe computer, desktop computer, workstation, portable computer, or electronic device (such as a portable device, mobile device, mobile phone, smartphone, tablet computer, television, TV set-top box, personal digital assistant (PDA), media player, game console, wearable device, etc.). In some implementations, device 700 includes a processor 702, memory 704, input / output (I / O) interface 706, and audio / video input / output device(s) 714 (e.g., a display screen, a touch screen, display goggles or glasses, audio speaker, headphones, a microphone, etc.).
[0087] Processor 702 may be one or more processors and / or processing circuits for executing program code and controlling the underlying operation of device 700. A "processor" includes any suitable hardware and / or software system, mechanism, or component for processing data, signals, or other information. A processor may include a general-purpose central processing unit (CPU), a system with multiple processing units, dedicated circuits for achieving functions, or other systems. Processing need not be limited to a particular geographic location or time limit. For example, a processor may perform its functions in "real-time," "offline," "batch mode," etc. Portions of processing may be performed at different times and in different locations by different (or the same) processing systems. A computer may be any processor in communication with a memory.
[0088] Memory 704 is typically provided in device 700 for access by processor 702 and may be any suitable processor-readable storage medium suitable for storing instructions for execution by the processor, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc., located remotely from processor 702 and / or integrated with processor 702. Memory 704 may store software executed by processor 702 on server device 700, including an operating system 708, software applications 710, and associated databases 712. In some implementations, applications 710 may include instructions that enable processor 702 to perform the functions described herein. Software applications 710 may include some or all of the functionality necessary to train a CNN filter to capture distinctive style features of an input shape and to manipulate the distinctive style features of the input shape to match the shape of an asset. In some implementations, portions of one or more software applications 710 may be implemented in dedicated hardware such as an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), a machine learning processor, etc. In some implementations, portions of one or more software applications 710 may be implemented in a general-purpose processor such as a central processing unit (CPU) or a graphics processing unit (GPU). In various implementations, an appropriate combination of dedicated and / or general-purpose processing hardware may be used to implement the software application 710.
[0089] For example, software application 710 stored in memory 704 may include instructions necessary to train a CNN filter to obtain distinctive style features of an input shape, and to manipulate the distinctive style features of the input or software such as machine learning component 130 to match the shape of an asset. Any of the software in memory 704 may alternatively be stored in any other suitable storage location or computer-readable medium. Additionally, memory 704 (and / or other connected storage devices) may store instructions and data used for the features described herein. Memory 704 and any other storage type (such as magnetic disk, optical disk, magnetic tape, or other tangible medium) may be considered "storage" or "storage device."
[0090] The I / O interface 706 may provide functionality to allow other systems and devices to interface with the server device 700. For example, network communication devices, storage devices (e.g., memory and / or data store 108), and input / output devices may communicate through the interface 706. In some implementations, the I / O interface may connect to interface devices including input devices (keyboards, pointing devices, touchscreens, microphones, cameras, scanners, etc.) and / or output devices (display devices, speaker devices, printers, motors, etc.).
[0091] For ease of explanation, FIG. 7 shows one block for each of processor 702, memory 704, I / O interface 706, operating system 708, software applications 710, and database 712. These blocks may represent one or more processors or processing circuits, operating systems, memories, I / O interfaces, applications, and / or software modules. In other implementations, device 700 may not have all the components shown and / or may have other elements, including other types of elements instead of or in addition to the types of elements shown herein. While virtual experience server 102 is described as performing operations as described in some implementations herein, any suitable component or combination of components of virtual experience server 102, or a similar system, or processor or any suitable processor associated with such a system, may perform the described operations.
[0092] Additionally, user devices may implement and / or be used in conjunction with features described herein. An exemplary user device may be a computing device including several components similar to device 700, e.g., processor 702, memory 704, and / or I / O interface 706. Appropriate operating systems, software, and applications for the client device may be provided in the memory and used by the processor. The I / O interface for the client device may be connected to a network communication device and input / output devices, such as a microphone for capturing sound, a camera for capturing images or video, an audio speaker device for outputting sound, a display device for outputting images or video, or other output devices. A display device within audio / video input / output devices 714 may be connected to (or included in) device 700 to, for example, display the image pre-processing and post-processing described herein; such a display device may include any suitable display device, e.g., an LCD, LED, or plasma display screen, a CRT, a television, a monitor, a touchscreen, a 3D display screen, a projector, or other visual display device. Some implementations may provide an audio output device, for example, sound output, or text-to-speak synthesis.
[0093] Methods, blocks, and / or operations described herein may, where appropriate, be performed in an order different from that shown or described, and / or concurrently (partially or fully) with other blocks or operations. Some blocks or operations may be performed for a portion of data and then performed again, e.g., for a different portion of data. Not all of the blocks and operations described need be performed in various implementations. In some implementations, blocks and operations may be performed multiple times, in a different order, and / or at different times in the method.
[0094] In some implementations, some or all of the methods may be performed on a system, such as one or more client devices. In some implementations, one or more methods described herein may be performed on a server system and / or both a server system and a client system. In some implementations, different components of one or more servers and / or clients may perform different blocks, operations, or other portions of a method.
[0095] One or more methods described herein (e.g., method 600) may be implemented by computer program instructions or code that can be executed on a computer. For example, the code may be implemented by one or more digital processors (e.g., microprocessors or other processing circuits) and stored in a computer program product that includes a non-transitory computer-readable medium (e.g., storage medium), such as a magnetic, optical, electromagnetic, or semiconductor storage medium, including semiconductor or solid-state memory, magnetic tape, removable computer diskettes, random access memory (RAM), read-only memory (ROM), flash memory, rigid magnetic disks, optical disks, solid-state memory drives, etc. The program instructions may also be contained in or provided as electronic signals, for example, in the form of software as a service (SaaS) delivered from a server (e.g., a distributed system and / or a cloud computing system). Alternatively, one or more methods may be implemented in hardware (e.g., logic gates) or a combination of hardware and software. Exemplary hardware may be a programmable processor (e.g., a field programmable gate array (FPGA), a complex programmable logic device), a general-purpose processor, a graphics processor, an application-specific integrated circuit (ASIC), etc. One or more of the methods may be implemented as a component of or part of an application running on the system, or as an application or software running in conjunction with other applications and the operating system.
[0096] One or more methods described herein may be operated in a standalone program that may be run on any type of computing device, a program run in a web browser, or a mobile application ("app") running on a mobile computing device (e.g., a mobile phone, a smartphone, a tablet computer, a wearable device (such as a watch, armband, jewelry, hat, goggles, glasses, etc.), a laptop computer, etc.). In one embodiment, a client / server architecture may be used, e.g., a mobile computing device (as a client device) sends user input data to a server device and receives live feedback data from the server for output (e.g., for display). In another embodiment, computing may be split between the mobile computing device and one or more server devices.
[0097] Although specific implementations have been described in this description, these specific implementations are merely illustrative and not limiting, and the concepts shown in the examples may be applied to other examples and implementations.
[0098] It should be noted that the functional blocks, operations, features, methods, devices, and systems described in this disclosure may be integrated or separated into various combinations of systems, devices, and functional blocks as known to those skilled in the art. Any suitable programming language and programming techniques may be used to implement the routines of a particular implementation. Various programming techniques, e.g., procedural or object-oriented, may be used. The routines may be executed on a single processing device or on multiple processors. While steps, operations, or computations may be presented in a particular order, the order may change in different particular implementations. In some implementations, multiple steps or operations shown sequentially in this specification may be performed simultaneously. [Explanation of symbols]
[0099] 100 Network Environment 102 Online Virtual Experience Server 104 Virtual Experience Engine 105 Virtual Experiences 108 Datastores 110 client devices 112 Virtual Experience Application 114 I / O interfaces 118 Third Party Servers 122 Network 130 Machine Learning Components A set of 202 positive input samples A set of 204 negative input samples 302 Feature Map 304a Horizontal feature map 304b Vertical feature map 306 Gram Matrix 402 Serif Decoration 404 Change 404a Horizontal feature map 404b Vertical feature map 502 series 504 Assets 600 ways 700 computing devices 702 processor 704 memory 706 Input / Output (I / O) Interface 708 Operating Systems 710 Software Applications 712 databases 714 Video Input / Output Device
Claims
1. 1. A computer-implemented method comprising: receiving the visual asset by a processor; receiving, by said processor, an indicator of a visual style; identifying, by the processor, a set of distinctive style features for the visual style using a loss function calculated based on a first sum of differences between a first set of Gram matrix pairs, a first set of Gram matrix pairs associated with a set of positive input samples of the same visual style, and a second set of Gram matrix pairs associated with a set of negative input samples of a different visual style, and a second sum of differences between the first set of Gram matrix pairs and a second set of Gram matrix pairs associated with a set of negative input samples of a different visual style; modifying, by the processor, the shape of the visual asset to match the set of unique style features identified for the visual style; and rendering, by the processor, the visual assets after the modification to match the set of distinctive style characteristics.
2. receiving, by the processor, the set of positive input samples of the same visual style and the set of negative input samples of the different visual style; generating, by the processor, a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples; obtaining, by the processor, the first set of Gram matrix pairs associated with the set of positive input samples based on the first set of feature maps and the second set of Gram matrix pairs associated with the set of negative input samples based on the second set of feature maps; 2. The method of claim 1 , further comprising: calculating, by the processor, the loss function based on a first sum of the differences of the first set of Gram matrix pairs and a second sum of the differences of the second set of Gram matrix pairs.
3. the set of positive input samples associated with the same visual style includes a first set of shape pairs of the same style; The method of claim 2 , wherein the set of negative input samples associated with the different visual styles comprises a second set of shape pairs of different styles.
4. The method of claim 3 , wherein one shape of the second set of shape pairs has the same style as a style of a first set of shape pairs.
5. generating the first set of feature maps associated with the set of positive input samples and the second set of feature maps associated with the set of negative input samples, applying, by the processor, a set of convolution filters to the set of positive input samples to generate the first set of feature maps; and applying, by the processor, the set of convolution filters to the set of negative input samples to generate the second set of feature maps.
6. The method of claim 5 , wherein the set of convolution filters includes kernels of different sizes.
7. obtaining the first set of Gram matrix pairs associated with the set of positive input samples based on the first set of feature maps and the second set of Gram matrix pairs associated with the set of negative input samples based on the second set of feature maps; determining, by the processor, first correlations between a first set of the feature maps associated with a set of positive input samples for obtaining the first set of Gram matrix pairs; and determining, by the processor, second correlations between a second set of the feature maps associated with the set of negative input samples for obtaining the second set of Gram matrix pairs.
8. 1. A computing device comprising: a processor; When instructions stored in memory are executed by the processor, Receiving visual assets; Receiving visual style indicators and identifying a set of distinctive style features for the visual style using a loss function calculated based on a first sum of differences between a first set of Gram matrix pairs, where a first set of Gram matrix pairs is associated with a set of positive input samples of the same visual style, and a second set of Gram matrix pairs is associated with a set of negative input samples of a different visual style, and a second sum of differences between the second set of Gram matrix pairs; modifying the shape of the visual asset to match the set of unique style features identified for the visual style; rendering the visual assets after said modification to match the set of distinctive style characteristics; a memory coupled to the processor that causes the processor to perform operations comprising:
9. receiving the set of positive input samples of the same visual style and the set of negative input samples of the different visual style; generating a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples; obtaining the first set of Gram matrix pairs associated with the set of positive input samples based on the first set of feature maps and the second set of Gram matrix pairs associated with the set of negative input samples based on the second set of feature maps; 9. The computing device of claim 8, further comprising: calculating the loss function based on a first sum of the differences between the first set of Gram matrix pairs and a second sum of the differences between the second set of Gram matrix pairs.
10. the set of positive input samples associated with the same visual style includes a first set of shape pairs of the same style; The computing device of claim 9 , wherein the set of negative input samples associated with the different visual styles comprises a second set of shape pairs of different styles.
11. The computing device of claim 10 , wherein a shape in the second set of shape pairs has the same style as a style in the first set of shape pairs.
12. generating the first set of feature maps associated with the set of positive input samples and the second set of feature maps associated with the set of negative input samples, applying a set of convolution filters to the set of positive input samples to generate the first set of feature maps; and applying the set of convolution filters to the set of negative input samples to generate the second set of feature maps.
13. The computing device of claim 12 , wherein the set of convolution filters includes kernels of different sizes.
14. obtaining the first set of Gram matrix pairs associated with the set of positive input samples based on the first set of feature maps and the second set of Gram matrix pairs associated with the set of negative input samples based on the second set of feature maps; determining a first correlation between the first set of feature maps associated with a set of positive input samples for obtaining the first set of Gram matrix pairs; and determining second correlations between a second set of the feature maps associated with the set of negative input samples to obtain the second set of Gram matrix pairs.
15. A non-transitory computer-readable medium that, when executed by a processor, contains instructions stored thereon: Receiving visual assets; Receiving visual style indicators and identifying a set of distinctive style features for the visual style using a loss function calculated based on a first sum of differences between a first set of Gram matrix pairs, where a first set of Gram matrix pairs is associated with a set of positive input samples of the same visual style, and a second set of Gram matrix pairs is associated with a set of negative input samples of a different visual style, and a second sum of differences between the second set of Gram matrix pairs; modifying the shape of the visual asset to match the set of unique style features identified for the visual style; rendering the visual assets after said modification to match the set of distinctive style characteristics; A non-transitory computer-readable medium that causes the processor to perform operations comprising:
16. receiving the set of positive input samples of the same visual style and the set of negative input samples of the different visual style; generating a first set of feature maps associated with the set of positive input samples and a second set of feature maps associated with the set of negative input samples; obtaining the first set of Gram matrix pairs associated with the set of positive input samples based on the first set of feature maps and the second set of Gram matrix pairs associated with the set of negative input samples based on the second set of feature maps; and calculating the loss function based on a first sum of differences between the first set of Gram matrix pairs and a second sum of differences between the second set of Gram matrix pairs.
17. the set of positive input samples associated with the same visual style includes a first set of shape pairs of the same style; 17. The non-transitory computer-readable medium of claim 16, wherein the set of negative input samples associated with the different visual styles comprises a second set of shape pairs of different styles.
18. 20. The non-transitory computer-readable medium of claim 17, wherein a shape in the second set of shape pairs has the same style as a style in the first set of shape pairs.
19. generating the first set of feature maps associated with the set of positive input samples and the second set of feature maps associated with the set of negative input samples, applying a set of convolution filters to the set of positive input samples to generate the first set of feature maps; and applying the set of convolution filters to the set of negative input samples to generate the second set of feature maps.
20. 20. The non-transitory computer-readable medium of claim 19, wherein the set of convolution filters includes kernels of different sizes.
Citation Information
Patent Citations
Image synthesis device and method for embedding watermark
US20220156873A1