Asset creation using generative artificial intelligence
Generative AI is used to iteratively generate and edit attributes of video game characters, addressing the time-consuming nature of character development and enhancing the efficiency of game creation.
Patent Information
- Application Number
- JP2025009652
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2025-01-23
- Publication Date
- 2025-08-04
AI Technical Summary
The development of characters for video games is time-consuming, leading to potential delays in the release of the game, as the creation process from initial conceptual representations to detailed final versions can take several days to several months.
A method using generative artificial intelligence to iteratively generate and edit attributes of a target object, including decomposition, selection, and blending phases to create multiple versions of the object, facilitated by an image generation AI system performing latent diffusion.
This approach significantly reduces the time required to develop characters by enabling dynamic and intuitive creation of visual assets, allowing for rapid iteration and adjustment of attributes, thereby accelerating the game development process.
Smart Images

Figure 2025114008000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to creating visual assets using generative artificial intelligence, and more specifically, to creating custom prompts for execution by a generative artificial intelligence system to generate a target object through an iterative process.
Background Art
[0002] Video games and / or game applications, and their related industries (e.g., video games) are very popular and occupy a large proportion of the world's entertainment market. Video games are played anytime, anywhere using various types of platforms including game consoles, desktop computers, laptop computers, mobile phones, etc.
[0003] The development of characters for video games can be time-consuming. One or more visual representations of the character are created and modified by the creation team from concept to final version. For example, each of the visual representations can be generated manually. At the start of the development cycle, the representations are conceptual and may be created without much consideration for details. These initial representations can be generated relatively quickly. On the other hand, as the development cycle nears its end, the character's representation becomes very detailed. These later representations require more time to generate. The entire development cycle can last from several days to several months, or even longer, until the creation team is satisfied with the final version of the character. The time taken for character development affects the development cycle of the complete video game, and as a result, if character development takes too long, the release date of the video game will be delayed.
[0004] Embodiments of the present disclosure arise in such circumstances.
Summary of the Invention
Problems to be Solved by the Invention
[0005] Embodiments of the present disclosure relate to the creation of visual assets using generative artificial intelligence, and more specifically, one or more versions of a target object can be generated through an iterative process by selecting and editing determined attributes of the target object.
[0006] In one embodiment, a method is disclosed. The method includes collecting one or more inputs each describing a target object. The method includes generating a plurality of images of the target object using an image generation artificial intelligence system configured to perform latent diffusion based on the one or more inputs. The method includes decomposing the target object into a first plurality of attributes each including one or more variations based on the plurality of images of the target object. The method includes receiving a selection of one or more of the plurality of variations of the plurality of attributes. The method includes blending one or more of the selected plurality of variations of the plurality of attributes into one or more options of the target object.
[0007] In another embodiment, a non-transitory computer-readable medium storing a computer program for implementing the method is disclosed. The computer-readable medium includes program instructions for collecting one or more inputs each describing a target object. The computer-readable medium includes program instructions for generating a plurality of images of the target object using an image generation artificial intelligence system configured to perform latent diffusion based on the one or more inputs. The computer-readable medium includes program instructions for decomposing the target object into a first plurality of attributes each including one or more variations based on the plurality of images of the target object. The computer-readable medium includes program instructions for receiving a selection of one or more of the plurality of variations of the plurality of attributes. The method includes blending one or more of the selected plurality of variations of the plurality of attributes into one or more options of the target object.
[0008] In yet another embodiment, a computer system is disclosed. The computer system includes a processor and a memory coupled to the processor and storing internally instructions that, when executed by the computer system, cause the computer system to execute a method. The method includes collecting one or more inputs each describing a target object. The method includes generating a plurality of images of the target object using an image generation artificial intelligence system configured to perform potential diffusion based on the one or more inputs. The method includes decomposing the target object into a first plurality of attributes each including one or more variations based on the plurality of images of the target object. The method includes receiving a selection of one or more of the plurality of variations of the plurality of attributes. The method includes blending one or more of the selected plurality of variations of the plurality of attributes into one or more options of the target object.
[0009] Other aspects of the present disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The present disclosure may be best understood by reference to the following description taken in conjunction with the accompanying drawings.
[0011]
Figure 1A
[0012]
Figure 1B
[0013]
Figure 2A
[0014]
Figure 2B
[0015]
Figure 2C
[0016]
Figure 3
[0017]
Figure 4A
[0018]
Figure 4B
[0019]
Figure 4C
[0020]
Figure 5
[0021]
Figure 6
DETAILED DESCRIPTION OF THE INVENTION
[0022] The following detailed description includes many specific details for illustrative purposes, but those skilled in the art will understand that many variations and modifications to the following details are within the scope of the present disclosure. Accordingly, aspects of the present disclosure are presented without any loss of generality with respect to the claims that follow this description and without imposing any limitation on those claims.
[0023] Generally speaking, various embodiments of the present disclosure describe systems and methods for creating visual assets using generative artificial intelligence (AI), and one or more versions of a target object or asset can be iteratively generated through the selection and editing of attributes of the AI-generated target object. The techniques proposed in the embodiments of the present disclosure enable the dynamic generation of original and visually usable content and can be applied to game elements and other visual elements (e.g., game characters, use of animations, visual elements / assets) within video games or other applications. Multiple phases may be executed to create one or more versions of an asset, and these phases may also be executed iteratively. For example, the phases include an input phase configured to generate a custom prompt for use by a generative AI system, a decomposition phase to identify one or more attributes of the asset, an iterative phase to select, edit, and / or adjust one or more attributes and / or variations of those attributes, and a merge / blend phase to generate different permutations of the asset based on the selected attributes and / or selected variations of one or more selected attributes.
[0024] Advantages of embodiments of the present disclosure include providing an intuitive and / or visual way to create target objects or assets, such as characters for video games, during development. The process used to create the target object / asset allows a user to selectively view different variations of the attributes of the target object / asset for purposes of selection, editing, and adjustment. For example, the user interface may be configured to provide one or more visual interfaces that allow selection, editing, or adjustment of each variation of the corresponding attribute. In one embodiment, the user interface allows locking or approval of attributes and / or variations of attributes. By providing the selection and / or change to the generative AI system for the variation of the attribute, another iteration of one or more variations of another set of attributes for the target object may be generated. Thus, the user interface presents one or more variations of the attributes of the target object and, further, is configured to provide the user with the ability to select and / or edit and / or adjust those one or more variations of the attributes in a single or iterative process that each generates a new set of attributes and their corresponding variations for the target object. In addition, when the generated attributes are approved, one or more permutations and / or versions of the target object are generated and are made visible in the user interface.
[0025] Throughout this specification, references to "game" or "video game" or "game application" are meant to represent any type of interactive application that is directed through the execution of input commands. For illustrative purposes only, interactive applications include applications for games, document processing, video processing, video game processing, and the like. Also, the terms "virtual world" or "virtual environment" or "metaverse" are meant to represent any type of environment that is generated by one or more corresponding applications for interaction among multiple users in a multiplayer session or a multiplayer game session. Further, the terms introduced above are compatible with each other.
[0026] Based on the foregoing general understanding of the various embodiments, the following describes exemplary details of the embodiments with reference to the various drawings.
[0027] FIG. 1A shows a system 100 configured for creating assets and / or target objects using a generative artificial intelligence (AI) and / or an AI model, according to one embodiment of the present disclosure. In particular, visual assets may be created using a generative artificial intelligence (AI). Here, one or more versions of a target object or asset may be iteratively generated by selecting and editing the attributes of the AI-generated target object. For example, an asset or target object may be a character in a video game, and / or the use of animation, and / or visual elements / assets that may occur within a corresponding video game. In general, embodiments of the present disclosure enable the dynamic generation of original and visual usable content as a target object, and are applicable to game elements, the use of animation, and other visual elements / assets (e.g., game characters, non-player character assets, backgrounds, texture elements, world / landscape elements such as furniture or trees, etc., and other visual objects) that occur within a video game or other application. For illustrative purposes only, throughout this specification, the asset created may be a fictional dragon. The system 100 may be implemented at a backend cloud service or as an intermediate layer third-party service remote from the client device.
[0028] As shown, an original input 101 is provided to a target object construction unit 105. In part, the target object construction unit 105 implements a generative AI configured to construct an asset or target object. More specifically, the target object construction unit 105 is configured to generate a plurality of attributes for the asset, each of which may include one or more variations. By combining different variations of the plurality of attributes, one or more permutations of the asset may be generated by the target object construction unit 105.
[0029] For example, since the target object construction unit 105 includes attribute generation, it may include one or more artificial intelligence processing engines and / or models for creating target objects and / or assets. For example, an image generation AI (IGAI) system may be configured to implement a generative AI used to generate one or more output images, graphics, and / or 3D representations of an asset. Additional artificial intelligence may be implemented to identify attributes of the one or more output images (e.g., via one or more AI models). Additionally, since there may be multiple output images generated by the IGAI system, each attribute may include one or more variations. For example, the tail of a dragon (commonly used as a representation of a desired target object) may include multiple characteristics or variations such as a long tail, a short tail, a stubby tail, etc. Therefore, artificial intelligence may be used to further identify each variation of the attribute.
[0030] The target object construction unit 105 may perform multiple phases to create visual assets using generative AI, and these phases may be repeatedly performed to output one or more permutations of the target object using user input. In particular, four phases may be performed. These phases may include an input phase 110, a decomposition phase 120 for generating an array of attributes, a selection and / or adjustment and / or iteration phase 130, and a merge / blend phase 150.
[0031] In the input phase 110, a custom prompt is generated based on the original input for implementation by the generative AI system. This prompt is targeted at the generation of the target object. The original input may be in any format, including text, voice annotations, visual images, and / or sequences of images, etc. Generally, the original input may describe the desired target object and may further include parameters defining the target object.
[0032] Furthermore, in the decomposition phase 120, the target object is decomposed into one or more attributes or components. This can be achieved by prompting the generative AI system to generate selectable components / attributes for the target object. The prompt may further define a specific artistic style of the target object, such as an artistic style inferred from the original input. In one embodiment, the generative AI system generates one or more representations of the target object. The generative AI system also generates different variations (e.g., one or more features or characteristics or samples, etc.) for each of the components / attributes based on the one or more representations. The attributes and their variations may be presented within an array. For example, one or more AI models may be configured to identify the attributes and / or variations of each attribute.
[0033] In the iteration phase 130, the user can select, edit, and / or adjust each of the one or more variations of each component / attribute. In particular, the user can focus on a specific component / attribute, more specifically, the variation of that component / attribute, to make additional modifications using the prompt. In particular, during the iteration phase, it becomes possible to select and / or edit and / or adjust the variations of the attributes of the target object. Furthermore, the modification of the variation of the attribute is provided back to the target object constructor 105, thereby generating another iteration of the attribute (i.e., a new set of attributes). Each of these attributes includes one or more variations, and these variations are also generated using generative AI. For example, a plurality of arrays of attributes 400 are generated in different iterations, including array 400A in the first iteration, array 400B in the second iteration, and array 400N in the Nth and final iteration.
[0034] For example, there may be an iterative interface 140 configured to facilitate user interactions for the purpose of selecting, editing, and / or adjusting variations of attributes. The iterative interface 140 includes a user interface 145 configured to enable interaction by the user. Different functions may be provided by the user interface. Such functions include a selection interface 146 configured for selection of attributes and / or variations of corresponding attributes, an adjustment interface 147 configured for adjustment of attributes and / or variations of corresponding attributes, and a user response interface 148 configured to enable the user to perform desired actions by the target object constructor 105, such as executing another iteration or generating one or more permutations of the target object based on the selected attributes and / or selected variations of the corresponding attributes.
[0035] Therefore, each variation of a component / attribute can be made unique in its arrangement (e.g., an array) and / or made selectable such that it generates different permutations of the target object in the merge blend phase 150. In this way, using generative AI and / or an AI model, target objects / assets with different permutations or versions can be created. For example, at the end of an iteration, a merge and / or blend phase 150 is performed to generate one or more permutations and / or variations of the target object 160. Thus, the target object can include a first variation or permutation, a second variation or permutation, up to an Nth variation or permutation. The storage 190 can be configured to store each of the permutations of the target object, as well as and / or the representations of the target object generated during processing and / or the attributes of the iteration, and / or each of the variations of the representations of the target object and / or the attributes of the iteration.
[0036] Figure 1B provides more detailed information about the system 100 as introduced in Figure 1A, and will explain in more detail the multiple phases of asset creation executed by the target object construction unit 101 according to an embodiment of the present disclosure. To create visual assets using generative AI, the generative AI is executed by the system 100. Here, one or more versions of the target object or asset may be iteratively generated through the selection and editing of the AI-generated attributes of the target object. For example, the system 100 is configured to dynamically generate original and visual available content and can be applied to game elements and other visual elements (e.g., game characters, use of animations, visual elements / assets) within video games or other applications.
[0037] A custom prompt is generated in the input phase 110 based on the original input 101 provided to the target object construction unit 105. As previously explained, the original input may be provided by the user and describes the target object and the desired characteristics of the target object. For example, the user may provide the original input using any communication means such as text, audio, photos, etc. For illustrative purposes only, the original input may be directed to a desired target object. This target object may be a dragon. The original input may provide details regarding the characteristics of the dragon, such as the overall pose of the dragon, desirable facial features, and other body features. The original input may also include an artistic style, such as a dragon influenced by the Far East or a dragon influenced by Europe. More specifically, the prompt generation unit 115 receives the original input 101 and generates a custom prompt in a format suitable for use by the target object construction unit 105 that implements the generation AI service.
[0038] In some embodiments, the prompt generation unit 115 is incorporated into a generation AI system such as IGAI121. For example, IGAI121 can be customized to accept a unique description language statement for setting the style regarding the requested output image or content. The description language statement can be text or other sensor inputs, such as inertial sensor data, input speed, emphasis statements, and other data that can be compiled during the input request. Images, videos, or a set of images can also be provided to IGAI to define the context of the input request. In one embodiment, the input can be text describing the desired output along with one or more images for conveying the scene of the desired context requested as the output.
[0039] The custom prompt is used in the decomposition phase 120 configured for attribute generation, and the target object includes one or more attributes. The custom prompt can be provided to an IGAI 121 configured to perform generative AI to generate one or more output images, graphics, and / or 3D representations of the target object. The IGAI may include one or more artificial intelligence processing engines and / or models that are trained and / or curated for a particular desired output. In some cases, the training dataset can include a wide range of general-purpose data that is available from a number of sources on the Internet. By way of example, the IGAI needs to access large amounts of data such as images, videos, and 3D data. The general-purpose data is used by the IGAI to gain an understanding of the type of content desired by the input. For example, if the input requests the generation of a dragon, the dataset needs to have various images of dragons to access and draw during the processing of the output image. On the other hand, the curated dataset may be specialized for a type of content, such as, for example, video game-related art, videos, and other asset-related content. More specifically, the curated dataset can include images related to a particular scene of the game, or action sequences including game assets such as, for example, unique avatar characters.
[0040] In one embodiment, an IGAI 121 is provided to enable generation from text to image. The image generation is configured to perform a latent diffusion process for synthesizing text into image processing in a latent space. In one embodiment, the conditioning process helps to shape the output towards a desired usage output, for example, using structured metadata. The structured metadata may include information obtained from user input to guide a machine learning model to sequentially and progressively denoise using cross-attention until the finally processed denoised one is decoded back into the pixel space. In the decoding stage, upscaling is applied to achieve higher quality images, videos or 3D assets. Thus, IGAI is a custom tool designed to process a specific type of input and render a specific type of output. When IGAI is customized, machine learning and deep learning algorithms are adjusted to achieve a specific custom output, such as in game technology, unique image assets used in a specific game title and / or movie.
[0041] In another configuration, the IGAI 121 can be a third-party processor, such as a processor provided by Stable Difffusion, or other processors such as OpenAI's GLIDE, DALL-E, MidJourney or Imagen. In some configurations, the IGAI can be used online via one or more application programming interface (API) calls. It should be understood that the references to the available IGAI are for information reference only. For additional information related to IGAI technology, reference may be made to the paper "High-Resolution Image Synthesis with Latent Diffusion Models" by Robin Rombach, et al., published by Ludwig Maximilian University of Munich, pp. 1-45. This paper is incorporated by reference.
[0042] In particular, the IGAI 121 generates, for example, one or more representations and / or versions of the target object in each iteration of constructing the target object. The respective latent space representations 123 of those representations / versions of the target object are stored in the cache 122. In that way, each iteration of the target object can be constructed based on the previous iteration of the target object by using the latent space representation or a part thereof.
[0043] The decomposition phase 120 includes the decomposition 125 of attributes. Here, additional AI is implemented (e.g., via one or more AI models) to identify the attributes of one or more output images previously generated for the target object by the IGAI 121. For example, the AI model 126 is configured to classify and / or identify the respective attributes of the generated representations of the target object by using artificial intelligence and / or deep / machine learning. The AI model 126 may be assisted by a target object classification unit 127 configured to identify a general class or type of the target object. The general class may include a basic set of attributes. For example, the target object may be classified as a dragon, and the basic set of attributes for the dragon may include a head, a body, wings, arms and legs, and a tail.
[0044] Therefore, the attribute construction unit 129 is configured to construct a detailed set of attributes identified from the representations of the target object generated through the IGAI 121, and this detailed set may include more attributes than the basic set. That is, the attribute construction unit may arrange different variations (e.g., one or more features or characteristics or samples, etc.) for each of the components / attributes of the target object in the array. The array may be constructed for each iteration of constructing the target object.
[0045] One example of the array is provided in FIG. 4A. This figure shows an array 400N of attributes of a target object (e.g., a dragon) where each attribute includes one or more variations, according to one embodiment of the present disclosure. In particular, the array includes a column 401 of one or more attributes and a column showing one or more variations 405 of each of the attributes. For example, the array 400N may include attributes 1 to N. In this array, each attribute may include one or more variations. For example, attribute 1 includes variations 1 to N (i.e., v1, v2,... vN). The legend 410 shows various user interactions, including a "like" interaction, a "dislike" interaction, and a "lock" interaction. Also, other interactions may be supported.
[0046] Furthermore, the locked attribute incorporation engine 128 is used to generate the next iteration of the attributes. In particular, within the array of attributes and their variations, the user may lock the desired attributes and / or the corresponding variations of the attributes. In so doing, the locked variations of the corresponding attributes are provided again as input to the IGAI 121 in the next iteration of constructing the target object. As a result, each newly constructed version and / or representation of the target object will include the locked variations of the corresponding attributes, or at least a version of the locked variations that matches the other attributes of the corresponding representation of the target object.
[0047] As previously described, the iteration phase 130 is configured such that the selection, editing, and / or adjustment of one or more variations of one or more attributes of the target object is possible (e.g., in each iteration). The iteration phase 130 may include a filter engine 131, a selection engine 132, an adjustment engine, and a prompt generation unit 137 configured to generate a prompt used in the next iteration of constructing one or more representations of the target object.
[0048] The filter engine 131 may be implemented during the decomposition phase 120, such as during the generation of the array, and / or during the iteration phase 130. The filter engine 131 is configured to exclude attributes and / or any variations of the corresponding attributes based on defined parameters. For example, the filter engine 131 may exclude unpleasant variations of an attribute, or variations that are inconsistent with a particular context (e.g., a cheerful atmosphere for a dragon that is desired to be terrifying and intimidating).
[0049] The selection engine 132 is configured to generate and provide one or more interfaces for the user to select, edit, and / or adjust each of one or more variations of one or more attributes of the target object. In that way, the variations of the corresponding attributes can be edited and / or manipulated via the selection engine 132. The one or more interfaces generated by the selection engine 132 are presented via an iteration interface 140 configured to facilitate user interaction. In particular, the user interface 146 enables user interaction and includes a selection part interface 146 that operates in cooperation with the selection engine 132 and an adjustment part interface 147 that operates in cooperation with the adjustment engine 135.
[0050] In particular, the selection engine 132 is configured to generate an attribute highlight display section 133. For example, one or more attributes and / or variations of the corresponding attributes may be highlighted for further interaction by the user via, for example, the selection interface 146. The highlighting may be performed to draw the user's attention to specific attributes of the user and / or variations of those attributes, or may be performed in response to a user interaction indicating that the user desires to view and / or interact with an attribute or variation of an attribute. For purposes of illustration, in the interface including the array 400N shown in FIG. 4A, a variation N of attribute 3 is highlighted within block 420. This may enable further interaction by the user.
[0051] For purposes of explanation, the auto-rotation interface 133C is configured to automatically present one or more attributes and their variations to the user via, for example, the selection interface 146. That is, the user is presented with images of possible iterations of the target object on the display. The target object includes one or more attributes, each with a variation that has been selected either by the user or automatically. During auto-rotation, the variations of one or more attributes may rotate and switch. In that way, multiple iterations of the target object may be shown to the user. In one embodiment, the user may select which attributes and / or their variations to rotate, thereby enabling the user to quickly view the selected items.
[0052] Also, for the purpose of explanation, the manual interface 133A is configured to enable a user to select one or more attributes and their variations, and may be implemented via the selection interface 146. For example, an array including the attributes and their variations may be presented to the user, such that the user can select and visually identify one or more variations of the corresponding attributes. That is, a plurality of iterations of the target object, each including a different set of variations of the attributes, may be shown to the user. For example, FIG. 4C provides an example of the attributes and their corresponding variations used for a particular iteration of the target object that may be presented to the user, as further described below.
[0053] Further, for the purpose of explanation, the mix-and-match interface 133B is configured to enable a user to select and visually identify a desired iteration of the target object, and may be implemented via the selection interface 146. For example, when the user has narrowed down the selection of attributes and / or variations of the corresponding attributes to a manageable number, the mix-and-match interface 133B may present an overview of those attributes and their variations (e.g., via thumbnails) to the user, and may present a function for visually identifying different permutations or versions of the target object through the selection of a particular combination of variations of the attributes. That is, a dragon may be shown with a first set of attributes, each with the selected variation. The user can select a variation and replace that variation with another variation, and the result is immediately reflected in the displayed dragon. The user may continue to mix and match between variations of the attributes in order to quickly visually identify multiple iterations of the target object.
[0054] Furthermore, the selection engine 132 includes a user preference indicator 134. This enables a user to indicate whether corresponding variations of corresponding attributes are preferred or not, etc., via, for example, the selection interface 146. For example, information regarding user preferences for one or more variations of corresponding attributes may be used to provide the user with one or more sample iterations of a target object (e.g., a dragon). In another example, user preferences may be used in a next iteration process to generate a next iteration of an array of attributes and their variations. That is, the user preference for each of the variations of the corresponding attributes may be used for a next iteration of the process for creating the target object.
[0055] For the purpose of explanation, the like / dislike function / interface 134A is configured to enable a user to provide "like" and / or "dislike" for one or more variations of the corresponding attribute via a selection interface 146 or the like. For example, this interface may provide an array of attributes and their variations for user interaction. Within the array, the user can select a variation of the corresponding attribute and provide a "like" or "dislike" indication (e.g., a check mark for "like" or an "x mark" for dislike). This may be performed one or more times for one or more attributes and / or one or more variations of the corresponding attribute. For the purpose of explanation, in the interface including the array 400N shown in FIG. 4A, a plurality of variations of the corresponding attribute are selected and user preferences are given. For example, variations v1, v2, and vN of attribute 1 are selected by the user and marked as "liked" as indicated by the corresponding check marks. Also, variation 3 (v3) is selected by the user and marked as "disliked" as indicated by the corresponding x mark. Other user preferences are also shown for one or more attributes shown in the array (e.g., attribute N has a "like" indication for variations v1 and v2 and a "dislike" indication for variations v4 and v5). In addition, the attribute and / or the variation of the corresponding attribute may have no user preference as shown for variations v2, v3, v5, and vN of attribute 5.
[0056] For purposes of explanation, the deletion function / interface 134C of the selection engine 132 is configured to enable a user to indicate a strong dislike for an attribute and / or a variation of a corresponding attribute via, for example, the selection unit interface 146. For example, the interface shown in FIG. 4A may provide an array 400N of attributes and their iterations, and further may enable the user to select and delete a variation of a corresponding attribute. This may be performed one or more times for one or more attributes and / or one or more variations of corresponding attributes. For example, variation 5 of attribute 4 is selected and deleted.
[0057] For purposes of explanation, the lock selection unit function / interface 134B of the selection engine 132 is configured to enable a user to lock an attribute and / or a variation of a corresponding attribute via, for example, the selection unit interface 146. In one embodiment, an attribute may have only one variation that is locked. In other embodiments, an attribute may have one or more variations that are lockable. For example, the interface shown in FIG. 4A may provide an array 400N of attributes and their iterations, and further may enable the user to select and lock a variation of a corresponding attribute. This may be performed multiple times for multiple attributes and / or variations of corresponding attributes. For example, variation 2 of attribute 2 is locked, variation 5 of attribute 3 is locked, and variation 1 of attribute 4 is locked.
[0058] In one embodiment, the user may choose to visually inspect iterations or permutations of the target object. In so doing, the user can determine whether to continue with the approach for creating the target object or move to a different approach. For purposes of explanation, the manual interface 133A may be configured to enable the user to select one or more attributes and their variations, and may be implemented via the selection interface 146. For example, an array including the attributes and their variations may be presented to the user, whereby the user can select and visually inspect one or more variations of the corresponding attribute. That is, multiple iterations of the target object, each including a different set of variations of the attribute, may be shown to the user. Additionally, one or more iterations may be automatically presented for visual inspection based on one or more of the variations of the corresponding attribute preferentially selected and / or edited by the user. The iterations may be selected for visual inspection via the iteration viewing interface 148A of the user response interface 148.
[0059] For example, FIG. 4C shows a visual presentation of one or more variations of one or more attributes for an iterative view 440 of a target object (e.g., a dragon) according to one embodiment of the present disclosure, and the variations of the attributes can be selected by a user via an iterative view interface 148A or the like. The attributes and their variations may follow an array 400N of target objects. That is, the target object includes N attributes, each with a corresponding variation selected by and / or affected by a user selection. In particular, the iterative view 440 includes locked attributes. For example, the iterative view 440 includes variation 2 of attribute 2, variation 5 of attribute 3, and variation 1 of attribute 4, which have been locked by the user as previously described in array 400N. In addition, the remaining attributes for the iterative view 440 include variation 1 of attribute 1, variation 1 of attribute 5, and variation 2 of attribute N. The remaining attributes for the target object shown in the iterative view 440 can be selected by the user or affected by the user's selection. That is, the user may actively select and view the variations of the corresponding attributes. Also, for attributes that the user has not actively selected to view in the iteration 440, the variations of the corresponding attributes may be automatically selected based on user preferences. For example, assume that the user has not locked the variations of attribute 1 but likes a plurality of variations (e.g., variations 1, 2, 5,..., and N) and has deleted one variation (e.g., variation 5). Since the user has not selected the variation of attribute 1 for the iterative view 440, a variation may be automatically selected (e.g., variation 1). That is, a preferred variation for attribute 1, such as variation 1 of attribute 1 that matches the array 400N (e.g., "liked", excluding the deleted variation), may be randomly selected for the iterative view 440.
[0060] For purposes of explanation, the voting system function / interface 134D of the selection engine 132 is configured to enable a user to vote on attributes and / or variations of corresponding attributes in a group setting via, for example, the selection interface 146. For purposes of illustration only, attributes and / or variations of corresponding attributes that have been “liked” by the corresponding user may correspond to votes on those “liked” selections, such as those shown in the array 400N of FIG. 4A. For example, other ways of selecting attributes and / or variations of corresponding attributes may be supported, such as through annotations. In this way, once a vote is taken by members of a group (e.g., a design group dealing with characters for a video game), the preferences of that group can be determined based on the popularity of attributes and / or variations of corresponding attributes. One or more versions of a target object can be constructed and viewed, each version including a unique set of one or more variations of corresponding attributes (e.g., variations for each attribute). Additionally, one or more versions of the target object constructed based on group preferences may be presented via the previously described mix-and-match function / interface 133B.
[0061] The adjustment engine 135 is configured to generate and provide one or more interfaces for enabling user editing for each of one or more variations of one or more attributes of a target object. The one or more interfaces generated by the adjustment engine 135 are presented via an adjustment interface 147 configured for user interaction.
[0062] For example, the adjustment engine 135 includes an attribute adjustment unit 136. This attribute adjustment unit is configured to enable editing by the user via an adjustment unit interface 147. In one embodiment, the attribute adjustment unit 136 includes a slider 136A configured to provide selection and adjustment of variations of the corresponding attribute. For example, the selection of a variation of the corresponding attribute, such as variation N of attribute 3 shown within the highlighted block 420 for the purpose of further user interaction, may be performed via an interface showing the array 400N of FIG. 4A. For illustrative purposes only, FIG. 4B illustrates editing the variation of an attribute of a target object using a slider 430 in accordance with one embodiment of the present disclosure. As shown, the object 425 is a representation of variation N of attribute 3. The slider 430 enables the user to modify variation N of attribute 3, where the modification is reflected in the visible object 425. For example, the slider 430 may enable modification of one or more parameters of variation N of attribute 3. For illustrative purposes, the modification may include making the object 425 thinner or thicker in the vertical direction, making it smaller or larger uniformly, making it darker or lighter, and so on.
[0063] A plurality of methods for modifying the variation of the corresponding attribute are supported. For example, after selection of the variation of the corresponding attribute, modification of the aspect of the variation of the corresponding attribute may be performed using a narration / text input modification unit 136B. That is, the user may provide an annotation describing the desired modification to the variation of the corresponding attribute, such as a modification to the object 425 shown within the block 420 highlighting variation N of attribute 3 (i.e., via text or narration). In one embodiment, the modification to the variation of the corresponding attribute may be performed by an AI model, such as a model implementing generative AI. Thus, the annotation may be encoded (e.g., using an encoder) into a text prompt supported by the AI model.
[0064] As shown, the iterative interface 140 includes a user response interface 148 configured to enable selection of one or more actions to be performed by the system 100 when creating a target object, which includes a plurality of interfaces, namely, an iterative view 148A, a selection return 148B, a reiteration 148C, and an approval 148D.
[0065] In particular, the iterative view interface 148A provides a visual representation of the target object including attributes and their corresponding attributes. That is, a representation of the target object (e.g., a dragon) is generated and displayed, and the representation includes one or more attributes and their corresponding variations. The interface 148A may also enable the user to select attributes and their corresponding variations. For example, the attributes and their corresponding variations may be selected by the user or influenced by user preferences as shown in the corresponding iterative view (e.g., iterative view 440 of FIG. 4C).
[0066] Also, the selection return interface 148B enables the user to return to the iterative phase 130 of the system 100. In particular, the user may select to return to the selection engine 132 to enable selection, deletion, and / or editing of attributes and / or variations of attributes, such as may be possible via, for example, the attribute highlighting section 133 and / or the user preference indicator section.
[0067] Furthermore, the re-iteration interface 148C enables the user to initiate another iteration of the target object. Specifically, user actions including selection of attributes and / or variations of attributes (e.g., lock, unselected, disinterested, etc.), deletion and / or editing are used to generate another set of attributes and one or more variations for each of those attributes. That is, user actions regarding attributes and / or variations of attributes are provided again to the IGAI engine 121 to generate another iteration of one or more representations of the target object. In particular, the next iteration prompt generation unit 137 is configured to consider user actions regarding attributes and / or variations of attributes received from the previous iteration and generate a prompt suitable for input to the IGAI engine 121 to generate one or more representations of the target for the next iteration. Therefore, another iteration of the array of attributes and their corresponding variations (i.e., a new set of attributes and their variations) is generated using the generative AI (e.g., one of the arrays 400A - 400N in FIG. 1A).
[0068] Also, the approval interface 148D is configured to generate a final output for the target object. Specifically, when various selections of attributes and / or corresponding variations of attributes (e.g., lock, unselected, ignore, etc.), deletion and / or editing are approved by the user, the next phase of creation of the target object is initiated. Therefore, the merge / blend phase 150 is executed after user approval. For example, the final version 160 of the target object may include one or more representations or variations such as representation / variation 1, representation / variation 2, …, representation / variation N. Each of these representations / variations may be presented for the user to select, for example, to incorporate into a video game.
[0069] The save engine 151 is configured to save each of the attributes and their corresponding variations. One or more of the corresponding variations of the attributes may be shared with other users via the shared engine 152. Additionally, one or more variations of the corresponding attributes may be exported via the export engine 153 to other parties within or to dedicated services and / or applications, or to third-party services and / or applications. For example, the attributes and their corresponding attributes used to construct one or more representations of the final version of the target object may be used by other services. In that way, the representation of the final version of the target object may be used by these other services.
[0070] FIG. 2A is a general representation of an image generation AI (IGAI) processing sequence, such as implemented by an IGAI processing engine 121 that implements, for example, generative AI, according to one embodiment. As shown, input 206 is configured to receive input in the form of data, such as a semantic description or a text description with keywords. The text description can be in the form of a sentence having, for example, at least a noun and a verb. The text description can also be in the form of a fragment or a simple single word. The text can also be in the form of multiple sentences describing a scene, or an action, or a characteristic. In one configuration, the input text can be entered in a specific order such that one word acts more prominently than other words, or even without emphasizing words, characters, or sentences. Furthermore, the text input can be in any form including characters, emoticons, ions, foreign language characters (e.g., Japanese, Chinese, Korean, etc.). In one embodiment, the text description is enabled by contrastive learning. The basic idea is to map the text corresponding to an image to a region within the same latent space as that image by embedding both the image and the text into the latent space. Thereby, a structure that means being a dog, for example, is abstracted from both the visual representation and the text representation. In one embodiment, the goal of contrastive representation learning is to learn the embedding space such that similar sample pairs are close to each other while dissimilar sample pairs are far apart in that space. Contrastive learning can be applied to both supervised and unsupervised settings. When dealing with unsupervised data, contrastive learning is one of the most powerful approaches in self-supervised learning.
[0071] In addition to text, the input can also include other content, such as an image, or even an image having descriptive content itself. The image can be interpreted using image analysis to identify objects, colors, intentions, characteristics, shadows, textures, three-dimensional representations, depth data, and combinations thereof. Broadly speaking, the input 206 is configured to convey the intention of a user who desires to generate certain digital content using IGAI. From the perspective of game technology, the generated target content can be a game asset for use in a specific game scene. In such a scenario, by using the dataset used to train the IGAI 121 and the input 206, the way in which an artificial intelligence, such as a deep neural network, processes data to manipulate and adjust the desired output image, data, or three-dimensional digital asset can be customized.
[0072] Input 206 is then passed to IGAI 121, where encoder 208 captures the input data and / or pixel space data and converts it into latent space data. The concept of "latent space" is at the core of deep learning. The reason is that for the purpose of discovering patterns and using those patterns, feature data is reduced to a simplified data representation. Thus, latent space processing 210 is performed on the compressed data. This results in a significant reduction in processing overhead compared to the processing of learning algorithms in pixel space, which is more resource-intensive and requires greater processing power and time to analyze and produce the desired image. The latent space is simply a representation of the compressed data, where similar data points are closer together in the space. In the latent space, the processing is configured such that the machine learning system learns the relationships between the learned data points that can be derived from the supplied information, e.g., the dataset used to train IGAI. In latent space processing 210, a diffusion model is used to compute the diffusion process. The latent diffusion model relies on an autoencoder that learns a lower-dimensional representation of the pixel space. The latent representation is passed through a diffusion process that adds noise at each step, e.g., in multiple stages. The output is then fed into a denoising network based on the U-Net architecture with a cross-attention layer. A conditioning process is also applied to remove the noise and guide the machine learning model to reach an image with a representation close to what was requested via user input. Decoder 212 then converts the result output from the latent space back to pixel space. Output 214 can then be processed to improve the resolution. Output 214 is then delivered as a result. This may be an image, graphics, 3D data, or data that can be rendered in physical or digital form.
[0073] Figure 2B shows, in one embodiment, additional processing that can be performed on input 206. User interface tool 220 can be used to enable a user to provide input request 204. Input request 204 can be an image, text, structured text, or general-purpose data, as described above. In one embodiment, before the input request is provided to encoder 208, the input can be processed by a machine learning process that generates machine learning model 232 and learns from training data set 234. By way of example, the input data can be processed via context analysis unit 226 to understand the context of the request. For example, if the input is "a space rocket for flying to Mars", the input can be analyzed to determine that the context is related to space and planets (226). Context analysis can use machine learning model 232 and training data set 234 to find an image related to this context or to identify a particular library of art, images, or videos. If the input request also includes an image of a rocket, feature extraction unit 228 can function to automatically identify characteristics of features within the rocket image, such as fuel tanks, length, color, position, edges, lettering, flames, and the like. Also, feature classification unit 230 can be used to classify the features and improve machine learning model 232. In one embodiment, input data 207 can be generated to produce structured information that can be encoded by encoder 208 into a latent space. Additionally, it is possible to extract structured metadata 222 from the input request. Structured metadata 222 can be, for example, descriptive text used to instruct, e.g., IGAI121, to make modifications to properties, or changes to the input image, or changes to color, texture, or combinations thereof. For example, input request 204 can include an image of a rocket, and the text can state "make the rocket wider" or "add more flames" or "make the rocket more robust", or some other modifier intended by the user (e.g., semantically provided and contextually analyzed).The structured metadata 222 can then be used in subsequent latent space processing to adjust the output to approximate the user's intent. In one embodiment, the structured metadata may be in a data format designed to represent a semantic map, text, an image, or the user's intent regarding what changes or modifications should be made to the input image or content.
[0074] Figure 2C shows a system in which the output of the encoder 208 is then fed into the latent space processing 210, in accordance with one embodiment. The diffusion process is carried out by the diffusion process stage 240. Here, the input is processed through a plurality of stages to add noise to the input image or an image related to the input text. This is a step-by-step process in which noise is added at each stage, for example, 10 to 50 or more stages. Next, the noise removal process is carried out through the noise removal stage 242. Similar to the noise stage, a reverse process is carried out in which noise is removed step by step at each stage, and at each stage, machine learning is used to predict how the output image or content should be in light of the input requirements. In one embodiment, the structured metadata 222 can be used by the machine learning model 244 at each stage of noise removal to predict how the resulting denoised image should look and how it should be corrected. During these predictions, the machine learning model 244 uses the training dataset 246 and the structured metadata 222 to get closer and closer to the output that is most similar to what was requested in the input. In one embodiment, a U-Net architecture with cross-attention layers can be used during noise removal to improve the predictions. After the final noise removal stage, the output is provided to the decoder 212, which converts the output into pixel space. In one embodiment, the output is also upscaled to improve the resolution. The output of the decoder can, in one embodiment, optionally be run through the context conditioning unit 236. The context conditioning unit is a process that can use machine learning to examine the resulting output and make adjustments to make the output more realistic or to remove unrealistic or unnatural outputs. For example, if the input requests a "teenager pushing a lawnmower" and the output shows a teenager with three legs, the context conditioning unit can make adjustments by inpainting or overlaying to correct or block the conflicting or undesirable output.However, as the machine learning model 244 becomes smarter with more training over time, it will require less context conditioning 23 before the output is rendered within the user interface tool 220.
[0075] In light of the detailed description of the system 100 of FIGS. 1A - 1B, flowchart 300 of FIG. 3 discloses a method for creating an asset or target object using generative artificial intelligence, according to one embodiment of the present disclosure. In particular, the operations performed in the flowchart can be implemented by one or more of the previously described components through sections and editing of the attributes of the target object and their corresponding variations. In some embodiments, the method of flowchart 300 enables the dynamic generation of original and visually available content and can be applied to game elements and other visual elements (e.g., game characters, use of animations, visual elements / assets) within a video game or other application.
[0076] At 310, the method includes collecting one or more inputs, each of which describes a target object. This input can include text, annotations, images, etc. By using the collected input, a custom prompt can be generated to direct the generative AI system to create an asset or target object, such as a character used within a video game. More specifically, the custom prompt is generated in a format suitable for use by the generative AI system to perform one or more iterations of creating the target object.
[0077] At 320, the method includes generating a plurality of images of a target object using an image generation artificial intelligence system configured to perform latent diffusion based on one or more collected inputs. In particular, the generative AI system is configured to generate a plurality of images of the target object based on the collected inputs. For example, instead of outputting one image or representation, the previously generated prompt is input into the generative AI system to generate a plurality of images and / or representations of the target object. Those representations are used to generate attributes of the target object.
[0078] At 330, the method includes decomposing the target object into a plurality of attributes, each including one or more variations, based on the plurality of images of the target object. In particular, the plurality of images / representations of the target object are input into an AI model configured to extract a plurality of attributes. More specifically, each of the representations of the target object output by the generative AI system includes attributes. For example, a dragon representative of the target object being created may include attributes such as a tail, a face, a mouth capable of breathing fire, a belly, arms, wings, and the like. Thus, the plurality of representations of the target object may include a set of similar attributes (e.g., in that case, the same attributes are included in each set), or a set of slightly different attributes (e.g., in that case, each set includes a basic set of attributes and may include one or more additional attributes that are unique to the corresponding representation).
[0079] The AI model is configured to identify a plurality of attributes and their variations based on a plurality of images of the target object to be generated. That is, the AI model is configured to classify one or more attributes of the representation of the target object. Here, one attribute may include a plurality of variations based on different representations. For example, the AI model may classify the tail as an attribute of the target object with one or more variations of the tail. In particular, relevant features useful for classifying the attributes of the target object are extracted from the image. Further, based on the extracted features, the AI model applies machine learning and / or deep learning to classify the attributes of various representations of the target object. Here, machine learning is a subclass of artificial intelligence, and deep learning is a subclass of machine learning. As previously explained, the attributes and their variations may be arranged within an array of attributes (e.g., array 400N of FIG. 4A).
[0080] For purely illustrative purposes, the AI model 126 implementing deep learning / machine learning may be configured as a neural network. Generally, a neural network represents a network of interconnected nodes that respond to inputs (e.g., extracted features) and generate outputs (e.g., classify or identify or predict the intent of an executed gesture). In one embodiment, the AI neural network includes a hierarchy of nodes. For example, there may be an input layer of nodes, an output layer of nodes, and an intermediate or hidden layer of nodes. The input nodes are interconnected to hidden nodes within the hidden layer, and the hidden nodes are interconnected to output nodes. The interconnections between nodes may have numerical weights that can be used to link multiple nodes to each other between the input and the output, such as when defining the rules of the AI model 160. More specifically, the AI model 126 of FIG. 1B is configured to apply rules (e.g., a length corresponding to a particular tail attribute) that define the relationship between features and outputs, and the features may be defined within one or more nodes located at one or more hierarchical levels of the AI model 126. By this rule, features (as defined by the nodes) are linked between the layers of the hierarchy so that a given input set of data is associated with a particular output (e.g., an attribute classification) of the AI model. For example, the rules may link one or more features or nodes between the input and the output (e.g., using relationship parameters including weights) across the AI model 160 (e.g., at the hierarchical level), such that one or more features result in rules being learned through the training of the AI model 160. That is, each feature may be linked to one or more features of other layers, in which case one or more relationship parameters (e.g., weights) define the interconnections between features of other layers of the AI model 160. Therefore, each rule or set of rules corresponds to a classified output. In that way, the output obtained according to the rules of the AI model 126 may classify and / or label and / or identify and / or predict the attributes of the target object.
[0081] At 340, the method includes receiving a selection of one or more of a plurality of variations of a plurality of attributes. In so doing, the user can select, edit, and / or adjust each of one or more variations of each of the previously generated attributes. As previously described, by selecting an attribute and / or a variation of an attribute, the user can further modify and / or indicate preferences for that component. For example, a particular variation of a corresponding attribute may be selected by the user (e.g., via an interface). The user may indicate a preference by positively selecting a variation of that attribute (e.g., indicating a "like" preference), or negatively selecting a variation (e.g., indicating a "dislike" preference). Also, the user may indicate a positive preference by locking a variation of an attribute. Further, the user may indicate a negative preference by completely deleting a variation. During adjustment, a variation of a corresponding attribute can be further modified through user interaction. For purposes of illustration, the user may provide an edit input (e.g., a text command, movement of a slider, etc.), and when that input is acted upon, the variation of the attribute is adjusted to be included in a list of variations for that corresponding attribute.
[0082] In one embodiment, attributes can be automatically filtered. For example, variations of the corresponding attributes can be filtered based on at least one filtering parameter. The filtering parameter may be automatically generated or manually set by the user. The filter may be automatically applied or specified by the user. Thus, one or more variations of the corresponding attributes can be filtered through modification and / or deletion. For example, to avoid the use of unfavorable materials in creating the target object or the use of materials that affect the creation of the target object, the attributes and / or variations of the corresponding attributes may be filtered (e.g., deleted).
[0083] In another example, the user may prevent the modification of their own character in an unpleasant form (e.g., modifying the character to have an unpleasant tattoo) through filtering. In one embodiment, filtering may be enabled by presenting the user with a small toolset for modifying attributes and their variations, such as when modifying a character or NPC in a video game. This small toolset may be simplified compared to the tools presented to developers. Thus, the functions for generating and / or modifying characters are limited.
[0084] In yet another embodiment, a user acting as an administrator controls the generation of a target object (e.g., a character for a video game). For example, the character can be created by one or more developers. In that case, the administrator can introduce filters that restrict the use of AI tools used to generate and / or edit one or more attributes and their variations. These filters may be implemented automatically. In this way, the administrator can guide the development of the target object and, for illustrative purposes only, maximize the prevention of the development of video game characters that may be offensive, or undesirable (e.g., inappropriately so, targeted at the context of an adult game, racially discriminatory characters, war criminals, etc.).
[0085] In one embodiment, the creation of the target object may be performed collaboratively by a group of developers. Each developer may act independently to generate different versions of the target object. During collaboration, different versions may be simultaneously displayed within the interface, such as by displaying the versions within one or more sandboxes. The sandboxes are displayed simultaneously, showing side by side the development status of the target object by various designers. In that way, features within a design (e.g., within a particular sandbox) can be incorporated into another design (within another sandbox) in a manner similar to the previously described mix-and-match functionality. Thus, the version agreed upon by the group of designers can be created using the agreed-upon attributes. The final version may include similar attributes corresponding to features found within each version provided by each designer, and further unique attributes corresponding to unique features that may be found in a particular version of the corresponding designer. In one embodiment, the sandbox is implemented through the use of a shared spreadsheet. For example, the spreadsheet may include one or more attributes and their corresponding variations. In other embodiments, designers may label those versions of the target object, and this label may limit the degree of control that other developers may have to edit the corresponding attributes and / or variations of the corresponding attributes. Thus, the target object can be dynamically created in real time using a generation process by multiple developers.
[0086] In some embodiments, the prompts used by the designers may also be displayed. In that way, other designers may provide comments to further refine one or more of the prompts used in another iteration of target object generation.
[0087] At 350, the method includes blending one or more of a plurality of variations of a plurality of selected attributes into one or more options of a target object. In particular, one or more iterations of creating the target object may be performed. For example, based on user interactions (such as selection, editing, modification input, input, etc.) with the attributes and their variations, the IGAI system may be tasked to generate a second plurality of attributes for the target object using the previously described process. For example, the IGAI system may generate one or more versions and / or representations of the target object based on the selected and / or edited attributes and / or corresponding attribute variations. These representations of the target object are used to generate a second plurality of attributes and their variations for the current iteration of the target object. This process may be continuously repeated in a series of iterations of developing the target object until a final iteration is performed that outputs a final version of the target object that includes one or more options of the final version.
[0088] When constructing the final version of the target object, the last iteration of user interactions (such as selection, editing, modification input, input, etc.) with the attributes and their variations is considered. In particular, one or more options of the final version are generated based on the attributes and their variations. That is, each option includes most, if not all, of the attributes, and each option includes a unique set of variations for those attributes. For example, the first option may include variation 1 of attribute 1, while the second option may include variation 2 of attribute 1, and so on for each attribute. For each of the different options, the unique sets of variations of the corresponding attributes are blended with each other. That is, the generative AI is not used for the blending to generate the corresponding options of the final version of the target object.
[0089] In one embodiment, it is determined that the variations of the attributes are not preferentially selected. In that case, the selected variations of the corresponding attributes are automatically selected for use when performing a blend of the selected variations of the corresponding attributes used to generate the corresponding options for the final version of the target object. The selection of the variations can be performed randomly or in some predefined order.
[0090] In one embodiment, an array including the variations of the corresponding attributes is saved. In this way, the final version and their options are saved and can be exported for use in other services or applications. Additionally, the attributes and the corresponding variations may be saved and exported for use in other services or applications. This can shorten the development time for other target objects, such as the same video game or other characters within other video games.
[0091] FIG. 5 is a flow diagram 500 showing the flow of data for generating one or more options for the final version of a target object over one or more iterations of a process according to one embodiment of the present disclosure. The operations performed in the flow diagram may be implemented by one or more entities of the components previously described, and further by the system 100 described in FIGS. 1A - 1B. The process shown in FIG. 5 is intended to exemplify one way, but not to limit, for performing the next iteration of generating the target object.
[0092] In particular, the latent diffusion method is used to generate one or more representations 570 of the target object for an iteration of the entire process (e.g., the previous iteration) used to generate the target object. For example, one or more representations 570 of the target object can be provided as the output (e.g., output image) by the IGAI processing model that implements the generation AI, such as during the previous iteration of the entire process. During the next iteration of the entire process, a new iteration of one of the representations 570 of the target object (e.g., the selected output image 570x) is generated taking into account user preferences, such as a locked attribute that the user absolutely wants to retain in the final version of the target object to be created or a locked variation of the corresponding attribute. During the next iteration, latent diffusion can be performed on one or more of the representations 570 of the target object generated during the previous iteration of the entire process to generate a new set of representations of the target object, which new set may include the same number, fewer number, or greater number of representations as the representations provided by the representation 570 in the previous iteration.
[0093] As described previously, latent diffusion is a process that adds and removes noise to generate an image (e.g., the output image of a target object for the corresponding iteration of the process used to create the target object). For example, a desired image (e.g., a target object including one or more options of the final version of the target object) can be generated from a noise patch concatenated with a vector (e.g., text encoded in a latent vector) for conditioning. This vector defines the parameters when using latent diffusion to form an image. When generating one of the representations 570 of the target object (e.g., at each iteration of the process of creating the target object), multiple steps of noise addition and noise removal may be performed sub-iteratively by the diffusion model. In particular, at each sub-iteration step during one iteration of the overall process, the diffusion model 550 outputs a latent space representation 555 for sub-iteration of the previously generated output image 570x that can be automatically selected for the next iteration of the overall process for creating the final version of the target object and its options. Throughout the implementation of latent diffusion by the diffusion model 550, one or more latent space representations 555 of the selected output image 570x, such as those generated when removing noise from the noise patch based on the vector, can be generated during the current iteration of the overall process (e.g., at each sub-iteration step). This latent space representation can be stored in the cache 565. The last sub-iteration executed by the diffusion model generates the last latent space representation, which is then decoded by the decoder 560 to generate one of the representations (e.g., the output image) of the target object provided as the output in the next iteration of the overall process. This process can be executed to generate each of the representations of the target object in the next iteration of the overall process.
[0094] As described above, the user may provide user input 501 that targets attributes found in the representation 570 of the target object generated during the previous iteration of the overall process and / or variations of the corresponding attributes. In particular, user input targeting one or more variations of the corresponding one or more attributes may, in part, be selected to indicate user preferences (e.g., like and / or dislike, etc.), and / or edited (e.g., modified, deleted, etc.), and / or adjusted to modify the selected variations of the corresponding attributes. User input 501 may not be applied to locked attributes and / or locked variations of the corresponding attributes, whereby the portion of the image corresponding to the locked feature is retained when the next iteration is performed. For illustrative purposes only, user input 501 may be visualized within an array such as array 400N of FIG. 4A.
[0095] Accordingly, at least a portion of the user input may target the identified portion of the selected output image 570x corresponding to the user input. For example, the tagging unit 520 may be implemented to automatically identify the portion 525 of the selected output image 570x corresponding to one or more variations of the corresponding attributes for which selection and / or editing and / or adjustment, etc. has been made.
[0096] As shown, user input 501 (e.g., array 400N) is encoded by encoder 510 into text prompt 515 suitable for use by the IGAI system. Additionally, encoder 510 may convert the text prompt into a latent vector for the purpose of performing latent diffusion. In one embodiment, noise addition unit 530 is configured to process identified portion 525 of selected output image 570x and generate noise patch 535 and / or a noisy version of identified portion 525. In another embodiment, the noise patch is generated randomly. In another embodiment, the portion of the final latent space representation of selected output image 570x is identified as corresponding to identified portion 525 and is used when performing latent diffusion. For example, noise addition unit 530 may be configured to identify the corresponding portion of the final latent space representation of selected output image 570x, or diffusion model 550 may be configured to perform the identification. Thus, noise patch 535 may be generated based on identified portion 525 of selected output image 570x (i.e., in image form), or based on the final latent space representation (e.g., the corresponding portion of the final latent space representation of identified output image 570x).
[0097] Furthermore, noise patch 535 corresponding to identified portion 525 of selected output image 570x is concatenated by conditioning unit 540 with text prompt 515 (i.e., latent vector) as the first set of conditioning factors and provided as an input to diffusion model 550. Latent diffusion is performed to process and / or generate (e.g., encode or denoise) a modified or updated portion 527 of selected output image 570x based on the first set of conditioning factors. The modified or updated portion of original image 575 is encoded into a latent space representation, etc. Thus, the encoded, modified or updated portion 527 of selected output image 570x reflects at least a portion of the feedback provided by the user corresponding to one or more variations of user input 501 (e.g., corresponding attributes for which selection and / or editing and / or adjustment, etc. have been made).
[0098] Rather than encoding, decoding the modified or updated portion 527 of the selected output image 570x, the modified or updated portion 527 is provided back to the conditioning unit 540 to generate a second set of conditioning factors. In particular, changes made using latent diffusion to the remaining portion of the selected output image 570x (i.e., corresponding to the locked attributes and / or locked variations of the corresponding attributes) are conditioned on or based on the result of conditioning the identified portion 525 of the selected output image 570x (i.e., the modified or updated portion 527) using a concatenated prompt. For example, the modified or updated portion 527 is partially concatenated with a text prompt 515 (or latent vector) that caused a change or modification to the identified portion 525 of the selected output image 570x to generate a second set of conditioning factors (e.g., a second latent vector).
[0099] In one embodiment, this second set of conditioning factors is then provided to the diffusion model 550 such that latent diffusion is performed on the final latent space representation of the selected output image 570x (i.e., the version that is decoded to generate the selected output image 570x), thereby causing other portions of the selected output image 570x (i.e., corresponding to the locked attributes and / or locked variations of the corresponding attributes) to be partially changed and / or modified to match the changes and / or modifications made to the identified portion 525 (i.e., corresponding to the encoded and updated portion 527). For example, the diffusion model 550 may add noise to the final latent space representation of the selected output image 570x (i.e., the version that is decoded to generate the selected output image 570x) to perform latent diffusion (i.e., denoising) based on the second set of conditioning factors.
[0100] In another embodiment, the second set of conditioning factors may include a noisy version of the final latent space representation of the selected output image 570x (i.e., the one decoded to generate the selected output image 570x), which can be generated by the noise addition unit 530. As a technical overview, the text prompt 515 (i.e., the latent vector) that causes a change or modification to the selected portion 525 is provided to an updated object (e.g., encoded and updated portion 527) for the purpose of performing latent diffusion by the diffusion model 550 on at least the remaining portion of the selected output image 570x (i.e., corresponding to the locked attributes and / or the locked variations of the corresponding attributes).
[0101] In some embodiments, the diffusion model 550 performs latent diffusion on the entire selected output image 570x (i.e., the one encoded), but makes minimal changes or no changes to the already modified selected portion 525. In that way, the remaining portion can be aligned with or take into account the changes and / or modifications made to the selected portion 525 (i.e., the encoded and updated portion 527) of the selected output image 570x. In other embodiments, the diffusion model 550 performs latent diffusion on the entire selected output image 570x (i.e., the one encoded), but makes minimal changes or no changes to the remaining portion of the selected output image 570x (i.e., corresponding to the locked attributes and / or the locked variations of the corresponding attributes). Thereafter, the remaining portion can be blended with the changes and / or modifications made to the selected portion 525 (i.e., the encoded and updated portion 527) of the selected output image 570x.
[0102] Therefore, the diffusion model 550 generates another latent space representation of the image that has been corrected this time. After decoding, the decoder 560 outputs the corrected image of the selected output image 570x. This corrected image is one representation of the target object generated in the next iteration of the entire process. For example, one or more corrected output images 575 are those generated in the current iteration and are representations of the target object for that iteration. This process may be performed on some or all of the representations 570 generated in the previous iteration (i.e., to construct the corrected output image 575, which is a newly generated representation of the target object), or may be performed on at least those representations, including all of the user input corresponding to one or more variations of the corresponding attributes on which selection and / or editing and / or adjustment, etc. have been made. This process may be repeatedly performed for one or more iteration cycles throughout the process to achieve the desired version and / or the final version of the target object, and this final version may include one or more options.
[0103] FIG. 6 shows the components of an exemplary device 600 that can be used to implement aspects of various embodiments of the present disclosure. This block diagram shows a device 600 that can be incorporated into or be such a personal computer, video game console, personal digital assistant, server, or other digital device, and includes a central processing unit (CPU) 602 for executing software applications and optionally an operating system. The CPU 602 can be composed of one or more homogeneous or heterogeneous processing cores. Further embodiments can be implemented using one or more CPUs having a microprocessor architecture particularly adapted for highly parallel and computationally intensive applications.
[0104] In particular, the CPU 602 may be configured to implement a target object construction unit 105 configured to construct a target object by performing generative AI through an iterative process that includes user input provided as feedback for the next iteration. For example, the target object construction unit 105 may generate a plurality of attributes for the target object, each of these attributes may include one or more variations. By combining different variations of corresponding attributes, one or more permutations of the target object may be generated by the target object construction unit 105.
[0105] The memory 604 stores applications and data for use by the CPU 602. The storage 606 provides non-volatile storage and other computer-readable media for applications and data, and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROMs, DVDs, Blu-ray (registered trademark), HD-DVD, or other optical storage devices, as well as signal transmission and storage media. The user input device 608 communicates user input from one or more users to the device 600. Examples of this user input device 608 may include a keyboard, mouse, joystick, touchpad, touch screen, still recorder / camera or video recorder / camera, a tracking device for recognizing gestures, and / or a microphone. The network interface 614 enables the device 600 to communicate with other computer systems via an electronic communication network, and may include wired or wireless communication via wide area networks such as local area networks and the Internet. The audio processor 612 is adapted to generate analog or digital audio output from instructions and / or data provided by the CPU 602, memory 604 and / or storage 606. The components of the device 600 are connected via one or more data buses 622.
[0106] The graphics subsystem 620 is further connected to the data bus 622 and the components of the device 600. The graphics subsystem 620 includes a graphics processing unit (GPU) 616 and a graphics memory 618. The graphics memory 618 includes a display memory (e.g., frame buffer) used to store pixel data for each pixel of the output image. The pixel data can be provided directly from the CPU 602 to the graphics memory 618. Alternatively, the CPU 602 provides data and / or instructions defining the desired output image to the GPU 616, from which the GPU 616 generates pixel data for one or more output images. The data and / or instructions defining the desired output image can be stored in the memory 604 and / or the graphics memory 618. In an embodiment, the GPU 616 includes a 3D rendering function for generating pixel data for the output image from instructions and data defining geometry, lighting, shading, texturing, motion, and / or camera parameters for the scene. The GPU 616 can further include one or more programmable execution units capable of executing shader programs. In one embodiment, the GPU 616 may be implemented within an AI engine (e.g., machine learning engine 190) to provide additional processing power for, e.g., AI, machine learning functions, or deep learning functions.
[0107] The graphics subsystem 620 periodically outputs pixel data for the image from the graphics memory 618 for display on the display device 610. The display device 610 can be any device capable of displaying visual information in response to signals from the device 600.
[0108] In other embodiments, the graphics subsystem 620 includes multiple GPU devices, which are combined to perform graphics processing for a single application running on the CPU. For example, multiple GPUs can perform alternative forms of frame rendering, including different GPUs rendering different frames at different times, different GPUs performing different shader operations, having a master GPU perform main rendering and compositing the output from a slave GPU that performs selected shader functions (e.g., smoke, rivers, etc.), different GPUs rendering different objects or different parts of a scene, etc. In the above embodiments and implementations, these operations can be performed (simultaneously in parallel) within the same frame period or (sequentially in parallel) within different frame periods.
[0109] Accordingly, in various embodiments, the present disclosure describes systems and methods configured to implement generative AI to construct target objects through an iterative process that includes user input provided as feedback for a next iteration.
[0110] Note that access services distributed over a wide geographic area, such as providing access to the games of the present embodiment, often use cloud computing. Cloud computing is a computing paradigm in which dynamically scalable and often virtualized resources are provided as services over the Internet. For example, cloud computing services often provide common applications (e.g., video games) accessible from a web browser online, while software and data are stored on servers within the cloud.
[0111] In some embodiments, the game server can be used to perform operations for video game players who play video games via the Internet. In a multiplayer game session, a dedicated server application collects data from players and distributes it to other players. The video game may be executed by a distributed game engine that includes a plurality of processing entities (PEs) that function as nodes, and as a result, each PE executes a given functional segment of the game engine on which the video game is executed. For example, the game engine implements game logic and performs game calculations, physics, geometry transformation, rendering, lighting, shading, audio, and additional in-game or game-related services. Additional services may include, for example, messaging, social utilities, voice communication, game play replay functionality, help functionality, and the like. The PEs may be virtualized by a hypervisor of a particular server, or the PEs may be present on different server units of a data center. Each processing entity for performing operations may be a server unit, virtual machine or container, GPU, or CPU, depending on the needs of each game engine segment. By distributing the game engine, the game engine is provided with flexible computing characteristics that are not restricted by the capabilities of physical server units. Instead, the game engine is provisioned with more or fewer computing nodes as needed to meet the requirements of the video game.
[0112] The user accesses the remote service using a client device (e.g., a PC, a mobile phone, etc.) that includes at least a CPU, a display, and I / O and can communicate with the game server. It should be understood that a given video game may be developed for a specific platform and related controller device. However, when such a game becomes available via a game cloud system, the user can access the video game using different controller devices, such as when accessing a game designed for a game console from a personal computer using a keyboard and mouse. In such a scenario, the input parameter settings define the mapping from the input that can be generated by the controller devices available to the user to the input that is acceptable for the execution of the video game.
[0113] In another example, the user may access the cloud game system via a tablet computing device, a touch screen smartphone, or other touch screen-driven device. In this case, the client device and the controller device are integrated, and the input is provided by the detected touch screen input / gesture. For such a device, the input parameter settings may define specific touch screen inputs (e.g., buttons, direction pads, gestures or swipes, touch motions, etc.) corresponding to the game inputs for the video game.
[0114] In some embodiments, the client device functions as a connection point for the controller device. That is, the controller device communicates with the client device via a wireless or wired connection to send inputs from the controller device to the client device. Next, after processing these inputs, the client device may send the input data to the cloud game server via the network. For example, these inputs may include video or audio captured from the game environment, which may be processed by the client device before being sent to the cloud game server. Additionally, inputs from the controller's motion detection hardware may be processed by the client device in conjunction with the video that has been captured to detect the position and motion of the controller before being sent to the cloud game server.
[0115] In other embodiments, the controller can be a network device itself and has the function of directly communicating inputs to the cloud game server via the network without first communicating such inputs through the client device. As a result, the input latency can be reduced. For example, inputs that do not depend on detection of any additional hardware or processing remote from the controller itself can be sent directly from the controller to the cloud game server. Such inputs may include button inputs, joystick inputs, built-in motion detection inputs (e.g., accelerometer, magnetometer, gyroscope), and the like.
[0116] Access by a client device to a cloud gaming network can be realized through a network implementing one or more communication technologies. In some embodiments, the network can include fifth generation (5G) wireless network technology including a cellular network serving small geographical cells. Analog signals representing sound and images are digitized at the client device and transmitted as a bitstream. 5G wireless devices within a cell communicate by radio waves with a local antenna array and a low-power transceiver. The local antenna is connected to the telephone network and the Internet by a high-bandwidth optical fiber or a wireless backhaul connection. Mobile devices crossing between cells are automatically transferred to a new cell. The 5G network is just one communication network, and embodiments of the present disclosure may utilize communication networks of earlier generations as well as wired or wireless technologies of generations subsequent to 5G.
[0117] In one embodiment, various technical examples can be implemented using a virtual environment via a head-mounted display (HMD) which may also be referred to as a virtual reality (VR) headset. As used herein, this term generally refers to user interaction with a virtual space / virtual environment, including viewing a virtual space through an HMD that responds in real time to the movement of the HMD (controlled by the user) to provide the user with a sense of being in a virtual space or metaverse. The HMD can be worn in the same way as glasses, goggles or a helmet and is configured to display video games or other metaverse content to the user. The HMD can provide a very immersive experience in a virtual environment with three-dimensional depth and perspective.
[0118] In one embodiment, the HMD may include an eye-tracking camera configured to capture an image of the user's eyes while the user is interacting with the VR scene. The gaze information captured by the eye-tracking camera(s) may include information related to the user's gaze direction and specific virtual objects and content items within the VR scene that the user is looking at or is interested in interacting with.
[0119] In some embodiments, the HMD may include an outward-facing camera(s) configured to capture an image of the user's real-world space, such as the user's body movements, and an image of any real-world objects that may be located in the real-world space. In some embodiments, the images captured by the outward-facing camera can be analyzed to determine the position / orientation of the real-world object relative to the HMD. Using the known position / orientation of the HMD, the real-world object, and inertial sensor data from the HMD, the user's gestures and movements can be continuously monitored and tracked while the user is interacting with the VR scene. For example, while interacting with a scene in a game, the user may perform various gestures (commands, communication, pointing at and walking towards a specific content item within the scene, etc.). In one embodiment, the gestures can be tracked and processed by the system to generate a prediction of the interaction with a specific content item within the game scene. In some embodiments, machine learning may be used to facilitate or assist with that prediction.
[0120] During use of the HMD, various types of single-handed and two-handed controllers can be used. In some embodiments, the controller itself can be tracked by tracking the lights included in the controller or by tracking the shape, sensors, and inertial data associated with the controller. By using these various types of controllers or by using simple hand gestures created and captured by one or more cameras, it becomes possible to interface with, control, operate, interact with, and participate in the virtual reality environment or metaverse rendered on the HMD. In some cases, the HMD can be wirelessly connected to cloud computing and game systems via a network such as the Internet or cellular. In one embodiment, the cloud computing and game systems maintain and execute the video game being played by the user. In some embodiments, the cloud computing and game systems are configured to receive inputs from the HMD and / or interface objects via the network. The cloud computing and game systems are configured to process the inputs and affect the game state of the running video game. Outputs from the running video game, such as video data, audio data, and tactile feedback data, are sent to the HMD and interface objects.
[0121] In addition, while embodiments of the present disclosure may be described with reference to an HMD, in other embodiments, for example, the screen of a portable device (e.g., a tablet, smartphone, laptop, etc.), or rendering video and / or providing the display of an interactive scene or virtual environment, or any other type of display capable of being configured to do so, other than an HMD, may be substituted. It should be understood that the various embodiments defined herein may be combined or assembled in a particular embodiment using the various features disclosed herein. Accordingly, the examples provided are only some of the possible examples and do not limit the various embodiments that can define more embodiments by combining the various elements.
[0122] Embodiments of the present disclosure may be implemented using a variety of computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and the like. Embodiments of the present disclosure may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked through a wired or wireless network.
[0123] Although the operations of the method have been described in a particular order, other housekeeping operations may be performed during the operations, or the operations may be adjusted to occur at slightly different times, or the operations may be distributed over a system that allows the processing operations to occur at various intervals related to the processing, as long as the telemetry and game state data processing for generating the modified game state is performed in the desired manner.
[0124] With the above embodiments in mind, it should be understood that embodiments of the present disclosure can employ various computer-implemented operations that require data stored in a computer system. These operations are operations that require physical manipulation of physical quantities. Any of the operations described herein in embodiments of the present disclosure are useful mechanical operations. Embodiments of the present disclosure also relate to devices or apparatuses for performing these operations. The apparatus can be specially configured for the required purpose, or the apparatus can be a general-purpose computer selectively actuated or configured by a computer program stored in the computer. In particular, various general-purpose machines can be used with a computer program written according to the teachings herein, or it may be more convenient in some cases to construct a more specialized apparatus for performing the required operations.
[0125] One or more embodiments can also be implemented as computer-readable code on a computer-readable medium. The computer-readable medium can be any data storage device capable of storing data. This data can then be read by a computer system. Examples of computer-readable media include hard drives, network attached storage (NAS), read-only memory, random access memory, CD-ROM, CD-R, CD-RW, magnetic tape, and other optical and non-optical data storage devices. The computer-readable medium can include computer-readable tangible media in which the computer-readable code is stored in a distributed manner and is distributed via a network-coupled computer system so as to be executed.
[0126] In one embodiment, a video game is executed locally on a game console, a personal computer, or on a server or by one or more servers of a data center. When the video game is executed, some instances of the video game may be simulations of the video game. For example, the video game may be executed by an environment or server that generates a simulation of the video game. The simulation is, in some embodiments, an instance of the video game. In other embodiments, the simulation may be created by an emulator that emulates a processing system.
[0127] The foregoing embodiments have been described in some detail for purposes of clarity of understanding, but it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims. Accordingly, the embodiments are to be regarded as illustrative rather than restrictive, and the embodiments should not be limited to the details given herein, but may be modified within the spirit and equivalents of the appended claims.
Claims
1. collecting one or more inputs each describing a target object; generating a plurality of images of the target object using an image generation artificial intelligence (AI) system configured to perform potential diffusion based on the one or more inputs; decomposing the target object into a first plurality of attributes each including one or more variations based on the plurality of images of the target object; receiving one or more selections from among the plurality of variations of the plurality of attributes; blending one or more of the selected plurality of variations of the plurality of attributes into one or more options of the target object, a method comprising.
2. The method according to claim 1, further comprising providing a prompt to the image generation artificial intelligence system for generating the plurality of images of the target object.
3. Said decomposing the target object comprises providing the plurality of images of the target object to an AI model configured to extract the plurality of attributes from the plurality of images, the method according to claim 2.
4. editing one or more of the plurality of variations of the plurality of attributes; using the image generation AI system to generate a second plurality of attributes for the target object based on one or more of the edited plurality of variations of the plurality of attributes and the one or more inputs collected, the method according to claim 1.
5. Said editing one or more of the plurality of variations of the plurality of attributes comprises receiving a selection of an attribute variation; receiving an editing input; adjusting the attribute variation based on the editing input to be included in the plurality of variations of the plurality of attributes, the method according to claim 4.
6. Said editing one or more of the plurality of variations of the plurality of attributes comprises filtering an attribute variation based on at least one filtering parameter, the method according to claim 4.
7. Said editing one or more of the plurality of variations of the plurality of attributes comprises receiving a selection of an attribute variation; receiving an editing input; The method according to claim 4, comprising selectively and favorably or negatively selecting a variation of the attribute based on the editing input.
8. Editing one or more of the plurality of variations of the plurality of attributes comprises: Receiving a selection of a variation of an attribute; Receiving an editing input; The method according to claim 4, comprising locking a variation of the attribute based on the editing input.
9. Determining that no variation of the attribute has been selected; The method according to claim 1, further comprising automatically selecting a variation of the attribute to perform blending the plurality of selected variations of the plurality of attributes.
10. The method according to claim 1, further comprising saving at least one of the plurality of variations of the plurality of attributes selected for use in constructing a second target object.
11. A computer system, comprising: A processor; A memory coupled to the processor and storing therein instructions that, when executed by the computer system, cause the computer system to perform a method, the method comprising: Collecting one or more inputs each describing a target object; Generating a plurality of images of the target object using an image generation artificial intelligence (AI) system configured to perform potential diffusion based on the one or more inputs; Decomposing the target object based on the plurality of images of the target object into a first plurality of attributes each including one or more variations; Receiving a selection of one or more of the plurality of variations of the plurality of attributes; The memory, comprising blending one or more of the plurality of selected variations of the plurality of attributes into one or more options of the target object.
12. The method comprises: Editing one or more of the plurality of variations of the plurality of attributes; The computer system according to claim 11, further comprising using the image generation AI system to generate a second plurality of attributes for the target object based on one or more of the plurality of variations of the edited plurality of attributes and the one or more collected inputs.
13. In the method, editing one or more of the plurality of variations of the plurality of attributes comprises receiving a selection of a variation of an attribute, receiving an editing input, and adjusting the variation of the attribute based on the editing input to be included in the plurality of variations of the plurality of attributes. The computer system according to claim 12.
14. In the method, editing one or more of the plurality of variations of the plurality of attributes comprises receiving a selection of a variation of an attribute, receiving an editing input, and selecting the variation of the attribute positively or negatively based on the editing input. The computer system according to claim 12.
15. In the method, editing one or more of the plurality of variations of the plurality of attributes comprises receiving a selection of a variation of an attribute, receiving an editing input, and locking the variation of the attribute based on the editing input. The computer system according to claim 12.
16. A non-transitory computer-readable storage medium storing a computer program executable by a processor-based system, program instructions for collecting one or more inputs each describing a target object, program instructions for generating a plurality of images of the target object using an image generation artificial intelligence (AI) system configured to perform latent diffusion based on the one or more inputs, program instructions for decomposing the target object into a first plurality of attributes each including one or more variations based on the plurality of images of the target object, program instructions for receiving a selection of one or more of the plurality of variations of the plurality of attributes, Program instructions for blending one or more of the plurality of variations of the selected plurality of attributes into one or more options of the target object, and a non-transitory computer-readable storage medium containing the same. **Claim 17** Program instructions for editing one or more of the plurality of variations of the plurality of attributes, Program instructions for using the image generation AI system to generate a second plurality of attributes for the target object based on one or more of the plurality of variations of the edited plurality of attributes and the one or more collected inputs. The non-transitory computer-readable storage medium according to claim 16, further comprising the same. **Claim 18** Said editing one or more of the plurality of variations of the plurality of attributes comprises Program instructions for receiving a selection of a variation of an attribute, Program instructions for receiving an editing input, Program instructions for adjusting the variation of the attribute based on the editing input so as to be included in the plurality of variations of the plurality of attributes. The non-transitory computer-readable storage medium according to claim 17, further comprising the same. **Claim 19** Said editing one or more of the plurality of variations of the plurality of attributes comprises Program instructions for receiving a selection of a variation of an attribute, Program instructions for receiving an editing input, Program instructions for positively or negatively selecting the variation of the attribute based on the editing input. The non-transitory computer-readable storage medium according to claim 17, further comprising the same. **Claim 20** Said editing one or more of the plurality of variations of the plurality of attributes comprises Program instructions for receiving a selection of a variation of an attribute, Program instructions for receiving an editing input, Program instructions for locking the variation of the attribute based on the editing input. The non-transitory computer-readable storage medium according to claim 17, further comprising the same.
Citation Information
Patent Citations
Program, information storage medium and image generation system
JP2010029397A