Using Game State Data for Semantic Understanding with AI Image Generative Models

The system addresses the challenge of user intent alignment in AI image generation by using game state data and dynamic interfaces to refine image output, ensuring better alignment with user preferences in video game contexts.

JP2026502519APending Publication Date: 2026-01-23SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025540464
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-10
Filing Date
2023-12-21
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing AI image generation models often fail to accurately understand user intent and generate images that align with user preferences, particularly in the context of video game interactions and streaming.

Method used

A system that utilizes game state data and dynamic user interfaces to provide feedback for AI image generation, allowing users to modify and train the model based on semantic understanding of game scenes, including depth, object arrangement, and stylistic elements, using features like feature analyzers and correction logic to refine image generation.

Benefits of technology

Enhances the accuracy and alignment of AI-generated images with user intent by incorporating game state data and dynamic user feedback, improving the quality of generated images through semantic understanding and user preference learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502519000001_ABST
    Figure 2026502519000001_ABST
Patent Text Reader

Abstract

A method is provided. [Solution] The method includes receiving game images, the game images being captured from gameplay of a video game, the game images depicting a video game scene; receiving game state data describing attributes of the video game scene to be depicted in the game images; receiving modification data from a client device over a network describing changes to the game images, the modification data being defined from user input received at the client device; applying the game images, game state data, and user input by an image generation artificial intelligence (AI) to generate an AI-generated image; and transmitting the AI-generated image over the network to the client device for rendering on a display.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure generally relates to methods, systems, and devices for using game state data for semantic understanding by AI image generation models, and to dynamic interfaces for providing feedback regarding output by the AI ​​image generation models. [Background technology]

[0002] The video game industry has undergone many changes over the years. As technology advances, video games continue to achieve greater immersion through sophisticated graphics, realistic sounds, compelling soundtracks, haptic feedback, and more. Players can enjoy immersive gaming experiences that involve participating in and engaging with virtual environments, demanding new methods of interaction. Additionally, players can stream videos of their gameplay for spectators to watch, allowing others to share their gameplay experiences.

[0003] It is in these situations that embodiments of the present disclosure arise. Summary of the Invention

[0004] Embodiments of the present disclosure include methods, systems, and devices for using game state data for semantic understanding by AI image generation models, and dynamic interfaces for providing feedback regarding outputs by the AI ​​image generation models.

[0005] In some embodiments, a method is provided that includes receiving game images, the game images being captured from gameplay of a video game, the game images depicting a video game scene; receiving game state data describing attributes of the video game scene to be depicted in the game images; receiving, over a network from a client device, modification data describing changes to the game images, the modification data being defined from user input received at the client device; applying, by an image generation artificial intelligence (AI), the game images, the game state data, and the user input to generate an AI-generated image; and transmitting the AI-generated image over the network to the client device for rendering on a display.

[0006] In some implementations, applying the game state data enables the image generation AI to semantically understand the scene depicted in the game image.

[0007] In some implementations, a semantic understanding of the scene is applied by the image generation AI to perform the changes described in the correction data.

[0008] In some implementations, the game state data identifies one or more elements in a scene that are to be depicted in the game image.

[0009] In some implementations, the game state data describes the depth of one or more virtual objects in a scene.

[0010] In some implementations, the correction data describes a change in location of a given virtual object within the scene, and generating the AI-generated image is configured to implement the change in location of the given virtual object described by the correction data using a depth of the one or more virtual objects.

[0011] In some implementations, the correction data describes an arrangement of given virtual objects within the scene, and generating the AI-generated image is configured to implement the arrangement of the given virtual objects described by the correction data using a depth of one or more virtual objects.

[0012] In some implementations, the depth of the one or more virtual objects is configured such that performing the changes described by the modification data allows for proper occlusion of or by the one or more virtual objects.

[0013] In some implementations, the game state data describes the three-dimensional structure of one or more virtual objects in a scene.

[0014] In some implementations, the modification data is defined by a word or phrase generated by user input received at the client device.

[0015] In some embodiments, a non-transitory computer-readable medium having program instructions embodied thereon, the program instructions being configured, when executed by at least one server computer, to cause the at least one server computer to perform a method, the method including receiving game images, the game images captured from gameplay of a video game, the game images depicting video game scenes; receiving game state data describing attributes of the video game scenes to be depicted in the game images; receiving, over a network from a client device, modification data describing changes to the game images, the modification data defined from user input received at the client device; applying, by an image generation artificial intelligence (AI), the game images, the game state data, and the user input to generate AI-generated images; and transmitting the AI-generated images over the network to the client device for rendering on a display.

[0016] Other aspects and advantages of the present disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, illustrating by way of example the principles of the disclosure.

[0017] The present disclosure may be better understood by reference to the following description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0018] [Figure 1] 1 conceptually illustrates an image generation service that provides a user interface (UI) for modifying an image, according to an embodiment of the present disclosure. [Figure 2] 1 conceptually illustrates the generation of an image by an image generation AI based on game images and associated game state information, according to an embodiment of the present disclosure. [Figure 3] 1 conceptually illustrates the extraction of features from a game generation scene for use as input to an image generation AI, according to an embodiment of the present disclosure. [Figure 4] 1 conceptually illustrates a system for storing profiles for interpreting user input for AI image generation, according to an embodiment of the present disclosure. [Figure 5] 1 illustrates conceptually crowdsourcing of a topic related to image generation by image generation AI according to an embodiment of the present disclosure. [Figure 6A] 1 illustrates an overview of an image generation AI (IGAI) processing sequence according to an embodiment of the present disclosure. [Figure 6B] 1 illustrates an overview of an image generation AI (IGAI) processing sequence according to an embodiment of the present disclosure. [Figure 6C] 1 illustrates an overview of an image generation AI (IGAI) processing sequence according to an embodiment of the present disclosure. [Figure 7] 1 illustrates components of an exemplary device that can be used to perform aspects of various embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0019] The following embodiments of the present disclosure provide methods, systems, and devices for a dynamic interface to provide feedback regarding output by an AI image generation model using game state data for semantic understanding by the AI ​​image generation model.

[0020] When an image is automatically generated by an AI model in response to user input, the resulting image may not be the user's favorite. A UI is dynamically generated for the image to modify the image or to provide dynamic feedback to train the AI ​​model to generate images consistent with the user's intent. The UI has a selection interface that dynamically locks to specific features shown in the image. For example, if a person is present in the scene of the generated image, the user interface may automatically identify the person and suggest options to modify them, and feedback on such modifications is returned to the AI ​​model for training, which then generates additional images that are more consistent with that feedback. In one embodiment, the image is broken down into layers, which can be automatically identified by the AI ​​analysis model. Layers may include, for example, background, foreground, midground, and isolation of specific objects in the image for selective modification or removal. In one embodiment, feedback provided by the UI may allow the user to provide instructions on how the image should be modified. Modifications may include not only the content but also the angle at which the image is taken, similar to how a photographer is instructed to capture different images from different perspectives. This feedback is then processed by the AI ​​model as training information to generate new images that more closely match the user's desired intent. In some embodiments, the user's intent itself can be analyzed to provide a profile of the user, for example, identifying the user's preferences and favorites for future revisions to automatically generated images.

[0021] With the above summary in mind, the following provides several illustrative figures to facilitate understanding of example embodiments.

[0022] FIG. 1 conceptually illustrates an image generation service that provides a user interface (UI) for modifying an image, according to an embodiment of the present disclosure.

[0023] In the illustrated implementation, image generation service 100 includes image generation artificial intelligence 102 configured to generate images in response to user input. Image generation service 100 is accessible over network 108 (e.g., including the Internet) by client device 110. By way of non-limiting example, client device 100 may be a personal computer, laptop, tablet, mobile phone, gaming console, set-top box, streaming box, or any other type of computing device capable of performing the functions attributed thereto in this disclosure. Client device 110 executes application 112 that accesses image generation service 100 over network 108. In some implementations, application 112 is a web browser, and image generation service 100 is accessible as a website on the Internet. In other implementations, application 112 is a dedicated application or app executed by the client device that communicates with image generation service 100, such as by accessing an application programming interface (API) exposed by image generation service 100.

[0024] The application 112 renders a user interface 114, through which a user 116 interfaces with the image generation service 100. For example, via the user interface 114, the user 116 can provide user input, such as descriptive text or an image, that is used by the image generation AI 102 to generate an image. It will be appreciated that when an image is automatically generated by the image generation AI 102 in response to user input, the resulting image may not be to the user's liking. Therefore, a correction UI is dynamically generated for the image and presented as part of the UI 114 to correct the image or to provide dynamic feedback to train the AI ​​model to generate an image consistent with the user's intent. In some implementations, the image generation service 100 further includes a feature analyzer 104 configured to analyze the image to identify features of the image that the user may wish to correct. For example, the feature analyzer 104 may identify various elements or objects within the image, and based on that identification, the correction logic 106 determines possible corrections that can be suggested to the user via the correction UI. In some implementations, the feature analyzer 104 uses a recognition model to identify features of the image.

[0025] In this manner, the modification UI can provide a selection interface that is dynamically locked to specific features shown in the image. For example, if a person 120 is present in the scene of the generated image 118, the feature analyzer 104 may automatically identify the person 120, and the modification logic 106 may suggest modification options, such as making the person taller or shorter, adjusting the person's clothing, etc. In some implementations, the system may identify the user's head 122 and suggest specific modifications, such as changing the person's facial expression, hair color, etc. In some implementations, the system may identify a tree 124 in the image 118 and suggest modifications, such as making the tree shorter or taller, wider, more or less green, more or less foliage, adding flowers, fruit, etc.

[0026] In some implementations, a selection tool is provided in the modification UI whereby a user may identify an area of ​​the image to modify, such as by using a paintbrush tool to shade the area, drawing a box with a drawing tool or predefined shape tool, or surrounding the area, etc. The feature analyzer 104 analyzes the identified area to determine the content of the area, and the modification logic 106 suggests modifications based on the identified content of the area.

[0027] Thus, the user 116 can select one or more of the suggested corrections as appropriate. Additionally, the correction UI may allow the user 116 to input additional corrections or instructions on how to modify the image, such as by entering text describing the additional corrections. The selected or entered corrections are returned as feedback to the image generation AI for training, which then generates additional images that more closely match the feedback.

[0028] In one embodiment, an image is analyzed and broken down into layers, which can then be automatically identified by an AI analysis model. Layers may include, for example, background, foreground, midground, and isolation of specific objects within the image for selective modification or removal. Thus, a user can identify objects relative to their position in the scene and use this information to issue modification commands (e.g., move an object from the foreground to the background), or modify entire regions of the scene, such as the background, foreground, midground, or any identified layer.

[0029] Additionally, in one embodiment, modifications may include not only the content but also the angle at which the image is taken, similar to how a photographer is instructed to capture different images from different viewpoints. For example, user input may include instructions to create an image from a different viewpoint or with different optics: closer, farther away, rotated, lower, higher, overhead, left, right, wider / narrower angle, zoom, etc. This feedback is then processed by the image generation AI model as training information, which then generates new images that more closely match the user's desired intent.

[0030] In some embodiments, the user's intent itself can be analyzed to provide a profile of the user, for example, to identify the user's preferences and favorites for future modifications to automatically generated images.

[0031] FIG. 2 conceptually illustrates the generation of an image by an image generation AI based on game images and associated game state information, according to an embodiment of the present disclosure.

[0032] In some implementations, a user 200 playing a video game 204 executed by a gaming console 202 may capture game images 206 from their gameplay. In local gaming implementations, the gaming console 202 is a local device (e.g., a computer, gaming console, etc.), while in cloud gaming implementations, the gaming console 202 is a remote device, such as a server computer / blade, that runs the video game 204 and streams gameplay video to the user's local client device (not shown). In some implementations, the game images 206 are still images from the user's gameplay. In other implementations, the game images 206 may be video clips from the user's gameplay.

[0033] The game image 206 can be provided as input to the image generation AI 102 to generate an image. Additionally, a user can provide user input in the form of modification data 210 indicating how the user wants to modify the game image 206, or can otherwise use the game image 206 as a seed image for generating a new AI-generated image 212. In some implementations, game state data 208 is also provided as input to the image generation AI 102 to provide further semantic understanding of the content of the game image 206. The game state data 208 includes data describing the state of the virtual environment of the video game 204 at the time the game image 206 was captured during gameplay.

[0034] It will be appreciated that the game state data 208 may describe various aspects of the scene depicted in the game image 206, such as the identity of objects / elements, the depth of objects, object movements occurring in the scene, sounds played at the time of image capture, sounds associated with particular objects, words spoken by characters, the 3D structure of objects, information about occluded objects in the scene, lighting information, physics information, etc. Using such game state information as the image generation AI 102, this may result in improved understanding of the content of the game image, and consequently improved semantic understanding of user-described modifications / alterations.

[0035] By way of example and not limitation, the game state data 208 may include depth information for objects in the game image 206. Thus, the depth information allows for the understanding of the relative positioning of objects in a scene in three dimensions. Accordingly, when the user modification data 210 includes user instructions to move an object, such movement can be understood relative to other objects in three dimensions. For example, a user may input a statement such as "put a plant behind the couch," and in response, the image generation AI understands the depth of the couch in the scene and places the plant at the correct depth relative to the couch and relative to the depths of other objects in the scene. Or, when a user inputs a direction to move an object, the object can be moved and replaced with the correct depth relative to other objects or elements. For example, a command to "move that tree to the left" can be understood with the correct depth information, such that the movement and placement of the tree does not interfere with a person walking a dog in the scene, which is at a closer depth than the tree.

[0036] In related implementations, additional game state information can provide further improvements in image generation. For example, by using information describing occluded objects or occluded portions of objects in game image 206, when an object is moved, the previously occluded object or occluded portion can be revealed and included by image generation AI 102 in AI-generated image 212. As another example, by using information describing the 3D structure of an object, the object can be moved or positioned while taking into account the depth boundaries of other objects, so that, for example, an object does not appear too close in front of or behind another object in the scene. Thus, by providing additional semantic understanding of what is in the scene in game image 206, image generation AI 102 can better process user prompts or inputs describing modifications or changes the user wants to make.

[0037] FIG. 3 conceptually illustrates the extraction of features from a game generation scene for use as input to an image generation AI, according to an embodiment of the present disclosure.

[0038] In some implementations, the game image 300 is provided by the image generation AI 102 to provide style or art information for generating the AI-generated image 306. To facilitate understanding of the style / art elements of the game image 300, the image analyzer 304 is configured to analyze the game image 300 to determine the style / art elements of the game image 300. For example, in some implementations, the image analyzer 304 is configured to analyze the lighting of the scene depicted in the game image 300, which may include analyzing light sources, light source positions, color temperature, intensity, contrast, etc. In some implementations, the image analyzer 304 is configured to analyze other artistic aspects of the game image 300, such as the color palette employed, the type of lines defining object boundaries, the type of texture or shading, etc.

[0039] The extracted art / style information can be provided to the image generation AI 102 as additional input used to influence the generation of the AI-generated image 306. For example, the correction data 302 can reference style elements of the game image 300, which can be determined by the image analyzer 304 and used as input for the image generation AI 306. For example, the correction data 302 can include instructions to generate an image that includes lighting similar to that of the game image 300. The image analyzer 304 can analyze the lighting of the game image 300 and thereby generate lighting information describing the lighting of the game image 300, which lighting information can be used by the image generation AI to generate the AI-generated image 306 with similar lighting. In this example, the AI-generated image 306 may have lighting of a similar color temperature to that of the game image 300, a similarly positioned light source, etc.

[0040] In some implementations, the image analyzer 304 is triggered to analyze a given stylistic aspect of the game image 300 in response to a reference to such stylistic aspect of the game image in the modification data. For example, in the embodiment described above involving lighting, the image analyzer 304 can be triggered to analyze the lighting of the game image 300 in response to a user input indicating a reference to the lighting of the game image 300. In some implementations, the image analyzer 304 is a recognition model configured to recognize stylistic or artistic elements of an image.

[0041] In this manner, the game image 300 is used as a type of reference image that provides stylistic input for purposes of image generation by the image generation AI 102. It will be appreciated that a user may not wish to have the AI-generated image 306 simply complete in the style of the game image 300, but rather wish to apply only certain style elements. Thus, the present embodiment allows for selective use of style elements from the game image 300 to be applied to the image generation, as specified by the modification data 302.

[0042] In some implementations, the image generation AI 102 generates an AI-generated image 306 based on input provided in the modification data 302, incorporating style elements from the game image 300 as described herein. In other implementations, the image generation AI 102 uses another image, as described above, with the modification data 302 describing modifications / alterations to the image to generate the AI-generated image 306, incorporating style elements from the game image 300 as described herein.

[0043] In some implementations, the image analyzer 304 can also analyze the style / art elements of the game image 300 using the game state data and information described above.

[0044] Although the present embodiment describes a game image 300, in other embodiments, other types of images can be utilized and analyzed for style / art elements according to implementations of the present disclosure.

[0045] FIG. 4 conceptually illustrates a system for storing profiles for interpreting user input for AI image generation, according to an embodiment of the present disclosure.

[0046] In some implementations, the system is configured to learn a user's understanding and intent using words used as input to the image generation AI 102. This understanding can define a user profile that is used to interpret the user's input words. In the illustrated implementation, profile storage 402 is provided in which user profiles are stored. It will be appreciated that a given user may have two or more profiles, such as profile 404 and profile 406, to allow for different understandings of the user's intent to be used. It will be appreciated that a user may wish to create different profiles to facilitate the generation of images with different styles or elements based on different learned understandings of user preferences associated with the different profiles. In some implementations, a given profile maps words to one or more other words and / or data that define the semantic understanding of the words as determined for the particular profile.

[0047] In some embodiments, onboard logic 400 is configured to provide an onboarding process in which a user can indicate their preferred usage or understanding of particular words. For example, in some embodiments, onboard logic 400 presents a number of images via UI 114, and the user may be asked to describe the images, or the user may be asked to associate the images with specific predefined terms as understood by the user. In this manner, the user's understanding of the language used to describe the images can be learned. In some embodiments, the user's description or indication of their understanding of the images is mapped to known input words or phrases used to generate the images. This learned information regarding the user's preferences or understanding of word or phrase usage is stored in a given profile, such as profile 404 in the illustrated embodiment. While onboarding logic 400 is useful for the initial setup of a given profile, it can also be used at any time as a training tool to provide explicit training regarding the user's understanding and intent regarding words and images.

[0048] Next, when a user enters user input 412 for generating an image, such input is processed by interpreter 408 based on, in the illustrated embodiment, the currently active profile 404, and user input 412 is interpreted based on profile 404. In some embodiments, profile 404 is used by interpreter 408 to convert user input 412 into converted input 410 that is provided to image generation AI 102. For example, in some embodiments, profile 404 is used to map words or phrases included in user input 412 to other words or phrases that are included in converted input 410. For example, user input 412 may include the word "dark," and based on the user's active profile 404, the word "dark" is mapped to additional words / phrases such as "fantasy," "HRGiger," etc. Accordingly, the interpreter generates converted input 410 to include one or more of these additional words / phrases.

[0049] It will be appreciated that because many words or phrases are subjective or open to interpretation, the profile of the present embodiment provides a way to learn, remember, and apply a user's subjective understanding of the meaning of such words or phrases so as to achieve results from the image generation AI 102 that better match the user's expectations. Additionally, a user's profile can further learn over time through use of the system. For example, the profile logic 416 can be configured to analyze the user input 412 and associate words / phrases used by the user with the user's profile, such as words used repeatedly by the user, or words that tend to be clustered or used together by the user.

[0050] In some implementations, the user provides modification input 414 in response to the image 418 generated from the image generation AI 102. The modification input 414 may indicate changes the user wants to make to the image 418, as described above, and may provide insight into the user's original intent with the original user input 412 used to generate the image 418. Accordingly, in some implementations, the profile logic 416 analyzes the modification input 414 to further determine the user's understanding of the words provided in the user input 412, which understanding is stored in the active profile 404. For example, in some implementations, the words provided in the modification input 414 may be mapped or associated with the words provided in the user input 412 and, as such, stored in the active profile 404.

[0051] In some implementations, the system can be configured to suggest words or phrases as the user is generating their user input 412. In some implementations, in response to a given word or phrase provided by the user, multiple potentially related words or phrases are suggested, and the user can select one or more of the suggestions. Based on the user's selections in such cases, the user's active profile 404 can be updated over time, such as by correlating or mapping words based on such selections.

[0052] In some implementations, a given profile can define a learning model that is trained to predict or infer a user's preferred words / phrases based on given supplied words or phrases. The learning model is trained using any of the techniques described herein and data describing the user's understanding of words, terms, phrases, etc. In some implementations, the learning model is configured to associate, map, or cluster various words or phrases, and these associations are strengthened or weakened over time as a result of training. The trained learning model is used by the interpreter 408 to generate predicted words based on the user input 412, which can be appended to the user input 412 or otherwise included to generate the transformed input 410 that is provided to the image generation AI 102.

[0053] In some implementations, a given profile is configured to calibrate degree terms when used by a user. For example, a user's use of the term "tall" may be equivalent to "actually tall" as applied by the image generation AI 102 to achieve a preferred result for the user. Thus, the profile system of this implementation can be configured to learn user preferences in this regard.

[0054] In some implementations, additional signals can be used by the profile logic 416 to further refine a given profile. For example, in some implementations, the image generation AI 102 can generate multiple images based on a given user input, and the user can select which one best matches what they intended. Such a selection by the user can be used as feedback to adjust the user's profile. In further implementations, selection of additional features following image generation, such as choosing an image to upscale or re-running image generation based on a given selected image, can also be used as feedback.

[0055] It will be recognized that a challenge in using image generation AI systems is that users struggle to provide the correct input that will achieve the desired result. Therefore, by implementing a system that learns users' preferences and understanding of input terms, users can more efficiently achieve improved image generation.

[0056] In some embodiments, a theme can be defined for a given user or a given profile. A theme can be configured to define a particular style and accordingly include particular words / phrases or other types of acceptable input that, when applied to an image generation AI, cause the image generation AI to generate images in a particular style. In some embodiments, a theme is editable, allowing a user to specify particular words / phrases or other particular input to be part of the theme's definition. Then, when user input is entered to generate an image, the theme is applied by appending the words / phrases / input stored in the theme.

[0057] FIG. 5 conceptually illustrates crowdsourcing of a theme related to image generation by an image generation AI according to an embodiment of the present disclosure.

[0058] In some embodiments, inputs that tend to be used or are popular for generating images by the image generation AI are determined and used to crowdsource themes that can be applied by subsequent users of the image generation AI. In the illustrated embodiment, various users 500a, 500b, 500c, 500d, etc. each generate user inputs 502a, 502b, 502c, 502d, etc. The user inputs are provided for the purpose of generating images by the image generation AI, as has been described.

[0059] In some implementations, the user input is analyzed by a trend analyzer 504, which identifies popular or trending inputs based on the user input. For example, the trend analyzer 504 may identify popular or trending words, terms, phrases, or other inputs being entered by users of the system. Based on these trending inputs, the theme generator 506 is configured to generate one or more themes that include or are otherwise defined by the set of popular or trending inputs. For example, a given theme may be defined to include a particular collection of words that users tend to use together.

[0060] It will be appreciated that by analyzing input terms across many users, certain relationships between various input terms may be discovered that would not otherwise be apparent. For example, some words that may not appear to be related to each other may be found to occur in conjunction with each other in highly regular user input.

[0061] It will be appreciated that there may be a library of themes, such as themes 510, 512, stored in theme storage 508 from which a user may select a given theme to apply to their image generation instance. Implementing selectable themes in accordance with embodiments of the present disclosure allows users to more quickly generate images with a particular style or look.

[0062] In one embodiment, generation of output images, graphics, and / or 3D representations by an Image Generating AI (IGAI) may involve one or more artificial intelligence processing engines and / or models. Generally, AI models are generated using training data from a dataset. The dataset selected for training can be custom curated for a specific desired output, or in some cases, the training dataset may include a wide range of general-purpose data available from numerous sources over the internet. For example, the IGAI may have access to a large amount of data, such as images, videos, and 3D data. The general-purpose data is used by the IGAI to gain an understanding of the type of content desired by the input. For example, if the input calls for the generation of a tiger in the Sahara Desert, the dataset should have a variety of images of tigers and deserts to access and render during processing of the output image. Alternatively, a curated dataset may be more specific to a type of content, such as video game-related art, videos, and other asset-related content. More specifically, a curated dataset may include images related to a particular scene in a game or action sequences involving game assets, such as a unique avatar character. As described above, the IGAI can be customized to allow input of unique descriptive language statements to set styles for requested output images or content. The descriptive language statements can be text or other sensor input, such as inertial sensor data, input velocity, emphasis statements, and other data that can be generated during the input request. The IGAI can also provide an image, video, or set of images to define the context of the input request. In one embodiment, the input can be text describing the desired output along with one or more images to convey the desired contextual scene being requested as output.

[0063] In one embodiment, an IGAI is provided to enable text-to-image generation. The image generation is configured to perform a latent diffusion process in a latent space to synthesize text to image processing. In one embodiment, a conditioning process assists in tuning the output toward a desired usage output, for example, using structured metadata. The structured metadata may include information obtained from user input to teach the machine learning model to gradually denoise in stages using cross-attention until the processed denoising is decoded back into pixel space. During the decoding stage, upscaling is applied to achieve higher quality images, videos, or 3D assets. Thus, the IGAI is a custom tool designed to process a specific type of input and render a specific type of output. When the IGAI is customized, the machine learning and deep learning algorithms are tuned to achieve a specific custom output, such as a gaming technique, a specific game title, and / or a unique image asset used in a movie.

[0064] In another configuration, the IGAI may be a third-party processor, such as a processor provided by Stable Diffusion, or other processors such as OpenAI's GLIDE, DALL-E, MidJourney, or Imagen. In some configurations, the IGAI may be available online via one or more application programming interface (API) calls. It should be understood that references to available IGAIs are for informational reference only. For additional information related to IGAI technology, reference may be made to a paper by Robin Rombach, et al., entitled "High-Resolution Image Synthesis with Latent Diffusion Models," published by the Ludwig Maximilian University of Munich (pp. 1-45). This paper is incorporated by reference.

[0065] FIG. 6A illustrates an overview of the processing sequence of an image generating AI (IGAI) 602 according to one embodiment. As shown, an input 606 is configured to receive input in the form of data, such as a text description having semantic descriptions or keywords. The text description may be in the form of a sentence having at least a noun and a verb. The text description may also be in the form of a fragment or a simple single word. The text may also be in the form of multiple sentences describing a scene, some actions, or some characteristics. In some configurations, the input text may be entered in a specific order to emphasize some words over others, or even to de-emphasize words, characters, or sentences. Furthermore, the text input may be in any form, including characters, emojis, icons, and foreign language characters (e.g., Japanese, Chinese, Korean, etc.). In one embodiment, the text description is enabled by contrastive learning. The basic idea is to embed both the image and the text into a latent space so that the text corresponding to the image is mapped to the same region in the latent space as the image. This abstracts structure, e.g., what it means to be a dog, from both the visual and textual representations. In one embodiment, the goal of contrastive representation learning is to learn an embedding space in which similar sample pairs are close to each other while dissimilar sample pairs are far apart. Contrastive learning can be applied in both supervised and unsupervised settings. When dealing with unsupervised data, contrastive learning is one of the most powerful approaches in self-supervised learning.

[0066] In addition to text, the input may also include other content, such as images, or even images that themselves have descriptive content. Images can be interpreted using image analysis to identify objects, colors, shapes, characteristics, shading, textures, three-dimensional representations, depth data, and combinations thereof. Broadly speaking, the input 606 is configured to convey a user's intent to generate some digital content using the IGAI. In the context of gaming technologies, the target content to be generated may be game assets for use in a particular game scene. In such a scenario, the dataset and input 606 used to train the IGAI can be used to customize how the artificial intelligence, e.g., a deep neural network, processes the data to manipulate and adjust the desired output image, data, or three-dimensional digital asset.

[0067] The input 606 is then passed to the IGAI, where an encoder 608 takes the input data and / or pixel-space data and converts it into latent-space data. The concept of "latent space" is central to deep learning because feature data is reduced to a simplified data representation for the purpose of discovering patterns and using those patterns. Thus, latent-space processing 610 is performed on compressed data. This significantly reduces processing overhead compared to processing learning algorithms in pixel space, which consumes significantly more resources and requires significantly more processing power and time to analyze and produce the desired image. Latent space is simply a representation of compressed data, in which similar data points are closer together in space. In latent space, processing is configured to allow the machine learning system to learn learned relationships between data points that could be derived from supplied information, e.g., the dataset used to train the IGAI. In latent-space processing 610, a diffusion process is calculated using a diffusion model. A latent-space diffusion model relies on an autoencoder to learn a lower-dimensional representation of pixel space. The latent representation is passed through a diffusion process that adds noise at each step, e.g., multiple stages. The output is then fed to a denoising network based on a U-Net architecture with cross-attention layers. A conditioning process is also applied to train the machine learning model to remove noise and arrive at an image that closely represents what was requested via user input. A decoder 612 then converts the resulting output from the latent space back to pixel space. The output 614 can then be processed to improve resolution. The output 614 is then passed as a result. This can be an image, graphics, 3D data, or data that can be rendered into a physical or digital format.

[0068] FIG. 6B illustrates additional processing that may be performed on input 206 in one embodiment. A user interface tool 620 may be used to allow a user to provide an input request 604. The input request 604 may be an image, text, structured text, or general-purpose data, as described above. In one embodiment, before the input request is provided to encoder 608, the input may be processed by a machine learning process that generates a machine learning model 632 and learns from a training dataset 634. As an example, the input data may be processed using a context analyzer 626 to understand the context of the request. For example, if the input is "space rocket for flight to Mars," the input may be analyzed by context analyzer 626 to determine that the context is related to outer space and planets. The context analysis may use machine learning model 632 and training dataset 634 to find images related to this context or to identify a specific library of art, images, or videos. If the input request also includes an image of a rocket, feature extractor 628 can function to automatically identify characteristic features of the rocket image, such as fuel tanks, length, color, position, edges, lettering, flames, etc. A feature classifier 630 can also be used to classify the features and improve machine learning model 632. In one embodiment, input data 607 can be generated to produce structured information that can be encoded into a latent space by encoder 608. Additionally, structured metadata 622 can be extracted from the input request. Structured metadata 622 can be, for example, descriptive text used to instruct IGAI 602 to modify characteristics or alter the input image, or change the color, texture, or a combination thereof. For example, input request 604 can include an image of a rocket, and the text can say "make the rocket wider," or "add more flames," or "make the rocket sturdier," or some other modifier intended by the user (e.g., semantically provided and contextually analyzed).The structured metadata 622 can then be used in subsequent latent space processing to tailor the output to more closely match the user's intent. In one embodiment, the structured metadata can be in the form of a semantic map, text, image, or other data designed to represent the user's intent regarding what changes or modifications should be made to the input image or content.

[0069] FIG. 6C illustrates how the output of the encoder 208 is then provided to the latent space processing 210, according to one embodiment. A diffusion process is performed by the diffusion process stage 640, where the input is processed through multiple stages to add noise to the input image, or an image associated with the input text. This is a step-by-step process, with noise being added at each stage, e.g., 10 to 50 or more stages. Next, a denoising process is performed by the denoising stage 642. Similar to the noise stage, an inverse process is performed, where noise is progressively removed at each stage, using machine learning to predict what the output image or content should be, given the intent of the input request. In one embodiment, the structured metadata 622 can be used by the machine learning model 644 at each stage of denoising to predict what the resulting denoised image will look like and how that image should be modified. During these predictions, the machine learning model 644 uses the training dataset 646 and the structured metadata 622 to increasingly approximate an output that most closely resembles what was requested in the input. In one embodiment, a U-Net architecture with a cross-attention layer may be used to improve prediction during denoising. After the final denoising stage, the output is provided to a decoder 612, which converts the output to pixel space. In one embodiment, the output is also upscaled to improve resolution. The decoder output, in one embodiment, may optionally be run through a context conditioner 636. The context conditioner is a process that uses machine learning to examine the resulting output and may make adjustments to make the output more realistic or to remove unrealistic or unnatural outputs. For example, if the input asks for "a boy pushing a lawnmower" and the output shows a boy with three legs, the context conditioner may make adjustments through inpainting or overlay to correct or block inconsistent or undesirable outputs.However, as the machine learning model 644 becomes smarter over time with more training, there is less need for the context conditioner 636 before the output is rendered in the user interface tool 620.

[0070] FIG. 7 illustrates components of an exemplary device 700 that can be used to perform aspects of various embodiments of the present disclosure. The block diagram illustrates device 700, which can incorporate or be a personal computer, video game console, personal digital assistant, server, or other digital device suitable for implementing embodiments of the present disclosure. Device 700 includes a central processing unit (CPU) 702 for executing software applications and, optionally, an operating system. CPU 702 can be comprised of one or more homogeneous or heterogeneous processing cores. For example, CPU 702 can be one or more general-purpose microprocessors having one or more processing cores. Further embodiments can be implemented using one or more CPUs with a microprocessor architecture particularly adapted for highly parallel and computationally intensive applications, such as interpreting queries, identifying contextually relevant resources, and immediately implementing and rendering contextually relevant resources in a video game. The device 700 may be localized to the player playing the game segment (e.g., a gaming console), or may be remote from the player (e.g., a back-end server processor), or may be one of many servers that use virtualization in a gaming cloud system for remote streaming of gameplay to clients.

[0071] Memory 704 stores applications and data used by CPU 702. Storage 706 provides non-volatile storage and other computer-readable media for applications and data and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other optical storage devices, as well as signal transmission and storage media. User input devices 708 communicate user input from one or more users to device 700; examples of user input devices 708 may include a keyboard, mouse, joystick, touchpad, touchscreen, still or video recorder / camera, tracking device for gesture recognition, and / or microphone. Network interface 714 enables device 700 to communicate with other computer systems over electronic communications networks, which may include wired or wireless communications over local area networks and wide area networks such as the Internet. The audio processor 712 is adapted to generate analog or digital audio output from instructions and / or data provided by the CPU 702, memory 704, and / or storage 706. The components of the device 700, including the CPU 702, memory 704, data storage 706, user input device 708, network interface 710, and audio processor 712, are connected via one or more data buses 722.

[0072] A graphics subsystem 720 is further connected to a data bus 722 and the components of device 700. The graphics subsystem 720 includes a graphics processing unit (GPU) 716 and a graphics memory 718. The graphics memory 718 includes a display memory (e.g., a frame buffer) used to store pixel data for each pixel of an output image. The graphics memory 718 may be integrated into the same device as the GPU 708, connected as a separate device from the GPU 716, and / or implemented within memory 704. Pixel data may be provided to the graphics memory 718 directly from the CPU 702. Alternatively, the CPU 702 may provide data and / or instructions defining a desired output image to the GPU 716, and the GPU 716 may generate pixel data for one or more output images from the data and / or instructions. The data and / or instructions defining a desired output image may be stored in memory 704 and / or the graphics memory 718. In one embodiment, GPU 716 includes 3D rendering functionality for generating pixel data for output images from instructions and data defining scene geometry, lighting, shading, texturing, motion, and / or camera parameters. GPU 716 may also include one or more programmable execution units capable of executing shader programs.

[0073] Graphics subsystem 714 periodically outputs pixel data of an image from graphics memory 718 for display on display device 710. Display device 710 may be any device capable of displaying visual information in response to signals from device 700, including CRT displays, LCD displays, plasma displays, and OLED displays. Device 700 may provide analog or digital signals to display device 710, for example.

[0074] It should be noted that access services distributed over wide geographic areas, such as providing access to games in this embodiment, often use cloud computing. Cloud computing is a style of computing in which dynamically scalable, often virtualized resources are provided as a service over the Internet. Users do not need to be experts in the "cloud" technical infrastructure that supports them. Cloud computing can be categorized into different services, such as infrastructure as a service (IaaS), platform as a service (PaaS), and software as a service (SaaS). Cloud computing services often provide common applications, such as online video games, accessed from a web browser, while software and data are stored on servers in the cloud. The term cloud is used as a metaphor for the Internet based on how the Internet is depicted in computer network diagrams and is an abstract concept that hides a complex infrastructure.

[0075] In some embodiments, a game server may be used to operate a persistent information platform for video game players. Most video games played over the Internet operate through a connection to a game server. Typically, the game uses a dedicated server application that collects data from players and distributes it to other players. In other embodiments, a video game may be executed by a distributed game engine. In these embodiments, the distributed game engine may run on multiple processing entities (PEs), such that each PE performs a functional segment of the given game engine on which the video game is executed. Each processing entity appears to the game engine as simply a computational node. A game engine typically performs a functionally diverse set of operations to execute a video game application along with additional services experienced by the user. For example, a game engine implements game logic and performs game calculations, physics analysis, geometry transformations, rendering, lighting, shading, audio, and additional in-game or game-related services. The additional services may include, for example, messaging, social utilities, voice communication, gameplay playback functions, help functions, etc. The game engine may sometimes run on an operating system virtualized by a hypervisor on a particular server, but in other embodiments the game engine itself may be distributed among multiple processing entities, each of which may reside on a different server unit in a data center.

[0076] According to this embodiment, each processing entity for performing an operation may be a server unit, a virtual machine, or a container, depending on the needs of each game engine segment. For example, if a game engine segment is responsible for camera transformation, that particular game engine segment may be provisioned with a virtual machine associated with a graphics processing unit (GPU) because it performs a large number of relatively simple mathematical operations (e.g., matrix transformations). Other game engine segments, requiring fewer but more complex operations, may be provisioned with processing entities associated with one or more more powerful central processing units (CPUs).

[0077] By distributing the game engine, the game engine is given flexible computational characteristics that are not constrained by the capabilities of a physical server unit. Instead, the game engine is provisioned with as many or as few computational nodes as needed to meet the demands of the video game. From the perspective of the video game and the video game player, a game engine that is distributed across multiple computational nodes is indistinguishable from a non-distributed game engine running on a single processing entity. This is because a game engine manager or supervisor distributes the workload and seamlessly integrates the results to provide the video game output component to the end user.

[0078] Users access the remote service through a client device that includes at least a CPU, a display, and I / O. The client device may be a PC, a mobile phone, a netbook, a PDA, etc. In one embodiment, a network running on the game server recognizes the type of device used by the client and adjusts the communication method employed. In other cases, the client device accesses the application on the game server over the Internet using a standard communication method, such as HTML. It should be recognized that a given video game or game application may be developed for a specific platform and a specific associated controller device. However, when such a game is made available through a game cloud system such as that presented herein, users may access the video game with a different controller device. For example, a game may be developed for a gaming console and its associated controller, while a user may access a cloud-based version of the game from a personal computer using a keyboard and mouse. In such a scenario, the input parameter configuration may define a mapping from inputs that can be generated by the user's available controller device (in this case, the keyboard and mouse) to inputs that are acceptable for execution of the video game.

[0079] In another example, a user may access the cloud gaming system via a tablet computing device, a touchscreen smartphone, or other touchscreen-driven device. In this case, the client device and the controller device are integrated together on the same device, and input is provided by detected touchscreen input / gestures. For such a device, the input parameter settings may define specific touchscreen inputs that correspond to game inputs for the video game. For example, buttons, directional pads, or other types of input elements may be displayed or overlaid while the video game is running to indicate locations on the touchscreen that the user can touch to generate game input. Gestures, such as swipes in specific directions or specific touch actions, may also be detected as game inputs. In one embodiment, to familiarize the user with control operations on the touchscreen, a tutorial showing how to input gameplay via the touchscreen may be provided to the user, for example, before beginning gameplay of the video game.

[0080] In some embodiments, the client device serves as a connection point for the controller device. That is, the controller device communicates with the client device via a wireless or wired connection and transmits inputs from the controller device to the client device. The client device, in turn, may process these inputs and then transmit the input data to a cloud gaming server over a network (e.g., accessed through a local network device such as a router). However, in other embodiments, the controller itself may be a network device capable of communicating inputs directly to the cloud gaming server over a network without requiring such inputs to first be communicated via the client device. For example, the controller may connect to a local network device (such as the aforementioned router) and send data to and receive data from the cloud gaming server. Thus, while the client device may still be required to receive video output from the cloud-based video game and render it on a local display, input latency may be reduced by allowing the controller to send inputs directly over the network to the cloud gaming server, bypassing the client device.

[0081] In one embodiment, networked controllers and client devices can be configured to transmit certain types of input directly from the controller to the cloud gaming server, while transmitting other types of input via the client device. For example, inputs whose detection does not rely on any additional hardware or processing, aside from the controller itself, can be transmitted directly from the controller to the cloud gaming server over the network, bypassing the client device. Such inputs can include button inputs, joystick inputs, built-in motion-sensing inputs (e.g., accelerometers, magnetometers, gyroscopes), etc. However, inputs that utilize additional hardware or require processing by the client device can be transmitted by the client device to the cloud gaming server. These may include video or audio captured from the game environment, which may be processed by the client device before transmission to the cloud gaming server. Additionally, inputs from the controller's motion-sensing hardware may be processed by the client device in conjunction with the captured video to detect the position and movement of the controller, which will then be communicated by the client device to the cloud gaming server. It should be appreciated that controller devices according to various embodiments can also receive data (e.g., feedback data) from the client device or directly from the cloud gaming server.

[0082] In one embodiment, various example technologies can be implemented using a virtual environment via a head-mounted display (HMD). An HMD may also be referred to as a virtual reality (VR) headset. As used herein, the term "virtual reality" (VR) generally refers to user interaction with a virtual space / environment, including viewing the virtual space via an HMD (or VR headset) in a way that responds in real time to the HMD's movements (controlled by the user) to provide the user with the sensation of being in the virtual space or metaverse. For example, when facing a given direction, the user may see a three-dimensional (3D) view of the virtual space, and when the user turns to the side, a view of the virtual space to that side is rendered on the HMD. The HMD can be worn similarly to glasses, goggles, or a helmet and is configured to display video games or other metaverse content to the user. The HMD can provide a highly immersive experience to the user by providing a display mechanism in close proximity to the user's eyes. Thus, an HMD can provide each of the user's eyes with a display area that occupies most, or even the entire, of the user's field of view, and can also provide viewing with three-dimensional depth and perspective.

[0083] In one embodiment, the HMD may include an eye-tracking camera configured to capture images of the user's eyes while the user interacts with the VR scene. The gaze information captured by the eye-tracking camera(s) may include information related to the user's gaze direction and particular virtual objects and content items in the VR scene that the user is focusing on or interested in interacting with. Thus, based on the user's gaze direction, the system may detect particular virtual objects and content items, such as game characters, game objects, game items, etc., that may be of potential interest to the user if the user is interested in interacting and participating.

[0084] In some embodiments, the HMD may include outward-facing camera(s) configured to capture images of the user's real-world space, such as the user's body movements, and images of any real-world objects that may be located in the real-world space. In some embodiments, images captured by the outward-facing cameras may be analyzed to determine the location / orientation of the real-world objects relative to the HMD. Using the known position / orientation of the HMD, the real-world objects, and inertial sensor data, the user's gestures and movements may be continuously monitored and tracked while the user interacts with the VR scene. For example, while interacting with a game scene, the user may perform various gestures, such as pointing at and walking toward a particular content item in the scene. In one embodiment, the gestures may be tracked and processed by the system to generate a prediction of an interaction with a particular content item in the game scene. In some embodiments, machine learning may be used to facilitate or assist the prediction.

[0085] Various types of single-handed and dual-handed controllers can be used during use of the HMD. In some implementations, the controllers themselves can be tracked by tracking lights included on the controllers or by tracking shape, sensor, and inertial data associated with the controllers. These various types of controllers, or even simple hand gestures recognized and captured by one or more cameras, can be used to interface, control, manipulate, interact with, and participate in the virtual reality environment or metaverse rendered in the HMD. In some cases, the HMD can be wirelessly connected to a cloud computing and gaming system through a network. In one embodiment, the cloud computing and gaming system maintains and executes the video game being played by the user. In some embodiments, the cloud computing and gaming system is configured to receive input from the HMD and interface objects over the network. The cloud computing and gaming system is configured to process the input to affect the game state of the running video game. Output from the running video game, such as video data, audio data, and haptic feedback data, is transmitted to the HMD and interface objects. In other embodiments, the HMD may communicate with the cloud computing and gaming system wirelessly via alternative mechanisms or channels, such as a cellular network.

[0086] Additionally, while embodiments of the present disclosure may be described with reference to a head-mounted display, it will be appreciated that in other embodiments, a non-head-mounted display may be used instead, including, but not limited to, a portable device screen (e.g., a tablet, smartphone, laptop, etc.) or any other type of display that can be configured to render video and / or provide a display of an interactive scene or virtual environment in accordance with the present embodiments. It should be understood that the various embodiments defined herein may be combined or assembled into specific embodiments using various features disclosed herein. Thus, the examples provided are merely some possible examples and are not intended to be limiting of various embodiments, where many more embodiments may be defined by combining various elements. In some examples, some embodiments may include fewer elements without departing from the spirit of the disclosed embodiments or equivalent embodiments.

[0087] Embodiments of the present disclosure may be practiced with a variety of computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. Embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a wire-based or wireless network.

[0088] Although the operations of the method have been described in a particular order, it should be understood that other housekeeping operations may be performed between operations, or operations may be coordinated to occur at slightly different times, or operations may be distributed in a system that allows processing operations to occur at various intervals associated with the processing, so long as the processing of the telemetry data and game state data to generate the modified game state is performed in the desired manner.

[0089] One or more embodiments may also be fabricated as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device that can store data, which can thereafter be read by a computer system. Examples of computer-readable media include hard drives, network-attached storage (NAS), read-only memory, random-access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tape, and other optical and non-optical data storage devices. The computer-readable medium may include computer-readable tangible media distributed across network-connected computer systems, whereby the computer-readable code is stored and executed in a distributed fashion.

[0090] In one embodiment, a video game is executed either locally on a gaming console, personal computer, or on a server. In some cases, the video game is executed by one or more servers in a data center. When a video game is executed, some instances of the video game may be a simulation of the video game. For example, the video game may be executed by an environment or server that generates a simulation of the video game. A simulation, in some embodiments, is an instance of the video game. In other embodiments, a simulation may be generated by an emulator. In either case, when a video game is represented as a simulation, the simulation may be executed to render interactive content that can be interactively streamed, executed, and / or controlled by user input.

[0091] Although the foregoing embodiments have been described in some detail for purposes of clarity of understanding, it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims. The present embodiments are therefore to be considered as illustrative and not restrictive, and the present embodiments are not limited to the details given herein, but can be modified within the scope of the appended claims and their equivalents.

Claims

1. receiving a game image, the game image being captured from gameplay of a video game, the game image depicting a scene from the video game; receiving game state data describing attributes of the scene of the video game to be depicted in the game image; receiving, over a network from a client device, modification data describing changes to the game image, the modification data being defined from user input received at the client device; applying the game image, the game state data, and the user input to generate an AI-generated image by an image generating artificial intelligence (AI); transmitting the AI-generated image over the network to the client device for rendering on a display; The method comprising:

2. The method of claim 1 , wherein the applying the game state data enables a semantic understanding of the scene depicted in the game image by the image generation AI.

3. The method of claim 2 , wherein the semantic understanding of the scene is applied by the image generation AI to perform the changes described in the modification data.

4. The method of claim 1 , wherein the game state data identifies one or more elements in the scene to be depicted in the game image.

5. The method of claim 1 , wherein the game state data describes the depth of one or more virtual objects in the scene.

6. 6. The method of claim 5, wherein the correction data describes a change in location of a given virtual object in the scene, and wherein generating the AI-generated image is configured to use the depth of the one or more virtual objects to implement the change in location of the given virtual object described by the correction data.

7. 6. The method of claim 5, wherein the correction data describes a placement of a given virtual object in the scene, and wherein generating the AI-generated image is configured to use the depth of the one or more virtual objects to implement the placement of the given virtual object described by the correction data.

8. 6. The method of claim 5, wherein the depth of the one or more virtual objects is configured such that performing the changes described by the modification data allows for proper occlusion of or by the one or more virtual objects.

9. The method of claim 1 , wherein the game state data describes a three-dimensional structure of one or more virtual objects in the scene.

10. The method of claim 1 , wherein the correction data is defined by a word or phrase generated by the user input received at the client device.

11. A non-transitory computer-readable medium having program instructions embodied therein, the program instructions being configured, when executed by at least one server computer, to cause the at least one server computer to perform a method, the method comprising: receiving a game image, the game image being captured from gameplay of a video game, the game image depicting a scene from the video game; receiving game state data describing attributes of the scene of the video game to be depicted in the game image; receiving, over a network from a client device, modification data describing changes to the game image, the modification data being defined from user input received at the client device; applying the game image, the game state data, and the user input to generate an AI-generated image by an image generating artificial intelligence (AI); transmitting the AI-generated image over the network to the client device for rendering on a display; The non-transitory computer-readable medium comprising:

12. 12. The non-transitory computer-readable medium of claim 11, wherein the applying the game state data enables the image generation AI to semantically understand the scene depicted in the game image.

13. 13. The non-transitory computer-readable medium of claim 12, wherein the semantic understanding of the scene is applied by the image generation AI to perform the changes described in the modification data.

14. The non-transitory computer-readable medium of claim 11 , wherein the game state data identifies one or more elements in the scene to be depicted in the game image.

15. The non-transitory computer-readable medium of claim 11 , wherein the game state data describes a depth of one or more virtual objects in the scene.

16. 16. The non-transitory computer-readable medium of claim 15, wherein the correction data describes a change in location of a given virtual object in the scene, and wherein generating the AI-generated image is configured to use the depth of the one or more virtual objects to implement the change in location of the given virtual object described by the correction data.

17. 16. The non-transitory computer-readable medium of claim 15, wherein the correction data describes a placement of a given virtual object in the scene, and wherein generating the AI-generated image is configured to use the depth of the one or more virtual objects to implement the placement of the given virtual object described by the correction data.

18. 16. The non-transitory computer-readable medium of claim 15, wherein the depth of the one or more virtual objects is configured such that implementing the changes described by the modification data allows for proper occlusion of or by the one or more virtual objects.

19. The non-transitory computer-readable medium of claim 11 , wherein the game state data describes a three-dimensional structure of one or more virtual objects in the scene.

20. The non-transitory computer-readable medium of claim 11 , wherein the correction data is defined by a word or phrase generated by the user input received at the client device.

Citation Information

Patent Citations

  • Method and system for low-latency transport protocols

    JP2013506348A

  • Hardware acceleration and event decisions for late latch and warp in interactive computer products

    US20210106912A1

  • Depth-Aware Photo Editing

    US20210304431A1

  • Image generation using one or more neural networks

    US20220068037A1