Use of game state data for semantic understanding of AI image generation models

Through the combination of dynamic user interface and game state data, image features are analyzed and modified, the problem that AI image generation model generation does not meet user intentions is solved, and more efficient image generation consistency and user experience are achieved.

CN120529948APending Publication Date: 2025-08-22SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380092103.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-10
Filing Date
2023-12-21
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

When the existing AI image generation model responds to user input, the generated images often do not meet the user's intentions, and lacks effective feedback mechanisms and training methods to improve the consistency of image generation.

Method used

By providing a dynamic user interface, analyzing image features and allowing users to modify, combining game status data and user input, AI models are trained to generate images that are more in line with user intentions.

Benefits of technology

It improves the user experience of the AI ​​image generation model, ensures that the generated images are more in line with user expectations and preferences, and enhances the model's self-learning and adjustment capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120529948A_ABST
    Figure CN120529948A_ABST
Patent Text Reader

Abstract

There is provided a method comprising: receiving a game image, the game image being captured from a game play of a video game, and the game image depicting a scene of the video game; receiving game state data describing attributes of the scene of the video game depicted in the game image; receiving modification data describing a change to the game image from a client device over a network, the modification data being defined according to user input received at the client device; generating, by an image generation artificial intelligence (AI), an AI generated image by applying the game image, the game state data, and the user input; the AI generated image is transmitted over the network to the client device for rendering to a display.
Need to check novelty before this filing date? Find Prior Art

Description

1. Technical Field

[0002] The present disclosure generally relates to methods, systems, and apparatus for using game state data for semantic understanding by an AI image generation model, and a dynamic interface for providing feedback on the output of the AI ​​image generation model. Background Art

[0003] 2. Description of Related Technologies

[0004] The video game industry has undergone many changes over the years. As technology advances, video games continue to achieve greater immersion through sophisticated graphics, realistic sounds, captivating soundtracks, haptics, and more. Players are able to enjoy immersive gaming experiences in which they engage with and become part of the virtual environment, exploring new ways to interact. Furthermore, players can stream their gameplay for viewers to watch, allowing others to share in the gameplay experience.

[0005] It is in this context that the implementation of the present disclosure is proposed. Summary of the Invention

[0006] Implementations of the present disclosure include methods, systems, and apparatus for using game state data for semantic understanding by an AI image generation model, and a dynamic interface for providing feedback on the output of the AI ​​image generation model.

[0007] In some implementations, a method is provided, comprising: receiving a game image captured from gameplay of a video game, the game image depicting a scene of the video game; receiving game state data describing attributes of the scene of the video game depicted in the game image; receiving modification data describing changes to the game image from a client device over a network, the modification data being defined based on user input received at the client device; applying the game image, the game state data, and the user input by image-generating artificial intelligence (AI) to generate an AI-generated image; and transmitting the AI-generated image to the client device over the network for rendering to a display.

[0008] In some implementations, the application of the game state data enables the image generation AI to perform semantic understanding of the scene depicted in the game image.

[0009] In some implementations, the semantic understanding of the scene is applied by the image generation AI to perform the changes described in the modification data.

[0010] In some implementations, the game state data identifies one or more elements of the scene depicted in the game image.

[0011] In some implementations, the game state data describes a depth of one or more virtual objects in the scene.

[0012] In some implementations, the modification data describes a change in position of a given virtual object within the scene, and wherein generating the AI-generated image is configured to use the depth of the one or more virtual objects to perform the change in position of the given virtual object described by the modification data.

[0013] In some implementations, the modification data describes the placement of a given virtual object within the scene, and wherein generating the AI-generated image is configured to use the depth of the one or more virtual objects to perform the placement of the given virtual object described by the modification data.

[0014] In some implementations, the depth of the one or more virtual objects is configured to enable the one or more virtual objects to properly occlude or be properly occluded when performing the change described by the modification data.

[0015] In some implementations, the game state data describes a three-dimensional structure of one or more virtual objects in the scene.

[0016] In some implementations, the modification data is defined by a word or phrase generated by the user input received at the client device.

[0017] In some implementations, a non-transitory computer-readable medium having program instructions embodied thereon is provided, the program instructions being configured to, when executed by at least one server computer, cause the at least one server computer to perform a method comprising: receiving a game image, the game image being captured from gameplay of a video game and depicting a scene of the video game; receiving game state data describing attributes of the scene of the video game depicted in the game image; receiving modification data describing changes to the game image from a client device over a network, the modification data being defined based on user input received at the client device; generating an AI-generated image by image-generating artificial intelligence (AI) applying the game image, the game state data, and the user input; and transmitting the AI-generated image to the client device over the network for rendering to a display.

[0018] Other aspects of the present disclosure will become apparent from the following detailed description read in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present disclosure may be better understood with reference to the following description in conjunction with the accompanying drawings, in which:

[0020] Figure 1 An image generation service providing a user interface (UI) for modifying an image according to an implementation of the present disclosure is conceptually illustrated.

[0021] Figure 2 The diagram conceptually illustrates the generation of an image based on a game image and related game state information by an image generation AI according to an implementation of the present disclosure.

[0022] Figure 3 Conceptually illustrates the extraction of features from a game-generated scene for use as input to an image-generating AI, according to implementations of the present disclosure.

[0023] Figure 4 A system for storing configuration files for interpreting user input for AI image generation according to implementations of the present disclosure is conceptually illustrated.

[0024] Figure 5 The present invention conceptually illustrates crowdsourcing topics for image generation AI to generate images according to an implementation of the present disclosure.

[0025] Figure 6A 、 Figure 6B and Figure 6C Shown is a general representation of an Image Generation AI (IGAI) processing sequence according to an implementation of the present disclosure.

[0026] Figure 7 Components of an exemplary apparatus that may be used to perform aspects of various embodiments of the present disclosure are shown. DETAILED DESCRIPTION

[0027] The following implementations of the present disclosure provide methods, systems, and apparatus for a dynamic interface for using game state data for semantic understanding by an AI image generation model, and for providing feedback on the output of the AI ​​image generation model.

[0028] When AI models automatically generate images in response to user input, the resulting images may not match the user’s preferences. In order to modify images or provide dynamic feedback to train AI models to generate images that match user preferences, Figure 1A UI is dynamically generated for a consistent image. The UI features a selection interface that is dynamically targeted to specific features shown in the image. For example, if a person is in the scene of the generated image, the user interface can automatically identify the person and suggest options for modification. This modification feedback is then fed back to the AI ​​model for training and subsequent generation of additional images that are more consistent with the feedback. In one embodiment, the image is broken down into layers so that the AI ​​analysis model can automatically identify the layers. For example, the layers may include background, foreground, midground, and isolation of specific objects in the image for selective modification or removal. In one embodiment, the feedback provided by the UI enables the user to provide guidance on how the image should be modified. Modifications can include not only content but also the angle from which the image was taken, similar to telling a photographer to take different images from different perspectives. The AI ​​model then processes this feedback as training information to generate new images that are more consistent with the user's requested intent. In some embodiments, the user's intent itself can be analyzed to provide a user profile, such as to identify the user's preferences and preferences, to facilitate future modifications to the automatically generated image.

[0029] With the above overview in mind, several example figures are provided below to facilitate an understanding of example embodiments.

[0030] Figure 1 An image generation service providing a user interface (UI) for modifying an image according to an implementation of the present disclosure is conceptually illustrated.

[0031] In the illustrated implementation, image generation service 100 includes image generation artificial intelligence 102 configured to generate images in response to user input. Image generation service 100 is accessible by client device 100 via a network 108 (e.g., including the Internet). By way of example, but not limitation, client device 100 may be a personal computer, laptop, tablet, mobile device, cell phone, game console, set-top box, streaming media box, or any other type of computing device capable of performing the functions described herein. Client device 110 executes application 112 that accesses image generation service 100 via network 108. In some implementations, application 112 is a web browser, and image generation service 100 is accessible as a website on the Internet. In other implementations, application 112 is a dedicated application or app executed by the client device that communicates with image generation service 100, such as by accessing an application programming interface (API) exposed by image generation service 100.

[0032] The application 112 renders a user interface 114 through which a user 116 interacts with the image generation service 100. For example, through the user interface 114, the user 116 can provide user input such as descriptive text or an image, which the image generation AI 102 uses to generate an image. It will be understood that when the image generation AI 102 automatically generates an image in response to user input, the resulting image may not be to the user's liking. Therefore, in order to modify the image or provide dynamic feedback to train the image generation AI to generate an image that is consistent with the user's wishes, the image generation AI 102 can be used to generate an image that is consistent with the user's wishes. Figure 1 In some implementations, the image generation service 100 dynamically generates a modification UI for the image and presents it as part of the UI 114. In some implementations, the image generation service 100 further includes a feature analyzer 104 configured to analyze the image to identify features of the image that the user may wish to modify. For example, the feature analyzer 104 may identify various elements or objects within the image, and based on the identification, the modification logic 106 determines possible modifications that can be suggested to the user via the modification UI. In some implementations, the feature analyzer 104 uses a recognition model to identify features in the image.

[0033] In this way, the modification UI can provide a selection interface that is dynamically locked to a specific feature shown in the image. For example, if a person 120 is in the scene of the generated image 118, the feature analyzer 104 can automatically identify the person 120, and the modification logic 106 can suggest options for modification, such as making the person taller, shorter, adjusting the person's clothing, etc. In some implementations, the system can recognize the user's head 122 and suggest specific modifications, such as changing the expression on the person's face, hair color, etc. In some implementations, the system can recognize the tree 124 in the image 118 and suggest modifications, such as making the tree shorter or taller, wider, more or less green, more or less leaves, with flowers, with fruit, etc.

[0034] In some implementations, a selection tool is provided in the modification UI, which allows the user to identify an area in the image to be modified, such as by drawing a box or circling the area using a drawing tool or a predefined shape tool, shading the area using a brush tool, etc. The feature analyzer 104 analyzes the identified area to determine the content of the area, and the modification logic 106 suggests modifications based on the identified content of the area.

[0035] The user 116 may then select one or more of the suggested modifications. Furthermore, the modification UI may enable the user 116 to enter additional modifications or instructions for how to change the image, such as by entering text describing the additional modifications. The selected or entered modifications are returned as feedback to the image generation AI for training and subsequent generation of additional images that are more consistent with the feedback.

[0036] In one embodiment, an image is analyzed and broken down into layers, which can be automatically identified by an AI analysis model. For example, layers may include background, foreground, midground, and isolation of specific objects in the image for selective modification or removal. Thus, a user can identify an object relative to its location in the scene and use this information to issue modification commands (e.g., move an object from the foreground to the background), or modify entire areas of the scene, such as the background, foreground, midground, or any identified layer.

[0037] Furthermore, in one embodiment, the modification may include not only the content, but also the angle from which the image was taken, similar to telling a photographer to take a different image from a different perspective. For example, the user input may include instructions to make an image from a different perspective or using different optics, such as closer, farther, rotated, lower, higher, overhead, left, right, wider / narrower angles, zoom, etc. The image generation AI model then processes this feedback as training information to subsequently generate new images that are more consistent with the user's request intent.

[0038] In some embodiments, the user's intent itself may be analyzed to provide a profile of the user, such as to identify the user's preferences and likes for future modifications to automatically generated images.

[0039] Figure 2 The diagram conceptually illustrates the generation of an image based on a game image and related game state information by an image generation AI according to an implementation of the present disclosure.

[0040] In some implementations, a user 200 playing a video game 204 executed by a gaming machine 202 can capture a game image 206 from their game play. In local gaming implementations, the gaming machine 202 is a local device (e.g., a computer, game console, etc.), while in cloud gaming implementations, the gaming machine 202 is a remote device, such as a server computer / blade that executes the video game 204 and streams the game play video to the user's local client device (not shown). In some implementations, the game image 206 is a still image from the user's game play. In other implementations, the game image 206 may be a video clip from the user's game play.

[0041] Game image 206 may be provided as input to image generation AI 102 to generate an image. Further, a user may provide user input in the form of modification data 210 indicating how the user wishes to modify game image 206 or otherwise use game image 206 as a seed image to generate a new AI-generated image 212. In some implementations, to provide further semantic understanding of the content of game image 206, game state data 208 is also provided as input to image generation AI 102. Game state data 208 includes data describing the state of the virtual environment of video game 204 at the time game image 206 was captured during game play.

[0042] It will be appreciated that the game state data 208 may describe various aspects of the scene depicted in the game image 206, such as identification of objects / elements, depth of objects, movement of objects occurring in the scene, audio playing at the time of image capture, audio associated with specific objects, words spoken by characters, 3D structure of objects, information about occluded objects in the scene, lighting information, physics information, etc. By using such game state information as input to the image generation AI 102, this provides an improved understanding of the content of the game image, and therefore an improved semantic understanding of the modifications / changes described by the user.

[0043] By way of example, but not limitation, the game state data 208 may include depth information about objects in the game image 206. Depth information enables understanding the relative positioning of objects in the scene in three dimensions. Consequently, when the user modification data 210 includes user instructions for moving an object, this movement can be understood relative to other objects in three dimensions. For example, a user may input a statement such as "put a plant behind the couch," and the image generation AI will understand the depth of the couch in the scene and place the plant at the correct depth relative to the couch and other objects in the scene. Alternatively, if the user inputs a direction to move an object, the object can be moved and repositioned at the correct depth relative to other objects or elements. For example, the instruction "move that tree to the left" can be understood using correct depth information so that the tree's movement and placement does not obscure a person walking their dog in the scene, who is at a closer depth than the tree.

[0044] In a related implementation, additional game state information can provide further improvements in image generation. For example, by using information describing occluded objects or occluded portions of objects in the game image 206, when the object moves, the previously occluded object or occluded portion can be presented by the image generation AI 102 and included in the AI-generated image 212. As another example, by using information describing the 3D structure of an object, the object can then be moved or placed while respecting the depth boundaries of other objects, for example, so that the object does not appear too close in front of or behind another object in the scene. Thus, by providing additional semantic understanding of what is in the scene in the game image 206, the image generation AI 102 is better able to handle user prompts or input describing modifications or changes that the user wishes to make.

[0045] Figure 3 Conceptually illustrates the extraction of features from a game-generated scene for use as input to an image-generating AI, according to implementations of the present disclosure.

[0046] In some implementations, the game image 300 is provided to provide stylistic or artistic information to the image generation AI 102 for generating the AI-generated image 306. To facilitate understanding the stylistic / artistic elements of the game image 300, the image analyzer 304 is configured to analyze the game image 300 to determine the stylistic / artistic elements of the game image 300. For example, in some implementations, the image analyzer 304 is configured to analyze the lighting in the scene depicted in the game image 300, which may include analyzing the source of light, the location of the light source, the color temperature, the intensity, the contrast, etc. In some implementations, the image analyzer 304 is configured to analyze other artistic aspects of the game image 300, such as the color palette used, the type of lines used to delineate the boundaries of objects, the type of texture or shading, etc.

[0047] The extracted artistic / stylistic information may be provided to the image generation AI 102 as additional input for influencing the generation of the AI-generated image 306. For example, the modification data 302 may reference stylistic elements of the game image 300, and such stylistic elements may be determined by the image analyzer 304 and used as input to the image generation AI 306. For example, the modification data 302 may include instructions to generate an image with lighting similar to that of the game image 300. The image analyzer 304 may analyze the lighting of the game image 300 and thereby generate lighting information describing the lighting of the game image 300, and such lighting information may be used by the image generation AI to generate the AI-generated image 306 to have similar lighting. In this example, the AI-generated image 306 may have lighting with a similar color temperature to the game image 300, or a light source in a similar position, etc.

[0048] In some implementations, the image analyzer 304 is triggered to analyze a given stylistic aspect of the game image 300 in response to a reference to such a stylistic aspect of the game image in the modification data. For example, in the above-described embodiment involving lighting, the image analyzer 304 can be triggered to analyze the lighting of the game image 300 in response to a user input indicating a reference to the lighting of the game image 300. In some implementations, the image analyzer 304 is a recognition model configured to recognize stylistic or artistic elements of an image.

[0049] In this manner, the game image 300 is used as a reference image to provide stylistic input for the image generation AI 102. It will be appreciated that a user may not wish to simply have the AI ​​generate an image 306 entirely in the style of the game image 300, but rather may wish to apply only certain stylistic elements. Therefore, the current implementation is capable of selectively using stylistic elements from the game image 300 for image generation, as specified by the modification data 302.

[0050] In some implementations, the image generation AI 102 generates the AI-generated image 306 based on input provided in the modification data 302 and incorporates stylistic elements from the game image 300 as described herein. In other implementations, the image generation AI 102 generates the AI-generated image 306 using another image, such as previously described, and modification data 302 describing modifications / changes to that image and incorporates stylistic elements from the game image 300 as described herein.

[0051] In some implementations, the image analyzer 304 may also analyze the stylistic / artistic elements of the game image 300 using the game state data and information described above.

[0052] While game image 300 is described in the current embodiment, in other embodiments, other types of images may be utilized and analyzed for stylistic / artistic elements according to implementations of the present disclosure.

[0053] Figure 4 A system for storing configuration files for interpreting user input for AI image generation according to implementations of the present disclosure is conceptually illustrated.

[0054] In some implementations, the system is configured to use words used as input to the image generation AI 102 to learn the user's understanding and intent. This understanding can define a profile for the user, which is used to interpret the user's input words. In the illustrated implementation, a profile store 402 is provided in which user profiles are stored. It will be understood that a given user can have more than one profile, such as profile 404 and profile 406, to enable different understandings of the user's intent to be used. It will be understood that a user may wish to create different profiles to generate images with different styles or elements based on different learned understandings of user preferences associated with different profiles. In some implementations, a given profile maps a word to one or more other words and / or data that defines the semantic understanding of a word determined for a particular profile.

[0055] In some implementations, onboarding logic 400 is configured to provide an onboarding process in which a user can indicate their preferred usage or understanding of certain words. For example, in some implementations, onboarding logic 400 presents multiple images via UI 114 and may ask the user to describe the images, or may ask the user to associate the images with certain predefined terms that the user understands, and in this way, may learn the user's understanding of the language used to describe the images. In some implementations, the description or indication of the user's understanding of the images is mapped to known input words or phrases that were used to generate the images. This learned information about the user's preferred usage or understanding of words or phrases is stored in a given profile, such as profile 404 in the illustrated implementation. It will be understood that onboarding logic 400 is useful for the initial setup of a given profile, but can also be used as a training tool at any time to provide explicit training on the user's understanding and intent with respect to words and images.

[0056] Then, when the user enters user input 412 to generate an image, in the illustrated implementation, such input is processed by the interpreter 408 based on the currently active profile 404, and the user input 412 is interpreted based on the profile 404. In some implementations, the interpreter 408 uses the profile 404 to translate the user input 412 into translated input 410, which is fed to the image generation AI 102. For example, in some implementations, the profile 404 is used to map words or phrases found in the user input 412 to other words or phrases, thereby including them in the translated input 410. For example, the user input 412 may include the word "dark", and based on the user's active profile 404, the word "dark" is mapped to additional words / phrases such as "fantasy", "HR Giger", etc., and thus the interpreter generates the translated input 410 to include one or more of these additional words / phrases.

[0057] It will be appreciated that because many words or phrases are subjective or open to interpretation, the profiles of the current implementation provide a method for learning, storing, and applying the user's subjective understanding of the meaning of these words or phrases in order to obtain results from the image generation AI 102 that are more consistent with the user's expectations. In addition, the user's profile can be further learned through the use of the system over time. For example, the profile logic 416 can be configured to analyze the user input 412 and can associate the words / phrases used by the user with their profile, such as words that the user uses repeatedly, or words that the user uses in clusters or tends to use in groups, etc.

[0058] In some implementations, the user provides modification input 414 in response to the generated image 418 from the image generation AI 102. As described above, the modification input 414 can indicate changes that the user wishes to make to the image 418 and can provide insight into the user's original intent in generating the image 418 using the original user input 412. Accordingly, in some implementations, the profile logic 416 analyzes the modification input 414 to further determine the user's understanding of the words supplied in the user input 412 and stores that understanding in the active profile 404. For example, in some implementations, the words provided in the modification input 414 can be mapped or associated with the words provided in the user input 412 and stored in the active profile 404.

[0059] In some implementations, the system can be configured to suggest words or phrases when the user generates their user input 412. In some implementations, in response to a given word or phrase provided by the user, multiple possible related words or phrases are then suggested, and the user can select one or more of the suggestions. Based on the user's selections in such instances, the user's activity profile 404 can be updated over time, such as by associating or mapping words with each other based on such selections.

[0060] In some implementations, a given profile may define a learning model that is trained to predict or infer a user's preferred words / phrases based on a given supplied word or phrase. The learning model is trained using any of the techniques described herein and data describing the user's understanding of words, terms, phrases, etc. In some implementations, the learning model is configured to associate or map or cluster various words or phrases, and these associations strengthen or weaken over time as a result of training. The interpreter 408 uses the trained learning model to generate a predicted word based on the user input 412, which may be appended to the user input 412 or otherwise included to generate the translated input 410 fed to the image generation AI 102.

[0061] In some implementations, a given profile is configured to calibrate terms of degree as they are used by the user. For example, a user's use of the term "tall" may be equivalent to the image generation AI 102 applying "really tall" to achieve the user's preferred results. Thus, the profile system of the current implementation may be configured to learn the user's preferences in this regard.

[0062] In some implementations, profile logic 416 may use additional signals to further refine a given profile. For example, in some implementations, image generation AI 102 may generate multiple images based on a given user input, and the user may select which one most closely matches their intent. Such user selections may be used as feedback to adjust the user's profile. In further implementations, selections of additional features after image generation (such as selecting an image to upgrade, or re-running image generation based on a given selected image) may also be used as feedback.

[0063] It will be appreciated that a challenge in using image generation AI systems is that it is difficult for users to provide the correct input that achieves their desired outcome. Therefore, by implementing a system that learns the user's preferences and understanding of input terms, users are able to achieve improved image generation more efficiently.

[0064] In some implementations, a theme can be defined for a given user or a given profile. A theme can be configured to define a specific style and, therefore, include certain words / phrases or other types of acceptable input that, when applied to the image generation AI, will cause the image generation AI to generate an image in that specific style. In some implementations, the theme is editable, allowing the user to specify specific words / phrases or other specific inputs as part of the theme definition. Then, when user input is entered to generate an image, the theme is applied by appending the words / phrases / inputs stored in the theme.

[0065] Figure 5 The present invention conceptually illustrates crowdsourcing topics for image generation AI to generate images according to an implementation of the present disclosure.

[0066] In some implementations, trending or popular inputs for image generation by the image generation AI are determined and used to crowdsource topics that can be applied by subsequent users of the image generation AI. In the illustrated implementation, individual users 500a, 500b, 500c, 500d, etc. generate user inputs 502a, 502b, 502c, 502d, etc., respectively. The user inputs are provided for the purpose of generating an image by the image generation AI, as described.

[0067] In some implementations, the trend analyzer 504 analyzes user input and identifies popular or trending inputs based on the user inputs. For example, the trend analyzer 504 can identify popular or trending words, terms, phrases, or other inputs being input by users of the system. Based on these trending inputs, the topic generator 506 is configured to generate one or more topics that include or are otherwise defined by a set of popular or trending inputs. For example, a given topic can be defined to include a particular set of words that users tend to use together.

[0068] It will be appreciated that by analyzing input terms across many users, certain relationships among the various input terms may be discovered that would not otherwise be apparent. For example, some words that may appear unrelated to each other may be found to appear in conjunction with each other with a high degree of regularity in user input.

[0069] It will be appreciated that there may be a library of themes such as themes 510, 512, etc. stored in theme storage 508 from which users may select a given theme to apply to their image generation instance. By enabling selectable themes according to implementations of the present disclosure, users may be able to more quickly generate images having a particular style or appearance.

[0070] In one embodiment, an image generation AI (IGAI) that generates output images, graphics, and / or three-dimensional representations may include one or more artificial intelligence processing engines and / or models. Generally, AI models are generated using training data from a dataset. The dataset selected for training may be custom-curated for a specific desired output, and in some cases, the training dataset may include a wide range of general-purpose data available from a variety of sources on the internet. For example, the IGAI may have access to a vast amount of data, such as images, videos, and three-dimensional data. The IGAI uses this general-purpose data to understand the type of content expected by the input. For example, if the input requests the generation of a tiger in the Sahara Desert, the dataset should contain various images of tigers and deserts to access and draw upon during the processing of the output image. Alternatively, a curated dataset may be more specifically targeted at a particular type of content, such as video game-related art, videos, and other asset-related content. Even more specifically, the curated dataset may include images relevant to a particular scene or action sequence in a game, including game assets such as unique avatar characters. As described above, the IGAI can be customized to allow the input of unique descriptive language statements to set the style of the requested output image or content. The descriptive language statement can be text or other sensory input, such as inertial sensor data, input speed, emphasis, and other data that can form an input request. IGAI can also provide an image, video, or image set to define the context of the input request. In one embodiment, the input can be text describing the desired output and one or more images to convey the desired context of the requested output.

[0071] In one embodiment, IGAI is provided to implement text-image generation. Image generation is configured to implement potential diffusion processing to synthesize text into image processing in a latent space. In one embodiment, the adjustment process helps, for example, to shape the output into the desired output using structured metadata. The structured metadata may include information obtained from user input to guide the machine learning model to use cross-attention to gradually denoise in stages until the processed denoising is decoded back into pixel space. In the decoding stage, scaling is used to obtain higher quality images, videos, or 3D assets. Therefore, IGAI is a customized tool that is engineered to process specific types of inputs and render specific types of outputs. When IGAI is customized, machine learning and deep learning algorithms are tuned to achieve specific customized outputs, for example, unique image assets to be used in gaming technology, specific game titles, and / or movies.

[0072] In another configuration, IGAI may be a third-party processor, such as one provided by Stable Diffusion or other companies such as OpenAI's GLIDE, DALL-E, MidJourney, or Imagen. In some configurations, IGAI may be used online via one or more application programming interface (API) calls. It should be understood that references to available IGAI are for informational purposes only. For additional information about IGAI technology, reference may be made to the paper "High-Resolution Image Synthesis with Latent Diffusion Models" published by Ludwig Maximilian University of Munich (authors Robin Rombach et al., pages 1-45). This paper is incorporated by reference.

[0073] Figure 6A 602 processing sequence according to an embodiment. As shown in the figure, input 606 is configured to receive input in the form of data, such as text descriptions or keywords with semantic descriptions. The text description can be in the form of a sentence, for example, having at least one noun and one verb. The text description can also be in the form of a fragment, or just a single word. The text can also be in the form of multiple sentences describing a scene, an action, or a feature. In some configurations, the input text can also be input in a specific order to influence the focus on a word, or even weaken the emphasis on a word, letter, or sentence. Furthermore, the text input can be in any form, including characters, emoticons, symbols, foreign language characters (such as Japanese, Chinese, Korean, etc.). In one embodiment, text description is achieved through contrastive learning. The basic idea is to embed the image and text into a latent space so that the text corresponding to the image is mapped to the same area in the latent space as the image. For example, this abstracts the structure of what it means to be a dog from the visual representation and text representation. In one embodiment, the goal of contrastive representation learning is to learn an embedding space where similar pairs of samples remain close to each other, while dissimilar pairs of samples are far apart. Contrastive learning can be applied to both supervised and unsupervised settings. When working with unsupervised data, contrastive learning is one of the most powerful methods in self-supervised learning.

[0074] In addition to text, input may also include other content, such as images, or even images that themselves have descriptive content. Image analysis can be used to interpret images to identify objects, colors, intent, characteristics, shadows, textures, three-dimensional representations, depth data, and combinations thereof. Broadly speaking, input 606 is configured to convey the user's intention to generate certain digital content using IGAI. In the context of gaming technology, the target content to be generated may be game assets for use in specific gaming scenarios. In such scenarios, the data sets used to train IGAI and input 606 can be used to customize the way artificial intelligence (e.g., deep neural networks) processes data to manipulate and tune the desired output images, data, or three-dimensional digital assets.

[0075] Input 606 is then passed to the IGAI, where an encoder 608 takes the input data and / or pixel-space data and converts it into latent space data. The concept of "latent space" is central to deep learning, as feature data is reduced to a simplified data representation with the goal of finding and exploiting patterns. Therefore, latent space processing 610 is performed on the compressed data, significantly reducing processing overhead compared to processing learning algorithms in pixel space, which are much more resource-intensive and require significantly more processing power and time to analyze and produce the desired image. The latent space is simply a representation of the compressed data, where similar data points are spatially closer. In the latent space, the processing is configured to learn relationships between data points that the machine learning system has been able to derive from information it has access to (e.g., the dataset used to train the IGAI). In the latent space processing 610, a diffusion model is used to calculate the diffusion process. The latent diffusion model relies on an autoencoder to learn a lower-dimensional representation of the pixel space. The latent representation undergoes a diffusion process, with noise added at each step (e.g., multiple stages). The output is then fed into a denoising network based on a U-Net architecture with a cross-attention layer. A conditioning process is also applied to guide the machine learning model to remove noise and arrive at an image that represents something close to what was requested via the user input. Decoder 612 then transforms the resulting output from the latent space back into pixel space. Output 614 can then be processed to improve resolution. Output 614 is then transmitted as a result, which can be an image, a graphic, 3D data, or data that can be rendered in physical or digital form.

[0076] Figure 6BIn one embodiment, additional processing that can be performed on input 606 is shown. User interface tools 620 can be used to enable a user to provide input request 604. As described above, input request 604 can be an image, text, structured text, or general data. In one embodiment, before the input request is provided to encoder 608, the input can be processed by a machine learning process that generates a machine learning model 632 and learns from a training data set 634. For example, the input data can be processed via a context analyzer 626 to understand the context of the request. For example, if the input is "space rockets for flying to Mars," the context analyzer 626 can analyze the input to determine that the context is related to outer space and planets. The context analysis can use the machine learning model 632 and training data set 634 to find relevant images for this context or identify specific art, images, or video libraries. If the input request also includes an image of a rocket, a feature extractor 628 can function to automatically identify characteristic features in the rocket image, such as fuel tanks, length, color, position, edges, letters, flames, etc. Feature classifier 630 can also be used to classify features and improve machine learning model 632. In one embodiment, input data 607 can be generated to produce structured information that can be encoded into the latent space by encoder 608. In addition, it is possible to extract structured metadata 622 from the input request. Structured metadata 622 can be, for example, descriptive text that instructs IGAI 602 to modify a feature, make a change to the input image, or change the color, texture, or a combination thereof. For example, input request 604 may include an image of a rocket, and the text may say "make the rocket wider," "add more flames," or "make it stronger," or some other modifier desired by the user (e.g., semantically provided and contextually analyzed). Structured metadata 622 can then be used in subsequent latent space processing to tune the output toward the user's intent. In one embodiment, the structured metadata can take the form of a semantic graph, text, image, or data engineered to represent the user's intent regarding what changes or modifications should be made to the input image or content.

[0077] Figure 6CThe output of encoder 608 is then fed into latent space processing 610, according to one embodiment. A diffusion process is performed by diffusion process stage 640, where the input is processed through several stages to add noise to the input image or image associated with the input text. This is a gradual process, where noise is added at each stage (e.g., 10-50 or more stages). Next, a denoising process is performed by denoising stage 642. Similar to the denoising stage, the reverse process is performed, where noise is gradually removed at each stage, and at each stage, machine learning is used to predict what the output image or content should be, given the input request intent. In one embodiment, a machine learning model 644 can use structured metadata 622 at each stage of denoising to predict what the resulting denoised image should look like and how it should be modified. During these predictions, machine learning model 644 uses training dataset 646 and structured metadata 622 to move closer and closer to an output that is most similar to the output requested in the input. In one embodiment, a U-Net architecture with a cross-attention layer can be used to improve predictions during denoising. After the final denoising stage, the output is provided to decoder 612, which transforms that output to pixel space. In one embodiment, the output is also scaled to improve resolution. In one embodiment, the output of the decoder can optionally be run by context adjuster 636. Context adjuster is a process that uses machine learning to check the resulting output to adjust so that the output is more realistic or eliminates unreal or unnatural output. For example, if the input requires "a boy pushing a lawnmower" and the output shows a boy with three legs, the context adjuster can be adjusted with a process or overlay in the drawing to correct or prevent inconsistent or undesirable output. However, as machine learning model 644 utilizes more training to become more intelligent over time, the demand for context adjuster 636 will be less and less before rendering output in user interface tool 620.

[0078] Figure 7Shown are components of an exemplary device 700 that can be used to perform various aspects of the various embodiments of the present disclosure. This block diagram shows a device 700 that can be combined with or can be a personal computer, video game console, personal digital assistant, server, or other digital device suitable for practicing the embodiments of the present disclosure. The device 700 includes a central processing unit (CPU) 702 for running software applications and optionally an operating system. The CPU 702 can be composed of one or more homogeneous or heterogeneous processing cores. For example, the CPU 702 is one or more general-purpose microprocessors with one or more processing cores. Further embodiments can be implemented using one or more CPUs with a microprocessor architecture that is particularly suitable for highly parallel and computationally intensive applications, such as processing interpretation queries, identifying context-related resources, and immediately implementing and rendering context-related resources in a video game. The device 700 can be local to the player playing the game segment (e.g., a game console), remote to the player (e.g., a back-end server processor), or one of many servers that use virtualization to remotely stream gameplay to clients in a game cloud system.

[0079] Memory 704 stores applications and data for use by CPU 702. Storage 706 provides non-volatile storage and other computer-readable media for applications and data, and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other optical storage devices, as well as signal transmission and storage media. User input device 708 transmits user input from one or more users to device 700. Examples of such devices may include a keyboard, mouse, joystick, touchpad, touch screen, still or video recorder / camera, tracking device for gesture recognition, and / or microphone. Network interface 714 allows device 700 to communicate with other computer systems via an electronic communication network, and may include wired or wireless communication over local area networks and wide area networks (such as the Internet). Audio processor 712 is adapted to generate analog or digital audio output from instructions and / or data provided by CPU 702, memory 704, and / or storage 706. The components of device 700 , including CPU 702 , memory 704 , data storage 706 , user input device 708 , network interface 710 , and audio processor 712 , are connected via one or more data buses 722 .

[0080] Graphics subsystem 720 is further connected to data bus 722 and components of device 700. Graphics subsystem 720 includes a graphics processing unit (GPU) 716 and graphics memory 718. Graphics memory 718 includes display memory (e.g., a frame buffer) for storing pixel data for each pixel of an output image. Graphics memory 718 may be integrated into the same device as GPU 708, connected to GPU 716 as a separate device, and / or implemented within memory 704. Pixel data may be provided directly from CPU 702 to graphics memory 718. Alternatively, CPU 702 provides data and / or instructions defining desired output images to GPU 716, which generates pixel data for one or more output images based on the data and / or instructions. The data and / or instructions defining the desired output images may be stored in memory 704 and / or graphics memory 718. In one embodiment, the GPU 716 includes 3D rendering capabilities that generate pixel data for an output image from instructions and data defining the scene's geometry, lighting, shading, textures, motion, and / or camera parameters. The GPU 716 may further include one or more programmable execution units capable of executing shader programs.

[0081] Graphics subsystem 714 periodically outputs pixel data for an image to be displayed on display device 710 from graphics memory 718. Display device 710 may be any device capable of displaying visual information in response to a signal from device 700, including CRT, LCD, plasma, and OLED displays. For example, device 700 may provide analog or digital signals to display device 710.

[0082] It should be noted that access services delivered over a wide geographic area (such as providing access to games in the current implementation scheme) typically use cloud computing. Cloud computing is a computing method in which dynamically scalable and usually virtualized resources are provided as a service over the Internet. Users do not need to be experts in the technical infrastructure in the "cloud" that supports them. Cloud computing can be divided into different services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS) and Software as a Service (SaaS). Cloud computing services typically provide common applications (such as video games) that are accessed online from a web browser, while the software and data are stored on servers in the cloud. Based on the way the Internet is depicted in computer network diagrams, the term cloud is used as a metaphor for the Internet and is an abstraction of the complex infrastructure that it hides.

[0083] In some embodiments, a game server can be used to operate a platform for video game player duration information. Most video games played over the internet operate through a connection to a game server. Typically, games utilize dedicated server applications that collect data from players and distribute it to other players. In other embodiments, video games may be executed by a distributed game engine. In these embodiments, a distributed game engine may be executed on multiple processing entities (PEs), with each PE performing a functional fragment of the given game engine on which the video game is running. The game engine simply views each processing entity as a compute node. Game engines typically perform a variety of operations to execute the video game application and provide additional services for the user experience. For example, a game engine implements game logic, performs game calculations, physics effects, geometric transformations, rendering, lighting, shading, audio, and additional in-game or game-related services. Additional services may include, for example, messaging, social utilities, audio communication, game replay functionality, help functions, and the like. While game engines may sometimes execute on an operating system virtualized by a specific server's hypervisor, in other embodiments, the game engine itself is distributed across multiple processing entities, each of which may reside on a different server unit in a data center.

[0084] According to this embodiment, the corresponding processing entity used to perform operations can be a server unit, a virtual machine, or a container, depending on the needs of each game engine fragment. For example, if a game engine fragment is responsible for camera transformations, this particular game engine fragment can be provided with a virtual machine associated with a graphics processing unit (GPU) because it will be performing a large number of relatively simple mathematical operations (e.g., matrix transformations). Other game engine fragments that require fewer but more complex operations can be provided with processing entities associated with one or more higher-powered central processing units (CPUs).

[0085] By distributing the game engine, it is provided with elastic computing properties that are not constrained by the capabilities of physical server units. Instead, the game engine is provisioned with more or fewer computing nodes as needed to meet the demands of the video game. From the perspective of the video game and the video game player, a game engine distributed across multiple computing nodes is no different from a non-distributed game engine executed on a single processing entity, as the game engine manager or supervisor distributes the workload and seamlessly integrates the results to provide the video game output components to the end user.

[0086] The user uses a client device to access the remote service, and the client device includes at least a CPU, a display and an I / O. The client device can be a PC, a mobile phone, a notebook computer, a PDA, etc. In one embodiment, the network identification client used on the game server adjusts the communication method adopted. In other cases, the client device uses a standard communication method such as html to access the application on the game server through the Internet. It should be understood that a given video game or game application can be developed for a specific platform and a specifically associated controller device. However, when making such a game available via the game cloud system presented herein, the user may access the video game with different controller devices. For example, a game may have been developed for a game console and its associated controller, and the user may use a keyboard and mouse to access a cloud-based version of the game from a personal computer. In such scenarios, the input parameter configuration can define the mapping of the input generated from the user's available controller device (in this case, i.e., keyboard and mouse) to the acceptable input for the execution of the video game.

[0087] In another example, the user can access the cloud gaming system via a tablet computing device, a touch screen smart phone or other touch screen driven device. In this case, the client device and the controller device are integrally formed in the same device, wherein input is provided by means of detected touch screen input / gesture. For such devices, the input parameter configuration can define a specific touch screen input corresponding to the game input of the video game. For example, during the running of the video game, buttons, direction pads or other types of input elements may be displayed or overlaid to indicate the position on the touch screen that the user can touch to generate game input. Gestures (such as swipes or specific touch movements in a specific direction) can also be detected as game input. In one embodiment, guidance can be provided to the user, indicating, for example, how to provide input for game play via the touch screen before starting the game play of the video game, so as to adapt the user to the operation of the controls on the touch screen.

[0088] In some embodiments, the client device serves as a connection point for the controller device. That is, the controller device communicates with the client device via a wireless or wired connection to transmit input from the controller device to the client device. The client device can process these inputs in turn and then transmit the input data to the cloud gaming server via a network (e.g., accessed via a local networking device such as a router). However, in other embodiments, the controller itself can be a networked device with the ability to transmit input directly to the cloud gaming server via a network without first transmitting such input through a client device. For example, the controller can be connected to a local networking device (such as the above-mentioned router) to send data to and receive data from the cloud gaming server. Therefore, although the client device may still be required to receive video output from a cloud-based video game and render it on a local display, input latency can be reduced by allowing the controller to send input directly to the gaming cloud server over the network, thereby bypassing the client device.

[0089] In one embodiment, the networked controller and client device may be configured to send certain types of input directly from the controller to the cloud gaming server and to send other types of input via the client device. For example, inputs whose detection does not rely on any additional hardware or processing other than the controller itself may be sent directly from the controller to the cloud gaming server via the network, thereby bypassing the client device. Such inputs may include button inputs, joystick inputs, embedded motion detection inputs (e.g., accelerometers, magnetometers, gyroscopes), and the like. However, inputs that utilize additional hardware or require processing by the client device may be sent by the client device to the cloud gaming server. These may include video or audio captured from the gaming environment, which may be processed by the client device before being sent to the cloud gaming server. In addition, input from the controller's motion detection hardware may be processed by the client device in conjunction with the captured video to detect the position and motion of the controller, which the client device then transmits to the cloud gaming server. It should be understood that the controller device according to various embodiments may also receive data (e.g., feedback data) from the client device or directly from the cloud gaming server.

[0090] In one embodiment, various technical examples can be implemented using a virtual environment via a head-mounted display (HMD). The HMD may also be referred to as a virtual reality (VR) headset. As used herein, the term "virtual reality" (VR) generally refers to a user's interaction with a virtual space / environment, which involves viewing the virtual space through an HMD (or VR headset) in a manner that responds in real time to the HMD's movements (controlled by the user) to provide the user with a sense of being in the virtual space or metaverse. For example, a user may see a three-dimensional (3D) view of the virtual space when facing a given direction. When the user turns to one side, thereby rotating the HMD, the view of that side in the virtual space is rendered on the HMD. The HMD can be worn similarly to glasses, goggles, or a helmet and is configured to display video games or other metaverse content to the user. The HMD can provide a highly immersive experience by placing its display mechanism close to the user's eyes. Thus, the HMD can provide each eye with a display area that occupies most or even the entire field of view of the user, and can also provide viewing with three-dimensional depth and perspective.

[0091] In one embodiment, the HMD may include a gaze tracking camera configured to capture images of the user's eyes as the user interacts with the VR scene. The gaze information captured by the gaze tracking camera may include information related to the user's gaze direction and specific virtual objects and content items in the VR scene that the user is focused on or interested in interacting with. Thus, based on the user's gaze direction, the system can detect specific virtual objects and content items that may be of potential focus to the user (with which the user is interested in interacting and engaging with them), such as game characters, game objects, game props, etc.

[0092] In some embodiments, the HMD may include an outward-facing camera configured to capture images of the user's real-world space, such as the user's body movements and any real-world objects that may be located in the real-world space. In some embodiments, the images captured by the external camera may be analyzed to determine the position / orientation of real-world objects relative to the HMD. Using the known position / orientation of the HMD, real-world objects, and inertial sensor data from the user's posture and movement, the user's posture and movement during user interaction with the VR scene may be continuously monitored and tracked. For example, when interacting with a scene in a game, the user may make various gestures, such as pointing and walking towards specific content items in the scene. In one embodiment, the system may track and process gestures to generate predictions of interactions with specific content items in the game scene. In some embodiments, machine learning may be used to facilitate or assist in the predictions.

[0093] During use of the HMD, various single-handed and two-handed controllers can be used. In some implementations, the controllers themselves can be tracked by tracking lights included in the controllers or by tracking shapes, sensors, and inertial data associated with the controllers. Using these different types of controllers, or even simply using gestures made and captured by one or more cameras, one can interface with, control, manipulate, interact with, and participate in a virtual reality environment or metaverse rendered on the HMD. In some cases, the HMD can be wirelessly connected to a cloud computing and gaming system via a network. In one embodiment, the cloud computing and gaming system maintains and executes the video game played by the user. In some embodiments, the cloud computing and gaming system is configured to receive input from the HMD and interface objects via the network. The cloud computing and gaming system is configured to process the input to affect the game state of the executing video game. Output from the executing video game (such as video data, audio data, and haptic feedback data) is transmitted to the HMD and interface objects. In other implementations, the HMD can wirelessly communicate with the cloud computing and gaming system via alternative mechanisms or channels, such as a cellular network.

[0094] In addition, although the implementations in the present disclosure may be described with reference to a head-mounted display, it should be understood that in other implementations, a non-head-mounted display may be substituted, including but not limited to a portable device screen (e.g., a tablet, a smartphone, a laptop computer, etc.) or any other type of display that can be configured to render video and / or provide a display of an interactive scene or virtual environment according to the present implementation. It should be understood that the various embodiments defined herein can be combined or assembled into specific implementations using the various features disclosed herein. Therefore, the examples provided are only some possible examples and are not limited to various implementations that can define more implementations by combining various elements. In some examples, some implementations may include fewer elements without departing from the spirit of the disclosed or equivalent implementations.

[0095] The embodiments of the present disclosure may be practiced with various computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. The embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a wired or wireless network.

[0096] While the method operations are described in a particular order, it should be understood that other housekeeping operations may be performed between the operations, or that the operations may be adjusted so that they occur at slightly different times, or that the operations may be distributed throughout a system that allows processing operations to occur at various intervals associated with the processing so long as the processing of telemetry and game state data used to generate the modified game state is performed in the desired manner.

[0097] One or more embodiments may also be manufactured as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device that can store data that can subsequently be read by a computer system. Examples of computer-readable media include hard drives, network attached storage (NAS), read-only memory, random access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. Computer-readable media may include computer-readable tangible media distributed across a network of coupled computer systems so that computer-readable code is stored and executed in a distributed manner.

[0098] In one embodiment, the video game is executed locally on a game console, a personal computer, or on a server. In some cases, the video game is executed by one or more servers in a data center. When executing the video game, some instances of the video game may be simulations of the video game. For example, the video game may be executed by an environment or server that generates a simulation of the video game. In some embodiments, the simulation is an instance of the video game. In other embodiments, the simulation may be generated by an emulator. In either case, if the video game is represented as a simulation, the simulation can be executed to render interactive content that can be interactively streamed, executed, and / or controlled by user input.

[0099] Although the foregoing embodiments have been described in some detail for purposes of clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims. The present embodiments are therefore to be considered as illustrative and not restrictive, and the embodiments are not to be limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.

Claims

1. A method comprising: receiving a game image, the game image being captured from gameplay of a video game, the game image depicting a scene of the video game; receiving game state data describing attributes of the scene of the video game depicted in the game image; receiving modification data describing changes to the game image from a client device over a network, the modification data being defined based on user input received at the client device; generating, by an image-generating artificial intelligence (AI), an AI-generated image by applying the game image, the game state data, and the user input; The AI-generated image is transmitted over the network to the client device for rendering to a display.

2. The method of claim 1, wherein the application of the game state data enables the image generation AI to perform semantic understanding of the scene depicted in the game image.

3. The method of claim 2, wherein the semantic understanding of the scene is applied by the image generation AI to perform the changes described in the modification data.

4. The method of claim 1, wherein the game state data identifies one or more elements of the scene depicted in the game image. The method of claim 1 , wherein the game state data describes a depth of one or more virtual objects in the scene.

6. A method as claimed in claim 5, wherein the modification data describes a change in position of a given virtual object within the scene, and wherein generating the AI-generated image is configured to use the depth of the one or more virtual objects to perform the change in position of the given virtual object described by the modification data.

7. The method of claim 5, wherein the modification data describes placement of a given virtual object within the scene, and wherein generating the AI-generated image is configured to use the depth of the one or more virtual objects to perform the placement of the given virtual object described by the modification data.

8. The method of claim 5, wherein the depth of the one or more virtual objects is configured to enable the one or more virtual objects to properly occlude or be properly occluded when performing the change described by the modification data.

9. The method of claim 1, wherein the game state data describes a three-dimensional structure of one or more virtual objects in the scene.

10. The method of claim 1, wherein the modification data is defined by a word or phrase generated by the user input received at the client device.

11. A non-transitory computer-readable medium having program instructions embodied thereon, the program instructions being configured to, when executed by at least one server computer, cause the at least one server computer to perform a method comprising: receiving a game image, the game image being captured from gameplay of a video game, the game image depicting a scene of the video game; receiving game state data describing attributes of the scene of the video game depicted in the game image; receiving modification data describing changes to the game image from a client device over a network, the modification data being defined based on user input received at the client device; generating, by an image-generating artificial intelligence (AI), an AI-generated image by applying the game image, the game state data, and the user input; The AI-generated image is transmitted over the network to the client device for rendering to a display.

12. The non-transitory computer-readable medium of claim 11, wherein the application of the game state data enables the image generation AI to perform semantic understanding of the scene depicted in the game image.

13. The non-transitory computer-readable medium of claim 12, wherein the changes described in the modification data are performed by the image generation AI applying the semantic understanding of the scene.

14. The non-transitory computer-readable medium of claim 11, wherein the game state data identifies one or more elements of the scene depicted in the game image.

15. The non-transitory computer-readable medium of claim 11, wherein the game state data describes a depth of one or more virtual objects in the scene.

16. A non-transitory computer-readable medium as described in claim 15, wherein the modification data describes a change in position of a given virtual object within the scene, and wherein generating the AI-generated image is configured to use the depth of the one or more virtual objects to perform the change in position of the given virtual object described by the modification data.

17. The non-transitory computer-readable medium of claim 15, wherein the modification data describes placement of a given virtual object within the scene, and wherein generating the AI-generated image is configured to use the depth of the one or more virtual objects to perform the placement of the given virtual object described by the modification data.

18. The non-transitory computer-readable medium of claim 15, wherein the depth of the one or more virtual objects is configured to enable the one or more virtual objects to properly occlude or be properly occluded when performing the change described by the modification data.

19. The non-transitory computer-readable medium of claim 11, wherein the game state data describes a three-dimensional structure of one or more virtual objects in the scene.

20. The non-transitory computer-readable medium of claim 11, wherein the modification data is defined by a word or phrase generated by the user input received at the client device.