Voice driven 3D static asset creation in computer simulation
The use of neural networks and AI engines to convert natural language into 2D and 3D assets addresses limitations in conventional asset creation, enabling efficient and realistic asset generation and modification for computer simulations.
Patent Information
- Application Number
- JP2025071200
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-04
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2042-04-22
AI Technical Summary
Existing computer game asset creation methods are limited by the use of conventional assets, lacking flexibility and efficiency in generating or modifying 2D and 3D assets for computer simulations.
A method and system that utilizes natural language input to generate 2D and 3D assets through neural networks, allowing for the conversion of text or speech into 2D images and further into 3D assets, with artist modifications and integration into computer simulations, using AI engines like GANs and physics engines for refinement.
Enables efficient and customizable creation of 2D and 3D assets for computer simulations, enhancing the creative process by allowing artists to modify and integrate assets seamlessly, while ensuring physical accuracy and realism.
Smart Images

Figure 2025105784000001_ABST
Abstract
Description
Technical Field
[0001] This application relates to a technically inventive non-conventional solution that necessarily results from computer technology and brings about specific technical improvements.
Background Art
[0002] As understood herein, the apparent appeal of computer games is enhanced using commonly used computer game assets such as common background objects.
Summary of the Invention
[0003] This principle enables content creators to describe the assets they desire as natural language input and create 2D or 3D assets from that (voice) input. It also facilitates the creation of initial prototype assets for artists who use them repeatedly.
[0004] Accordingly, the method includes receiving text from speech conversion or the like and processing the text using at least one neural network to render a two-dimensional (2D) image of a computer simulation asset. This method also includes converting the 2D image into a three-dimensional (3D) asset. The method includes presenting the 3D asset in at least one computer simulation.
[0005] The text can be input from a keyboard or speech and can indicate at least one position, and the 3D asset is consistent with this position. The text / speech can indicate at least a plurality of objects, and the 3D asset is consistent with the plurality of objects. This method may include using an artist computer to modify the 3D asset before presenting the 3D asset. A microphone can be used to input the modification of the 3D asset into the artist computer.
[0006] In another aspect, the device includes at least one computer memory that is not a transient signal, the computer memory including instructions executable by at least one processor for receiving a photograph of a two-dimensional (2D) image. The instructions are executable for converting the 2D image into a 3D asset and presenting the 3D asset in at least one computer simulation.
[0007] In another aspect, the apparatus includes at least one processor and at least one computer output device configured to be controlled by the processor. The processor is programmed with instructions to identify a two-dimensional (2D) image, convert the 2D image into a 3D asset, and use the 3D asset as an object in a computer simulation.
[0008] The details of the present application can be best understood with reference to the accompanying drawings, both as to its structure and operation, in which like reference numerals refer to like parts.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
DETAILED DESCRIPTION OF THE INVENTION
[0010] The present disclosure generally relates to a computer ecosystem including, but not limited to, aspects of a home appliance (CE) device network such as a computer game network. The systems herein may include server components and client components that may be connected through a network, whereby data may be exchanged between the client components and the server components. The client components may include one or more computing devices including game consoles such as Sony PlayStation®, or game consoles made by Microsoft®, Nintendo®, or other manufacturers, virtual reality (VR) headsets, augmented reality (AR) headsets, portable televisions (e.g., smart TVs, Internet-enabled TVs), portable computers such as laptop and tablet computers, and other mobile devices including smartphones and additional examples described below. These client devices may operate in various operating environments. For example, some of the client computers may use, by way of example, the Linux® operating system, the Microsoft® operating system, or the Unix® operating system, or an operating system made by Apple® or Google®. Using these operating environments, one or more browsing programs such as browsers made by Microsoft®, Google®, or Mozilla®, or other browser programs that can access websites hosted by the Internet servers described below may be executed. Also, using an operating environment according to the present principles, one or more computer game programs may be executed.
[0011] The server and / or gateway may include one or more processors that execute instructions to configure a server that receives and transmits data through a network such as the Internet. Alternatively, the client and server can be connected through a local intranet or a virtual private network. The server or controller can be instantiated by a gaming machine such as a Sony PlayStation (registered trademark), a personal computer, or the like.
[0012] Information can be exchanged between the client and the server through a network. For this purpose and for security, the server and / or client may include a firewall, a load balancer, a temporary storage, and a proxy, as well as other network infrastructure for reliability and security. One or more servers may form an apparatus that implements a method of providing a secure community such as an online social website to network members.
[0013] The processor can be a single-chip processor or a multi-chip processor that can execute logic by various lines such as address lines, data lines, and control lines, as well as registers and shift registers.
[0014] The components included in one embodiment can be used in any suitable combination in other embodiments. For example, any of the various components described herein and / or shown in the figures can be combined, exchanged, or excluded from other embodiments.
[0015] A "system having at least one of A, B, and C" (similarly, "a system having at least one of A, B, or C" and "a system having at least one of A, B, C") includes a system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together.
[0016] Referring specifically to FIG. 1 here, an exemplary system 10 is shown, which may include one or more of the exemplary devices described above and detailed below in accordance with the present principle. The first exemplary device included in the system 10 is, but not limited to, a consumer electronics (CE) device such as an audio-video device (AVD) 12 like an Internet-enabled television having a TV tuner (similarly, a set-top box for controlling the TV). Alternatively, the AVD 12 can also be a computer-controlled Internet-enabled (“smart”) phone, a tablet computer, a notebook computer, an HMD, a wearable computer-controlled device, a computer-controlled Internet-enabled music player, a computer-controlled Internet-enabled headset, an implantable device for the skin, or other implantable devices that are computer-controlled and Internet-enabled. Anyway, it should be understood that the AVD 12 is configured to implement the present principle (e.g., communicate with other CE devices to implement the present principle, execute the logic described herein, and perform any other functions and / or operations described herein).
[0017] Therefore, in order to implement such a principle, the AVD 12 can be established by some or all of the components shown in FIG. 1. For example, the AVD 12 may include one or more displays 14, and the one or more displays 14 may be implemented by a high-resolution or ultra-high-resolution "4K" or higher-resolution flat screen, and may be touch-responsive to receive user input signals by touching the display. The AVD 12 may include one or more speakers 16 for outputting audio according to this principle, and at least one additional input device 18, such as an audio receiver / microphone, etc., for inputting audible commands to the AVD 12 to control the AVD 12. An exemplary AVD 12 may also include one or more network interfaces 20 for communicating through at least one network 22, such as the Internet, WAN, LAN, etc., under the control of one or more processors 24. It may also include a graphics processor. Therefore, the interface 20 may be, but is not limited to, a Wi-Fi (registered trademark) transceiver, and the Wi-Fi (registered trademark) transceiver is an example of a wireless computer network interface, such as a mesh network transceiver. It should be understood that the processor 24 controls the AVD 12 to implement the principle including other elements of the AVD 12 described herein, such as controlling the display 14 to present an image there and receiving an input therefrom. Further, it should be noted that the network interface 20 may be a wired or wireless modem or router, or other suitable interfaces such as a wireless telephone transceiver or the above-described Wi-Fi (registered trademark) transceiver.
[0018] In addition to the above, the AVD12 may also include one or more input ports 26, such as a high-definition multimedia interface (HDMI (registered trademark)) port or a USB port for physically connecting to another CE device, and / or a headphone port for connecting headphones to the AVD12 to present audio to the user via the headphones. For example, the input port 26 may be connected wired or wirelessly to a cable of audio-video content or a satellite source 26a. Thus, the source 26a may be a separate or integrated set-top box, or a satellite receiver. Or, the source 26a may be a game console or a disc player containing content. When implemented as a game console, the source 26a may include some or all of the components described below in relation to the CE device 44.
[0019] The AVD12 may further include one or more computer memories 28, such as disk-based storage or solid-state storage, which are not temporary signals, and in some cases, these storages are implemented as stand-alone devices, or as a personal video recording device (PVR) or a video disc player either inside or outside the chassis of the AVD for playing AV programs, or as removable memory media, within the chassis of the AVD. Also, in some embodiments, the AVD12 may include a position receiver or location receiver, such as a mobile phone receiver, a GPS receiver, and / or an altimeter 30, which is configured to receive geographical location information from a satellite base station or a mobile phone base station, provide the information to the processor 24, and / or determine the altitude at which the AVD12 is disposed together with the processor 24. The component 30 may also be realized by an inertial measurement unit (IMU) typically including a combination of an accelerometer, a gyroscope, and a magnetometer to determine the position and orientation of the AVD12 in three dimensions.
[0020] Continuing the description of the AVD12, in some embodiments, the AVD12 may include one or more cameras 32, and the one or more cameras 32 may be digital cameras such as thermal imaging cameras, web cameras, etc., and / or cameras integrated into the AVD12 and controllable by the processor 24 to collect photos / images and / or videos according to this principle. Also, included in the AVD12 may be a Bluetooth (registered trademark) transceiver 34 and other NFC elements 36 for communicating with other devices using Bluetooth (registered trademark) and / or near-field communication (NFC) technologies respectively. An exemplary NFC element may be a radio frequency identification (RFID) element.
[0021] Furthermore, the AVD12 may include one or more auxiliary sensors 38 that provide inputs to the processor 24 (for example, motion sensors such as accelerometers, gyroscopes, cyclometers, etc., or magnetic sensors, infrared (IR) sensors, optical sensors, speed sensors and / or cadence sensors, gesture sensors (for example, sensors for detecting gesture commands)). The AVD12 may include a wireless television broadcast port 40 for receiving over-the-air (OTA) TV broadcasts that provide inputs to the processor 24. In addition to the above, it should be noted that the AVD12 may also include an IR transmitter and / or IR receiver and / or IR transceiver 42 such as an infrared (IR) data association (IRDA) device. A battery (not shown) may be provided to power the AVD12, and a motion energy harvester that can convert motion energy into electricity to charge the battery and / or power the AVD12 may be provided. A graphics processing unit (GPU) 44 and a field programmable gate array 46 may also be included.
[0022] Referring further to FIG. 1, in addition to the AVD 12, the system 10 may include one or more other CE device types. In one example, the first CE device 48 can be a computer game console that can be used to send computer game audio and video to the AVD 12 via commands sent directly to the AVD 12 and / or via the server described below, while the second CE device 50 may include components similar to those of the first CE device 48. In the example shown, the second CE device 50 can be configured as a computer game controller operated by a player or a head-mounted display (HMD) worn by the player. It should be understood that in the example shown, only two CE devices are shown and a fewer or greater number of devices may be used. The devices herein may implement some or all of the components shown for the AVD 12. Any of the components shown in the next figure may incorporate some or all of the components shown for the AVD 12 case.
[0023] Referring now to at least one of the servers 52 described above, the server 52 includes at least one server processor 54, at least one tangible computer-readable storage medium 56 such as disk-based storage or solid-state storage, and at least one network interface 58 that enables communication with other devices in FIG. 1 through the network 22 under the control of the server processor 54 and that can actually facilitate communication between the server and client devices in accordance with this principle. Note that the network interface 58 can be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface such as, for example, a wireless telephone transceiver.
[0024] Thus, in some embodiments, server 52 can be an Internet server or an entire server "farm", can include "cloud" functionality, can perform "cloud" functionality, whereby the devices of system 10 can access a "cloud" environment via server 52, for example, in an exemplary embodiment regarding a network gaming application. Or, server 52 can be implemented by one or more gaming machines or other computers in the same room as or near the other devices shown in FIG. 1.
[0025] The components shown in the following figures can include some or all of the components shown in FIG. 1.
[0026] FIGS. 2 and 3 show techniques for enabling a game designer to create and / or modify three-dimensional (3D) assets, typically assets that are not common characters, for computer simulations such as computer games, either from scratch or by adapting assets pre-stored in an asset library.
[0027] As shown in FIG. 2, user interface 200 is presented on a display 202 such as any display described herein, and at 204 can prompt the designer to speak the name of the desired asset, for example, the name of a chair in the example shown.
[0028] FIG. 3 shows that at block 300, the designer's next speech (e.g., "a brown chair with armrests, four legs, a cushioned surface, and a backrest") is received and at block 302 is converted to text. Block 303 shows that a keyword extraction module is used to extract keywords from the text to extract keywords. In that example, the output of the keyword extraction can be as follows. Object: Chair Color: Brown Legs: Four legs Surface: Cushioned Back: With backrest
[0029] The text can be input into an artificial intelligence (AI) engine such as one or more neural networks in block 304 to generate a 2D image of the requested asset. The image can be generated from scratch or selected by accessing a library of assets. The library search can first be performed on images that match the keywords, and only if no match is found, the AI engine can generate an image of the asset using the text for 2D or 3D generation models based on supervised or unsupervised training in human language.
[0030] When proceeding from block 304 to block 306, the 2D image is converted to a 3D asset of the asset using a 2D-to-3D conversion system that uses other techniques such as layer stacking or creating a 3D anaglyph stereogram, pseudo-height resolution, etc. A 2D-to-3D reconstruction model can be used. It can include an encoder-decoder neural architecture, where the encoder takes the 2D image as input, generates an encoding, and the 3D decoder generates a 3D object based on the encoding. Thus, a 3D object or asset can be generated using 2D-to-3D reconstruction, generate a 3D object using a generative neural model, and then convert it to meet the specifications or convert an existing 3D model according to the desired specifications. Further details are described in FIGS. 5 and 12-15.
[0031] The 3D asset can be presented, for example, on the display shown in FIG. 2, and in block 308, artist modifications to the asset can be received using other inputs such as audio or graphic manipulation inputs of a point-and-click device. The modifications can include changes such as the size, shape, color, style (but not all parts of the asset), texture of the surface of the asset, etc. of a specific part of the asset. In block 310, the final 3D asset after modification is generated for use in computer simulations.
[0032] Figure 4 shows a UI 400 that can be presented on a display 402 such as any display disclosed herein to prompt a user to input a photo of a desired asset at 404. The photo is depicted in 2D form at 406 and can be uploaded for the process of FIG. 5 by selecting an upload selector 408.
[0033] Figure 5 shows that at block 500, a 2D image of an asset within a photo is received. When moving to block 502, the 2D image is converted into a 3D asset. When proceeding to block 504, the 3D asset can be modified as described herein by an artist or other user for use in a computer simulation. Additional details of 3D asset generation are shown in FIGS. 12 - 15 described below.
[0034] Figure 6 shows exemplary logic for specifying a plurality of assets and their desired relative positions with respect to each other in a computer simulation. Starting from block 600, text is received from direct text input or speech - to - text conversion, and the text describes the assets by name and their desired relative positions with respect to each other.
[0035] Proceeding to block 602, optionally, a description of only a part of an asset that does not apply to the entire asset can also be received. If the description is received as voice input, at block 604, it is converted into text. At block 606, an AI engine such as a generative adversarial network (GAN) can be used to generate a 2D image based on the previously received asset description and position, and the image is converted into a 3D scene at block 608 according to the principles described herein. 3D assets can be generated directly without going through the 2D phase.
[0036] Shown in Figure 7 is the following. UI700 can be presented on a display 702 such as any display described in this specification. UI700 can include a prompt 704 for a person to speak a description of a desired asset scene that can be presented in text form after speech-to-text conversion at 706. In the example shown, the person is specifying a scene where there is a couch diagonally in front of the left of a chair formed as a Gaudi style chair.
[0037] Figure 8 shows an exemplary result of the process of Figure 7. Continuing with the example described in Figure 7, in front of the left of the 3D asset 802 of the chair, a 3D model 800 of a couch is shown, and the back 804 of the chair is in the Gaudi style drawn by frills 806. A label 808 can also be presented by each image indicating what the image is depicting so that an artist can confirm whether the GAN has correctly executed the desired task.
[0038] One approach for verifying the label is to render the 3D model into a 2D image and use a similarity metric to compare the similarity between the 2D image generated from the text and the 2D image rendered from the 3D model.
[0039] Figure 9 shows a UI900 that can be presented on a display 902 such as any display described in this specification. UI900 can include text 904, which shows, for example, a speech-to-text conversion from an artist's voice input for modifying the chair shown in Figure 8, in the example shown, from the Gaudi style to the Louis XIV style. As a result, the frills on the back of the chair shown in Figure 8 change to a more decorative and elegant style, giving rise to the example given.
[0040] Figure 10 shows another principle related to the above disclosure. At block 1000, text indicating a desired modification to an asset, for example, text that can be converted from speech, is received. Based on the desired modification, at block 1002, portions of the related asset are appropriately synthesized together to satisfy the requested modification. This can be done by varying the weights of pixels interpolated along the boundary regions in the asset identified as related to the desired modification.
[0041] In addition to the asset, the artist can also articulate a desired background terrain, such as "mud" or "marble chamber", or other terrains. Also, as described above, the size of the asset can be specified by the artist. For example, the artist can specify a chair that is 20 feet tall. Thereby, when the asset incorporated into the simulation game space interferes with other assets such as the top part of an object, the top part can be made to automatically appear as if it is deformed to accommodate the chair. Thereby, a collaborative method between humans and AI may be required. Using an AI-only approach, more qualitative requirements such as a wide seat or a tall chair can be met.
[0042] Figure 11 shows additional aspects. When a 3D asset is created as described herein, at block 1100, it can be input into a physics engine. Proceeding to block 1102, the geometry of the asset can be modified, for example, by a GAN, to maintain a certain inertia tensor calculated by the physics engine such that the asset tends to move or deform. Thus, the inertia tensor can be solved by the physics engine to describe the behavior of the asset in response to forces. For example, the physics engine can determine whether the generated 3D asset will fall over when pushed with a specific force based on its current structural characteristics.
[0043] In other words, the AI engine can examine the physical characteristics of the structure of the asset, predict how the structure will react physically, and determine how to maintain the physical ratios of previous objects. Constraints can be imposed. For example, if the asset is furniture, no matter what weight value the 3D asset is emulated at, it is necessary to generate it using attributes that prevent the furniture from tipping over. This can be achieved, for example, by appropriately changing the dimensions and weights of the parts of the asset, such as by maintaining the total torque of the various parts of the asset at zero. In other words, the rule-based approach can be combined with AI to generate the object itself. In block 1104, the updated asset (or its physical determination) is fed back to the AI engine.
[0044] In addition to visual characteristics, the techniques described herein can be used to modify the acoustic and material characteristics of an asset using each separate AI engine such as a GAN. For example, a GAN can be used to define the properties of an asset regarding how the asset absorbs force. For example, it is determined whether the asset shatters or cracks, or absorbs a bullet when hit by a bullet. An asset representing a shuriken can be designed to produce different types of explosions in the presence of different assets.
[0045] Referring now to FIG. 12, an overview of techniques for 2D to 3D graphic asset generation is shown. The technique of FIG. 12 is useful for new assets or when it is not possible to convert an existing 3D model. This technique supports generation and conversion.
[0046] Starting from block 1200, to implement the above example, a representation 1202 such as a photograph of a real 2D object like a chair is input into a conditional generation neural model for 2D synthesis. The resulting output 1204 is a 2D representation of the synthesized chair. The output 1204 is sent to an optional 2D transformation model 1206 for interpolation and feature editing. The model 1206 can be fully AI-based, or the model 1206 can be interactive between an AI model and a human operator.
[0047] In the example shown, the 2D transformation model 1206 outputs a transformed synthetic representation 1208 of the chair in 2D. The representation 1208 is included in an asset library, used for artist input, and can be used for 3D reconstruction.
[0048] In practice, the transformed synthetic representation 1208 of a 2D object like a chair and / or the representation 1202 of a 2D real asset can be input into a neural model 1210. The neural model 1210 transforms the 2D representation into a 3D shape and outputs a reconstruction mesh 1212 of the asset. The neural model 1210 appropriately includes implicit functions and mesh deformations. Optionally, the reconstruction mesh 1212 can be input into a texture transformation model 1214 for neural rendering of the texture of the 3D asset.
[0049] Figure 13 shows the control of feature transformation. Starting from block 1300, a 2D generation model (such as a generative adversarial network (GAN), etc.) is trained in each asset class such as tables and chairs to generate assets. The training can be supervised, semi-supervised, or unsupervised.
[0050] When an asset is requested, a properly trained model for the specified asset in this specification is selected. For example, if there are separate models for generating chairs, tables, etc., the model is selected based on the specified asset.
[0051] Artists typically specify the characteristics of the assets to be transformed, such as texture, color, and shape (geometry). In block 1302, the generation is adjusted based on keywords (e.g., attributes) extracted from the description, which can be regarded as annotated features (y - labels), in order to transform the generated assets to meet the specifications within the input description. In one example, five characteristics of a chair, namely, armrests, legs, back, surface, and view (e.g., front or rear) can be used.
[0052] Moving on to block 1304, the encoding can be generated for the annotated chair using different weights, and the weights can be interpolated to best fit the artist's specifications. The encoding is sent to train the supervised classifier 1306 to discover the feature axes F(i). In block 1308, since the features can be edited with the feature axes for a new chair, the unique features are interactively controlled, the attributes are transformed (collaboration between human and AI), for example, an existing chair asset is changed to a chair with a backrest. Thus, the encoding W’ for the new chair is the sum of the encoding W of the existing chair and the product of alpha and the feature axis F(i), where alpha can be determined or discovered empirically.
[0053] Figure 14 shows a further approach. The real or synthetic chair representation 1400 in 2D is sent to a 2D encoder - decoder neural model 1402 for shape encoding. The 2D encoder model 1402 can be a convolutional network or a similar deep neural network. The input 1400 to the encoder model 1402 can be an image generated (and optionally transformed) in Figure 13 that meets the description of the desired asset. Optionally, a texture encoder 1404 can also be provided to encode the texture of the object.
[0054] The 3D decoder 1406 obtains the input encoding and generates a 3D object. The 3D decoder 1406 can also be a convolutional network or a similar DNN. The output of the 3D decoder is a reconstructed mesh 1408 representing the 3D asset.
[0055] To train the network, the 3D output can be rendered into a 2D image and compared with the input image. Training can be repeated continuously until the input and output closely match. Alternatively, mesh deformation can be used.
[0056] The encoder-decoder model can be adapted to incorporate additional encodings (e.g., texture encoding) to transform the 3D object to meet the specifications in the description.
[0057] Referring to FIG. 15 for an alternative approach to generating a 3D asset, at block 1500, the 3D GAN model is trained to generate a 3D object. At block 1502, partial encodings for each part of the asset, e.g., encodings for the armrest, legs, back, etc. for a chair, are extracted. Proceeding to block 1504, the partial encodings are transformed based on the shape description 1506 of the desired asset. Proceeding to block 1508, the generation of the 3D asset is adjusted based on appearance descriptions 1510 such as non-shape descriptions like style or size or color. The reconstructed mesh 1512 of the 3D asset is output with or without texturing as required. That is, the 3D asset model can be rendered based on the specified texture. 3D variations can be generated based on the specified attributes.
[0058] Although the principles have been described with reference to some exemplary embodiments, it is recognized that these are not intended to be limiting and that various alternative arrangements can be used to implement the subject matter claimed herein.
Claims
Claim 1 Receiving a photograph of a two-dimensional (2D) image, rather than a transient signal, converting the 2D image into a 3D asset, presenting the 3D asset in at least one computer simulation, and at least one computer memory comprising instructions executable by at least one processor for the foregoing. A device comprising the same. Claim 2 The device of claim 1, wherein the instructions are executable to associate audio with the 3D asset, at least in part based on text. Claim 3 The device of claim 1, wherein the instructions are executable to receive speech indicating at least one position, and the 3D asset is consistent with the position. Claim 4 The device of claim 1, wherein the instructions are executable to receive speech indicating at least a plurality of objects, and the 3D asset is consistent with the plurality of objects. Claim 5 The device of claim 1, wherein the instructions are executable to modify the 3D asset using an artist computer prior to presenting the 3D asset. Claim 6 The device of claim 1, wherein the instructions are executable to present on a display a user interface (UI) having a selector for uploading the photograph. Claim 7 The device of claim 1, wherein the instructions are executable to present on a display a user interface (UI) having a prompt for uttering a desired asset scene. Claim 8 The device of claim 1, wherein the instructions are executable to identify the 2D image, at least in part based on an input of the photograph of the 2D image. Claim 9 The device of claim 1, wherein the instructions are executable to modify the 3D asset, at least in part based on a physical modeling of environmental effects on the 3D asset.
Citation Information
Patent Citations
Image processor and processing method, and program
JP2006163871A
Pixel-region-to-be-changed extraction device, image processing system, pixel-region-to-be-changed extraction method, image processing method, and program
JP2020057037A
System And Method For Generating 3D Scenes
US20070146360A1
Natural Language Based Computer Animation
US20180293050A1