Voice-driven creation of 3D static assets in computer simulations
By receiving natural language input and using neural network to generate 2D images, combined with a 2D to 3D conversion system, the problem of generating and modifying computer simulation assets in the prior art is solved, and the rapid and efficient asset generation and flexible modification capabilities are achieved.
Patent Information
- Application Number
- JP2023564623
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-04
- Filing Date
- 2022-04-22
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2042-04-22
AI Technical Summary
The prior art is difficult to quickly and efficiently generate 2D or 3D computer simulation assets that meet a particular description from natural language input, and existing methods lack flexibility and automation in generating and modifying assets.
By receiving natural language input, text is processed using a neural network to generate 2D images and converted to 3D assets through a 2D to 3D conversion system. The system also supports the generation of assets through voice input and image processing, and allows real-time modification and optimization of assets to suit different computer simulation environments.
It realizes a fast and efficient process from natural language input to generating 3D computer simulation assets, improves the flexibility and automation of asset generation and modification, and is suitable for a variety of computer simulation applications.
Smart Images

Figure 0007673241000001 
Figure 0007673241000002 
Figure 0007673241000003
Abstract
Description
[Technical field]
[0001] The present application relates to technically inventive and unconventional solutions which are necessarily due to computer technology and give rise to concrete technical improvements. [Background technology]
[0002] As understood herein, commonly used computer game assets, such as common background objects, are used to enhance the visual appeal of the computer game. Summary of the Invention
[0003] The principles allow content creators to describe their desired assets as natural language input and create 2D or 3D assets from that (voice) input, and also facilitate the creation of initial prototype assets for artists to iterate on.
[0004] Thus, the method includes receiving text, such as from a speech converter, and processing the text using at least one neural network to render a two-dimensional (2D) image of a computer simulation asset. The method also includes converting the 2D image into a three-dimensional (3D) asset. The method includes presenting the 3D asset in the at least one computer simulation.
[0005] The text can be input from a keyboard or speech and can indicate at least one location, and the 3D asset is consistent with the location. The text / speech can indicate at least a number of objects, and the 3D asset is consistent with the number of objects. The method may include using an artist computer to modify the 3D asset prior to presenting the 3D asset. A microphone can be used to input the modifications to the 3D asset into the artist computer.
[0006] In another aspect, the device includes at least one computer memory that is not a transitory signal, the computer memory including instructions executable by at least one processor for receiving a photograph of a two-dimensional (2D) image, the instructions being executable for converting the 2D image into a 3D asset, and presenting the 3D asset in at least one computer simulation.
[0007] In another aspect, an apparatus includes at least one processor and at least one computer output device configured to be controlled by the processor, the processor being programmed with instructions to identify a two-dimensional (2D) image, convert the 2D image into a 3D asset, and use the 3D asset as an object in a computer simulation.
[0008] The details of the present application, both as to its structure and operation, can best be understood in reference to the accompanying drawings, in which like reference numerals refer to like parts, and in which: [Brief description of the drawings]
[0009] [Figure 1] 1 is a block diagram of an exemplary system including an embodiment in accordance with the present principles; [Diagram 2] 1 shows an example screenshot prompting a person to input speech for textual identification of a computer simulation asset. [Diagram 3] 1 illustrates example logic in an example flowchart format for converting speech to text for a 3D asset. [Figure 4] 1 shows an exemplary screenshot prompting a person to input an image to generate a computer simulation asset. [Diagram 5] 1 illustrates exemplary logic in an exemplary flowchart format for converting an image into a 3D asset. [Figure 6]1 illustrates exemplary logic in an exemplary flowchart format for converting text to speech relative to locations and portions of 3D assets. [Figure 7] Related example screenshots are shown in FIG. [Figure 8] Related example screenshots are shown in FIG. [Figure 9] For modifying a portion of an asset, an example screenshot is shown in relation to FIG. [Figure 10] An exemplary logic is presented in an exemplary flow chart format for modifying a portion of an asset. [Figure 11] 1 illustrates exemplary logic in an exemplary flowchart format for closed loop processing between 3D assets and a physics engine. [Figure 12] Provides an overview of technologies for 2D to 3D asset generation. [Figure 13] We present techniques for controlled feature transformation. [Figure 14] A 2D to 3D reconstruction approach is presented. [Figure 15] We present a technique for generating 3D assets without using 2D input. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] The present disclosure relates generally to computer ecosystems, including, but not limited to, aspects of consumer electronics (CE) device networks, such as computer gaming networks. The systems herein may include server and client components that may be connected through a network, whereby data may be exchanged between the client and server components. The client components may include one or more computing devices, including gaming consoles such as Sony PlayStation® or gaming consoles made by Microsoft® or Nintendo® or other manufacturers, virtual reality (VR) headsets, augmented reality (AR) headsets, portable televisions (e.g., smart televisions, Internet-enabled televisions), portable computers such as laptops and tablet computers, and smartphones and other mobile devices, including additional examples described below. These client devices may operate in a variety of operating environments. For example, some of the client computers may use, by way of example, a Linux® operating system, a Microsoft® operating system, or a Unix® operating system, or an operating system manufactured by Apple® or Google®. These operating environments may be used to run one or more browsing programs, such as browsers made by Microsoft® or Google® or Mozilla®, or other browser programs that can access web sites hosted by the Internet servers described below. Operating environments consistent with the present principles may also be used to run one or more computer game programs.
[0011] The server and / or gateway may include one or more processors that execute instructions to configure the server to receive and transmit data over a network such as the Internet. Alternatively, the clients and servers may be connected through a local intranet or a virtual private network. The server or controller may be instantiated by a gaming console such as a Sony PlayStation®, a personal computer, or the like.
[0012] Information may be exchanged between the clients and the servers over a network. For this purpose and for security, the servers and / or clients may include firewalls, load balancers, temporary storage, and proxies, as well as other network infrastructure for reliability and security. One or more servers may form an apparatus for implementing a method for providing a secure community, such as an online social website, for network members.
[0013] The processor may be a single-chip processor or a multi-chip processor capable of implementing logic through various lines, such as address lines, data lines, and control lines, as well as registers and shift registers.
[0014] Components included in one embodiment may be used in other embodiments in any suitable combination, for example, any of the various components described herein and / or illustrated in the figures may be combined, interchanged, or eliminated from other embodiments.
[0015] A "system having at least one of A, B, and C" (similarly, "a system having at least one of A, B, or C" and "a system having at least one of A, B, C") includes systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.
[0016] Referring now specifically to FIG. 1, an exemplary system 10 is shown, which may include one or more of the exemplary devices described above and in detail below in accordance with the present principles. A first exemplary device included in the system 10 is a consumer electronics (CE) device, such as an audio-video device (AVD) 12, such as, but not limited to, an Internet-enabled television having a television tuner (as well as a set-top box that controls the television). Alternatively, the AVD 12 may also be a computer-controlled Internet-enabled ("smart") phone, a tablet computer, a notebook computer, an HMD, a wearable computer-controlled device, a computer-controlled Internet-enabled music player, a computer-controlled Internet-enabled headphones, a computer-controlled Internet-enabled implantable device such as an implantable skin device, and the like. Regardless, it should be understood that the AVD 12 is configured to implement the present principles (e.g., to communicate with other CE devices to implement the present principles, to execute the logic described herein, and to perform any other functions and / or operations described herein).
[0017] Accordingly, to implement such principles, an AVD 12 may be established by some or all of the components shown in FIG. 1. For example, the AVD 12 may include one or more displays 14, which may be implemented by high-resolution or super-resolution "4K" or higher resolution flat screens and may be touch-enabled to receive user input signals by touching the display. The AVD 12 may include one or more speakers 16 for outputting audio in accordance with the present principles, and at least one additional input device 18, such as, for example, an audio receiver / microphone, for inputting audible commands to the AVD 12 and controlling the AVD 12. The exemplary AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22, such as the Internet, a WAN, a LAN, etc., under the control of one or more processors 24. It may also include a graphics processor. Thus, the interface 20 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It should be understood that processor 24 controls AVD 12 to implement the present principles, including other elements of AVD 12 described herein, such as controlling display 14 to present images thereon and receiving input therefrom. It should further be noted that network interface 20 may be a wired or wireless modem or router, or other suitable interface, such as a wireless telephony transceiver or the Wi-Fi transceiver discussed above.
[0018] In addition to the above, AVD 12 may also include one or more input ports 26, such as a High-Definition Multimedia Interface (HDMI) port or a USB port for physically connecting to another CE device, and / or a headphone port for connecting headphones to AVD 12 to present audio to a user from AVD 12 via headphones. For example, input port 26 may be wired or wirelessly connected to a cable or satellite source 26a of audio-video content. Thus, source 26a may be a separate or integrated set-top box, or a satellite receiver. Alternatively, source 26a may be a game console or disc player containing content. When implemented as a game console, source 26a may include some or all of the components described below in connection with CE device 44.
[0019] AVD 12 may further include one or more computer memories 28, such as non-transitory, disk-based or solid-state storage, which in some cases are embodied within the AVD's chassis, either as a standalone device or as a personal video recording device (PVR) or video disk player, or as a removable memory medium, either internal or external to the AVD's chassis for playing AV programs. In some embodiments, AVD 12 may also include a position or location receiver, such as, but not limited to, a cellular receiver, a GPS receiver, and / or an altimeter 30, configured to receive geographic location information from a satellite or cellular base station, provide the information to processor 24, and / or determine the altitude at which AVD 12 is located in conjunction with processor 24. Component 30 may also be implemented with an inertial measurement unit (IMU), typically including a combination of accelerometers, gyroscopes, and magnetometers, to determine the position and orientation of AVD 12 in three dimensions.
[0020] Continuing with the description of the AVD 12, in some embodiments, the AVD 12 may include one or more cameras 32, which may be digital cameras such as thermal imaging cameras, webcams, and / or cameras integrated into the AVD 12 and controllable by the processor 24 to collect pictures / images and / or videos in accordance with the present principles. Also included in the AVD 12 may be a Bluetooth transceiver 34 and other NFC elements 36 for communicating with other devices using Bluetooth and / or Near Field Communication (NFC) technologies, respectively. An exemplary NFC element may be a Radio Frequency Identification (RFID) element.
[0021] Furthermore, the AVD 12 may include one or more auxiliary sensors 38 (e.g., motion sensors such as accelerometers, gyroscopes, cyclometers, or magnetic sensors, infrared (IR) sensors, optical sensors, speed and / or cadence sensors, gesture sensors (e.g., sensors for detecting gesture commands)) that provide input to the processor 24. The AVD 12 may include a wireless television broadcast port 40 for receiving over-the-air (OTA) TV broadcasts that provide input to the processor 24. In addition to the above, it is noted that the AVD 12 may also include an IR transmitter and / or an IR receiver and / or an IR transceiver 42, such as an infrared (IR) data association (IRDA) device. A battery (not shown) may be provided to power the AVD 12, and may be a kinetic energy harvester that may convert kinetic energy into electrical power to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field programmable gate array 46 may also be included.
[0022] With further reference to FIG. 1, in addition to the AVD 12, the system 10 may include one or more other CE device types. In one example, the first CE device 48 may be a computer game console that can be used to send computer game audio and video to the AVD 12 via commands sent directly to the AVD 12 and / or via a server as described below, while the second CE device 50 may include similar components as the first CE device 48. In the example shown, the second CE device 50 may be configured as a computer game controller operated by a player or a head mounted display (HMD) worn by a player. It should be understood that in the example shown, only two CE devices are shown and a fewer or greater number of devices may be used. The devices herein may implement some or all of the components shown for the AVD 12. Any of the components shown in the following figures may incorporate some or all of the components shown in the case of the AVD 12.
[0023] Referring now to the at least one server 52 mentioned above, the server 52 includes at least one server processor 54, at least one tangible computer-readable storage medium 56, such as disk-based or solid-state storage, and at least one network interface 58 that, under the control of the server processor 54, enables communication with other devices of FIG. 1 over the network 22, and may, in effect, facilitate communications between the server and client devices in accordance with the present principles. It should be noted that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface, such as, for example, a wireless telephony transceiver.
[0024] Thus, in some embodiments, server 52 may be an entire Internet server or server "farm" and may include or perform "cloud" functions such that devices of system 10 may access a "cloud" environment via server 52, for example in an exemplary embodiment directed to a network gaming application, or server 52 may be implemented by one or more gaming consoles or other computers in the same room or nearby as the other devices shown in FIG.
[0025] The components illustrated in the following figures may include some or all of the components illustrated in FIG.
[0026] 2 and 3 show techniques for enabling game designers to create and / or modify three-dimensional (3D) assets for computer simulations such as computer games, typically assets that are not common characters, either from scratch or by adapting assets pre-stored in an asset library.
[0027] As shown in FIG. 2, a user interface 200 may be presented on a display 202, such as any display described herein, and may prompt the designer at 204 to speak the name of a desired asset, such as the name of a chair in the example shown.
[0028] 3 shows that in block 300, the designer's next speech (e.g., "A brown chair with arms, four legs, a cushioned surface, and a backrest") is received and converted to text in block 302. Block 303 shows that keywords are extracted from the text using a text processing module to extract keywords. In the example, the output of keyword extraction may be: Object: Chair Color:Brown Legs: 4 legs Surface: Cushion Back: With backrest
[0029] The text may be input to an artificial intelligence (AI) engine, such as one or more neural networks, to generate a 2D image of the requested asset in block 304. The image may be generateable ab initio or may be selected by accessing a library of assets. A search of the library may first be done for images matching the keywords, and only if no match is found may the AI engine generate an image of the asset using the text in a 2D or 3D generative model based on supervised or unsupervised training in human language.
[0030] Proceeding from block 304 to block 306, the 2D image is converted to a 3D asset using a 2D to 3D conversion system using, for example, layer stacking or other techniques such as creating a 3D anaglyph stereogram, false height resolution, etc. The 2D to 3D reconstruction model may be used. It may include an encoder-decoder neural architecture, where the encoder takes the 2D image as input and generates an encoding, and the 3D decoder generates a 3D object based on the encoding. Thus, a 3D object or asset may be generated using 2D to 3D reconstruction, using a generative neural model to generate a 3D object and then convert it to specs or convert an existing 3D model according to desired specs. Further details are described in Figures 5 and 12-15.
[0031] The 3D asset may be presented, for example, on a display as shown in Figure 2 and may receive artist modifications to the asset at block 308 using voice or other input, such as point-and-click device graphic manipulation input, that may include changing the size, shape, color, style of certain parts of the asset (but not all parts of the asset), the texture of the asset's surface, etc. At block 310, the final 3D asset after modifications is generated for use in the computer simulation.
[0032] 4 illustrates a UI 400 that may be presented on a display 402, such as any display disclosed herein, to prompt a user to input a photograph of a desired asset at 404. The photograph is depicted in 2D format at 406 and can be uploaded for processing in FIG.
[0033] 5 shows that at block 500, a 2D image of an asset in a photograph is received. Moving to block 502, the 2D image is converted into a 3D asset. Proceeding to block 504, the 3D asset may be modified as described herein by an artist or other user for use in a computer simulation. Additional details of 3D asset generation are provided in FIGS. 12-15, described below.
[0034] 6 illustrates exemplary logic for specifying multiple assets and their desired relative positions to one another in a computer simulation. Beginning at block 600, text from direct text input or speech-to-text conversion is received that describes the assets by name and their desired relative positions to one another.
[0035] Proceeding to block 602, if desired, a description of only a portion of the asset may be received that does not apply to the entire asset. If the description is received as a voice input, it is converted to text in block 604. In block 606, an AI engine such as a generative adversarial network (GAN) may be used to generate a 2D image based on the previously received asset description and location, and the image is converted to a 3D scene in block 608 according to the principles described herein. The 3D asset may be generated directly without going through a 2D phase.
[0036] 7 shows: A UI 700 may be presented on a display 702, such as any of the displays described herein. The UI 700 may include a prompt 704 for a person to speak a description of a desired asset scene, which may be presented in text form after speech-to-text conversion at 706. In the example shown, the person is specifying a scene with a couch to the left front of a chair configured as a Gaudi style chair.
[0037] Figure 8 shows an example result of the process of Figure 7. Continuing with the example described in Figure 7, to the left front of the chair 3D asset 802, a 3D model of a couch 800 is shown, with the back of the chair 804 in Gaudi style rendered with ruffles 806. Labels 808 may also be presented with each image indicating what the image is trying to depict, so that the artist can verify if the GAN has performed the desired task correctly.
[0038] One approach to verifying the labels is to render the 3D model into a 2D image and use a similarity metric to compare the similarity between the 2D image generated from the text and the 2D image rendered from the 3D model.
[0039] Figure 9 illustrates a UI 900 that may be presented on a display 902, such as any of the displays described herein. The UI 900 may include text 904, which indicates a speech-to-text conversion from an artist's voice input, for example, to modify the chair shown in Figure 8, in the illustrated example, from a Gaudi style to a Louis XIV style. As a result, the frills on the back of the chair shown in Figure 8 change to a more decorative and elegant style, resulting in the given example.
[0040] 10 illustrates another principle related to the above disclosure. At block 1000, text, e.g., text that may be converted from speech, indicating a desired modification to an asset is received. Based on the desired modification, at block 1002, portions of the associated asset are appropriately composited together to satisfy the requested modification. This may be done by varying weights of interpolated pixels along boundary regions in the asset that are identified as being associated with the desired modification.
[0041] In addition to the assets, the artist may also vocalize the desired background terrain, such as "mud" or "marble palace," or other terrain. Also, as mentioned above, the size of the asset may be specified by the artist. For example, the artist may specify a chair that is 20 feet tall. This may cause the top of an object to automatically deform to accommodate the chair when the asset is embedded in the simulation's game space and interferes with other assets, such as the top of the object. This may require a human-AI collaborative method. More qualitative requirements, such as a wide seat or a tall chair, may be met using an AI-only approach.
[0042] 11 illustrates an additional embodiment. Once a 3D asset is created as described herein, in block 1100, it may be input to a physics engine. Proceeding to block 1102, the geometry of the asset may be modified, for example by a GAN, to maintain a constant inertia tensor calculated by the physics engine that tends to move or deform the asset. Thus, the inertia tensor may be solved by the physics engine to describe the behavior of the asset in response to forces. For example, the physics engine may determine whether the generated 3D asset will tip over when pushed with a certain force based on the current structural features of the asset.
[0043] In other words, the AI engine can look at the physical properties of the asset's structure, predict how the structure will react to physics, and determine how to maintain the physical ratios of the previous object. Constraints can be imposed, for example, if the asset is a piece of furniture, it needs to be generated with attributes that prevent the furniture from tipping over, no matter what weight values the 3D asset can emulate. This can be achieved, for example, by appropriately modifying the dimensions and weights of the parts of the asset, for example, by maintaining the total torque of the various parts of the asset at zero. In other words, the rule-based approach can be combined with AI to generate the object itself. In block 1104, the updated asset (or its physical determination) is fed back to the AI engine.
[0044] In addition to visual properties, the techniques described herein may be used to modify the acoustic and material properties of an asset using separate respective AI engines such as GANs. For example, GANs may be used to define an asset's properties regarding how it absorbs forces. For example, if hit by a bullet, the asset will shatter or break apart, or absorb the bullet. An asset representing a grenade may be designed to produce different types of explosions in the presence of different assets.
[0045] Referring now to Fig. 12, an overview of a technique for 2D to 3D graphic asset generation is shown. The technique of Fig. 12 is useful for new assets or when it is not possible to convert an existing 3D model. This technique supports generation and conversion.
[0046] Beginning at block 1200, to implement the above example, a representation 1202, such as a photo, of a real 2D object, such as a chair, is input to a conditional generative neural model for 2D synthesis. The resulting output 1204 is a 2D representation of the synthesized chair. The output 1204 is sent to an optional 2D transformation model 1206 for interpolation and feature editing. The model 1206 can be fully AI-based, or the model 1206 can be interactive between the AI model and a human operator.
[0047] The 2D transformation model 1206 outputs a transformed composite representation 1208 of the chair in 2D, in the example shown. The representation 1208 can be included in an asset library, used for artist input, and used in 3D reconstruction.
[0048] In practice, the 2D transformed synthetic representation 1208 of a chair, etc. and / or the 2D real asset representation 1202 may be input to a neural model 1210, which converts the 2D representation to a 3D shape and outputs a reconstructed mesh 1212 of the asset. The neural model 1210 includes implicit functions and mesh deformations, as appropriate. Optionally, the reconstructed mesh 1212 may be input to a texture transformation model 1214 for neural rendering of the 3D asset's texture.
[0049] 13 shows the control of feature transformation. Starting at block 1300, a 2D generative model (such as a generative adversarial network (GAN)) is trained with each asset class, such as tables and chairs, to generate assets. Training can be supervised, semi-supervised, or unsupervised.
[0050] When an asset is requested, the appropriate trained model is selected for the asset specified herein. For example, if separate models exist for generating chairs, tables, etc., a model is selected based on the specified asset.
[0051] An artist typically specifies the characteristics of the asset to be transformed, such as texture, color, and shape (geometry). To transform the generated asset to meet the specifications in the input description, in block 1302 the generation is adjusted based on keywords (e.g., attributes) extracted from the description, which may be considered as annotated features (y-labels). In one example, five features of a chair may be used: arms, legs, back, surface, and scenery (e.g., front or back).
[0052] Proceeding to block 1304, an encoding may be generated for the annotated chair using different weights, which may be interpolated to best fit the artist's specifications. The encodings are sent to train a supervised classifier 1306 to find the feature axes F(i). In block 1308, the features may be edited along with the feature axes for the new chair to interactively control the characteristic features and transform attributes (human-AI collaboration), e.g., to change an existing chair asset to a chair with a backrest. Thus, the encoding W' for the new chair is the encoding W of the existing chair plus the product of alpha and the feature axes F(i), where alpha may be empirically determined or discovered.
[0053] Figure 14 shows a further approach. A representation 1400 of a real or synthetic chair in 2D is sent to a 2D encoder-decoder neural model 1402 for shape encoding. The 2D encoder model 1402 may be a convolutional network or similar deep neural network. The input 1400 to the encoder model 1402 may be the image generated in Figure 13 and (optionally) transformed to fulfill the description of the desired asset. Optionally, a texture encoder 1404 may also be provided to encode the texture of the object.
[0054] The 3D decoder 1406 takes the input encoding and generates a 3D object. The 3D decoder 1406 can also be a convolutional network or similar DNN. The output of the 3D decoder is a reconstructed mesh 1408 that represents the 3D asset.
[0055] To train the network, the 3D output can be rendered into a 2D image and compared to the input image. Training can continue iteratively until the input and output closely match. Alternatively, mesh deformation can be used.
[0056] The encoder-decoder model can be adapted to incorporate additional encodings (e.g., texture encodings) that transform the 3D object to meet the specifications in the description.
[0057] Referring to FIG. 15 for an alternative approach to generating a 3D asset, in block 1500 a 3D GAN model is trained to generate a 3D object. In block 1502, a part encoding for each part of the asset is extracted, e.g., encodings for arms, legs, back, etc. for a chair. Proceeding to block 1504, the part encoding is transformed based on a shape description 1506 of the desired asset. Proceeding to block 1508, the generation of the 3D asset is adjusted based on an appearance description 1510, such as style or non-shape descriptions such as size or color. A reconstructed mesh 1512 of the 3D asset is output with or without texturing, as needed. That is, the 3D asset model can be rendered based on a specified texture. 3D variations can be generated based on specified attributes.
[0058] While the present principles have been described with reference to certain illustrative embodiments, it will be recognized that these are not intended to be limiting and that various alternative arrangements may be used to practice the subject matter claimed herein.
Claims
1. Receiving a text; and processing the text using at least one neural network to render a two-dimensional (2D) image of a computer simulation asset; converting the 2D image into a three-dimensional (3D) asset; modifying the 3D asset using an artist computer by at least partially varying weights of interpolated pixels along at least one boundary region in the 3D asset where a modification is identified; presenting the modified 3D asset in at least one computer simulation; and A method comprising:
2. The method of claim 1 , wherein the text is received from a speech conversion.
3. The method of claim 1 , comprising associating audio with the 3D asset based at least in part on the text.
4. The method of claim 2 , wherein the speech conversion indicates at least one location, and the 3D assets are consistent with the location.
5. The method of claim 2 , wherein the speech conversion is indicative of at least a plurality of objects, and the 3D assets are consistent with the plurality of objects.
6. The method of claim 1 , comprising using a microphone to input modifications to the 3D asset into the artist computer.
7. At least one processor; at least one computer output device configured to be controlled by said processor; An apparatus comprising: The processor, Identifying a two-dimensional (2D) image; converting the 2D images into 3D assets; modifying the 3D asset using an artist computer by at least partially varying weights of interpolated pixels along at least one boundary region in the 3D asset where a modification is identified; using the modified 3D asset as an object in a computer simulation; , programmed with instructions to The apparatus.
8. The instruction: The apparatus of claim 7 , wherein the apparatus is operable to identify the 2D image based at least in part on a text input describing the 2D image.
9. The instruction: The apparatus of claim 8 , wherein the apparatus is operable to derive the text input from a speech input.
10. The instruction:
10. The apparatus of claim 8, wherein the apparatus is operable to generate the 2D image based at least in part on a text input describing the 2D image using at least one neural network.
11. The instruction: and operable to associate audio with the 3D asset based at least in part on the text input.
8. The apparatus of claim 7.
12. The instruction: The apparatus according to claim 7 , which is executable to modify the 2D images based on text and / or voice input prior to 3D reconstruction.
Citation Information
Patent Citations
Object generation device, method, and program
JP2015056132A
Data composing apparatus and method
JP2019045984A
System And Method For Generating 3D Scenes
US20070146360A1
Natural Language Based Computer Animation
US20180293050A1
System and method for improved neural network training
US20190130266A1