Voice modification of sub-parts of assets in computer simulation

The method enables content creators to describe assets in natural language and convert them into 2D or 3D assets using neural networks, addressing the inefficiencies in current asset creation processes and enhancing the speed and flexibility of asset development for computer simulations.

JP7676579B2Active Publication Date: 2025-05-14SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023561232
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-04
Filing Date
2022-04-28
Publication Date
2025-05-14
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

Current computer game development processes require significant time and expertise to create 2D or 3D assets, limiting the ability of content creators to efficiently generate and modify assets for computer simulations.

Method used

A method that allows content creators to describe desired assets in natural language, using neural networks to convert these descriptions into 2D or 3D assets, and modify existing assets without altering their entire structure, facilitating the creation of early prototype assets and reusable artist assets.

Benefits of technology

Enables content creators to quickly generate and modify 2D or 3D assets for computer simulations, reducing the time and expertise required, and allowing for more intuitive and efficient asset creation and modification processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007676579000001
    Figure 0007676579000001
  • Figure 0007676579000002
    Figure 0007676579000002
  • Figure 0007676579000003
    Figure 0007676579000003
Patent Text Reader

Abstract

A computer-simulated object, such as a chair, is described with voice (302) or photographic input in order to render a 2D image. Machine learning may be used to convert audio input into a 2D image, which is converted (304) into a 3D asset, and the 3D asset, or a portion thereof, is used as an object, such as a chair, in a computer simulation (310), such as a computer game.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present application relates to technically original and unconventional solutions that are necessarily rooted in computer technology and provide concrete technical improvements. [Background technology]

[0002] As understood herein, commonly used computer game assets, such as common background objects, are used to enhance the visual appeal of the computer game. Summary of the Invention

[0003] The principles of the present invention allow content creators to describe the assets they want as natural language input and create 2D or 3D assets from that (voice) input, and also facilitate the creation of early prototype assets for artists to iterate on.

[0004] Thus, the method includes receiving at least one two-dimensional (2D) image of a computer-simulated object. The method also includes converting the 2D image into a three-dimensional (3D) asset. The method includes using at least one neural network to modify a first portion of the 3D asset and leave a second portion of the 3D asset unchanged to create a modified asset, and presenting the modified asset in the at least one computer simulation.

[0005] In some exemplary embodiments, the 2D image is generated from text, which may be created from voice recognition or from a camera image of the object.

[0006] In an exemplary implementation, the method can include blending the first and second portions of the object at least partially into one another by modifying an interpolation weight of at least one of the portions. The speech can indicate at least one location, and the 3D asset is consistent with the location. The speech can indicate at least a plurality of objects, and the 3D asset is consistent with the plurality of objects.

[0007] In another aspect, the device includes at least one computer memory that is not a transitory signal and includes instructions executable by at least one processor to sequentially identify a two-dimensional (2D) object and convert the 2D object into a 3D object, the instructions being executable to modify a first portion of the 3D object but not a second portion to generate a modified 3D object, and use the modified 3D object in a computer game.

[0008] In another aspect, an apparatus includes at least one processor and at least one computer output device configured to be controlled by the processor. The processor is programmed with instructions to identify a two-dimensional (2D) image, convert the 2D image to a 3D asset, and modify a first portion of the 3D asset but not a second portion of the 3D asset. The instructions are executable to use the 3D asset as an object in a computer simulation.

[0009] The details of the present application, both as to its structure and operation, can best be understood in reference to the accompanying drawings, in which like reference numerals refer to like parts, and in which: [Brief description of the drawings]

[0010] [Figure 1] 1 is a block diagram of an exemplary system including an example according to the present principles; [Diagram 2]1 shows an exemplary screenshot prompting a person to input voice for textual identification of assets of a computer simulation. [Diagram 3] 1 illustrates exemplary logic in the form of an exemplary flowchart for converting speech to text into a 3D asset. [Figure 4] 13 shows an exemplary screenshot prompting a person to input an image for generating a computer simulation asset. [Diagram 5] 1 illustrates exemplary logic in the form of an exemplary flowchart for converting an image into a 3D asset. [Figure 6] 1 illustrates exemplary logic in the form of an exemplary flowchart for converting speech to text into locations and portions of 3D assets. [Figure 7] Related example screenshots are shown in FIG. [Figure 8] Related example screenshots are shown in FIG. [Figure 9] An exemplary screenshot is shown in relation to FIG. 6 for modifying a portion of an asset. [Figure 10] 1 illustrates exemplary logic in the form of an exemplary flow chart for modifying a portion of an asset. [Figure 11] 1 illustrates example logic in the form of an example flowchart for closed-loop processing between 3D assets and a physics engine. [Figure 12] Provides an overview of technologies for 2D to 3D asset generation. [Figure 13] We present techniques for controlled feature transformation. [Figure 14] A 2D to 3D reconstruction approach is presented. [Figure 15] We present a technique for generating 3D assets without using 2D input. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] The present disclosure relates generally to computer ecosystems, including, but not limited to, aspects of consumer electronics (CE) device networks, such as computer gaming networks. The systems herein may include server and client components that may be connected through a network such that data may be exchanged between the client and server components. The client components may include one or more computing devices, including gaming consoles such as Sony PlayStation®, or gaming consoles made by Microsoft® or Nintendo® or other manufacturers, virtual reality (VR) headsets, augmented reality (AR) headsets, portable televisions (e.g., smart televisions, Internet-enabled televisions), portable computers such as laptops and tablet computers, and smartphones and other mobile devices, including additional examples described below. These client devices may operate in a variety of operating environments. For example, some of the client computers may employ, by way of example, the Linux® operating system, Microsoft® operating system, or Unix® operating system, or operating systems produced by Apple, Inc.® or Google®. These operating environments may be used to run one or more browsing programs, such as browsers created by Microsoft® or Google® or Mozilla®, or other browser programs that can access web sites hosted by the Internet servers discussed below. Also, operating environments according to the present principles may be used to run one or more computer game programs.

[0012] The server and / or gateway may include one or more processors that execute instructions that configure the server to receive and transmit data over a network such as the Internet. Alternatively, the clients and servers may be connected via a local intranet or a virtual private network. The server or controller may be instantiated by a gaming console such as a Sony PlayStation®, a personal computer, or the like.

[0013] Information can be exchanged between the clients and the servers over a network. For this purpose and for security, the servers and / or clients can include firewalls, load balancers, temporary storage, and proxies, as well as other network infrastructure for reliability and security. One or more servers can form an apparatus for implementing a method for providing a secure community, such as an online social website, for network members.

[0014] The processor may be a single-chip processor or a multi-chip processor capable of implementing logic through various lines, such as address lines, data lines and control lines, as well as registers and shift registers.

[0015] Components included in one embodiment may be used in other embodiments in any suitable combination. For example, any of the various components described herein and / or illustrated in the figures may be combined, interchanged, or excluded from other embodiments.

[0016] "A system having at least one of A, B, and C" (and similarly "a system having at least one of A, B, or C" and "a system having at least one of A, B, C") includes systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B and C together, etc.

[0017] Referring now specifically to FIG. 1, an exemplary system 10 is shown that may include one or more of the exemplary devices described above and further below in accordance with the present principles. A first of the exemplary devices included in the system 10 is a consumer electronics (CE) device, such as an audio-video device (AVD) 12, such as, but not limited to, an Internet-enabled TV with a TV tuner (equivalently, a set-top box that controls the TV). Alternatively, the AVD 12 may also be a computer-controlled Internet-enabled ("smart") phone, a tablet computer, a notebook computer, an HMD, a wearable computer-controlled device, a computer-controlled Internet-enabled music player, a computer-controlled Internet-enabled headphones, a computer-controlled Internet-enabled implantable device such as an implantable skin device, and the like. In any case, it should be understood that the AVD 12 is configured to implement the present principles (e.g., to communicate with other CE devices to implement the present principles, to execute the logic described herein, and to perform any other functions and / or operations described herein).

[0018] Thus, to implement such principles, the AVD 12 may be established by some or all of the components shown in FIG. 1. For example, the AVD 12 may include one or more displays 14, which may be implemented by a high-resolution flat screen or an ultra-high-resolution flat screen of "4K" or higher, and may be touch-enabled for receiving user input signals via touch on the display. The AVD 12 may include one or more speakers 16 for outputting audio in accordance with the principles of the present invention, and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands to the AVD 12 to control the AVD 12. The exemplary AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22, such as the Internet, a WAN, a LAN, etc., under the control of one or more processors 24. A graphics processor may also be included. Thus, the interface 20 may be, without limitation, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, without limitation, a mesh network transceiver. It should be understood that processor 24 controls AVD 12, including the other elements of AVD 12 described herein, such as controlling display 14 to present images and receiving input therefrom, so as to implement the present principles. It should further be noted that network interface 20 may be a wired or wireless modem or router, or other suitable interface, such as a wireless telephony transceiver or Wi-Fi transceiver as discussed above.

[0019] In addition to the above, the ADV 12 may also include one or more input ports 26, such as, for example, a high-definition multimedia interface (HDMI) port or a USB port for physically connecting to another CE device, and / or a headphone port for connecting headphones to the ADV 12 to provide audio from the ADV 12 to a user through the headphones. For example, the input port 26 may be wired or wirelessly connected to a cable or satellite source 26a of audio-video content. Thus, the source 26a may be a separate or integrated set-top box, or a satellite receiver. Alternatively, the source 26a may be a game console or disc player containing the content. When implemented as a game console, the source 26a may include some or all of the components described below in connection with the CE device 44.

[0020] AVD 12 may further include one or more computer memories 28, such as non-transitory, disk-based or solid-state storage devices, in some cases embodied in the AVD's chassis as a stand-alone device or as a personal video recording device (PVR), or as a video disk player or removable memory medium either inside or outside the AVD's chassis for playing AV programs. In some embodiments, AVD 12 may also include a position or location receiver, such as, but not limited to, a cellular receiver, a GPS receiver, and / or an altimeter 30, configured to receive geographic location information from a satellite or cellular base station, provide the information to processor 24, and / or determine the altitude at which AVD 12 is located in conjunction with processor 24. Component 30 may also be implemented by an inertial measurement unit (IMU), typically including a combination of accelerometers, gyroscopes, and magnetometers, to determine the position and orientation of AVD 12 in three dimensions.

[0021] Continuing with the description of the AVD 12, in one embodiment, the AVD 12 may include one or more cameras 32, which may be digital cameras such as thermal imaging cameras, webcams, and / or cameras integrated into the AVD 12 and controllable by the processor 24 to collect pictures / images and / or videos in accordance with the present principles. The AVD 12 may also include a Bluetooth transceiver 34 and other NFC elements 36 for communication with other devices using Bluetooth and / or Near Field Communication (NFC) technologies, respectively. An exemplary NFC element may be a Radio Frequency Identification (RFID) element.

[0022] Furthermore, the AVD 12 may include one or more auxiliary sensors 38 (e.g., motion sensors such as accelerometers, gyroscopes, cyclometers, or magnetic sensors, infrared (IR) sensors, optical sensors, speed and / or cadence sensors, gesture sensors (e.g., sensors for detecting gesture commands) that provide input to the processor 24. The AVD 12 may include an over-the-air TV broadcast port 40 for receiving over-the-air (OTA) TV broadcasts that provide input to the processor 24. In addition to the above, it is noted that the AVD 12 may also include an infrared (IR) transmitter and / or an IR receiver and / or an IR transceiver 42, such as an Infrared Data Association (IRDA) device. A battery (not shown) may be provided to power the AVD 12, or may be a kinetic energy harvester that can convert kinetic energy into electrical power to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field programmable gate array 46 may also be included.

[0023] With further reference to FIG. 1, in addition to the AVD 12, the system 10 may include one or more other CE device types. In one embodiment, the first CE device 48 may be a computer game console that may be used to transmit computer game audio and video to the AVD 12 via commands sent directly to the AVD 12 and / or through a server, as described below, while the second CE device 50 may include similar components to the first CE device 48. In the illustrated example, the second CE device 50 may be configured as a computer game controller operated by a player or a head mounted display (HMD) worn by a player. In the illustrated example, only two CE devices are shown, but it will be understood that fewer or more devices may be used. The devices herein may implement some or all of the components shown for the AVD 12. Any of the components shown in the following figures may incorporate some or all of the components shown for the AVD 12.

[0024] Referring now to the at least one server 52 described above, the server includes at least one server processor 54, at least one tangible computer-readable storage medium 56, such as a disk-based or solid-state storage device, and at least one network interface 58 that, under the control of the server processor 54, enables communication with other devices of FIG. 1 over the network 22, and may, in fact, facilitate communication between the server and client devices in accordance with the present principles. It should be noted that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi® transceiver, or other suitable interface, such as, for example, a wireless telephony transceiver.

[0025] Thus, in some embodiments, server 52 may be an entire Internet server or server "farm" and may include "cloud" functionality and perform such functionality such that devices of system 10 may access the "cloud" environment via server 52, for example, in the exemplary embodiment of a network gaming application. Alternatively, server 52 may be implemented by one or more game consoles, or other computers, in the same room or nearby as the other devices shown in FIG.

[0026] The components illustrated in subsequent figures may include some or all of the components illustrated in FIG.

[0027] 2 and 3 illustrate techniques that enable game designers to create and / or modify three-dimensional (3D) assets for computer simulations such as computer games, typically common assets that are not characters in the first place, or by adapting assets that have been pre-stored in an asset library.

[0028] As shown in FIG. 2, a user interface 200 may be presented on a display 202, such as any of the displays described herein, and in step 204, the designer may be prompted to state the name of a desired asset, such as the name of a chair in the illustrated example.

[0029] 3 shows that the designer's next speech (e.g., "A brown chair with arms, four legs, a cushioned surface and a bar back") is received in block 300 and converted to text in block 302. Block 303 shows that keywords are extracted from the text using a text processing module to extract keywords. In this example, the output of the keyword extraction can be: Object: Chair Color:Brown Legs: 4 legs Surface: Cushioned Back: back bar

[0030] This text may be input to an artificial intelligence (AI) engine, such as one or more neural networks, to generate a 2D image of the requested asset in block 304. The image may be generated from scratch or selected by accessing a library of assets. A search of the library may first be performed for images matching the keywords, and only if no match is found can the AI ​​engine generate an image of the asset using the text in a 2D or 3D generative model based on supervised or unsupervised training in human language.

[0031] Proceeding from block 304 to block 306, the 2D images are converted to 3D assets of the asset using a 2D to 3D conversion system, for example, using layer stacking or other techniques, such as creating 3D anaglyph stereograms, false height relief, etc. A 2D to 3D reconstruction model can be used. An encoder-decoder neural architecture can be included, where an encoder receives a 2D image as input and generates an encoding, and a 3D decoder generates a 3D object based on the encoding. Thus, a 3D object or asset can be generated using 2D to 3D reconstruction, a generative neural model can be used to generate a 3D object and then transform it to the specifications, or an existing 3D model can be transformed according to the desired specifications. Further details are described in Figures 5 and 12-15.

[0032] The 3D asset may be presented, for example, on a display as shown in Figure 2 and may receive artist modifications to the asset at block 308 using voice or other input, such as point-and-click device graphic manipulation input, which may include changing the size, shape, color, style of certain parts of the asset (but not all parts of the asset), the texture of the asset's surface, etc. The final modified 3D asset is generated at block 310 for use in the computer simulation.

[0033] 4 illustrates a UI 400 that may be presented on a display 402, such as any display specified herein, to prompt a user to input a photograph of a desired asset at 404. The photograph is depicted in 2D format at 406 and can be uploaded for processing in FIG.

[0034] 5 shows that a 2D image of an asset in a photograph is received at block 500. Moving to block 502, the 2D image is converted into a 3D asset. Proceeding to block 504, the 3D asset can be modified as described herein by an artist or other user for use in a computer simulation. Additional details of the generation of the 3D asset are provided in FIGS. 12-15, which are described below.

[0035] 6 shows exemplary logic for specifying multiple assets and their desired relative positions to one another in a computer simulation. Beginning at block 600, text is received, either from direct text input or from a voice-to-text conversion, describing the assets by name and their desired relative positions to one another.

[0036] Proceeding to step 602, a description of only a portion of an asset that does not apply to the entire asset may be received if desired. If the description is received as voice input, it is converted to text in block 604. An AI engine such as a generative adversarial network (GAN) may be used in block 606 to generate a 2D image based on the previously received asset description and location, which is converted to a 3D scene in block 608 according to the principles described herein. The 3D asset may be generated directly without going through a 2D phase.

[0037] 7 shows: A UI 700 may be presented on a display 702, such as any display described herein. The UI 700 may include a prompt 704 for a person to speak a description of a desired asset scene, which may be presented in text form after speech-to-text conversion at 706. In the illustrated example, the person is specifying a scene with a couch in front and to the left of a chair shaped as a Gaudi-style chair.

[0038] Figure 8 shows an example result of the process of Figure 7. Continuing with the example shown in Figure 7, to the left and in front of the chair 3D asset 802, a 3D model of a couch 800 is shown, with the back of the chair 804 in Gaudi style depicted with ruffles 806. Labels 808 can also be presented with each asset indicating what the asset is intended to depict, allowing the artist to verify if the GAN has performed the desired task correctly.

[0039] One approach to verifying the labels is to render the 3D model into a 2D image and use a similarity metric to compare the similarity between the 2D image generated from the text and the 2D image rendered from the 3D model.

[0040] 9 illustrates a UI 900 that may be presented on a display 902, such as any of the displays described herein. The UI 900 may include text 904 that illustrates a speech-to-text conversion from an artist's voice input to modify, for example, the chair shown in FIG. 8 from a Gaudi style to a Louis XIV style in the illustrated example, which results in the frills on the back of the chair shown in FIG. 8 changing to a more decorative and elegant style following the example given.

[0041] 10 illustrates further principles related to the above disclosure. At block 1000, text, for example convertible from speech, is received to indicate a desired modification to an asset. Based on the desired modification, at block 1002, portions of the relevant asset are appropriately composited together to satisfy the requested modification. This can be done by varying weights of the interpolated pixels along boundary regions in the asset that are identified as being associated with the desired modification.

[0042] In addition to the assets, the artist may also vocally describe the desired background terrain, e.g., "mud" or "marble palace" or other terrain. Also, as mentioned above, the size of the asset may be specified by the artist. For example, the artist may specify a chair that is 20 feet tall. This may result in the ceiling automatically appearing to deform to accommodate the chair if the asset, once incorporated into the simulation's game space, interferes with other assets, such as the roof of an object. This may require a collaborative human-AI method. More qualitative requirements, such as a wide seat or a tall chair, may be met using an AI-only approach.

[0043] 11 illustrates an additional embodiment. Once a 3D asset is created as described herein, it may be input to a physics engine at block 1100. Proceeding to block 1102, the geometry of the asset may be modified, for example by a GAN, to maintain a constant inertia tensor calculated by the physics engine for the tendency to move or deform the asset. The inertia tensor may then be solved by the physics engine to describe how the asset reacts to forces. For example, the physics engine may determine whether a tip will point when pushed with a particular force based on the current structural features of the generated 3D asset.

[0044] In other words, the AI ​​engine can look at the physical properties of the asset's structure, predict how the structure will react to physics, and determine how to maintain the physical ratios of the preceding object. Constraints may be imposed for this purpose. For example, if the asset is a piece of furniture, it must be generated with attributes that prevent it from tipping over no matter how heavy the top of the 3D asset that may be emulated is, which may be achieved, for example, by keeping the sum of the torques of the various parts of the asset at zero, for example, by appropriately diversifying the dimensions and weights of the parts of the asset. In other words, a rule-based approach may be combined with AI to generate the object itself. The updated asset (or its physics determination) is fed back to the AI ​​engine in block 1104.

[0045] In addition to visual properties, the techniques described herein can be used to modify the acoustic and material properties of an asset using separate respective AI engines such as GANs. For example, GANs can be used to characterize an asset with respect to how it absorbs forces, such as whether the asset shatters or breaks when hit by a bullet, or absorbs the bullet. An asset representing a grenade can be designed to have different types of explosions in the presence of different assets.

[0046]

[0013] Referring now to Figure 12, an overview of a technique for 2D to 3D graphic asset generation is shown. The technique of Figure 12 is useful for new assets or when it is not feasible to convert existing 3D models. This technique supports generation and conversion.

[0047] Beginning at step 1200, to implement the above example, a representation 1202, such as a photo, of a real 2D object, such as a chair, is input to a conditional generative neural model for 2D synthesis. The resulting output 1204 is a 2D synthesized representation of the chair. The output 1204 is sent to an optional 2D transformation model 1206 for interpolation and feature editing. The model 1206 may be fully AI-based, or it may be interactive between the AI ​​model and a human operator.

[0048] The 2D transformation model 1206, in the illustrated example, outputs a transformed composite representation 1208 of the chair in 2D. The representation 1208 can be included in an asset library and used for artist input and for 3D reconstruction.

[0049] In practice, the transformed synthetic representation 1208 and / or representation 1202 of a real asset in 2D, such as a chair, can be input to a neural model 1210, which converts the 2D representation to a 3D shape and outputs a reconstructed mesh 1212 of the asset. The neural model 1210 optionally includes implicit functions and mesh deformations. If desired, the reconstructed mesh 1212 can be input to a texture transformation model 1214 for neural rendering of the 3D asset's texture.

[0050] 13 illustrates controlled feature transformation. Starting at block 1300, a 2D generative model (such as a generative adversarial network (GAN)) is trained on individual asset classes, such as tables and chairs, to generate assets. Training can be supervised, semi-supervised, or unsupervised.

[0051] When an asset is requested, the appropriate trained model is selected for the asset specified in the description. For example, if there are separate models for generating chairs, tables, etc., the model is selected based on the specified asset.

[0052] Typically, an artist specifies the characteristics of the asset to be transformed, such as texture, color, and shape (geometry). To transform the generated asset to meet the specifications of the input description, the generation is adjusted based on keywords (e.g., attributes), which can be considered as annotated features (y-labels), extracted from the description in block 1302. In the example, five features of a chair can be used: arms, legs, back, surface, and view (e.g., front or rear).

[0053] Proceeding to block 1304, an encoding can be generated for the annotated chair with various weights, which can be interpolated to best fit the artist's specifications. The encoding is sent to train a supervised classifier 1306 to find the feature axes F(i). In block 1308, the features can be edited along with the feature axes for the new chair, allowing certain features to be interactively controlled and attributes transformed (human AI collaboration). For example, an existing chair asset is changed to a chair with a backrest. Thus, the encoding W' for the new chair is the product of alpha and the feature axes F(i), plus the encoding W of the pre-existing chair, where alpha can be empirically determined or discovered.

[0054] Figure 14 shows another approach. A representation 1400 of a real or synthetic chair in 2D is sent to a 2D encoder-decoder neural model 1402 for shape encoding. The 2D encoder model 1402 may be a convolutional network or a similar deep neural network. The input 1400 to the encoder model 1402 may be the image generated in Figure 13 and (optionally) transformed to satisfy the description of the desired asset. If desired, a texture encoder 1404 may also be provided to encode the texture of the object.

[0055] The 3D decoder 1406 receives the input encoding and generates a 3D object. The 3D decoder 1406 can also be a convolutional network or similar DNN. The output of the 3D decoder is a reconstructed mesh 1408 that represents the 3D asset.

[0056] To train the network, the 3D output can be rendered into a 2D image and compared to the input image. Training can continue iteratively until the input and output closely match. Alternatively, mesh deformation can be used.

[0057] The encoder-decoder model can be adapted to incorporate additional encoding (eg, texture coding) that transforms the 3D object to meet the specifications in the description.

[0058] Turning to FIG. 15 for an alternative approach to generating a 3D asset, a 3D GAN model is trained in block 1500 to generate a 3D object. Partial encodings for each part of the asset are extracted in block 1502, e.g., arms, legs, back, etc. for a chair. Proceeding to block 1504, the part encodings are transformed based on a shape description 1506 of the desired asset. Proceeding to block 1508, the generation of the 3D asset is adjusted based on an appearance description 1510, such as style or non-shape descriptions such as size or color. The reconstructed mesh 1512 of the 3D asset is output as desired, with or without texturing. That is, the 3D asset model can be rendered based on the specified texture. 3D variations can be generated based on the specified attributes.

[0059] While the present principles have been described with reference to certain illustrative embodiments, it will be understood that these are not intended to be limiting and that a variety of alternative configurations may be used to implement the subject matter claimed herein.

Claims

1. 1. A method comprising: receiving at least one two-dimensional (2D) image of a computer simulation object; converting said 2D images into three-dimensional (3D) assets; creating a modified asset by modifying a first portion of the 3D asset and leaving a second portion of the 3D asset unchanged using at least one neural network; and presenting the modified asset in at least one computer simulation; Including, the modification includes modifying one or more of a color of the first portion, a style of the first portion, and a texture of a surface of the first portion; The method further comprising inputting the 3D asset into a physics engine to modify the geometry of the 3D asset according to an inertia tensor calculated by the physics engine for tending to move or deform the 3D asset.

2. A method comprising: receiving at least one two-dimensional (2D) image of a computer simulation object; converting said 2D images into three-dimensional (3D) assets; creating a modified asset by modifying a first portion of the 3D asset and leaving a second portion of the 3D asset unchanged using at least one neural network; and presenting the modified asset in at least one computer simulation; Including, the modification includes modifying one or more of a color of the first portion, a style of the first portion, and a texture of a surface of the first portion; The method, wherein the modifying comprises at least partially blending a first portion and a second portion of the computer simulation object with each other by modifying a weight of at least one interpolated pixel of the portion.

3. The method of claim 1 or 2, wherein the text used to generate the 2D image is generated from speech recognition.

4. The method of claim 1 or 2, wherein the 2D image is generated from a camera image of the computer simulation object.

5. The method of claim 1 or 2, wherein the 2D image is generated from audio, the audio indicating at least one location, and the 3D asset is positioned in a computer simulation at the location indicated by the audio.

6. A device, comprising: at least one computer memory containing instructions executable by at least one processor that are not transitory signals, said instructions comprising: Identifying two-dimensional (2D) objects; Converting the 2D object into a 3D object; modifying a first portion of the 3D object but not a second portion to generate a modified 3D object; using the modified 3D object in a computer game, the modification modifying a color of the first portion; The device, wherein the instructions are executable to maintain a sum of torques of various portions of the 3D object at zero by diversifying dimensions and weights of the portions of the 3D object.

7. A device comprising: at least one computer memory containing instructions executable by at least one processor that are not transitory signals, said instructions comprising: Identifying two-dimensional (2D) objects; Converting the 2D object into a 3D object; modifying a first portion of the 3D object but not a second portion to generate a modified 3D object; using the modified 3D object in a computer game, the modification modifying a color of the first portion; The device, wherein the instructions are executable to define characteristics of how the 3D object absorbs forces.

8. The 2D object is at least partially receiving a voice description of the 2D object; and using artificial intelligence (AI) to generate the 2D object from the voice description; 8. The device according to claim 6 or 7, characterized by:

9. 8. A device according to claim 6 or 7, wherein the instructions are executable to associate audio with the 3D object based at least in part on a voice description.

10. 8. A device according to claim 6 or 7, wherein the instructions are executable to receive an audio indicative of at least one location, the 3D object being consistent with the location.

11. 8. The device of claim 6 or 7, wherein the instructions are executable to receive an audio indicative of at least a plurality of objects, and the 3D object is consistent with the plurality of objects.

12. An apparatus comprising: At least one processor; at least one computer output device configured to be controlled by said processor; wherein the processor: Identifying a two-dimensional (2D) image; converting said 2D images into 3D assets; modifying a first portion of the 3D asset but not modifying a second portion of the 3D asset; programmed with instructions to use the 3D asset as an object in a computer simulation after the modification, the modification including changing a style of the first portion; The apparatus inputs the 3D asset into a physics engine to modify a geometry of the 3D asset according to an inertia tensor calculated by the physics engine for tending to move or deform the 3D asset.

13. An apparatus comprising: At least one processor; at least one computer output device configured to be controlled by said processor; wherein the processor: Identifying a two-dimensional (2D) image; converting said 2D images into 3D assets; modifying a first portion of the 3D asset but not modifying a second portion of the 3D asset; programmed with instructions to use the 3D asset as an object in a computer simulation after the modification, the modification including changing a style of the first portion; The apparatus, wherein the modifying includes at least partially blending the first and second portions of the 3D asset into one another by modifying a weight of at least one interpolated pixel of the portions.

14. An apparatus comprising: At least one processor; at least one computer output device configured to be controlled by said processor; wherein the processor: Identifying a two-dimensional (2D) image; converting said 2D images into 3D assets; modifying a first portion of the 3D asset but not modifying a second portion of the 3D asset; programmed with instructions to use the 3D asset as an object in a computer simulation after the modification, the modification including changing a style of the first portion; The apparatus, wherein the modification includes maintaining a sum of torques of various portions of the 3D asset at zero by diversifying dimensions and weights of the portions of the 3D asset.

15. An apparatus comprising: At least one processor; at least one computer output device configured to be controlled by said processor; wherein the processor: Identifying a two-dimensional (2D) image; converting said 2D images into 3D assets; modifying a first portion of the 3D asset but not modifying a second portion of the 3D asset; programmed with instructions to use the 3D asset as an object in a computer simulation after the modification, the modification including changing a style of the first portion; The device, wherein the modification includes defining characteristics of how the 3D asset absorbs forces.

16. The instruction:

16. Apparatus according to any one of claims 12 to 15, operable to identify the 2D image based at least in part on input of a photograph of the 2D image.

17. The instruction:

16. An apparatus according to any one of claims 12 to 15, wherein the apparatus is operable to identify the 2D image based at least in part on a text input describing the 2D image.

18. The instruction: The apparatus of claim 17 , wherein the apparatus is operable to derive the text input from a voice input.

19. The instruction:

20. The apparatus of claim 17, wherein the apparatus is operable to generate the 2D image based at least in part on a text input describing the 2D image using at least one neural network.

20. The instruction:

16. The apparatus of claim 12, wherein the apparatus is operable to associate audio with the 3D asset based at least in part on text input.

Citation Information

Patent Citations

  • System, server, and user computer for image control, and storage medium

    JP2002288689A

  • Information processing device

    JP2017182129A

  • Machine learning for estimating 3D modeled object

    JP2020115336A

  • JPP6843409B