Hyper-personalized game items

Neural Radiance Field technology with CLIP model facilitates rapid generation of hyper-personalized game items from text, addressing inefficiencies in character and equipment creation for video games.

JP2026053447APending Publication Date: 2026-03-25SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Creating characters and equipment for video games can be time-consuming and inefficient.

Method used

Utilizing a Neural Radiance Field (NeRF) technology combined with a Contrastive Language-Image Pretraining (CLIP) model to generate hyper-personalized game items from text inputs, enabling rapid creation of virtual characters and equipment in computer simulations.

Benefits of technology

Enables easy, rapid, and intuitive generation of hyper-personalized game items, allowing for real-time conversion of text descriptions into 3D virtual equipment within two minutes, enhancing game development efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053447000001_ABST
    Figure 2026053447000001_ABST
Patent Text Reader

Abstract

This invention provides a method, system, and non-temporary computer-readable recording medium for rapidly generating hyper-personalized game items. [Solution] A two-dimensional image is converted into a 3D neural luminance field (NeRF) (302), modified (402) based on text personalized to the player, and input to resemble the character equipment (1700) requested by the text. The model scores how well the lines of text match the image (404) to produce the final 3D NeRF, which can be converted into a polygon mesh (408) and imported into computer simulations such as computer games.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to rapidly generating hyper-personalized game items.

Background Art

[0002] As understood herein, creating characters such as non-player characters (NPCs) for computer simulations such as computer games and their equipment can sometimes take time.

Summary of the Invention

[0003] As further understood herein, it is desirable to enable game developers to create characters and equipment for video games in an easy, rapid, and intuitive manner.

[0004] Thus, the device includes at least one computer storage that is not a temporary signal and includes instructions executable by at least one processor to generate a base mesh or a Neural Radiance Field (NeRF) from multiple images. Techniques from text to 2D images (e.g., stable diffusion) or from text to meshes can be used directly, but note that in all embodiments, this technique is applied to the real-time creation of hyper-personalized in-game items.

[0005] In the context of NeRF, the instructions are executable to generate a modified NeRF from a base NeRF using a text input to a Contrastive Language-Image Pretraining (CLIP) model and to convert the modified NeRF into a polygon mesh representing a virtual character to present the equipment in at least one computer simulation.

[0006] In some examples, the CLIP model evaluates the match of an image to text.

[0007] The text is preferably derived from player information that players share, such as the title of at least one computer simulation. The text may describe equipment for game characters, such as masks. Instructions may be executable to generate text from a starting phrase using learned subsequent phrases.

[0008] In another embodiment, the device includes at least one processor programmed with instructions to receive a text description of equipment, which is personalized to player data. The instructions are executable to generate a virtual three-dimensional (3D) equipment less than two minutes after receiving the text description, based at least in part on the text description, and to present the virtual equipment on a display.

[0009] In another embodiment, the method includes receiving text based on data relating to a player of a computer simulation and generating a neural luminance field based on text starting from a base model.

[0010] Details of this application, both in terms of its structure and operation, can be best understood by referring to the attached drawings, where similar reference numbers in the drawings refer to the same parts. [Brief explanation of the drawing]

[0011] [Figure 1] This is a block diagram of an exemplary system based on this principle. [Figure 2] An exemplary initial logic is shown in an exemplary flowchart format. [Figure 3] This demonstrates the creation of a 3D neural luminance field (NERF) representing a human head from a 2D image. [Figure 4] An example of text-based NERF customization logic is shown in an illustrative flowchart format. [Figure 4A] This is another way of expressing the exemplary logic. [Figure 5]The first descriptive phrase is used to indicate a customized NeRF. [Figure 6] The first descriptive phrase is used to indicate a customized NeRF. [Figure 7] The first descriptive phrase is used to indicate a customized NeRF. [Figure 8] The first descriptive phrase is used to indicate a customized NeRF. [Figure 9] The second descriptive phrase is used to indicate a customized NeRF. [Figure 10] A customized NeRF is shown using a third descriptive phrase. [Figure 11] This document outlines a series of image generation steps to learn how to customize a general-purpose NeRF. [Figure 12] This shows the data structure for generating text. [Figure 13] Additional example text-based NeRF customization logic is shown in an illustrative flowchart format. [Figure 14] An example of text-based equipment customization logic is shown in an illustrative flowchart format. [Figure 15] An example of a screenshot matching Figure 14 is shown. [Figure 16] Figure 14 shows an example of player-customized game equipment. [Figure 17] This screenshot shows an example of equipment customized for a specific player. [Figure 18] This screenshot shows an example of equipment customized for a specific player. [Modes for carrying out the invention]

[0012] This disclosure generally relates to a computer ecosystem including, but not limited to, a computer game network or a consumer electronics (CE) device network. The system herein may include server and client components that can be connected via a network so that data can be exchanged between the client and server components. The client components may include one or more computing devices, including a game console such as a Sony PlayStation®, or a game console from Microsoft, Nintendo, or another manufacturer, an extended reality (XR) headset such as a virtual reality (VR) headset or an augmented reality (AR) headset, a portable television (e.g., a smart TV or an internet-enabled television), a portable computer such as a laptop or tablet computer, and other mobile devices, including smartphones and additional examples described below. These client devices may operate in a variety of operating environments. For example, some client computers may use, for example, a Linux® operating system, a Microsoft operating system, or a Unix® operating system, or an operating system from Apple or Google, or a Berkeley Software Distribution or a Berkeley Standard Distribution of a BSD-based OS. These operating environments may be used to run one or more browsing programs, such as browsers from Microsoft, Google, or Mozilla, or other browser programs that can access websites hosted by the Internet servers described below. Furthermore, operating environments based on these principles may be used to run one or more computer game programs.

[0013] A server and / or gateway that may include one or more processors for executing instructions to configure a server to transmit and receive data via a network such as the Internet may be used. Alternatively, the client and server can be connected via a local intranet or virtual private network. The server or controller may be instantiated by a game console such as Sony PlayStation (registered trademark), a personal computer, or the like.

[0014] Information can be exchanged between the client and server via the network. For this purpose and for security, the server and / or client can include a firewall, load balancer, temporary storage, and proxy, as well as other network infrastructure for reliability and security. One or more servers can form an apparatus for implementing a method of providing a secure community such as an online social website or gamer network to network members.

[0015] The processor may be a single-chip processor or a multi-chip processor that can execute logic by various lines such as address lines, data lines, and control lines, as well as registers and shift registers. A processor including a digital signal processor (DSP) can be an embodiment of the circuit.

[0016] Components included in one embodiment can be used in any suitable combination in other embodiments. For example, any of the various components described and / or illustrated herein can be integrated, substituted, or excluded from other embodiments.

[0017] "A system having at least one of A, B, and C" (similarly, "a system having at least one of A, B, or C" and "a system having at least one of A, B, and C") includes systems having A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, and C together.

[0018] Referring here to Figure 1, an exemplary system 10 is shown, which may include one or more of the exemplary devices described above, and may include one or more of the exemplary devices further described below, which operate according to the present principle. The first exemplary device included in system 10 is a consumer electronics (CE) device such as an audio-video device (AVD) 12, which may be, for example, a theater display system, an internet-enabled television with a projector base or TV tuner (equivalently, a set-top box that controls the TV). The AVD 12 may instead be a computer-controlled internet-enabled ("smart") phone, a tablet computer, a notebook computer, a head-mounted device (HIVID) and / or a headset such as smart glasses or a VR headset, another wearable computer device, a computer-controlled internet-enabled music player, a computer-controlled internet-enabled headphones, a computer-controlled internet-enabled implantable device such as an implantable skin device, etc. In any case, it should be understood that the AVD 12 is configured to implement the present principle (e.g., to communicate with other CE devices to implement the present principle, to execute the logic described herein, and to perform any other functions and / or operations described herein).

[0019] Therefore, to implement such a principle, AVD12 can be established by some or all of the components shown. For example, AVD12 may include one or more touch-enabled displays 14, which can be implemented by high-resolution or ultra-high-resolution "4K" or higher flat screens. The touch-enabled display(s) 14 may include, for example, a capacitive or resistive touch-sensing layer with a grid of electrodes for touch sensing consistent with the present principle.

[0020] The AVD12 may also include one or more speakers 16 for outputting audio in accordance with this principle, and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands to the AVD12 to control it. The exemplary AVD12 may also include one or more network interfaces 20 for communicating over at least one network 22, such as the Internet, WAN, LAN, etc., under the control of one or more processors 24. Thus, the interface 20 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as a mesh network transceiver. It should be understood that the processor 24 controls the AVD12 to implement this principle, including other elements of the AVD12 described herein, such as controlling a display 14 to present images and receiving input therefrom. Furthermore, it should be noted that the network interface 20 may be a wired or wireless modem or router, or other suitable interface such as a wireless telephone transceiver or a Wi-Fi transceiver as described above.

[0021] In addition to the above, the AVD12 may also include one or more input ports and / or output ports 26, such as a High Definition Multimedia Interface (HDMI®) port or a Universal Serial Bus (USB) port for physically connecting to another CE device, and / or a headphone port for connecting headphones to the AVD12 in order to present audio from the AVD12 to the user via headphones. For example, the input port 26 may be wired or wirelessly connected to a cable or satellite source 26a of audio video content. Thus, the source 26a may be a separate or integrated set-top box or satellite receiver. Alternatively, the source 26a may be a game console or disc player containing content. When implemented as a game console, the source 26a may include some or all of the components described below in relation to the CE device 48.

[0022] The AVD12 may further include one or more computer memory / computer-readable storage media 28, such as disk-based or solid-state storage, which are not transient signals, and may be embodied in the AVD chassis as a standalone device, or as a personal video recording device (PVR) or video disc player, either inside or outside the AVD chassis, for playing AV programs, or as a removable storage medium or a server as described below. In some embodiments, the AVD12 may also include a location or place receiver, such as a cell phone receiver, a GPS receiver, and / or an altimeter 30, configured to receive geographic location information from a satellite or cell phone base station and provide that information to the processor 24, and / or to determine the altitude at which the AVD12 is located in cooperation with the processor 24.

[0023] Continuing the description of AVD12, in some embodiments, AVD12 may include one or more cameras 32, which may be digital cameras such as infrared cameras and webcams, IR sensors, event-based sensors, and / or cameras integrated into AVD12 and controllable by processor 24 for collecting photographs / images and / or videos in accordance with the present principle. AVD12 may also include a Bluetooth transceiver 34 and other NFC elements 36 for communicating with other devices using Bluetooth® and / or Near Field Communication (NFC) technology, respectively. An exemplary NFC element may be a radio frequency identification (RFID) element.

[0024] Furthermore, the AVD12 may also include one or more auxiliary sensors 38 that provide input to the processor 24. For example, one or more of the auxiliary sensors 38 may include, but are not limited to, one or more pressure sensors that form a layer of the touch-enabled display 14 itself, such as piezoelectric pressure sensors, capacitive pressure sensors, piezoresistive strain gauges, optical pressure sensors, and electromagnetic pressure sensors. Examples of other sensors include motion sensors such as pressure sensors, accelerometers, gyroscopes, cyclometers, or magnetic sensors, infrared (IR) sensors, optical sensors, velocity and / or cadence sensors, event-based sensors, and gesture sensors (for example, for sensing gesture commands). Thus, the sensors 38 may be implemented by an inertial measurement unit (IMU) which includes one or more motion sensors such as individual accelerometers, gyroscopes, and magnetometers, and / or usually a combination of accelerometers, gyroscopes, and magnetometers, or by an event-based sensor such as an event detection sensor (EDS) to determine the position and orientation of the AVD12 in three dimensions. An EDS consistent with this disclosure provides an output indicating a change in light intensity sensed by at least one pixel of a photosensing array. For example, if the light sensed by the pixel is decreasing, the output of the EDS may be -1, and if it is increasing, the output of the EDS may be +1. No change in light intensity below a certain threshold may be indicated by an output binary signal of 0.

[0025] The AVD12 may also include an over-the-air TV broadcast port 40 for receiving OTA TV broadcasts that provide input to the processor 24. In addition to the above, it should be noted that the AVD12 may also include an infrared (IR) transmitter and / or IR receiver and / or IR transceiver 42, such as an infrared data association (IRDA) device. A battery (not shown) may be provided to power the AVD12, which may be a kinetic energy harvester that converts kinetic energy into power to charge the battery and / or power to power the AVD12. A graphics processing unit (GPU) 44 and a field-programmable gate array 46 may also be included. One or more tactile / vibration generators 47 may be provided to generate tactile signals that can be sensed by a person holding the device or a person in contact with the device. Therefore, the tactile generator 47 may vibrate all or part of the AVD 12 using an electric motor connected to an uneven weight and / or unbalanced weight via a rotatable shaft of the motor, so that the shaft can rotate under the control of the motor (which may be controlled by a processor such as the processor 24), and can create vibrations of various frequencies and / or amplitudes, and simulations of forces in various directions.

[0026] Light sources such as infrared (IR) projectors may also be included.

[0027] In addition to AVD12, System 10 may include one or more other CE device types. For example, the first CE device 48 may be a computer game console that can be used to transmit audio and video of a computer game to AVD12 via commands sent directly to AVD12 and / or via a server described later, while the second CE device 50 may include components similar to the first CE device 48. In the illustrated example, the second CE device 50 may be configured as a computer game controller operated by a player, or as a head-mounted display (HMD) worn by the player. The HMD may include a transparent head-up display or an opaque head-up display that presents AR / MR content or VR content (more broadly, Extended Reality (XR) content), respectively. The HMD may be configured as glasses-type displays or as large VR displays sold by computer game equipment manufacturers.

[0028] Although only two CE devices are shown in the illustrated example, it should be understood that fewer or more devices may be used. The devices described herein may implement some or all of the components shown for AVD12. Any of the components shown in the following diagrams may incorporate some or all of the components shown in the AVD12 example.

[0029] Referring here to the at least one server 52 described above, it includes at least one server processor 54, at least one tangible computer-readable storage medium 56 such as disk-based or solid-state storage, and at least one network interface 58 that, under the control of the server processor 54, enables communication with other illustrated devices via the network 22 and facilitates communication between the server and client devices in accordance with this principle. Note that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface such as a wireless telephone transceiver.

[0030] Therefore, in some embodiments, server 52 may be an internet server or an entire server "farm," and may include, or perform, a "cloud" function such that, for example, in an exemplary embodiment for a network game application, devices of system 10 can access the "cloud" environment via server 52. Alternatively, server 52 may be implemented by one or more game consoles or other computers in the same room or near the other illustrated devices.

[0031] The components shown in the following diagrams may include some or all of the components shown herein. Any user interface (UI) described herein may be aggregated and / or expanded, and UI elements may be mixed and matched between UIs.

[0032] This principle can utilize a variety of machine learning models, including deep learning models. Machine learning models consistent with this principle can use a variety of algorithms trained in ways including supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature representation learning, self-learning, and other forms of learning. Examples of such algorithms that can be implemented by computer circuits include one or more neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and a type of RNN known as a long-short-term memory (LSTM) network. Support vector machines (SVMs) and Bayesian networks may also be considered examples of machine learning models. In addition to the types of networks described above, the models herein may be implemented by classifiers.

[0033] Therefore, as understood herein, performing machine learning may involve training a model by accessing training data so that the model can further process the data and perform inference. Thus, an artificial neural network / artificial intelligence model trained through machine learning may include input layers, output layers, and multiple hidden layers between them that are configured and weighted to make estimates about appropriate outputs.

[0034] Refer to Figures 2 and 3 here. The initial logic, which can be executed by any or more processors herein, begins in block 200 by receiving a set of images of the same base head (labeled 300 in Figure 3), which may be 2D images of the same base head, from different viewpoints, if generated by one or more artists. From these images, a 3D base head (302 in Figure 3) is generated in block 202. The 3D base head 302 may also be a 3D neural luminance field (NeRF), which can be thought of as a 3D volume stored in a machine learning (ML) model. In certain examples, the base head 302 may be an NeRF that can be rapidly trained in less than two minutes, and in some examples less than one minute, after the following text description is input, to generate a modified NeRF, which will be further described below. An exemplary NeRF for this purpose is an NeRF encoded in a multiresolution hash table, which stores the positions of a sparse multiresolution 3D grid to speed up the training and rendering of the NeRF.

[0035] It should be noted that, instead of an image, in some embodiments, the base head 302 in Figure 3 may be generated using a head scan of a real-world head.

[0036] Refer to Figure 4 here. Once the base head 302 is generated, the text description of the modified base head may be received in block 400 of the desired head. The text may be entered from a text input device such as a keyboard, and / or from speech-to-text conversion of an oral description, and / or from an initial introductory phrase followed by learned additional descriptive phrases, as further described herein. Alternatively, the text may be generated in a Madrib style using more primitive algorithmic techniques so that the survey response can be inserted into the text describing the target head.

[0037] Moving to block 402, the text is input to a NeRF modification engine, which may include a fully connected (non-convolutional) deep network. The input to the modification engine can be a single continuous 5D input containing values ​​representing three spatial dimensions and two visual dimensions, while the engine's output may include volume density and view-dependent radiance at spatial locations represented by associated 3D values.

[0038] Moving on to block 404, the output of the NeRF correction engine, which can be considered a corrected NeRF, is assessed to determine how closely the output matches the input text description. In one example, this assessment may be performed by a contrasting language-image pre-trained (CLIP) model that evaluates the image match to the input text string. CLIP is an open-source model that scores how well lines of text match images. The CLIP model can be trained on image-text pairs, such as captioned photographs from the internet, using cosine similarity to score the goodness of the match. Thus, the correction to the base NeRF is essentially derived using a text input-dependent loss function to the CIP model.

[0039] Decision diamond 406 indicates that if, after the current iteration of the correction, it is determined that the loss has not reached a threshold (target small) loss, then in block 408, the current iteration of the corrected NeRF is corrected again and the logic loops back to block 408. However, if it is determined that the loss is sufficiently small, the logic proceeds to block 410 to output the final corrected NeRF. Note that instead of a loss threshold, a set number of loops, e.g., 100 loops, may be performed. Blocks 402-408 may use gradient descent techniques.

[0040] In some embodiments, the logic then moves to block 412, where the modified NeRF can be converted into a polygon mesh, which can then be imported into a computer simulation such as a computer game in block 414 and presented on a display during gameplay in block 416. Block 412 can also generate material properties for the mesh, such as albedo, roughness, and normal textures. The modified NeRF can be converted into a polygon mesh having the same topology as the modified NeRF. This may be done using the marching cubes method. Alternatively, the modified NeRF can be converted into a polygon mesh by essentially shrink-wrapping the modified NeRF by placing a mesh on top of the modified NeRF and recording the parts of the mesh that come into contact with the surface of the modified NeRF when emulating the mesh being vacuum-evacuated.

[0041] Figure 4A is an alternative diagram. Starting in block 420, the NeRF is generated from parameters that determine how the head looks and may include values ​​that are adjusted so that the final NeRF looks the same as the input image. Moving to block 422, the NeRF may be rendered from multiple camera angles, as well as extensions such as color, rotation, and affine applied in block 424.

[0042] Moving on to block 426, in block 428 we can obtain and use the embedding of the CLIP image, and compare the current iteration of NeRF with the embedding of the text image to obtain a loss value. In block 430, the loss value is added to other conditions such as how "cloudy" NeRF is by obtaining the sum of the depth variances of the rendered image. If parameter modifications are desired in block 420, the sum of losses is backpropagated using gradient descent in block 432.

[0043] Figures 5–8 show four perspective views of the modified NeRF500 generated from the base NeRF302 in Figure 3 using the input text "photorealistic head of a robust old man". The modified NeRF500 was generated in less than two minutes after the text was entered into the model.

[0044] Figures 9 and 10 show the modified NeRF900 and NeRF1000 generated from the base NeRF302, respectively, using the input texts "photorealistic head of a sea monster" and "photorealistic head of a demon."

[0045] Figure 11 shows that the above ML model, particularly for use with the mesh method, can be trained with a chain of causal relationships from initial parameters (variables) controlling the vertices of an object to pixels rendered on the screen. Starting from vertex 1100, the chain of causal relationships can proceed to primitive 1102, which is based on vertex 1100, and then further to fragment 1104, which is at least partially filled with, for example, a texture. From fragment 1104, a processed fragment 1106 can be generated to which color can be added, and the processed fragment 1106 is merged into a pixel output 1108 presented on the screen. Arrow 1110 shows that this chain can be traced back to the original vertex 1100 to see how certain values ​​in the chain affect the final pixel image 1108, which can be used to train the ML model.

[0046] Note that if a head mesh is used instead of a luminance field, the ML model can perform differentiable rendering, allowing you to modify parameters and see in real time how the variables affect the output image. The ML model uses a loss function to minimize loss as new faces begin to match the text-defined target.

[0047] As described above, part or all of the input text describing the desired modified NeRF can be generated from the starting phrase using learned subsequent phrases. Figure 12 shows this. The prompt template 1200 may be presented on the developer's computer screen with multiple phrases that the player / developer can input in response. Additional text 1202 may be presented for each prompt response. The additional text 1202 may be generated by an ML model trained to complete sentences based on the input phrases.

[0048] Figure 13 illustrates additional technology. Beginning in block 1300, text is received indicating the character's history and the gear the character is using to start the game. A text-image generator 1302 receives the text and generates an image consistent with the principles herein, the image is imported into the computer game in block 1304 and presented while playing the game in block 1306. Block 1302 may instead use a text-3D mesh generator (such as one from NeRF) or an image such as a decal on the player's gear, but it should be noted that the more dominant use case is generating actual 3D assets in mesh / material form.

[0049] Please understand here that the modified 3D NeRF images from text are generated in real time and are not drawn, processed in Photoshop®, discovered, or pre-planned. Instead, they are generated based on other text-based image generation.

[0050] Therefore, this principle provides a technique for generating consistent 3D heads from text in minutes, e.g., less than 2 minutes or less than 1 minute. By having a base model at the start, a new NeRF is generated for text of a specific domain. Using hashtable encoding of the NeRF facilitates rapid editing (-1 minute).

[0051] Here, refer to Figures 14-18, which illustrate the use of base models other than heads. For example, a base model of an in-game item such as a hat or weapon may be used to produce a hyper-personalized game item.

[0052] Starting from block 1400 in Figure 14, player data for a specific individual player is retrieved. To generate a single type of in-game item, the data is preferably non-confidential player data, such as previously played first-party games.

[0053] Moving to block 1402, text is created based on the data from block 1100. For example, if player data indicates that the player is a fan of game X, the system may create a text prompt containing or related to "game X" to be used in block 1404, consistent with this principle, and present it on the player's display in block 1406, and, if necessary, generate images of equipment such as different in-game masks that can be presented on the player's character during gameplay in block 1408, for example, based on the player's consent. The images may include materials such as texture data.

[0054] The text generator can be quite primitive. For example, the final text prompt 1500 shown in Figure 15 (presented on the screen along with the subsequent image 1600 in Figure 16) could be "Kabuki mask" and could include "based on [character Y] from [game X]". A typical prompt could be "[style] mask based on [character] from [game] ([release year])".

[0055] Figures 17 and 18 further illustrate. The user interface (UI) 1700 may be presented on the player's display 1702, prompting the player (1704) whether they wish to customize equipment for their character (PC), in this case a mask. One or more selectors 1706 may be presented, allowing the player to accept or decline. If the player accepts, text input may be automatically generated in accordance with the above principle to produce an image of the equipment, which may be presented to the player for viewing in 1800 of Figure 18.

[0056] Figure 18 shows that selector 1802 may be presented to a player to give them specific options, including the option to purchase custom-made equipment, and that the same equipment will not be presented to other players. A player may also wish to share equipment with other players.

[0057] The tools and techniques described herein may be provided on an end-user game computing device, such as a computer game console, so that an end-user game player can create and / or modify game objects in the game (i.e., as part of playing a computer game) using the tools described herein.

[0058] While this specification illustrates and describes specific embodiments, it should be understood that the subject matter covered by the present invention is limited only by the claims.

Claims

1. A method performed by a computer, Using the first model, a base three-dimensional (3D) asset is generated based on multiple two-dimensional (2D) images. Using the second model, receive text input indicating modifications to the base 3D asset, Using the first model described above, a modified 3D asset is generated by at least modifying the base 3D asset according to the text input. Using the second model described above, the text input and the modified 3D asset are compared to determine a similarity index. A method comprising converting the modified 3D asset into a polygon mesh based on the similarity index.

2. The method according to claim 1, wherein the base 3D asset is a neural luminance field (NeRF) encoded with a multi-resolution hash table.

3. The method according to claim 1, wherein the similarity index is determined using a contrasting language-image pre-trained (CLIP) model.

4. The method according to claim 1, further comprising repeatedly modifying the base 3D asset until the similarity index reaches a predetermined loss threshold.

5. The method according to claim 1, wherein the similarity index is calculated using the cosine similarity between the embedded image of the modified 3D asset and the embedded text of the text input.

6. The method according to claim 1, wherein the text input is derived from user information, and the user information includes at least the title of a computer simulation.

7. The method according to claim 1, wherein generating the modified 3D asset comprises rendering the modified 3D asset from multiple camera angles and applying enhancements including at least one of color, rotation, and affine transformation.

8. The method according to claim 1, further comprising presenting a user interface that prompts the user to select which base 3D asset to equip for modification, and generating the text input in accordance with the selection.

9. The method according to claim 1, further comprising storing the modified 3D assets in a library for later retrieval and modification.

10. The method according to claim 1, further comprising presenting the modified 3D asset on a display.

11. Memory configured to store computer-executable instructions, The system includes a processor that accesses the memory, and the processor executes the computer-executable instructions. Using the first model, a base three-dimensional (3D) asset is generated based on multiple two-dimensional (2D) images. Using the second model, receive text input indicating modifications to the base 3D asset, Using the first model described above, a modified 3D asset is generated by at least modifying the base 3D asset according to the text input. Using the second model described above, the text input and the modified 3D asset are compared to determine a similarity index. A system configured to convert the modified 3D asset into a polygon mesh based on the aforementioned similarity index.

12. The system according to claim 11, wherein the base 3D asset is a neural luminance field (NeRF) encoded with a multi-resolution hash table.

13. The system according to claim 11, wherein the similarity index is determined using a contrasting language-image pre-trained (CLIP) model.

14. The system according to claim 11, wherein the processor is further configured to iteratively modify the base 3D asset until the similarity index reaches a predetermined loss threshold.

15. The system according to claim 11, wherein the similarity index is calculated using the cosine similarity between the embedded image of the modified 3D asset and the embedded text of the text input.

16. One or more non-temporary computer-readable recording media for storing computer-readable instructions, wherein, when the instructions are executed by one or more processors, the system... Using the first model, a base three-dimensional (3D) asset is generated based on multiple two-dimensional (2D) images. Using the second model, receive text input indicating modifications to the base 3D asset, Using the first model described above, a modified 3D asset is generated by at least modifying the base 3D asset according to the text input. Using the second model described above, the text input and the modified 3D asset are compared to determine a similarity index. A non-temporary computer-readable recording medium that causes an operation to be performed, including converting the modified 3D asset into a polygon mesh based on the similarity index.

17. The non-temporary computer-readable recording medium according to claim 16, wherein the base 3D asset is a neural luminance field (NeRF) encoded with a multi-resolution hash table.

18. The non-temporary computer-readable recording medium according to claim 16, wherein the similarity index is determined using a contrasting language-image pre-trained (CLIP) model.

19. The operation further comprises repeatedly modifying the base 3D asset until the similarity index reaches a predetermined loss threshold, according to claim 16, a non-temporary computer-readable recording medium.

20. The non-temporary computer-readable recording medium according to claim 16, wherein the similarity index is calculated using the cosine similarity between the embedded image of the modified 3D asset and the embedded text of the text input.

21. The non-temporary computer-readable recording medium according to claim 16, wherein the text input is derived from user information, and the user information includes at least the title of a computer simulation.