Hyper-personalized game items

The combination of NeRF and CLIP model facilitates quick and intuitive creation of personalized characters and equipment in video games by generating hyper-personalized in-game items from text input, addressing the inefficiencies of traditional methods.

JP7793110B2Active Publication Date: 2025-12-26SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025519645
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-05
Filing Date
2023-10-02
Publication Date
2025-12-26
Estimated Expiration
2043-10-02

AI Technical Summary

Technical Problem

Creating characters and equipment for computer simulations, such as video games, is time-consuming and inefficient.

Method used

Utilizing a neural radiance field (NeRF) and a contrastive language-image pre-training (CLIP) model to generate hyper-personalized in-game items from text input, converting the NeRF into a polygon mesh for real-time character and equipment creation.

Benefits of technology

Enables rapid and intuitive generation of personalized characters and equipment in computer simulations, reducing the time required to less than two minutes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007793110000001
    Figure 0007793110000001
  • Figure 0007793110000002
    Figure 0007793110000002
  • Figure 0007793110000003
    Figure 0007793110000003
Patent Text Reader

Abstract

The two-dimensional image is converted (302) into a 3D neural luminance field (NeRF) and modified (402) based on player-personalized text to resemble the character's outfit (1700) that the text requests. The model scores (404) how well the lines of text match the image to produce the final 3D NeRF, which is converted (408) into a polygon mesh that can be imported into a computer simulation, such as a computer game.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates generally to rapidly generating hyper-personalized game items. [Background technology]

[0002] As understood herein, creating characters, such as non-player characters (NPCs), and their equipment for computer simulations, such as computer games, can be time consuming. Summary of the Invention

[0003] As will be further understood herein, it is desirable to enable game developers to create characters and equipment for video games in an easy, fast, and intuitive manner.

[0004] Thus, the device includes at least one computer storage that is not a transient signal and includes at least one computer storage that includes instructions executable by at least one processor to generate a base mesh or neural radiance field (NeRF) from a plurality of images. Note that while text-to-2D image techniques (e.g., stable diffusion) or text-to-mesh techniques may be used directly, in all embodiments, the technique is applied to the real-time creation of hyper-personalized in-game items.

[0005] In the context of a NeRF, the instructions are executable to generate a modified NeRF from a base NeRF using text input to a contrastive language-image pre-training (CLIP) model, and to convert the modified NeRF into a polygon mesh representing a virtual character for presenting the rig in at least one computer simulation.

[0006] In some examples, the CLIP model evaluates the match of an image to text.

[0007] The text is preferably derived from player information with which the players have commonality, such as the title of at least one computer simulation. The text may describe equipment for a game character, such as a mask. The instructions may be executable to generate text from a starting phrase using learned subsequent phrases.

[0008] In another aspect, the device includes at least one processor programmed with instructions to receive a text description of equipment personalized to player data, the instructions being executable to generate a virtual three-dimensional (3D) equipment based at least in part on the text description in less than two minutes after receiving the text description, and present the virtual equipment on a display.

[0009] In another aspect, a method includes receiving text based on data related to a player of a computer simulation and generating a neural luminance field based on the text starting from a base model.

[0010] The details of the present application, both as to its structure and operation, can best be understood in reference to the accompanying drawings, in which like reference numerals refer to like parts, and in which: [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram of an exemplary system in accordance with the present principles; [Figure 2] 1 illustrates exemplary initialization logic in exemplary flow chart form. [Figure 3] We demonstrate the creation of a 3D neural luminance field (NERF) representing the human head from a 2D image. [Figure 4] 1 illustrates exemplary text-based NERF customization logic in exemplary flowchart form. [Figure 4A] 1 is another representation of exemplary logic. [Figure 5]The first descriptive phrase is used to indicate a customized NeRF. [Figure 6] The first descriptive phrase is used to indicate a customized NeRF. [Figure 7] The first descriptive phrase is used to indicate a customized NeRF. [Figure 8] The first descriptive phrase is used to indicate a customized NeRF. [Figure 9] A second descriptive phrase is used to indicate a customized NeRF. [Figure 10] A third descriptive phrase is used to indicate a customized NeRF. [Figure 11] We present a series of image generation steps to learn how to customize generic NeRF. [Figure 12] 1 shows a data structure for generating text. [Figure 13] 10 illustrates additional exemplary text-based NeRF customization logic in exemplary flow chart form. [Figure 14] 1 illustrates exemplary text-based equipment customization logic in exemplary flow chart form. [Figure 15] An example screenshot is shown corresponding to Figure 14. [Figure 16] 14 shows an example of player-customized gaming gear consistent with FIG. [Figure 17] 10 shows example screenshots that provide customized gear for a specific player. [Figure 18] 10 shows example screenshots that provide customized gear for a specific player. DETAILED DESCRIPTION OF THE INVENTION

[0012] The present disclosure generally relates to computer ecosystems, including aspects of consumer electronics (CE) device networks, such as, but not limited to, computer gaming networks. The systems herein may include server and client components that may be connected via a network such that data may be exchanged between the client and server components. The client components may include one or more computing devices, including game consoles such as Sony PlayStation®, or game consoles from Microsoft, Nintendo, or other manufacturers; extended reality (XR) headsets such as virtual reality (VR) headsets, augmented reality (AR) headsets; portable televisions (e.g., smart TVs, Internet-enabled televisions); portable computers such as laptops and tablet computers; and smartphones and other mobile devices, including additional examples described below. These client devices may operate in a variety of operating environments. For example, some of the client computers may utilize, by way of example, the Linux® operating system, a Microsoft operating system, or a Unix® operating system, or an operating system from Apple or Google, or the Berkeley Software Distribution or Berkeley Standard Distribution of BSD-based operating systems. These operating environments may be used to run one or more browsing programs, such as browsers from Microsoft, Google, or Mozilla, or other browser programs capable of accessing websites hosted by the Internet servers described below. An operating environment according to present principles may also be used to run one or more computer game programs.

[0013] Servers and / or gateways may be used, which may include one or more processors that execute instructions that configure the server to send and receive data over a network such as the Internet. Alternatively, clients and servers may be connected via a local intranet or a virtual private network. The server or controller may be instantiated by a game console such as a Sony PlayStation®, a personal computer, or the like.

[0014] Information may be exchanged between the client and the server over a network. For this purpose and for security, the server and / or client may include firewalls, load balancers, temporary storage, and proxies, as well as other network infrastructure for reliability and security. One or more servers may form an apparatus that implements a method for providing a secure community, such as an online social website or a gamer network, to network members.

[0015] The processor may be a single-chip processor or a multi-chip processor capable of performing logic through various lines, such as address lines, data lines, and control lines, as well as registers and shift registers. A processor including a digital signal processor (DSP) may be an embodiment of a circuit.

[0016] Components included in one embodiment may be used in other embodiments in any suitable combination. For example, any of the various components described and / or illustrated herein may be combined, substituted, or excluded from other embodiments.

[0017] A "system having at least one of A, B, and C" (and similarly, "a system having at least one of A, B, or C" and "a system having at least one of A, B, and C") includes systems having A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, and C together.

[0018] Referring now to FIG. 1 , an exemplary system 10 is shown, which may include one or more of the exemplary devices described above and may include one or more of the exemplary devices further described below in accordance with the present principles. The first exemplary device included in system 10 is a consumer electronics (CE) device such as, but not limited to, an audio-video device (AVD) 12, which may be, for example, a theater display system, which may be a projector-based or Internet-enabled television with a TV tuner (equivalently, a set-top box that controls the TV). AVD 12 may alternatively be a computer-controlled Internet-enabled (“smart”) phone, a tablet computer, a notebook computer, a head-mounted device (HIVID) and / or a headset such as smart glasses or a VR headset, another wearable computing device, a computer-controlled Internet-enabled music player, computer-controlled Internet-enabled headphones, a computer-controlled Internet-enabled implantable device such as an implantable skin device, etc. In any event, it should be understood that AVD 12 is configured to implement the present principles (e.g., to communicate with other CE devices to implement the present principles, execute the logic described herein, and perform any other functions and / or operations described herein).

[0019] Accordingly, to implement such principles, AVD 12 may be established by some or all of the components shown. For example, AVD 12 may include one or more touch-enabled displays 14, which may be implemented by high-resolution or ultra-high-resolution "4K" or higher flat screens. Touch-enabled display(s) 14 may include, for example, a capacitive or resistive touch-sensing layer with a grid of electrodes for touch sensing consistent with the present principles.

[0020] The AVD 12 may also include one or more speakers 16 for outputting audio in accordance with the present principles and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands to the AVD 12 to control it. The exemplary AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22, such as the Internet, a WAN, or a LAN, under the control of one or more processors 24. Accordingly, the interface 20 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It should be understood that the processor 24 controls the AVD 12 to implement the present principles, including other elements of the AVD 12 described herein, such as controlling the display 14 to present images and receive input therefrom. Furthermore, it should be noted that the network interface 20 may be a wired or wireless modem or router, or other suitable interface, such as a wireless telephone transceiver or a Wi-Fi transceiver as described above.

[0021] In addition to the above, AVD 12 may also include one or more input and / or output ports 26, such as a High-Definition Multimedia Interface (HDMI) port or a Universal Serial Bus (USB) port, for physically connecting to another CE device, and / or a headphone port for connecting headphones to AVD 12 to present audio from AVD 12 to a user via headphones. For example, input port 26 may be wired or wirelessly connected to a cable or satellite source 26a of audio-video content. Thus, source 26a may be a separate or integrated set-top box or satellite receiver. Alternatively, source 26a may be a game console or disc player containing content. When implemented as a game console, source 26a may include some or all of the components described below in connection with CE device 48.

[0022] AVD 12 may further include one or more computer memory / computer-readable storage media 28, such as non-transitory disk-based or solid-state storage, possibly embodied in the AVD's chassis as a standalone device, or as a personal video recording device (PVR) or video disc player either internal or external to the AVD's chassis for playing AV programs, or as a removable storage medium or server as described below. In some embodiments, AVD 12 also includes a location or position receiver, such as, but not limited to, a cellular receiver, a GPS receiver, and / or an altimeter 30, configured to receive geographic location information from a satellite or cellular tower and provide that information to processor 24 and / or to determine the altitude at which AVD 12 is located in cooperation with processor 24.

[0023] Continuing with the description of AVD 12, in some embodiments, AVD 12 may include one or more cameras 32, which may be an infrared camera, a digital camera such as a webcam, an IR sensor, an event-based sensor, and / or a camera integrated into AVD 12 and controllable by processor 24 to collect pictures / images and / or video in accordance with the present principles. AVD 12 may also include a Bluetooth transceiver 34 and other NFC elements 36 for communicating with other devices using Bluetooth and / or near field communication (NFC) technology, respectively. An exemplary NFC element may be a radio frequency identification (RFID) element.

[0024] Furthermore, AVD 12 may include one or more auxiliary sensors 38 that provide input to processor 24. For example, one or more of auxiliary sensors 38 may include one or more pressure sensors that form a layer of touch-enabled display 14 itself, and may be, without limitation, piezoelectric pressure sensors, capacitive pressure sensors, piezoresistive strain gauges, optical pressure sensors, electromagnetic pressure sensors, etc. Other example sensors include pressure sensors, motion sensors such as accelerometers, gyroscopes, cyclometers, or magnetic sensors, infrared (IR) sensors, light sensors, speed and / or cadence sensors, event-based sensors, and gesture sensors (e.g., for sensing gesture commands). Thus, sensors 38 may be implemented by one or more motion sensors such as individual accelerometers, gyroscopes, and magnetometers, and / or an inertial measurement unit (IMU), which typically includes a combination of accelerometers, gyroscopes, and magnetometers, or an event-based sensor such as an event detection sensor (EDS), to determine the position and orientation of AVD 12 in three dimensions. An EDS consistent with the present disclosure provides an output indicative of a change in light intensity sensed by at least one pixel of the light-sensing array. For example, if the pixel senses a decrease in light, the output of the EDS may be −1; if it senses an increase, the output of the EDS may be +1. No change in light intensity below a certain threshold may be indicated by an output binary signal of 0.

[0025] The AVD 12 may also include an over-the-air TV broadcast port 40 for receiving an OTA TV broadcast, which provides input to the processor 24. In addition to the above, it should be noted that the AVD 12 may also include an infrared (IR) transmitter and / or an IR receiver and / or an IR transceiver 42, such as an infrared data association (IRDA) device. A battery (not shown) may be provided to power the AVD 12, which may be a kinetic energy harvester that converts kinetic energy into power to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field-programmable gate array 46 may also be included. One or more haptic / vibration generators 47 may be provided to generate haptic signals that can be sensed by a person holding or in contact with the device. Thus, the haptic generator 47 may vibrate all or part of the AVD 12 using an electric motor connected to biased and / or unbalanced weights via the motor's rotatable shaft, so that the shaft may rotate under the control of the motor (which may be controlled by a processor such as processor 24), creating simulations of vibrations of various frequencies and / or amplitudes, and forces in various directions.

[0026] A light source such as a projector, such as an infrared (IR) projector, may also be included.

[0027] In addition to the AVD 12, the system 10 may include one or more other CE device types. In one example, the first CE device 48 may be a computer game console that can be used to transmit computer game audio and video to the AVD 12 via commands sent directly to the AVD 12 and / or via a server, as described below, while the second CE device 50 may include similar components to the first CE device 48. In the illustrated example, the second CE device 50 may be configured as a computer game controller operated by a player or a head-mounted display (HMD) worn by a player. The HMD may include a transparent or opaque head-up display that presents AR / MR content or VR content (or, more broadly, extended reality (XR) content), respectively. The HMD may also be configured as a glasses-type display or a large VR-type display sold by a computer game console manufacturer.

[0028] In the illustrated example, only two CE devices are shown, but it should be understood that a fewer or greater number of devices may be used. The devices herein may implement some or all of the components shown for AVD 12. Any of the components shown in the following figures may incorporate some or all of the components shown in the AVD 12 example.

[0029] Referring now to the aforementioned at least one server 52, it includes at least one server processor 54, at least one tangible computer-readable storage medium 56, such as disk-based or solid-state storage, and at least one network interface 58 that, under the control of the server processor 54, enables communication with the other devices shown over the network 22, and indeed may facilitate communication between the server and client devices in accordance with the present principles. It should be noted that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface, such as, for example, a wireless telephone transceiver.

[0030] Thus, in some embodiments, server 52 may be an entire Internet server or server "farm" and may include or perform "cloud" functionality such that, in an exemplary embodiment, for example, for a network gaming application, devices of system 10 may access the "cloud" environment via server 52. Alternatively, server 52 may be implemented by one or more game consoles or other computers in the same room or nearby as the other devices shown.

[0031] The components shown in the following figures may include some or all of the components shown herein. Any user interfaces (UIs) described herein may be aggregated and / or expanded, and UI elements may be mixed and matched between UIs.

[0032] The present principles may utilize various machine learning models, including deep learning models. Machine learning models consistent with the present principles may use various algorithms trained using methods including supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature representation learning, self-learning, and other forms of learning. Examples of such algorithms that may be implemented by computer circuitry include one or more neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and a type of RNN known as a long short-term memory (LSTM) network. Support vector machines (SVMs) and Bayesian networks may also be considered examples of machine learning models. In addition to the types of networks described above, the models herein may be implemented by classifiers.

[0033] Thus, as understood herein, performing machine learning may include accessing training data to train a model so that the model can further process the data and make inferences. Thus, an artificial neural network / artificial intelligence model trained through machine learning may include an input layer, an output layer, and multiple hidden layers in between that are configured and weighted to make inferences regarding appropriate outputs.

[0034] Reference is now made to FIGS. 2 and 3. Initial logic, which may be executed by any processor or processors herein, begins in block 200 by receiving a set of head images, which may be 2D images of the same base head (labeled 300 in FIG. 3), from different viewpoints if generated by one or more artists. From these images, a 3D base head (302 in FIG. 3) is generated in block 202. The 3D base head 302 may be a 3D neural intensity field (NeRF), which may be thought of as a 3D volume stored in a machine learning (ML) model. In a particular example, the base head 302 may be a NeRF that can be quickly trained in less than two minutes, and in some examples, less than one minute, after receiving the following text description as input to generate a modified NeRF, described further below. An exemplary NeRF for this purpose is a multi-resolution hash table-encoded NeRF, which stores the locations of a sparse multi-resolution 3D grid to speed up training and rendering of the NeRF.

[0035] Note that instead of an image, in some embodiments, the base head 302 of FIG. 3 may be generated using a head scan of a real-world head.

[0036] Referring now to Figure 4, once the base head 302 has been generated, a text description of the modified base head may be received in block 400 of the desired head. The text may be entered from a text input device such as a keyboard, and / or from a speech-to-text conversion of a spoken description, and / or from an initial starting phrase followed by additional learned descriptive phrases as further described herein. Alternatively, the text may be generated in a Madlib style using more primitive algorithmic techniques so that survey responses can be inserted into the text describing the target head.

[0037] Moving to block 402, the text is input to a NeRF retrieval engine, which may include a fully connected (non-convolutional) deep network. The input to the retrieval engine may be a single continuous 5D input including values ​​representing three spatial dimensions and two visual dimensions, while the output of the engine may include volumetric density and view-dependent radiance at spatial locations represented by associated 3D values.

[0038] Proceeding to block 404, the output of the NeRF modification engine, which can be thought of as a modified NeRF, is evaluated to determine how closely the output matches the input text description. In one example, this evaluation may be performed by a contrastive language-image pre-training (CLIP) model, which evaluates the match of an image to an input text string. CLIP is an open-source model that scores how well a line of text matches an image. A CLIP model can be trained on image-text pairs, such as captioned photos from the internet, using cosine similarity to score the goodness of the match. Thus, modifications to the base NeRF are essentially derived using a loss function that depends on the text input to the CLIP model.

[0039] Decision diamond 406 indicates that if, after the current iteration of modification, it is determined that the loss has not reached a threshold (target small) loss, then in block 408, the current iteration of the modified NeRF is modified again and the logic loops back to block 408. However, if the loss is determined to be small enough, the logic proceeds to block 410 to output the final modified NeRF. Note that instead of a loss threshold, a set number of loops may be performed, e.g., 100 loops. Blocks 402-408 may use gradient descent techniques.

[0040] In some embodiments, the logic may then move to block 412 and convert the modified NeRF into a polygon mesh that can be imported into a computer simulation, such as a computer game, in block 414 and presented on a display during gameplay in block 416. Block 412 may also generate materials for the mesh, such as albedo, roughness, and normal textures. The modified NeRF can be converted into a polygon mesh that has the same topology as the modified NeRF. This may be done using the marching cubes method. Alternatively, the modified NeRF can be converted into a polygon mesh by essentially shrink-wrapping the mesh by placing a mesh over the modified NeRF and recording the portions of the mesh that touch the surface of the modified NeRF when emulating a vacuum evacuation of the mesh.

[0041] 4A is an alternative diagram. Beginning in block 420, a NeRF is generated from parameters that determine what the head will look like and may include values ​​that are adjusted so that the final NeRF looks similar to the input image. Moving to block 422, the NeRF may be rendered from multiple camera angles, as well as color, rotation, and extensions such as affine applied in block 424.

[0042] Proceeding to block 426, the embedding of the CLIP image can be obtained and used in block 428, and the current iteration of NeRF can be compared to the embedding of the text image to obtain a loss value. In block 430, the loss value is added to other conditions, such as how "cloudy" the NeRF is, by obtaining the sum of the variance of the depth of the image being rendered. If a parameter modification is desired in block 420, the sum of the losses is backpropagated using gradient descent in block 432.

[0043] Figures 5-8 show four perspective views of the modified NeRF 500 generated from the base NeRF 302 in Figure 3 using the input text "Photorealistic head of a burly old man." The modified NeRF 500 was generated in less than two minutes after inputting the text into the model.

[0044] 9 and 10 show modified NeRF900, NeRF1000, respectively, generated from base NeRF302 using the input texts "Photorealistic head of a sea monster" and "Photorealistic head of a devil", respectively.

[0045] FIG. 11 illustrates that the above-described ML model, particularly for use in meshing methods, can be trained with a chain of causality from the initial parameters (variables) controlling the vertices of an object to the pixels rendered on the screen. Starting at a vertex 1100, the causality chain can proceed to a primitive 1102 based on the vertex 1100, and then further to a fragment 1104 where the primitive is at least partially filled, for example, with a texture. From the fragment 1104, a processed fragment 1106 can be generated, to which color can be added, and the processed fragment 1106 is merged into a pixel output 1108 presented on the screen. Arrow 1110 indicates that this chain can be traced back to the original vertex 1100 to see how particular values ​​in the chain affect the final pixel image 1108, which can be used to train the ML model.

[0046] Note that if a head mesh is used instead of a luminance field, the ML model can perform differentiable rendering, where parameters can be changed to see in real time how the variables affect the output image. The ML model uses a loss function to minimize the loss as new faces begin to match the targets defined in the text.

[0047] As described above, some or all of the input text describing the desired modified NeRF can be generated from the starting phrase using learned subsequent phrases. Figure 12 shows a prompt template 1200 that can be presented on a developer's computer screen with multiple phrases that can be entered into a player / developer's response. Additional text 1202 can be presented for each prompt response. The additional text 1202 can be generated by an ML model trained to complete sentences based on the input phrase.

[0048] Figure 13 illustrates an additional technique. Starting at block 1300, text is received that indicates a character's biography and the gear the character is starting the game with. A text-to-image generator 1302 receives the text and generates an image consistent with the principles herein, which is imported into the computer game at block 1304 and presented while playing the game at block 1306. Note that block 1302 may alternatively use a text-to-3D mesh generator (such as that of NeRF) or an image such as a decal on the player's gear, but the more predominant use case is to generate actual 3D assets in the form of meshes / materials.

[0049] It should be understood here that the modified 3D NeRF images from text are generated in real time and are not drawn, Photoshopped, or otherwise discovered or preconceived. Instead, they are generated based on other text-based image generation.

[0050] Thus, the present principles provide a technique for generating consistent 3D heads from text in a few minutes, e.g., under 2 minutes or even under 1 minute. By having a starting base model, a new NeRF is generated with text from a specific domain. The use of a hash table encoding of the NeRF facilitates rapid editing (-1 minute).

[0051] 14-18, which illustrate the use of base models other than heads. For example, base models of in-game items such as hats or weapons may be used to produce hyper-personalized game items.

[0052] Beginning at block 1400 of Figure 14, player data for a particular individual player is obtained. To generate a type of in-game item, the data is preferably non-confidential player data, such as previously played first-party games.

[0053] Moving to block 1402, text is created based on the data from block 1100. For example, if the player data indicates that the player is a fan of Game X, the system creates a text prompt containing or related to "Game X" for use in block 1404 consistent with the present principles, which can be presented on the player's display in block 1406, and optionally, based on the player's consent, generates images of equipment such as different in-game masks that can be presented on the player's character during game play in block 1408. The images may include materials such as texture data.

[0054] The text generator can be fairly primitive. For example, the final text prompt 1500 shown in Figure 15 (presented on the screen along with the subsequent image 1600 of Figure 16) is "Kabuki mask" and can include "based on [character Y] from [game X]." A general prompt could be "Mask of [style] based on [character] from [game] ([release year])."

[0055] 17 and 18 are further shown. A user interface (UI) 1700 may be presented on a player's display 1702 prompting the player 1704 whether they wish to customize equipment, in this case a mask, for the player's character (PC). One or more selectors 1706 may be presented, allowing the player to accept or decline. If the player accepts, text input may be automatically generated consistent with the principles described above to generate an image of the equipment, which may be presented at 1800 in FIG. 18 for the player to view.

[0056] 18 further illustrates that a selector 1802 may be presented to give a player certain options, including the option to purchase custom gear, and that other players will not be presented with the same gear. A player may also wish to share gear with other players.

[0057] The above tools and techniques may be provided in an end-user game computing device, such as a computer game console, so that end-user game players can create and / or modify game objects within the game (i.e., as part of playing a computer game) and with each other using the tools described herein.

[0058] Although particular embodiments are shown and described in detail herein, it should be understood that the subject matter encompassed by the present invention is limited only by the claims.

Claims

1. at least one computer storage device that is not a transitory signal; by at least one processor, generating a base three-dimensional (3D) asset using a neural luminance field (NeRF) from a plurality of two-dimensional (2D) images; generating a text input based on at least one of user information about a user of the base 3D asset or game information about a video game presenting the base 3D asset; generating a modified 3D asset from the base 3D asset using the text input to a contrastive language-image pre-training (CLIP) model; converting the modified 3D assets into a polygon mesh representing the equipment of a virtual character to represent the equipment of the virtual character in at least one computer simulation; a device including the at least one computer storage including instructions executable to cause the modified 3D asset to be presented on a display;

2. The device of claim 1 , wherein the CLIP model evaluates the match of an image to the text input.

3. The device of claim 2 , wherein the text input is derived from player information.

4. The device of claim 3 , wherein the player information includes the title of at least one computer simulation.

5. The device of claim 1 , wherein the text input describes a character's equipment.

6. The device of claim 5 , wherein the character's equipment includes a mask.

7. The instruction: The device of claim 1 , wherein the device is executable to generate the text input from a starting phrase using learned subsequent phrases.

8. The device of claim 1 comprising the at least one processor.

9. The device described in claim 1, wherein the text input is based on dialogue information regarding a user's response to a question regarding the base 3D asset.

10. A method for generating a base three-dimensional (3D) asset from a plurality of two-dimensional (2D) images using a neural luminance field (NeRF), generating a text input based on at least one of user information about a user of the base 3D asset or game information about a video game presenting the base 3D asset; generating a modified 3D asset from the base 3D asset using the text input to a contrastive language-image pre-training (CLIP) model; converting the modified 3D assets into a polygon mesh representing the equipment of a virtual character to represent the equipment of the virtual character in at least one computer simulation; The method causes the modified 3D asset to be presented on a display.

11. The method of claim 10, wherein the CLIP model evaluates the match of an image to the text input.

12. The method of claim 11, wherein the text input is derived from player information.

13. The method of claim 12, wherein the player information includes the title of at least one computer simulation.

14. The method of claim 10, wherein the text input describes a character's equipment.

15. The method of claim 14, wherein the character's equipment includes a mask.

16. The instruction: The method of claim 10 , wherein the method is executable to generate the text input from a starting phrase using learned subsequent phrases.

17. The method of claim 10, comprising the at least one processor.

18. The method described in claim 10, wherein the text input is based on dialogue information regarding a user's response to a question about the base 3D asset.

19. A non-transitory computer-readable medium storing a plurality of instructions, the plurality of instructions, when executed by one or more processors of a computing device, causing the one or more processors to: generating a base three-dimensional (3D) asset using a neural luminance field (NeRF) from a plurality of two-dimensional (2D) images; generating a text input based on at least one of user information about a user of the base 3D asset or game information about a video game presenting the base 3D asset; generating a modified 3D asset from the base 3D asset using the text input to a contrastive language-image pre-training (CLIP) model; converting the modified 3D assets into a polygon mesh representing the equipment of a virtual character to represent the equipment of the virtual character in at least one computer simulation; A non-transitory computer-readable medium operable to cause the modified 3D asset to be presented on a display.

20. The non-transitory computer-readable medium of claim 19, wherein the operations further include the CLIP model evaluating a match of an image to the text input.

Citation Information

Patent Citations

  • Method and system for generating 3D human reconstructions

    JP2021536613A

  • Pixel-aligned volumetric avatars

    US20220198731A1

  • Real-time system for generating 4d spatio-temporal model of a real world environment

    WO2021099778A1