Fast generation of 3D heads using natural language fields
By using multiple images to generate neural fields and contrast language-image pre-trained models, using text input to generate modified NeRFs and convert them into polygon mesh, solving the problem of time-consuming and professional knowledge in the prior art creation of roles and equipment, achieving fast and simple virtual 3D header generation.
Patent Information
- Application Number
- CN202380071025.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-14
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is time-consuming and requires expertise in creating characters and equipment for computer games, making it difficult to achieve a fast, simple and intuitive process.
By using multiple image generation neural field (NeRF) and contrast language-image pre-training (CLIP) models, the modified NeRF is generated using text input and converted into a polygon mesh representing the virtual human head to present the virtual human head in a computer simulation.
It realizes the generation of virtual coherent 3D headers within less than two minutes after receiving the text description, simplifying the process of game developers and end users when creating characters and equipment.
Smart Images

Figure CN120019384A_ABST
Abstract
Description
Technical Field
[0001] The present application as a whole relates to the rapid generation of 3D heads using natural language. Background Art
[0002] As understood herein, creating characters, such as non-player characters (NPCs), and their equipment for a computer simulation, such as a computer game, can be time consuming and require specialized knowledge. Summary of the invention
[0003] As further understood herein, it is desirable to enable game developers to create characters and equipment for video games in a simple, fast and intuitive manner. It is further desirable to enable end users to create characters using text or their voice (or similar text-based interfaces, such as questionnaires).
[0004] Thus, an apparatus includes at least one computer memory that is not a transient signal and includes instructions that are executable by at least one processor to generate neural fields, such as a base three-dimensional (3D) neural radiance field (NeRF) (including NeRF encoded in a multi-resolution hash table) from a plurality of images. The instructions are executable to use text input of a contrastive language-image pre-training (CLIP) model to generate a modified NeRF from the base NeRF, and convert the modified NeRF into a polygonal mesh representing a virtual human head for rendering the virtual human head in at least one computer simulation. It is noted that the base head can be derived from a 3D model or image of a real human head.
[0005] In some examples, the CLIP model rates the match of an image to text, and can be trained on image-text pairs using cosine similarity to rate the goodness of the match. The CLIP model rates text-image similarity, which is used to rate how well the text matches the rendering of the image of the head.
[0006] In some embodiments, the instructions may be executable to use a machine learning (ML) model to minimize the loss indication in matching text. The instructions may be executable to train the ML model based on a causal chain from initial image parameters controlling the vertices of an object to the pixels rendered by the object on the screen.
[0007] In an example, the ML model includes at least one fully connected (non-convolutional) deep network.
[0008] In some implementations, the input to the ML model can include values representing three spatial dimensions and two viewing dimensions, and the output of the ML model can include volume density and viewing-dependent emitted radiation.
[0009] If desired, the instructions may be executable to generate text from the starting phrase using the learned subsequent phrase.
[0010] In another aspect, an apparatus includes at least one processor programmed with instructions to receive a text description of a person and, based at least in part on the text description, generate a virtual coherent three-dimensional (3D) head in less than two minutes after receiving the text description. The instructions are executable to present the virtual coherent 3D head on a display.
[0011] In another aspect, a method includes receiving text; and generating a neural radiance field based on the text starting from a base model.
[0012] The details of the present application, both as to its structure and operation, may best be understood with reference to the accompanying drawings, in which like reference numerals refer to like parts, and in which: BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a block diagram of an example system according to the present principles;
[0014] Figure 2 Example initial logic is illustrated in an example flow chart format;
[0015] Figure 3 Demonstrates the creation of a 3D Neural Radiance Field (NeRF) representing a human head from a 2D image;
[0016] Figure 4 An example text-based NeRF customization logic is illustrated in an example flow chart format;
[0017] Figure 4A is another representation of the example logic;
[0018] Figures 5 to 8 A custom NeRF using a first descriptive phrase is illustrated;
[0019] Fig. 9 A custom NeRF using a second descriptive phrase is illustrated;
[0020] Fig.10 A custom NeRF using a third descriptive phrase is illustrated;
[0021] Fig.11 illustrates a sequence of image generation steps using differential rendering;
[0022] Fig.12 illustrates the data structures used to generate text;
[0023] Fig.13 Additional example text-based NeRF customization logic is illustrated in example flow chart format;
[0024] Fig.14 Example text-based equipment customization logic is illustrated in an example flow chart format;
[0025] Fig.15 Illustrated with Fig.14 consistent example screenshots;
[0026] Fig.16 Illustrated with Fig.14 Consistent sample player customized gaming gear; and
[0027] Fig.17 and Fig.18 An example screenshot illustrating providing custom gear for a specific player. DETAILED DESCRIPTION
[0028] The present disclosure as a whole relates to computer ecosystems including aspects of consumer electronics (CE) device networks such as, but not limited to, computer gaming networks. The systems herein may include server and client components that may be connected via a network such that data may be exchanged between the client and server components. The client components may include one or more computing devices, including gaming consoles (such as Sony or game consoles manufactured by Microsoft or Nintendo or other manufacturers), extended reality (XR) headsets (such as virtual reality (VR) headsets, augmented reality (AR) headsets), portable televisions (e.g., smart TVs, Internet-enabled televisions), portable computers (such as laptops and tablet computers), and other mobile devices including smart phones and additional examples discussed below. These client devices can operate in a variety of operating environments. For example, as examples, some of the client computers may use a Linux operating system, an operating system from Microsoft, or a Unix operating system, or an operating system produced by Apple or Google, or a Berkeley Software Distribution or Berkeley Standard Distribution (BSD) OS including descendants of BSD. These operating environments can be used to execute one or more browsing programs, such as browsers manufactured by Microsoft or Google or Mozilla or other browser programs that can access websites hosted by Internet servers discussed below. In addition, an operating environment according to the present principles can be used to execute one or more computer game programs.
[0029] A server and / or gateway may be used, which may include one or more processors executing instructions that configure the server to receive and send data over a network such as the Internet. Alternatively, the client and server may be connected via a local intranet or a virtual private network. The server or controller may be provided by, for example, Sony Personal computers and other game consoles to instantiate.
[0030] Information may be exchanged between the client and the server over the network. To this end, for security purposes, the server and / or the client may include firewalls, load balancers, temporary storage and proxies, as well as other network infrastructure for reliability and security. One or more servers may form a device that implements the method of providing a secure community such as an online social networking site or a gamer network to network members.
[0031] The processor may be a single chip or a multi-chip processor that may execute logic with the aid of various lines such as address lines, data lines, control lines, and registers and shift registers.A processor including a digital signal processor (DSP) may be an embodiment of the circuit.
[0032] The components included in one embodiment may be used in other embodiments in any appropriate combination. For example, any of the various components described herein and / or depicted in the accompanying drawings may be combined, interchanged or excluded from other embodiments.
[0033] “A system having at least one of A, B, and C” (also “a system having at least one of A, B, or C” and “a system having at least one of A, B, C”) includes systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B and C together.
[0034] Reference now Figure 1, an example system 10 is shown, which may include one or more of the example devices mentioned above and further described below according to the present principles. A first example device among the example devices included in the system 10 is a consumer electronics (CE) device, such as an audio video device (AVD) 12, such as but not limited to a projector-based theater display system, or an Internet-enabled TV with a TV tuner (equivalently a set-top box that controls the TV). Alternatively, the AVD 12 may also be a computerized Internet-enabled ("smart") phone, a tablet computer, a notebook computer, a head-mounted device (HMD) and / or a headset such as smart glasses or a VR headset, another wearable computerized device, a computerized Internet-enabled music player, a computerized Internet-enabled headset, a computerized Internet-enabled implantable device such as an implantable skin device, etc. In any case, it should be understood that the AVD 12 is configured to implement the present principles (e.g., communicate with other CE devices to implement the present principles, execute the logic described herein, and perform any other functions and / or operations described herein).
[0035] Thus, to implement such principles, AVD 12 may be built from some or all of the components shown. For example, AVD 12 may include one or more touch-enabled displays 14 that may be implemented as high-definition or ultra-high-definition "4K" or higher flat screens. Touch-enabled displays 14 may include, for example, a capacitive or resistive touch sensing layer with an electrode grid for touch sensing consistent with the present principles.
[0036] AVD 12 may also include one or more speakers 16 for outputting audio according to the present principles, and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands to AVD 12 to control AVD 12. Example AVD 12 may also include one or more network interfaces 20 for communicating through at least one network 22 (such as the Internet, WAN, LAN, etc.) under the control of one or more processors 24. Therefore, interface 20 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It should be understood that processor 24 controls AVD 12 to implement the present principles, including other elements of AVD 12 described herein, such as controlling display 14 to present images thereon and receiving input therefrom. In addition, it should be noted that network interface 20 may be a wired or wireless modem or router, or other suitable interfaces, such as a wireless telephone transceiver or a Wi-Fi transceiver as mentioned above, etc.
[0037] In addition to the foregoing, AVD 12 may also include one or more input and / or output ports 26, such as a high-definition multimedia interface (HDMI) port or a universal serial bus (USB) port for physically connecting to another CE device and / or a headphone port for connecting headphones to AVD 12 to present audio from AVD 12 to the user through headphones. For example, input port 26 may be connected to a cable or satellite source 26a of audio video content via wire or wireless. Therefore, source 26a may be a separate or integrated set-top box, or a satellite receiver. Or source 26a may be a game console or disk player containing content. When implemented as a game console, source 26a may include some or all of the components described below with respect to CE device 48.
[0038] AVD 12 may also include one or more computer memories / computer readable storage media 28 (such as disk-based or solid-state memory that is not a transient signal), which in some cases is implemented as a stand-alone device in the AVD chassis or a personal video recording device (PVR) or a video disk player (located inside or outside the AVD chassis) for playing back AV programs, or as a removable storage medium or a server described below. In addition, in some embodiments, AVD 12 may include a position receiver or a positioning receiver (such as, but not limited to, a cellular phone receiver, a GPS receiver, and / or an altimeter 30) that is configured to receive geographic location information from a satellite or cellular phone base station and provide the information to the processor 24 and / or determine the altitude at which AVD 12 is located in conjunction with the processor 24.
[0039] Continuing with the description of the AVD 12, in some embodiments, the AVD 12 may include one or more cameras 32, which may be thermal imaging cameras, digital cameras such as webcams, IR sensors, event-based sensors, and / or cameras integrated into the AVD 12 and controllable by the processor 24 to collect pictures / images and / or videos in accordance with the present principles. The AVD 12 may also include a Bluetooth and / or NFC technology for communicating with other devices, respectively. Transceiver 34 and other near field communication (NFC) elements 36. An example NFC element may be a radio frequency identification (RFID) element.
[0040] Further, the AVD 12 may include one or more auxiliary sensors 38 that provide input to the processor 24. For example, one or more of the auxiliary sensors 38 may include one or more pressure sensors that form a layer of the touch-enabled display 14 itself, and may be, but are not limited to, piezoelectric pressure sensors, capacitive pressure sensors, piezoresistive strain gauges, optical pressure sensors, electromagnetic pressure sensors, and the like. Other sensor examples include pressure sensors, motion sensors such as accelerometers, gyroscopes, registers, or magnetic sensors, infrared (IR) sensors, optical sensors, speed and / or cadence sensors, event-based sensors, gesture sensors (e.g., for sensing gesture commands). Thus, the sensor 38 may be implemented by one or more motion sensors (such as separate accelerometers, gyroscopes, and magnetometers) and / or an inertial measurement unit (IMU), which typically includes a combination of accelerometers, gyroscopes, and magnetometers to determine the location and orientation of the AVD 12 in three dimensions, or by an event-based sensor (such as an event detection sensor (ED)). An EDS consistent with the present disclosure provides an output indicating a change in light intensity sensed by at least one pixel of a light sensing array. For example, if the light sensed by the pixel is decreasing, the output of the EDS may be -1; if it is increasing, the output of the EDS may be + 1. An output binary signal of 0 may indicate no change in light intensity below a certain threshold.
[0041] The AVD 12 may also include an air TV broadcast port 40 for receiving OTA TV broadcasts that provide input to the processor 24. In addition to the foregoing, it is noted that the AVD 12 may also include an infrared (IR) transmitter and / or an IR receiver and / or an IR transceiver 42, such as an IR Data Association (IRDA) device. A battery (not shown) may be provided for powering the AVD 12, and a kinetic energy collector may also be provided that converts kinetic energy into electricity to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field programmable gate array 46 may also be included. One or more tactile / vibration generators 47 may be provided for generating tactile signals that can be sensed by a person holding the device or in contact with the device. Thus, the tactile generator 47 can vibrate all or part of the AVD 12 using an electric motor connected to an eccentric and / or unbalanced weight via a rotatable shaft of the motor so that the shaft can rotate under the control of the motor (which in turn can be controlled by a processor such as processor 24) to produce vibrations of various frequencies and / or amplitudes and force simulations in various directions.
[0042] A light source such as a projector, such as an infrared (IR) projector, may also be included.
[0043] In addition to the AVD 12, the system 10 may include one or more other CE device types. In one example, the first CE device 48 may be a computer game console that can be used to transmit computer game audio and video to the AVD 12 via commands transmitted directly to the AVD 12 and / or through a server described below, while the second CE device 50 may include components similar to the first CE device 48. In the example shown, the second CE device 50 may be configured as a computer game controller manipulated by a player or a head-mounted display (HMD) worn by a player. The HMD may include a transparent or non-transparent head-up display for presenting AR / MR content or VR content (more generally, extended reality (XR) content), respectively. The HMD may be configured as a glasses-type display or a larger VR-type display sold by a computer game device manufacturer.
[0044] In the example shown, only two CE devices are shown, and it should be understood that fewer or more devices may be used. The devices herein may implement some or all of the components shown for AVD 12. Any of the components shown in the following figures may be combined with some or all of the components shown in the context of AVD 12.
[0045] Reference is now made to the aforementioned at least one server 52, which includes at least one server processor 54, at least one tangible computer-readable storage medium 56 (such as disk-based or solid-state memory), and at least one network interface 58, which allows communication with other exemplified devices over the network 22 under the control of the server processor 54, and can actually facilitate communication between the server and client devices according to the present principles. It is noted that the network interface 58 can be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface, such as, for example, a wireless telephone transceiver.
[0046] Thus, in some embodiments, server 52 may be an Internet server or an entire "farm" of servers, and may include and perform "cloud" functionality, such that the devices of system 10 may access a "cloud" environment for, for example, online gaming applications, via server 52 in an example embodiment. Alternatively, server 52 may be implemented by one or more game consoles or other computers in the same room or nearby as the other devices shown.
[0047] The components shown in the following figures may include some or all of the components shown herein.Any user interfaces (UIs) described herein may be merged and / or extended, and UI elements may be mixed and matched between UIs.
[0048] The present principle can adopt various machine learning models, including deep learning models. The machine learning model consistent with the present principle can use various algorithms trained in the manner of learning including supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature learning, self-learning and other forms. The example of such an algorithm that can be implemented by a computer circuit includes one or more neural networks, such as convolutional neural networks (CNN), recursive neural networks (RNN) and RNN types known as long short-term memory (LSTM) networks. Support vector machines (SVM) and Bayesian networks can also be considered as examples of machine learning models. In addition to the network types set forth above, the model herein can be implemented by a classifier.
[0049] As understood herein, performing machine learning may thus involve accessing training data and then using that training data to train a model so that the model can process further data to make inferences. Thus, an artificial neural network / artificial intelligence model trained by machine learning may include an input layer, an output layer, and multiple hidden layers therebetween, the multiple hidden layers being configured and weighted to make inferences about an appropriate output.
[0050] Reference now Figure 2 and Figure 3 The initial logic that can be executed by any one or more processors herein begins at block 200, receiving a set of head images (in Figure 3 300), these images may be 2D images of the same base head from corresponding different perspectives that may be generated by one or more designers. Based on these images, a 3D base head ( Figure 3 302 in ). The 3D base head 302 can be a 3D neural radiance field (NeRF), which can be considered as a 3D volume stored in a machine learning (ML) model. In a specific example, the base head 302 can be a NeRF, which can be quickly trained in less than two minutes and in some examples less than one minute from the time of inputting the text description described below to produce a modified NeRF discussed further below. To this end, the example NeRF is a NeRF encoded in a multi-resolution hash table, which serves to store the locations of sparse multi-resolution 3D meshes to speed up the training and rendering of NeRF.
[0051] Note that in some embodiments, a head scan of a real head may be used instead of an image to generate the Figure 3 The base header 302 in .
[0052] Now go to Figure 4Once the base header 302 has been generated, a text description of the modified base header of the desired header may be received at 400. The text may be input from a text input device such as a keyboard and / or from a speech-to-text conversion of a spoken description, and / or from an initial starting phrase followed by learned additional descriptive phrases, as discussed further herein. Alternatively, a more primitive algorithmic technique may be used to generate madlib-style text so that answers to a questionnaire may be inserted into the text describing the target header.
[0053] Moving to block 402, text is input to a NeRF modification engine, which may include a fully connected (non-convolutional) deep network. The input to the modification engine may be a single continuous 5D input including values representing three spatial dimensions and two viewing dimensions, and the output of the engine may include volume density and viewing-dependent emitted radiation at a spatial location represented by associated 3D values.
[0054] Proceeding to block 404, the output of the NeRF modification engine, which may be considered a modified NeRF, is evaluated to determine how well the output matches the input text description. In one example, the evaluation may be performed by a comparative language-image pre-trained (CLIP) model that rates the matching of an image to an input text string. CLIP is an open source model that rates the matching of a line of text to an image. The CLIP model may be trained on image-text pairs (such as photos with captions from the Internet) using cosine similarity to rate the goodness of the match. Thus, the modification of the base NeRF is essentially manipulated using a loss function that is dependent on the text input of the CLIP model.
[0055] Decision diamond 406 indicates that if it is determined that the loss has not reached the threshold (target small) loss after the modified current iteration, the current iteration of the modified NeRF is modified again at box 408, and the logic loops back to box 404. However, once it is determined that the loss is small enough, the logic moves to box 410 to output the final modified NeRF. It is noted that instead of a loss threshold, a set number of loops may be performed, such as one hundred times. Boxes 402-408 may use a gradient descent technique.
[0056] In some examples, the logic may next move to box 412 to convert the modified NeRF to a polygonal mesh that may be imported into a computer simulation such as a computer game at box 414 and presented on a display during game play at box 416. Box 412 may also generate materials for the mesh, such as albedo, roughness, and normal textures. The modified NeRF may be converted to a polygonal mesh having the same topology as the modified NeRF. This may be accomplished using a marching cube technique. Alternatively, the modified NeRF may be converted to a polygonal mesh by essentially shrink-wrapping the modified NeRF by placing the mesh over the modified NeRF and recording which portions of the mesh contact the surface of the modified NeRF when the mesh is simulated as evacuated.
[0057] Figure 4A is an alternative illustration. Beginning at block 420, a NeRF is generated from parameters, which may include values that determine what a head looks like and are tuned to make the final NeRF look like the input image. Moving to block 422, the NeRF may be rendered from multiple camera angles, and enhancements such as color, rotation, and affine are applied at block 424.
[0058] Proceeding to block 426, at block 428, the CLIP image embedding may be retrieved and used to compare the current iteration of NeRF to the text image embedding to obtain a loss value. At block 430, the loss value is summed with other terms such as the degree of NeRF "blurriness" by obtaining the total variance of the depth of the rendered image. If the parameters need to be modified at block 420, the total loss is back-propagated using gradient descent at block 432.
[0059] Figures 5 to 8 The example uses the input text "realistic head of old tough guy" from Figure 3 302 in. The modified NeRF 500 is generated in less than two minutes from the time the text is input to the model.
[0060] Fig. 9 and Fig.10 Modified NeRFs 900, 1000 generated from the base NeRF 302 using corresponding input texts "realistic sea monster head" and "realistic devil head", respectively, are illustrated.
[0061] Fig.111108. The ML model described above, particularly for use with mesh techniques, is illustrated as being able to be trained based on a chain of causal relationships from initial parameters (variables) controlling the vertices of an object to the pixels rendered on the screen. Starting from vertex 1100, the chain of causality can proceed to primitives 1102 based on vertex 1100, and then further to fragments 1104, where the primitives are at least partially populated with, for example, textures. From fragments 1104, processing fragments 1106 can be generated, in which color can be added, which are merged into pixel output 1108 rendered on the screen. Arrow 1110 illustrates that the chain can be traversed back down to the original vertex 1100 to see how a particular value in the chain affects the final pixel image 1108, which can be used to train the ML model.
[0062] Note that if a mesh of the head is used instead of a radiance field, the ML model can perform differentiable rendering, where parameters can be changed to see in real time how these changes affect the output image. When the new face starts to match the target defined by the text, the ML model uses a loss function to minimize the loss.
[0063] As mentioned above, some or all of the input text describing the desired modified NeRF may be generated from the starting phrase using the learned subsequent phrases. Fig.12 1. A prompt template 1200 may be presented on a developer's computer screen with a number of phrases that may be filled in based on the player / developer's answers. Additional text 1202 may be presented for each prompt answer. Additional text 1202 may be generated by an ML model that is trained to complete sentences based on input phrases.
[0064] Fig.13 Additional techniques are illustrated. Beginning at block 1300, text is received indicating a character's biography and the equipment the character begins the game with. A text-to-image generator 1302 receives the text and generates an image consistent with the principles herein, which is imported into a computer game at block 1304 and rendered during game play at block 1306. Note that block 1302 may alternatively use a text-to-3D mesh generator (such as a NeRF generator) or an image (such as a decal on a player's equipment), but the more primary use case is to generate actual 3D assets in the form of meshes / materials.
[0065] It can now be appreciated that the modified 3D NeRF image from text is generated in real time, rather than being drawn, processed by image processing software or looked up or pre-designed. Instead, the modified 3D NeRF image is generated based on other text-based images.
[0066] Therefore, the present principles provide techniques for generating a coherent 3D head from text in a few minutes (e.g., two minutes or less or one minute or less). By having a starting base model, new NeRF is generated using text in a specific domain. Using hash table encoding of NeRF facilitates fast editing (~1 minute).
[0067] Now turn to Figures 14 to 18 , which illustrates the use of base models other than heads. For example, base models of game items (such as hats or weapons) can be used to produce hyper-personalized game items.
[0068] exist Fig.14 Beginning at block 1400 in the example, player data for a specific individual player is obtained. Preferably, the data is non-sensitive player data, such as first-party games previously played, to generate unique game items.
[0069] Moving to block 1402, text is created based on the data from block 1100. For example, if the player data indicates that the player is a fan of Game X, the system creates a text hint containing or related to "Game X" that is used at block 1404 to generate images of equipment such as different in-game masks that can be presented on the player's display at block 1406 and, if desired, presented on the player's character during game play at block 1408 based on, for example, player acceptance. The image may include material such as texture data.
[0070] The text generator can be relatively simple. For example, Fig.15 The final text prompt 1500 shown in Fig.16 A subsequent image 1600 in the image 1600 (presented on screen together) is a "Kabuki mask" and may include "based on [character Y] from [game X]." A generic hint may be "a [style] mask based on [character] from [game] ([release year])."
[0071] Fig.17 and 18 As further illustrated. A user interface (UI) 1700 may be presented on the player's display 1702, prompting (1704) the player whether he wishes to customize equipment, in this case a mask, for the player's character (PC). One or more selectors 1706 may be presented allowing the player to accept or not accept. If the player accepts, a text input may be automatically generated according to the principles described above to generate an image of the equipment, which may be displayed on the Fig.18 1800 of them are presented for players to view.
[0072] Fig.18It is further illustrated that selector 1802 may be presented to allow the player certain options, including the option to purchase customized equipment so that the same equipment is not provided to other players. A player may also wish to share equipment with other players.
[0073] While particular embodiments have been shown and described in detail herein, it is to be understood that the subject matter which is encompassed by the present invention is limited only by the claims.
Claims
1. A device comprising: at least one computer memory, the computer memory being not a transient signal and comprising instructions executable by at least one processor to: Generates the underlying three-dimensional (3D) neural radiance field (NeRF) from multiple images; generating a modified NeRF from the base NeRF using text input of a contrastive language-image pre-training (CLIP) model; and The modified NeRF is converted into a polygonal mesh representing a virtual human head for use in rendering the virtual human head in at least one computer simulation.
2. The apparatus of claim 1, wherein the CLIP model rates matches between images and text.
3. The apparatus of claim 2, wherein the CLIP model is trained on image-text pairs using cosine similarity to score goodness of match.
4. The apparatus of claim 1, wherein the instructions are executable to: A machine learning (ML) model is used on the base NeRF to minimize the loss indication in matching the text.
5. The apparatus of claim 4, wherein the instructions are executable to train the ML model based on a causal chain from initial image parameters controlling vertices of an object to pixels rendered by the object on a screen.
6. The apparatus of claim 4, wherein the ML model comprises at least one fully connected (non-convolutional) deep network.
7. The apparatus of claim 4, wherein the input to the ML model comprises values representing three spatial dimensions and two viewing dimensions.
8. The apparatus of claim 7, wherein outputs of the ML model include volume density and view-dependent emitted radiation.
9. The apparatus of claim 1, wherein the instructions are executable to: The text is generated from the starting phrase using the learned subsequent phrase.
10. The apparatus of claim 1, comprising the at least one processor.
11. An apparatus comprising: at least one processor programmed with instructions to: A text description of the recipient; generating a virtual coherent three-dimensional (3D) head in less than two minutes after receiving the textual description based at least in part on the textual description; and The virtual coherent 3D head is presented on a display.
12. The apparatus of claim 11, wherein the instructions are executable to: The virtual coherent 3D head is generated in less than one minute after receiving the text description.
13. The apparatus of claim 11, wherein the virtual coherent 3D head comprises a modified neural radiance field (NeRF).
14. The apparatus of claim 13, wherein the modified NeRF comprises a modified NeRF encoded in a multi-resolution hash table.
15. The apparatus of claim 13, wherein the instructions are executable to: using a text input of a contrastive language-image pre-training (CLIP) model to generate the modified NeRF from a base NeRF; and The modified NeRF is converted into a polygonal mesh representing a virtual human head for use in rendering the virtual human head in at least one computer simulation.
16. The apparatus of claim 15, wherein the CLIP model rates matches between images and text.
17. The apparatus of claim 11, wherein the instructions are executable to: A machine learning (ML) model is used to generate the virtual coherent 3D head by minimizing a loss indication in matching descriptive text.
18. The apparatus of claim 17, wherein the ML model comprises at least one fully connected deep network.
19. The apparatus of claim 17, wherein the input to the ML model comprises values representing three spatial dimensions and two viewing dimensions, and the output of the ML model comprises volume density and viewing-dependent emitted radiation.
20. A method comprising: Receive text; as well as A neural radiation field is generated based on the text starting from a base model.