Training method, 3D object generation method, device, equipment and medium
By optimizing the parameters of the neural radiation field and using the neural radiation field to generate optimized 3D objects, the problem of difficulty in generating high-quality 3D objects in the existing technology is solved, and semantic-related 3D object generation is achieved.
Patent Information
- Application Number
- CN202310714650.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-06-15
AI Technical Summary
It is difficult for the prior art to automatically generate 3D objects with good results.
By obtaining the 2D picture rendered by the neural radiation field, adding random noise, performing noise prediction, and optimizing the parameters of the neural radiation field based on the random noise and predicting noise, the optimized neural radiation field is generated for 3D object generation.
It realizes the generation of semantic-related 3D objects based on the prompt text, which improves the quality and semantic consistency of 3D object generation.
Smart Images

Figure CN116883587B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically computer vision, augmented reality, virtual reality, deep learning and other technical fields, and can be applied to scenes such as the metaverse and digital humans, and specifically to a training method for a neural radiation field for 3D object generation, a 3D object generation method, device, equipment and medium. Background Art
[0002] With the development of artificial intelligence technology, existing technologies are able to automatically generate images based on text, which not only provides great convenience for relevant workers in their work, but also provides ordinary users with many interesting applications.
[0003] However, current technology still makes it difficult to automatically generate 3D objects with good effects. Summary of the invention
[0004] The present disclosure provides a training method for a neural radiation field for 3D object generation, a 3D object generation method, an apparatus, a device and a medium.
[0005] According to one aspect of the present disclosure, a method for training a neural radiation field for 3D object generation is provided, comprising:
[0006] Obtaining a 2D image rendered by a neural radiation field, wherein the neural radiation field is created based on a signed distance function;
[0007] Adding random noise to the 2D picture to obtain a 2D noise picture;
[0008] Perform noise prediction on the 2D noise picture according to the prompt text to obtain predicted noise;
[0009] According to the random noise and the predicted noise, the parameters of the neural radiation field are optimized to obtain an optimized neural radiation field, wherein the optimized neural radiation field is used to generate a 3D object according to the prompt text.
[0010] According to another aspect of the present disclosure, a 3D object generation method is provided, comprising:
[0011] According to the signed distance function, extracting target sampling points from the optimized neural radiation field; wherein the optimized neural radiation field is obtained by training using the neural radiation field training method for 3D object generation described in any embodiment of the present disclosure;
[0012] Rendering a 3D object according to the target sampling point.
[0013] According to another aspect of the present disclosure, a training device for a neural radiation field generated by a 3D object is provided, comprising:
[0014] An image acquisition module, used to acquire a 2D image rendered by a neural radiation field, wherein the neural radiation field is created based on a signed distance function;
[0015] A noise adding module, used for adding random noise to the 2D picture to obtain a 2D noise picture;
[0016] A noise prediction module, used for performing noise prediction on the 2D noise picture according to the prompt text to obtain predicted noise;
[0017] An optimization module is used to optimize the parameters of the neural radiation field according to the random noise and the predicted noise to obtain an optimized neural radiation field, wherein the optimized neural radiation field is used to generate a 3D object according to the prompt text.
[0018] According to another aspect of the present disclosure, there is provided a 3D object generating device, comprising:
[0019] A sampling point extraction module, used to extract target sampling points from the optimized neural radiation field according to the signed distance function; wherein the optimized neural radiation field is obtained by training the neural radiation field training device for 3D object generation described in any embodiment of the present disclosure;
[0020] A rendering module is used to render a 3D object according to the target sampling point.
[0021] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0022] at least one processor; and
[0023] a memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training method of the neural radiation field for 3D object generation or the 3D object generation method described in any embodiment of the present disclosure.
[0025] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the neural radiation field training method for 3D object generation or the 3D object generation method described in any embodiment of the present disclosure.
[0026] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0028] Figure 1 is a flowchart of a method for training a neural radiation field for 3D object generation according to an embodiment of the present disclosure;
[0029] Figure 2 It is a schematic diagram of a process of rendering a 2D picture by a neural radiation field in a training method of a neural radiation field for generating a 3D object according to an embodiment of the present disclosure;
[0030] Figure 3 is a schematic diagram of a training process of a diffusion model according to an embodiment of the present disclosure;
[0031] Figure 4 is a flowchart of a 3D object generation method according to an embodiment of the present disclosure;
[0032] Figure 5 is a schematic structural diagram of a neural radiation field training device for 3D object generation according to an embodiment of the present disclosure;
[0033] Figure 6 is a structural schematic diagram of a 3D object generating device according to an embodiment of the present disclosure;
[0034] Figure 7 It is a block diagram of an electronic device used to implement the training method of the neural radiation field for 3D object generation according to the embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0036] Figure 1 This is a flowchart of a method for training a neural radiation field for 3D object generation according to an embodiment of the present disclosure. This embodiment can be applied to training a neural radiation field to generate a 3D object based on an optimized neural radiation field, and relates to the field of artificial intelligence technology, specifically computer vision, augmented reality, virtual reality, deep learning and other technical fields. The method can be performed by a training device for a neural radiation field for 3D object generation, which is implemented in software and / or hardware, and is preferably configured in an electronic device, such as a server, any computer device or smart terminal, etc. Figure 1 As shown, the method specifically includes the following:
[0037] S101. Obtain a 2D image rendered by a neural radiation field, wherein the neural radiation field is created based on a signed distance function.
[0038] S102: Add random noise to the 2D image to obtain a 2D noise image.
[0039] S103 . Predict the noise of the 2D noise image according to the prompt text to obtain predicted noise.
[0040] S104. Optimize the parameters of the neural radiation field according to the random noise and the predicted noise to obtain an optimized neural radiation field, wherein the optimized neural radiation field is used to generate a 3D object according to the prompt text.
[0041] Specifically, the neural radiation field is randomly initialized before training, and then the neural radiation field will render a 2D image under any camera pose. In one embodiment, the neural radiation field can obtain the 2D image by voxel rendering. Among them, the neural radiation field is created based on the signed distance function, and the signed distance function is used to perform surface constraints on the 3D object generated by the neural radiation field, so as to achieve the purpose of smoothing the surface of the object.
[0042] After obtaining a 2D image rendered by the neural radiation field, random noise, such as Gaussian noise, is added to the 2D image to obtain a 2D noise image, and then the noise of the 2D noise image is predicted according to the prompt text to obtain the predicted noise. Then, the parameters of the neural radiation field can be optimized according to the random noise and the predicted noise. For example, a loss function is constructed by comparing the random noise and the predicted noise, that is, a first loss is calculated according to the random noise and the predicted noise, and the parameters of the neural radiation field are optimized using the first loss.
[0043] Among them, a neural network model can be used to realize noise prediction, and the network model can generate a denoised image according to the prompt text by means of noise prediction. The parameters of the neural radiation field are optimized according to the random noise and the predicted noise, and the difference between the predicted noise and the random noise is made smaller and smaller during the iteration process, so as to ensure that the distribution of the 2D image rendered by the neural radiation field is consistent with the distribution of the image generated by the network model according to the prompt text, that is, the image semantics or features are consistent, thereby achieving the purpose of generating a 3D object according to the prompt text using the optimized neural radiation field, and making the generated 3D object semantically related to the prompt text. It should be noted that the prompt text can come directly from the text input by the user, or from the text extracted from the voice, or from any other form of input, and the present disclosure does not make any limitation on this.
[0044] In one embodiment, the total loss can be used to optimize the parameters of the neural radiation field, and the total loss includes the first loss described above, and can also include the reconstruction loss calculated according to the 2D picture and the prompt picture, that is, the parameters of the neural radiation field are optimized using the first loss and the reconstruction loss, for example, the first loss is added to the reconstruction loss to construct the total loss. Among them, the reconstruction loss can be calculated by comparing the 2D picture and the prompt picture. Specifically, since the neural radiation field can render the corresponding 2D picture under any camera pose, the camera perspective corresponding to the prompt picture can be the same as the camera perspective corresponding to the 2D picture for comparison, for example, both are front camera poses, or both are back camera poses, etc. Optimizing the parameters of the neural radiation field through reconstruction loss can make the 3D object finally generated by the neural radiation field semantically related to the prompt picture. For example, when the user enters the prompt text and the prompt picture, after optimization, the neural radiation field can finally generate a 3D object semantically related to the prompt text and the prompt picture, thereby improving the semantic confidence of the 3D object generation, and the generated 3D object has semantic consistency no matter which camera perspective is viewed from, ensuring that the generated 3D object is accurate. In this scenario, the prompt image can be used to represent the object or scene described by the prompt text and to depict the object or scene in more detail, thereby achieving a multimodal 3D generation task of generating 3D objects based on text and images.
[0045] In another embodiment, the total loss may also include a second loss for supervising the clustering of the voxel density of the neural radiation field, that is, the parameters of the neural radiation field are optimized using the first loss, the second loss and the reconstruction loss. For example, the first loss, the second loss and the reconstruction loss are added to construct the total loss. Exemplarily, the second loss may be an entropy loss calculated based on the visible density of the sampling point. Specifically, the visible density of the sampling point on the pixel ray corresponding to each pixel on the object is normalized, and then the entropy loss is calculated based on the entropy function and the normalized visible density. By optimizing the parameters of the neural radiation field through the second loss, the surface of the generated 3D object can be made sharper. The visible density of the sampling point is calculated based on the voxel density and the signed distance of the sampling point.
[0046] In another embodiment, the total loss may also include a third loss for supervising the correctness of the signed distance function, that is, the parameters of the neural radiation field are optimized using the first loss, the second loss, the third loss and the reconstruction loss. For example, the first loss, the second loss, the third loss and the reconstruction loss are added to construct the total loss. Since the neural radiation field is constructed based on the signed distance function, the relevant parameters of the signed distance function need to be optimized during the optimization process to ensure that the optimized signed distance function can truly represent the signed distance of the sampling points. Therefore, a third loss for supervising the correctness of the signed distance function can be added to the total loss. For example, the third loss can be an Eikonal loss. By optimizing the parameters of the neural radiation field through the third loss, the accuracy of the signed distance of the sampling points output by the neural radiation field can be improved, thereby determining a more accurate object surface, thereby improving the accuracy of the generated 3D object. At the same time, the object surface can be constrained during rendering to improve the smoothness of the object surface.
[0047] Figure 2 This is a flow chart of rendering a 2D image with a neural radiation field in a training method for generating a 3D object according to an embodiment of the present disclosure. This embodiment further optimizes how the neural radiation field renders a 2D image based on the above embodiment. Figure 2 As shown, the method specifically includes the following:
[0048] S201. Output the color and sign distance of the sampling point on each pixel ray under each camera pose according to any camera pose.
[0049] Among them, the neural radiation field constructed based on the signed distance function is used to output the color and signed distance of the sampling point. The signed distance is used to determine the sampling point located on the surface of the object. For example, the sampling point with a signed distance of zero is the sampling point located on the surface of the object. During the rendering process, the sampling point located on the surface of the object has a greater weight on the color of the pixel in the 2D image than other sampling points, that is, the sampling point located on the surface of the object contributes more to the color of the pixel than other sampling points during the rendering process. In this way, the surface of the rendered object will be smoother, improving the rendering effect.
[0050] S202 , performing voxel rendering according to the color of the sampling point on each pixel ray and a weight function to obtain the color of each pixel point corresponding to each pixel ray.
[0051] S203, obtaining a 2D image at each camera position according to the color of each pixel point; wherein the value of the weight function of the sampling point located on the surface of the object in the 2D image determined according to the symbol distance is the largest.
[0052] Specifically, the weight function can be calculated based on the transparency and visible density of the sampling point, and the visible density is calculated based on the voxel density and signed distance of the sampling point. Therefore, the weight function of the sampling point is different for different signed distances. In this embodiment, the value of the weight function of the sampling point located on the surface of the object in the 2D picture determined by the signed distance is the largest. Therefore, the surface constraint of the object can be achieved through the signed distance to improve the smoothness of the object surface. It should be noted that for how the neural radiation field calculates the color of the sampling point, please refer to the introduction in the prior art, which will not be repeated here.
[0053] In one embodiment, a diffusion model capable of generating images from text may be used to realize noise prediction. That is, a pre-trained diffusion model is used to predict noise on a 2D noise image according to the prompt text to obtain predicted noise. The diffusion model is a pre-trained model, and its training process may be to first perform initial training on the diffusion model, and then fine-tune the diffusion model after initial training based on a given image. The fine-tuned diffusion model is used to generate an image that is semantically related to the given image. Specifically, Figure 3 As shown, Figure 3 Schematic diagram of the training process of the diffusion model according to the embodiment of the present disclosure. This embodiment further optimizes how to train the diffusion model based on the above embodiment. The method specifically includes the following:
[0054] S301, performing initial training on the diffusion model.
[0055] Among them, the initial training can be implemented by sampling the training method in the prior art, which will not be described here.
[0056] S302: Generate a picture model using the picture to generate training pictures of the given picture at different viewing angles.
[0057] S303: Train the diffusion model using the given prompt words and training pictures used in the initial training to fine-tune the diffusion model.
[0058] The picture-to-picture model can be implemented using an existing model, which will not be described here. The training pictures generated by the picture-to-picture model can be multiple pictures, and they are pictures of the given picture at other perspectives. That is to say, the picture-to-picture model can be used to construct pictures of the same object at different perspectives. Then, the diffusion model obtained after the initial training is trained again using given prompt words and these training pictures, so as to achieve the purpose of fine-tuning. When the prompt words are given, the pictures generated by the fine-tuned diffusion model have a high degree of semantic consistency with the training pictures used during training. Therefore, the fine-tuned diffusion model is used to predict noise, so that the parameters of the neural radiation field are optimized according to random noise and predicted noise, so that the 3D pictures generated by the optimized neural radiation field at different perspectives have semantic consistency, thereby improving the quality of 3D object generation.
[0059] The technical solution of the disclosed embodiment, which models the neural radiation field based on the signed distance function, can achieve unbiased rendering, improve the smoothness of the object surface, and enable the surface material of the object to play the greatest role in rendering. Moreover, by adding reconstruction loss to the total loss function for optimizing the neural radiation field, the correlation between the rendered image and the provided image can be improved, thereby improving semantic consistency, and multi-modal high-quality 3D object generation can also be achieved. In addition, by fine-tuning the diffusion model, the semantic consistency of the images generated by the diffusion model with the prompt words and the images used in training is further improved, thereby improving the semantic consistency of 3D objects rendered by the neural radiation field from different perspectives.
[0060] Figure 4 This is a flow chart of a 3D object generation method according to an embodiment of the present disclosure. This embodiment can be applied to training a neural radiation field to generate a 3D object based on an optimized neural radiation field, and relates to the field of artificial intelligence technology, specifically computer vision, augmented reality, virtual reality, deep learning and other technical fields. The method can be executed by a 3D object generation device, which is implemented in software and / or hardware, and is preferably configured in an electronic device, such as a server, any computer device or smart terminal. Figure 4 As shown, the method specifically includes the following:
[0061] S401. Extract target sampling points from the optimized neural radiation field according to the signed distance function; wherein the optimized neural radiation field is trained using the training method for the neural radiation field for 3D object generation as described in any embodiment of the present disclosure.
[0062] S402: Rendering a 3D object according to the target sampling point.
[0063] Among them, sampling points with a signed distance of zero can be extracted from the optimized neural radiation field as target sampling points, that is, points located on the surface of the object, and the color of the target sampling points calculated by the neural radiation field can be obtained. The mesh of the 3D object can be determined according to the color of the target sampling point, so as to render the 3D object according to the mesh. Exemplarily, a surface drawing algorithm can be used to extract an isosurface composed of sampling points with a signed distance of zero from the optimized neural radiation field to obtain the mesh of the 3D object. The method of extracting sampling points using signed distance and rendering 3D objects can more accurately obtain the isosurface located on the surface of the object, making the generated 3D object more accurate and improving the rendering effect.
[0064] Figure 5 This is a schematic diagram of the structure of a training device for a neural radiation field for 3D object generation according to an embodiment of the present disclosure. This embodiment can be applied to training a neural radiation field to generate a 3D object based on an optimized neural radiation field, and relates to the field of artificial intelligence technology, specifically computer vision, augmented reality, virtual reality, deep learning and other technical fields. The device can implement the training method for a neural radiation field for 3D object generation described in any embodiment of the present disclosure. Figure 5 As shown, the device 500 specifically includes:
[0065] An image acquisition module 501 is used to acquire a 2D image rendered by a neural radiation field, wherein the neural radiation field is created based on a signed distance function;
[0066] A noise adding module 502, configured to add random noise to the 2D picture to obtain a 2D noise picture;
[0067] A noise prediction module 503 is used to perform noise prediction on the 2D noise picture according to the prompt text to obtain predicted noise;
[0068] The optimization module 504 is used to optimize the parameters of the neural radiation field according to the random noise and the predicted noise to obtain an optimized neural radiation field, wherein the optimized neural radiation field is used to generate a 3D object according to the prompt text.
[0069] Optionally, the 2D image is obtained by voxel rendering of the neural radiation field.
[0070] Optionally, the neural radiation field is used to output the color and signed distance of the sampling point, and the signed distance is used to determine the sampling point located on the surface of the object.
[0071] Optionally, during the rendering process, the sampling point located on the surface of the object has a greater weight on the color of the pixel in the 2D image than other sampling points.
[0072] Optionally, the 2D image is rendered by the neural radiation field as follows:
[0073] Output the color and sign distance of the sampling points on each pixel ray under each camera pose according to any camera pose;
[0074] Perform voxel rendering according to the color of the sampling point on each pixel ray and a weight function to obtain the color of each pixel point corresponding to each pixel ray;
[0075] Obtaining a 2D image at each camera position according to the color of each pixel;
[0076] Among them, the value of the weight function of the sampling point located on the surface of the object in the 2D picture determined according to the symbol distance is the largest.
[0077] Optionally, the optimization module includes:
[0078] The first optimization unit is used to calculate a first loss according to the random noise and the predicted noise, and optimize the parameters of the neural radiation field using the first loss.
[0079] Optionally, the optimization module further includes:
[0080] A second optimization unit is used to calculate a reconstruction loss based on the 2D image and the prompt image, and optimize the parameters of the neural radiation field using the reconstruction loss; wherein the 3D object is semantically related to the prompt image.
[0081] Optionally, the optimization module further includes:
[0082] The third optimization unit is used to optimize the parameters of the neural radiation field using a second loss, wherein the second loss is used to supervise the clustering of the voxel density of the neural radiation field.
[0083] Optionally, the optimization module further includes:
[0084] The fourth optimization unit is used to optimize the parameters of the neural radiation field using a third loss, wherein the third loss is used to supervise the correctness of the signed distance function.
[0085] Optionally, the noise prediction module is specifically used for:
[0086] Using a pre-trained diffusion model, noise prediction is performed on the 2D noise image according to the prompt text to obtain predicted noise.
[0087] Optionally, the diffusion model is trained in the following manner:
[0088] The diffusion model after initial training is fine-tuned based on a given picture; wherein the fine-tuned diffusion model is used to generate pictures semantically related to the given picture.
[0089] Optionally, the fine-tuning process includes:
[0090] Generate a picture model using the picture to generate training pictures of the given picture at different viewing angles;
[0091] The diffusion model is trained using the given prompt words used in the initial training and the training pictures to fine-tune the diffusion model.
[0092] Figure 6 This is a schematic diagram of the structure of a 3D object generation device according to an embodiment of the present disclosure. This embodiment can be applied to training a neural radiation field to generate a 3D object based on an optimized neural radiation field, and relates to the field of artificial intelligence technology, specifically computer vision, augmented reality, virtual reality, deep learning and other technical fields. The device can implement the 3D object generation method described in any embodiment of the present disclosure. Figure 6 As shown, the device 600 specifically includes:
[0093] A sampling point extraction module 601 is used to extract target sampling points from the optimized neural radiation field according to the signed distance function; wherein the optimized neural radiation field is obtained by training the neural radiation field training device for 3D object generation as described in any embodiment of the present disclosure;
[0094] The rendering module 602 is used to render a 3D object according to the target sampling point.
[0095] Optionally, the sign distance of the target sampling point is zero.
[0096] Optionally, the rendering module includes:
[0097] A grid determination unit, configured to determine a grid of the 3D object according to the color of the target sampling point;
[0098] A rendering unit is used to render the 3D object according to the grid.
[0099] The above-mentioned product can execute the method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0100] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0101] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0102] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0103] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0104] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0105] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the XX method. For example, in some embodiments, the XX method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the XX method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the XX method in any other appropriate manner (e.g., by means of firmware).
[0106] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0107] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0108] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0109] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0110] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0111] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services. The server may also be a server of a distributed system, or a server combined with a blockchain.
[0112] Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, knowledge graph technology, and other major directions.
[0113] Cloud computing refers to a technology system that uses network access to elastically scalable shared physical or virtual resource pools. Resources can include servers, operating systems, networks, software, applications, and storage devices, and can be deployed and managed on demand and in a self-service manner. Cloud computing technology can provide efficient and powerful data processing capabilities for technical applications such as artificial intelligence and blockchain, as well as model training.
[0114] In addition, according to an embodiment of the present disclosure, the present disclosure also provides another electronic device, another readable storage medium and another computer program product, which are used to execute one or more steps of the 3D object generation method described in any embodiment of the present disclosure. The specific structure and program code thereof can be found in Figure 7 The description of the contents of the illustrated embodiment will not be repeated here.
[0115] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution provided by this disclosure can be achieved, and this document does not limit this.
[0116] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for training a neural radiation field for 3D object generation, comprising: Obtaining a 2D image rendered by a neural radiation field, wherein the neural radiation field is created based on a signed distance function, the neural radiation field is used to output the color and signed distance of a sampling point, the signed distance is used to determine a sampling point located on a surface of an object, and during the rendering process, the sampling point located on the surface of the object has a greater weight on the color of a pixel point in the 2D image than other sampling points, the weight is calculated based on the transparency and visible density of the sampling point, and the visible density is calculated based on the voxel density and signed distance of the sampling point; Adding random noise to the 2D picture to obtain a 2D noise picture; Perform noise prediction on the 2D noise picture according to the prompt text to obtain predicted noise; Optimizing the parameters of the neural radiation field according to the random noise and the predicted noise to obtain an optimized neural radiation field, wherein the optimized neural radiation field is used to generate a 3D object according to the prompt text; The step of optimizing the parameters of the neural radiation field according to the random noise and the predicted noise includes: The first loss, the second loss, the third loss and the reconstruction loss are added to construct a total loss, and the total loss is used to optimize the parameters of the neural radiation field; Among them, the first loss is calculated based on the random noise and the predicted noise; the reconstruction loss is calculated by comparing the 2D picture and the prompt picture; the second loss is used to supervise the clustering of the voxel density of the neural radiation field; the third loss is used to supervise the correctness of the signed distance function; the 3D object is semantically related to the prompt picture; the camera perspective corresponding to the prompt picture is the same as the camera perspective corresponding to the compared 2D picture, and the prompt picture is used to represent the object or scene described by the prompt text.
2. The method according to claim 1, wherein: The 2D image is obtained by voxel rendering of the neural radiation field.
3. The method according to claim 1, wherein: The 2D image is rendered using the neural radiance field as follows: Output the color and sign distance of the sampling points on each pixel ray under each camera pose according to any camera pose; Perform voxel rendering according to the color of the sampling point on each pixel ray and a weight function to obtain the color of each pixel point corresponding to each pixel ray; Obtaining a 2D image at each camera position according to the color of each pixel; Among them, the value of the weight function of the sampling point located on the surface of the object in the 2D picture determined according to the symbol distance is the largest.
4. The method according to claim 1, wherein: The performing noise prediction on the 2D noise picture to obtain predicted noise includes: Using a pre-trained diffusion model, noise prediction is performed on the 2D noise image according to the prompt text to obtain predicted noise.
5. The method according to claim 4, wherein: The diffusion model is trained in the following way: The diffusion model after initial training is fine-tuned based on a given picture; wherein the fine-tuned diffusion model is used to generate pictures semantically related to the given picture.
6. The method according to claim 5, wherein: The fine-tuning process includes: Generate a picture model using the picture to generate training pictures of the given picture at different viewing angles; The diffusion model is trained using the given prompt words used in the initial training and the training pictures to fine-tune the diffusion model.
7. A 3D object generation method, comprising: According to the signed distance function, extracting target sampling points from the optimized neural radiation field; wherein the optimized neural radiation field is obtained by training using the method described in any one of claims 1 to 6; Rendering a 3D object according to the target sampling point.
8. The method according to claim 7, wherein: The sign distance of the target sampling point is zero.
9. The method according to claim 8, wherein: The rendering of the 3D object according to the target sampling point comprises: Determining a grid of the 3D object according to the color of the target sampling point; The 3D object is rendered according to the mesh.
10. A training device for a neural radiation field generated by a 3D object, comprising: An image acquisition module, used to acquire a 2D image rendered by a neural radiation field, wherein the neural radiation field is created based on a signed distance function, the neural radiation field is used to output the color and signed distance of a sampling point, the signed distance is used to determine a sampling point located on the surface of an object, and during the rendering process, the sampling point located on the surface of the object has a greater weight on the color of a pixel in the 2D image than other sampling points, the weight is calculated based on the transparency and visible density of the sampling point, and the visible density is calculated based on the voxel density and signed distance of the sampling point; A noise adding module, used for adding random noise to the 2D picture to obtain a 2D noise picture; A noise prediction module, used for performing noise prediction on the 2D noise picture according to the prompt text to obtain predicted noise; an optimization module, used for optimizing the parameters of the neural radiation field according to the random noise and the predicted noise to obtain an optimized neural radiation field, wherein the optimized neural radiation field is used to generate a 3D object according to the prompt text; Wherein, the optimization module is specifically used for: The first loss, the second loss, the third loss and the reconstruction loss are added to construct a total loss, and the total loss is used to optimize the parameters of the neural radiation field; Among them, the first loss is calculated based on the random noise and the predicted noise; the reconstruction loss is calculated by comparing the 2D picture and the prompt picture; the second loss is used to supervise the clustering of the voxel density of the neural radiation field; the third loss is used to supervise the correctness of the signed distance function; the 3D object is semantically related to the prompt picture; the camera perspective corresponding to the prompt picture is the same as the camera perspective corresponding to the compared 2D picture, and the prompt picture is used to represent the object or scene described by the prompt text.
11. The device according to claim 10, wherein: The 2D image is obtained by voxel rendering of the neural radiation field.
12. The device according to claim 10, wherein: The 2D image is rendered using the neural radiance field as follows: Output the color and sign distance of the sampling points on each pixel ray under each camera pose according to any camera pose; Perform voxel rendering according to the color of the sampling point on each pixel ray and a weight function to obtain the color of each pixel point corresponding to each pixel ray; Obtaining a 2D image at each camera position according to the color of each pixel; Among them, the value of the weight function of the sampling point located on the surface of the object in the 2D picture determined according to the symbol distance is the largest.
13. The device according to claim 10, wherein: The noise prediction module is specifically used for: Using a pre-trained diffusion model, noise prediction is performed on the 2D noise image according to the prompt text to obtain predicted noise.
14. The device according to claim 13, wherein: The diffusion model is trained in the following way: The diffusion model after initial training is fine-tuned based on a given picture; wherein the fine-tuned diffusion model is used to generate pictures semantically related to the given picture.
15. The device according to claim 14, wherein: The fine-tuning process includes: Generate a picture model using the picture to generate training pictures of the given picture at different viewing angles; The diffusion model is trained using the given prompt words used in the initial training and the training pictures to fine-tune the diffusion model.
16. A 3D object generation device, comprising: A sampling point extraction module, used to extract target sampling points from the optimized neural radiation field according to the signed distance function; wherein the optimized neural radiation field is obtained by training the device according to any one of claims 10 to 15; A rendering module is used to render a 3D object according to the target sampling point.
17. The device according to claim 16, wherein: The sign distance of the target sampling point is zero.
18. The device according to claim 17, wherein: The rendering module includes: A grid determination unit, configured to determine a grid of the 3D object according to the color of the target sampling point; A rendering unit is used to render the 3D object according to the grid.
19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the neural radiation field training method for 3D object generation described in any one of claims 1-6 or the 3D object generation method described in any one of claims 7-9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the neural radiation field training method for 3D object generation according to any one of claims 1-6 or the 3D object generation method according to any one of claims 7-9.
21. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the training method of a neural radiation field for 3D object generation according to any one of claims 1-6 or the 3D object generation method according to any one of claims 7-9.
Citation Information
Patent Citations
Image processing method and device based on neural radiation field
CN114972632A
High dynamic range view synthesis from noisy raw images
WO2023086194A1