Endogenous watermark adding method and device for neural radiation field, equipment and storage medium

By endogenous watermarks during NeRF generation, the problem of delay in adding watermarks after NeRF generation in the prior art is solved, and effective ownership protection and flexible development capabilities of NeRF are achieved.

CN120070143APending Publication Date: 2025-05-30SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510024234.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has a delay in adding a watermark to the generated neural radiation field (NeRF), and a malicious user may acquire NeRF and claim ownership before the watermark is added.

Method used

A method for adding endogenous watermarks to the neural radiation field is proposed. By adding watermarks in the generation process of digital 3D content, combined with the 3D generation architecture of fractional distillation sampling, the endogenous watermarks are achieved during training to generate NeRF, so as to achieve good ownership protection of NeRF.

Benefits of technology

It realizes the direct addition of watermarks during NeRF generation, avoids the delay in adding watermarks after generation, ensures that NeRF has watermark functions during generation, effectively prevents NeRF from being stolen, and increases flexibility for future development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070143A_ABST
    Figure CN120070143A_ABST
Patent Text Reader

Abstract

The invention provides an endogenous watermark adding method and device for a neural radiation field, equipment and a storage medium, and belongs to the technical field of computers, and the method comprises the steps: generating a trigger view angle set according to given watermark information; performing iterative training on the NeRF network based on a randomly selected target view angle to obtain a target NeRF added with a watermark; in at least one iteration training, when the target view angle belongs to the trigger view angle set, fractional distillation sampling loss of a rendered picture of the target view angle is calculated, a binary cross entropy function of decoding watermark information and given watermark information of the rendered picture is calculated, and network parameters of NeRF are optimized based on the fractional distillation sampling loss and the binary cross entropy function. Wherein the decoded watermark information is obtained by analyzing the rendered picture based on a preset watermark decoder. According to the method and the device, the watermark can be added in the generation process of the digital 3D content, and the high-quality NeRF with the watermark is generated in combination with a fractional distillation sampling type 3D generation architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to an in - built watermark adding method, device, equipment, and storage medium for neural radiance fields. Background Art

[0002] Digital 3D content has become an essential part of the metaverse and virtual and augmented reality, enabling visualization, understanding, and interaction with complex scenes representing our real life. To generate high - quality digital 3D assets, a large amount of time, computing resources, and skilled expertise are required. Therefore, protecting the ownership of the generated digital 3D content has become even more important.

[0003] In related technologies, for digital 3D content - neural radiance fields (NeRFs), watermarks are mostly added to NeRFs in a post - generation watermarking manner to protect the ownership of NeRFs. For example, both CopyRNeRF and WateNeRF first create the NeRF and then embed the watermark into the NeRF. Among them, CopyRNeRF adds an additional message feature field to the network structure of the NeRF, and by fine - tuning the NeRF, watermarks can be extracted from the rendered images of the NeRF at any angle; WateNeRF embeds the watermark in the frequency domain of the NeRF's rendered images, without requiring additional modification to the NeRF's network structure, and WateNeRF can also extract the watermark from the frequency domain of the NeRF's rendered images at any angle by fine - tuning the NeRF.

[0004] However, due to the delay in the post - generation watermarking method before adding the watermark to the generated NeRF and embedding the watermark after generating the NeRF, malicious users have the opportunity to obtain the NeRF before watermarking it and claim ownership of it. Summary of the Invention

[0005] The main purpose of the embodiments of this application is to propose an in - built watermark adding method, device, equipment, and storage medium for neural radiance fields. By adding a watermark during the generation process of digital 3D content and combining it with a 3D generation architecture such as fractional distillation sampling, a high - quality NeRF with a built - in watermark is obtained when training to generate the NeRF, achieving good ownership protection for the NeRF.

[0006] To achieve the above object, the first aspect of the embodiments of this application proposes an in - built watermark adding method for neural radiance fields, and the method includes:

[0007] Generate a trigger view set according to the given watermark information;

[0008] Iteratively train the network of the Neural Radiance Field (NeRF) based on randomly selected target viewpoints to obtain a target NeRF with a watermark added.

[0009] Among them, at least one iteration training in the iterative training process performs the following operations: when the target viewpoint belongs to the set of trigger viewpoints, calculate the score distillation sampling loss of the rendered image of the target viewpoint, and calculate the binary cross-entropy function between the decoded watermark information of the rendered image and the given watermark information; wherein, the rendered image is obtained from the NeRF, and the decoded watermark information is obtained by parsing the rendered image based on a preset watermark decoder; and optimize the network parameters of the NeRF based on the score distillation sampling loss and the binary cross-entropy function.

[0010] In some embodiments, the method further includes:

[0011] Obtain image data and original watermark information.

[0012] Train a preset deep neural network model based on the image data and the original watermark information to obtain the watermark decoder.

[0013] In some embodiments, the deep neural network model includes an encoder, a noise layer, and a decoder.

[0014] The training of the preset deep neural network model based on the image data and the original watermark information to obtain the watermark decoder includes:

[0015] Input the image data and the original watermark information into the encoder to obtain a watermark image.

[0016] Add noise to the watermark image based on the noise layer to obtain a noisy watermark image.

[0017] Input the noisy watermark image into the decoder to obtain the parsed watermark information.

[0018] Update the deep neural network model based on the binary cross-entropy function between the parsed watermark information and the original watermark information to obtain the watermark decoder.

[0019] In some embodiments, the adding noise to the watermark image based on the noise layer to obtain a noisy watermark image includes:

[0020] Calculate the residual of the noise to be added and remove the gradient to obtain a residual calculation result; wherein, the noise to be added belongs to non-differentiable noise.

[0021] Based on the noise layer, add the residual calculation result to the watermark image to obtain a noisy watermark image.

[0022] In some embodiments, the generating a set of trigger perspectives according to given watermark information includes:

[0023] Obtain the hash value of the given watermark information;

[0024] Initialize a pseudo-random number generator based on the hash value;

[0025] Generate a set of trigger perspectives based on the initialized pseudo-random number generator.

[0026] In some embodiments, at least one iteration training in the iterative training process further performs the following operations:

[0027] In the case where the target perspective does not belong to the set of trigger perspectives, calculate the score distillation sampling loss of the rendered image of the target perspective; and, optimize the network parameters of the NeRF based on the score distillation sampling loss.

[0028] To achieve the above object, a second aspect of the embodiments of the present application proposes an inborn watermark addition device for a neural radiance field, the device includes:

[0029] A trigger perspective generation module, configured to generate a set of trigger perspectives according to given watermark information;

[0030] A NeRF network training module, configured to iteratively train the network of the neural radiance field NeRF based on randomly selected target perspectives to obtain a target NeRF with a watermark added;

[0031] Wherein, the NeRF network training module performs the following operations in at least one iteration training in the iterative training process: in the case where the target perspective belongs to the set of trigger perspectives, calculate the score distillation sampling loss of the rendered image of the target perspective, and calculate the binary cross-entropy function of the decoded watermark information of the rendered image and the given watermark information; wherein, the rendered image is obtained from the NeRF, and the decoded watermark information is obtained by parsing the rendered image based on a preset watermark decoder; and, optimize the network parameters of the NeRF based on the score distillation sampling loss and the binary cross-entropy function to add a watermark to the NeRF during the generation of the NeRF.

[0032] To achieve the above object, a third aspect of the embodiments of the present application proposes a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method described in the first aspect above when executing the computer program.

[0033] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a computer-readable storage medium storing a computer program, which when executed by a processor implements the method described in the first aspect above.

[0034] To achieve the above object, a fifth aspect of the embodiments of the present application proposes a computer program product storing a computer program, which when executed by a processor implements the method described in the first aspect above.

[0035] The method, device, computer device, computer-readable storage medium, and computer program product for adding endogenous watermarks to a neural radiance field proposed in the embodiments of the present application generate a set of trigger viewpoints that only depend on a given secret message (watermark information), and then iteratively train the network of the neural radiance field NeRF based on randomly selected target viewpoints. When the target viewpoint belongs to the set of trigger viewpoints, calculate the score distillation sampling loss of the rendered image of the target viewpoint, and calculate the binary cross-entropy function between the decoded watermark information of the rendered image and the given watermark information, so as to optimize the network parameters of NeRF based on the score distillation sampling loss and the binary cross-entropy function. Among them, the rendered image is obtained from NeRF, and the decoded watermark information is obtained by parsing the rendered image based on a preset watermark decoder. Thus, watermarks are added to NeRF during the process of score distillation sampling, and the target NeRF with added watermarks is directly generated.

[0036] In this way, the embodiments of the present application implement a solution for adding watermarks during the generation of text-to-3D (text generation 3D model). Combined with a 3D generation architecture of the score distillation sampling type, it can endogenously watermark the NeRF during training to generate a high-quality target NeRF with a built-in watermark, achieving good ownership protection for NeRF. And, compared with the method of adding watermarks to NeRF after generation, the method of adding watermarks to NeRF by endogenous watermarking in the embodiments of the present application can directly generate a target NeRF with watermarks and does not require changing the NeRF architecture, thus increasing the flexibility for future further development in 3D generation (this target NeRF). BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic flowchart of the steps of the method for adding endogenous watermarks to a neural radiance field provided by the embodiments of the present application in some embodiments;

[0038] Figure 2 It is a schematic flowchart of the steps of the method for adding endogenous watermarks to a neural radiance field provided by the embodiments of the present application in other embodiments;

[0039] Figure 3 ForFigure 2 Schematic diagram of the refined step flow of step S202;

[0040] Figure 4 Schematic diagram of the complete watermarking process involved in the method for adding endogenous watermarks to a neural radiance field proposed in an embodiment of the present application;

[0041] Figure 5 Schematic diagram of the logical flow of the training watermark decoder stage involved in the method for adding endogenous watermarks to a neural radiance field proposed in an embodiment of the present application in some embodiments;

[0042] Figure 6 Schematic diagram of the logical flow of the watermark addition stage involved in the method for adding endogenous watermarks to a neural radiance field proposed in an embodiment of the present application in some embodiments;

[0043] Figure 7 Schematic diagram of the structure of the device for adding endogenous watermarks to a neural radiance field provided in an embodiment of the present application;

[0044] Figure 8 Schematic diagram of the hardware structure of the computer device provided in an embodiment of the present application. Detailed implementation manners

[0045] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0046] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0048] First, several terms involved in the present application are analyzed:

[0049] Neural Radiance Field (NeRF).

[0050] NeRF is a representation form of digital 3D content. 3D representations also include meshes, etc. The neural network of NeRF consists of MLPf σ and f cComposed, given 3D coordinates and the observation direction f σ and f c Map them into the color of points in the following way and the geometric density

[0051] σ,z = f σ (γ x (x));

[0052] c = f c (z,γ d (d)).

[0053] where γ x and γ d are the encoding functions of the coordinates and the observation direction respectively. To render the image of the given viewing angle p, the rendering process defines the observation ray of each pixel in 3D space and calculates the weighted sum of the colors c j of the sampling points along each ray as the color of the current observation ray:

[0054]

[0055] where δ k is the distance between adjacent sampling points. For convenience of representation, it can be directly represented by g(θ,p) ∈ [0,1] H×W to represent the above rendering process. Among them, θ represents the network parameters of NeRF, and g takes the viewing direction p as the input and outputs a normalized image.

[0056] Diffusion model.

[0057] The diffusion model is a technical route for text-to-image. The diffusion model involves a forward process {q t} t∈[0,1] (gradually adding noise to the data point x 0 ~q 0 (x 0 ))), and a reverse process {p t} t∈[0,1] (denoising, generating data). The forward process is defined by and, q t (x t ): = ∫q t (x t |x 0 )q 0 (x 0 )dx 0 where, α t ,σ tis a hyperparameter. The reverse process is defined as: using the parameterized noise prediction network ∈ φ (x t ,t) to denoise in order to predict the noise added to the clean data x 0 . This network is trained by minimizing the following conditions:

[0058]

[0059] where w(t) is a time-dependent weighting function. After training, we have p t ≈q t ; thus, samples can be drawn from p 0 ≈q 0 . One of the most important applications is to generate images from text, where the noise prediction model ∈ φ (x t ,t,y) is conditioned on the text prompt y.

[0060] Score Distillation Sampling.

[0061] Score Distillation Sampling is a widely used technique for text-to-3D model generation that elevates 2D information into a 3D NeRF model by distilling a pre-trained diffusion model. Given a pre-trained text-to-image diffusion model p t (x t |y) and a noise prediction network ∈ φ (x t ,t,y), Score Distillation Sampling optimizes the parameters θ of a single NeRF using the Score Distillation Sampling loss function. Specifically, given a camera view p, a prompt y, and a differentiable rendering map g(θ,p), Score Distillation Sampling optimizes θ by minimizing the following terms:

[0062]

[0063] where x t =α t g(θ,p)+σ t ∈. In practical applications, to save computing power, Score Distillation Sampling does not directly optimize the above formula but instead uses the gradient fitted by the above formula:

[0064]

[0065] Next, the overall concept of the embodiments of this application will be described.

[0066] Digital 3D content has become an indispensable part of the metaverse and virtual and augmented reality, enabling visualization, understanding, and interaction with complex scenarios representing our real lives. To generate high-quality digital 3D assets, a significant amount of time, computing resources, and skilled expertise are required. Therefore, protecting the ownership of the generated digital 3D content becomes even more important.

[0067] Text-to-3D generation has become a focus in 3D content modeling. Current popular 3D generation algorithms can generate 3D representations such as meshes and NeRFs. The embodiments of this application focus on the generation of NeRFs because NeRFs can represent 3D models more compactly. Recent text-to-3D generation methods can generate NeRFs by refining pre-trained diffusion models (such as Stable Diffusion) through given text descriptions. The basis for this remarkable progress is the use of Score Distillation Sampling (SDS). Using SDS, NeRF training can be performed without real images. Therefore, the main problem to be addressed in the embodiments of this application is: how to protect the ownership of the neural radiance fields generated by score distillation sampling.

[0068] In related technologies, for NeRFs, watermarks are mostly added to NeRFs in a post-generation watermarking manner to protect the ownership of NeRFs. For example, both CopyRNeRF and WateNeRF create NeRFs first and then embed watermarks in them. Among them, CopyRNeRF adds an additional message feature field to the network structure of NeRF, and by fine-tuning NeRF, watermarks can be extracted from the rendered images of NeRF at any angle; WateNeRF embeds watermarks in the frequency domain of the rendered images of NeRF, and no additional modification to the network structure of NeRF is required. WateNeRF can also extract watermarks from the frequency domain of the rendered images of NeRF at any angle by fine-tuning NeRF.

[0069] However, due to the delay in adding watermarks to the generated NeRFs in the post-generation watermarking method and the fact that watermarks are embedded after NeRF generation, malicious users have the opportunity to obtain the NeRF and claim ownership of it before adding watermarks to the NeRF. Secondly, CopyRNeRF also incurs additional watermarking costs, and because CopyRNeRF needs to fine-tune the additional message feature field in the NeRF structure, it will limit the flexibility of the generation task.

[0070] In view of the limitations of the method of adding watermarks to NeRFs in related technologies, the embodiments of this application propose a technical concept of embedding watermarks during the generation process without modifying the NeRF structure, that is, watermarks are embedded in NeRFs when they are generated.

[0071] Therefore, the embodiments of the present application propose an endogenous watermark addition method, device, equipment, and storage medium for neural radiance fields. By adding watermarks during the generation of text-to-3D (also known as the generative watermarking scheme), the embodiments of the present application are combined with 3D generation architectures such as fractional distillation sampling, and can endogenously watermark during the process of training to generate NeRF, so as to obtain a high-quality NeRF with built-in watermarks. Different from the post-generation NeRF watermarking method, the embodiments of the present application can directly obtain the target NeRF with the watermark added based on the endogenous watermarking method during the process of generating NeRF, and there is no need to change the NeRF architecture, thus increasing the flexibility for future further development tasks on 3D generation (target NeRF).

[0072] In order to inject a backdoor into NeRF during the generation process, the embodiments of the present application first generate a set of trigger viewpoints that only depend on a given secret message (watermark information), and then perform fractional distillation sampling so that the secret message can be extracted from the images rendered from any trigger viewpoint. In order to extract the secret message from the rendered images, the embodiments of the present application use a pre-trained watermark decoder from HiDDeN. All NeRFs generated by the scheme proposed by the embodiments of the present application can verify the watermark through this pre-trained decoder.

[0073] Based on the overall concept of the above embodiments of the present application, specific embodiments of the endogenous watermark addition method, device, computer equipment, computer-readable storage medium, and computer program product for neural radiance fields provided by the embodiments of the present application are proposed. First, each specific embodiment of the endogenous watermark addition method for neural radiance fields in the embodiments of the present application is described in detail.

[0074] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0075] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0076] It should be noted that in each specific embodiment of this application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.

[0077] In addition, the method for adding endogenous watermarks to neural radiance fields provided by the embodiments of this application relates to the field of computer technology. The method for adding endogenous watermarks to neural radiance fields provided by the embodiments of this application can be applied to terminals, can also be applied to server sides, or can also be software running on terminals or server sides. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application for implementing the method for adding endogenous watermarks to neural radiance fields, etc., but is not limited to the above forms.

[0078] Alternatively, the embodiments of this application can also be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0079] For the convenience of understanding and elaboration, in the following text, the method for adding an endogenous watermark to a neural radiance field provided by the embodiments of the present application will be taken as an example for detailed description. The implementation of the method for adding an endogenous watermark to a neural radiance field provided by the embodiments of the present application for any of the above forms of the subject can refer to the process of the method for adding an endogenous watermark to a neural radiance field by a terminal device described later.

[0080] Please refer to Figure 1 , Figure 1 which is a schematic flow chart of the steps of the method for adding an endogenous watermark to a neural radiance field provided by the embodiments of the present application in some embodiments. It should be understood that although Figure 1 shows the execution order of some method steps, based on different design requirements of actual applications, the method for adding an endogenous watermark to a neural radiance field provided by the embodiments of the present application can of course adopt an execution order different from that shown in the figure. That is, Figure 1 the order of the method steps shown does not constitute a limitation on the execution logic order of the method for adding an endogenous watermark to a neural radiance field provided by the embodiments of the present application, and any reasonable change based on Figure 1 the order of the method steps shown should be included in the protection scope of the method for adding an endogenous watermark to a neural radiance field provided by the embodiments of the present application.

[0081] As Figure 1 shown, in some embodiments, the method for adding an endogenous watermark to a neural radiance field provided by the embodiments of the present application may include but is not limited to steps S101 to S102.

[0082] Step S101: Generate a set of trigger viewpoints according to the given watermark information.

[0083] During the training stage of the NeRF by the terminal device, a set of trigger viewpoints that only depends on the given secret message (watermark information) is randomly generated.

[0084] In some embodiments, when the given watermark information is the secret message m, the terminal device can generate a set of trigger viewpoints from the given secret message m In this way, a copy of the set of trigger viewpoints does not need to be retained when verifying the watermark of the generated target NeRF later.

[0085] In some embodiments, considering that a constant set of trigger viewpoints is easy to predict, which may lead to potential vulnerabilities in protecting the ownership of the finally generated target NeRF, the terminal device can generate different sets of trigger viewpoints for different secret messages.

[0086] In some embodiments, the terminal device may use a pseudo-random number generator (PRNG) to generate a set of trigger perspectives that depends on the secret message m. Based on this, step S101 above: generating a set of trigger perspectives according to the given watermark information may include:

[0087] Obtain the hash value of the given watermark information;

[0088] Initialize the pseudo-random number generator based on the hash value;

[0089] Generate a set of trigger perspectives based on the initialized pseudo-random number generator.

[0090] For a given watermark information (secret message), the terminal device first uses the hash algorithm SHA256 to obtain the hash value of the secret message m. Subsequently, the terminal device uses the hash value of the secret message m to initialize the pseudo-random number generator. Finally, the terminal device uses the initialized pseudo-random number generator to generate a set of trigger perspectives that depends on the secret message m.

[0091] In some embodiments, when the terminal device randomly generates a set of trigger perspectives according to the given watermark information, it may establish a NeRF network.

[0092] Exemplarily, the network structure of the NeRF established by the terminal device may at least include: MLP f c and f σ . Both of these two MLPs are composed of a single layer of fully connected layers, and the channel size of the fully connected layer is 64. Among them, the input of f c is the 3D coordinate x ∈ [0, 1] N×3 , and the output is the color c ∈ [0, 1] of the current 3D coordinate N×3 ; the input of f σ is the 3D coordinate x ∈ [0, 1] N×3 , and the output is the geometric density of the current 3D coordinate

[0093] Step S102: Iteratively train the network of the neural radiance field NeRF based on a randomly selected target perspective to obtain a target NeRF with a watermark added; wherein, at least one iteration training in the iterative training process performs the following operations: when the target perspective belongs to the set of trigger perspectives, calculate the score distillation sampling loss of the rendered image of the target perspective, and calculate the binary cross-entropy function between the decoded watermark information of the rendered image and the given watermark information; wherein, the rendered image is obtained from the NeRF, and the decoded watermark information is obtained by parsing the rendered image based on a preset watermark decoder; and optimize the network parameters of the NeRF based on the score distillation sampling loss and the binary cross-entropy function.

[0094] After the terminal device constructs the NeRF network structure, it can directly perform the operation of iteratively training the NeRF network. Among them, when the terminal device iteratively trains the NeRF network each time, it randomly selects a view from the trigger view set as the target view, or selects a random view that does not belong to the trigger view set as the target view, then obtains the rendered image of the target view from NeRF, and then uses score distillation sampling to train NeRF to generate the target NeRF.

[0095] During the process of the terminal device iteratively training the NeRF network, if the target view selected in the current iterative training is a view in the trigger view set, the terminal device calculates the score distillation sampling loss of the rendered image of the target view and calculates the binary cross-entropy function of the decoded watermark information of the rendered image and the given watermark information during the current iterative training, so as to update the network parameters of NeRF by backpropagation based on the score distillation sampling loss and the binary cross-entropy function. In this way, after the terminal device completes the NeRF training based on score distillation sampling, it can finally generate the target NeRF with a watermark added.

[0096] It should be noted that during the process of the terminal device iteratively training the NeRF network, when the target view is a view in the trigger view set, the terminal device can parse the rendered image of the target view through a pre-trained watermark decoder to obtain the decoded watermark information, and then calculate the binary cross-entropy function of the decoded watermark and the original watermark information (the previously given watermark information).

[0097] In some embodiments, the terminal device can obtain the rendered image of the target view from NeRF through the following rendering process g(θ,p).

[0098] The input of the rendering process is the view the camera intrinsic matrix K and the rendered image resolution size H×W. Among them, the view p is composed of the rotation matrix and the translation vector which respectively represent the view orientation and view position in the world coordinate. The view p can transform the camera coordinate x c into the world coordinate x w = p r x c + p t . The intrinsic matrix can transform the camera coordinate x c into the image coordinate x i = Kx c . H×W image pixels require H×W rays. The terminal device first generates a ray set Then transform the ray set to the camera coordinates Finally, transform the ray set to the world coordinates Thus, since there is already a ray set, the terminal device can calculate the color of each ray according to the following formula:

[0099]

[0100] In some embodiments, when the target view is a view in the trigger view set, when optimizing the network parameters of NeRF based on the score distillation sampling loss and the binary cross-entropy function, the terminal device can optimize the network parameters by minimizing the score distillation sampling loss and the binary cross-entropy function:

[0101]

[0102] where W D is a pre-trained watermark decoder.

[0103] It should be noted that the terminal device needs to lock the parameters of the watermark decoder W D at this stage, so that any target NeRF obtained by optimizing the network parameters from the above formula can be decoded for watermark information by the same watermark decoder W D

[0104] In some embodiments, during the iterative training of the NeRF network by the terminal device, the following operations are also performed for at least one iterative training during the iterative training process:

[0105] When the target view does not belong to the trigger view set, calculate the score distillation sampling loss of the rendered image of the target view; and optimize the network parameters of NeRF based on the score distillation sampling loss.

[0106] During the iterative training of the NeRF network by the terminal device, if the target view selected in the current iterative training is not a view in the trigger view set but a random view outside the trigger view set, the terminal device only calculates the score distillation sampling loss of the rendered image of the target view during the current iterative training and directly updates the network parameters of NeRF based on the score distillation sampling loss through backpropagation.

[0107] ​In the embodiments of the present application, the terminal device first generates a set of trigger perspectives that only depend on a given secret message (watermark information), and establishes a NeRF network. Then, during the iterative training of the NeRF network, the NeRF network is iteratively trained based on randomly selected target perspectives. When the target perspective belongs to the set of trigger perspectives, the score distillation sampling loss of the rendered image of the target perspective is calculated, and the binary cross-entropy function between the decoded watermark information of the rendered image and the given watermark information is calculated, so as to optimize the network parameters of NeRF based on the score distillation sampling loss and the binary cross-entropy function. Thus, a watermark is added to NeRF during the score distillation sampling process, and the target NeRF with the watermark added is directly generated.

[0108] In this way, the embodiments of the present application achieve adding a watermark during the generation of text-to-3D. Combined with the 3D generation architecture of the score distillation sampling type, it is possible to endogenously add a watermark during the process of training and generating NeRF, so as to obtain a high-quality target NeRF with a built-in watermark, realizing good ownership protection for NeRF. Moreover, compared with the method of adding a watermark to NeRF after generation, the embodiments of the present application directly generate the target NeRF with a watermark and do not need to change the NeRF architecture, thus increasing the flexibility for further development in 3D generation in the future.

[0109] In addition, compared with the traditional method of separating the NeRF generation stage and the watermark stage, the method of adding a 3D watermark during the generation process in the embodiments of the present application eliminates the delay between NeRF generation and watermark addition, ensuring that no NeRF version without a watermark is generated, thus effectively preventing NeRF from being stolen. And, in the embodiments of the present application, the terminal device optimizes the network parameters of NeRF based on the score distillation sampling loss and the binary cross-entropy function, so as to inject a backdoor during the score distillation sampling to endogenously add a watermark during the process of generating NeRF, and then complete the process of adding a watermark to NeRF, which can ensure that the secret information can be extracted from the images rendered from any trigger perspective for verification.

[0110] In some embodiments, the terminal device may pre-train a common watermark decoder before calculating the binary cross-entropy function between the decoded watermark information of the rendered image and the given watermark information, for extracting the hidden watermark information from the rendered images of NeRF.

[0111] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the step flow of the method for adding endogenous watermark to the neural radiance field provided by the embodiments of the present application in other embodiments.

[0112] As Figure 2As shown, in some embodiments, the method for adding an endogenous watermark to a neural radiance field provided by the embodiments of the present application may further include steps S201 and S202 as shown below.

[0113] Step S201: Obtain picture data and original watermark information.

[0114] During the stage of training the watermark decoder by the terminal device, first obtain the picture data for pre-training the watermark decoder, and obtain the original watermark information to be added to the picture.

[0115] It should be noted that the image data contains four dimensions (B, C, H, W), where B is the batch size, C is the number of picture channels, and H and W are the height and width of the picture. In addition, the terminal device records the length of the original watermark information added to the picture as L.

[0116] Step S202: Train a preset deep neural network model based on the picture data and the original watermark information to obtain the watermark decoder.

[0117] It should be noted that the preset deep neural network model can be an end-to-end deep neural network model. The terminal device can establish the deep neural network model before obtaining the picture data and the original watermark information.

[0118] After the terminal device obtains the picture data and the original watermark information for pre-training the watermark decoder, use the picture data and the original watermark information as model inputs to train the previously established deep neural network model until it converges, and then use the converged deep neural network model as the watermark decoder for parsing the rendered picture to obtain the decoded watermark information.

[0119] In some embodiments, the end-to-end deep neural network model established by the terminal device includes an encoder, a noise layer, and a decoder. Among them, the encoder receives two parts of input. The first part is the picture data I or , and the second part is the original watermark information m. The encoder outputs the watermark picture I en ; the noise layer randomly adds different noises to the watermark picture I en to obtain the noisy watermark picture I no ; the decoder receives the noisy watermark picture I no , and outputs the hidden watermark information m ′ .

[0120] Please refer to Figure 3 , Figure 3 For Figure 2 the detailed step flow schematic diagram of step S202 in

[0121] As Figure 3As shown, in some embodiments, the above step S202: training a preset deep neural network model based on the picture data and the original watermark information to obtain the watermark decoder may include steps S301 to S304 as shown below.

[0122] Step S301: Input the picture data and the original watermark information into the encoder to obtain a watermarked picture.

[0123] During the process of the terminal device training the deep neural network model based on the picture data and the original watermark information, first, the picture data I or and the original watermark information are input into the encoder. The encoder will preprocess the picture data I or using a convolutional layer, then expand the dimension of the original watermark information and repeat the same information on the two expanded dimensions to obtain preprocessed watermark data of size (B, L, H, W). And the encoder combines the two parts of preprocessed data with the picture data I or in the channel dimension and obtains a watermark residual picture I r of size (B, C, H, W) through convolution. Finally, the encoder multiplies the watermark residual picture by a certain ratio and adds it to the original picture data I or to obtain the watermarked picture I en .

[0124] Step S302: Add noise to the watermarked picture based on the noise layer to obtain a noisy watermarked picture.

[0125] After the terminal device obtains the watermarked picture I no through the encoder, it further adds noise to the watermarked picture I no based on the noise layer to obtain a noisy watermarked picture I no .

[0126] It should be noted that the noise added by the terminal device to the watermarked picture based on the noise layer is divided into two parts: differentiable noise and non-differentiable noise. Among them, differentiable noise refers to the noise that does not affect the gradient transmission between deep neural networks after being added, and non-differentiable noise refers to the noise that the gradient transmission will be disconnected at the noise layer after being added. For differentiable noise, the terminal device can directly add it to the deep neural network for training. For non-differentiable noise, the terminal device can calculate the residual picture of the noise and add it to the watermarked picture in a gradient-free manner.

[0127] In some embodiments, when the noise to be added (the noise to be added to the watermarked picture) belongs to non-differentiable noise (non-differentiable noise), the above step S302: adding noise to the watermarked picture based on the noise layer to obtain a noisy watermarked picture may include the following steps:

[0128] Calculate the residual of the noise to be added and remove the gradient to obtain the residual calculation result;

[0129] Based on the noise layer, add the residual calculation result to the watermark image to obtain a noise watermark image.

[0130] For the watermark image I to be added en For the noise to be added that belongs to non-differentiable noise, the terminal device first calculates the image I' after adding real noise to the watermark image I en Then, calculate the residual of the noise to be added and remove the gradient I no = NoGrad(I' diff - I no ) Then, based on the noise layer, add the residual back to the watermark image to obtain the noise watermark image I en = I no + I en diff .

[0131] In this way, the terminal device can prevent non-differentiable noise from participating in gradient transmission, enabling the gradient of the deep neural network to be normally backpropagated.

[0132] Step S303: Input the noise watermark image into the decoder to obtain the parsed watermark information.

[0133] After the terminal device obtains the noise watermark image I no , it receives the noise watermark image I through the decoder no . The decoder first performs multi-layer convolutional layer processing on the watermark noise image I no to obtain watermark information data of size (B, L, H, W), and then through the global average layer and a fully connected layer, processes the watermark information data into hidden watermark information m of size (B, L) ′ (parsed watermark information).

[0134] Step S304: Update the deep neural network model based on the binary cross-entropy function of the parsed watermark information and the original watermark information to obtain the watermark decoder.

[0135] During the training process of the deep neural network model by the terminal device, after obtaining the parsed watermark information based on the decoder, calculate the binary cross-entropy function (Binary CrossEntropy, BCE) of the parsed watermark information and the original watermark information, and thus use this binary cross-entropy function BCE as the loss function to update the deep neural network model. In this way, by repeatedly executing the above training process from the encoder to the decoder until the neural network converges, the above watermark decoder is obtained.

[0136] ​It should be noted that different from the watermarking method which focuses on two objectives: invisibility and robustness. The terminal device only focuses on the robustness metric during the stage of training the watermark decoder, because after obtaining the pre-trained watermark decoder, the terminal device will discard this part of the watermark encoder. Therefore, the terminal device uses the binary cross-entropy function BCE as the loss function:

[0137]

[0138] In this embodiment, the terminal device first pre-trains a common watermark decoder to extract the hidden watermark information from the rendered images of NeRF. Then, during the iterative training of NeRF, watermarks are endogenously added during the score distillation sampling process to add watermarks to NeRF, so that the target NeRF generated by the score distillation sampling satisfies: for any image rendered from a trigger view, the pre-trained watermark decoder can extract the hidden watermark information from it.

[0139] Next, a complete embodiment of the method for adding endogenous watermarks to a neural radiance field proposed in the embodiments of the present application is presented.

[0140] Please refer to Figure 4 , Figure 4 which is a schematic diagram of the complete watermarking process involved in the method for adding endogenous watermarks to a neural radiance field proposed in the embodiments of the present application.

[0141] As Figure 4 shown, in a complete embodiment of the method for adding endogenous watermarks to a neural radiance field proposed in the embodiments of the present application, the terminal device implements the steps of adding watermarks to NeRF in two stages. The first stage is (a) pre-training a common watermark decoder, and the watermark decoder obtained in this stage can be used to extract the hidden watermark information from the rendered images of NeRF. The second stage is (b) adding watermarks by score distillation sampling, and through the operations in this stage, the target NeRF generated by the score distillation sampling can be made to satisfy: for any image rendered from a trigger view, the pre-trained watermark decoder can extract the hidden watermark information from it. Among them, in the first stage, the terminal device first pre-trains the watermark encoder W E to embed the watermark into the image, and then trains the watermark decoder W D to decode the watermark from the image. In addition, in the second stage, the terminal device generates a set of trigger views from the given secret message m and optimizes NeRF during the subsequent score distillation sampling, so that the secret message can be decoded from any image rendered from the trigger view rendered.

[0142] Please refer to Figure 5 , Figure 5Schematic diagram of the logical process involved in the training watermark decoder phase for the method of adding endogenous watermarks to neural radiance fields proposed in the embodiments of the present application.

[0143] It should be noted that the picture data I used for pre-training the watermark decoder or contains four dimensions (B, C, H, W), where B is the batch size, C is the number of picture channels, and H and W are the height and width of the picture. Here, the terminal device records the length of the watermark information added to the picture as L. The terminal device first establishes an end-to-end deep neural network model, and its structure is divided into three parts: an encoder noise layer and a decoder The encoder receives two parts of input. The first part is the picture data I or , and the second part is the watermark information m, and outputs the watermark picture I en . The noise layer randomly adds different noises to the watermark picture. The decoder receives the noisy watermark picture I no , and outputs the hidden watermark information m ′ .

[0144] As Figure 5 shown, when the terminal device pre-trains the watermark decoder, it first inputs the original picture and the watermark information into the encoder to obtain the watermark picture. At this time, the encoder will preprocess the picture information I or using a convolutional layer, then expand the dimension of the watermark information and repeat the same information on the two expanded dimensions to obtain preprocessed watermark data of size (B, L, H, W). Combine the two parts of preprocessed data with I or in the channel dimension, and obtain the watermark residual picture I of size (B, C, H, W) through convolution r , and add the watermark residual picture multiplied by a certain ratio to the original picture to obtain the watermark picture I en .

[0145] After that, the terminal device adds noise to the watermark picture. At this time, the terminal device needs to determine whether the noise is differentiable. If the judgment is yes, the terminal device directly adds the noise. If the judgment is no, the terminal device calculates the residual picture of the noise and adds it to the watermark picture in a gradient-free manner. For differentiable noise, the terminal device can directly add it to the deep neural network for training. For non-differentiable noise, the terminal device will be processed in steps 1 to 3: 1. Calculate the picture I' after adding real noise to the watermark picture I en , 2. Calculate the residual of the noise and remove the gradient I no =NoGrad(I' diff -I no ), 3. Add the residual back to the watermark picture I en no= I en + I diff In this way, the non-differentiable noise of the terminal device does not participate in the gradient transmission, so that the gradient of the deep neural network can be normally backpropagated.

[0146] After the terminal device adds noise to the watermark image to obtain the noisy watermark image, it further inputs the noisy watermark image into the decoder to obtain the parsed watermark information. Among them, the decoder first performs multi-layer convolutional layer processing on the noisy watermark image to obtain watermark information data of size (B, L, H, W), and then through the global average layer and a fully connected layer, processes the watermark information data into hidden watermark information m' of size (B, L).

[0147] Finally, the terminal device updates the network through backpropagation by calculating the binary cross-entropy loss function of the parsed watermark information and the original watermark information. In this way, the terminal device repeats the above steps 1 to 3 until the neural network converges, thereby obtaining the watermark decoder.

[0148] Different from the watermark method that focuses on two objectives, invisibility and robustness, at this stage, the terminal device only focuses on the robustness index, because after obtaining the pre-trained watermark decoder , this part of the watermark encoder ε will be discarded. The terminal device uses the binary cross-entropy function as the loss function:

[0149]

[0150] Please refer to Figure 6 , Figure 6 which is a schematic logical flow diagram of the watermark addition stage involved in the method for adding an endogenous watermark to a neural radiance field proposed in an embodiment of this application in some embodiments.

[0151] As Figure 6 shown, when the terminal device generates the target NeRF based on score distillation sampling, it first randomly generates a trigger view set of size N according to the given watermark information and establishes a NeRF network.

[0152] Among them, the network structure of the NeRF constructed by the terminal device consists of two MLPs f c and f σ . Both MLPs are composed of a single layer of fully connected layers, and the channel size of the fully connected layer is 64. Among them, the input of f c is the 3D coordinate x ∈ [0, 1] N×3 , and the output is the color c ∈ [0, 1] of the current 3D coordinate N×3 ; the input of f σ is the 3D coordinate x ∈ [0, 1] N×3 , and the output is the geometric density of the current 3D coordinate

[0153] In addition, when the terminal device generates a set of trigger perspectives, the terminal device hopes to generate a set of trigger perspectives from the secret message m. In this way, there is no need to keep a copy of the set of trigger perspectives during watermark verification. In addition, different messages should generate different sets of perspectives because a constant set of trigger perspectives is easy to predict, leading to potential vulnerabilities. Therefore, the terminal device uses a pseudo-random number generator (PRNG) to generate a set of trigger perspectives that depends on m. That is: based on the given secret message m, the terminal device first uses the hash algorithm SHA256 to obtain the hash value of m, then uses the hash value of m to initialize the pseudo-random number generator, and finally uses the pseudo-random generator to generate the set of trigger perspectives.

[0154] After the terminal device generates the set of trigger perspectives and establishes the NeRF network, it randomly selects one perspective from the set of trigger perspectives or a random perspective for each training iteration. Then, the terminal device determines whether the currently selected perspective belongs to the set of trigger perspectives, and in both cases where the determination is yes and where the determination is no, it obtains the rendered image of the current perspective from NeRF. However, in the case where it is determined that the currently selected perspective belongs to the set of trigger perspectives, the terminal device calculates the score distillation sampling loss of the rendered image and the two-dimensional cross-entropy loss function between the watermark information after watermark parsing and the original watermark information, and updates the NeRF parameters through backpropagation. In the case where it is determined that the currently selected perspective does not belong to the set of trigger perspectives, the terminal device calculates the score distillation sampling loss of the rendered image and updates the NeRF parameters through backpropagation.

[0155] Among them, when the terminal device obtains the rendered image of the current perspective from NeRF, the rendering process that can be adopted is as follows:

[0156] The input of the rendering process is the perspective the camera intrinsic matrix K and the size of the rendered image resolution H×W. Among them, the perspective p is composed of the rotation matrix and the translation vector which respectively represent the perspective orientation and perspective position in the world coordinate system. The perspective p can transform the camera coordinate x c into the world coordinate x w = p r x c + p t . The intrinsic matrix can transform the camera coordinate x c into the image coordinate x i = Kx c . H×W image pixels require H×W rays. The terminal device first generates a ray set Then transform the ray set into the camera coordinate system Finally, transform the ray set into the world coordinate system Thus, since there is already a ray set, the terminal device can calculate the color of each ray according to the following formula:

[0157]

[0158] In addition, when the current view does not belong to the set of trigger views, the terminal device optimizes the network parameters by minimizing the loss function of fractional distillation sampling:

[0159]

[0160] If the current view belongs to the set of trigger views, the terminal device optimizes the network parameters by minimizing the fractional distillation sampling loss and the binary cross-entropy function:

[0161]

[0162] where, W D is a pre-trained watermark decoder. The terminal device needs to lock the parameters of W D so that any target NeRF optimized by the above formula can be decoded by the same W D to decode the watermark.

[0163] To verify the effectiveness of the proposed method for adding endogenous watermarks to neural radiance fields in this application embodiment, the terminal device selects two works on post-generated watermarks for NeRF, namely CopyRNeRF[1] and WateNeRF[2], to compare with an embodiment of the proposed method for adding endogenous watermarks to neural radiance fields in this application embodiment. The watermark embedding processes of CopyRNeRF[1] and WateNeRF[2] are as follows: First, generate NeRF using fractional distillation sampling, and then add watermarks to it using the post-generated watermark work.

[0164] Since the two key evaluations of the watermark algorithm are invisibility and robustness, the terminal device uses bit accuracy to evaluate the robustness under various image distortions (such as Gaussian noise, rotation, scaling, Gaussian blur, cropping, and brightness adjustment). For the evaluation of invisibility, different from the previous post-generated watermark algorithms, because in the generation context, there is no concept of the original NeRF. Therefore, the typical evaluation metric peak signal-to-noise ratio (PSNR) is not applicable to evaluate the proposed method for adding endogenous watermarks to neural radiance fields in this application embodiment. To follow the comparison method of the previous 2D generative watermark algorithms, the terminal device uses CLIP-Score to evaluate the deviation introduced by the watermark algorithm.

[0165] Thus, the comparison results of each specific index are shown in Table 1 and Table 2 below. Among them, Table 1 below is the comparison of bit precision and CLIP score with the post-generation method, and Table 2 below is the comparison of robustness with the post-generation method. In Table 1, "None" reports the performance when no watermark is applied, so bit precision is not applicable in this row. In addition, in Table 1 and Table 2, "ours" reports the performance when watermarking is performed using an embodiment of the endogenous watermark addition method for the neural radiance field proposed in the embodiments of the present application.

[0166] Method Bit accuracy (%) CLIP / 16 CLIP / 32 None N / A 0.3156 0.2859 SDS+CopyRNeRF 91.16 0.3152 0.2831 SDS+WateRF 94.24 0.3164 0.2823 Ours 98.93 0.3218 0.2943

[0167] Table 1

[0168]

[0169] Table 2

[0170] Based on the same inventive concept, the embodiments of the present application also provide an endogenous watermark addition device for a neural radiance field, which can implement the above-mentioned endogenous watermark addition method for the neural radiance field.

[0171] Please refer to Figure 7 , the endogenous watermark addition device for a neural radiance field provided by the embodiments of the present application includes:

[0172] A trigger view generation module, configured to generate a set of trigger views according to the given watermark information;

[0173] A NeRF network training module, configured to iteratively train the network of the neural radiance field NeRF based on randomly selected target views to obtain a target NeRF with a watermark added;

[0174] Wherein, the NeRF network training module performs the following operations in at least one iteration training during the iterative training process: when the target view belongs to the set of trigger views, calculate the score distillation sampling loss of the rendered image of the target view, and calculate the binary cross-entropy function between the decoded watermark information of the rendered image and the given watermark information; wherein, the rendered image is obtained from the NeRF, and the decoded watermark information is obtained by parsing the rendered image based on a preset watermark decoder; and optimize the network parameters of the NeRF based on the score distillation sampling loss and the binary cross-entropy function to add a watermark to the NeRF during the generation of the NeRF.

[0175] In some embodiments, the endogenous watermark addition device for a neural radiance field provided by the embodiments of the present application further includes:

[0176] A watermark decoder training module, configured to obtain image data and original watermark information; and train a preset deep neural network model based on the image data and the original watermark information to obtain the watermark decoder.

[0177] In some embodiments, the deep neural network model includes an encoder, a noise layer, and a decoder; the watermark decoder training module is further configured to input the image data and the original watermark information into the encoder to obtain a watermarked image; add noise to the watermarked image based on the noise layer to obtain a noisy watermarked image; input the noisy watermarked image into the decoder to obtain the parsed watermark information; and update the deep neural network model based on the binary cross-entropy function of the parsed watermark information and the original watermark information to obtain the watermark decoder.

[0178] In some embodiments, the watermark decoder training module is further configured to calculate the residual of the noise to be added and remove the gradient to obtain a residual calculation result; wherein the noise to be added belongs to non-differentiable noise; and add the residual calculation result to the watermarked image based on the noise layer to obtain a noisy watermarked image.

[0179] In some embodiments, the trigger view generation module is further configured to obtain the hash value of the given watermark information; initialize a pseudo-random number generator based on the hash value; and generate a set of trigger views based on the initialized pseudo-random number generator.

[0180] In some embodiments, the NeRF network training module is further configured to calculate the score distillation sampling loss of the rendered image of the target view in the case that the target view does not belong to the set of trigger views; and optimize the network parameters of the NeRF based on the score distillation sampling loss.

[0181] It should be noted that the specific implementation of the endogenous watermark addition device of the neural radiance field provided in the embodiments of the present application is basically the same as the specific embodiments of the above-mentioned endogenous watermark addition method of the neural radiance field, and will not be elaborated here.

[0182] The embodiments of the present application further provide a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned endogenous watermark addition method of the neural radiance field. The computer device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0183] Please refer to Figure 8 , Figure 8 which schematically shows the hardware structure of a computer device in another embodiment. The computer device includes:

[0184] The processor 801 can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0185] The memory 802 can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802 and are called by the processor 801 to execute the method for adding an endogenous watermark to a neural radiance field in the embodiments of the present application;

[0186] The input / output interface 803 is used to implement information input and output;

[0187] The communication interface 804 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0188] The bus 805 transmits information between various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);

[0189] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 achieve communication connections with each other inside the device through the bus 805.

[0190] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the method for adding an endogenous watermark to a neural radiance field described above is implemented.

[0191] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0192] An embodiment of the present application also provides a computer program product. The computer program product stores a computer program, and when the computer program is executed by a processor, the above-mentioned method for adding an endogenous watermark to a neural radiance field is implemented.

[0193] The method for adding an endogenous watermark to a neural radiance field, the device for adding an endogenous watermark to a neural radiance field, the computer device, the computer-readable storage medium, and the computer program product provided by the embodiments of the present application generate a set of trigger viewpoints that only depend on a given secret message (watermark information), and then iteratively train the network of the neural radiance field (NeRF) based on randomly selected target viewpoints. When the target viewpoint belongs to the set of trigger viewpoints, calculate the score distillation sampling loss of the rendered image of the target viewpoint, and calculate the binary cross-entropy function between the decoded watermark information of the rendered image and the given watermark information, so as to optimize the network parameters of NeRF based on the score distillation sampling loss and the binary cross-entropy function. Among them, the rendered image is obtained from NeRF, and the decoded watermark information is obtained by parsing the rendered image based on a preset watermark decoder. Thus, a watermark is added to NeRF during the process of score distillation sampling, and the target NeRF with the added watermark is directly generated.

[0194] In this way, a scheme for endogenous watermarking in the process of text-to-3D (generating a 3D model from text) is realized. It is combined with a 3D generation architecture of the score distillation sampling type, and can add an endogenous watermark when training to generate NeRF, so as to obtain a high-quality target NeRF with a built-in watermark, realizing good ownership protection for NeRF. Moreover, compared with the method of adding a watermark to NeRF after generation, the method of adding a watermark to NeRF by endogenous watermarking in the embodiments of the present application can directly generate the target NeRF with a watermark, and there is no need to change the NeRF architecture, thereby increasing the flexibility for future further development on 3D generation (the target NeRF).

[0195] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0196] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0197] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0198] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0199] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0200] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one)" or a similar expression thereof refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0201] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0202] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0203] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0204] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0205] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.

Claims

1. A method for adding endogenous watermarks to neural radiation fields, characterized in that: The method comprises: Generate a trigger perspective set according to the given watermark information; The network of Neural Radiance Field (NeRF) is iteratively trained based on the randomly selected target viewpoint to obtain the target NeRF with watermark added. At least one iterative training in the iterative training process performs the following operations: when the target perspective belongs to the trigger perspective set, calculating the fractional distillation sampling loss of the rendered image of the target perspective, and calculating the binary cross entropy function of the decoded watermark information of the rendered image and the given watermark information; wherein the rendered image is obtained from the NeRF, and the decoded watermark information is obtained by parsing the rendered image based on a preset watermark decoder; and optimizing the network parameters of the NeRF based on the fractional distillation sampling loss and the binary cross entropy function.

2. The method according to claim 1, characterized in that The method further comprises: Get image data and original watermark information; A preset deep neural network model is trained based on the image data and the original watermark information to obtain the watermark decoder.

3. The method according to claim 2, characterized in that The deep neural network model includes an encoder, a noise layer and a decoder; The step of training a preset deep neural network model based on the image data and the original watermark information to obtain the watermark decoder includes: Inputting the picture data and the original watermark information into the encoder to obtain a watermarked picture; Adding noise to the watermark image based on the noise layer to obtain a noise watermark image; Inputting the noise watermark image into the decoder to obtain the parsed watermark information; The deep neural network model is updated based on the binary cross entropy function of the parsed watermark information and the original watermark information to obtain the watermark decoder.

4. The method according to claim 3, characterized in that The adding noise to the watermark image based on the noise layer to obtain a noise watermark image includes: Calculating the residual of the noise to be added and removing the gradient to obtain the residual calculation result; wherein the noise to be added is non-differentiable noise; The residual calculation result is added to the watermark image based on the noise layer to obtain a noise watermark image.

5. The method according to claim 1, characterized in that The generating a triggering view set according to given watermark information includes: Get the hash value of the given watermark information; Initialize a pseudo-random number generator based on the hash value; Generate a trigger perspective set based on the initialized pseudo-random number generator.

6. The method according to claim 1, characterized in that At least one iterative training in the iterative training process further performs the following operations: When the target perspective does not belong to the trigger perspective set, the fractional distillation sampling loss of the rendered image of the target perspective is calculated; and the network parameters of the NeRF are optimized based on the fractional distillation sampling loss.

7. A device for adding endogenous watermarks to a neural radiation field, characterized in that: The device comprises: A trigger view generation module, used to generate a trigger view set according to given watermark information; The NeRF network training module is used to iteratively train the network of the neural radiation field NeRF based on the randomly selected target viewpoint to obtain the target NeRF with a watermark added; Among them, the NeRF network training module performs the following operations in at least one iterative training in the iterative training process: when the target perspective belongs to the trigger perspective set, calculating the fractional distillation sampling loss of the rendered image of the target perspective, and calculating the binary cross entropy function of the decoded watermark information of the rendered image and the given watermark information; wherein the rendered image is obtained from the NeRF, and the decoded watermark information is obtained by parsing the rendered image based on a preset watermark decoder; and, optimizing the network parameters of the NeRF based on the fractional distillation sampling loss and the binary cross entropy function to add a watermark to the NeRF in the process of generating the NeRF.

8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method for adding an endogenous watermark to a neural radiation field as described in any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for adding an endogenous watermark to a neural radiation field according to any one of claims 1 to 6 is implemented.

10. A computer program product, wherein the computer program product stores a computer program, characterized in that: When the computer program is executed by a processor, the method for adding an endogenous watermark to a neural radiation field according to any one of claims 1 to 6 is implemented.