Method and apparatus for extracting watermark

By identifying the target file type and using the corresponding decoder combined with the lighting model to extract the watermark, the problem of neural network extracting watermarks after the three-dimensional model is transformed into two-dimensional data is solved, which improves the extraction success rate and accuracy and enhances robustness.

WO2025161328A1PCT designated stage Publication Date: 2025-08-07HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2024/109946
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2024-08-06
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing neural networks are difficult to effectively extract watermarks after being transformed from a three-dimensional model to two-dimensional data. Attacks on attacks on network simulations are usually targeted at three-dimensional models, and it is difficult to cope with the changes in watermark shape when the three-dimensional model is transformed into two-dimensional data.

Method used

By identifying the type of the target file, the watermark is extracted using the corresponding three-dimensional model decoder or two-dimensional data decoder, and the watermark is extracted in a targeted manner to enhance robustness.

Benefits of technology

It improves the success rate and accuracy of extracting watermarks from the attacked three-dimensional model and its rendered two-dimensional data, and enhances its resistance to different attack methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109946_07082025_PF_FP_ABST
    Figure CN2024109946_07082025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of computers, and in particular to a method and apparatus for extracting a watermark. Embedding a digital watermark in a three-dimensional model can reduce the risk of illegitimate use of the three-dimensional model. However, a three-dimensional model may be converted into two-dimensional data, and the attacks simulated by a neural network during training are usually attacks on the three-dimensional model. Thus, when the three-dimensional model is converted into the two-dimensional data, the shape of a watermark is difficult to predict, and a neural network trained on the basis of the described attack network is difficult to extract an effective watermark from the two-dimensional data. In the present application, an electronic device acquires a target file, and then on the basis of the type of the target file, uses a corresponding decoder in a neural network to extract a watermark, for example, using a three-dimensional model decoder to process a three-dimensional model file, and using a two-dimensional data decoder to process a two-dimensional data file, thereby eliminating in a targeted manner damage to the watermark caused by attack means, and improving the success rate of extracting a watermark from an attacked three-dimensional model.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for extracting watermark

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on February 4, 2024, with application number 202410160130.5 and application name “Watermark generation method, device, computing device cluster and storage medium”, and the Chinese patent application filed with the State Intellectual Property Office on April 30, 2024, with application number 202410559391.4 and application name “Method and device for extracting watermarks”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a method and device for extracting a watermark. Background Art

[0003] A 3D digital model, often referred to as a 3D model, is a digital representation of the form represented in drawings of engineering or product designs, created using 3D modeling software. Compared to 2D data (such as images), 3D models visualize and visualize abstract spatial information, providing a richer display space.

[0004] A digital watermark, also known as a watermark, is information that can be embedded in a digital signal, such as audio, images, video, or a data structure. One important application of digital watermarks is copyright protection. For example, embedding a digital watermark in a 3D model or 2D data can reduce the risk of unauthorized use of the data.

[0005] Once publicly available, 3D models may be exposed to various attacks, such as geometric transformations, affine transformations, retopology, and mesh simplification. One approach to extracting watermarks is through neural networks. During neural network training, an attack network is added to simulate various potential attacks on 3D models, hoping to extract a valid watermark from the attacked 3D model. However, 3D models can be converted into 2D data, and the attacks simulated by the attack network are typically targeted at 3D models. When 3D models are converted into 2D data, the watermark's morphology is difficult to predict. Therefore, neural networks trained using this attack network struggle to extract a valid watermark from 2D data.

[0006] Summary of the Invention

[0007] Embodiments of the present application provide a method, apparatus, computer-readable storage medium, and computer program product for extracting watermarks, as well as a method, apparatus, computer-readable storage medium, and computer program product for training neural networks, which can improve the success rate of extracting watermarks from attacked three-dimensional models.

[0008] In a first aspect, an embodiment of the present application provides a method for extracting a watermark, wherein the execution subject of the method can be an electronic device or a chip applied to an electronic device. The following description is based on the execution subject being an electronic device. The method comprises: obtaining a target file, wherein the target file is a target three-dimensional model or target two-dimensional data, wherein the target two-dimensional data is an image or video generated by rendering based on a basic three-dimensional model and a two-dimensional map, wherein the target three-dimensional model carries a target watermark, wherein the target two-dimensional data carries a target watermark, wherein the target watermark carried by the target two-dimensional data is a watermark carried by the basic three-dimensional model and / or the two-dimensional map; determining the type of the target file, wherein the type includes a three-dimensional model or two-dimensional data; when the type is a three-dimensional model, extracting the target watermark carried by the target file through a three-dimensional model decoder, wherein the three-dimensional model decoder is a sub-network of a neural network; or, when the type is two-dimensional data, extracting the target watermark carried by the target file through at least one two-dimensional data decoder, wherein at least one two-dimensional data decoder is a sub-network of a neural network.

[0009] After obtaining the target file, the type of the target file is first identified, and then the corresponding decoder in the neural network is used to extract the watermark based on the type of the target file. The watermark can be specifically extracted from the attacked 3D model or its rendered 2D data, and the tampering of the data caused by the attack means can be identified, thereby improving the success rate of extracting watermarks from various types of attacked 3D models.

[0010] In an optional embodiment of the first aspect, before extracting the target watermark carried by the target file through at least one two-dimensional data decoder, the method also includes: determining at least one illumination model corresponding to the target file; determining at least one two-dimensional data decoder based on the at least one illumination model, and the at least one two-dimensional data decoder is used to extract the target watermark carried by the target file based on its corresponding illumination model.

[0011] Two-dimensional data, for example, is data obtained after rendering a basic three-dimensional model and a two-dimensional map. Different lighting models are used in the rendering process. By determining the lighting model used by the target file during the rendering process, it is possible to specifically analyze the impact of different lighting models on the target watermark, thereby improving the success rate and accuracy of extracting the target watermark from the target file.

[0012] In an optional embodiment of the first aspect, determining at least one illumination model corresponding to a target file includes: when illumination model indication information is received, determining at least one illumination model based on the illumination model indication information, the illumination model indication information being used to indicate the illumination model used when rendering to generate target two-dimensional data; or, when no illumination model indication information is received, and when watermark extraction indication information is received, determining at least one illumination model to be a preset plurality of illumination models, the illumination model indication information being used to indicate the illumination model used when rendering to generate target two-dimensional data, and the watermark extraction indication information being used to indicate extraction of a target watermark carried by the target file.

[0013] When the user determines the illumination model corresponding to the target file, extracting the watermark based on the illumination model indicated by the user can improve the accuracy of watermark extraction. When the user is unsure of the illumination model corresponding to the target file, extracting the watermark based on multiple preset illumination models can improve the success rate of watermark extraction.

[0014] In an optional embodiment of the first aspect, when at least one illumination model is a preset plurality of illumination models, at least one two-dimensional data decoder is a plurality of two-dimensional data decoders, and the plurality of two-dimensional data decoders correspond one-to-one to the preset plurality of illumination models. The target file is processed by the at least one two-dimensional data decoder, including: using the plurality of two-dimensional data decoders to process the target file respectively, and determining a plurality of intermediate watermarks; comparing the plurality of intermediate watermarks with the reference watermark respectively, and determining a plurality of similarities corresponding to the plurality of intermediate watermarks, the plurality of similarities being used to indicate the degree of similarity between the plurality of intermediate watermarks and the reference watermark; when at least one of the plurality of similarities is greater than or equal to a similarity threshold, determining a target watermark from the intermediate watermarks corresponding to the at least one similarity, wherein the target watermark is the intermediate watermark corresponding to the highest similarity value among the at least one similarity.

[0015] Using two-dimensional data decoders corresponding to multiple illumination models to extract watermarks can improve the success rate of watermark extraction when the user is not sure of the illumination model of the target file.

[0016] In an optional embodiment of the first aspect, before obtaining the target file, the method includes: obtaining an original watermark and an original three-dimensional model; encoding the original watermark and the original three-dimensional model through a three-dimensional model encoder to generate three-dimensional training data, where the three-dimensional training data is a three-dimensional model carrying the original watermark, and the three-dimensional model encoder is a subnetwork of a neural network; attacking the three-dimensional training data to generate attacked three-dimensional training data; and using the attacked three-dimensional training data to train a three-dimensional model decoder, where the three-dimensional model decoder is used to extract a first watermark from the attacked three-dimensional training data.

[0017] The above-mentioned attacks on three-dimensional training data can be implemented by a three-dimensional model attack library, which can simulate possible attacks on the target file. Training the neural network with the participation of the three-dimensional model attack library can improve the robustness of the three-dimensional model decoder.

[0018] In an optional embodiment of the first aspect, before obtaining the target file, the method includes: obtaining an original watermark, an original three-dimensional model, and an original two-dimensional map; encoding the original watermark and the original two-dimensional map through a two-dimensional data encoder to generate a two-dimensional training map, where the two-dimensional training map is a two-dimensional map carrying the original watermark, and the two-dimensional data encoder is a subnetwork of the neural network; attacking a three-dimensional model composed of the original three-dimensional model and the two-dimensional training map to generate an attacked three-dimensional model; rendering the attacked three-dimensional model using at least one differentiable renderer to obtain at least one two-dimensional training data, where the at least one two-dimensional training data is a picture or a video; and training at least one two-dimensional data decoder using the at least one two-dimensional training data, where the at least one two-dimensional data decoder is used to extract a second watermark from the at least one two-dimensional training data.

[0019] The above-mentioned attack on the three-dimensional model composed of the original three-dimensional model and the two-dimensional training map can be implemented by the model attack library. The model attack library can simulate the attacks that the target file may be subjected to. The model attack library can be implemented by the two-dimensional model attack library or the three-dimensional model attack library. The two-dimensional model attack library can attack the two-dimensional map in the three-dimensional model, change the color and texture of the two-dimensional map, etc. The three-dimensional model attack library can attack the original three-dimensional model. The above two attack methods can also be implemented by the same model attack library. Training the neural network with the participation of the two-dimensional model attack library can improve the robustness of at least one two-dimensional data decoder and enhance the decoder's watermark extraction ability for two-dimensional data under different attack methods.

[0020] In an optional embodiment of the first aspect, at least one two-dimensional model decoder is trained using at least one two-dimensional training data, including: when the training degree of the three-dimensional model decoder meets the requirements, training at least one two-dimensional data decoder through at least one two-dimensional training data.

[0021] The 3D model decoder is a relatively independent sub-network. Training the 3D model decoder first requires fewer neural network parameters to be adjusted, which enables the local neural network to converge as quickly as possible. In this way, when training the neural network using 3D training data and 2D training data, the overall convergence speed of the neural network can be accelerated.

[0022] In an optional implementation of the first aspect, the at least one two-dimensional data decoder includes multiple two-dimensional data decoders, the differentiable renderer includes multiple lighting models, and the multiple two-dimensional data decoders correspond one-to-one to the multiple lighting models.

[0023] Multiple two-dimensional data decoders correspond one-to-one to multiple illumination models in the differentiable renderer. When the user is not sure about the illumination model of the target file, multiple two-dimensional data decoders can be used to extract watermarks, thereby improving the success rate of watermark extraction.

[0024] In a second aspect, embodiments of the present application provide a device for extracting a watermark. The device may include an input module and a processing module, configured to execute any of the methods described in the first aspect and its optional embodiments, wherein the input module may execute the acquisition step under the control of the processing module.

[0025] The input module is used to: obtain a target file, where the target file is a target three-dimensional model or target two-dimensional data, where the target two-dimensional data is an image or video generated based on a basic three-dimensional model and two-dimensional map rendering, where the target three-dimensional model carries a target watermark, where the target two-dimensional data carries a target watermark, and where the target watermark carried by the target two-dimensional data is a watermark carried by the basic three-dimensional model and / or two-dimensional map; the processing module is used to: determine the type of the target file, where the type includes a three-dimensional model or two-dimensional data; when the type is a three-dimensional model, the processing module is further used to extract the target watermark carried by the target file through a three-dimensional model decoder, where the three-dimensional model decoder is a sub-network of a neural network; or, when the type is two-dimensional data, the processing module is further used to extract the target watermark carried by the target file through at least one two-dimensional data decoder, where at least one two-dimensional data decoder is a sub-network of a neural network.

[0026] In an optional embodiment of the second aspect, before extracting the target watermark carried by the target file through at least one two-dimensional data decoder, the processing module is also used to: determine at least one illumination model corresponding to the target file; determine at least one two-dimensional data decoder based on the at least one illumination model, and the at least one two-dimensional data decoder is used to extract the target watermark carried by the target two-dimensional data based on its corresponding illumination model.

[0027] In an optional embodiment of the second aspect, the processing module is specifically used to: when illumination model indication information is received, determine at least one illumination model based on the illumination model indication information, the illumination model indication information is used to indicate the illumination model used when rendering and generating target two-dimensional data; or, when no illumination model indication information is received, and when watermark extraction indication information is received, determine at least one illumination model to be a preset plurality of illumination models, the illumination model indication information is used to indicate the illumination model used when rendering and generating target two-dimensional data, and the watermark extraction indication information is used to indicate the extraction of the target watermark carried by the target file.

[0028] In an optional embodiment of the second aspect, when at least one illumination model is a preset plurality of illumination models, at least one two-dimensional data decoder is a plurality of two-dimensional data decoders, and the plurality of two-dimensional data decoders correspond one-to-one to the preset plurality of illumination models, and the processing module is specifically used to: use the plurality of two-dimensional data decoders to process the target file respectively to determine a plurality of intermediate watermarks; compare the plurality of intermediate watermarks with the reference watermark respectively to determine a plurality of similarities corresponding to the plurality of intermediate watermarks, and the plurality of similarities are used to indicate the degree of similarity between the plurality of intermediate watermarks and the reference watermark; when at least one of the plurality of similarities is greater than or equal to a similarity threshold, determine a target watermark from the intermediate watermarks corresponding to the at least one similarity, wherein the target watermark is the intermediate watermark corresponding to the highest similarity value among the at least one similarity.

[0029] In an optional implementation of the second aspect, before obtaining the target file, the input module is further used to: obtain the original watermark and the original three-dimensional model; the processing module is further used to: encode the original watermark and the original three-dimensional model through a three-dimensional model encoder to generate three-dimensional training data, where the three-dimensional training data is a three-dimensional model carrying the original watermark, and the three-dimensional model encoder is a sub-network of the neural network; attack the three-dimensional training data to generate attacked three-dimensional training data; and use the attacked three-dimensional training data to train a three-dimensional model decoder, where the three-dimensional model decoder is used to extract the first watermark from the attacked three-dimensional training data.

[0030] In an optional embodiment of the second aspect, before obtaining the target file, the input module is further used to: obtain the original watermark, the original three-dimensional model and the original two-dimensional map; the processing module is further used to: encode the original watermark and the original two-dimensional map through a two-dimensional data encoder to generate a two-dimensional training map, where the two-dimensional training map is a two-dimensional map carrying the original watermark, and the two-dimensional data encoder is a subnetwork of the neural network; attack the three-dimensional model composed of the original three-dimensional model and the two-dimensional training map to generate an attacked three-dimensional model; use at least one differentiable renderer to render the attacked three-dimensional model to obtain at least one two-dimensional training data, where the at least one two-dimensional training data is a picture or a video; use the at least one two-dimensional training data to train at least one two-dimensional data decoder, where the at least one two-dimensional data decoder is used to extract the second watermark from the at least one two-dimensional training data.

[0031] In an optional implementation of the second aspect, the processing module is specifically configured to: when the training degree of the three-dimensional model decoder meets the requirements, train at least one two-dimensional data decoder using at least one two-dimensional training data.

[0032] In an optional implementation of the second aspect, the at least one two-dimensional data decoder includes multiple two-dimensional data decoders, the differentiable renderer includes multiple lighting models, and the multiple two-dimensional data decoders correspond one-to-one to the multiple lighting models.

[0033] In a third aspect, embodiments of the present application provide a device for extracting a watermark, which may be an electronic device or a chip used in an electronic device. The device may include a processor configured to execute any of the methods in the first aspect and its optional embodiments.

[0034] Optionally, when the device is an electronic device, the processor is, for example, a system on chip (SoC) or a central processor unit (CPU); when the device is a chip, the processor is, for example, a core, which may include at least one execution unit, such as an arithmetic and logic unit (ALU).

[0035] Optionally, the device may further include a transceiver. When the device is an electronic device, the transceiver may be a transceiver circuit, an antenna, etc.; when the device is a chip, the transceiver may be an input / output interface, a pin, a circuit, etc.

[0036] Optionally, the device may further include a memory for storing a computer program or instructions, and the processor executes the computer program or instructions stored in the memory, so that the device performs any of the methods in the first aspect and its optional embodiments. When the device is an electronic device, the memory may be a read-only memory, a random access memory, or the like; when the device is a chip, the memory may be a register, a cache, or the like.

[0037] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed on a device for extracting a watermark, enables the device to execute: any one of the methods in the first aspect and its optional embodiments.

[0038] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions, the computer program product comprising: computer program code or computer program instructions, which, when the computer program code or computer program instructions are run by a device for extracting a watermark, enables the device to execute: any one of the methods in the first aspect and its optional embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] FIG1 is a schematic diagram of the architecture of a neural network provided in an embodiment of the present application;

[0040] FIG2 is a schematic diagram of a neural network training method provided in an embodiment of the present application;

[0041] FIG3 is a schematic diagram of a method for extracting watermarks based on a neural network according to an embodiment of the present application;

[0042] FIG4 is a schematic diagram of a scenario in which a three-dimensional model is attacked, provided in an embodiment of the present application;

[0043] FIG5 is a schematic diagram of a method for extracting a watermark provided in an embodiment of the present application;

[0044] FIG6 is a schematic diagram of a method for obtaining a target file provided in an embodiment of the present application;

[0045] FIG7 is a schematic diagram of a method for determining the type of a target file provided in an embodiment of the present application;

[0046] FIG8 is a schematic diagram of a method for determining a lighting model of a target file provided in an embodiment of the present application;

[0047] FIG9 is a schematic diagram of another method for determining a lighting model of a target file provided by an embodiment of the present application;

[0048] FIG10 is a schematic diagram of a method for training a neural network provided in an embodiment of the present application;

[0049] FIG11 is a schematic diagram of a reasoning process of a neural network provided in an embodiment of the present application;

[0050] FIG12 is a schematic structural diagram of a device for extracting watermarks provided in an embodiment of the present application;

[0051] FIG13 is a schematic structural diagram of a computer device for extracting watermarks provided in an embodiment of the present application;

[0052] FIG14 is a schematic diagram of the structure of a computer device cluster for extracting watermarks provided in an embodiment of the present application;

[0053] FIG15 is a schematic structural diagram of another computer device cluster for extracting watermarks provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] To facilitate understanding of the technical solution of this application, we first briefly introduce some terms and concepts involved in this application. It should be noted that the descriptions of these terms and concepts are examples rather than limitations.

[0055] 1. Neural network (NN).

[0056] A neural network, also known as an artificial neural network (ANN) or a neural network-like network, is a mathematical or computational model in the fields of machine learning and cognitive science that mimics the structure and function of biological neural networks and is used to estimate and approximate functions. Neural networks can include multilayer perceptrons (MLPs), convolutional neural networks (CNNs), and deep neural networks (DNNs).

[0057] 1.1. Architecture of neural network.

[0058] A neural network can include multiple neural network layers. As shown in Figure 1, the neural network 100 can include n neural network layers, where n is a positive integer, wherein the first layer is the input layer, the nth layer is the output layer, and the remaining layers are hidden layers. The arrows between the layers represent the direction of data transmission. In the embodiment of the present application, the neural network layer is a logical probability. A neural network layer can refer to a set of computing nodes corresponding to a neural network calculation. The computing nodes can be called neurons. In practical applications, the neural network layer can include convolutional layers, pooling layers, and fully connected layers.

[0059] The working principle of each neural network layer can be expressed mathematically To describe, among them, Represents a sample (i.e., input data). A sample generally has multiple attributes, so a sample is usually a vector; w represents the weight value of each neuron in the neural network; b is the bias; a is the activation function; Is the output vector. From a physical perspective, the work of each layer in a neural network can be understood as completing the transformation from input space to output space (i.e., a set of output vectors) through five operations on the input space (i.e., a set of input vectors), among which, Determines the transformation from input space to output space, that is, w of each layer controls how to transform the space. The above five operations include: 1. Dimensionality increase / decrease; 2. Zoom in / out; 3. Rotation; 4. Translation; 5. "Bending". Among them, operations 1, 2 and 3 are determined by Operation 4 is implemented by +b, and operation 5 is implemented by a().

[0060] 1.2. Training of neural networks.

[0061] Because the goal is to ensure that the output of a neural network is as close to the actual situation as possible, the difference between the predicted value and the target value can be determined. The neural network's parameters (e.g., weights) are then updated based on this difference. Before the first parameter update, an initialization process typically occurs, preconfiguring the parameters for each layer of the neural network. During training, if the predicted value is too high, the parameters are adjusted to lower the predicted value; if the predicted value is too low, the parameters are adjusted to increase the predicted value. This process continues until the neural network's output meets the requirements. Therefore, it is necessary to predefine how to compare the difference between the predicted value and the target value. Loss functions and objective functions are commonly used to measure the difference between the predicted value and the target value. For example, a larger loss function output (loss) indicates a greater difference, while a smaller loss function output indicates a smaller difference. Therefore, training a neural network becomes a process of minimizing this output value.

[0062] As shown in Figure 2, the starting point corresponds to the output value of the loss function when the neural network parameters have not been updated. As the neural network parameters are updated, the output value of the loss function gradually decreases. The optimal point corresponds to the minimum output value of the loss function. When the output value of the loss function reaches the minimum value, the neural network training is completed.

[0063] The loss function is usually a multivariable function. The gradient can reflect the rate of change of the output value of the loss function when the variable changes. The larger the absolute value of the gradient, the greater the rate of change of the output value of the loss function. The gradient of the loss function when updating different parameters can be calculated, and the parameters can be updated in the direction with the fastest gradient descent to reduce the output value of the loss function as quickly as possible.

[0064] The output value of the neural network's loss function can be obtained through forward propagation (FP). For example, the output of the previous layer of the neural network can be input into the next layer until the output of the output layer is obtained. The output value of the loss function is then calculated based on this output and the target value. Subsequently, back propagation (BP) calculation is performed based on the output value of the loss function to obtain the gradient of each layer. The neural network parameters are adjusted along the direction of the fastest gradient descent until the output value of the loss function reaches the minimum value.

[0065] 2. Three-dimensional model.

[0066] A 3D model is a polygonal representation of an object, typically displayed on a computer or other video device. 3D models can take a variety of shapes, including spheres, cones, curved and flat objects, and objects with irregular surfaces. The objects represented by a 3D model can be real-world entities or imaginary objects. 3D models can be created using specialized software such as 3D modeling tools, but other methods are also possible.

[0067] The basic structure of a 3D model is a mesh. A mesh can be composed of triangles, quadrilaterals, or other polygons, each of which is called a facet. The intersections of the edges of a facet are called vertices. Facets can be divided based on information such as the 3D model's material. A material is a collection of properties such as an object's appearance and optical characteristics, including its color, reflectivity (such as diffuse and specular reflection), transparency, and refractive index.

[0068] Materials define how objects interact with light, determining their appearance when rendered. For example, depending on the material, the interaction between a patch and light can be categorized into three types: refraction, specular reflection, and diffuse reflection. Light refracts when it hits a transparent material, reflects when it hits a smooth, opaque material, and diffusely reflects when it hits a rough, opaque material.

[0069] For example, a material can be expressed by a texture to simulate the properties of an object's surface, which may include color information, geometric information, process information, and the like.

[0070] Textures that contain color information are called color textures, which are used to express patterns, designs, and text on the surface of objects, such as marble walls, calligraphy and paintings on the walls, and patterns on utensils.

[0071] A texture containing geometric information is called a geometric texture, which is used to represent the microscopic geometric shape of an object's surface, such as an uneven surface.

[0072] Textures that contain process information are called process textures, which are used to express various regular or irregular dynamically changing scenes, such as water waves, clouds, fire, smoke, etc.

[0073] The process of applying a texture to a model's surface is called texture mapping. Textures mapped to a 3D model's surface are also called texture maps. Texture mapping matches the texture to the vertices or pixels of a 3D model, overlaying the texture on the model's surface. During rendering, the texture's coordinate information determines the model's surface color, geometry, and process. Mapping a texture to a model's surface can give the 3D model a more realistic and detailed appearance.

[0074] 3. Two-dimensional data.

[0075] Two-dimensional data includes images or videos, which can be obtained by rendering a three-dimensional model.

[0076] In rendering, different material parameters can be specified for different objects. If the same effect needs to be achieved at different locations on an object, the values ​​of all material parameters can be constants. If different effects need to be achieved at different locations on an object, textures can be used to specify material parameters for the object. There is a mapping relationship (i.e., UV unfolding) between the object and the texture, which is the unfolding of the solid geometric surface to the two-dimensional plane. U and V refer to the horizontal and vertical axes of the two-dimensional plane. Specifying different values ​​for different locations in the unfolded two-dimensional plane is equivalent to covering the surface of the object with different textures. Optionally, the texture can include color, surface unevenness, metalness, roughness, opacity, etc.

[0077] Color can be expressed through a color map. A color map can be a pixel feature component on a two-dimensional image, such as red, green, and blue (RGB), which is used to represent the color or refraction and absorption coefficient of a three-dimensional model surface.

[0078] The degree of surface unevenness can be expressed using normal maps and bump maps. Normal maps affect the orientation of the surface normals of a 3D model, including surface angle information. Bump maps are primarily used to express surface heights, potholes, and other aspects, creating an uneven surface effect.

[0079] Metalness can be expressed using a metalness map. A metalness map represents the specular properties of an object's surface. A metalness map can define whether a 3D model is metal or non-metal.

[0080] Roughness can be expressed through a roughness map. A roughness map represents the diffuse properties of an object's surface.

[0081] Transparency can be expressed through a transparency map. A transparency map represents the degree of transparency of an object's surface.

[0082] In addition to rendering, 3D models can also be processed by shading, ray tracing, and path tracing to make the final generated 2D data more realistic.

[0083] Shading is the process of calculating the pixel value of each pixel in the image based on the direction of the object relative to the light and its distance from the light source. The pixel value determines which graphics and light and dark effects are displayed in the final image.

[0084] The main idea of ​​ray tracing is to emit a primary ray from the viewpoint to each pixel on the imaging plane, find the intersection point with the nearest object that intersects the primary ray, and calculate the color of this intersection point based on the properties of the object, the properties of the light source, and the lighting model. At the intersection point, the primary ray is reflected and / or refracted to form secondary rays. From this intersection point, the secondary rays are traced along all reflection and / or refraction directions to determine the intersection point of the secondary ray with the next surface, and the color of this intersection point is calculated. This recursive process continues until the ray escapes the scene or reaches a light source.

[0085] The idea of ​​path tracing is similar to that of ray tracing. A ray of light is emitted from the viewpoint. When the ray encounters a surface point and needs to be reflected, a hemisphere is made with the surface point as the center. Several beams of light are drawn in several directions on the hemisphere, and then the Monte Carlo algorithm is used to calculate the illumination contribution of each of these beams of light to the surface point.

[0086] 4. Lighting model.

[0087] The lighting model is a mathematical model in computer graphics that is used to simulate and calculate the color and brightness of an object's surface under lighting.

[0088] When light strikes a surface, it may be absorbed, reflected, or transmitted. Some of the incident light energy is absorbed and converted into heat, while the remainder is reflected or transmitted. It is this reflected or transmitted portion of light that makes the object visible. The amount of light energy absorbed, reflected, or transmitted depends on the wavelength of the light. If approximately equal amounts of all wavelengths of incident light are absorbed, the object appears gray under white light. If only a small amount is absorbed, the object appears white. If certain wavelengths of light are selectively absorbed, the reflected and projected light leaving the surface will have different energy distributions, giving the object its appearance of color. The object's color depends on the wavelengths of light it selectively absorbs. The amount of light reflected or transmitted from an object's surface depends on the composition of the light source, its direction, the geometry of the light source, and the orientation and properties of the surface.

[0089] Lighting models are based on the laws of optics and describe the reflection, refraction, and scattering of light on surfaces. They can be categorized into two types: those based on physical theory and those based on experience. Physically based models use physical measurements and statistical methods to simulate real-world lighting effects, but they are computationally complex and difficult to implement. Empirical models, on the other hand, use specific probabilistic formulas tailored to a set of surface types. They are generally simpler and tend to produce idealized results.

[0090] Common illumination models based on empirical models include the Lambert model, the Phong model, and the Blinn-Phong model.

[0091] 4.1. Lambert model.

[0092] The Lambert model simulates the intensity of diffuse light based on the angle between the incident light and the surface normal. According to Lambert's cosine law, when light strikes a surface, the intensity of the diffuse light is proportional to the cosine of the angle. The larger the angle, the smaller the diffuse light intensity.

[0093] 4.2. Phong model.

[0094] In the real world, the light reaching the observer includes not only diffuse light, but also ambient light and specular light.

[0095] The Phong model simulates the intensity of specularly reflected light based on the Fresnel equations, the angle of incident light, the wavelength of incident light, and the material properties of the reflecting surface. The intensity of specularly reflected light depends on the angle of incident light. For an ideal mirror, the angle of reflection equals the angle of incidence, so only an observer located exactly in the direction of the reflected light can see the reflected light. For imperfect mirrors, the amount of light reaching the observer depends on the spatial distribution of the specularly reflected light. Smooth surfaces have a narrower or more focused spatial distribution of reflected light, while rough surfaces have a more diffused reflection.

[0096] 4.3. Blinn-Phong model.

[0097] The Blinn-Phong model is a modified version of the Phong model. It introduces the concept of a half-angle vector, denoted by H. Half-angle vector H is obtained by adding the light source direction L and the view direction V and then normalizing them. The Phong model calculates the intensity of specular light using the angle between the reflection direction R and the view direction V, while the Blinn-Phong model uses the angle between the half-angle vector H and the view direction V. Compared to the Phong model, the Blinn-Phong model simplifies the calculation process while achieving similar results.

[0098] 5. Digital watermark.

[0099] With the widespread adoption of smart devices and the internet, a large number of users have begun using these devices to create digital works. Digital works (such as the aforementioned 3D models and 2D data) are often disseminated online and used for commercial purposes, such as advertising and sales. Failure to prove ownership of digital works can harm the rights holder.

[0100] Digital watermarking is a technology that embeds specific digital signals into digital works to protect their copyright or integrity. It exploits the limitations of human senses to tightly integrate digital signals—images, text, symbols, numbers, and any other information that can serve as markers or identifiers—with the original data (e.g., 3D models and 2D data) and hide them within the original data. This technology can survive certain attacks.

[0101] Digital watermarks have many application scenarios. Based on the different characteristics of digital watermarks, watermark algorithms suitable for different scenarios can be designed. For example, robust invisible watermarks can be used for copyright protection, copy control, fingerprint recognition, etc.; fragile invisible watermarks can be used for content authentication, file modification detection, etc. Generally speaking, for copyright protection, digital watermarks should have the following basic characteristics:

[0102] 1) Provability: Watermarks should be able to provide complete and reliable evidence of the ownership of copyrighted digital works.

[0103] 2) Imperceptibility: Imperceptibility has two meanings. One refers to visual imperceptibility (the same requirement applies to hearing), that is, the changes in the image caused by the embedded watermark should be imperceptible to the observer's visual system. The ideal situation is that the watermarked image is visually identical to the original image, which is the requirement that most watermark algorithms should achieve. On the other hand, the watermark cannot be restored by statistical methods. For example, for a large number of information products that have been processed with the same method and watermark, it is impossible to extract the watermark or determine its existence even if statistical methods are used.

[0104] 3) Robustness: Robustness generally refers to the ability of a watermark to resist malicious attempts to remove it. A robust watermark should be able to withstand a wide range of attacks, such as vertex rearrangement, face rearrangement, rotation, scaling, translation, noise addition, smoothing, quantization, mesh simplification, retopology, local cropping, local deformation, and affine transformations.

[0105] Traditional 3D model watermarking algorithms can be categorized as spatial domain watermarking and transform domain watermarking based on the watermark embedding method. Spatial domain watermarking algorithms embed the watermark by modifying the model's vertex coordinates, edge lengths, triangle areas, and vertex order. Transform domain watermarking algorithms apply the watermark to transform domain coefficients, such as wavelet coefficients or spherical coordinate coefficients. 3D model digital watermarking algorithms are further categorized as blind or non-blind, depending on whether the original 3D model is required as input for watermark extraction.

[0106] Traditional 3D model watermarking algorithms offer significant advantages in watermark capacity, but due to the limitations of manual design, an algorithm is often only resistant to a small number of attacks. In recent years, with the development of deep learning, a growing number of deep learning-based watermarking techniques have emerged for embedding and extracting watermarks from 3D models. Deep learning-based digital watermarking techniques employ deep neural networks to embed digital watermarks, demonstrating strong robustness. However, most work follows the paradigm of first embedding the watermark into the 3D model and then extracting it from the model when authentication is required, as shown in Figure 3.

[0107] The original 3D model is a 3D model that does not contain a watermark. The original watermark may be, for example, the author's name or company logo. A deep neural network acquires the original 3D model and the original watermark, embeds the original watermark into the original 3D model, and generates a 3D model containing the watermark. After the 3D model containing the watermark is attacked, an attacked 3D model is generated. The attacked 3D model may still contain the watermark. The attacked 3D model is then processed using a deep neural network to extract the attacked watermark. After the 3D model is attacked, the watermark may contain missing information. During the training process of the deep neural network, an attack network is added to simulate various possible attacks on the 3D model in the hope of recovering the missing information and extracting a watermark that can prove the ownership of the 3D model.

[0108] The above method is effective in most scenarios. However, in some scenarios, the 3D model may be converted into 2D data. Attacks against network simulations usually target 3D models, making it difficult to deal with such scenarios.

[0109] FIG4 is a schematic diagram of a process for converting a three-dimensional model into two-dimensional data.

[0110] As shown in Figure 4, texture mapping is performed on the 3D model and then a watermark is embedded, generating a 3D model containing the watermark. The watermarked 3D model is then attacked to obtain the attacked 3D model. The attacked 3D model is then converted into 2D data through rendering, shading, and ray tracing. After the 3D model is converted into 2D data, the shape of the watermark is difficult to predict, making it difficult for a deep neural network trained using the attack network to extract a valid watermark from the 2D data.

[0111] The following describes a method for extracting watermarks provided in an embodiment of the present application.

[0112] FIG5 is a schematic diagram of a method 500 for extending a kernel according to an embodiment of the present application. The method 500 may be executed by an electronic device or a chip applied to the electronic device.

[0113] The electronic device may be a terminal, a server, or a chip used in a terminal or a server.

[0114] As an optional example, the terminal can be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wearable device, a wireless terminal in industrial control, a whole vehicle, a wireless communication module in the whole vehicle, a telematics box (T-box), a roadside unit (RSU), a wireless terminal in unmanned driving, a smart speaker in the Internet of Things (IoT), a wireless terminal device in remote medical, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, or a wireless terminal device in a smart home, etc. The embodiments of the present application are not limited to this.

[0115] As an optional example, the server may be a tower server, a blade server, or a rack server.

[0116] The embodiments of the present application do not limit the specific form of the electronic device and the specific technology adopted.

[0117] As shown in FIG5 , the method 500 includes:

[0118] S510, obtaining a target file, where the target file is a target three-dimensional model or target two-dimensional data, where the target two-dimensional data is an image or video generated based on a basic three-dimensional model and two-dimensional texture rendering, where the target three-dimensional model carries a target watermark, where the target two-dimensional data carries a target watermark, and where the target watermark carried by the target two-dimensional data is a watermark carried by the basic three-dimensional model and / or two-dimensional texture.

[0119] The target file may be a target 3D model, which may be a 3D model without texture data (i.e., a bare model) or a 3D model with texture data. For example, the target file may be an OSGB file, a PLY file, an OBJ file, or an S3MB file.

[0120] The target file may also be target two-dimensional data, for example, various image files or video files. The target two-dimensional data is an image or video generated by rendering based on a basic three-dimensional model and a two-dimensional texture. The basic three-dimensional model may be a bare model. Optionally, the basic three-dimensional model may be subjected to shading, ray tracing, path tracing, and other processing to generate the target two-dimensional data. The embodiments of the present application do not limit the specific method of generating the target two-dimensional data based on the basic three-dimensional model and the two-dimensional texture. The basic three-dimensional model and the two-dimensional texture in the target two-dimensional data may carry the target watermark at the same time, or only one of them may carry the target watermark.

[0121] The target watermark may be in the form of one or more of audio, image, video, and data structures. The target watermark may be embedded in the target three-dimensional model or target two-dimensional data in the form of binary message bits, where the data structure may be, for example, the type of a single facet or the arrangement of multiple facets. The target watermark may also be in the form of other data that can be embedded in a three-dimensional model. The content of the target watermark may be a company logo or a personal name, or other information that can prove the ownership of the target file. The embodiments of the present application do not limit the form and content of the target watermark.

[0122] The electronic device can obtain the target file locally. As shown in Figure 6, the user can drag the target file from the local folder to the watermark detection area. In this way, the electronic device can determine the location of the target file and read the target file from this location.

[0123] The electronic device may also receive the target file from the cloud via the network. The embodiments of the present application do not limit the specific manner in which the electronic device obtains the target file.

[0124] After obtaining the target file, the electronic device can perform the following steps.

[0125] S520: Determine the type of the target file, which may include a three-dimensional model or two-dimensional data.

[0126] The electronic device can determine the type of the target file based on the format of the target file. For example, if the format of the target file is an OBJ file, the electronic device can determine that the type of the target file is a three-dimensional model; if the format of the target file is a JPG file, the electronic device can determine that the type of the target file is two-dimensional data. The electronic device can also open the target file and determine the type of the target file based on the content in the target file. For example, if the electronic device determines that the target file contains three-dimensional model content (such as a mesh file) after opening the target file, the electronic device can determine that the type of the target file is a three-dimensional model; if the electronic device determines that the target file does not contain three-dimensional model content (such as a mesh file) after opening the target file, the electronic device can determine that the type of the target file is two-dimensional data.

[0127] Electronic devices can also determine the target file type based on user instructions. As shown in Figure 7, after obtaining the target file, the electronic device can pop up a "Please Select File Type" dialog box, allowing the user to select the target file type. After clicking the "3D Model" option or the "2D Data" option, the user can click the "Next" option to instruct the electronic device to perform subsequent processing. Due to the diverse types of 3D models and 2D data, the electronic device may not be able to recognize all types, or may incorrectly identify the type. Using the user's experience to determine the target file type can improve the accuracy of type determination.

[0128] If the user is unsure of the type of the target file, he or she can directly click the "Next" option, and the electronic device will determine the type of the target file based on the format or content of the 3D model. The embodiments of this application do not limit the specific method of determining the type of the target file.

[0129] After determining the type of the target file, the electronic device may perform the following steps.

[0130] Optionally, the method 500 further includes:

[0131] Determining at least one lighting model corresponding to the target file;

[0132] At least one two-dimensional data decoder is determined according to at least one illumination model, and the at least one two-dimensional data decoder is used to extract a target watermark carried by a target file according to its corresponding illumination model.

[0133] Two-dimensional data is the data obtained after rendering basic three-dimensional models and two-dimensional maps. The illumination model may be used in the rendering process. The illumination model used in the target file during the rendering process is determined. The two-dimensional data decoder corresponding to the illumination model can be used to specifically eliminate the impact of the illumination model on the target watermark, thereby improving the success rate of extracting the target watermark from the target file.

[0134] The at least one illumination model corresponding to the target file may be a Lambert model, a Phong model, a Blinn-Phong model, or other illumination models. The embodiments of the present application do not limit the type of the at least one illumination model.

[0135] Optionally, determining at least one lighting model corresponding to the target file includes:

[0136] When illumination model indication information is received, at least one illumination model is determined based on the illumination model indication information, and the illumination model indication information is used to indicate the illumination model used when rendering and generating the target two-dimensional data; or, when no illumination model indication information is received, and when watermark extraction indication information is received, at least one illumination model is determined to be a preset plurality of illumination models, and the illumination model indication information is used to indicate the illumination model used when rendering and generating the target two-dimensional data.

[0137] The illumination model indication information and the watermark extraction indication information may be indication information generated by the user clicking an icon on the interface.

[0138] As shown in Figure 8, after the electronic device determines the type of the target file, a dialog box "Please select an illumination model" may pop up. This dialog box contains one or more illumination models used by the neural network during training. If the user determines the illumination model of the target file, one of the models can be selected, such as clicking the "Blinn-Phong model" option. This click action triggers the electronic device to generate illumination model indication information. The electronic device determines that at least one illumination model is a Blinn-Phong model based on the illumination model indication information. Subsequently, the user can click the "Extract Watermark" option to trigger the electronic device to generate watermark extraction indication information, instructing the electronic device to extract the watermark in the target model.

[0139] If the user is not sure about the illumination model of the target file, the user can select multiple illumination models that may be used by the target file, or can directly click the "Extract Watermark" option, as shown in Figure 9. At this time, the electronic device receives the watermark extraction indication information without receiving the illumination model indication information. The electronic device determines that at least one illumination model is a preset multiple illumination models, and extracts the watermark in the target model based on the preset multiple illumination models.

[0140] After determining to extract the watermark in the target model, the electronic device may execute S530 or S540.

[0141] S530: When the type is a three-dimensional model, extract the target watermark carried by the target file through a three-dimensional model decoder, where the three-dimensional model decoder is a sub-network of the neural network.

[0142] S540: When the data type is two-dimensional data, extract the target watermark carried by the target file through at least one two-dimensional data decoder, where the at least one two-dimensional data decoder is a sub-network of a neural network.

[0143] Optionally, when the at least one illumination model is a plurality of preset illumination models, the at least one two-dimensional data decoder is a plurality of two-dimensional data decoders, and the plurality of two-dimensional data decoders correspond one-to-one to the plurality of preset illumination models, the electronic device may perform the following steps during the process of extracting the target watermark:

[0144] Using multiple two-dimensional data decoders to process target files respectively to determine multiple intermediate watermarks;

[0145] Comparing the plurality of intermediate watermarks with the reference watermark respectively to determine a plurality of similarities, wherein the plurality of similarities are used to indicate a degree of similarity between the plurality of intermediate watermarks and the reference watermark;

[0146] When at least one similarity among the multiple similarities is greater than or equal to a similarity threshold, a target watermark is determined from the intermediate watermarks corresponding to the at least one similarity, wherein the target watermark is the intermediate watermark corresponding to a value with the highest similarity among the at least one similarity.

[0147] For example, the preset multiple illumination models are the Lambert model, the Phong model, and the Blinn-Phong model. The two-dimensional data decoders corresponding to these three illumination models are decoder A (corresponding to the Lambert model), decoder B (corresponding to the Phong model), and decoder C (corresponding to the Blinn-Phong model). The neural network can use these three decoders to process the target file and extract three intermediate watermarks respectively. Assuming that the three intermediate watermarks are watermark A, watermark B, and watermark C, the neural network can respectively compare these three intermediate watermarks with the watermarks registered in the database (i.e., the reference watermark) to obtain three similarities, namely similarity A, similarity B, and similarity C. If similarity B and similarity C among the three similarities are greater than or equal to the similarity threshold, then the value with the highest similarity can be determined from similarity B and similarity C, such as similarity C, and the intermediate watermark corresponding to similarity C (watermark C) can be displayed to the user as the target watermark.

[0148] The reference watermark can be a user-specified watermark, allowing the user to identify the watermark embedded in their 3D model. The target watermark determined based on the user-specified watermark has a high degree of accuracy. The reference watermark can also be a watermark determined independently by the electronic device. In some cases, if the user is unsure which watermark is embedded in their 3D model, the electronic device can determine the reference watermark based on the type of intermediate watermark. For example, if watermarks A, B, and C are all text, the electronic device can determine a text watermark as the reference watermark. If watermarks A, B, and C are all images, the electronic device can determine an image watermark as the reference watermark.

[0149] Using two-dimensional data decoders corresponding to multiple illumination models to extract watermarks can improve the success rate of watermark extraction when the user is not sure of the illumination model of the target file.

[0150] The above details the watermark extraction methods provided by the embodiments of this application. The watermark forms differ between 3D models and 2D data. For example, a 3D watermark may appear as an arrangement of facets, while a 2D watermark may appear as pixels. Using a decoder tailored to the target file type to extract the watermark can specifically mitigate damage to the watermark caused by attacks, improving the success rate of extracting watermarks from attacked 3D models.

[0151] The neural network in S530 and S540 can be a deep neural network. Before executing S510, the deep neural network can be pre-trained using a three-dimensional model, two-dimensional data, and different attack methods to enable it to extract the target watermark from the three-dimensional model and two-dimensional data. The neural network in S530 and S540 can also be other neural networks capable of extracting the target watermark from the three-dimensional model and two-dimensional data. The embodiments of this application do not limit the specific type of the neural network in method 500.

[0152] In addition, the neural networks in S530 and S540 can be the same neural network or different neural networks. When the neural networks in S530 and S540 are the same neural network, the neural network has the ability to extract watermarks from three-dimensional models and two-dimensional data; when the neural networks in S530 and S540 are different neural networks, different neural networks can have different capabilities. For example, one of the neural networks has the ability to extract watermarks from three-dimensional models, and the other neural network has the ability to extract watermarks from two-dimensional data. Of course, both neural networks can also have the ability to extract watermarks from three-dimensional models and two-dimensional data.

[0153] The following describes a method for training a neural network provided by an embodiment of the present application.

[0154] As shown in FIG10 , the neural network provided by an embodiment of the present application may include a 3D model encoder, an attack library, a 3D model decoder, a 2D data encoder, a differentiable renderer, and at least one 2D data decoder, wherein:

[0155] The 3D model encoder is used to encode the original watermark and the original 3D model (e.g., a bare model) to generate 3D training data. The 3D training data is the 3D model carrying the original watermark. The 3D model encoder is a sub-network of the neural network.

[0156] Attack library 1 is used to simulate possible attack methods of the target file, and attack the three-dimensional training data based on the attack methods to generate attacked three-dimensional training data. The attack method simulated by attack library 1 can be an attack on the three-dimensional model;

[0157] The 3D model decoder is used to extract the first watermark from the attacked 3D training data;

[0158] The two-dimensional data encoder is used to encode the original watermark and the original two-dimensional map (e.g., image) to generate a two-dimensional training map. The two-dimensional training map is a two-dimensional map carrying the original watermark. The two-dimensional data encoder is a sub-network of the neural network.

[0159] Attack Library 2 is used to simulate possible attack methods against the target file, and attacks the 3D model composed of the original 3D model and the 2D training map based on the attack methods to generate an attacked 3D model. The attack methods simulated by Attack Library 2 can be attacks on the 3D model, attacks on the 2D data, or attacks that do not distinguish between the 3D model and the 2D data.

[0160] The differentiable renderer is used to render the attacked three-dimensional model into at least one two-dimensional training data, where the at least one two-dimensional training data is an image or a video;

[0161] The at least one two-dimensional data decoder is used to extract a second watermark from the at least one two-dimensional training data.

[0162] The original watermark can be one or more of text, images, audio, video, and data structures. During training, the 3D model encoder and decoder can be trained using 3D training data. Alternatively, the decoder can be trained separately to facilitate faster neural network convergence. Attack Library 1 and Attack Library 2 can simulate different attacks, such as vertex rearrangement, face rearrangement, rotation, scaling, translation, noise addition, smoothing, quantization, mesh simplification, retopology, local cropping, local deformation, and affine transformation. Training with the attack library can improve the robustness of the 3D model decoder.

[0163] After the training degree of the 3D model encoder and the 3D model decoder meets the requirements, the 2D data encoder and at least one 2D data decoder are trained using the 2D training data, or, optionally, the 2D data decoder is trained separately. During the training process of the 2D data encoder and the at least one 2D data decoder, the 3D model encoder and the 3D model decoder can also be fine-tuned. Among them, at least one 2D data decoder corresponds one-to-one to the illumination model in the differentiable renderer. For example, the illumination model in the differentiable renderer includes the Lambert model, the Phong model and the Blinn-Phong model, then the at least one 2D data decoder includes a 2D data decoder corresponding to the Lambert model, a 2D data decoder corresponding to the Phong model and a 2D data decoder corresponding to the Blinn-Phong model, and these three 2D data decoders are trained separately during the training process.

[0164] In other embodiments provided herein, the 3D model encoder / decoder and the 2D model encoder / decoder may also be trained separately, that is, the 2D data encoder and at least one 2D data decoder may be trained directly using 2D training data, or, optionally, the 2D data decoder may be trained separately. This application does not limit the order in which the encoders and decoders involved in the embodiments are trained.

[0165] The first and second watermarks are the predicted values ​​of the neural network, and the original watermark is the true value. The first, second, and original watermarks are input into a first loss function to determine the output value of the first loss function. Backpropagation is performed based on the output value of the first loss function to obtain the gradients of each neural network layer in the 3D model encoder and the 3D model decoder. The neural network parameters are adjusted along the direction of fastest gradient descent until the output value of the first loss function reaches a minimum. Similarly, the second watermark and the original watermark are input into a second loss function to determine the output value of the second loss function. Backpropagation is performed based on the output value of the second loss function to obtain the gradients of each neural network layer in the 2D data encoder and at least one 2D data decoder. The neural network parameters are adjusted along the direction of fastest gradient descent until the output value of the second loss function reaches a minimum.

[0166] The neural network may also be trained based on other methods, and the embodiments of the present application do not limit the training method of the neural network.

[0167] In the above training method, the 3D model encoder and the 2D data encoder are both sub-networks to be trained. Optionally, when both the 3D model encoder and the 2D data encoder are already trained sub-networks, the parameters of the 3D model encoder and the 2D data encoder can be fixed during training, thereby improving the training efficiency of the neural network.

[0168] After the neural network training is completed, it can be used for inference operations.

[0169] As shown in Figure 11, in the inference stage, a three-dimensional model containing a watermark or two-dimensional data containing a watermark can be input. Optionally, the three-dimensional model decoder and the two-dimensional data decoder can share an input layer and an output layer. More neural network layers can be included between the input layer and the three-dimensional model decoder, between the output layer and the three-dimensional model decoder, between the input layer and the two-dimensional data decoder, and between the output layer and the two-dimensional data decoder.

[0170] After the input layer obtains the three-dimensional model containing the watermark, it passes the three-dimensional model containing the watermark to the three-dimensional model decoder. The three-dimensional model decoder extracts the watermark from the three-dimensional model containing the watermark and outputs the watermark through the output layer.

[0171] After the input layer obtains the two-dimensional data containing the watermark, it passes the two-dimensional data containing the watermark to the two-dimensional data decoder. The two-dimensional data decoder extracts the watermark from the two-dimensional data containing the watermark and outputs the watermark through the output layer.

[0172] The above describes in detail the method examples provided by the embodiments of the present application. It is understandable that, in order to implement the above functions, the corresponding device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0173] 12 is a schematic structural diagram of a watermark extraction apparatus 1200 provided in an embodiment of the present application, wherein the apparatus 1200 includes a processing module 1210 and an input module 1220. The input module 1220 performs a receiving step or an input step under the control of the processing module 1210.

[0174] The input module 1220 is used to: obtain a target file, where the target file is a target three-dimensional model or target two-dimensional data, where the target two-dimensional data is an image or video generated based on a basic three-dimensional model and two-dimensional map rendering, where the target three-dimensional model carries a target watermark, and where the target two-dimensional data carries a target watermark, where the target watermark carried by the target two-dimensional data is a watermark carried by the basic three-dimensional model and / or two-dimensional map; the processing module 1210 is used to: determine the type of the target file, where the type includes a three-dimensional model or two-dimensional data; when the type is a three-dimensional model, the processing module 1210 is further used to extract the target watermark carried by the target file through a three-dimensional model decoder, where the three-dimensional model decoder is a sub-network of a neural network; or, when the type is two-dimensional data, the processing module 1210 is further used to extract the target watermark carried by the target file through at least one two-dimensional data decoder, where at least one two-dimensional data decoder is a sub-network of a neural network.

[0175] Among them, the module is an example of a software functional unit, and the processing module 1210 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the above-mentioned computing instance may be one or more. For example, the processing module 1210 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Among them, usually a region can include multiple AZs.

[0176] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Inter-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0177] As an example of a hardware functional unit, processing module 1210 may include at least one computing device, such as a server. Alternatively, processing module 1210 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0178] The multiple computing devices included in processing module 1210 can be distributed in the same zone or in different zones. The multiple computing devices included in processing module 1210 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in processing module 1210 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0179] It should be noted that, in other embodiments, the processing module 1210 can be used to execute the processing steps in the method 500, and the input module 1220 can be used to execute the input steps, receiving steps or acquisition steps in the method 500. The steps that the processing module 1210 and the input module 1220 are responsible for implementing can be specified as needed. The full functions of the device 1200 for extracting watermarks can be realized by respectively implementing different steps in the method 500 through the processing module 1210 and the input module 1220.

[0180] This application also provides a computing device 100. As shown in FIG13 , computing device 100 includes a bus 102, a processor 104, a memory 106, and a communication interface 108. Processor 104, memory 106, and communication interface 108 communicate with each other via bus 102. Computing device 100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 100.

[0181] Bus 102 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG13 shows only one line, but this does not imply a single bus or type of bus. Bus 102 may include a path for transmitting information between various components of computing device 100 (e.g., memory 106, processor 104, and communication interface 108).

[0182] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0183] The memory 106 may include volatile memory, such as random access memory (RAM). The processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0184] The memory 106 stores executable program codes, and the processor 104 executes the executable program codes to respectively implement the functions of the aforementioned processing module 1210 and input module 1220 , thereby implementing the method 500 . That is, the memory 106 stores instructions for executing the method 500 .

[0185] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.

[0186] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0187] 14 , the computing device cluster includes at least one computing device 100. The same instructions for executing the method 500 may be stored in the memory 106 of one or more computing devices 100 in the computing device cluster.

[0188] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also respectively store some instructions for executing the method 500. In other words, the combination of one or more computing devices 100 may jointly execute the instructions for executing the method 500.

[0189] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster may store different instructions, each for executing part of the functions of the watermark extraction apparatus 1200. That is, the instructions stored in the memory 106 in different computing devices 100 may implement the functions of one or more modules in the processing module 1210 and the input module 1220.

[0190] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), among others. FIG. 15 illustrates a possible implementation. As shown in FIG. 15 , two computing devices 100A and 100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 106 in the computing device 100A stores instructions for executing the functions of the processing module 1210. Simultaneously, the memory 106 in the computing device 100B stores instructions for executing the functions of the input module 1220.

[0191] The connection method between the computing device clusters shown in Figure 15 may be that considering that the method 500 provided in this application requires a large amount of data storage, it is considered to hand over the functions implemented by the input module 1220 to the computing device 100B for execution.

[0192] It should be understood that the functions of the computing device 100A shown in FIG15 may also be completed by multiple computing devices 100. Similarly, the functions of the computing device 100B may also be completed by multiple computing devices 100.

[0193] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes method 500.

[0194] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a high-density digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute method 500.

[0195] It should be understood that the “embodiment” mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the various embodiments in the entire specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the terminal device and / or the network device can perform some or all of the steps in each embodiment. These steps or operations are merely examples, and the embodiments of the present application can also perform other operations or variations of various operations. In addition, the various steps can be performed in a different order than presented in the various embodiments, and it is possible that not all operations in the embodiments of the present application need to be performed. Moreover, the size of the sequence number of each of the above processes does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0196] It should also be understood that in this application, "when", "if" and "if" all mean that the executing subject will take corresponding measures under certain objective circumstances, and do not limit the time. It does not require the executing subject to make judgments when implementing it, nor does it mean that there are other limitations.

[0197] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0198] It should be understood that in each embodiment of the present application, "B corresponding to A" means that B is associated with A and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A, but B can also be determined based on A and / or other information.

[0199] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for extracting a watermark, characterized in that: The method comprises: Obtaining a target file, where the target file is a target three-dimensional model or target two-dimensional data, where the target two-dimensional data is an image or video generated by rendering based on a basic three-dimensional model and a two-dimensional map, where the target three-dimensional model carries a target watermark, and the target two-dimensional data carries a target watermark, where the target watermark carried by the target two-dimensional data is the watermark carried by the basic three-dimensional model and / or the two-dimensional map; Determining the type of the target file, wherein the type includes a three-dimensional model or two-dimensional data; When the type is a three-dimensional model, the target watermark carried by the target file is extracted by a three-dimensional model decoder, wherein the three-dimensional model decoder is a sub-network of a neural network; or When the type is two-dimensional data, the target watermark carried by the target file is extracted by at least one two-dimensional data decoder, and the at least one two-dimensional data decoder is a sub-network of a neural network.

2. The method according to claim 1, characterized in that Before extracting the target watermark carried by the target file through at least one two-dimensional data decoder, the method further includes: Determining at least one lighting model corresponding to the target file; The at least one two-dimensional data decoder is determined according to the at least one illumination model, and the at least one two-dimensional data decoder is used to extract the target watermark carried by the target file according to its corresponding illumination model.

3. The method according to claim 2, characterized in that Determining at least one illumination model corresponding to the target file includes: When receiving illumination model indication information, determining the at least one illumination model according to the illumination model indication information, wherein the illumination model indication information is used to indicate the illumination model used when rendering and generating the target two-dimensional data; or When no illumination model indication information is received, and when watermark extraction indication information is received, it is determined that the at least one illumination model is a preset plurality of illumination models, the illumination model indication information is used to indicate the illumination model used when rendering and generating the target two-dimensional data, and the watermark extraction indication information is used to indicate the extraction of the target watermark carried by the target file.

4. The method according to claim 3, characterized in that When the at least one illumination model is the plurality of preset illumination models, the at least one two-dimensional data decoder is a plurality of two-dimensional data decoders, and the plurality of two-dimensional data decoders correspond one-to-one to the plurality of preset illumination models. Processing the target file by using the at least one two-dimensional data decoder includes: Processing the target file using the multiple two-dimensional data decoders respectively to determine multiple intermediate watermarks; Comparing the plurality of intermediate watermarks with a reference watermark, respectively, to determine a plurality of similarities corresponding to the plurality of intermediate watermarks, the plurality of similarities being used to indicate a degree of similarity between the plurality of intermediate watermarks and the reference watermark; When at least one similarity among the multiple similarities is greater than or equal to a similarity threshold, the target watermark is determined from the intermediate watermarks corresponding to the at least one similarity, wherein the target watermark is the intermediate watermark corresponding to a value with the highest similarity among the at least one similarity.

5. The method according to any one of claims 1 to 4, characterized in that Before obtaining the target file, the method includes: Obtain the original watermark and original 3D model; Encoding the original watermark and the original three-dimensional model through a three-dimensional model encoder to generate three-dimensional training data, wherein the three-dimensional training data is a three-dimensional model carrying the original watermark, and the three-dimensional model encoder is a subnetwork of the neural network; Attacking the three-dimensional training data to generate attacked three-dimensional training data; The 3D model decoder is trained using the attacked 3D training data, and the 3D model decoder is used to extract a first watermark from the attacked 3D training data.

6. The method according to any one of claims 1 to 4, characterized in that Before obtaining the target file, the method includes: Obtain the original watermark, original 3D model and original 2D texture; Encoding the original watermark and the original two-dimensional map through a two-dimensional data encoder to generate a two-dimensional training map, wherein the two-dimensional training map is a two-dimensional map carrying the original watermark, and the two-dimensional data encoder is a subnetwork of the neural network; Attacking a three-dimensional model composed of the original three-dimensional model and the two-dimensional training map to generate an attacked three-dimensional model; Rendering the attacked three-dimensional model using at least one differentiable renderer to obtain at least one two-dimensional training data, where the at least one two-dimensional training data is a picture or a video; The at least one two-dimensional data decoder is trained using the at least one two-dimensional training data, and the at least one two-dimensional data decoder is used to extract a second watermark from the at least one two-dimensional training data.

7. The method according to claim 6, characterized in that The using the at least one two-dimensional training data to train the at least one two-dimensional model decoder comprises: When the training degree of the three-dimensional model decoder meets the requirement, the at least one two-dimensional data decoder is trained using the at least one two-dimensional training data.

8. The method according to claim 6 or 7, characterized in that The at least one two-dimensional data decoder includes a plurality of two-dimensional data decoders, the differentiable renderer includes a plurality of lighting models, and the plurality of two-dimensional data decoders correspond one-to-one to the plurality of lighting models.

9. A device for extracting watermarks, characterized in that: The device comprises: An input module is configured to obtain a target file, wherein the target file is a target three-dimensional model or target two-dimensional data, wherein the target two-dimensional data is an image or video generated by rendering based on a basic three-dimensional model and a two-dimensional map, wherein the target three-dimensional model carries a target watermark, and the target two-dimensional data carries a target watermark, wherein the target watermark carried by the target two-dimensional data is the watermark carried by the basic three-dimensional model and / or the two-dimensional map; A processing module, configured to determine a type of the target file, wherein the type includes a three-dimensional model or two-dimensional data; When the type is a three-dimensional model, the processing module is further configured to extract the target watermark carried by the target file through a three-dimensional model decoder, wherein the three-dimensional model decoder is a sub-network of a neural network; or When the type is two-dimensional data, the processing module is further used to extract the target watermark carried by the target file through at least one two-dimensional data decoder, and the at least one two-dimensional data decoder is a sub-network of a neural network.

10. The device according to claim 9, characterized in that Before extracting the target watermark carried by the target file through at least one two-dimensional data decoder, the processing module is further configured to: Determining at least one lighting model corresponding to the target file; The at least one two-dimensional data decoder is determined according to the at least one illumination model, and the at least one two-dimensional data decoder is used to extract the target watermark carried by the target two-dimensional data according to its corresponding illumination model.

11. The device according to claim 10, characterized in that The processing module is specifically used for: When receiving the illumination model indication information, determining the at least one illumination model according to the illumination model indication information, wherein the illumination model indication information is used to indicate the illumination model used when rendering and generating the target two-dimensional data; or, When no illumination model indication information is received, and when watermark extraction indication information is received, it is determined that the at least one illumination model is a preset plurality of illumination models, the illumination model indication information is used to indicate the illumination model used when rendering and generating the target two-dimensional data, and the watermark extraction indication information is used to indicate the extraction of the target watermark carried by the target file.

12. The device according to claim 11, characterized in that When the at least one illumination model is the plurality of preset illumination models, the at least one two-dimensional data decoder is a plurality of two-dimensional data decoders, and the plurality of two-dimensional data decoders correspond one-to-one to the plurality of preset illumination models, the processing module is specifically configured to: Processing the target file using the multiple two-dimensional data decoders respectively to determine multiple intermediate watermarks; Comparing the plurality of intermediate watermarks with a reference watermark, respectively, to determine a plurality of similarities corresponding to the plurality of intermediate watermarks, the plurality of similarities being used to indicate a degree of similarity between the plurality of intermediate watermarks and the reference watermark; When at least one similarity among the multiple similarities is greater than or equal to a similarity threshold, the target watermark is determined from the intermediate watermarks corresponding to the at least one similarity, wherein the target watermark is the intermediate watermark corresponding to a value with the highest similarity among the at least one similarity.

13. The device according to any one of claims 9 to 12, characterized in that Before obtaining the target file, The input module is further used to: obtain the original watermark and the original three-dimensional model; The processing module is further configured to: encode the original watermark and the original 3D model through a 3D model encoder to generate 3D training data, wherein the 3D training data is a 3D model carrying the original watermark, and the 3D model encoder is a subnetwork of the neural network; Attacking the three-dimensional training data to generate attacked three-dimensional training data; The 3D model decoder is trained using the attacked 3D training data, and the 3D model decoder is used to extract a first watermark from the attacked 3D training data.

14. The device according to any one of claims 9 to 12, characterized in that Before obtaining the target file, The input module is also used to: obtain the original watermark, the original three-dimensional model and the original two-dimensional map; The processing module is further configured to: encode the original watermark and the original two-dimensional map using a two-dimensional data encoder to generate a two-dimensional training map, wherein the two-dimensional training map is a two-dimensional map carrying the original watermark, and the two-dimensional data encoder is a subnetwork of the neural network; attack the three-dimensional model composed of the original three-dimensional model and the two-dimensional training map to generate an attacked three-dimensional model; Render the attacked three-dimensional model using at least one differentiable renderer to obtain at least one two-dimensional training data, wherein the at least A two-dimensional training data is a picture or video; The at least one two-dimensional data decoder is trained using the at least one two-dimensional training data, and the at least one two-dimensional data decoder is used to extract a second watermark from the at least one two-dimensional training data.

15. The device according to claim 14, characterized in that The processing module is specifically used for: When the training degree of the three-dimensional model decoder meets the requirement, the at least one two-dimensional data decoder is trained using the at least one two-dimensional training data.

16. The device according to claim 14 or 15, characterized in that The at least one two-dimensional data decoder includes a plurality of two-dimensional data decoders, the differentiable renderer includes a plurality of lighting models, and the plurality of two-dimensional data decoders correspond one-to-one to the plurality of lighting models.

17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.

18. A computer-readable storage medium, characterized in that The method comprises computer program instructions, which, when executed by a computing device cluster, perform the method according to any one of claims 1 to 8.

19. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Three-dimensional geographic model digital watermarking method with copyright protection service orientation

    CN103377455A

  • Encoder and decoder training method, embedding method and detection method of audio blind watermark

    CN117012211A

  • Three-dimensional model watermark authentication method and system based on hidden space

    CN117495648A

  • 3D digital watermark embedding and detecting method and device based on virtual optics

    CN1487421A

  • Optimized authentication of graphic authentication code

    EP3242239A1

Cited By

  • Cross-medium physical image watermarking method and system based on micro-physical simulation

    CN121746153A