Diffusion model watermarking method and apparatus
By using the encoder and decoder of the teacher model for segmented adjustments in the diffusion model, the problem of deteriorated image quality after fine-tuning of the diffusion model was solved, achieving high-quality watermark embedding and improved model usability.
Patent Information
- Application Number
- PCT/CN2025/074068
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-19
- Filing Date
- 2025-01-22
- Publication Date
- 2026-01-22
AI Technical Summary
Using watermark extraction networks to fine-tune the parameters of the diffusion model leads to a deterioration in image quality and a decrease in the usability of the fine-tuned diffusion model.
The first diffusion model is segmented and adjusted using the encoder and decoder in the teacher model to learn the features of different stages, generate a 3D model with the target watermark, and extract the watermark from the 3D model to obtain the second diffusion model.
It improves the usability and robustness of the diffusion model, ensures the quality of the generated 3D model and the watermark matching degree, and meets users' customization needs.
Smart Images

Figure CN2025074068_22012026_PF_FP_ABST
Abstract
Description
Diffusion model watermark embedding method and device
[0001] The present application claims priority from the Chinese patent application No. 202410979812.9 filed on July 19, 2024, and entitled "Diffusion model watermark embedding method and device", the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The present application relates to the field of artificial intelligence, and in particular to a diffusion model watermark embedding method and device. BACKGROUND
[0003] Model watermarking refers to a technology of embedding specific marks or information in a neural network model, so as to identify the owner of the model or prove the source of the model. The model parameters of a diffusion model used for generating images are fine-tuned by a watermark extraction network (decoder), so that the images generated by the diffusion model can obtain corresponding watermarks through the processing of the watermark extraction network.
[0004] However, fine-tuning the parameters of the diffusion model by the watermark extraction network will cause more model parameters in the diffusion model to be adjusted for generating watermarks, and thus the quality of the images generated by the diffusion model will be deteriorated, and the availability of the fine-tuned diffusion model will be deteriorated. SUMMARY
[0005] The present application provides a diffusion model watermark embedding method and device to solve the problem that fine-tuning the parameters of the diffusion model by the watermark extraction network will cause more model parameters in the diffusion model to be adjusted for generating watermarks, and thus the quality of the images generated by the diffusion model will be deteriorated, and the availability of the fine-tuned diffusion model will be deteriorated.
[0006] The present application adopts the following scheme.
[0007] In a first aspect, the present application provides a diffusion model watermark embedding method. The method can be executed by a computing device, a chip (such as a processor) in the computing device, or a computing device cluster composed of multiple computing devices. The diffusion model watermark embedding method executed by the computing device includes: the computing device obtains a first diffusion model, and trains the first diffusion model according to an encoder in a teacher model to obtain an intermediate diffusion model, and trains the intermediate diffusion model according to a decoder in the teacher model to obtain a second diffusion model. The first diffusion model is used to generate a three-dimensional model, the encoder is used to generate a three-dimensional model with a target watermark, the decoder is used to extract the watermark in the three-dimensional model, the intermediate diffusion model can generate a three-dimensional model with a first watermark, and the second diffusion model can generate a three-dimensional model with a second watermark. The matching degree between the second watermark and the target watermark is greater than the matching degree between the first watermark and the target watermark.
[0008] In the present application, the computing device adjusts the first diffusion model through the encoder and the decoder in the teacher model, learns the characteristics of different stages such as the stage of generating a three-dimensional model with a target watermark and the stage of extracting the watermark in the three-dimensional model, and further obtains the second diffusion model which can preserve the model parameters of generating a three-dimensional model as much as possible and update the model parameters of generating a watermark, thereby improving the usability (robustness) of the second diffusion model.
[0009] In a possible case, the computing device is provided by a cloud service provider, that is, the diffusion model watermark embedding method can be applied to a cloud computing platform and executed by one or more computing devices included in the cloud computing platform.
[0010] In a possible example, the diffusion model watermark embedding method can be a service on the cloud computing platform, and a user calls the service to embed a watermark in the first diffusion model to obtain the second diffusion model.
[0011] In a possible case, the computing device can be a computing device located on the user side.
[0012] In a possible case, the target watermark can be a watermark predefined by the computing device. The predefined watermark can be randomly generated by the computing device, or generated by the computing device according to user information.
[0013] For example, the computing device takes an identity document (ID) in the user information as the target watermark.
[0014] In a possible case, the first diffusion model is a pre-trained diffusion model.
[0015] In a possible case, the three-dimensional model is represented by any one of the following: a point cloud, a polygon mesh, and a voxel grid.
[0016] In a possible implementation, before the computing device trains the first diffusion model according to the encoder in the teacher model to obtain the intermediate diffusion model, the diffusion model watermark embedding method further includes: the computing device providing a user interface, the user interface including a first control component, and in response to a triggering operation of the first control component by the user, obtaining the target watermark.
[0017] In this application, the computing device obtains the target watermark determined by the user by responding to the triggering operation of the first control component on the display interface, realizes visualization while meeting the customization needs of the user for the target, and improves the user experience.
[0018] In a possible example, the computing device displays the user interface at a front end connected thereto, which can be a display or a terminal located at the user side and accessing the user interface provided by the computing device through an application programming interface (API) provided by the computing device.
[0019] In a possible case, the target watermark includes one or more of the following: a picture, text, audio, video, and a bit stream.
[0020] In a possible example, the watermark in the three-dimensional model is in the form of a bit stream. In other words, after obtaining the target watermark in the form of a picture, text, audio, or video, the computing device converts the target watermark in the form of a picture, text, audio, or video into a target watermark in the form of a bit stream, and then embeds the target watermark in the form of a bit stream into the first diffusion model to obtain a second diffusion model.
[0021] Two possible implementations are provided below for the manner in which the computing device obtains the first diffusion model.
[0022] In a possible implementation, the computing device obtains the first diffusion model, including: the computing device providing a user interface, the user interface including a second control component, and in response to a triggering operation of the second control component by the user, obtaining the first diffusion model.
[0023] In a possible example, the computing device displays the user interface at a front end connected thereto.
[0024] In the present application, the computing device obtains the first diffusion model determined by the user by responding to the triggering operation of the second control component on the display interface, and meets the watermark injection of the user-specified first diffusion model while realizing the visualization, realizes the targeted processing of the user-specified first diffusion model, and realizes that the obtained second diffusion model is easier to trace.
[0025] In another possible implementation, the computing device obtains the first diffusion model, including: the computing device obtains the first diffusion model from a plurality of reference diffusion models, each reference diffusion model in the plurality of reference diffusion models being used to generate a three-dimensional model.
[0026] In the present application, in the case where the user does not provide the first diffusion model, the computing device can obtain the first diffusion model from a plurality of reference diffusion models, and then embed the watermark in the first diffusion model, so as to obtain the user's own second diffusion model, and when the second diffusion model of the user is used by others, the three-dimensional model generated according to the second diffusion model can be traced according to the watermark in the three-dimensional model, thereby reducing the difficulty of evidence collection when the second diffusion model of the user is used by others.
[0027] In a possible example, before the computing device executes the diffusion model watermark embedding method provided in the present application, a plurality of reference diffusion models have been deployed in the computing device, and therefore the computing device can select one of the plurality of reference diffusion models as the first diffusion model.
[0028] In a possible case, the reference diffusion model is a pre-trained diffusion model.
[0029] The selection manner can be random selection, selection according to the user's instruction, selection according to the user-determined label, and each reference diffusion model in the plurality of reference diffusion models has one or more labels.
[0030] In a possible case, the difference between the plurality of reference diffusion models can be different use scenarios and different diffusion model structures.
[0031] For example, a first reference diffusion model in the plurality of reference diffusion models is used to generate a three-dimensional model of an office appliance, and a second reference diffusion model is used to generate a three-dimensional model of a transportation tool.
[0032] The network structure of a third reference diffusion model in the plurality of reference diffusion models adds a convolution layer, a pooling processing layer, a feature extraction network layer, etc., compared with the network structure of a fourth reference diffusion model.
[0033] In a possible implementation, the teacher model is trained by using a training data set and the target watermark. The training data set includes three-dimensional models of different types, including one or more of the following: vehicles, office supplies, and buildings.
[0034] In the present application, the teacher model is trained by using a training data set including three-dimensional models of different types and the target watermark, which can improve the robustness of the encoder in the teacher model in generating three-dimensional models with the target watermark. Then, the encoder in the teacher model is used to train the first diffusion model, which can improve the robustness of the final second diffusion model.
[0035] In a possible example, the above types can further include animals, plants, clothing, and the like.
[0036] In a possible example, the computing device jointly trains the encoder and the decoder in the teacher model by using the training data set and the target watermark.
[0037] In a possible implementation, the computing device trains the first diffusion model according to the encoder in the teacher model to obtain an intermediate diffusion model, including: the computing device inputs a first three-dimensional model in the fine-tuning data set into the encoder and the first diffusion model, respectively outputs corresponding first reconstructed three-dimensional models and second reconstructed three-dimensional models, and updates the model parameters of the first diffusion model according to the loss between the first reconstructed three-dimensional models and the second reconstructed three-dimensional models to obtain the intermediate diffusion model. The first three-dimensional model is any one of the plurality of three-dimensional models included in the fine-tuning data set.
[0038] In the present application, the model parameters of the first diffusion model are updated according to the loss between the first reconstructed three-dimensional models and the second reconstructed three-dimensional models, which can realize the learning ability of the obtained intermediate diffusion model in generating three-dimensional models with the target watermark, so that the matching degree between the first reconstructed three-dimensional models and the second reconstructed three-dimensional models generated by the first diffusion model is continuously increased, the quality of the three-dimensional models generated by the intermediate model is improved, and the robustness of the final second diffusion model is further improved.
[0039] In a possible example, the above first three-dimensional model has a watermark.
[0040] In another possible example, the above first three-dimensional model does not have a watermark.
[0041] In a possible case, the three-dimensional models in the fine-tuning data set are discrete point clouds, i.e., without fixed shapes.
[0042] In a possible implementation, the computing device trains the intermediate diffusion model according to the decoder in the teacher model to obtain the second diffusion model, including: the computing device inputs the second three-dimensional model in the fine-tuning dataset into the intermediate diffusion model to obtain a third reconstructed three-dimensional model, and inputs the third reconstructed three-dimensional model into the decoder to obtain the first watermark, and further updates the model parameters of the intermediate diffusion model according to the loss between the first watermark and the target watermark to obtain the second diffusion model. The third reconstructed three-dimensional model includes the first watermark, and the second three-dimensional model is any one of the plurality of three-dimensional models included in the fine-tuning dataset.
[0043] In the present application, the computing device optimizes the model parameters of the intermediate diffusion model according to the loss between the first watermark and the target watermark, so that the watermark in the three-dimensional model generated by the finally obtained second diffusion model has a higher matching degree with the target watermark, thereby improving the quality of the watermark in the three-dimensional model generated by the second diffusion model.
[0044] In a possible implementation, the encoder includes a local feature extraction network and a global feature extraction network.
[0045] In a possible example, the local feature extraction network includes a graph convolutional neural network; and / or, the global feature extraction network includes a graph convolutional neural network.
[0046] In a possible implementation, if the three-dimensional model generated by the second diffusion model can extract the watermark in the three-dimensional model by using the decoder in the aforementioned teacher model, the traceability of the second diffusion model can be realized quickly, that is, the owner of the second diffusion model can be determined, and the difficulty of forensics on the use of the aforementioned second diffusion model by others is reduced.
[0047] In a second aspect, the present application provides a diffusion model watermark embedding device. The diffusion model watermark embedding device is applied to a computer system (such as a computer cluster) or a computing device supporting the computer system to implement a diffusion model watermark embedding method. The diffusion model watermark embedding device includes various modules for executing the diffusion model watermark embedding method in the first aspect or any optional implementation of the first aspect. For example, the diffusion model watermark embedding device includes a first acquisition module, a first training module, and a second training module. Wherein,
[0048] The first acquisition module is configured to acquire a first diffusion model, and the first diffusion model is used to generate a three-dimensional model.
[0049] The first training module is configured to train the first diffusion model according to an encoder in a teacher model to obtain an intermediate diffusion model, and the encoder is used to generate a three-dimensional model with a target watermark, and the intermediate diffusion model is capable of generating a three-dimensional model with a first watermark.
[0050] The second training module is configured to train the intermediate diffusion model according to a decoder in the teacher model to obtain a second diffusion model; the decoder is configured to extract the watermark in the three-dimensional model, and the second diffusion model is capable of generating a three-dimensional model with a second watermark, and the matching degree between the second watermark and the target watermark is greater than the matching degree between the first watermark and the target watermark.
[0051] For more detailed implementation of the diffusion model watermark embedding device, refer to the description of any of the implementation manners of the first aspect above, and the content of the following specific embodiments, which will not be repeated here.
[0052] In a third aspect, the present application provides a chip, comprising: a processor and a power supply circuit; the power supply circuit is configured to supply power to the processor, and the processor is configured to execute the method in the first aspect or any possible implementation manner of the first aspect; and / or the processor is configured to execute the method in the second aspect or any possible implementation manner of the second aspect.
[0053] In a fourth aspect, the present application provides a computing device. The computing device comprises a memory and a processor, the memory is configured to store computer instructions; when the processor executes the computer instructions, the method in the first aspect or any possible implementation manner of the first aspect is implemented; and / or when the processor executes the computer instructions, the method in the second aspect or any possible implementation manner of the second aspect is implemented.
[0054] In a fifth aspect, the present application provides a computing device cluster. The computing device cluster comprises at least one computing device, and the computing device comprises a memory and a processor, the memory is configured to store computer instructions; when the processor executes the computer instructions, the method in the first aspect or any possible implementation manner of the first aspect is implemented.
[0055] In a sixth aspect, the present application provides a computer readable storage medium, the storage medium stores computer programs or instructions; when the computer programs or instructions are executed by a processing device, the method in the first aspect or any possible implementation manner of the first aspect is implemented; and / or when the computer programs or instructions are executed by the processing device, the method in the second aspect or any possible implementation manner of the second aspect is implemented.
[0056] In a seventh aspect, the present application provides a computer program product, the computer program product comprises computer programs or instructions; when the computer programs or instructions are executed by a processing device, the method in the first aspect or any possible implementation manner of the first aspect is implemented; and / or when the computer programs or instructions are executed by the processing device, the method in the second aspect or any possible implementation manner of the second aspect is implemented.
[0057] The beneficial effects of the second aspect to the seventh aspect above can refer to the first aspect or any possible implementation manner of the first aspect, and will not be repeated here. On the basis of the implementation manners provided by the application in the above aspects, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF DRAWINGS
[0058] FIG. 1 is a schematic diagram of the architecture of a computer system provided by the application;
[0059] FIG. 2 is a schematic diagram of the structure of a chip provided by the application;
[0060] FIG. 3 is a schematic diagram of the flow of a diffusion model watermark embedding method provided by the application;
[0061] FIG. 4 is a schematic diagram of a user interface provided by the application;
[0062] FIG. 5 is a schematic diagram of the model structure of a teacher model provided by the application;
[0063] FIG. 6 is a schematic diagram of the flow of a teacher model training method provided by the application;
[0064] FIG. 7 is a schematic diagram of a user interface provided by the application;
[0065] FIG. 8a is a schematic diagram of the flow of a distillation learning method provided by the application;
[0066] FIG. 8b is a schematic diagram of the flow of a diffusion model verification method provided by the application;
[0067] FIG. 9 is a schematic diagram of the structure of a diffusion model watermark embedding device provided by the application;
[0068] FIG. 10 is a schematic diagram of the structure of a diffusion model watermark embedding device provided by the application;
[0069] FIG. 11 is a schematic diagram of the structure of a computing device cluster provided by the application;
[0070] FIG. 12 is a schematic diagram of the connection between computing devices provided by the application. DETAILED DESCRIPTION
[0071] To solve the problem that the parameters of the diffusion model are adjusted to generate the watermark when the network fine-tuning diffusion model is used, resulting in that more model parameters of the fine-tuned diffusion model are adjusted to generate the watermark, and the quality of the image generated by the fine-tuned diffusion model is poor, and the availability of the fine-tuned diffusion model is poor. The present application provides a diffusion model watermark embedding method. The computing device adjusts the first diffusion model by the encoder and the decoder in the teacher model, realizes that the first diffusion model learns the features of different stages, such as the features of the stage that the encoder generates the three-dimensional model with the target watermark, and the features of the stage that the decoder extracts the watermark in the three-dimensional model, and then obtains the second diffusion model which can preserve the model parameters of generating the three-dimensional model as much as possible and update the model parameters of generating the watermark, thereby improving the availability of the second diffusion model, i.e., the robustness.
[0072] In order to facilitate understanding, first, the technical terms involved in the present application are introduced.
[0073] Watermark: content reflecting the source of the diffusion model. The watermark can include but is not limited to one or more of the following: picture, text, audio, video, and bit stream. The watermark can be content specified by the owner / provider / generator / producer of the diffusion model.
[0074] Diffusion model: a generative model that learns the underlying data distribution by adding noise and iterative denoising. In the present application, the diffusion model can generate a three-dimensional model according to the text.
[0075] Model parameters: parameters that can be learned and adjusted inside the model, used to describe the relationship between data features and target variables.
[0076] The following will illustrate the scenarios to which the embodiments of the present application can be applied with reference to the accompanying drawings.
[0077] FIG. 1 is a schematic diagram of the architecture of a computer system provided by the present application. The computer system 100 includes a computing device 110. The computing device 110 can be an electronic device with computing capability or a virtual device with computing capability. In the case where the computing device 110 is an electronic device with computing capability, the computing device 110 can be a server, a personal computer, a tablet computer, etc. In the case where the computing device 110 is a virtual device with computing capability, the computing device 110 can be a virtual machine, a container, etc.
[0078] In one possible example, a user can input a first diffusion model and a target watermark to the computing device 110, and the computing device 110 trains a teacher model using training data and the target watermark, and then fine-tunes the first diffusion model using the trained teacher model, and the computing device 110 outputs a second diffusion model (i.e., the fine-tuned first diffusion model).
[0079] In another possible example, the user can input a target watermark to the computing device 110, the computing device 110 trains the teacher model by using the training data and the target watermark, and then fine-tunes the first diffusion model in the plurality of reference diffusion models by using the trained teacher model, and the computing device 110 outputs the second diffusion model.
[0080] The computing device 110 includes a communication interface 114, a processor 111, and a memory 112. The communication interface 114 is configured to communicate with devices located outside the computing device 110. For example, the computing device 110 receives the first diffusion model or the target watermark through the communication interface 114, and the computing device 110 processes (e.g., trains) the teacher model by using the received target watermark, and then outputs the processing result (e.g., the second diffusion model) through the communication interface 114. The communication interface 114 can be an input / output (I / O) interface.
[0081] The processor 111 is the operation core and control core of the computing device 110, and can include a central processing unit (CPU), a specific integrated circuit, other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, and the like. In actual applications, the computing device 110 can also include multiple processors. The processor 111 can include one or more processor cores. The processor 111 is installed with an operating system and other software programs, so that the processor 111 can access the memory 112 and various peripheral component interconnect express (PCIe) devices.
[0082] The processor 111 is connected to the memory 112 through a double data rate (DDR) bus or other types of buses. The memory 112 is the main memory of the computing device 110. The memory 112 is usually used to store various running software in the operating system, received input data, and output results obtained by processing the input data, etc. In order to improve the access speed of the processor 111, the memory 112 needs to have the advantage of fast access speed. In the conventional computer device, a dynamic random access memory (DRAM) is usually used as the memory 112. In addition to the DRAM, the memory 112 can also be other random access memories, such as a static random access memory (SRAM), etc. In addition, the memory 112 can also be a read only memory (ROM). For the read only memory, for example, it can be a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), etc. The embodiment does not limit the number and type of the memory 112.
[0083] Optionally, in order to store the data persistently, the computer system 100 is also provided with a data storage system 113, which can be located outside the computing device 110 (as shown in FIG. 1) and exchange data with the computing device 110 through a network. Optionally, the data storage system 113 can also be located inside the host, such as the data storage system 113 exchanges data with the processor 111 through the bus 116. At this time, the data storage system 113 is a hard disk.
[0084] Optionally, the computer system 100 can also include a client device 120. The user can input the first diffusion model or the target watermark to the computing device 110 through the client device 120, and the computing device 110 sends the processing result (such as the second diffusion model) to the user through the client device 120. The client device 120 can be a terminal device, including but not limited to a personal computer, a server, a mobile phone, a tablet computer, or a smart car, etc.
[0085] Optionally, the computer system 100 can also include an acceleration device 115. The acceleration device 115 is configured to perform the watermark embedding task of the first diffusion model. The processor 111 sends the received watermark embedding task to the acceleration device 115, and the acceleration device 115 sends the processing result (e.g., the second diffusion model) to the processor 111 after completing the watermark embedding task according to the training data. As shown in FIG. 1, the acceleration device 115 can be directly inserted into a card slot on the mainboard of the computing device 110, and exchange data with the processor 111 through the bus 116. It should be noted that the bus 116 in FIG. 1 can also be replaced by a bus acceleration device 115 of a compute express link (CXL) protocol, a universal serial bus (USB) protocol, or other protocols for data transmission.
[0086] In addition, the acceleration device 115 described above can not be directly inserted into the card slot on the mainboard of the computing device 110, but can be located in an acceleration device. For example, the acceleration device is a device independent of the computing device 110, such as an acceleration card. At this time, the computing device 110 can be connected with the acceleration device 115 through a wired network such as a network cable, or can be connected with the acceleration device 115 through a wireless network such as a wireless hotspot or Bluetooth. For example, the acceleration device 115 is configured to complete the watermark embedding task of the first diffusion model, such as training a teacher model using training data, and then fine-tuning the first diffusion model using the teacher model. The acceleration device 115 can be implemented by one or more chips. For example, the chip includes any one of a CPU, a graphics processing unit (GPU), a neural-network processing units (NPU), a tensor processing unit (TPU), an FPGA, and an ASIC. The GPU is also called a display core, a visual processor, or a display chip, which is a microprocessor specially designed for image operation in personal computers, workstations, game consoles, and some mobile devices (such as tablet computers and smart phones). The NPU simulates human neurons and synapses at the circuit layer, and directly processes large-scale neurons and synapses with a deep learning instruction set, and a single instruction completes the processing of a group of neurons. The ASIC is suitable for single-purpose integrated circuit products.
[0087] For example, the processor 111 in FIG. 1 can be implemented by a chip, as shown in FIG. 2, which is a structural schematic diagram of a chip provided in the present application. For example, the chip 200 includes a core 201, a CPU 202, a system buffer 203, a DDR 204, and an input / output (I / O) device 205.
[0088] The CPU 202 is configured to accept an AI task (such as a watermark embedding task of a first diffusion model) and call the core 201 to execute the task. In the case where the chip 200 has multiple cores 201, the CPU 202 is further configured to undertake a scheduling task. For example, the CPU 202 can be implemented by an advanced RISC machine (ARM) processor, which has a small size, low power consumption, adopts a 32-bit RISC, and has simple and flexible addressing. Of course, in some embodiments, the CPU 202 can also be implemented by other processors.
[0089] The core 201 is configured to provide the computing power required in the watermark embedding task of the first diffusion model. In an optional case, the core 201 includes a load / store unit (LSU), a cube computing unit, a scalar computing unit, a vector computing unit, and a buffer. The LSU is configured to load data to be processed and store processed data, and can also be used for read / write management of internal data in the core between different buffers, and for performing some format conversion operations. The cube computing unit is configured to provide the core computing power of matrix multiplication. The scalar computing unit is a single instruction single data (SISD) processor, which processes only one data (usually an integer or a floating point number) at the same time. The vector computing unit, also known as an array processor, is a processor that can directly operate a group of arrays or vectors for computation. The number of buffers can be one or more. For example, the buffer mainly refers to a level 1 buffer (L1 buffer). The buffer is used to temporarily store some data that needs to be repeatedly used by the core 201, so as to reduce read / write from the bus. In addition, the implementation of some data format conversion functions also requires that the source data be located in the buffer. In the present embodiment, since the buffer is located in the core, the distance between the cube computing unit in the core and the storage area where the data is located is shortened, the access of the cube computing unit to the DDR 204 is reduced, and thus the data access delay and the data processing delay of the core are reduced.
[0090] System buffer 203, mainly refers to a level 2 buffer (L1 buffer or L2 cache), which is used to temporarily store input data (such as training data) passing through the chip, intermediate results or final results (such as a second diffusion model).
[0091] DDR 204 is an off-chip memory, which can also be replaced by high bandwidth memory (HBM) or other off-chip memories. DDR 204 is located between the chip and the external memory, overcoming the access speed limitation when the computing resource shares the memory for reading and writing.
[0092] The I / O device 205 included in the chip 200 refers to hardware for data transmission, which can also be understood as a device connected to the I / O interface. Common I / O devices include network cards, printers, keyboards, mice, etc. All external storage can also be used as I / O devices, such as hard disks, floppy disks, optical disks, etc.
[0093] The core 201, CPU 202, system buffer 203, DDR 204, and I / O device 205 are connected through a bus. The bus can include a path for transmitting information between the above components (such as CPU 202, system buffer 203). In addition to the data bus, the bus can also include a power bus, a control bus, and a status signal bus, etc. However, for the purpose of clear illustration, the bus can be a PCIe bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. For example, the core 201 can access these I / O devices 205 through a PCIe bus. The core 201 is connected to the system buffer 203 through a DDR bus. Here, different system buffers 203 can use different data buses to communicate with the core 201, so the DDR bus can also be replaced by other types of data buses, and the bus type is not limited in the present application.
[0094] For example, after the CPU 202 loads data to be processed by an AI task (e.g., a first diffusion model) into the DDR 204, the LSU in the core 201 reads (loads) the data from the DDR 204, and obtains a processing result (e.g., a second diffusion model) by training the first diffusion model. After the processing result is obtained, the LSU stores the processing result in the DDR 204, and the network interface card sends the processing result to the client device 120, or to the data storage system 113 for persistent storage.
[0095] It is worth noting that the acceleration device 115 shown in FIG. 1 can also be implemented by the chip 200 shown in FIG. 2, which is not limited in the present application.
[0096] It can be understood that the structure shown in the embodiment does not constitute a specific limitation on the computing device. In other embodiments, the computing device and the chip can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0097] In the embodiments of the present application, the teacher model can be deployed in the computing device or the processor (or chip) in the computing device, and other neural network models or algorithm models with three-dimensional model generation function can also be deployed, such as the reference diffusion model, which is not limited in the present application. In this case, the computing device receives the first diffusion model and the target watermark input by the user to the computing device, and then trains the encoder and the decoder in the teacher model using the target watermark and the training data, so as to segmentally fine-tune the first diffusion model using the encoder and the decoder in the teacher model, and obtain the second diffusion model.
[0098] The implementation of the diffusion model watermark embedding method provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0099] FIG. 3 is a flowchart of a diffusion model watermark embedding method provided by the present application, which can be executed by a computing device or a chip or processor in the computing device. The computing device can be the computing device 110, the client device 120, or the acceleration device 115 shown in FIG. 1. The chip or processor in the computing device can be the chip shown in FIG. 2, etc. The hardware implementation of the computing device can refer to the description of FIG. 1, and the implementation of the chip or processor can refer to the description of FIG. 2, which is not described here. In some optional examples, the diffusion model watermark embedding method can also be executed by other computing devices or computer clusters, and the hardware implementation of the computing device can refer to the description of FIG. 1 and FIG. 2, which is not described here. The computer cluster includes a plurality of computing devices.
[0100] In a possible scenario, the computing device 110 can be one or more of N computing devices managed by a cloud computing platform, N being an integer greater than or equal to 2.
[0101] In a possible scenario, the computing device 110 can be a terminal, a server, or the like.
[0102] Here, taking the computing device 110 as an example, the diffusion model watermark embedding method provided by the embodiment includes the following steps S310-S330, as shown in FIG. 3.
[0103] S310, the computing device 110 acquires a first diffusion model.
[0104] The first diffusion model is used to generate a three-dimensional model.
[0105] In a possible scenario, the first diffusion model is a pre-trained diffusion model. That is, the first diffusion model can generate a corresponding three-dimensional model according to a prompt word.
[0106] In a possible example, the prompt word is an airplane, and the first diffusion model can generate a three-dimensional model of an airplane according to the prompt word.
[0107] In a possible example, the prompt word is a train, and the first diffusion model can generate a three-dimensional model of a train according to the prompt word.
[0108] It is worth noting that the first diffusion model being a pre-trained diffusion model is only a possible scenario provided by the embodiment, and should not be construed as a limitation on the present application. In other embodiments of the present application, the first diffusion model can also be an untrained diffusion model.
[0109] In a possible scenario, the three-dimensional model can be represented by any of the following: a point cloud, a polygon mesh, a voxel grid.
[0110] Two possible implementation manners are provided below for the content of the computing device 110 acquiring the first diffusion model.
[0111] In the first possible implementation manner, the computing device 110 provides a user interface including a second control component, and acquires the first diffusion model in response to a user triggering operation on the second control component.
[0112] In a possible example, the client device 120 accesses the computing device 110 through an API provided by the computing device 110, and displays the user interface on the terminal. Further, the computing device 110 acquires the user triggering operation on the second control component on the user interface through various input devices (keyboard, mouse, touch screen, etc.) connected with the client device 120.
[0113] As shown in FIG. 4, FIG. 4 is a schematic diagram of a user interface provided by the present application. The user interface includes a second control component for uploading the first diffusion model, and a control component for exporting the second diffusion model.
[0114] For the specific implementation of the trigger operation, three possible examples are provided below.
[0115] Example 1, the trigger operation can be the user's confirmation of the second control component through the keyboard, such as the user triggering the enter key.
[0116] Example 2, the trigger operation can be the user's click on the second control component through the mouse.
[0117] For example, the user drags a file including the first diffusion model to a designated position on the user interface, so that the terminal sends the aforementioned file including the first diffusion model to the computing device 110, and the computing device 110 directly obtains the first diffusion model.
[0118] For another example, after the user clicks on the second control component on the user interface, the computing device 110 provides a floating window on the user interface, which includes the storage paths of the first diffusion models (such as the storage paths of the first diffusion model 1, the first diffusion model 2, the first diffusion model 3, …, the first diffusion model n). After the user selects a designated first diffusion model (such as the first diffusion model 1) on the floating window, the user clicks on the "Confirm" control component through the mouse, so that the terminal sends the storage path of the aforementioned first diffusion model 1 to the computing device 110, and the computing device 110 obtains the first diffusion model 1 from the storage path of the first diffusion model 1. The storage path indicates the storage location of the first diffusion model.
[0119] Example 3, the trigger operation can be the user's input operation in the second control component through the keyboard, etc., such as inputting the corresponding storage path.
[0120] The above examples are only optional contents provided by the present embodiment, and should not be understood as a limitation of the present application. In other embodiments of the present application, the trigger operation can also be a hand-in-air operation, voice control, knuckle tapping, etc.
[0121] In a second possible implementation, the computing device 110 obtains the first diffusion model from a plurality of reference diffusion models, each of the plurality of reference models being used to generate a three-dimensional model.
[0122] In one possible case, the plurality of reference diffusion models are all pre-trained diffusion models.
[0123] The plurality of reference diffusion models differ in that the use scenarios are different and the model structures are different.
[0124] In one possible example, a first reference diffusion model in the plurality of reference diffusion models is used to generate a three-dimensional model of an office appliance, and a second reference diffusion model is used to generate a three-dimensional model of a transportation tool. In other words, different types of training sets are used when training the reference diffusion models.
[0125] The network structure of a third reference diffusion model in the plurality of reference diffusion models includes a convolution layer, a pooling processing layer, and a feature extraction network layer, in addition to the network structure of a fourth reference diffusion model.
[0126] Three possible examples are provided below for the computing device 110 to obtain the content of the first diffusion model from the plurality of reference diffusion models.
[0127] Example 1: The computing device 110 randomly determines the first diffusion model from the plurality of reference diffusion models.
[0128] Example 2: The computing device 110 determines the first diffusion model from the plurality of reference diffusion models according to the user's instructions.
[0129] For example, the computing device 110 provides a control component on the user interface that includes information of the plurality of reference diffusion models, and in response to the user's triggering operation on the aforementioned control component, the first diffusion model in the plurality of reference diffusion models is obtained.
[0130] Example 3: The computing device 110 determines the first diffusion model from the plurality of reference diffusion models according to the label determined by the user.
[0131] Each of the plurality of reference diffusion models has one or more labels.
[0132] The computing device 110 determines at least one reference diffusion model from the plurality of reference diffusion models that meets the label determined by the user, and the label of the at least one reference diffusion model matches the label determined by the user. If the at least one reference diffusion model is only one reference diffusion model, the computing device 110 takes the reference diffusion model as the first diffusion model. If the at least one reference diffusion model has multiple reference diffusion models, the computing device 110 determines the content of the first diffusion model, which can refer to the descriptions of the aforementioned examples 1 and 2, and will not be described here.
[0133] S320: The computing device 110 trains the first diffusion model according to the encoder in the teacher model to obtain an intermediate diffusion model.
[0134] The aforementioned encoder is used to generate a three-dimensional model with a target watermark, and the intermediate diffusion model can generate a three-dimensional model with a first watermark.
[0135] In a possible implementation, the teacher model is trained by using a training data set and the target watermark. The training data set includes three-dimensional models of different types, and the types include one or more of the following: a transportation tool, an office appliance, and a building.
[0136] In a possible implementation, the types also include an animal, a plant, a piece of clothing, a toy, and a household item.
[0137] For details of the implementation, refer to the description of FIG. 6 below, which will not be repeated here.
[0138] In a possible implementation, the computing device 110 trains the first diffusion model according to the encoder in the teacher model to obtain an intermediate diffusion model, including:
[0139] The computing device 110 migrates the ability of the encoder to generate a three-dimensional model with a target watermark to the first diffusion model by model distillation, adjusts the model parameters of the first diffusion model, and obtains an intermediate diffusion model (the first diffusion model after the model parameter adjustment).
[0140] Model distillation, also known as knowledge distillation, refers to fixing the model parameters of the trained teacher model, setting a loss, and making the first diffusion model gradually approach the performance characteristics (generating a three-dimensional model with a target watermark) of the teacher model during training. After the training reaches a preset round or the loss converges, the intermediate diffusion model is obtained.
[0141] The loss refers to the loss between different output results obtained by the encoder in the teacher model and the first diffusion model for the same input.
[0142] For details of the implementation, refer to the description of FIG. 8a below, which will not be repeated here.
[0143] S330, the computing device 110 trains the intermediate diffusion model according to the decoder in the teacher model to obtain a second diffusion model.
[0144] The decoder is used to extract a watermark in a three-dimensional model, and the second diffusion model can generate a three-dimensional model with a second watermark. The matching degree between the second watermark and the target watermark is greater than the matching degree between the first watermark and the target watermark.
[0145] In a possible implementation, the computing device 110 trains the intermediate diffusion model according to the decoder in the teacher model to obtain a second diffusion model, including:
[0146] The computing device 110 extracts the first watermark from the output of the intermediate diffusion model by using the decoder, and further updates the parameters of the intermediate diffusion model according to the loss between the first watermark and the target watermark, to obtain a second diffusion model.
[0147] For the content of the present implementation, reference can be made to the description of the following FIG. 8a, which will not be repeated here.
[0148] In one possible embodiment, the present application provides a network structure of a teacher model. First, as shown in FIG. 5, FIG. 5 is a schematic diagram of a model structure of a teacher model provided by the present application.
[0149] The teacher model 500 shown in FIG. 5 includes an encoder 510 and a decoder 520.
[0150] The encoder 510 includes a local feature extraction network 511, a global feature extraction network 512, a convolutional layer 513, a convolutional layer 514, a pooling layer 515, and a multi-layer perceptron (MLP) 516.
[0151] The decoder 520 includes a global feature extraction network 521, a convolutional layer 522, a pooling layer 523, and an MLP 524.
[0152] And, the convolutional layer 530 and the convolutional layer 540 are respectively arranged before the encoder 510 and the decoder 520.
[0153] In one possible case, the local feature extraction network 511 includes a graph convolutional network (GCN).
[0154] In one possible case, the global feature extraction network 512 includes a GCN.
[0155] In one possible case, the global feature extraction network 521 includes a GCN.
[0156] In one possible case, the pooling layer 515 is a max pool layer.
[0157] In one possible case, the pooling layer 523 is a max pool layer.
[0158] It is worth noting that f l represents a vector obtained after the convolutional layer 513 in the encoder 510, f1 g represents a vector obtained after the pooling layer 515 in the encoder 510, represents a vector obtained after the pooling layer 523 in the decoder 520, and m represents a target watermark. predicates a predicted watermark.
[0159] The convolutional layer 513, the convolutional layer 514, the convolutional layer 530, and the convolutional layer 540 can have the same or different sizes (sizes of convolution kernels), and the present application is not limited in this regard.
[0160] In a possible embodiment, based on the model structure of the teacher model shown in FIG. 5, a training method of the teacher model is provided as follows. As shown in FIG. 6, FIG. 6 is a flowchart of the training method of the teacher model provided by the present application. The method includes the following steps S610-S630.
[0161] S610, the computing device 110 acquires a training data set and a target watermark.
[0162] The training data set includes three-dimensional models of different types.
[0163] In a possible case, the three-dimensional models in the training data set do not have watermarks.
[0164] In view of the content of the training data set acquired by the computing device 110, three possible implementation manners are provided as follows.
[0165] In a first possible implementation manner, the training data set is stored in the data storage system 113, and the computing device 110 acquires the training data set from the data storage system 113.
[0166] In a second possible implementation manner, the training data set is stored in the memory 112, and the computing device 110 acquires the training data set from the memory 112.
[0167] In a third possible implementation manner, the computing device 110 acquires the training data set sent by the client device 120.
[0168] In view of the content of the target watermark acquired by the computing device 110, three possible implementation manners are provided as follows.
[0169] In a first possible implementation manner, the computing device 110 uses a preset watermark as the target watermark.
[0170] In a possible case, the user interface provided by the computing device 110 includes the preset target watermark.
[0171] In a second possible implementation manner, the computing device 110 uses a randomly generated watermark as the target watermark.
[0172] In a possible case, the user interface provided by the computing device 110 includes the randomly generated target watermark.
[0173] In a third possible implementation, the computing device 110 acquires the target watermark sent by the client device 120.
[0174] In a possible case, as shown in FIG. 7, which is a schematic diagram of a second user interface provided by the present application, the user interface shown in FIG. 7 adds a first control component for receiving user input of the target watermark to the user interface shown in FIG. 4, and the computing device 110 acquires the target watermark sent by the client device 120 in response to the triggering operation of the user on the first control component.
[0175] In other words, the user interface provided by the computing device 110 includes the first control component and the second control component.
[0176] In a possible example, the user interface described above further includes a control component for specifying the watermark, if the user triggers “yes” in the control component, the user needs to input the target watermark in the first control component; if the user triggers “no” in the control component, the computing device 110 acquires the target watermark in the first possible implementation or the second possible implementation described above.
[0177] For example, if the client device 120 accesses the computing device 110 and sends user information to the computing device 110, the computing device 110 takes the ID in the user information as the target watermark, and the user interface provided by the computing device 110 includes the target watermark.
[0178] It is worth noting that the target watermark acquired by the computing device 110 can be in the form of a picture, text, audio, video, or bit stream.
[0179] In S620, the computing device 110 inputs the training data set and the target watermark into the teacher model, and the encoder outputs a three-dimensional model with the watermark, and the decoder outputs a predicted watermark.
[0180] The computing device 110 divides the data in the training data set into multiple batches of data, and then inputs the multiple batches of data into the teacher model for training in sequence.
[0181] In a possible case, the computing device 110 converts the target watermark from a picture, text, audio, or video into a bit stream before inputting the target watermark into the teacher model.
[0182] In a possible example, the computing device 110 can encode the picture, text, audio, or video to obtain the corresponding bit stream.
[0183] The following describes an example of inputting one batch of data in the multiple batches of data into the teacher model for training, and the batch of data includes a three-dimensional model A.
[0184] The computing device 110 inputs the three-dimensional model A into the teacher model 500. The computing device 110 extracts the feature A in the three-dimensional model through the convolutional layer 530, and then inputs the feature into the local feature extraction network 511 and the global feature extraction network 512 respectively to obtain the corresponding features B and C. The computing device 110 processes the feature B by using the convolutional layer 513 to obtain f l , and processes the feature B by using the convolutional layer 514 to obtain the feature D, and the feature D is processed by the pooling layer 515 to obtain f1 g . The computing device 110 inputs f l , f1 g and the target watermark fused feature into the MLP, and outputs the three-dimensional model with the watermark.
[0185] The computing device 110 further inputs the three-dimensional model with the watermark into the convolutional layer 540 to extract the feature E of the three-dimensional model with the watermark. The feature E is processed by the global feature extraction network 521 to obtain the feature F. The computing device 110 inputs the feature F into the convolutional layer 522 to obtain the feature G, and inputs the feature G into the pooling layer 523 to obtain f2 . The computing device 110 inputs f into the MLP to output the predicted watermark.
[0186] It is worth noting that the target watermark input into the teacher model is in the form of a bit stream, such as 01001011.
[0187] It is worth noting that in the actual training process, each batch of data will include a large number of three-dimensional models.
[0188] S630, the computing device 110 updates the model parameters of the encoder according to the loss A between the three-dimensional model in the training data set and the corresponding three-dimensional model with the watermark, and updates the model parameters of the encoder and the model parameters of the decoder according to the loss B between the predicted watermark and the target watermark, so as to obtain the trained teacher model.
[0189] The trained teacher model is the teacher model in the content shown in FIG. 3.
[0190] In a possible case, the above loss A can be calculated by using the Euclidean distance, the chamfer distance (CD), the mean square error, the cross-entropy loss function, etc., which are not limited by the present application.
[0191] The following takes the example of calculating the loss between the three-dimensional model in the training data set and the corresponding three-dimensional model with the watermark by using the Euclidean distance by the computing device 110.
[0192] The three-dimensional model in the training data set and the three-dimensional model with watermark are in the form of point cloud. The three-dimensional model in the training data set includes a plurality of point clouds 1, and the three-dimensional model with watermark includes a plurality of point clouds 2. The computing device 110 determines the distance between the point cloud 1 and the point cloud 2 by using the Euclidean distance, and then determines the loss A according to the distance between the plurality of point clouds 1 and the plurality of point clouds 2. The point cloud 1 and the point cloud 2 that determine the distance have the following corresponding relationship: the position of the point cloud 1 in the three-dimensional model in the training data set corresponds to the position of the point cloud 2 in the three-dimensional model with watermark.
[0193] For example, the position of the point cloud 1 in the three-dimensional model in the training data set is the outermost corner on the right upper corner, and the position of the point cloud 2 in the three-dimensional model with watermark is the outermost corner on the right upper corner.
[0194] In a possible example, the reciprocal of the sum of the distances between the plurality of point clouds 1 and the plurality of point clouds 2 is the loss A.
[0195] In a possible example, the computing device 110 calculates the distance between the plurality of point clouds 1 and the plurality of point clouds 2 according to a preset formula to obtain the loss A.
[0196] In a possible example, the loss B can be calculated by using a hamming distance, an edit distance, a cross-entropy loss function, etc., which is not limited herein.
[0197] The computing device 100 can update the parameters of the encoder 510 by using the loss A (L recon ) between the three-dimensional model A and the three-dimensional model with watermark, and update the parameters of the teacher model (the encoder and the decoder) by using the loss B (L w ) between the predicted watermark and the target watermark.
[0198] The above is only one round of training of the teacher model by the computing device 110 in the multiple rounds of training. The content of each round of training in the multiple rounds of training can be referred to the above description, which is not repeated here. The computing device 110 performs multiple rounds of training on the teacher model until the L recon and / or the L w converge, the data in the training data set is used up, or the set number of training rounds is reached, so as to obtain the trained teacher model.
[0199] For the content of obtaining the intermediate diffusion model in S320 and obtaining the second diffusion model in S330, a possible example is provided as follows. As shown in FIG. 8a, FIG. 8a is a flowchart of the distillation learning method provided by the present application. The content shown in FIG. 8a can include the following steps S810-S850.
[0200] S810, the computing device 110 inputs the first three-dimensional model in the fine-tuning dataset into the encoder and the first diffusion model, and outputs a first reconstructed three-dimensional model and a second reconstructed three-dimensional model.
[0201] The first three-dimensional model is any one of a plurality of three-dimensional models included in the fine-tuning dataset.
[0202] In a possible case, the three-dimensional model in the fine-tuning dataset does not have a watermark. The fine-tuning dataset in this case can refer to the content of the training dataset described above, which will not be repeated here.
[0203] In another possible case, the three-dimensional model in the fine-tuning dataset has a watermark. The watermark can be a bit stream of any content, such as 1010111, which is not limited in the present application.
[0204] In a possible implementation, the computing device 110 inputs the first three-dimensional model in the fine-tuning dataset into the encoder and the first diffusion model, and outputs the first reconstructed three-dimensional model and the second reconstructed three-dimensional model, including: the computing device 110 inputs the first three-dimensional model into the encoder, and outputs the first reconstructed three-dimensional model; and inputs the first three-dimensional model into the first diffusion model, and outputs the second reconstructed three-dimensional model.
[0205] The first reconstructed three-dimensional model and the second reconstructed three-dimensional model have a watermark.
[0206] S820, the computing device 110 updates the model parameters of the first diffusion model according to the loss between the first reconstructed three-dimensional model and the second reconstructed three-dimensional model, to obtain an intermediate diffusion model.
[0207] In a possible implementation, the computing device 110 updates the model parameters of the first diffusion model according to the loss between the first reconstructed three-dimensional model and the second reconstructed three-dimensional model, to obtain an intermediate diffusion model, including: the computing device 110 calculates the loss (L denoising ) between the first reconstructed three-dimensional model and the second reconstructed three-dimensional model, and then performs back propagation on the first diffusion model using the loss to update the model parameters of the first diffusion model, thereby obtaining the intermediate diffusion model.
[0208] L denoising It can be calculated by Euclidean distance, Chambolle distance, mean square error, cross-entropy loss function, etc., which is not limited in the present application.
[0209] It is worth noting that the above is only described by taking one three-dimensional model (the first three-dimensional model) as an example. In the actual training process, the computing device 110 will perform multiple rounds of training using the fine-tuning data set. In each round of training, the computing device 110 inputs a batch of data into the encoder and the first diffusion model, outputs a plurality of sets of reconstructed three-dimensional models, and then updates the model parameters of the first diffusion model according to the loss between the plurality of sets of reconstructed three-dimensional models, to obtain an intermediate diffusion model.
[0210] The above batch of data includes at least one three-dimensional model. The computing device 110 updates the model parameters of the first diffusion model once for each batch of data input into the encoder and the first diffusion model, that is, the model parameters of the first diffusion model are updated once for each round of training, until the loss between the reconstructed three-dimensional models (such as the first reconstructed three-dimensional model and the second reconstructed three-dimensional model) converges, reaches a set training round, or the three-dimensional models in the fine-tuning data set are all used up, to obtain the intermediate diffusion model.
[0211] S830, the computing device 110 inputs the second three-dimensional model in the fine-tuning data set into the intermediate diffusion model to obtain a third reconstructed three-dimensional model.
[0212] The third reconstructed three-dimensional model includes the first watermark, and the second three-dimensional model is any one of the plurality of three-dimensional models included in the fine-tuning data set.
[0213] In one possible example, the computing device 110 inputs a batch of data in the fine-tuning data set into the intermediate diffusion model to obtain a plurality of third reconstructed three-dimensional models. The batch of data in the fine-tuning data set includes a plurality of three-dimensional models, and the plurality of three-dimensional models include the second three-dimensional model.
[0214] For example, the fine-tuning data set includes a plurality of batches of data, and each batch of data in the plurality of batches of data includes at least one three-dimensional model.
[0215] S840, the computing device 110 inputs the third reconstructed three-dimensional model into the decoder in the teacher model to obtain the first watermark.
[0216] The computing device 110 extracts the watermark in the third reconstructed three-dimensional model through the decoder in the teacher model to obtain the first watermark.
[0217] S850, the computing device 110 updates the model parameters of the intermediate diffusion model according to the loss between the first watermark and the target watermark to obtain a second diffusion model.
[0218] In one possible implementation, the computing device 110 updates the model parameters of the intermediate diffusion model according to the loss between the first watermark and the target watermark to obtain the second diffusion model, including: the computing device 110 calculates the loss (L wm), and then the loss is used to back-propagate the intermediate diffusion model to update the model parameters of the intermediate diffusion model, so as to obtain the second diffusion model.
[0219] L wm The loss can be calculated by Hamming distance, edit distance, cross-entropy loss function, etc., which is not limited in the present application.
[0220] It is worth noting that the above is only described by taking one three-dimensional model (the second three-dimensional model) as an example. In the actual training process, the computing device 110 will perform multiple rounds of training using the fine-tuning data set. In each round of training, the computing device 110 inputs a batch of data into the intermediate diffusion model, outputs a plurality of reconstructed three-dimensional models, and then extracts the watermark in the plurality of reconstructed three-dimensional models using the decoder, updates the model parameters of the intermediate diffusion model using the loss between the plurality of watermarks and the target watermark, and obtains the second diffusion model.
[0221] The above batch of data includes at least one three-dimensional model, and each batch of data is processed by the intermediate diffusion model and the decoder. The computing device 110 will update the model parameters of the intermediate diffusion model once, that is, the model parameters of the intermediate diffusion model will be updated once for each round of training, until the loss between the first watermark and the target watermark converges, reaches a set number of training rounds, or the three-dimensional models in the fine-tuning data set are all used up, and the second diffusion model is obtained.
[0222] The above-mentioned second diffusion model can generate a three-dimensional model with a second watermark. Since the model parameters of the intermediate diffusion model are adjusted according to the loss between the first watermark and the target watermark during the training process, the matching degree between the second watermark and the target watermark is greater than the matching degree between the first watermark and the target watermark. The second diffusion model can both preserve the model parameters of the generated three-dimensional model as much as possible and update the model parameters of the generated watermark, thereby improving the usability (robustness) of the second diffusion model.
[0223] In one possible case, the matching degree between the second watermark and the target watermark is 100%, that is, the second watermark is the target watermark.
[0224] In one possible case, the matching degree between the second watermark and the target watermark is within a set interval. For example, greater than or equal to a first value and less than 1. The first value can be 70%. It is worth noting that the present application does not limit the first value, which can be greater than 70% or less than 70%.
[0225] In a possible embodiment, the computing device 110 can verify whether the model is watermarked. As shown in FIG. 8b, FIG. 8b is a flowchart of a diffusion model verification method provided by the present application. There are two suspected diffusion models 1 and 2 that may have stolen the second diffusion model, and the diffusion models 1 and 2 need to be watermarked.
[0226] In FIG. 8b, the user inputs the prompt words (such as airplane and train) into the diffusion models 1 and 2 on the client device 120, the diffusion model 1 outputs a three-dimensional model 1, and the diffusion model 2 outputs a three-dimensional model 2. Since the watermarks carried in the three-dimensional models 1 and 2 are implicit, the user can send the three-dimensional models 1 and 2 to the computing device 110 through the client device 120. The computing device 110 inputs the three-dimensional models 1 and 2 into the decoder in the teacher model to extract the watermarks 1 and 2, respectively. If the watermark 1 matches the target watermark of the second diffusion model of the user to a threshold, such as 70%, it can be determined that the diffusion model 1 is consistent with the second diffusion model, that is, the second diffusion model is stolen by others. If the watermark 2 does not match the target watermark of the second diffusion model of the user to the threshold, it can be determined that the diffusion model 2 is not consistent with the second diffusion model.
[0227] It can be understood that, in order to implement the functions in the above embodiments, the computing device includes corresponding hardware structures and / or software modules for performing various functions. Those skilled in the art should easily realize that, in combination with the units and method steps of the examples described in the embodiments disclosed in the present application, the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is implemented in hardware or computer software driven hardware depends on the specific application scenario and design constraints of the technical solution.
[0228] In the above, the diffusion model watermark embedding method provided by the embodiments of the present application is described in detail in combination with FIGS. 1 to 8b. In the following, the diffusion model watermark embedding device provided by the embodiments of the present application will be described in combination with FIG. 9.
[0229] FIG. 9 is a structural schematic diagram of a diffusion model watermark embedding device provided by the present application. The diffusion model watermark embedding device can be used to implement the functions of the computing device in the above diffusion model watermark embedding method embodiments, and thus can also achieve the beneficial effects possessed by the above method embodiments. In the present embodiment, the diffusion model watermark embedding device can be any device shown in FIG. 1, such as the computing device 110, the client device 120, and the acceleration device 115, or the computing device shown in subsequent embodiments, and can also be a module (such as a chip) applied to a device.
[0230] As shown in FIG. 9, the diffusion model watermark embedding device 900 includes a first acquisition module 910, a first training module 920, and a second training module 930. The diffusion model watermark embedding device 900 is configured to implement the functions of the computing device 110 in the method embodiments corresponding to FIGS. 1-8b. In one possible example, the diffusion model watermark embedding device 900 is configured to implement the specific processes of the diffusion model watermark embedding method described above, which include the following processes:
[0231] The first acquisition module 910 is configured to acquire a first diffusion model, and the first diffusion model is configured to generate a three-dimensional model.
[0232] The first training module 920 is configured to train the first diffusion model according to an encoder in a teacher model to obtain an intermediate diffusion model. The encoder is configured to generate a three-dimensional model with a target watermark, and the intermediate diffusion model is configured to generate a three-dimensional model with a first watermark.
[0233] The second training module 930 is configured to train the intermediate diffusion model according to a decoder in the teacher model to obtain a second diffusion model. The decoder is configured to extract a watermark in a three-dimensional model, and the second diffusion model is configured to generate a three-dimensional model with a second watermark, and the matching degree between the second watermark and the target watermark is greater than the matching degree between the first watermark and the target watermark.
[0234] To further implement the functions in the method embodiments shown in FIGS. 1-8b. The present application also provides a diffusion model watermark embedding device. As shown in FIG. 10, FIG. 10 is a structural schematic diagram of a diffusion model watermark embedding device provided by the present application. The diffusion model watermark embedding device 900 further includes a second acquisition module 940.
[0235] The second acquisition module 940 is configured to provide a user interface, the user interface includes a first control component, and in response to a triggering operation of the user on the first control component, the target watermark is acquired.
[0236] For more functions of the first acquisition module 910, the first training module 920, the second training module 930, and the second acquisition module 940, please refer to the description of the diffusion model watermark embedding method above, which will not be repeated here.
[0237] The diffusion model watermark embedding device 900 of the embodiments of the present application can be implemented by a software module. The diffusion model watermark embedding device 900 according to the embodiments of the present application can correspond to the execution of the diffusion model watermark embedding method described in the embodiments of the present application, and the above and other operations and / or functions of each module in the diffusion model watermark embedding device 900 are respectively for implementing the method flow in the foregoing figures. For the sake of brevity, they will not be repeated here.
[0238] It is worth noting that if the diffusion model watermark embedding device 900 is implemented by a software module, for example, the software module can be provided to users for use through a cloud service subscription mode, and users can select different subscription levels according to needs; for another example, the software module can also provide enterprise-level customization services with professional domain customization, interface personalization and extension functions according to the needs of users or enterprises.
[0239] The diffusion model watermark embedding device 900 of the embodiment of the present application can also be implemented by hardware, and the hardware refers to a computing device. The specific implementation of the computing device can refer to the description of FIG. 1, which will not be repeated here.
[0240] In addition, when the diffusion model watermark embedding device 900 is implemented by a diffusion model watermark embedding system, the diffusion model watermark embedding system can include the computing device 110 provided in FIG. 1 and a first diffusion model generation device. The first diffusion model generation device and the computing device communicate through wired or wireless connection. The first diffusion model generation device is used to generate a first diffusion model, and the computing device 110 is used to train the first diffusion model generated by the first diffusion model generation device to obtain a second diffusion model. For example, the computing device can be used to execute the diffusion model watermark embedding method provided in the foregoing embodiments.
[0241] The method steps in the embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in a random access memory (RAM), a flash memory, a ROM, a PROM, an EPROM, an electrically EPROM (EEPROM), a register, a hard disk, a mobile hard disk, a CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in the computing device 110. Of course, the processor and the storage medium can also exist as discrete components in a network device or a terminal device.
[0242] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device, which can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a desktop computer, a notebook computer, or a terminal device such as a smart phone.
[0243] As shown in FIG. 11, FIG. 11 is a structural diagram of a computing device cluster provided by the present application. The computing device cluster includes at least one computing device 110. The memory 112 in one or more computing devices 110 in the computing device cluster can store the same instructions for performing the diffusion model watermark embedding method.
[0244] In some possible implementation manners, the memory 112 in one or more computing devices 110 in the computing device cluster can also respectively store partial instructions for performing the diffusion model watermark embedding method. In other words, the combination of one or more computing devices 110 can collectively execute the instructions for performing the diffusion model watermark embedding method.
[0245] It should be noted that the memories 112 in different computing devices 110 in the computing device cluster can store different instructions, respectively used for performing partial functions of the diffusion model watermark embedding method. That is, the instructions stored in the memories 112 in different computing devices 110 can implement the functions of one or more of the first obtaining module 910, the first training module 920, and the second training module 930.
[0246] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network or a local area network, etc. FIG. 12 shows a possible implementation manner. As shown in FIG. 12, FIG. 12 is a connection diagram between computing devices provided by the present application, two computing devices 110A and 110B are connected through a network. Specifically, the communication interface in each computing device is connected with the network. In this type of possible implementation manner, the memory 112 in the computing device 110A stores instructions for performing the functions of the first obtaining module 910. Meanwhile, the memory 112 in the computing device 110B stores instructions for performing the functions of the first training module 920 and the second training module 930.
[0247] It should be understood that the functions of the computing device 110A shown in FIG. 12 can also be completed by multiple computing devices 110. Similarly, the functions of the computing device 110B can also be completed by multiple computing devices 110.
[0248] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions, capable of running on a computing device or stored in any available medium. When the computer program product runs on at least one computing device, the at least one computing device is caused to perform the diffusion model watermark embedding method described above.
[0249] The embodiments of the present application further provide a computer readable storage medium. The computer readable storage medium can be any available medium or data storage device that can be accessed by a computing device, such as a data center containing one or more available media. The available medium can be a magnetic medium, such as a floppy diskette, a hard disk drive, a magnetic tape, an optical medium, such as a digital video disc (DVD), or a semiconductor medium, such as a solid state drive, etc. The computer readable storage medium includes instructions that instruct the computing device to perform the diffusion model watermark embedding method.
[0250] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer programs or instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are performed. The computer can be a general purpose computer, a special purpose computer, a computer network, a network device, a user equipment or other programmable apparatus. The computer programs or instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer programs or instructions can be transferred from one website site, computer, server or data center to another website site, computer, server or data center through wired or wireless manner. The computer readable storage medium can be any available medium accessible by a computer or a data storage device, such as a server, data center, etc., integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, a magnetic tape, an optical medium, such as a digital video disc (DVD), or a semiconductor medium, such as a solid state drive (SSD).
[0251] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A diffusion model watermark embedding method characterized by, The method comprises: obtaining a first diffusion model, the first diffusion model being used to generate a three-dimensional model; training the first diffusion model according to an encoder in a teacher model to obtain an intermediate diffusion model; the encoder is used to generate a three-dimensional model with a target watermark, and the intermediate diffusion model can generate a three-dimensional model with a first watermark; training the intermediate diffusion model according to a decoder in the teacher model to obtain a second diffusion model; the decoder is used to extract a watermark in a three-dimensional model, and the second diffusion model can generate a three-dimensional model with a second watermark, and the matching degree between the second watermark and the target watermark is greater than the matching degree between the first watermark and the target watermark.
2. The method of claim 1, wherein, Before the training of the first diffusion model according to the encoder in the teacher model to obtain the intermediate diffusion model, the method further comprises: providing a user interface, the user interface comprising a first control component; in response to the triggering operation of the user on the first control component, obtaining the target watermark.
3. The method according to claim 1 or 2, characterized in that, The obtaining of the first diffusion model comprises: providing a user interface, the user interface comprising a second control component; in response to the triggering operation of the user on the second control component, obtaining the first diffusion model.
4. The method according to claim 1 or 2, characterized in that, The obtaining of the first diffusion model comprises: obtaining the first diffusion model from a plurality of reference diffusion models, each of the plurality of reference diffusion models being used to generate a three-dimensional model.
5. The method according to any one of claims 1 to 4, characterized in that, The teacher model is trained using a training data set and the target watermark, the training data set comprising three-dimensional models of different types, the types including one or more of the following: transportation tools, office supplies, buildings.
6. The method according to any one of claims 1 to 5, characterized in that, The training of the first diffusion model according to the encoder in the teacher model to obtain the intermediate diffusion model comprises: inputting a first three-dimensional model in a fine-tuning data set into the encoder and the first diffusion model to output a first reconstructed three-dimensional model and a second reconstructed three-dimensional model; the first three-dimensional model being any one of a plurality of three-dimensional models included in the fine-tuning data set; updating the model parameters of the first diffusion model according to the loss between the first reconstructed three-dimensional model and the second reconstructed three-dimensional model to obtain the intermediate diffusion model.
7. The method according to any one of claims 1 to 6, characterized in that, The training of the intermediate diffusion model according to the decoder in the teacher model to obtain the second diffusion model comprises: inputting a second three-dimensional model in a fine-tuning data set into the intermediate diffusion model to obtain a third reconstructed three-dimensional model; the third reconstructed three-dimensional model comprising the first watermark, and the second three-dimensional model being any one of a plurality of three-dimensional models included in the fine-tuning data set; inputting the third reconstructed three-dimensional model into the decoder to obtain the first watermark; updating the model parameters of the intermediate diffusion model according to the loss between the first watermark and the target watermark to obtain the second diffusion model.
8. The method according to any one of claims 1 to 7, characterized in that, The encoder comprises a local feature extraction network and a global feature extraction network.
9. The method of claim 8, wherein, The local feature extraction network comprises a graph convolutional neural network; and / or The global feature extraction network comprises the graph convolutional neural network.
10. A diffusion model watermark embedding apparatus characterized by comprising: The device comprises: The first acquisition module is configured to acquire a first diffusion model, the first diffusion model being configured to generate a three-dimensional model. The first training module is configured to train the first diffusion model according to an encoder in a teacher model to obtain an intermediate diffusion model, the encoder being configured to generate a three-dimensional model with a target watermark, and the intermediate diffusion model being capable of generating a three-dimensional model with a first watermark. The second training module is configured to train the intermediate diffusion model according to a decoder in the teacher model to obtain a second diffusion model, the decoder being configured to extract a watermark in a three-dimensional model, and the second diffusion model being capable of generating a three-dimensional model with a second watermark, a matching degree of the second watermark and the target watermark being greater than a matching degree of the first watermark and the target watermark.
11. The apparatus of claim 10, wherein, The apparatus further includes a second acquisition module. The second acquisition module is configured to provide a user interface, the user interface including a first control component, and in response to a triggering operation of the first control component by a user, to acquire the target watermark.
12. The apparatus of claim 10 or 11, wherein, The first acquisition module is specifically configured to provide a user interface, the user interface including a second control component, and in response to a triggering operation of the second control component by a user, to acquire the first diffusion model.
13. The apparatus of claim 10 or 11, wherein, The first acquisition module is specifically configured to acquire the first diffusion model from a plurality of reference diffusion models, each of the plurality of reference diffusion models being configured to generate a three-dimensional model.
14. The apparatus of any one of claims 10-13, wherein, The teacher model is trained using a training data set and the target watermark, the training data set including three-dimensional models of different types, the types including one or more of the following: a transportation tool, an office appliance, and a building.
15. The apparatus of any one of claims 10-14, wherein, The first training module is specifically configured to input a first three-dimensional model in a fine-tuning data set into the encoder and the first diffusion model, to output a first reconstructed three-dimensional model and a second reconstructed three-dimensional model, and to update model parameters of the first diffusion model according to a loss between the first reconstructed three-dimensional model and the second reconstructed three-dimensional model to obtain the intermediate diffusion model, the first three-dimensional model being any one of a plurality of three-dimensional models included in the fine-tuning data set.
16. The apparatus of any one of claims 10 to 15, wherein, The second training module is specifically configured to input a second three-dimensional model in the fine-tuning data set into the intermediate diffusion model to obtain a third reconstructed three-dimensional model, to input the third reconstructed three-dimensional model into the decoder to obtain the first watermark, to update model parameters of the intermediate diffusion model according to a loss between the first watermark and the target watermark to obtain the second diffusion model, and the third reconstructed three-dimensional model including the first watermark, the second three-dimensional model being any one of a plurality of three-dimensional models included in the fine-tuning data set.
17. The apparatus of any one of claims 10-16, wherein, The encoder includes a local feature extraction network and a global feature extraction network.
18. The apparatus of claim 17, wherein, The local feature extraction network includes a graph convolutional neural network, and / or the global feature extraction network includes the graph convolutional neural network.
19. A cluster of computing devices, characterized in that, The apparatus includes at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method of any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that, The storage medium has stored therein a computer program or instructions which, when executed by a computing device, implement the method of any one of claims 1 to 9.
21. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions, when executed by a computing device, implement the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Watermark generation, information processing and audio watermark generation model training methods and devices
CN116778935A
Watermark adding method and device, watermark identification method and device, equipment and readable storage medium
CN117615075A
Training and deployment of image generation models
US11995803B1
System and method for ai model watermarking
US20220300842A1
Network model training method and apparatus, and computer-readable storage medium
WO2023071743A1