Three-dimensional steganography method based on multi-modal large model and capable of realizing micro three-dimensional rendering
Through multimodal large model and differentiable three-dimensional rendering technology, an information decoder is built and information is embedded in the rendered perspective pictures of the three-dimensional model, solving the problems of small information embedding capacity and low extraction success rate in the existing technology, and achieving efficient and stable information hiding and extraction.
Patent Information
- Application Number
- CN202510347490.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-22
AI Technical Summary
The existing three-dimensional steganography technology has problems such as small information embedding capacity, low extraction success rate, poor adaptability of information decoder and long training time, especially in rendered viewing images, it is difficult to effectively hide and extract information.
Using multimodal large model and differentiable three-dimensional rendering technology, an information decoder is built. Through the feature representation and modal alignment capabilities of the multimodal large model, information is embedded in the rendered viewing picture of the three-dimensional model, and the gradient is returned to the three-dimensional model through the microroutable rendering pipeline to achieve efficient hiding and extraction of information.
It realizes efficient hiding and extraction of information, shortening the training time to within 5 minutes, with large embedding capacity and high extraction accuracy, adapting to various image degradation modes, and maintaining the stable structure of the three-dimensional model.
Smart Images

Figure CN120358308A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of three-dimensional steganography technology, and particularly relates to a three-dimensional steganography method based on a multimodal large model and differentiable three-dimensional rendering. Background Art
[0002] Three-dimensional steganography technology is a technology for hiding information based on three-dimensional model data and is an important technology in the field of information security. Its core purpose is to utilize the complexity and diversity of three-dimensional models without attracting the attention of third parties. Researchers have explored various three-dimensional steganography methods, including techniques based on geometric transformation, texture mapping, and color coding. These methods can embed secret information into the vertex coordinates, normals, texture coordinates, or color information of three-dimensional models, thereby achieving information hiding. Three-dimensional steganography technology can be used in various scenarios, such as digital copyright protection, military communication, digital watermarking, etc. In digital art and game development, artists and designers can use three-dimensional steganography technology to embed copyright information or creative intentions in their works. In medical imaging, steganography technology can be used to protect patients' privacy information.
[0003] Traditional three-dimensional steganography technology mainly includes the following: 1) directly hiding information in local areas of three-dimensional models. Hidden information is embedded by purposefully offsetting point clouds or reconstructing meshes in local areas of three-dimensional objects. However, this method often leads to problems such as unnatural point cloud structures and genus, and is easily detected; 2) adaptively hiding information in the overall shape of three-dimensional models. Although this approach is difficult to detect, it has certain requirements for the three-dimensional point cloud and mesh structures themselves, such as the number of point clouds being relatively consistent, etc., so its application will be restricted to a certain extent.
[0004] However, traditional three-dimensional steganography technology cannot hide information and protect property rights well because they have not considered embedding information into the rendered perspective pictures. In recent years, a series of three-dimensional steganography techniques based on differentiable three-dimensional reconstruction have emerged. They aim to embed hidden information into the feature expression or encoding of three-dimensional objects through a differentiable rendering pipeline, so that the corresponding hidden information can be extracted from the rendered perspective pictures through an information decoder. However, these methods have significant drawbacks. One is that an information decoder can usually only act on a tampered three-dimensional object model; the second is that optimizing the information decoder while fine-tuning the parameters of the three-dimensional model requires long-term training; the third is that their capacity and extraction success rate are relatively low. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a 3D steganography method based on a multimodal large model and differentiable 3D rendering. With the powerful feature representation ability and modal alignment ability of the multimodal large model, a general information extractor can be constructed to achieve efficient training of the tampering model.
[0006] The present invention is implemented through the following technical solutions: A 3D steganography method based on a multimodal large model and differentiable 3D rendering, comprising the steps:
[0007] (1) Based on the specific feature expression and description form of the multimodal large model, construct a corresponding information decoder, the purpose of which is to decode the feature description output by the multimodal large model into the corresponding information;
[0008] (2) Input the information to be hidden into the multimodal large model, then input the feature description output by the multimodal large model into the information decoder constructed in (1) to obtain the decoded information, and then use the information to be hidden to supervise and optimize the information decoder;
[0009] (3) Use existing 3D reconstruction techniques to generate the required 3D model, and use differentiable rendering techniques to render arbitrary perspective images;
[0010] (4) Input the perspective images obtained in (3) into the multimodal large model, and then use the cross-modal alignment ability of the multimodal large model to convert the information in the image modality into a feature description corresponding to the information to be hidden, and input it into the information decoder trained in (2);
[0011] (5) Use the information to be hidden to supervise the information decoded by the information decoder, and then the gradient can be backpropagated into the 3D model through the differentiable rendering pipeline, so as to hide the information to be hidden into the encoding or features of the 3D model;
[0012] (6) After the training in (5) is completed, the user can publish the 3D model with hidden information and provide the information decoder to the cooperation party or the property inspection party for extracting the hidden information in the 3D model, so as to achieve secret communication or property verification and protection.
[0013] Preferably, in the step (1), the multimodal large model includes but is not limited to models such as ChatGPT, LLAMA, and CLIP, etc., but different multimodal large models will have differences in the specific implementation details.
[0014] Preferably, in the step (1), the construction of the information decoder is not unique, and there will be different structural designs according to the specific modality of the information to be hidden and the selected multi-modal large model, but all have the following unified composition: 1) It can accept the specific feature expressions and description forms of the multi-modal large model as input; 2) The output is modal data corresponding to the information to be hidden; 3) A neural network structure that can extract multi-level features of the input image through the encoder, restore the resolution of the image through the decoder, and use skip connections to fuse the features of the encoder and decoder to achieve accurate image segmentation.
[0015] When constructing the information decoder, the implementer can regard the overall structure as an autoencoder using a multi-modal large model, where the multi-modal large model acts as the encoder of features, and the information decoder is the decoder of features, responsible for restoring information from the features.
[0016] Preferably, the loss function and training strategy for supervised optimization in the step (2) are not unique, and there will be different designs according to the specific modality of the information to be hidden; specifically, for binary strings, cross-entropy loss can be used, while for picture or audio information, loss functions for reconstruction such as Euclidean distance or cosine similarity can be used for optimization.
[0017] Preferably, the 3D reconstruction technology in the step (3) is not unique, as long as the rendering pipeline is differentiable, and methods such as neural radiance fields or 3D Gaussian sputtering can be used.
[0018] To achieve high-speed rendering and training, it is recommended to use a differentiable 3D reconstruction technology framework based on 3D Gaussian sputtering.
[0019] Preferably, in the step (4), if the output feature descriptions of the multi-modal large model are consistent among different modalities, the feature descriptions can be directly input into the information decoder; if not, an additional feature projection layer needs to be added in front of the information decoder for dimension alignment between different modality features and expressions.
[0020] Preferably, in the step (5), when backpropagating the gradient to the 3D model through the differentiable rendering pipeline, it is necessary to keep the 3D structure of the 3D model itself from changing too much while embedding the hidden information into the asset. Therefore, generally two loss functions need to be jointly optimized. One is the loss function used in the step (2), and the other is the Euclidean distance loss to ensure the rendering quality, ensuring that there is not too much deviation in the rendered picture after fine-tuning.
[0021] Moreover, the implementer needs to balance the influence between the two losses to ensure that the optimization route is towards the goal of both excellences.
[0022] Preferably, in step (6), the user can publish the 3D model with hidden information to the required place. It is difficult for ordinary users to extract the hidden information from this model, while the user's partner can use the information decoder shared by the user to extract the hidden information to achieve secret communication or property protection.
[0023] The present invention constructs a novel 3D steganography framework. By using the powerful feature representation ability and modality alignment ability of the multimodal large model, it can efficiently embed information of any modality into the rendering images of any perspective of the 3D model. Compared with the prior art, it has the following advantages and beneficial effects:
[0024] 1. Compared with the prior art which usually takes 12 hours and 2 hours respectively to train the information decoder and fine-tune the 3D model, the training efficiency of the present invention is higher. It only takes less than 5 minutes to train a decoder, and only 10 minutes to fine-tune the 3D model, enabling this technology to be effectively applied in real scenarios;
[0025] 2. With the support of the powerful feature representation ability of the multimodal large model, the present invention has a larger embedding capacity, higher extraction accuracy, and higher robustness, and can adapt to various image degradation modes to extract accurate and complete hidden information. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is the training flow chart of the information decoder in this embodiment;
[0027] Figure 2 is the fine-tuning flow chart of the 3D model in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The present invention will be further described in detail below with reference to the embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0029] Embodiment
[0030] The 3D steganography method based on the multimodal large model and differentiable 3D rendering in this embodiment mainly includes three steps: constructing an information decoder based on the multimodal large model, training the 3D model, and embedding the information to be hidden into the 3D model. The specific training framework for constructing the information decoder based on the multimodal large model can be referred to Figure 1 , and Figure 2 shows the fine-tuning process framework for embedding the information to be hidden into the 3D model. The above steps will be described in detail with examples below in conjunction with the accompanying drawings.
[0031] I. Constructing the information decoder:
[0032] In this step, it is necessary to construct and train an information decoder based on a multimodal large model.
[0033] 1.1 Before constructing the information decoder, it is necessary to determine the input and output forms and dimensions of the information decoder to design appropriate input and output layers. Among them, the input layer of the information decoder corresponds to the output result of the multimodal large model, and the output layer of the information decoder corresponds to the specific hidden information for reconstruction.
[0034] 1.2 To achieve a better reconstruction effect, the structure of the information decoder is preferably selected to be similar to the U-Net configuration, that is, a neural network structure that can extract multi-level features of the input image through an encoder, restore the resolution of the image through a decoder, and utilize skip connections to fuse the features of the encoder and decoder to achieve precise image segmentation.
[0035] 1.3 After constructing the information decoder, the implementer should Figure 1 train the information decoder according to the process framework shown. First, tokenize the input hidden information to be hidden, then input it into the multimodal large model to obtain the corresponding feature description; then input the feature description into the information decoder to reconstruct data consistent with the input information; finally, select an appropriate reconstruction loss function for supervision and optimization of the information decoder according to the actual situation. Note that the multimodal large model is frozen during the training process, and this training is only for the previously constructed information decoder itself.
[0036] 1.4 During the training process, the loss function and training strategy for supervision and optimization are not unique, and different designs will be available according to the specific modality of the hidden information to be hidden. Specifically, for binary strings, cross-entropy loss can be used, while for image or audio information, loss functions for reconstruction such as Euclidean distance or cosine similarity can be used for optimization.
[0037] II. Training a 3D model
[0038] This step is to allow the user to train and generate a 3D model carrier that needs to be protected by property rights or used for secret communication. The user can use any popular 3D reconstruction technology for generation, including but not limited to neural radiance fields and 3D Gaussian sputtering and other technologies. Note that the 3D reconstruction technology selected in the present invention needs to have a differentiable rendering pipeline so that subsequent fine-tuning can be performed, and the gradient can be backpropagated into the 3D model through the differentiable rendering pipeline, thereby embedding the hidden information to be hidden into the encoding or features of the 3D model.
[0039] III. Embedding hidden information
[0040] See Figure 2, this step intends to embed the information to be hidden into the 3D model trained in step 2 in a fine-tuning manner through a multimodal large model and the information decoder constructed in step 1.
[0041] 3.1 Arbitrarily select a 3D model and render any perspective images to be protected according to its rendering pipeline;
[0042] 3.2 Tokenize the rendered perspective images and then input them into the multimodal large model to obtain corresponding feature descriptions;
[0043] 3.3 Input the feature descriptions in 3.2 into the information decoder constructed in step 1 to reconstruct decoded information in the same form as the hidden information. It should be noted here that if the output feature descriptions of the multimodal large model are consistent among different modalities, the feature descriptions can be directly input into the information decoder; if not, an additional feature projection layer needs to be added in front of the information decoder for dimension alignment between different modality features and representations.
[0044] 3.4 Supervise and fine-tune the 3D model through a reconstruction loss function. When backpropagating the gradient to the 3D model through a differentiable rendering pipeline, it is necessary to keep the 3D structure of the 3D asset itself from changing too much while embedding the hidden information into the asset. Therefore, generally two loss functions need to be jointly optimized. One is the loss function used in step 1, and the other is the Euclidean distance loss to ensure the rendering quality, ensuring that there is not too much deviation in the rendered images after fine-tuning.
[0045] Note that during the entire fine-tuning process, the parameters of the information decoder and the multimodal large model are frozen, and the implementer only needs to fine-tune the parameters of the 3D model. Also, since existing 3D models can be roughly divided into implicit 3D models and explicit 3D models, their specific model files, parameters, and rendering methods are not consistent. Therefore, the specific fine-tuning process needs to be adjusted according to a specific rendering pipeline. Considering training efficiency and image rendering quality, it is recommended to use three-dimensional Gaussian sputtering as the differentiable 3D reconstruction technology of the present invention.
[0046] Implementers can implement the technologies described in the present invention through various means. For example, these technologies can be implemented in different multi-modal large models, including but not limited to ChatGPT, LLAMA, CLIP, Doubao, and ERNIE Bot, but there will be differences in the specific implementation details of different multi-modal large models. The technology can be implemented on any computing device. For the hardware implementation, the processing module can be implemented in one or more application-specific integrated circuits, digital signal processors, programmable logic devices, field-programmable gate arrays, processors, controllers, microcontrollers, electronic devices, other electronic units designed to execute the functions described in the present invention, or a combination thereof.
[0047] For the firmware and / or software implementation, the technologies can be implemented using modules (e.g., procedures, steps, processes, etc.) that execute the functions described herein. The firmware and / or software code can be stored in a memory and executed by a processor. The memory can be implemented inside or outside the processor.
[0048] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disks, or optical discs and other various media that can store program codes.
[0049] The above embodiments are preferred embodiments of the present invention, but the implementation manners of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention.
Claims
1. A 3D steganography method based on a multimodal large model and differentiable 3D rendering, characterized in that, Including the steps: (1) Based on the specific feature expressions and description forms of the multimodal large model, construct a corresponding information decoder, the purpose of which is to decode the feature descriptions output by the multimodal large model into corresponding information; (2) Input the information to be hidden into the multimodal large model, then input the feature descriptions output by the multimodal large model into the information decoder constructed in (1) to obtain the decoded information, and then use the information to be hidden to supervise and optimize the information decoder; (3) Use existing 3D reconstruction techniques to generate the required 3D model, and use differentiable rendering techniques to render arbitrary perspective images; (4) Input the perspective images obtained in (3) into the multimodal large model, and then use the cross-modal alignment ability of the multimodal large model to convert the information in the image modality into feature descriptions corresponding to the information to be hidden, and input them into the information decoder trained in (2); (5) Use the information to be hidden to supervise the information decoded by the information decoder, and then the gradient can be backpropagated into the 3D model through the differentiable rendering pipeline, so as to hide the information to be hidden into the encoding or features of the 3D model; (6) After the training in (5) is completed, the user can publish the 3D model with hidden information, and provide the information decoder to the cooperation party or the property inspection party for extracting the hidden information in the 3D model, so as to realize secret communication or property verification and protection.
2. The three-dimensional steganography method based on a multi-modal large model and differentiable three-dimensional rendering according to claim 1, wherein The construction of the information decoder in step (1) is not unique, and there will be different structural designs according to the specific modality of the information to be hidden; but all have the following unified composition: 1) It can accept the specific feature expressions and description forms of the multimodal large model as input; 2) The output is modal data corresponding to the information to be hidden; 3) A neural network structure that can extract multi-level features of the input image through an encoder, restore the resolution of the image through a decoder, and use skip connections to fuse the features of the encoder and decoder to achieve accurate image segmentation.
3. The three-dimensional steganography method based on a multimodal large model and differentiable three-dimensional rendering according to claim 1, characterized in that The loss function and training strategy for supervised optimization in step (2) are not unique, and there will be different designs according to the specific modality of the information to be hidden; specifically, cross-entropy loss can be used for binary strings, while Euclidean distance or cosine similarity and other loss functions for reconstruction can be used to optimize image or audio information.
4. The 3D steganography method based on a multimodal large model and differentiable 3D rendering according to claim 1, characterized in that, The 3D reconstruction technique in step (3) is not unique, as long as the rendering pipeline is differentiable, it can be methods such as neural radiance fields or 3D Gaussian sputtering.
5. The 3D steganography method based on a multimodal large model and differentiable 3D rendering according to claim 1, characterized in that In step (4), if the output feature descriptions of the multimodal large model are consistent between different modalities, the feature descriptions can be directly input into the information decoder; if not, an additional feature projection layer needs to be added in front of the information decoder for dimensional alignment between different modality features and representations.
6. The three-dimensional steganography method based on a multimodal large model and differentiable three-dimensional rendering according to claim 3, wherein In step (5), when backpropagating the gradient into the 3D model through the differentiable rendering pipeline, it is necessary to keep the 3D structure of the 3D model itself from changing too much while embedding the hidden information into the asset. Therefore, generally two loss functions need to be jointly optimized. One is the loss function used in step (2), and the other is the Euclidean distance loss to ensure the rendering quality, ensuring that there is not too much deviation in the rendered image after fine-tuning.
7. The 3D steganography method based on a multimodal large model and differentiable 3D rendering according to claim 1, wherein In step (6), the user can publish the 3D model with hidden information to the required place. It is difficult for ordinary users to extract the hidden information from this model, while the user's partner can use the information decoder shared by the user to extract the hidden information to achieve secure communication or property protection.
8. The 3D steganography method based on a multi-modal large model and differentiable 3D rendering according to claim 1, characterized in that, The multi-modal large model mentioned above includes but is not limited to models such as ChatGPT, LLAMA, and CLIP. However, there will be differences in the specific implementation details of different multi-modal large models.
Citation Information
Cited By
Digital watermarking method, device and equipment based on three-dimensional model, medium and product
CN121481820A
Three-dimensional model-based digital watermarking method, apparatus, device, medium and product
CN121481820B