This invention relates to the field of model watermarking technology and discloses a multimodal watermarking
processing method based on decoupled injection of 3D reconstruction models. The method involves acquiring a
monocular image input, extracting image coding features through a 3D reconstruction
backbone network, activating a progressive message
encoder, and using a unified semantic module to convert multimodal copyright messages into
watermark feature vectors. The
watermark feature vectors are then input into a feature injection
branch to generate a watermarked rendered image. A progressive decoupling strategy is applied to the rendered image, including a collaborative
training phase and an independent optimization phase. The progressive message
encoder is removed while the 3D reconstruction
backbone network is retained, achieving copyright
watermark embedding
verification with zero changes to the model structure. The rendering module outputs multi-view rendered images carrying the watermark, achieving zero structural changes during deployment or
inference. After training, auxiliary components are removed, and the deployed model maintains the same architecture as the original, without affecting reconstruction quality, balancing embedding strength and model performance.