A method and apparatus for tracing the origin of content generated by a visual language model

CN121168407BActive Publication Date: 2026-09-22CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511156973.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2026-09-22
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

一旦该类内容通过社交平台等渠道扩散,不仅可能损害用户隐私,还可能对社会造成广泛的不良影响

Benefits of technology

[0041]有益效果:本发明提供的结合对抗触发器与隐形文本标记的视觉语言模型生成内容的溯源方法及装置,通过在生成文本的语义空间中联合嵌入不可见标记,实现了对视觉语言模型生成内容的精确溯源,能够有效应对视觉语言模型滥用带来的安全风险;另外,本发明在保证生成内容语义一致性和用户体验不受影响的前提下,所提出的隐形标记构建能够在多种使用场景下可靠地标记与识别内容来源,提升了模型的责任归属能力与可监管性,为构建安全、可追责的多模态生成系统提供了技术支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168407B_ABST
    Figure CN121168407B_ABST
Patent Text Reader

Abstract

The application discloses a kind of visual language model generation content's traceability method and device, combined with the text generation capability of visual language model and adversarial trigger learning mechanism, by embedding invisible text mark in image, reliable traceability to generated content is realized;The method introduces target mark loss and non-target mark loss two antagonistic optimization targets, respectively ensure the consistency of mark content and original content in semantics, and enhance its distinguishability in text space, so as to realize accurate traceability identification;In addition, the application also designs a lightweight, no training traceability module, can effectively identify embedded mark, realize the rapid tracking of content source;The traceability technique proposed in the case can effectively mark VLM output under different use scenarios without affecting the quality of generation, and the experimental results show that excellent traceability ability and stable semantic retention performance are shown on multiple data sets and models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and apparatus for tracing the source of content generated by a visual language model that combines adversarial triggers and hidden text tags, belonging to the field of tracing the source of content generated by multimodal large language models. Background Technology

[0002] In recent years, Visual-Language Models (VLMs) have demonstrated powerful capabilities in multimodal understanding and generation tasks, and have been widely applied in areas such as image captioning, visual question answering (VQA), and human-computer interaction. With the emergence of advanced models such as GPT-4V, BLIP, and LLaVA, VLMs have made significant progress in vision-language alignment, multimodal reasoning, and context-aware generation. These models typically rely on large-scale image-text pre-training and instruction alignment techniques, significantly improving the system's performance in real-world tasks and driving the development of artificial intelligence systems towards generalization and interactivity. Thanks to the development of the open-source ecosystem, the training and deployment barriers of VLMs have been decreasing, prompting their rapid adoption in academic research and practical applications.

[0003] However, as the accessibility of Virtual Models (VLMs) continues to improve, their potential for abuse is also becoming increasingly prominent. The powerful visual analysis and language generation capabilities of these models can be exploited by malicious users to extract private information from images or generate misleading, discriminatory, or even illegal content. For example, attackers can use the model's inference process to extract sensitive content such as faces and identification documents, or automatically generate hate speech and fake news. Once such content spreads through social media platforms and other channels, it can not only harm user privacy but also have a wide-ranging negative impact on society. Although some current work has attempted to introduce tagging mechanisms to protect the copyright of model outputs, these efforts are mostly limited to model fine-tuning strategies and lack systematic research on the "content generation tracing" problem in complex malicious usage scenarios. Therefore, how to achieve accurate tracing of VLM-generated content without interfering with the quality of normal generation has become one of the key issues in ensuring its safe and controllable development. Summary of the Invention

[0004] Purpose of the invention: In order to overcome the shortcomings of the existing technology, the present invention provides a method and apparatus for tracing the source of content generated by a visual language model that combines adversarial triggers and hidden text tags, which can ensure that the content generated by the multimodal large language model can be successfully traced.

[0005] Technical solution: To achieve the above objectives, the technical solution adopted by this invention is as follows:

[0006] A method for tracing the origin of content generated by a visual language model is disclosed. The visual language model generates traceable text based on traceable triggers and latent markers. This tracing method embeds the traceable trigger T into the original image x through adversarial optimization to generate text containing the target marker w. TSL Source tracing images of relevant information Make the visual language model f θ While maintaining semantic consistency of the content, target tags w are embedded in the semantic space of the generated text. TSL This information enables reliable tracking of generated text across multiple scenarios; the tracing method specifically includes the following steps:

[0007] S1. Construct the set of invisible markers and the set of prompts: Construct the set of invisible markers W = {w} consisting of M invisible markers and the set of prompts P = {p} consisting of N prompts;

[0008] S2. Construct a set of non-target markers: Select one latent marker from the latent marker set W that corresponds to the original image x (usually an unmodified natural image) as the target marker w. TSL The remaining M-1 hidden markers are considered as non-target markers, forming the non-target marker set W. NTSL ={w NTSL}, W = w TSL ∪W NTsL The latent markers are divided into target markers and non-target markers to provide clear positive and negative supervision signals. Target markers are used to guide the visual language model f. θ By embedding specified source information (such as user information) and using non-target tags to suppress interference from irrelevant information, this contrastive training mechanism improves the discriminability and embedding accuracy of the hidden tags, enabling the visual language model f... θ It can still achieve stable and accurate content tracing in a multi-user environment;

[0009] S3. Construct and optimize the traceable trigger: Construct a traceable trigger T, and embed the traceable trigger T into the original image x to obtain the source image. The traceable trigger T allows for the tracing of the source image. Carrying and target marker w TSL Corresponding, specific, and traceable source information is used, and this source information is also used as an implicit identity representation of the original image x, to track the original image x and the user that generated the harmful content; using the non-target tag set W NTsL Adversarial-supervised training is performed on the traceable trigger T to enhance the pre-trained visual language model f. θ For non-target marker w NTSL The ability to distinguish;

[0010] S4. Test traceable triggers T: For a given set of hidden tags W = {w (1) w (2) ,…,w (m) ,…,w (M)}, and the original image x embedded with the latent marker w before and after passing through the visual language model f θ The generated text f θ (x, p), f θ ((T+x), p), the predicted hidden label w is: Based on the traceability information contained in the predicted hidden marker w, information such as the user can be traced back.

[0011] In this invention, the source image This is achieved by embedding a traceable trigger T in the original image x, through which the target marker w is triggered. TSL The source information is added to the original image x to obtain the source image. This process does not affect the visual effect and semantic content of the original image x, thus ensuring the visual language model f. θ Text f generated from the original image x θ (x, p) and source-tracing images The generated text It maintains semantic consistency while enabling effective traceability of user information, ensuring an undisturbed user experience. With current technology, without using a traceable trigger T, it is impossible to directly mark the target w. TSL Or target marker w TSL The source information contained therein is directly embedded into the original image x without being detected.

[0012] Specifically, in step S3, the non-target tag set W is used. NTSL Iterative optimization of the traceable trigger T includes the following steps:

[0013] (1) Set the learnable parameter λ, the iteration step size α, and the maximum number of iterations K, k = 1;

[0014] (2) From the prompt set P and the non-target tag set W respectively NTSL The prompt p and non-target marker w used in this iteration are selected by modulo index. NTSL ;

[0015] (3) Input the original image x and the prompt p into the visual language model f θ In the middle, the generated text f is obtained. θ (x, p);

[0016] (4) Add a traceable trigger T to the original image x to obtain the source image.

[0017] (5) Source tracing images The prompt p is input into the visual language model f θ In the process, the generated text is obtained.

[0018] (6) In generating text f θ Embed non-target marker w in (x, p) NTSL Calculate the non-target label loss L NTSL :

[0019]

[0020] Where: L(·) represents the visual language model f θ The loss function;

[0021] The non-target labeling loss L NTSL Guide the visual language model f θ Embed non-target tags w in the semantic space of the generated text. NTSL Enhance target marker w TSL Discriminability; the loss is achieved by maximizing the non-target label w. NTSL Text generation based on the original image x and text generation based on the source image Differences between generated texts, suppressing non-target tags w NTSL For visual language model f θ Perturbations in generated text ensure the visual language model f θ The generated text can be attributed to a unique target tag w TSL ;

[0022] (7) In generating text f θ Embed target marker w in (x, p) TSL Calculate the target label loss L TSL :

[0023]

[0024] The target label loss L TSL Guide the visual language model f θ Embed the target tag w in the semantic space of the generated text. TSL While maintaining semantic consistency of content; this loss is achieved by minimizing the loss with target label w TSL Text generation based on the original image x and text generation based on the source image The differences between the generated text and the target tag w TSL The traceable trigger T is naturally and semantically embedded into the visual language model f. θIn the generated text;

[0025] (8) Calculate the joint loss L ATL =L NTSL +λL TSL By reducing the combined loss L ATL It can suppress non-target markers w NTSL Activation, while increasing the target marker w TSL and non-target marker w NTSL Distinguishability;

[0026] (9) Based on the joint loss L ATL Calculate about the source image The gradient g is used to inversely optimize the traceable trigger T using the signed gradient, so that the source image updated by adding the updated traceable trigger T to the original image x is... Achieve source tracing in images The effect of embedding weak perturbations; by continuously optimizing the traceable trigger T, it can continuously trace the source image. Embedding perturbations to gradually guide the visual language model f θ Output with target marker w TSL But semantics and generated text f θ Generate text that maintains consistency between (x, p)

[0027] (10) Determine if k is equal to K: If yes, then the training of the traceable trigger T is complete, and the visual language model f θ Output with target marker w TSL But semantics and generated text f θ Generate text that maintains consistency between (x, p) Otherwise, k = k + 1, return to step (2), and continue iterating.

[0028] Specifically, multimodal large models (VLMs) are an important research direction in the current field of artificial intelligence. In this case, the visual language model adopts a typical VLM structure, including two core modules: a powerful visual encoder and an advanced large language model. The visual encoder usually uses a pre-trained model, such as the visual part of the CLIP model (ViT series), which is responsible for extracting visual features of the input image. The large language model adopts architectures such as GPT, LLaMA, or Vicuna, which have powerful text understanding and generation capabilities and are responsible for extracting features of the input image based on prompts. In order to achieve cross-modal fusion, a lightweight visual-language adaptation layer or a cross-modal attention mechanism is used to map visual features to the embedding space of the large language model, enabling the large language model to understand the input image and generate text.

[0029] In terms of training strategies, VLM utilizes large-scale image-text alignment data and multi-task instruction tuning, combining tasks such as visual question answering, image-text matching, multi-turn dialogue, and complex reasoning, and employs a multi-objective loss function to optimize VLM performance. To reduce training costs, many VLMs only fine-tune the visual-language adaptation layer and a small number of large language model parameters, keeping the main parameters of the visual encoder and large language model fixed. This ensures both the powerful multimodal understanding capabilities of VLMs and makes the training process more efficient. Representative VLMs include MiniGPT-4 and LLaVA 2, which achieve deep interaction between vision and language through lightweight visual-language adaptation layers and cross-modal attention mechanisms, respectively. They are widely used in multimodal tasks such as image description, visual question answering, and intelligent dialogue, promoting the development of AI in the field of vision and language fusion.

[0030] Specifically, the invisible marker is an invisible marker composed of invisible Unicode characters (Word Joiner and Zero WidthNo-Break Space). After being indirectly embedded into the original image x through a traceable trigger T, the invisible marker does not affect the visual effect or semantic content of the original image x, nor does it affect the semantic content of the text generated by the visual language model. However, it is embedded in the visual language model f. θ In the generated text, the identification and tracking of content generated by different users are achieved. Each latent tag is encoded from a randomly generated string. The string is first converted into a binary sequence and then mapped to the corresponding invisible Unicode character, forming a unique latent tag. As can be seen from the latent tag generation process, the latent tag itself does not carry semantic information. This mechanism ensures that it does not affect the semantic content of the embedding result, ensuring the naturalness and usability of the embedded content. This allows for the visualization language model f to be optimized without interfering with normal output. θ The generated text can be stably traced, which has good practicality and promotional value.

[0031] Specifically, the original image x is a high-quality image from the MS-COCO validation set. Diverse prompts p are automatically generated using GPT-4o and DeepSeek-V3. Simultaneously, visual-language association data from real-world scenarios are collected to further refine the original image set and prompt set, resulting in the final original image set {x} and prompt set P. These sets encompass three representative information types: Recognizable Information (ReI), Inferable Information (InI), and Associative Information (LiI). This allows for the simulation of different levels of content understanding and privacy leakage risks, ensuring comprehensive evaluation capabilities against various potential abuse scenarios. Furthermore, the visual-language association data from real-world scenarios further validates the effectiveness and robustness of the proposed method in practical applications.

[0032] A source tracing device for content generated by a visual language model includes an original image and prompt acquisition module, a latent marker generation module, a traceable trigger, a visual language model, an adversarial loss optimization module, and a latent marker prediction module.

[0033] The original image and prompt acquisition module is used to construct an original image set and a prompt set;

[0034] The stealth marker generation module first randomly generates a string, then converts the string into a binary sequence, and finally maps the binary sequence to the corresponding invisible Unicode character to form a unique stealth marker; all stealth markers constitute a stealth marker set; based on the original image, one stealth marker is selected from the stealth marker set as the target marker, and the rest constitute a non-target marker set;

[0035] The traceable trigger embeds the source information carried by the target marker into the original image to obtain the source image;

[0036] The visual language model generates text based on prompts obtained from the original image and the source image;

[0037] The adversarial loss optimization module uses a set of non-target labels to perform adversarial supervised training on the traceable triggers, thereby enhancing the ability of the pre-trained visual language model to distinguish non-target labels.

[0038] The latent marker prediction module predicts latent markers from a given set of latent markers by using the generated text of a visual language model before and after embedding the latent markers in the original image through a trained traceable trigger.

[0039] Specifically, the adversarial loss optimization module uses both target label loss and non-target label loss to optimize the traceable trigger. The target label loss minimizes the difference between the generated text based on the original image with target label and the generated text based on the source image, allowing the target label to be embedded into the semantic space of the text generated by the visual language model through the traceable trigger. The non-target label loss maximizes the difference between the generated text based on the original image with non-target label and the generated text based on the source image, suppressing the interference of non-target labels on the text generated by the visual language model, so that the text generated by the visual language model can be attributed to a unique target label.

[0040] Specifically, the occult marker prediction module identifies embedded occult markers by minimizing the difference between generated text based on the original image and generated text based on the source image.

[0041] Beneficial effects: The method and apparatus for tracing the source of content generated by a visual language model, which combines adversarial triggers and invisible text tags, provided by this invention, achieves accurate source tracing of content generated by the visual language model by jointly embedding invisible tags in the semantic space of the generated text. This effectively addresses the security risks caused by the abuse of visual language models. In addition, while ensuring the semantic consistency of the generated content and the user experience are not affected, the proposed invisible tag construction can reliably tag and identify the source of content in various usage scenarios, improving the model's ability to assign responsibility and its regulatory capacity. This provides technical support for building a safe and accountable multimodal generation system. Attached Figure Description

[0042] Figure 1 This is a schematic diagram illustrating the implementation process of the method of the present invention;

[0043] Figure 2 This is a schematic diagram of the signal flow of embedding a traceable trigger in the original image according to the present invention. Detailed Implementation

[0044] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0045] Currently, most methods for protecting images and text are limited to copyright protection and do not adequately cover the diverse risks of misuse faced by visual language models. To address this need, this paper proposes a robust and efficient source tracing technology for content generated by visual language models. Through adversarial learning using implicit text tags and traceable triggers, identifiable implicit tags are embedded in the text generated by the visual language model. This effectively tracks the source of content, enabling accountability for malicious behavior and technical supervision, promoting the safe and compliant use of visual language models, and possessing broad practical application value and security significance.

[0046] This paper first proposes a source tracing device for content generated by a visual language model. By embedding weak perturbations into the original image to construct a source image, and guiding the visual language model to generate semantically consistent output text with implicit markers, it achieves accurate source tracing of the generated content. The device includes an original image and prompt acquisition module, an implicit marker generation module, a traceable trigger, a visual language model, an adversarial loss optimization module, and an implicit marker prediction module. The original image and prompt acquisition module is used to construct a set of original images and a set of prompts. The implicit marker generation module first randomly generates a string, then converts the string into a binary sequence, and finally maps the binary sequence to the corresponding invisible Unicode character, forming a unique implicit marker. All implicit markers... The system consists of a set of latent markers; a target marker is selected from the latent marker set based on the original image, and the remaining markers form a set of non-target markers; the traceable trigger embeds the source information carried by the target marker into the original image to obtain the source image; the visual language model generates text from the original image and the source image based on the prompt; the adversarial loss optimization module uses the set of non-target markers to perform adversarial supervised training on the traceable trigger, enhancing the pre-trained visual language model's ability to distinguish non-target markers; and the latent marker prediction module, based on the original image and the generated text of the visual language model before and after embedding the latent marker through the trained traceable trigger, predicts the latent markers using the minimum generation loss, within the given latent marker set.

[0047] like Figure 1 , Figure 2 As shown, this is a method for tracing the source of visual language model-generated content based on the above-mentioned device. Through adversarial optimization, a traceable trigger T is embedded into the original image x to generate a target marker w. TSL Source tracing images of relevant information Make the visual language model f θ While maintaining semantic consistency of content, embed target tags w in the generated text. TSL This invention provides relevant information to enable reliable tracking of generated text in multiple scenarios. The following detailed description, in conjunction with specific implementation steps, further illustrates the invention.

[0048] PART 1: Data Preparation Stage

[0049] S1. Construct the original image set {x} and the prompt set.

[0050] Given the security risks and diverse abuse scenarios of VLM-generated content, constructing an evaluation dataset covering multiple information types is particularly important. In this case, the original image x is a high-quality image from the MS-COCO validation set. Diverse prompts p are automatically generated using GPT-4o and DeepSeek-V3. Simultaneously, visual-language association data from real-world scenarios are collected to further refine the original image set and prompt set, resulting in the final original image set {x} and prompt set P. The original image set {x} and prompt set P encompass identifiable information (ReI, such as directly sensitive content like faces and identification documents) and inferable information (InI, requiring further processing). The invention employs three representative information types: privacy obtained through reasoning (such as inferring occupation and health status through scenarios) and associative information (LiI, potential risks revealed through multimodal content association, such as text and image combinations implying geographical location and social relationships). This allows for the simulation of different levels of content understanding and privacy leakage risks, ensuring comprehensive assessment capabilities in combating various potential abuse scenarios. Furthermore, real-world visual-language association data can further validate the effectiveness and robustness of the proposed method in practical applications. These diverse and meticulously segmented datasets provide a solid testing foundation for the invention's method, ensuring reliable and accurate content tracing in complex real-world environments.

[0051] S2. Construct a set of hidden markers.

[0052] Latent tags are invisible tags composed of unseen Unicode characters. Each latent tag is encoded from a randomly generated string, which is first converted into a binary sequence and then mapped to the corresponding unseen Unicode character to form a unique latent tag. As can be seen from the generation process of latent tags, the tags themselves do not carry semantic information. This mechanism ensures that they do not affect the semantic content of the embedding result, guaranteeing the naturalness and usability of the embedded content. This allows for the embedding of the visual language model f without interfering with normal output. θ The generated text can be stably traced, which has good practicality and promotional value.

[0053] S3. Construct a set of non-target labels.

[0054] For a given original image x, select one hidden marker corresponding to the original image x from the set of hidden markers W as the target marker w. TSL The remaining M-1 hidden markers are considered as non-target markers, forming the non-target marker set W. NTSL ={w NTSL}, W = w TSL ∪W NTSL .

[0055] The latent markers are divided into target markers and non-target markers to provide clear positive and negative supervision signals. Target markers are used to guide the visual language model f. θ By embedding specified source information (such as user information) and using non-target tags to suppress interference from irrelevant information, this contrastive training mechanism improves the discriminability and embedding accuracy of the hidden tags, enabling the visual language model f... θ It can still achieve stable and accurate content tracing in a multi-user environment.

[0056] PART 2: Traceable Trigger Training Phase

[0057] S4. Initialize the traceable trigger T

[0058] Construct a traceable trigger T to embed a target marker w in the original image x. TSL Corresponding, specific, and traceable source information is used to obtain a source image. The source information also serves as an implicit identity representation of the original image x, and can be used to track the original image x and the user that generated the harmful content.

[0059] In subsequent optimization processes, by continuously optimizing the traceable trigger T, it is possible to gradually embed weak perturbations into the original image x, thereby achieving the effect of tracing the source image. Optimization, that is, optimization of the visual language model f θ Input optimization.

[0060] S5. Select the visual language model f θ

[0061] The visual language model adopts a typical VLM structure, comprising two core modules: a powerful visual encoder and an advanced large language model. The visual encoder, typically a pre-trained model, is responsible for extracting visual features from the input image. The large language model needs strong text understanding and generation capabilities, responsible for extracting features from the input image based on prompts. To achieve cross-modal fusion, a lightweight visual-language adaptation layer or a cross-modal attention mechanism is used to map visual features to the embedding space of the large language model, enabling the large language model to understand the input image and generate text. In this case, two typical visual language models, MiniGPT-4 and LLaVA2, are used.

[0062] The MiniGPT-4 combines a pre-trained visual encoder with a large language model. It uses a lightweight linear adaptation layer to convert visual features into the input embedding of the large language model, achieving cross-modal alignment between the input image and the generated text. Training MiniGPT-4 only requires fine-tuning the linear adaptation layer and a small number of parameters of the large language model. It utilizes large-scale image-text alignment data and multi-task instruction optimization to achieve powerful multimodal understanding and generation capabilities at low cost.

[0063] The LLaVA2, based on a more powerful pre-trained visual encoder and an optimized large language model, adopts a multi-task joint training strategy, combining visual question answering, instruction tuning and image-text matching tasks, and achieves deep fusion of vision and language through a cross-modal attention mechanism, thereby improving the VLM's multi-turn dialogue capabilities and complex visual reasoning performance.

[0064] S6, Optimize traceable trigger T

[0065] Using the non-target tag set W NTSL Adversarial-supervised training is performed on the traceable trigger T to enhance the pre-trained visual language model f. θ For non-target marker w NTSL The ability to distinguish; specifically, the iterative optimization process of the traceable trigger T includes the following steps:

[0066] (1) Set the learnable parameter λ, the iteration step size α, and the maximum number of iterations K, k = 1;

[0067] (2) From the prompt set P and the non-target tag set W respectively NTSL The prompt p and non-target marker w used in this iteration are selected by modulo index. NTSL ;

[0068] (3) Input the original image x and the prompt p into the visual language model f θ In the middle, the generated text f is obtained. θ (x, p);

[0069] (4) Add a traceable trigger T to the original image x to obtain the source image.

[0070] (5) Source tracing images The prompt p is input into the visual language model f θ In the process, the generated text is obtained.

[0071] (6) In generating text f θ Embed non-target marker w in (x, p) NTSL Calculate the non-target label loss L NTSL :

[0072]

[0073] Where: L(·) represents the visual language model f θ The loss function;

[0074] (7) In generating text f θ Embed target marker w in (x, p) TSL Calculate the target label loss L TSL :

[0075]

[0076] (8) Calculate the joint loss L ATL =L NTSL +λL TSL ; Loss L through non-target labeling NTSL and target label loss L TSL The synergistic effect not only ensures the target marker w TSL It exhibits natural integration in terms of semantic content and high precision in discriminative ability, thereby achieving a high degree of accuracy in visual language models. θ Robust, scalable, and user-specific source tracing capabilities for generated text;

[0077] (9) Based on the joint loss L ATL Calculate about the source image The gradient g is used to inversely optimize the traceable trigger T using the signed gradient, so that the source image updated by adding the updated traceable trigger T to the original image x is... Achieve source tracing in images The effect of embedding weak perturbations; by continuously optimizing the traceable trigger T, it can continuously trace the source image. Embedding perturbations to gradually guide the visual language model f θ Output with target marker w TSL But semantics and generated text f θ Generate text that maintains consistency between (x, p)

[0078] (10) Determine if k is equal to K: If yes, then the training of the traceable trigger T is complete, and the visual language model f θ Output with target marker w TSL But semantics and generated text f θ Generate text that maintains consistency between (x, p) Otherwise, k = k + 1, return to step (2), and continue iterating.

[0079] PART 3, Source Tracing Testing Phase

[0080] S7, Test Traceable Trigger T

[0081] To achieve effective source attribution for the generated text, this study designs a loss-based source attribution mechanism that can identify embedded latent markers without additional training. This mechanism does not rely on traditional classification networks; instead, it calculates the generation loss between each candidate latent marker and the generated text, selecting the best-matching latent marker as the source attribution result. Specifically, given a predefined set of latent markers W = {w (1) w (2) ,…,w (m) ,…,w (M)}, and through the visual language model f θ The process of generating text by predicting the embedded latent tag w can be represented as:

[0082]

[0083] Where: f θ (x, p) and f θ ((T+x), p) represent the original image x before and after embedding the latent tag w, respectively, through the visual language model f. θ The generated text; based on the traceability information contained in the predicted hidden marker w, information such as the user can be traced back.

[0084] In this case, when testing the traceable trigger T, no additional training is required; it is only necessary to minimize the generated text based on the original image x with hidden marker w and the text based on the source image. The differences between the generated text can be accurately identified to determine the embedded hidden tag w, thereby determining user information. This tracing mechanism does not rely on traditional classification networks. Instead, it selects the most matching tag as the tracing result by calculating the generation loss between each candidate tag and the generated content. This tracing mechanism is not only efficient and requires no training, but also has good scalability and practical deployment capabilities.

[0085] PART 4 ​​Hardware Environment Configuration

[0086] The hardware environment used in this case is an NVIDIA A40 GPU with 48GB of video memory and 8GB of RAM. The CPU model is a 20vCPU Intel(R) Xeon(R) Platinum 8470Q. The operating system is Ubuntu 20.04 with CUDA 10.0. The software versions are Python 3.8 and torch 2.1.2.

[0087] PART 5, Evaluation Indicators

[0088] This study employed four evaluation metrics to comprehensively assess the effectiveness of the source tracing technology: Traceability Accuracy (TA), BLEU score, Sentence Similarity (SS), and Word Similarity (WS). Traceability Accuracy determines whether the visual language model can accurately identify embedded latent markers, measuring source tracing capability. The BLEU score assesses the lexical accuracy and fluency of the generated text by calculating the n-gram overlap between the generated and reference texts. Sentence Similarity measures the overall semantic similarity between the generated and reference texts based on the CLIP encoder, reflecting the degree of semantic preservation. Word Similarity uses BERT embeddings to perform fine-grained comparisons between the generated and reference texts at the word level, capturing semantic consistency at the word level.

[0089] These metrics, from multiple dimensions such as source tracing ability, semantic preservation, and language naturalness, can comprehensively quantify the performance of the technology in this case.

[0090] Table 1. Comparison of different evaluation criteria for different multimodal large language models.

[0091]

[0092] Table 1 provides a systematic comparison of the proposed method (GhostMarking) with two baseline methods (Image Noise and PromptMark). The comparison results show that GhostMarking achieves perfect source tracing accuracy across multiple visual language models and application scenarios, significantly outperforming the image perturbation method (Image Noise) and the prompt word embedding method (PromptMark), which fail to achieve effective tracking. Furthermore, GhostMarking maintains high BLEU scores, sentence similarity, and word similarity scores while preserving the text generation capability of the VLM model. For example, on the LLaVA model, these scores reach 0.28, 0.91, and 0.83, respectively, validating that the proposed method possesses robust traceability and good generalization ability while maintaining semantic consistency. These results fully demonstrate that GhostMarking has superior content source tracing performance compared to traditional input perturbation methods, reflecting its practicality and effectiveness in multiple scenarios.

[0093] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any way, and all technical solutions obtained by equivalent substitution or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A method for tracing the source of content generated by a visual language model, characterized in that: The visual language model generates traceable text based on traceable triggers and implicit markers. This tracing method includes the following steps: S1. Construct a set of hidden markers and a set of prompts: Construct a set of hidden markers and a set of prompts. A set of invisible hidden markers and by A collection of prompts ; S2. Constructing a set of non-target markers: From the set of hidden markers... Select from the original image A corresponding hidden marker is used as the target marker. ,the remaining Each hidden marker is used as a non-target marker, forming a set of non-target markers. , ; S3. Construct and optimize the traceable trigger: Construct the traceable trigger. In the original image Embedded traceable trigger Obtain source image Through traceable triggers Let the source image Carrying and target markers Corresponding, specific, and traceable source information; utilizing non-target tag sets For traceable triggers Perform adversarial supervised training to enhance the pre-trained visual language model. Non-target markers The ability to distinguish; including the following steps: (1) Set learnable parameters , and maximum number of iterations , ; (2) From the set of prompts respectively Non-target tag set The prompt message for selecting the current iteration by modulo index is in the middle. Non-target tags ; (3) Transfer the original image and prompts Input to visual language model In the process, the generated text is obtained. ; (4) In the original image Add a traceable trigger Obtain source image ; (5) Source tracing images and prompts Input to visual language model In the process, the generated text is obtained. ; (6) In generating text Embedding non-target tags Calculate the non-target label loss : , in: Representation of visual language model The loss function; (7) In generating text Embedded target tags Calculate the target label loss : , (8) Calculate the combined loss ; (9) Based on joint losses Calculate about the source image gradient And back-optimize the traceable trigger using a signed gradient approach. This makes the original image Add updated traceable triggers The obtained source image is updated to To achieve source tracing images The effect of embedding perturbations; (10) Judgment Is it equal to If so, then the traceable trigger is complete. Training; otherwise, Return to step (2); S4, Test Traceable Trigger For a given set of implicit tags and the original image Embedded invisible tags Front and back visual language models Generate text , Predicting hidden markers for: ; The stealth mark is an invisible mark composed of invisible Unicode characters, and the stealth mark is triggered by a traceable trigger. Embedded into the original image in an indirect manner Then, it is embedded in the visual language model. In the generated text, each invisible token is encoded from a randomly generated string. The string is first converted into a binary sequence and then mapped to the corresponding invisible Unicode character to form a unique invisible token.

2. The method for tracing the source of content generated by the visual language model according to claim 1, characterized in that: The visual language model comprises two core modules: a visual encoder and a large language model. The visual encoder employs a pre-trained CLIP model and is responsible for extracting visual features from the input image. The large language model uses a GPT, LLaMA, or Vicuna architecture and is responsible for extracting features from the input image based on prompts. Through a lightweight visual-language adaptation layer or a cross-modal attention mechanism, visual features are mapped to the embedding space of the large language model, enabling the large language model to understand the input image and generate text.

3. The method for tracing the source of content generated by the visual language model according to claim 1, characterized in that: The original image Some of the prompts are derived from the MS-COCO validation set, and a variety of prompts are automatically generated using GPT-4o and DeepSeek-V3. Simultaneously, visual-language association data from real-world scenarios are collected to further refine the original image set and prompt set, resulting in the final original image set. and a collection of prompts Original image set and a collection of prompts It covers three representative information types: identifiable information, inferable information, and associative information.

4. A source tracing device for content generated by a visual language model, used to implement the source tracing method for content generated by a visual language model as described in any one of claims 1 to 3; characterized in that: It includes a raw image and prompt acquisition module, a stealth marker generation module, a traceable trigger, a visual language model, an adversarial loss optimization module, and a stealth marker prediction module; The original image and prompt acquisition module is used to construct an original image set and a prompt set; The stealth marker generation module first randomly generates a string, then converts the string into a binary sequence, and finally maps the binary sequence to the corresponding invisible Unicode character to form a unique stealth marker; all stealth markers constitute a stealth marker set; based on the original image, one stealth marker is selected from the stealth marker set as the target marker, and the rest constitute a non-target marker set; The traceable trigger embeds the source information carried by the target marker into the original image to obtain the source image; The visual language model generates text based on prompts obtained from the original image and the source image; The adversarial loss optimization module uses a set of non-target labels to perform adversarial supervised training on the traceable triggers, thereby enhancing the ability of the pre-trained visual language model to distinguish non-target labels. The latent marker prediction module predicts latent markers from a given set of latent markers by using the generated text of a visual language model before and after embedding the latent markers in the original image through a trained traceable trigger.

5. The source tracing device for content generated by a visual language model according to claim 4, characterized in that: The adversarial loss optimization module jointly uses target label loss and non-target label loss to optimize the traceable trigger. The target label loss minimizes the difference between the generated text based on the original image with target label and the generated text based on the source image, so that the target label is embedded into the semantic space of the text generated by the visual language model through the traceable trigger. The non-target label loss maximizes the difference between the generated text based on the original image with non-target label and the generated text based on the source image, suppressing the interference of non-target labels on the text generated by the visual language model, so that the text generated by the visual language model can be attributed to a unique target label.

6. The source tracing device for content generated by a visual language model according to claim 4, characterized in that: The occult marker prediction module identifies embedded occult markers by minimizing the difference between generated text based on the original image and generated text based on the source image.