Unified identification method and system for multi-modal artificial intelligence generated content

By employing a unified identification method for content generated by multimodal artificial intelligence and adopting a dual-mode training framework of fast and slow thinking, a full-modal interpretable anti-spoofing dataset is constructed. This solves the problem of insufficient cross-modal identification capabilities, realizes unified identification and interpretable reasoning of cross-modal content, and adapts to the needs of rapid detection and detailed explanation in multiple scenarios.

CN121935526APending Publication Date: 2026-04-28FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUDAN UNIVERSITY
Filing Date
2025-12-31
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing authentication models lack cross-modal unified authentication capabilities, struggle to handle mixed forgery scenarios, and lack clear reasoning and high-latency interpretability, making it difficult to meet the real-time requirements of low-latency scenarios such as social media and news dissemination.

Method used

A unified identification method for content generated by multimodal artificial intelligence is adopted. Through a dual-mode training framework of fast and slow thinking, a full-modal interpretable anti-counterfeiting dataset is constructed to achieve cross-modal authenticity judgment and interpretable reasoning. Combined with spatiotemporal positioning mechanism and collaborative annotation mechanism, an interpretable reasoning chain from evidence extraction to criterion induction is constructed.

Benefits of technology

It achieves unified identification of cross-modal content, can stably output judgment results across different modalities, accurately identify forged segments and regions, improve the credibility and interpretability of content review and source tracing, and adapt to the needs of rapid detection and detailed explanation in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935526A_ABST
    Figure CN121935526A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal artificial intelligence generation content unified identification method and system. For multi-modal to-be-identified data, directly identifying the multi-modal to-be-identified data by adopting a fast thinking mode or performing interpretability identification by adopting a slow thinking mode based on the scene demand identification large model; the large identification model is trained by adopting a fast and slow thinking dual-mode two-stage training framework, an interpretable authentic identification data set is constructed, and standardized labeling from evidence extraction, space-time positioning to thinking chain construction is carried out. In a dual-mode supervision fine tuning stage, for a fast thinking mode, supervision fine tuning is utilized to realize fast discrimination and multi-classification traceability in a high-concurrency and low-delay scene; for a slow thinking mode, learning a complete interpretable reasoning chain from evidence extraction to criterion induction to conclusion generation; then, the model output is optimized through a preference alignment stage. Compared with the prior art, the method can provide a unified and reliable technical scheme for authenticity verification and interpretable review in a complex content scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence content security and information traceability technology, and in particular to a method and system for identifying and tracing artificial intelligence-generated content. Background Technology

[0002] In recent years, generative artificial intelligence has made leaps and bounds in multimodal fields such as language, vision, and hearing. From images to videos, from text to speech, and then to multimodal joint generation, the realism and coherence of models have been continuously improved. Modern image generation models can reconstruct highly natural lighting, texture, and spatial details; video generation technology has made significant breakthroughs in long-term temporal consistency, action logic, and physical plausibility; speech and audio generation can accurately replicate human voices, emotions, and environmental sounds with very few samples; in terms of text generation, large models can now imitate human writing styles and generate multilingual news, commentary, and narrative content. In particular, the development of cross-modal systems has made "text-to-image," "text-to-video," and "image-to-animation" capabilities commonplace, greatly lowering the threshold for virtual content creation.

[0003] Meanwhile, traditional methods of content alteration and synthesis (such as face swapping, splicing, editing, and forging subtitles and audio dubbing) continue to evolve, overlapping with AI generation technologies to create new, more deceptive forms of forged content. This type of "hybrid forgery" poses risks to authenticity and a crisis of social trust in scenarios such as social media, news dissemination, judicial evidence collection, financial risk control, and education and training. For example, short videos and voice synthesis are being abused for identity theft and fabricating false evidence, while text and image generation are being used to fabricate public opinion or mislead the public, posing serious challenges to content security review and accountability mechanisms.

[0004] Existing models for identifying authenticity have made significant progress in each single modality—image and video detection can identify forgeries through texture residuals, lighting inconsistencies, or generative model features, while text-generated content can be distinguished through linguistic statistical features and style biases. However, these methods are mostly based on single-modal designs and lack a unified cross-modal identification model, making it difficult to cope with complex mixed forgery scenarios in reality. For example, a news video often contains AI-generated image clips, altered voice narration, and forged subtitle text simultaneously. Single-modal detection cannot reveal the logical connections between multiple sources of forgery, making it difficult to form an overall judgment. In addition, some models operate as black-box classifiers. Although they perform well in terms of discrimination accuracy, they fail to provide clear reasoning and cannot support the requirements for evidence chains and verifiability in content censorship and judicial evidence collection. Some studies have attempted to introduce interpretability enhancement or visualization localization mechanisms to improve the understandability of the results, but these methods often have high inference latency, making them difficult to apply in low-latency scenarios such as social media monitoring, public opinion emergency response, and real-time review. As a result, the current detection system still has significant gaps between accuracy, interpretability, and real-time performance, and has not yet formed a unified anti-counterfeiting solution that can simultaneously take into account full modal consistency and multi-scenario adaptability. Summary of the Invention

[0005] The purpose of this invention is to address the problems of current authenticity verification schemes, such as cross-modal unified verification capability, weak ability to provide clear reasoning basis, and high reasoning latency, and to provide a unified verification method and system for multimodal artificial intelligence generated content.

[0006] The objective of this invention can be achieved through the following technical solutions: As a first aspect of the present invention, a unified identification method for multimodal AI-generated content is provided. The method inputs multimodal data to be identified into a dialogic unified identification model. For fast-response identification scenarios, the identification model directly identifies the data using a fast-thinking mode. For detailed analysis identification scenarios, the identification model performs interpretable identification using a slow-thinking mode and outputs an interpretable reasoning chain from evidence extraction to criterion induction to conclusion generation. The large-scale discrimination model is trained using a two-stage training framework with both fast and slow thinking modes, and the steps are as follows: Construct a multimodal interpretable fake-detection dataset, and standardize the annotation of evidence extraction, spatiotemporal positioning, and thought chain construction by establishing a unified criterion system and collaborative annotation mechanism; In the dual-mode supervised fine-tuning stage: For the fast thinking mode, supervised fine-tuning is performed using true and false binary classification labels constructed based on the collected dataset and multi-class classification labels of the generation model, training the model to quickly complete true and false discrimination and generation source identification; for the slow thinking mode, the criteria label with spatiotemporal localization is used as the supervision signal, training the model to generate an interpretable reasoning chain from evidence extraction to criteria induction to conclusion generation. In the preference alignment phase: Targeting the slow thinking mode, we collect supervised fine-tuned model responses and construct preference pairs, then use reinforcement learning to optimize the model for preference alignment.

[0007] As a preferred technical solution, the process of performing spatiotemporal positioning and interpretability annotation on the data is as follows: Collect real and fake content samples of multiple modalities, including traditional fake content samples and AI-generated samples; The annotation work is completed collaboratively by human annotation experts and multiple multimodal large models: based on the established interpretability and counterfeit detection criteria for each modality, information used to determine whether the content is counterfeit or AI-generated is identified; based on the established description specifications for each modality, text descriptions are used to record evidence features, and the descriptions include time and space orientations to reflect the fragments or regions where counterfeiting occurred; Based on the text descriptions obtained from collaborative annotation, forgery signs are annotated in both time and space dimensions: time or structural segmentation is completed according to the characteristics of different modalities; for image and video modalities, multimodal models are used to locate relevant regions in the corresponding frames and generate bounding boxes or region markers based on the semantic information in the text descriptions. Based on criteria and spatiotemporal annotation, and leveraging the strong reasoning capabilities of closed-source large models, a sequential reasoning process is constructed for the content to be identified according to modal characteristics: videos and audios are unfolded in chronological order, images are scanned step by step according to spatial regions, and texts are analyzed sequentially according to paragraph structure; the evidence extraction process is constructed with the thought chain of clues to evidence extraction and then to criterion induction as the main line, and serves as an interpretable reasoning sample in the training process.

[0008] As a preferred technical solution, after the sequential reasoning is constructed, a joint human-model review is performed to eliminate erroneous samples. The retained labeled samples are divided into two levels, high confidence and low confidence, based on the quality and consistency of evidence, logic, and criteria. During the training process, high confidence samples are used to construct interpretable reasoning samples.

[0009] As a preferred technical solution, the dual-mode supervised fine-tuning stage constructs two types of training data to achieve dual-mode learning based on both fast and slow thinking: For the fast thinking mode, the true and false binary classification labels and the multi-class classification labels of the generation model source are constructed based on the collected dataset, and the labels are fixed in the results to train the model to quickly complete the true and false discrimination and the generation source identification. For the slow thinking mode, labeled samples in the multimodal interpretability authentication dataset are used as interpretable reasoning samples. The evidence retrieval and reasoning process in the reasoning samples are placed in the evidence extraction of the thinking chain, the inductively formed criteria are placed in the criteria induction of the thinking chain, and the final conclusion is written into the conclusion of the thinking chain. This trains the model to learn the complete thinking chain from evidence search to criteria induction with spatiotemporal positioning and then to result determination.

[0010] As a preferred technical solution, the preference alignment stage collects discriminative responses from the model oriented towards the slow thinking mode; Quality stratification is performed based on interpretability and spatiotemporal positioning accuracy: low-quality outputs that do not meet the set requirements in terms of length, argumentation, evidence items, and spatiotemporal positioning are used as negative examples, and the negative example set is expanded by pruning and / or constructing negative samples; data containing a standardized reasoning → evidence → answer thought chain that has been finely annotated and verified are used as positive examples; preference pairs are constructed based on negative and positive examples, and reinforcement learning is used to optimize the model for preference alignment.

[0011] As a second aspect of the present invention, a unified identification system for multimodal AI-generated content is provided. The system executes the unified identification method for multimodal AI-generated content as described above, specifically including: Fast Response Identification Module: In fast response identification scenarios, large-scale multimodal data to be identified are directly identified using a fast-thinking mode; Fine analysis and identification module: In the fine analysis and identification scenario, the large model of identification of multimodal data to be identified adopts the slow thinking mode to perform interpretability identification, and outputs an interpretable reasoning chain from evidence extraction to criterion induction to conclusion generation. The large-scale discrimination model is trained using a two-stage training framework with both fast and slow thinking modes, and the steps are as follows: Construct a multimodal interpretable fake-detection dataset, and standardize the annotation of evidence extraction, spatiotemporal positioning, and thought chain construction by establishing a unified criterion system and collaborative annotation mechanism; In the dual-mode supervised fine-tuning stage: For the fast thinking mode, supervised fine-tuning is performed using true and false binary classification labels constructed based on the collected dataset and multi-class classification labels of the generation model, training the model to quickly complete true and false discrimination and generation source identification; for the slow thinking mode, the criteria label with spatiotemporal localization is used as the supervision signal, training the model to generate an interpretable reasoning chain from evidence extraction to criteria induction to conclusion generation. In the preference alignment phase: Targeting the slow thinking mode, we collect supervised fine-tuned model responses and construct preference pairs, then use reinforcement learning to optimize the model for preference alignment.

[0012] As a preferred technical solution, the process of performing spatiotemporal positioning and interpretability annotation on the data is as follows: Collect real and fake content samples of multiple modalities, including traditional fake content samples and AI-generated samples; The annotation work is completed collaboratively by human annotation experts and multiple multimodal large models: based on the established interpretability and counterfeit detection criteria for each modality, information used to determine whether the content is counterfeit or AI-generated is identified; based on the established description specifications for each modality, text descriptions are used to record evidence features, and the descriptions include time and space orientations to reflect the fragments or regions where counterfeiting occurred; Based on the text descriptions obtained from collaborative annotation, forgery signs are annotated in both time and space dimensions: time or structural segmentation is completed according to the characteristics of different modalities; for image and video modalities, multimodal models are used to locate relevant regions in the corresponding frames and generate bounding boxes or region markers based on the semantic information in the text descriptions. Based on criteria and spatiotemporal annotation, and leveraging the strong reasoning capabilities of closed-source large models, a sequential reasoning process is constructed for the content to be identified according to modal characteristics: videos and audios are unfolded in chronological order, images are scanned step by step according to spatial regions, and texts are analyzed sequentially according to paragraph structure; the evidence extraction process is constructed with the thought chain of clues to evidence extraction and then to criterion induction as the main line, and serves as an interpretable reasoning sample in the training process.

[0013] As a preferred technical solution, after the sequential reasoning is constructed, a joint human-model review is performed to eliminate erroneous samples. The retained labeled samples are divided into two levels, high confidence and low confidence, based on the quality and consistency of evidence, logic, and criteria. During the training process, high confidence samples are used to construct interpretable reasoning samples.

[0014] As a preferred technical solution, the dual-mode supervised fine-tuning stage constructs two types of training data to achieve dual-mode learning based on both fast and slow thinking: For the fast thinking mode, the true and false binary classification labels and the multi-class classification labels of the generation model source are constructed based on the collected dataset, and the labels are fixed in the results to train the model to quickly complete the true and false discrimination and the generation source identification. For the slow thinking mode, labeled samples in the multimodal interpretability authentication dataset are used as interpretable reasoning samples. The evidence retrieval and reasoning process in the reasoning samples are placed in the evidence extraction of the thinking chain, the inductively formed criteria are placed in the criteria induction of the thinking chain, and the final conclusion is written into the conclusion of the thinking chain. This trains the model to learn the complete thinking chain from evidence search to criteria induction with spatiotemporal positioning and then to result determination.

[0015] As a preferred technical solution, the preference alignment stage collects discriminative responses from the model oriented towards the slow thinking mode; Quality stratification is performed based on interpretability and spatiotemporal positioning accuracy: low-quality outputs that do not meet the set requirements in terms of length, argumentation, evidence items, and spatiotemporal positioning are used as negative examples, and the negative example set is expanded by pruning and / or constructing negative samples; data containing a standardized reasoning → evidence → answer thought chain that has been finely annotated and verified are used as positive examples; preference pairs are constructed based on negative and positive examples, and reinforcement learning is used to optimize the model for preference alignment.

[0016] Compared with the prior art, the present invention has the following beneficial effects: 1) This invention proposes a method for identifying and tracing forged content based on a multimodal large model. It can simultaneously process multiple modalities of content, including images, videos, audio, and text, achieving cross-modal authenticity determination and interpretable reasoning within a unified model architecture. This unified system effectively breaks down the barriers between feature differences and criterion expression between different modalities, enabling the model to stably output consistent judgment results in single-input or multimodal fusion scenarios. By introducing a spatiotemporal localization mechanism at the criterion level, the system can not only identify the authenticity of content but also accurately pinpoint the segments and regions where forgery occurs, achieving integrated judgment, localization, and interpretation, significantly improving the credibility of content review and tracing.

[0017] 2) This invention establishes a unified, interpretable evidence annotation system across modalities. Through a multi-stage annotation mechanism involving human-model collaboration, combined with semantic deduplication, temporal segmentation, and spatial localization, interpretable annotation is achieved throughout the entire process from evidence extraction to interpretable authentication. This mechanism not only significantly improves annotation consistency and verifiability but also provides high-quality data support for subsequent interpretable reasoning and audit documentation of the model.

[0018] 3) This invention proposes a dual-mode model optimization framework for counterfeit detection tasks. During the training phase, a hybrid optimization strategy employing both fast and slow thinking modes is used to mutually promote rapid discrimination and interpretable reasoning capabilities. In deployment and inference, the fast-slow thinking dual-mode structure allows the system to adaptively switch working modes according to different application scenarios: in high-concurrency, low-latency scenarios, it can quickly complete authenticity screening and source tracing; in scenarios requiring detailed analysis, such as judicial evidence collection, media review, or brand compliance, it can output detailed explanatory reports with spatiotemporal positioning criteria and evidence chain support, demonstrating high versatility and scalability. Attached Figure Description

[0019] Figure 1 This is a flowchart of a unified identification method for multimodal artificial intelligence-generated content according to the present invention.

[0020] Figure 2 This is a schematic diagram illustrating data collection, annotation, and model training in this invention. Detailed Implementation

[0021] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0022] Example 1 This invention provides a unified method for identifying multimodal AI-generated content, enabling unified authentication of multi-source content such as images, videos, audio, and text. The technical solution comprises two core components: the construction and annotation of a full-modal interpretable authentication dataset, and supervised fine-tuning and preference alignment training of a dual-mode (fast and slow thinking) authentication model based on this dataset. By constructing an interpretable authentication dataset covering multiple modalities such as images, videos, audio, and text, a unified criterion system and collaborative annotation mechanism are established, achieving a standardized annotation process from evidence extraction and spatiotemporal localization to thought chain construction. Based on this, a two-stage training framework of "fast and slow thinking" dual modes is adopted: in the fast thinking mode, supervised fine-tuning (SFT) is used to achieve rapid discrimination and multi-class source tracing in high-concurrency, low-latency scenarios; simultaneously, in the slow thinking mode, a complete interpretable reasoning chain is learned from evidence extraction to criterion induction and conclusion generation. Subsequently, the model output is optimized through a preference alignment phase, which encourages the generation of results to be further improved in terms of expression diversity, evidence completeness and verifiability, so that the model can not only respond efficiently to actual audit requirements, but also output judgment explanations supported by evidence chains.

[0023] This invention introduces a dual-mode "fast and slow thinking" structure, enabling rapid forgery detection in real-time scenarios and generating highly interpretable reasoning results. This provides an efficient, unified, and verifiable technical solution for multimodal content security review and trusted information verification. Figure 1 As shown, the specific steps for verifying the authenticity of the content include: Step S1: Obtain multimodal data to be identified; Step S2: Based on the current identification scenario, choose to perform either rapid response identification or interpretable identification; Step S3: For fast-response identification scenarios, the large identification model adopts a fast-thinking mode for direct identification; Step S4: For detailed analysis and identification scenarios, the identification model adopts a slow thinking mode for interpretability identification and outputs an interpretable reasoning chain from evidence extraction to criterion induction to conclusion generation. Step S5: Output the identification results of the dialogic unified identification model.

[0024] The dialogic unified identification model employs a two-stage training framework with both fast and slow thinking modes, as detailed below: I. Construction and labeling of a multimodal interpretability fake-detection dataset.

[0025] like Figure 2 As shown, this invention addresses consistent multi-source modalities by establishing a standardized criterion and annotation system. The dataset construction comprises six main steps: data collection and cleaning, criterion specification formulation, collaborative annotation, spatiotemporal localization, thought chain construction, verification, and hierarchical screening. Through these processes, an interpretable anti-fakeness dataset covering multiple modalities such as images, videos, audio, and text is constructed. This achieves standardization across the entire process, from raw content acquisition to fine-grained annotation and quality verification, providing a unified and high-confidence data foundation for subsequent dual-mode model training and evaluation.

[0026] 1.1 Data collection and cleaning.

[0027] This step systematically collects both real and forged content samples across multiple modalities, including images, videos, audio, and text. Data sources include publicly available multimodal datasets, authorized authentic materials, and content synthesized by mainstream generative models. AI-generated samples encompass various types, such as image and video diffusion models and text and speech autoregressive generation models, while deepfake samples include diverse forms such as face swapping, splicing, editing, forged subtitles, and audio synthesis. A unified cleaning process is used to deduplicate and detect anomalies in the multi-source data, ensuring that the samples cover both AI-generated and traditionally tampered scenarios, laying a high-quality data foundation for subsequent annotation and model training.

[0028] 1.2. Formulation of judgment criteria and standards.

[0029] This step involves experts in forensics and content security conducting a systematic analysis of multimodal data. Combining existing forgery sample characteristics with typical forensic experience, they summarized interpretable forgery detection criteria and descriptive standards for each modality. These standards clarify the forgery signs, expression methods, and terminology standards to focus on in different modalities, ensuring consistency and verifiability in the annotation process. This criterion system also provides a unified judgment basis for both manual and model-based collaborative annotation, becoming a core reference standard for subsequent interpretable annotation and mind chain construction.

[0030] 1.3 Collaborative annotation.

[0031] Guided by established criteria, this step involves collaborative annotation work by human annotation experts and various multimodal large-scale models. Annotators identify key information that can be used to determine whether content is forged or AI-generated based on the criteria, and record the evidentiary features through detailed textual descriptions. The descriptions must include clear temporal and spatial indications to reflect the segment or area where the forgery occurred, but precise coordinates are not required. This collaborative process leverages the complementarity of human and model understanding methods, improving the coverage and diversity of the annotation results and providing rich expressive samples for subsequent inference training.

[0032] 1.4 Spatiotemporal positioning.

[0033] This step, based on the text descriptions obtained through collaborative annotation, precisely labels forgery signs in both temporal and spatial dimensions. First, it segments the text temporally or structurally according to the characteristics of different modalities: videos and audios are segmented by time segments, and text is divided by paragraph order. Then, for image and video modalities, a multimodal model is used to locate relevant regions in the corresponding frames and generate bounding boxes or region markers based on the semantic information in the labeled descriptions. Through this process, the original linguistic evidence descriptions are transformed into structured annotations with spatiotemporal directionality, providing a precise spatiotemporal localization foundation for subsequent interpretable reasoning and model training.

[0034] 1.5. Building a thought chain.

[0035] Building upon the aforementioned refined criteria and spatiotemporal annotations, and leveraging the powerful reasoning capabilities of a closed-source large-scale model, a sequential reasoning process is constructed based on the modal characteristics of the content to be identified: videos and audios are presented chronologically, images are scanned progressively by spatial region, and text is analyzed sequentially by paragraph structure. The thought chain follows the main thread of "clues → evidence extraction → criterion induction," ensuring that the model does not overlook details or key evidence during reasoning. The primary objective of this stage is to utilize the large-scale model's reasoning generation capabilities to construct a high-quality evidence extraction process and embed it as a reasoning example in subsequent training, enabling the model to generate criteria and reasoning paths more completely and systematically during autonomous judgment.

[0036] 1.6 Review and hierarchical screening.

[0037] This step is jointly completed by human experts and a multimodal model. Each thought chain and corresponding evidence generated in the initial annotation phase undergoes a final review to verify its sufficiency as valid evidence to determine whether the content is forged or AI-generated. Samples with sufficient evidence, clear logic, and explicit criteria are marked as high-confidence; samples with ambiguity or reliance on auxiliary information (such as unclear images, distorted speech, insufficient text context, etc.) are marked as low-confidence. This hierarchical mechanism ensures comprehensive coverage of annotated samples while preventing low-confidence samples from misleading reviewers, providing a clear reference hierarchy for subsequent model training and manual verification.

[0038] II. Dual-mode supervision fine-tuning and preference alignment optimization.

[0039] The dual-mode supervised fine-tuning and preference alignment optimization of this invention includes two stages: First, in the supervised fine-tuning (SFT) stage, the model is trained using both fast thinking and slow thinking modes to enable it to possess efficient true / false discrimination and interpretable reasoning capabilities, respectively; then, in the preference alignment (RL) stage, the model output is optimized based on human or model preference signals to further improve its diversity, verifiability, and rationality of expression, thereby forming a unified false-proofing model that combines accuracy and interpretability.

[0040] In the dual-mode supervised fine-tuning stage, two types of training data are constructed to achieve dual-mode learning with both fast and slow thinking. The first type consists of true / false binary classification labels constructed based on the collected dataset and multi-class classification labels generated by the model. The results are fixed in the inference process labels of the model's structured output. <answer> …< / answer> In the first category, the "fast thinking" mode is used to train models to quickly distinguish between true and false information and identify the source of the data. The second category constructs interpretable reasoning samples based on high-confidence samples, placing the evidence retrieval and reasoning process within the thought process. <think> …< / think> In this context, the inductively derived criteria are placed as evidence markers within the thought chain. <evidence>…in which the final conclusion is written into the conclusion marker of the thought process. <answer> …< / answer> In this context, the complete thought process used to train a model, from evidence searching to criterion induction with spatiotemporal positioning and then to result determination, is called the "slow thinking" mode.

[0041] After completing the supervised fine-tuning of the SFT, the model's discriminative responses were collected for the "slow thinking" scenario. Quality was stratified based on interpretability and spatiotemporal positioning accuracy: low-quality outputs such as those that were too short, weakly argued, had too few / missing evidence, or incorrect spatiotemporal positioning were designated as negative examples. A portion of "short responses" or "incorrect interpretations" were manually trimmed / constructed to expand the negative example set. Simultaneously, data that had undergone refined annotation and verification (including standardized...) was used... <think> → <evidence> → <answer>The chain of evidence is used as a positive example. Based on this, preference pairs (positive > negative) are constructed, and the model is optimized for preference alignment. This enables the model to generate a more complete chain of evidence, a more sufficient and accurate statement of criteria, and a clearer conclusion while maintaining the correctness of the judgment, thereby improving the verifiability and practical usability of the review.

[0042] This invention constructs a comprehensive, interpretable, and robust authentication model to achieve integrated recognition of multi-source content, including images, videos, audio, and text. It bridges the gap between AI-generated content and traditionally altered content, forming a unified framework for authenticity determination. Furthermore, this invention introduces a dual-mode "fast and slow thinking" structure. In fast mode, it enables rapid detection and source tracing in low-latency scenarios; in slow mode, it generates interpretable reasoning results based on evidence chains, meeting the verifiable requirements of review and evidence collection.

[0043] Especially in time-sensitive scenarios such as news and public communication, the model of this invention can serve as a key authentication component in a multimodal content review system, working in conjunction with the upper-level modal decomposition and fact-matching modules. The upper-level system first decomposes video, audio, and text content, extracting key visuals, audio, and textual clues. Subsequently, the model of this invention performs fine-grained analysis of these clues from the perspective of content authenticity, identifying potential signs of forgery or abnormal patterns, and generating an interpretable chain of evidence. This process, combined with external fact-checking and news source tracing modules, can provide credible content authentication support in situations involving the discovery of fake news videos and misleading reports, becoming a core component of a time-sensitive review system.

[0044] This invention can be widely applied in fields such as multimodal content security, media forensics, and intelligent review, and is suitable for various scenarios ranging from social media monitoring and news content verification to judicial evidence analysis, copyright protection, and content tracing. In platform-level content review, it can be embedded as a core authentication module in automated review processes, quickly identifying and labeling AI-generated or traditionally altered content, providing an interpretable chain of evidence for manual review. In judicial and evidence-gathering scenarios, it can be used to authenticate the authenticity of videos, audio, or text involved in a case, providing traceable evidence criteria. In news and public information dissemination, it can collaborate with fact-checking systems to identify mixed-modal fake news or misleading content. For example, in news verification of breaking events, when the system receives a report containing video, narration, and accompanying text, the model of this invention, as a media-layer content authenticity authentication module, can quickly identify possible signs of synthesis and alteration in the images, audio, or text, and generate corresponding interpretable criteria; simultaneously, the semantic-layer fact-checking system performs real-time comparison and verification of the event's factual content (such as time, location, people, and narrative logic) to determine the authenticity of the message itself. When these two work together, they can both enable a rapid response to multimodal false content and provide evidence-supported, interpretable results for review and tracing.

[0045] Example 2 As another embodiment of the present invention, this embodiment also provides a system for identifying and tracing the source of counterfeit content. This system executes the method for identifying and tracing the source of counterfeit content as described in Embodiment 1, specifically including: Fast Response Identification Module: In fast response identification scenarios, large-scale multimodal data to be identified are directly identified using a fast-thinking mode; Fine analysis and identification module: In the fine analysis and identification scenario, the large model of identification of multimodal data to be identified adopts the slow thinking mode to perform interpretability identification, and outputs an interpretable reasoning chain from evidence extraction to criterion induction to conclusion generation. The large-scale discrimination model is trained using a two-stage training framework with both fast and slow thinking modes, and the steps are as follows: Construct a multimodal interpretable fake-detection dataset, and standardize the annotation of evidence extraction, spatiotemporal positioning, and thought chain construction by establishing a unified criterion system and collaborative annotation mechanism; In the dual-mode supervised fine-tuning stage: For the fast thinking mode, supervised fine-tuning is performed using true and false binary classification labels constructed based on the collected dataset and multi-class classification labels of the generation model, training the model to quickly complete true and false discrimination and generation source identification; for the slow thinking mode, the criteria label with spatiotemporal localization is used as the supervision signal, training the model to generate an interpretable reasoning chain from evidence extraction to criteria induction to conclusion generation. In the preference alignment phase: Targeting the slow thinking mode, we collect supervised fine-tuned model responses and construct preference pairs, then use reinforcement learning to optimize the model for preference alignment.

[0046] Furthermore, the standardized annotation process for evidence extraction, spatiotemporal localization, and thought chain construction is as follows: Collect real and fake content samples of multiple modalities, including traditional fake content samples and AI-generated samples; The annotation work is completed collaboratively by human annotation experts and multiple multimodal large models: based on the established interpretability and counterfeit detection criteria for each modality, information used to determine whether the content is counterfeit or AI-generated is identified; based on the established description specifications for each modality, text descriptions are used to record evidence features, and the descriptions include time and space orientations to reflect the fragments or regions where counterfeiting occurred; Based on the text descriptions obtained from collaborative annotation, forgery signs are annotated in both time and space dimensions: time or structural segmentation is completed according to the characteristics of different modalities; for image and video modalities, multimodal models are used to locate relevant regions in the corresponding frames and generate bounding boxes or region markers based on the semantic information in the text descriptions. Based on criteria and spatiotemporal annotation, and leveraging the strong reasoning capabilities of closed-source large models, a sequential reasoning process is constructed for the content to be identified according to modal characteristics: videos and audios are unfolded in chronological order, images are scanned step by step according to spatial regions, and texts are analyzed sequentially according to paragraph structure; the evidence extraction process is constructed with the thought chain of clues to evidence extraction and then to criterion induction as the main line, and serves as an interpretable reasoning sample in the training process.

[0047] After the sequential reasoning is constructed, a joint human-model review is conducted to eliminate erroneous samples. The retained labeled samples are divided into two levels, high confidence and low confidence, based on the quality and consistency of evidence, logic, and criteria. During the training process, high confidence samples are used to construct interpretable reasoning samples.

[0048] Furthermore, in the dual-mode supervised fine-tuning stage, two types of training data are constructed to achieve dual-mode learning of fast and slow thinking: For the fast thinking mode, the true and false binary classification labels and the multi-class classification labels of the generation model source are constructed based on the collected dataset, and the labels are fixed in the results to train the model to quickly complete the true and false discrimination and the generation source identification. For the slow thinking mode, labeled samples in the multimodal interpretability authentication dataset are used as interpretable reasoning samples. The evidence retrieval and reasoning process in the reasoning samples are placed in the evidence extraction of the thinking chain, the inductively formed criteria are placed in the criteria induction of the thinking chain, and the final conclusion is written into the conclusion of the thinking chain. This trains the model to learn the complete thinking chain from evidence search to criteria induction with spatiotemporal positioning and then to result determination.

[0049] Furthermore, in the preference alignment phase, the model's discriminative responses are collected in response to slow thinking patterns; Quality stratification is performed based on interpretability and spatiotemporal positioning accuracy: low-quality outputs that do not meet the set requirements in terms of length, argumentation, evidence items, and spatiotemporal positioning are used as negative examples, and the negative example set is expanded by pruning and / or constructing negative samples; data containing a standardized reasoning → evidence → answer thought chain that has been finely annotated and verified are used as positive examples; preference pairs are constructed based on negative and positive examples, and reinforcement learning is used to optimize the model for preference alignment.

[0050] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0051] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.< / answer> < / evidence> < / think> < / evidence>

Claims

1. A unified identification method for multimodal artificial intelligence-generated content, characterized in that, The method inputs multimodal data to be identified into a dialogic unified identification model. For rapid response identification scenarios, the identification model adopts a fast thinking mode for direct identification. For detailed analysis identification scenarios, the identification model adopts a slow thinking mode for interpretable identification and outputs an interpretable reasoning chain from evidence extraction to criterion induction to conclusion generation. The large-scale discrimination model is trained using a two-stage training framework with both fast and slow thinking modes, and the steps are as follows: Construct a multimodal interpretable anti-spoofing dataset and perform spatiotemporal localization interpretability annotation on the data; In the dual-mode supervised fine-tuning stage: For the fast thinking mode, supervised fine-tuning is performed using true and false binary classification labels constructed based on the collected dataset and multi-class classification labels of the generation model, training the model to quickly complete true and false discrimination and generation source identification; for the slow thinking mode, the criteria label with spatiotemporal localization is used as the supervision signal, training the model to generate an interpretable reasoning chain from evidence extraction to criteria induction to conclusion generation. In the preference alignment phase: Targeting the slow thinking mode, we collect supervised fine-tuned model responses and construct preference pairs, then use reinforcement learning to optimize the model for preference alignment.

2. The method for unified identification of multimodal artificial intelligence-generated content according to claim 1, characterized in that, The process of performing spatiotemporal location interpretability annotation on the data is as follows: Collect real and fake content samples of multiple modalities, including traditional fake content samples and AI-generated samples; The annotation work is completed collaboratively by human annotation experts and multiple multimodal large models: based on the established interpretability and counterfeit detection criteria for each modality, information used to determine whether the content is counterfeit or AI-generated is identified; based on the established description specifications for each modality, text descriptions are used to record evidence features, and the descriptions include time and space orientations to reflect the fragments or regions where counterfeiting occurred; Based on the text descriptions obtained from collaborative annotation, forgery signs are annotated in both time and space dimensions: time or structural segmentation is completed according to the characteristics of different modalities; for image and video modalities, multimodal models are used to locate relevant regions in the corresponding frames and generate bounding boxes or region markers based on the semantic information in the text descriptions. Based on criteria and spatiotemporal annotation, and leveraging the strong reasoning capabilities of closed-source large models, a sequential reasoning process is constructed for the content to be identified according to modal characteristics: videos and audios are unfolded in chronological order, images are scanned step by step according to spatial regions, and texts are analyzed sequentially according to paragraph structure; the evidence extraction process is constructed with the thought chain of clues to evidence extraction and then to criterion induction as the main line, and serves as an interpretable reasoning sample in the training process.

3. The method for unified identification of multimodal artificial intelligence-generated content according to claim 2, characterized in that, After the sequential reasoning is constructed, a joint human-model review is conducted to eliminate erroneous samples. The retained labeled samples are divided into two levels, high confidence and low confidence, based on the quality and consistency of evidence, logic, and criteria. During the training process, high confidence samples are used to construct interpretable reasoning samples.

4. The method for unified identification of multimodal artificial intelligence-generated content according to claim 1, characterized in that, The dual-mode supervised fine-tuning stage constructs two types of training data to achieve dual-mode learning based on both fast and slow thinking: For the fast thinking mode, the true and false binary classification labels and the multi-class classification labels of the generation model source are constructed based on the collected dataset, and the labels are fixed in the results to train the model to quickly complete the true and false discrimination and the generation source identification. For the slow thinking mode, labeled samples in the multimodal interpretability authentication dataset are used as interpretable reasoning samples. The evidence retrieval and reasoning process in the reasoning samples are placed in the evidence extraction of the thinking chain, the inductively formed criteria are placed in the criteria induction of the thinking chain, and the final conclusion is written into the conclusion of the thinking chain. This trains the model to learn the complete thinking chain from evidence search to criteria induction with spatiotemporal positioning and then to result determination.

5. The method for unified identification of multimodal artificial intelligence-generated content according to claim 4, characterized in that, The preference alignment phase involves collecting discriminative responses from the model based on the slow thinking mode. Quality is stratified based on interpretability and spatiotemporal positioning accuracy: low-quality outputs that do not meet the set requirements in terms of length, argumentation, evidence, and spatiotemporal positioning are used as negative examples, and the negative example set is expanded by pruning and / or constructing negative samples; data that contains a standardized reasoning → evidence → answer thought chain after being carefully annotated and reviewed are used as positive examples. Preference pairs are constructed based on negative and positive examples, and reinforcement learning is used to optimize the model for preference alignment.

6. A unified identification system for multimodal artificial intelligence-generated content, characterized in that, The system executes the unified identification method for multimodal AI-generated content as described in any one of claims 1-5, including: Fast Response Identification Module: In fast response identification scenarios, large-scale multimodal data to be identified are directly identified using a fast-thinking mode; Fine analysis and identification module: In the fine analysis and identification scenario, the large model of identification of multimodal data to be identified adopts the slow thinking mode to perform interpretability identification, and outputs an interpretable reasoning chain from evidence extraction to criterion induction to conclusion generation. The large-scale discrimination model is trained using a two-stage training framework with both fast and slow thinking modes, and the steps are as follows: Construct a multimodal interpretable fake-detection dataset, and standardize the annotation of evidence extraction, spatiotemporal positioning, and thought chain construction by establishing a unified criterion system and collaborative annotation mechanism; In the dual-mode supervised fine-tuning stage: For the fast thinking mode, supervised fine-tuning is performed using true and false binary classification labels constructed based on the collected dataset and multi-class classification labels of the generation model, training the model to quickly complete true and false discrimination and generation source identification; for the slow thinking mode, the criteria label with spatiotemporal localization is used as the supervision signal, training the model to generate an interpretable reasoning chain from evidence extraction to criteria induction to conclusion generation. In the preference alignment phase: Targeting the slow thinking mode, we collect supervised fine-tuned model responses and construct preference pairs, then use reinforcement learning to optimize the model for preference alignment.

7. The unified identification system for multimodal artificial intelligence-generated content according to claim 6, characterized in that, The process of performing spatiotemporal location interpretability annotation on the data is as follows: Collect real and fake content samples of multiple modalities, including traditional fake content samples and AI-generated samples; The annotation work is completed collaboratively by human annotation experts and multiple multimodal large models: based on the established interpretability and counterfeit detection criteria for each modality, information used to determine whether the content is counterfeit or AI-generated is identified; based on the established description specifications for each modality, text descriptions are used to record evidence features, and the descriptions include time and space orientations to reflect the fragments or regions where counterfeiting occurred; Based on the text descriptions obtained from collaborative annotation, forgery signs are annotated in both time and space dimensions: time or structural segmentation is completed according to the characteristics of different modalities; for image and video modalities, multimodal models are used to locate relevant regions in the corresponding frames and generate bounding boxes or region markers based on the semantic information in the text descriptions. Based on criteria and spatiotemporal annotation, and leveraging the strong reasoning capabilities of closed-source large models, a sequential reasoning process is constructed for the content to be identified according to modal characteristics: videos and audios are unfolded in chronological order, images are scanned step by step according to spatial regions, and texts are analyzed sequentially according to paragraph structure; the evidence extraction process is constructed with the thought chain of clues to evidence extraction and then to criterion induction as the main line, and serves as an interpretable reasoning sample in the training process.

8. The unified identification system for multimodal artificial intelligence-generated content according to claim 7, characterized in that, After the sequential reasoning is constructed, a joint human-model review is conducted to eliminate erroneous samples. The retained labeled samples are divided into two levels, high confidence and low confidence, based on the quality and consistency of evidence, logic, and criteria. During the training process, high confidence samples are used to construct interpretable reasoning samples.

9. A unified identification system for multimodal artificial intelligence-generated content according to claim 6, characterized in that, The dual-mode supervised fine-tuning stage constructs two types of training data to achieve dual-mode learning based on both fast and slow thinking: For the fast thinking mode, the true and false binary classification labels and the multi-class classification labels of the generation model source are constructed based on the collected dataset, and the labels are fixed in the results to train the model to quickly complete the true and false discrimination and the generation source identification. For the slow thinking mode, labeled samples in the multimodal interpretability authentication dataset are used as interpretable reasoning samples. The evidence retrieval and reasoning process in the reasoning samples are placed in the evidence extraction of the thinking chain, the inductively formed criteria are placed in the criteria induction of the thinking chain, and the final conclusion is written into the conclusion of the thinking chain. This trains the model to learn the complete thinking chain from evidence search to criteria induction with spatiotemporal positioning and then to result determination.

10. A unified identification system for multimodal artificial intelligence-generated content according to claim 9, characterized in that, The preference alignment phase involves collecting discriminative responses from the model based on the slow thinking mode. Quality is stratified based on interpretability and spatiotemporal positioning accuracy: low-quality outputs that do not meet the set requirements in terms of length, argumentation, evidence, and spatiotemporal positioning are used as negative examples, and the negative example set is expanded by pruning and / or constructing negative samples; data that contains a standardized reasoning → evidence → answer thought chain after being carefully annotated and reviewed are used as positive examples. Preference pairs are constructed based on negative and positive examples, and reinforcement learning is used to optimize the model for preference alignment.