An AI-generated content authenticity separation identification method for a propagation distorted image
By identifying the image propagation state and calling the corresponding detection branch, the effective content region is extracted. Combined with multimodal features for dynamic fusion, the problem of metadata failure during image propagation is solved, and stable and accurate identification under different propagation states is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 赖海波
- Filing Date
- 2026-05-01
- Publication Date
- 2026-06-26
AI Technical Summary
In the process of image propagation, existing technologies are prone to the failure of original metadata and generation parameters. A single visual model is difficult to adapt to image distortion under different propagation conditions and cannot effectively distinguish between the authenticity of the acquisition and the authenticity of the content. There is a lack of targeted methods to handle scenarios such as screenshots, screen captures, and printed captures.
By identifying the propagation state of the image, calling the corresponding detection branch, extracting the effective content region, and dynamically fusing it with multimodal features, configuring weights, generating the realism separation result, and outputting the detection report.
It improves the stability and accuracy of recognition in scenarios with propagation distortion, and can handle complex scenarios such as real mobile phone shooting, real screenshots and real printed copies, adapting to image recognition under different propagation conditions.
Smart Images

Figure CN122289807A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image forensics, AI-generated content detection, computer vision, multimodal recognition, edge image analysis, and image credibility verification, and particularly to a method for separating and recognizing the authenticity of AI-generated content for the propagation of distorted images. Background Technology
[0002] With the development of generative artificial intelligence technology, AI-generated images are becoming increasingly common in scenarios such as social media, news dissemination, e-commerce displays, advertising, content creation, online communities, and public services. Unlike the original generated files, images in real-world dissemination are often not the original images downloaded directly from the generation platform, but rather images that have undergone processes such as being transferred from social media platforms, compressed using chat tools, saved from web pages, captured by mobile phone screenshots, photographed from the screen, printed and then photographed again, cropped, filtered, or re-encoded.
[0003] During the aforementioned propagation process, information such as EXIF, XMP, C2PA, JUMBF, generation parameters, and software fields in the original image may be deleted, overwritten, or become unreadable. Therefore, simply reading the image file information is prone to failure in real-world propagation scenarios.
[0004] On the other hand, existing AI image recognition methods based on image models typically input the image to be detected directly into the classification model and output either a real or AI-generated judgment result. This type of method has certain effectiveness on original images or standard test sets, but in scenarios such as screenshots, platform compression, screen captures, and printed images, additional factors such as screen display artifacts, camera noise, paper texture, printing dots, and compression damage are superimposed on the image, making it difficult for the model to make stable judgments.
[0005] It is particularly important to note that some distorted images may possess characteristics of genuine capture. For example, a user might use a real phone to photograph an AI-generated image on a computer screen, or a phone to photograph a printed AI poster. While these images are indeed captured by real devices, the main content they convey may still be AI-generated. Judging the authenticity of an image solely based on camera noise, capture information, or paper texture can easily lead to mistaking genuine capture for genuine content.
[0006] Therefore, the existing technology has the following main shortcomings: 1. The file information reading method is prone to failure after screenshotting, saving, compressing, or photographing; 2. A single vision model struggles to adapt to image distortion under different propagation conditions; 3. Existing solutions typically do not first identify the image propagation state before selecting the corresponding detection strategy; 4. Existing solutions rarely distinguish between "data collection authenticity" and "content authenticity"; 5. For scenarios such as screenshots, platform transfers, screen captures, and printed captures, there is a lack of targeted multimodal evidence fusion and authenticity separation methods.
[0007] Based on this, it is necessary to propose a method for separating and identifying the authenticity of AI-generated content in the context of disseminating distorted images. When the original file information is unavailable or its credibility is reduced, this method combines the dissemination status, image content, text anomalies, frequency domain noise, screen capture features, and printed imaging features to separate and judge whether the image acquisition method is authentic and whether the content carried by the image has the risk of being generated by AI. Summary of the Invention
[0008] (a) Technical problems to be solved The technical problem this invention aims to solve is that after an image is transferred, screenshotted, photographed, or printed, the original metadata, content credentials, and generation parameters become invalid. Traditional methods based on file information or single image classification models struggle to reliably identify AI-generated content and easily confuse real data collection behavior with content authenticity.
[0009] To address this, the present invention provides an AI-generated content authenticity separation and recognition method for images with propagation distortion. By first determining the propagation state, then calling the detection branch according to the scenario, and separating the collection authenticity and content authenticity output, the recognition adaptability in propagation distortion scenarios is improved.
[0010] (II) Technical Solution To address the aforementioned technical problems, this invention provides a method for separating and recognizing the authenticity of AI-generated content in the propagation of distorted images, comprising the following steps: S101, acquire the image to be detected, and read the image data of the image to be detected. The image data includes at least image pixel data, resolution information, compression features, boundary region features, local texture features, and readable file structure information.
[0011] S102, Extracting propagation state features based on image data. Propagation state features include platform transfer features, screenshot features, screen capture features, and print capture features. Platform transfer features characterize whether the image has been compressed, scaled, resampled, or re-encoded by the platform; screenshot features characterize whether the image was obtained from a screen capture; screen capture features characterize whether the image was obtained by photographing a display screen; and print capture features characterize whether the image was obtained by photographing a paper print.
[0012] S103, determine the propagation state of the image to be detected based on the propagation state characteristics. The propagation state includes at least the state of the image transferred from the platform, the state of the screenshot image, the state of the screen-captured image, and the state of the printed image.
[0013] S104: Invoke the corresponding detection branch according to the propagation status and determine the valid content area for content judgment. For screenshot image status, extract the valid content area within the screenshot; for screen copy image status, separate the screen content area and the external environment area; for printed copy image status, separate the paper content area and the paper acquisition area; for platform-transferred image status, determine the compression damage level.
[0014] S105, Extract content authenticity features from the effective content region. Content authenticity features include at least two of the following: spatial visual anomaly features, OCR text anomaly features, frequency domain anomaly features, noise residual features, and semantic consistency features.
[0015] S106, Extracting realistic features from regions or features related to the acquisition process in the image to be detected. Realistic features include at least one of the following: screen capture imaging features, printed imaging features, screenshot interface features, platform transfer compression features, and shooting noise features.
[0016] S107: Assign corresponding weights to the content authenticity features and collection authenticity features based on the dissemination status, and calculate the content authenticity risk result and collection authenticity judgment result respectively. Different weight configurations are used under different dissemination statuses, ensuring that this method does not simply apply a fixed AI image recognition model, but rather selects and modifies valid evidence based on the dissemination status.
[0017] S108 generates an authenticity separation result based on the content authenticity risk result and the acquisition authenticity judgment result. The authenticity separation result is used to distinguish whether the acquisition method of the image to be detected has authentic acquisition characteristics, and whether the main content carried by the image has the risk of AI generation.
[0018] S109 outputs a detection report that includes the propagation status, the results of the authenticity assessment, the results of the content authenticity risk assessment, the risk level generated by AI, evidence items, and an explanation of uncertainty.
[0019] The core of this invention is not simply using an image classification model to determine whether an image is generated by AI, but rather establishing a correspondence between propagation status recognition, effective content area extraction, content authenticity judgment, collection authenticity judgment, and authenticity separation output for propagation distortion scenarios such as platform transfer, screenshot, screen copying, and printed copying.
[0020] (III) Beneficial Effects Compared with the prior art, the present invention has at least the following beneficial effects: 1. This invention does not rely solely on metadata, content credentials, or generation parameters as the basis for judgment, and can adapt to situations where file information becomes invalid after an image is transferred, screenshotted, photographed from the screen, or photographed from a printed document. 2. This invention first identifies the propagation state, and then calls the corresponding detection branch, so that images under different propagation states can use different evidence combinations and weight configurations; 3. This invention dynamically fuses multimodal features such as spatial visual anomalies, frequency domain anomalies, noise residuals, OCR text anomalies, screen capture imaging, printed imaging, and semantic consistency to improve recognition stability in propagation distortion scenarios. 4. This invention separates the judgment of the authenticity of the data collection from the authenticity of the content, and can handle complex scenarios such as real mobile phone shooting of AI images, real screenshots of AI images, and real shooting of AI posters for printing. 5. This invention can be deployed on mini-programs, mobile devices, web pages, desktops, or servers. It can be used for local privacy protection detection, as well as for enterprise content review, image trust verification, and digital content risk identification. Attached Figure Description
[0021] Figure 1 is a flowchart of the overall method of the AI-generated content authenticity separation and recognition method for propagating distorted images according to the present invention; Figure 2 is a schematic diagram of the propagation state identification module of the present invention; Figure 3 is a flowchart of the branch detection process of the present invention; Figure 4 is a diagram of the multimodal dynamic fusion of the present invention; Figure 5 is a diagram showing the separation between the authenticity of the data collected and the authenticity of the content in this invention; Figure 6 shows the threshold determination and report output of the present invention. Detailed Implementation
[0022] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings and specific embodiments. The embodiments are used to illustrate the present invention, and not to limit the scope of protection of the present invention. Where there is no conflict, the technical features in the embodiments of the present invention can be combined with each other. Example 1: Overall Identification Process
[0023] As shown in Figure 1, in this embodiment, the system first acquires the image to be detected and reads the image data. The image to be detected can be a platform-transferred image, a chat screenshot, a webpage screenshot, a screenshot taken by a mobile phone, a printed image taken by a mobile phone, or an image that has been cropped or compressed. The image data read by the system can include image pixel data, file format, file size, resolution, aspect ratio, compression parameters, file structure information, readable metadata, image boundary regions, local texture, frequency domain information, and noise residual information. Subsequently, the system extracts propagation state features and determines whether a propagation state has been identified. If a propagation state has been identified, the corresponding detection branch is invoked to determine the effective content area, extract content authenticity features and collection authenticity features, and perform dynamic fusion to calculate the risk result, and finally output a detection report. If no clear propagation state has been identified, the system outputs an uncertain propagation state result and generates a low-confidence detection report. This low-confidence detection report may include information such as the reason for not being able to confirm the result, missing evidence items, and a suggestion for manual review. Example 2: Propagation State Recognition Module
[0024] As shown in Figure 2, the propagation state recognition module is used to determine the main propagation processes experienced by the image to be detected. Its inputs include basic image information, compression features, boundary features, texture features, noise features, and file structure features, and its outputs are the propagation state category and the propagation state confidence score. In one embodiment, the propagation state recognition module calculates the score for each propagation state as follows: T1 = a1×Fcompress + a2×Fresize + a3×Fjpeg + a4×Fexif T2 = b1×FscreenRatio + b2×FuiBorder + b3×FtextGrid + b4×FnoiseUniform T3 = c1×Fmoire + c2×FpixelGrid + c3×Freflection + c4×Fperspective T4 = d1×FpaperTexture + d2×FprintDot + d3×FinkSpread + d4×FpaperEdge Wherein, T1 represents the platform-transferred image status score, T2 represents the screenshot image status score, T3 represents the screen-retrieved image status score, and T4 represents the printed image status score; Fcompress represents compression trace features, Fresize represents resampling or scaling features, Fjpeg represents JPEG quantization table or double compression features, and Fexif represents missing or residual metadata status; FscreenRatio represents screen ratio features, FuiBorder represents UI boundary features, FtextGrid represents text edge rasterization features, and FnoiseUniform represents overall image noise consistency features; Fmoire represents moiré features, FpixelGrid represents screen pixel raster features, Freflection represents reflective area features, and Fperspective represents perspective distortion features; FpaperTexture represents paper texture features, FprintDot represents printed halftone features, FinkSpread represents ink spread features, and FpaperEdge represents paper edge features. In a specific embodiment, the above weights can be set as follows: a1=0.30, a2=0.25, a3=0.30, a4=0.15; b1=0.25, b2=0.30, b3=0.25, b4=0.20; c1=0.30, c2=0.25, c3=0.20, c4=0.25; d1=0.30, d2=0.30, d3=0.20, d4=0.20.
[0025] When the score of a certain propagation state is higher than a preset threshold, the system identifies it as the corresponding propagation state. The preset threshold can be set to 0.60. When the scores of two or more propagation states all exceed 0.60, or the difference between the scores of two propagation states is less than 0.10, the system can identify the image to be detected as a mixed propagation distortion image state. The above parameters are only one optional embodiment. In actual deployment, the weights and thresholds can be adjusted according to the terminal computing power, sample distribution, platform compression rules, and detection scenario. Example 3: Branch Detection Process
[0026] As shown in Figure 3, the system calls the corresponding detection branch based on the propagation state identification result.
[0027] When the propagation state is a screenshot image, the system calls the screenshot detection branch to extract the valid content area within the screenshot. The valid content area can be a chat image area, a webpage main image area, a platform content card area, a poster area, or an image main body area.
[0028] When the propagation state is a platform-transferred image state, the system calls the platform-transferred image detection branch to determine the compression impairment level. The compression impairment level can be determined based on JPEG double compression traces, quantization table features, resolution scaling features, chroma sampling changes, and local high-frequency detail loss.
[0029] When the propagation state is screen capture image state, the system calls the screen capture detection branch to separate the screen content area and the external environment area. The screen content area is used to determine the authenticity of the content, while the external environment area and screen capture imaging features are used to determine the authenticity of the data acquisition.
[0030] When the propagation state is "printed and photographed image," the system invokes the print and photographed image detection branch to separate the paper content area and the paper acquisition area. The paper content area is used to determine the authenticity of the content, while the paper acquisition area and print imaging features are used to determine the authenticity of the acquisition.
[0031] When the propagation state is a mixed propagation distortion image state, the system can call two or more corresponding detection branches and perform weighted fusion of the branch detection results based on the scores of each propagation state. For example, for an image that has been screenshotted and then transferred from the platform, the system can simultaneously call the screenshot detection branch and the platform-transferred image detection branch. Example 4: Screenshot Detection Module
[0032] The screenshot detection module is used to process images such as mobile phone screenshots, web page screenshots, chat screenshots, and platform page screenshots.
[0033] The screenshot detection module detects the following features: status bar area, such as the area containing time, battery level, signal strength, and carrier identification; navigation bar area, such as the back button, home button, and bottom gesture bar; UI component boundaries, such as buttons, cards, input boxes, icons, and list dividers; screen ratio features, such as the screen ratios of common mobile phones, tablets, or computers; overall image noise consistency, such as the lack of natural camera noise in screenshot areas; text edge rasterization features, such as text edges displaying screen display or anti-aliasing features; and local content areas, such as images, posters, product images, people images, or document images contained within the screenshot. When the system determines that the image to be detected is a screenshot, it does not directly use the entire screenshot as the object for AI-generated content recognition, but first extracts the effective content area within the screenshot.
[0034] In one embodiment, the fusion weights in the screenshot image state can be set as follows: OCR text anomaly feature weight is 0.25; semantic consistency feature weight is 0.25; effective content region lightweight visual model output weight is 0.30; frequency domain / noise residual feature weight is 0.10; and screenshot state confidence correction weight is 0.10.
[0035] When the OCR text anomalies, text-image semantic mismatches, and visual model outputs for effective content areas all exceed preset thresholds, the system will judge the content within the screenshot as having a high risk of AI generation; when the screenshot status is clear but the effective content area is small, severely obscured, or lacks clarity, the system will lower the confidence level of the final conclusion. Example 5: Platform Transfer Image Detection Module
[0036] The platform-transferred image detection module is used to process images that have been compressed, transferred, or re-encoded from social media platforms, chat tools, web pages, short video platforms, e-commerce platforms, etc.
[0037] The platform-transferred image detection module detects the following features: missing EXIF, XMP, or content credentials information; JPEG double compression traces; similarity between the JPEG quantization table and common platform compression parameters; resolution scaled to a common platform size; chroma sampling changes; enhanced compression block boundaries; loss of local high-frequency details; mismatch between image file size and resolution; and excessive smoothing of local textures.
[0038] The system calculates the compression impairment level based on the above characteristics. A higher compression impairment level indicates a more severe loss of detail in the original image, and the confidence level of the model output needs to be lowered accordingly. For images transferred from the platform, the system reduces the weight of metadata features and increases the weight of frequency domain anomaly features, noise residual features, compression trace features, and the output results of the lightweight vision model.
[0039] In one embodiment, the fusion weights in the platform-transferred image state can be set as follows: frequency domain anomaly feature weight 0.25; noise residual feature weight 0.20; compression trace feature weight 0.15; lightweight visual model output weight 0.30; semantic consistency feature weight 0.10. When the compression impairment level is high, the confidence level of the lightweight visual model output is multiplied by a correction factor of 0.80 to 0.90; when the image compression is severe and effective details are insufficient, the system outputs an uncertainty statement of "insufficient evidence" or "cannot be confirmed". Example 6: Screen Reproduction Detection Module
[0040] The screen capture detection module is used to handle scenarios where users use their mobile phones or cameras to capture images, web pages, chat pages, posters, or AI-generated images on the screen.
[0041] The screen capture detection module detects the following features: moiré patterns; screen pixel grid; screen border; local reflective areas; uneven screen brightness; screen refresh stripes; perspective distortion during shooting; noise difference between the screen area and the external environment area; and camera noise superimposed on the screen display content.
[0042] When the system determines that the image to be detected is a screen capture image, the system separates the screen content area from the external environment area. The external environment area, screen reflections, perspective distortion, and actual camera noise are used to determine the authenticity of the image; the people, objects, text, scenes, and textures in the screen content area are used to determine the authenticity of the content.
[0043] In one embodiment, the fusion weights in the screen capture image state can be set as follows: screen capture imaging feature weight is 0.20; lightweight visual model output weight for screen content area is 0.35; OCR text anomaly feature weight is 0.15; frequency domain anomaly feature weight is 0.15; and semantic consistency feature weight is 0.15.
[0044] If the area outside the screen displays real camera noise, and the content area on the screen has a high risk of being AI-generated content, the system outputs a judgment result of "real capture + suspected AI-generated content". This result indicates that the image capture behavior is real, but the main content in the image has the risk of being AI-generated. Example 7: Printed Copy Detection Module
[0045] The print copy detection module is used to handle scenarios where users photograph printed images, posters, certificates, reports, contracts, menus, brochures, etc.
[0046] The print and copy detection module detects the following features: paper texture; paper edges; CMYK halftone dots; print dots; ink diffusion; paper reflection; creases; shadows; and perspective distortion.
[0047] The system extracts content areas from the paper surface based on paper edge, paper texture, and perspective distortion features. Paper texture, creases, reflections, and perspective distortion are used to determine the authenticity of the data; spatial visual anomalies, OCR text anomalies, frequency domain anomalies, and semantic consistency features of the content area are used to determine the authenticity of the content.
[0048] In one embodiment, the fusion weights in the printed and reproduced image state can be set as follows: the weight of the printed imaging features is 0.20; the weight of the lightweight visual model output of the paper content area is 0.30; the weight of the OCR text anomaly features is 0.20; the weight of the frequency domain / noise residual features is 0.15; and the weight of the semantic consistency features is 0.15.
[0049] If the paper texture, paper edges, and perspective distortion indicate that the image is a real photograph of paper material, and there are abnormalities in the structure of people, text, or semantic inconsistencies between the image and text in the content area of the paper, then the system will output the judgment result of "real collection + suspected AI-generated content". Example 8: OCR Text Anomaly Detection Module
[0050] The OCR text anomaly detection module is used to identify text anomalies in images and uses these anomalies as one of the important pieces of evidence for AI-generated content recognition. This module is particularly suitable for images with text content, such as posters, menus, product packaging, certificates, chat screenshots, webpage screenshots, instruction manuals, contracts, and promotional images.
[0051] The OCR text anomaly detection module includes the following steps: performing OCR text detection on the image or content region to be detected, obtaining text boxes, character recognition results, and OCR confidence scores; analyzing whether there are anomalies in the character structure, including half-characters, broken strokes, overlapping characters, partial garbled characters, and incomplete character shapes; analyzing whether there are typos, repeated characters, semantic incoherence, or context mismatches in the text content; analyzing whether the size, direction, perspective, and edges of characters within the same text line are consistent; analyzing whether the text content in the image matches the subject matter and semantics of the scene; and generating OCR text anomaly scores based on OCR confidence score fluctuations, text structure anomalies, and semantic anomalies.
[0052] In one embodiment, the OCR text anomaly score can be calculated using the following formula: O = 0.30 × Ostruct + 0.25 × Osemantic + 0.20 × Oconf + 0.15 × Oedge + 0.10 × Olayout Where O represents OCR text anomaly score, Ostruct represents character structure anomaly score, Osemantic represents text semantic anomaly score, Oconf represents OCR confidence fluctuation score, Oedge represents text edge anomaly score, and Olayout represents text layout anomaly score.
[0053] OCR text anomaly scores do not solely determine the risk of AI generation, but rather serve as one type of evidence of content authenticity in dynamic fusion. For screenshot images and printed / photographed images, the weight of OCR text anomaly features can be appropriately increased. Example 9: Frequency Domain / Noise Residual Detection Module
[0054] The frequency domain / noise residual detection module is used to extract anomalous features related to AI-generated content or propagation distortion from the underlying image signal.
[0055] This module can include the following processing: performing Fourier transform on the image or content region to obtain spectral distribution characteristics; performing DCT transform on the image to obtain compressed frequency domain energy distribution; extracting noise residual maps using wavelet decomposition or high-pass filtering; analyzing local noise consistency and local noise discontinuity characteristics; analyzing high-frequency texture repetition, periodic texture anomalies, and compression block effects; analyzing the difference between real camera noise and content region noise; and outputting frequency domain anomaly scores, noise residual anomaly scores, and compression impairment levels.
[0056] In this invention, frequency domain / noise residual detection is not the sole criterion for judgment, but rather one of the multimodal pieces of evidence. Its weight is dynamically adjusted by the propagation state recognition module. For example, for platform-transferred image states, compression marks and frequency domain features have higher weights; for screen-captured image states, frequency domain features need to be used in conjunction with moiré patterns and screen capture imaging features; for severely compressed images, the system reduces the confidence level of frequency domain anomaly judgment. Example 10: Lightweight Visual Model
[0057] The lightweight vision model module is used to identify AI-generated visual features in image content. This module can be deployed on user terminals, mini-programs, mobile devices, desktops, or servers.
[0058] In one implementation, the lightweight vision model employs MobileViT, DeiT-Tiny, EfficientNet-Lite, MobileNetV3, a lightweight Transformer model, a TensorFlow Lite model, or a TensorFlow.js model.
[0059] The input to the lightweight vision model is not limited to the original RGB image, but can include: RGB thumbnails; edge maps; frequency domain maps; noise residual maps; OCR text region maps; screen content region maps; and paper content region maps. The output of the lightweight vision model includes AI generation probability, suspected source category, local visual anomaly score, content anomaly region heatmap, semantic anomaly score, and model confidence.
[0060] To adapt to edge operation, lightweight vision models can reduce their size using methods such as model pruning, quantization, distillation, and feature dimensionality reduction. The model can be trained using a teacher model and a student model. The teacher model provides high-dimensional features or soft labels during training, while the student model performs inference on the edge. Example 11: Multimodal Dynamic Fusion Module
[0061] As shown in Figure 4, the multimodal dynamic fusion module is used to fuse the propagation state recognition results, branch detection results, multimodal features, and lightweight visual model output results into the final AI-generated content risk results.
[0062] The inputs to this module include: propagation state probability; platform transfer detection results; screenshot detection results; screen capture detection results; printed capture detection results; OCR text anomaly score; frequency domain anomaly score; noise residual anomaly score; lightweight visual model AI generation probability; semantic consistency anomaly score; and acquired imaging features.
[0063] In one implementation, the comprehensive risk score can be calculated as follows: R = Σ Wi × Fi × Ci, where R represents the risk score of AI-generated content, Fi represents the score of the i-th modality feature, Wi represents the weight dynamically adjusted according to the propagation state, and Ci represents the credibility correction coefficient of the modality feature in the current propagation state.
[0064] For example, when the propagation state is a screenshot image, the weights of OCR text anomaly features and semantic consistency features are relatively high; when the propagation state is a platform-transferred image, the weights of frequency domain anomaly features, noise residual features, and lightweight visual model output results are relatively high; when the propagation state is a screen-captured image, the weights of screen capture imaging features and screen content region model output results are relatively high; when the propagation state is a printed image, the weights of printed imaging features, paper content region model output results, and OCR text anomaly features are relatively high.
[0065] The system can also adjust the confidence level of the final result based on the image compression impairment level, screenshot confidence, re-photograph confidence, and model confidence, so as to avoid outputting overly certain conclusions when there is insufficient evidence. Example 12: Separation Module for Data Acquisition Authenticity and Content Authenticity
[0066] As shown in Figure 5, the authenticity separation module is used to determine the authenticity of the acquisition and the authenticity of the content of the image to be detected.
[0067] Authenticity of acquisition refers to whether the method of obtaining the image possesses genuine acquisition characteristics, such as genuine photography, genuine screenshots, transfers from genuine platforms, or photography on genuine paper. Authenticity of content refers to whether the main content carried in the image carries the risk of AI generation. These two are not the same object of judgment and should not be substituted for each other.
[0068] In one implementation, the system divides the image to be detected into regions related to the acquisition process and regions related to the main content. For screen-captured images, the external environment of the screen, screen reflections, perspective distortion, and camera noise are used to determine the authenticity of the acquisition, while the screen content region is used to determine the authenticity of the content. For printed images, paper texture, paper edges, creases, paper reflections, and perspective distortion are used to determine the authenticity of the acquisition, while the paper content region is used to determine the authenticity of the content. For screenshots, the consistency of the status bar, navigation bar, UI component boundaries, and screenshot noise is used to determine the acquisition method, while the valid content region within the screenshot is used to determine the authenticity of the content. For images transferred from a platform, compression traces, resampling traces, and file structure changes are used to determine the propagation status, while the main image region is used to determine the authenticity of the content.
[0069] In one implementation, the authenticity score A can be determined based on camera noise, screen capture imaging features, printed imaging features, screenshot boundary features, and platform transfer features; the content authenticity risk score C can be determined based on lightweight visual model output, OCR text anomaly score, semantic consistency anomaly score, frequency domain anomaly score, and noise residual anomaly score.
[0070] The system can output one of the following categories based on the combination of A and C: Category A: High authenticity of data collection, low risk of content authenticity; Category B: High authenticity of data collection, but content may be generated by AI; Category C: Images are screenshots or transferred from platforms, and content may be generated by AI; Category D: High risk of content, but the dissemination status is complex, and the confidence level of the conclusion is moderate; Category E: Insufficient evidence, unable to confirm.
[0071] For example, if the screen capture detection branch outputs high screen capture imaging features, and the lightweight visual model of the screen content area outputs a high AI generation probability, while the OCR text anomaly score is high, then the system outputs a judgment result of "real capture + suspected AI-generated content". This judgment indicates that the image capture behavior itself may be real, but the main content carried by the image has the risk of being generated by AI.
[0072] For example, if the print and photograph detection branch outputs high paper texture features, print dot features, and perspective distortion features, while the paper content area has abnormal human structure, garbled text, image-text semantic mismatch, or frequency domain abnormal features, then the system outputs the judgment result of "real photographed paper medium + suspected AI-generated content".
[0073] This module avoids directly equating real camera noise, real screen copy features, real screenshot interfaces, or real paper textures with the authenticity of the content, thereby improving the accuracy of judgment in scenarios of distorted dissemination. Example 13: Threshold Determination and Report Output
[0074] As shown in Figure 6, to further illustrate the possible implementations of the present invention, in one embodiment, exemplary thresholds can be set for propagation status identification, content authenticity risk assessment, and collection authenticity assessment. These thresholds are merely one implementation method and can be adjusted in practical applications based on sample size, image type, device performance, and business scenario.
[0075] The propagation status recognition module calculates the status scores T1 for the platform-transferred image, T2 for the screenshot image, T3 for the screen-retrieved image, and T4 for the printed image.
[0076] When T1≥0.60, the system determines that the image to be detected meets the platform's image transfer status; When T2≥0.60, the system determines that the image to be detected meets the screenshot image state; When T3≥0.60, the system determines that the image to be detected meets the screen-retrieved image state; When T4≥0.60, the system determines that the image to be detected meets the requirements for printing and reproducing images.
[0077] When two or more propagation state scores are both greater than or equal to 0.60, the system marks the image to be detected as a mixed propagation distortion image state, and determines the subsequent detection branch based on the two propagation states with higher scores. For example, when T2=0.68 and T1=0.63, the system can determine that the image is a mixed propagation distortion image that has been captured and transferred from the platform, and simultaneously call the capture detection branch and the platform-transferred image detection branch.
[0078] When all propagation state scores are below 0.60, but at least one propagation state score is in the range of 0.45 to 0.60, the system can output a propagation state uncertainty prompt and reduce the confidence level of the final detection report.
[0079] The authenticity score A indicates whether the acquisition method of the image to be detected has authentic acquisition characteristics. The authentic acquisition characteristics may include real camera noise, screen capture imaging characteristics, printed paper imaging characteristics, screenshot interface structure characteristics, and platform transfer compression characteristics.
[0080] In one embodiment, the authenticity score A ranges from 0 to 1: When A ≥ 0.70, the data collection authenticity is judged to be high; When 0.45 ≤ A < 0.70, the data collection authenticity is judged as medium. When A < 0.45, the authenticity of the data collection is judged to be low or cannot be confirmed.
[0081] The content authenticity risk score C indicates whether the main content of the image to be detected has the risk of being generated by AI. The main content can be an image area within a screenshot, a screen content area, a paper content area, or a main image area transferred from a platform.
[0082] In one embodiment, the content authenticity risk score C ranges from 0 to 1: When C ≥ 0.75, the content is judged to have a high risk of being generated by AI. When 0.55 ≤ C < 0.75, the content is judged to have a moderate risk of being generated by AI. When 0.35≤C<0.55, the content is judged to have a low to medium risk of being generated by AI, and it is recommended to combine it with manual review; When C < 0.35, it is determined that no obvious AI generation risk was found.
[0083] The system outputs the authenticity separation result based on the data collection authenticity score A and the content authenticity risk score C.
[0084] When A≥0.70 and C<0.35, the system outputs "High authenticity of data collection, no obvious risks of AI-generated content found"; When A≥0.70 and C≥0.55, the system outputs "High authenticity of data collection, but content may be generated by AI". When 0.45≤A<0.70 and C≥0.55, the system outputs "The collection method has certain authenticity characteristics, but the content has the risk of being generated by AI. Manual review is recommended." When A < 0.45 and C ≥ 0.75, the system outputs "The content generated by AI has a high risk, but the authenticity of the data collection is insufficient or cannot be confirmed." When A < 0.45 and C < 0.55, the system outputs "Insufficient evidence, unable to confirm".
[0085] By using the above method, the present invention separates the image acquisition method from the image content, avoiding the direct equivalent of real shooting, real screenshots, real platform transfers, or real printed reproductions to the authenticity of the content. Example 14: Test Report Sample
[0086] Example 1: The image to be detected is a picture taken of a computer screen by a mobile phone. The system identifies the propagation state as a screen-retrieved image, T3=0.82; the collection authenticity score A=0.78; and the content authenticity risk score C=0.71. System output: Propagation state is a screen-retrieved image; collection authenticity is high; the content has a medium to high risk of AI generation; the main evidence includes screen pixel raster, perspective distortion, OCR text anomalies in the screen content area, and output anomalies of the lightweight visual model; the detection conclusion is "suspected AI-generated content that was actually collected."
[0087] Example 2: The image to be detected is a product image saved on a social media platform. The system identifies the propagation state as a platform-transferred image, T1=0.74; the compression impairment level is medium; the acquisition authenticity score A=0.58; and the content authenticity risk score C=0.63. System output: The propagation state is a platform-transferred image; the acquisition method has certain authenticity characteristics; the content has the risk of AI generation; the main evidence includes double compression traces, local high-frequency detail loss, abnormal product edge fusion, and semantic consistency anomalies; the detection conclusion is "suspected AI-generated content in the platform-transferred image".
[0088] Example 3: The image to be detected is a poster in a screenshot taken from a mobile phone. The system identifies the propagation state as a screenshot image, T2=0.79; the collection authenticity score A=0.72; and the content authenticity risk score C=0.68. System output: Propagation state is screenshot image; screenshot collection authenticity is high; the poster content in the screenshot has the risk of being AI-generated; the main evidence includes status bar and UI boundaries, garbled text on the poster, text edge adhesion, and semantic mismatch between image and text; the detection conclusion is "suspected AI-generated content in a real screenshot".
[0089] Example 4: The image to be detected is a printed poster photographed with a mobile phone. The system identifies the propagation state as a printed copy image, T4=0.81; the collection authenticity score A=0.76; and the content authenticity risk score C=0.73. System output: Propagation state is a printed copy image; collection authenticity is high; there is a risk of AI generation in the paper content; the main evidence includes paper texture, printing dots, perspective distortion, abnormal human structure in the paper content, and OCR text anomalies; the detection conclusion is "suspected AI-generated content in a genuine photograph of paper media". Example 15: Construction of Experimental Samples
[0090] To verify the adaptability of this invention in scenarios involving propagation distortion, the following set of test cases can be constructed. This set of test cases is used to illustrate the verification method of this invention and does not limit the scope of protection of this invention.
[0091] Example 1: Platform Transferred Image Example. Several AI-generated images and several real-shot images were selected and saved via social media platforms, chat tools, or web pages to form platform transferred images. The detection output of the method of this invention in the case of metadata loss was compared with that of traditional metadata detection methods.
[0092] Example 2: Screenshot Image Examples. Select AI-generated images and real images containing people, pets, products, posters, Chinese text, and webpage content, and create screenshot examples by taking screenshots from a mobile phone or computer. Detect screenshot status recognition, effective content area extraction, OCR text anomalies, and content authenticity output.
[0093] Example 3: Screen capture image example. An AI-generated image and a real image are displayed on a mobile phone screen, computer screen, or tablet screen respectively. A screen capture example is then created using a mobile phone. The screen capture imaging features, screen content area separation, and the separation of "capture authenticity" and "content authenticity" output are evaluated.
[0094] Example 4: Printed Image Examples. After printing AI-generated posters, real photos, product images, certificate templates, and text images, these were then photographed using a mobile phone to create printed image examples. The results were examined for paper texture, print dots, extraction of content areas on the paper, and the accuracy of the output content.
[0095] Example 5: Hybrid propagation distortion image example. This example involves compressing an AI-generated image using the platform before taking a screenshot, or printing the screenshot and then photographing it. The output includes confidence corrections and uncertainty explanations for the detection system under multiple propagation states superimposed.
[0096] In a suggested test table, you can record examples with the following fields: Sample number, original content type, propagation status, whether it was generated by AI, whether metadata was retained, propagation status identification result, collection authenticity result, content authenticity result, main evidence items, final risk level, and manual review conclusion.
[0097] The above examples are used to support the implementation of the present invention. During formal testing, the sample size can be expanded according to the actual product scenario, and the accuracy rates of propagation status recognition, AI-generated content recognition, false positive rate of real images, false negative rate of AI images, and uncertainty output ratio can be statistically analyzed respectively. Optional implementation methods
[0098] This invention is not limited to the above-described embodiments, and the following alternative solutions may also be adopted: 1. The propagation state recognition module can employ a rule-based model, machine learning model, deep learning model, or a hybrid model; 2. The OCR text anomaly detection module can use a local OCR model, a server-side OCR model, or a third-party OCR model; 3. Lightweight visual models can be deployed on the edge or the server. 4. Frequency domain / noise residual features can be extracted using FFT, DCT, wavelet transform, high-pass filtering, residual networks, or noise estimation algorithms; 5. The dynamic fusion module can employ weighted rules, logistic regression, gradient boosting trees, lightweight neural networks, attention fusion models, or Bayesian fusion models; 6. The test report can be output as a page report, structured JSON, PDF report, API return results, or enterprise audit records; 7. This invention can be applied to mini-programs, mobile apps, web pages, browser plugins, enterprise image verification systems, copyright verification systems, content security systems, news image verification systems, and trusted digital content platforms.
Claims
1. A method for separating and recognizing the authenticity of AI-generated content in the propagation of distorted images, characterized in that, include: Acquire the image to be detected and read the image data of the image to be detected. The image data includes at least image pixel data, resolution information, compression features, boundary region features, and local texture features. Based on the image data, propagation state features are extracted, including at least one of platform transfer features, screenshot features, screen capture features, and print capture features. Specifically, platform transfer features include at least one of compression traces, resampling traces, JPEG quantization table features, chroma sampling changes, or local high-frequency detail loss. Screenshot features include at least one of status bar area, navigation bar area, UI component boundaries, screen ratio, overall image noise consistency, or text edge rasterization features. Screen capture features include at least one of moiré patterns, screen pixel raster, screen borders, reflective areas, screen refresh stripes, or perspective distortion. Print capture features include at least one of paper texture, paper edges, print dots, ink diffusion, paper reflection, or shooting perspective distortion. The propagation state of the image to be detected is determined based on the propagation state characteristics. The propagation state includes at least one of the following: platform-transferred image state, screenshot image state, screen-retrieved image state, and printed image state. According to the propagation state, the detection branch corresponding to the propagation state is called, and the effective content area for content judgment is determined in the corresponding detection branch; wherein, the effective content area within the screenshot is determined in the screenshot image state, the screen content area and the screen external environment area are separated in the screen copy image state, the paper content area and the paper collection area are separated in the printed copy image state, and the compression damage level is determined in the platform transferred image state. Extract content authenticity features from the effective content region. The content authenticity features include at least two of the following: spatial visual anomaly features, OCR text anomaly features, frequency domain anomaly features, noise residual features, and semantic consistency features. Extract acquisition authenticity features from regions or features related to the acquisition process in the image to be detected. The acquisition authenticity features include at least one of screen capture imaging features, printed imaging features, screenshot interface features, platform transfer compression features, and shooting noise features. Based on the propagation state, corresponding weights are assigned to the content authenticity features and the acquisition authenticity features, and the content authenticity risk result and acquisition authenticity judgment result are calculated based on the corresponding weights respectively; wherein, in the screenshot image state, the weights of OCR text anomaly features and semantic consistency features are increased; in the platform transferred image state, the weights of frequency domain anomaly features, noise residual features and compression impairment level are increased; in the screen copy image state, the weights of content authenticity features and screen capture imaging features of the screen content area are increased; and in the printed copy image state, the weights of content authenticity features and printed imaging features of the paper content area are increased. Based on the content authenticity risk result and the acquisition authenticity judgment result, an authenticity separation result is generated. The authenticity separation result is used to distinguish whether the acquisition method of the image to be detected has authentic acquisition characteristics, and whether the main content carried by the image to be detected has AI generation risk. The output includes a detection report containing the propagation status, the results of the authenticity assessment, the results of the content authenticity risk assessment, the risk level generated by AI, evidence items, and an explanation of uncertainty.
2. The method according to claim 1, characterized in that, Determining the propagation state of the image to be detected includes: Based on at least one of the following: EXIF missing status, JPEG double compression traces, quantization table features, resampling traces, resolution scaling features, and local high-frequency detail loss, determine whether the image to be detected meets the platform's image transfer status. Based on at least one of the following: status bar area, navigation bar area, UI component boundary, screen ratio, overall image noise consistency, and text edge rasterization features, determine whether the image to be detected meets the screenshot image status. Based on at least one of moiré patterns, screen pixel grid, screen border, reflective area, screen refresh stripes, and perspective distortion, determine whether the image to be detected meets the screen re-photograph image state. Based on at least one of the following factors—paper texture, paper edge, printing dots, ink diffusion, paper surface reflection, creases, and perspective distortion during photography—it is determined whether the image to be detected meets the requirements for a printed or reproduced image.
3. The method according to claim 1, characterized in that, When the propagation state is a screenshot image state, the screenshot detection branch is invoked, and the screenshot detection branch includes: Detect the status bar area, navigation bar area, virtual button area, UI component boundaries, text edge rasterization features, and overall image noise consistency features; Extract the valid content region within the screenshot from the image to be detected; Spatial visual anomaly features, OCR text anomaly features, and semantic consistency features are extracted from the effective content region; Based on the aforementioned spatial visual anomaly features, OCR text anomaly features, and semantic consistency features, the risk score of AI-generated content in the screenshot image state is obtained.
4. The method according to claim 1, characterized in that, When the propagation state is a platform-transferred image state, the platform-transferred image detection branch is invoked, and the platform-transferred image detection branch includes: The detection of EXIF missing state, JPEG double compression traces, quantization table features, resolution scaling features, chroma sampling changes, compression block boundaries, and local high-frequency detail loss in the image to be detected; The compression impairment level is determined based on the JPEG double compression traces, quantization table features, resolution scaling features, and local high-frequency detail loss. The weights of the frequency domain anomaly features, noise residual features, and lightweight visual model output results are adjusted based on the compression damage level.
5. The method according to claim 1, characterized in that, When the propagation state is a screen re-capture image state, the screen re-capture detection branch is invoked, and the screen re-capture detection branch includes: Detect moiré patterns, screen pixel grid, screen border, reflective areas, local brightness unevenness, screen refresh stripes, and perspective distortion features; Based on the screen border, perspective distortion features, and brightness distribution, the screen content area and the screen external environment area are separated. The external environment area of the screen and the screen capture imaging features are used to determine the authenticity of the data. The spatial visual anomaly features, OCR text anomaly features, frequency domain anomaly features, and semantic consistency features of the screen content area are used to determine the authenticity of the content.
6. The method according to claim 1, characterized in that, When the propagation state is the print re-image state, the print re-image detection branch is invoked, and the print re-image detection branch includes: Detects paper texture, paper edges, print dots, CMYK halftone dots, ink spread, paper reflection, creases, shadows, and perspective distortion features. Extract the content area of the paper surface based on paper edge, paper texture and perspective distortion features; Paper texture, paper reflection, creases, and perspective distortion features are used to determine the authenticity of the data. The spatial visual anomaly features, OCR text anomaly features, frequency domain anomaly features, and semantic consistency features of the paper content area are used to determine the authenticity of the content.
7. The method according to claim 1, characterized in that, The abnormal features of the OCR text include at least one of the following: abnormal character structure, garbled characters, typos, repeated characters, broken strokes, overlapping text edges, inconsistent text perspective, abnormal character scale within the same text line, abnormal fluctuations in OCR recognition confidence, and mismatch between text and image semantics.
8. The method according to claim 1, characterized in that, The input to the lightweight visual model includes at least two of the following: RGB image, edge map, frequency domain map, noise residual map, OCR text region map, screen content region map, and paper content region map; The lightweight vision model includes one of MobileViT, DeiT-Tiny, EfficientNet-Lite, MobileNetV3, a lightweight Transformer model, a TensorFlow Lite model, or a TensorFlow.js model; The lightweight visual model outputs at least one of the following: AI generation probability, suspected generation source category, content anomaly area heatmap, semantic anomaly score, local visual anomaly score, and model confidence score.
9. The method according to claim 1, characterized in that, The weights assigned to the content authenticity feature and the acquisition authenticity feature based on the propagation status include: When the propagation state is the platform-transferred image state, reduce the weight of metadata features and increase the weight of frequency domain anomaly features, noise residual features, compression trace features, and lightweight visual model output results. When the propagation state is a screenshot image state, increase the weight of the lightweight visual model output results of OCR text anomaly features, semantic consistency features, and effective content areas within the screenshot; When the propagation state is screen capture image state, increase the weight of screen capture imaging features, lightweight visual model output of screen content area and frequency domain anomaly features; When the propagation state is a printed copy image, increase the weights of the output of the lightweight visual model of the printed imaging features, the paper content area, and the OCR text anomaly features.
10. A system for separating and recognizing the authenticity of AI-generated content for propagating distorted images, characterized in that, include: The image acquisition module is used to acquire the image to be detected and read the image data of the image to be detected; The propagation state recognition module is used to determine the propagation state of the image to be detected based on at least two of the following: resolution ratio, compression marks, boundary regions, moiré patterns, paper texture features, noise residual features, and file structure features. The branch detection module is used to call the screenshot detection branch, the platform-transferred image detection branch, the screen capture detection branch, or the print capture detection branch according to the propagation status. The multimodal feature extraction module is used to extract at least two of the following: spatial visual anomaly features, frequency domain anomaly features, noise residual features, OCR text anomaly features, screen capture imaging features, printed imaging features, and semantic consistency features. A lightweight visual model module is used to output AI generation probability and model confidence based on the multimodal features; The multimodal dynamic fusion module is used to adjust the weights of each modality feature according to the propagation status and generate a corrected AI-generated content risk result. The authenticity separation module is used to separately determine the authenticity judgment result of the data collection and the authenticity judgment result of the content; The report generation module is used to output a detection report that includes the spread status, collection authenticity, content authenticity, AI-generated risk level, evidence items, and uncertainty explanation.