Counterfeit identification method and device based on local image understanding

By employing a forgery recognition method based on local image understanding, this method utilizes a convolutional neural network for forgery recognition and a multi-scale attention mechanism to generate heatmaps of abnormal regions. Combined with structured textual prompts and a target expert model, it solves the problems of easy omission of local forgery features and untraceable overall recognition in existing technologies, achieving highly accurate and interpretable forgery recognition.

CN120953778AActive Publication Date: 2025-11-14BEIJING HISIGN TECH

Patent Information

Application Number
CN202511485152.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-11-14
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing technologies tend to overlook local forgery features when identifying forged images, leading to decreased recognition accuracy and a lack of observability and interpretability in the decision-making process, thus reducing user experience.

Method used

A forgery detection method based on local image understanding is adopted. Image features are extracted by forgery detection convolutional neural network and multi-scale attention mechanism to generate anomaly region heat map. Structured text prompt information is generated by combining with preset prompt information format, and target expert model is activated to analyze forgery region. Multi-scale attention mechanism and deformable convolutional block are used to improve the accuracy of forgery region localization.

Benefits of technology

It significantly improves the accuracy and interpretability of counterfeit identification, accurately locates suspicious counterfeit areas and provides detailed semantic descriptions, enhancing users' trust in the identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953778A_ABST
    Figure CN120953778A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a counterfeit recognition method and device based on local image understanding, and the method comprises the steps: receiving a to-be-recognized image, extracting the features of a counterfeit region of the to-be-recognized image, generating an abnormal region thermodynamic diagram of a suspicious counterfeit region corresponding to the to-be-recognized image based on the features of the counterfeit region, and carrying out the recognition of the abnormal region thermodynamic diagram. Under the condition that the spatial distribution of the abnormal region thermodynamic diagram is reasonable, generating structured text prompt information according to a preset prompt information format based on the abnormal region thermodynamic diagram and the to-be-recognized image, and encoding the text prompt information and the to-be-recognized image to obtain text embedding features, and activating a target expert model corresponding to the suspicious counterfeit region based on the text embedding feature, and processing the corresponding suspicious counterfeit region through the target expert model to obtain a counterfeit confidence coefficient and a counterfeit recognition result corresponding to each suspicious counterfeit region. According to the method, the defect that local counterfeit information is easy to omit is effectively overcome, and the accuracy of counterfeit identification is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and specifically to a forgery recognition method and apparatus based on local image understanding. Background Technology

[0002] With the development of digital media technology, deepfake technology has become increasingly realistic, which has brought about many security risks, such as the spread of false information, privacy violations, and identity theft. Deepfake technology can generate highly realistic fake face images and videos through deep learning algorithms, making it difficult for ordinary users to distinguish between the real and fake.

[0003] In existing technologies, methods for identifying forgery techniques are easily limited to specific datasets and tend to overlook or lack sensitivity to local forgery features, leading to a decrease in the accuracy of forgery identification. In addition, in the process of identifying forgery through overall identification or large models, the decision-making process lacks observability and interpretability, making it difficult for users to understand the judgment basis, reducing their trust in the forgery identification results, and lowering the user experience. Summary of the Invention

[0004] To address the problems in the prior art, this application provides a forgery identification method and apparatus based on local image understanding, which can effectively solve the shortcomings of traditional technologies such as easy omission of local forgery information, untraceable overall identification errors, and low accuracy of forgery identification, significantly improving the accuracy of forgery identification and enhancing the interpretability of forgery identification results.

[0005] To solve at least one of the above problems, this application provides the following technical solution: In a first aspect, this application provides a forgery recognition method based on local image understanding, comprising: The system receives an image to be identified, extracts the features of the forged regions of the image through a forgery recognition convolutional neural network and a multi-scale attention mechanism, and generates an abnormal region heatmap of the suspicious forged regions corresponding to the image to be identified based on the forgery region features. The tail of the forgery recognition convolutional neural network includes deformable convolutional blocks with a preset number of layers. When the spatial distribution of the heatmap of the abnormal region is reasonable, a structured text prompt information is generated based on the heatmap of the abnormal region and the image to be identified in accordance with the preset prompt information format. The text prompt information and the image to be identified are encoded to obtain text embedding features. The text prompt information is used to describe the abnormal features of at least one suspected forgery region and the abnormal confidence level corresponding to the abnormal features. The target expert model corresponding to the suspected forgery region is activated based on the text embedding feature. The target expert model processes the corresponding suspected forgery region to obtain the forgery confidence score of each suspected forgery region. The forgery recognition result of the image to be identified is determined based on the forgery confidence score.

[0006] Furthermore, it also includes: extracting multi-level features of the image to be identified based on a convolutional neural network for forgery detection. The features to be identified include low-level features and high-level features. The low-level features are used to obtain abnormal pixels in the image to be identified, and the high-level features are used to obtain semantic information of the image to be identified. By using a multi-scale attention mechanism, the features to be identified in the convolutional neural network for forgery detection are fused according to the feature pyramid network structure at a preset number of layers to obtain the features of the forgery region.

[0007] Furthermore, it also includes: performing pixel-level probability mapping on the features of the forged region, mapping the probability value of each pixel in the image to be identified belonging to the forged region to a heatmap; By coloring the heatmap using a pre-defined mapping relationship between color and probability value, a heatmap of abnormal areas is obtained.

[0008] Furthermore, the preset prompt message format is a triple structure, including region type, anomaly description, and anomaly confidence level, and also includes: Based on the heat map of abnormal areas, determine the area type and the anomaly confidence level of the suspected forgery area; By generating a corresponding abnormal description for each suspected forgery area through a preset abnormal description vocabulary, the area type, abnormal description and abnormal confidence level are combined according to a preset prompt information format to obtain text prompt information.

[0009] Furthermore, it also includes: receiving the forgery confidence scores output by each target expert model, and determining the identification decision logic based on the target expert model and the forgery confidence scores, wherein the identification decision logic is represented by a directed acyclic graph; The forgery identification result is obtained by processing the forgery confidence level based on the identification decision logic.

[0010] Furthermore, after generating an anomaly region heatmap of the suspected forged region corresponding to the image to be identified based on the forged region features, the process also includes: When the spatial distribution of the heatmap in the abnormal area is unreasonable, the image to be identified is determined to be an adversarial attack sample; Alarm information is generated based on adversarial attack samples and heatmaps of abnormal areas, and then sent to the terminal device.

[0011] Furthermore, it also includes: the spatial distribution of the heat map of the abnormal area is reasonable and simultaneously meets the following conditions: The distribution of high-response areas in the abnormal area heatmap matches the distribution of high-response areas corresponding to the preset forgery pattern. The target expert model to be activated is the same as the target expert model corresponding to the suspicious forged area described in the text prompt; Frequency domain analysis determined that there were no abnormal high-frequency components in the image to be identified.

[0012] Secondly, this application provides a forgery recognition device based on local image understanding, comprising: The first processing module is used to receive the image to be identified, extract the forgery region features of the image to be identified through a forgery recognition convolutional neural network and a multi-scale attention mechanism, generate an abnormal region heat map of the suspected forgery region corresponding to the image to be identified based on the forgery region features, and the tail of the forgery recognition convolutional neural network includes deformable convolutional blocks with a preset number of layers. The second processing module is used to generate structured text prompt information based on the heat map of the abnormal area and the image to be identified in a preset prompt information format, provided that the spatial distribution of the heat map of the abnormal area is reasonable. The text prompt information and the image to be identified are encoded to obtain text embedding features. The text prompt information is used to describe the abnormal features of at least one suspected forgery area and the abnormal confidence level corresponding to the abnormal features. The third processing module is used to activate the target expert model corresponding to the suspected forgery region based on the text embedding features, process the corresponding suspected forgery region through the target expert model, obtain the forgery confidence level corresponding to each suspected forgery region, and determine the forgery recognition result of the image to be recognized based on the forgery confidence level.

[0013] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the forgery recognition method based on local image understanding.

[0014] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the forgery recognition method based on local image understanding.

[0015] Fifthly, this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the forgery recognition method based on local image understanding.

[0016] As can be seen from the above technical solution, this application provides a forgery recognition method and apparatus based on local image understanding. It innovatively receives an image to be recognized, extracts forgery region features from the image using a forgery recognition convolutional neural network and a multi-scale attention mechanism, and generates an anomaly region heatmap of suspected forgery regions corresponding to the image to be recognized based on the forgery region features. The tail of the forgery recognition convolutional neural network includes a preset number of deformable convolutions. Given a reasonable spatial distribution of the anomaly region heatmap, structured text prompts are generated according to a preset prompt information format based on the anomaly region heatmap and the image to be recognized. The text prompts and the image to be recognized are encoded to obtain text embedding features. The text prompts describe the anomaly features of at least one suspected forgery region and the corresponding anomaly confidence level. A target expert model corresponding to the suspected forgery region is activated based on the text embedding features. The target expert model processes the corresponding suspected forgery regions to obtain the forgery confidence level for each suspected forgery region. The forgery recognition result of the image to be recognized is determined based on the forgery confidence level. This method can identify potential forgery regions in the image to be recognized, enhances sensitivity to local forgery information, enriches the semantic information corresponding to suspected forgery regions through text prompts, and improves the accuracy of forgery recognition. This method effectively addresses the shortcomings of traditional techniques, such as the easy omission of local forged information, the untraceability of overall identification errors, and the low accuracy of forgery identification. It significantly improves the accuracy of forgery identification and enhances the interpretability of forgery identification results. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the forgery identification method based on local image understanding in an embodiment of this application; Figure 2 This is a structural diagram of a forgery recognition device based on local image understanding in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of the electronic device in the embodiments of this application.

[0019] Figure label: Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] The acquisition, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.

[0022] In the existing technology, current forgery detection algorithms can detect forgeries through prompt text. However, prompt text is simple and has a fixed structure, which does not play a supervisory role, or it may miss local forgery information in the image to be identified. There are also problems with the overall identification of the image to be identified and the use of the preset model, such as the inability to trace and observe identification errors.

[0023] In view of the problems existing in the prior art, this application provides a forgery recognition method and apparatus based on local image understanding. By using an expert model framework to perform forgery recognition on the local / global aspects of the image to be recognized, it can focus on potential forgery areas in the image to be recognized, generate targeted analysis prompt text, guide the expert model framework to activate relevant expert networks, and achieve accurate location and description of forgery traces through multi-granular visual analysis and semantic reasoning. It mainly focuses on face analysis, activates different target expert models through dynamically generated prompt text, and feeds back the predicted time cost based on the number of activated target expert models, making the forgery recognition process interpretable.

[0024] To effectively address the shortcomings of traditional technologies, such as the easy omission of local forgery information, the untraceability of overall identification errors, and the low accuracy of forgery identification, and to significantly improve the accuracy of forgery identification and enhance the interpretability of the forgery identification results, this application provides an embodiment of a forgery identification method based on local image understanding. See [link to embodiment]. Figure 1 The forgery recognition method based on local image understanding specifically includes the following: Step S101: Receive the image to be identified, extract the forgery region features of the image to be identified through a forgery recognition convolutional neural network and a multi-scale attention mechanism, and generate an abnormal region heatmap of the suspected forgery region corresponding to the image to be identified based on the forgery region features.

[0025] The tail of the forgery detection convolutional neural network includes deformable convolutional blocks with a preset number of layers.

[0026] Optionally, this embodiment receives an image to be identified and extracts forgery region features from the image to be identified by combining a forgery recognition convolutional neural network with a multi-scale attention mechanism. The forgery region features can represent macroscopic forgery traces in the image to be identified, and can also represent microscopic anomalies. That is, abnormal pixels and semantic information in the image to be identified are obtained through low-level features and high-level features. Abnormal pixels are pixels in the image to be identified that have inconsistent textures and / or inconsistent gradients.

[0027] The tail of the forgery detection convolutional neural network includes a preset number of deformable convolutional blocks. For example, the tail of the fifth and sixth layers of the forgery detection convolutional neural network includes three deformable convolutional blocks to enhance the ability to locate irregular and suspicious forgery regions, making the location of suspicious forgery regions more accurate. The deformable convolutional blocks can be Deformable Convolutional Blocks.

[0028] Additionally, an anomaly heatmap of suspected forged regions in the image to be identified is generated based on the characteristics of the forged regions. Specifically, the anomaly heatmap is a pixel-level probability map used to display the probability that a pixel is a suspected forged region.

[0029] One approach is to insert a local binary pattern branch into the third layer of the forgery detection convolutional neural network to enhance the detection of texture inconsistencies in the image to be identified during the training of the forgery detection convolutional neural network.

[0030] This embodiment achieves accurate localization of suspicious forgery regions through a multi-scale attention mechanism and deformable convolutional blocks, thereby improving the accuracy of forgery identification.

[0031] Step S102: When the spatial distribution of the heat map of the abnormal area is reasonable, a structured text prompt information is generated based on the heat map of the abnormal area and the image to be identified in accordance with the preset prompt information format. The text prompt information and the image to be identified are encoded to obtain text embedding features.

[0032] The text prompt information is used to describe the abnormal features of at least one suspicious forged area and the abnormal confidence level corresponding to the abnormal features.

[0033] Optionally, in this embodiment, if the spatial distribution of the abnormal area heat map is reasonable, a structured text prompt message is generated according to the abnormal heat map and the original image to be identified, in accordance with a preset prompt message format. That is, the forgery clues in the abnormal heat map are converted into text form that is easy to understand and process, and the specific type of the suspected forgery area is determined in combination with the image to be identified to obtain the text prompt message.

[0034] The text prompt information is used to describe the abnormal features of at least one suspicious forged area and the abnormal confidence level corresponding to the abnormal features. For example, "unnatural hairline, abnormal confidence level 0.7", where "hairline" represents a suspicious forged area, "unnatural" represents an abnormal feature, and "abnormal confidence level 0.7" represents the abnormal confidence level corresponding to the abnormal feature. The text prompt information can also be expressed as "unnatural hairline, 0.7".

[0035] In addition, the text prompts and the image to be recognized are encoded separately and then fused into a single feature to obtain the text embedding feature. The text embedding feature can provide rich contextual information to activate the target expert model.

[0036] Specifically, the text prompt information can be encoded using the CLIP text encoder, and the image to be recognized can be encoded using the CLIP visual encoder. The encoding corresponding to the text prompt information and the encoding corresponding to the image to be recognized are then fused to obtain the text embedding features.

[0037] This embodiment realizes the generation of text prompt information based on the heat map of abnormal areas, and then obtains text embedding features based on the text prompt information and the image to be identified, which can provide a detailed and accurate semantic description of the suspected forged area and the image to be identified.

[0038] Step S103: Activate the target expert model corresponding to the suspected forgery region based on the text embedding features, process the corresponding suspected forgery region through the target expert model, obtain the forgery confidence of each suspected forgery region, and determine the forgery recognition result of the image to be recognized based on the forgery confidence.

[0039] Optionally, in this embodiment, based on text embedding features, a target expert model corresponding to the suspected forgery region is activated. The target expert model is used to focus on in-depth analysis of a specific forgery region. For example, the hairline region expert model focuses on identifying forgery features in the hairline region, and the lip region expert model focuses on identifying forgery features in the lip region.

[0040] Furthermore, the suspected forgery regions are processed using a target expert model to obtain the forgery confidence score for each region. The forgery recognition result of the image to be identified is then determined based on this confidence score. The determination of the forgery confidence score involves a comprehensive analysis of the output of the target expert model. The decision logic can be represented using a directed acyclic graph (DAG), and the forgery recognition result is determined by integrating the outputs of various expert models according to preset decision rules.

[0041] The forgery identification result includes whether the image to be identified is a forgery or not. In the case that the image to be identified is a forgery, the forgery identification result may also include the forgery area in the image to be identified.

[0042] Furthermore, a gating mechanism can be used to determine the target expert model that needs to be activated in order to process suspicious forged areas.

[0043] Furthermore, the target expert model can be deployed on multiple computing sub-nodes through a distributed computing architecture to enable parallel processing of the target expert model. After obtaining the forgery confidence scores corresponding to each suspected forgery area, the forgery identification results are summarized, which improves the processing speed and efficiency of forgery identification.

[0044] This embodiment activates the target expert model corresponding to the suspected forgery region through text embedding features, so that the target expert model processes the suspected forgery region in the image to be identified, determines the forgery identification result, focuses on the local image to be identified, prevents the omission of details, and improves the accuracy of forgery identification.

[0045] This embodiment achieves accurate localization of suspicious forgery regions through multi-scale attention mechanisms and deformable convolutional blocks, improving the accuracy of forgery identification. The structured text prompts can describe the suspicious forgery regions in detail, allowing users to intuitively understand the basis for the forgery identification results. Furthermore, the target expert model can be used to specifically identify local areas of the image to be identified, preventing omissions and improving the accuracy of forgery identification.

[0046] In some embodiments, forgery region features of the image to be identified are extracted using a forgery detection convolutional neural network and a multi-scale attention mechanism, including: Based on the forgery detection convolutional neural network, multi-level features to be identified in the image to be identified are extracted. The features to be identified include low-level features and high-level features. The low-level features are used to obtain abnormal pixels in the image to be identified, and the high-level features are used to obtain semantic information of the image to be identified. By using a multi-scale attention mechanism, the features to be identified in the convolutional neural network for forgery detection are fused according to the feature pyramid network structure at a preset number of layers to obtain the features of the forgery region.

[0047] Optionally, this embodiment uses a forgery detection convolutional neural network to process the image to be identified, extracting multi-level features to be identified. These features include, but are not limited to, low-level and high-level features. Low-level features are mainly used to capture abnormal pixels in the image to be identified, such as discontinuous edges or abnormal textures. High-level features are mainly used to obtain semantic information of the image to be identified, such as the layout of faces and irrelevant relative positions. By extracting features in a hierarchical manner, multi-level features are obtained, enabling a comprehensive understanding and analysis of suspicious forgery areas in the image to be identified from different perspectives.

[0048] In addition, by using a multi-scale attention mechanism, multi-level features to be identified are fused according to the Feature Pyramid Network (FPN). The Feature Pyramid Network structure can comprehensively consider the features to be identified at multiple scales, thereby more accurately locating suspicious forgery areas. This allows for full utilization of features to be identified at different levels and also enhances the ability to identify complex forgery methods.

[0049] For example, a fake image may have subtle pixel-level changes in local areas, while also having semantic inconsistencies in its overall structure. The fake region features fused through a multi-scale attention mechanism can simultaneously represent anomalies at different levels.

[0050] In this process, the features to be identified from the third layer output of the convolutional neural network for forgery detection can be fused to the features to be identified from the seventh layer output through a multi-scale attention mechanism. During the fusion process, edge details are preserved for low-level features to be identified, while semantic abnormal reactions are captured for high-level features to be identified, thereby obtaining forgery region features.

[0051] Furthermore, in the process of extracting low-level features to be identified, Local Binary Pattern (LBP) features can be introduced. LBP is a texture descriptor that can extract local texture information of the image to be identified. By introducing LBP features into the low layer of the forgery detection convolutional neural network, texture anomalies in the image can be captured more sensitively.

[0052] For example, forged images may exhibit texture discontinuities or unnatural transitions in local areas. LBP features can effectively describe anomalous textures, thereby improving the ability to detect local forgery traces.

[0053] Furthermore, a dynamic adjustment mechanism can be added to the multi-scale attention mechanism, thereby enabling the attention weights to be dynamically adjusted according to the content and features of different images to be recognized.

[0054] For example, when processing an image to be identified that mainly includes local forgeries, attention is increased to low-level features to be identified; when processing an image to be identified whose overall structure has been tampered with, attention is increased to high-level features to be identified. This improves adaptability and allows for optimization based on specific circumstances, thereby improving the efficiency and accuracy of forgery identification.

[0055] This embodiment extracts forged region features from the image to be identified through a forgery recognition convolutional neural network and a multi-scale attention mechanism. It can comprehensively understand and analyze forgery information in the image to be identified from different levels, and can more accurately locate suspicious forgery regions, thereby improving the accuracy of forgery recognition.

[0056] In some embodiments, generating an abnormal region heatmap of the suspected forged region corresponding to the image to be identified based on forged region features includes: Perform pixel-level probability mapping on the features of the forged region, mapping the probability value of each pixel in the image to be identified belonging to the forged region to a heatmap; By coloring the heatmap using a pre-defined mapping relationship between color and probability value, a heatmap of abnormal areas is obtained.

[0057] Optionally, this embodiment performs pixel-level probability mapping on the features of the forged region, that is, maps the probability value of each pixel in the image to be identified belonging to the suspected forged region to a heat map, so as to intuitively show the probability of each pixel in the image to be identified belonging to the suspected forged region.

[0058] For example, if a pixel has a high probability of belonging to a suspected forgery area, it will be assigned a high probability value in the heatmap to make it more visually prominent.

[0059] In addition, the heatmap is colored by a preset mapping relationship between color and probability value to obtain an abnormal area heatmap. The coloring process enhances the visual effect of the heatmap, making the difference between suspicious and non-suspicious forgery areas more obvious, which makes it easier for users to quickly identify and analyze.

[0060] For example, pixels with high probability values ​​can be colored red to indicate that they have a high probability of belonging to a suspected forgery area, while pixels with low probability values ​​can be colored blue to indicate that they have a low probability of belonging to a suspected forgery area, allowing users to intuitively see the suspected forgery areas in the image to be identified.

[0061] The anomaly region heatmap is the output of the forgery detection convolutional neural network processing the image to be identified. The process of obtaining the forgery detection convolutional neural network includes: Training dataset construction: The training dataset mainly includes two types of samples: real samples and fake human samples. Real samples are derived from sources including but not limited to actual photographed portraits or videos, and public datasets (such as CelebA, FFHQ, etc.). Fake samples can be generated through various algorithms and APIs, including but not limited to local fake samples and global fake samples.

[0062] Local forged samples can be generated using algorithms such as mask fusion, T-shaped glasses fusion, mouth fusion, and nose fusion, and corresponding binary masks can be generated to annotate the local forged regions. Global forged samples can be generated using algorithms such as face swapping and face replay, and the corresponding mask regions can be extracted using pixel difference comparison algorithms. In addition, publicly available forged datasets, such as FaceForensics++, can be collected to enhance the data diversity of the training dataset.

[0063] Design of the hybrid loss function: The hybrid loss function is used to supervise the training of the forgery detection convolutional neural network. The hybrid loss function consists of three parts: focus loss, multi-scale result similarity loss, and gradient sensitive loss.

[0064] Focus loss is used to alleviate the problem of uneven distribution of positive and negative samples in classification tasks, and to improve the overall performance of convolutional neural networks for forgery detection. The expression for focus loss is: ; in, Used to represent the predicted probability of a real sample, that is, the confidence level of a convolutional neural network for forgery detection in predicting that the current sample belongs to a real sample; Used for class balancing weights, that is, adjusting the weight balance between positive and negative samples in the training data to prevent the convolutional neural network for forgery detection from being biased towards the majority class; γ represents the adjustment factor, which can be set to γ≥0 to control the weight of easy and difficult samples. That is, when the value of γ is large, the convolutional neural network for forgery detection pays more attention to the training data that is difficult to classify.

[0065] Multi-scale result similarity loss (MS-SSIM Loss) is used to capture the differences in texture structure between fake and real regions, enabling the forgery detection convolutional neural network to better learn the subtle differences between fake and real samples. The expression for multi-scale result similarity loss is: L MS-SSIM =1−MS-SSIM(x,y); Where x represents a fake sample, y represents a real sample, and MS-SSIM stands for Multiscale Structural Similarity Measure, which measures the structural consistency of two training samples across multiple scales.

[0066] Gradient-sensitive loss is used to enhance the gradient difference in fake boundary regions. The expression for gradient-sensitive loss is: ; in, Used for region weight masks to emphasize the importance of forged regions. This indicates element-wise multiplication. Represents the gradient operator. This indicates a forged sample. Represents a real sample. This represents the L1 norm.

[0067] The expression for the hybrid loss function is: Where a, b, and c represent constant weight coefficients used to balance the influence of different loss terms. During the training of the convolutional neural network for forgery detection, the optimal model parameters are selected based on the performance on the validation set.

[0068] Furthermore, during the processing of dynamic image data, the heatmap of abnormal areas is updated in real time to reflect the instantaneous changes in forgery features in the image sequence to be identified. This enables continuous monitoring of forgery behavior in the video stream and timely detection of new forgery signs. For example, in a real-time surveillance video, an abnormal area heatmap can be generated and updated in real time, helping monitoring personnel to quickly identify and respond to forgery behavior.

[0069] This embodiment realizes the generation of anomaly region heatmaps based on the features of forged regions, which can intuitively show the forgery probability of each pixel in the image to be identified, and the anomaly region heatmap can reflect the forgery probability of each pixel, making the identification of suspicious forged regions more accurate.

[0070] In some embodiments, the preset prompt information format is a triple structure, including region type, anomaly description, and anomaly confidence level; Based on the heatmap of the abnormal area and the image to be identified, a structured text prompt message is generated according to a preset prompt message format, including: Based on the heat map of abnormal areas, determine the area type and the anomaly confidence level of the suspected forgery area; By generating a corresponding abnormal description for each suspected forgery area through a preset abnormal description vocabulary, the area type, abnormal description and abnormal confidence level are combined according to a preset prompt information format to obtain text prompt information.

[0071] Optionally, the preset prompt information format in this embodiment adopts a triple structure, including region type, anomaly description, and anomaly confidence. The structured prompt information format can accurately describe the characteristics of the suspicious forged region in the image to be identified.

[0072] By analyzing the probability distribution in the heatmap of abnormal regions, high-probability areas are identified and designated as abnormal regions. Based on these abnormal regions, their region type is determined within the image to be identified, such as hairline area, lip area, periorbital area, eyebrow area, nose area, hair area, or facial area. Simultaneously, based on the probability value corresponding to each abnormal region, an anomaly confidence level is determined, where the anomaly confidence level indicates the likelihood that the region has been forged.

[0073] In addition, by using a pre-defined anomaly description vocabulary, a corresponding anomaly description is generated for each suspected forgery area. The anomaly description vocabulary includes, but is not limited to, various possible forgery feature descriptions, such as “blurry”, “unnatural”, “inconsistent texture”, “jitter”, “incongruity”, etc.

[0074] In addition, based on the specific characteristics of the abnormal region, appropriate descriptive words are selected from the abnormal description vocabulary to generate the abnormal description corresponding to the abnormal region. The region type, abnormal description and abnormal confidence are combined according to the preset prompt information format to obtain a complete text prompt information.

[0075] For example, the generated text prompts could be "The hairline area is unnatural, with anomaly confidence level of 0.7; the lip area is blurred, with anomaly confidence level of 0.8." The structured text prompts are easy to understand and provide guidance for activating the expert model.

[0076] Furthermore, the process of converting visual information into text prompts in natural language can also be represented as follows: inputting an anomaly heatmap and the corresponding image to be identified, generating text prompts according to a preset prompt format, and inputting the anomaly heatmap, the image to be identified, and the text prompts into a multimodal large model to obtain the text prompts.

[0077] Furthermore, through a context-aware mechanism, more accurate and detailed anomaly descriptions can be generated based on the overall content and contextual information of the image to be identified.

[0078] For example, if a suspected forged area is located near the hairline area of ​​a face, and the texture of the suspected forged area is significantly inconsistent with the surrounding area, it can be described as "texture inconsistency". Combined with contextual information, it can be further described as "the texture of the hairline area is inconsistent with the texture of the surrounding hair, and there may be clipping traces". This makes the anomaly description more specific and accurate, and improves the reliability of forgery identification.

[0079] This embodiment enables the generation of structured text prompts through a preset prompt information format. It can accurately describe the region type, anomaly description, and anomaly confidence level of the suspicious forged region in the image to be identified, so as to clearly show the relevant information of the suspicious forged region to the user and provide guidance for activating the target expert model.

[0080] In some embodiments, determining the forgery identification result of the image to be identified based on the forgery confidence level includes: Receive the forgery confidence scores output by each target expert model, and determine the identification decision logic based on the target expert model and the forgery confidence scores. The identification decision logic is represented by a directed acyclic graph. The forgery identification result is obtained by processing the forgery confidence level based on the identification decision logic.

[0081] Optionally, this embodiment receives the forgery confidence scores output by each target expert model. Each target expert model focuses on the analysis of a specific suspected forgery region and outputs the forgery confidence score corresponding to the suspected forgery region. The forgery confidence score indicates the probability that the suspected forgery region has been forged.

[0082] For example, the forgery confidence score output by the target expert model for the hairline area is 0.9, indicating that the hairline area has a high probability of being forged, while the forgery confidence score output by the target expert model for the lip area is 0.4, indicating that the lip area has a low probability of being forged.

[0083] In addition, the identification decision logic is determined based on the target expert model and the forgery confidence. The identification decision logic can be represented by a directed acyclic graph (DAG). Different nodes in the DAG represent different target expert models or intermediate decision nodes, and the edges represent the dependencies between the nodes.

[0084] Furthermore, by using a Directed Acyclic Graph (DAG), the outputs of multiple target expert models can be comprehensively considered. This allows for the determination of the forgery identification result based on the forgery confidence score processed according to the identification decision logic, ensuring the logicality and interpretability of the decision-making process. For example, if multiple target expert models output high confidence scores, it may be determined that the image to be identified has been forged. If only a few target expert models output high confidence scores, it may be determined that the image to be identified has not been forged. Additionally, a comprehensive confidence score can be obtained based on the weights and rules in the DAG. This comprehensive confidence score can then be used to determine whether the image to be identified has been forged, and the corresponding forgery identification result can be output.

[0085] Furthermore, the target expert models can all be stored within the framework of the hybrid expert model, and the training process of the hybrid expert model framework includes: (1) Before putting each expert model into the framework of the hybrid expert model, each expert model is pre-trained separately according to a specific dataset, which is a region mask image that can be used to enhance the special ability of each expert model in a specific region. For example, if the expert model focuses on detecting the fake hairline region, its training dataset includes a large number of training samples with the hairline region labeled so that the expert model can learn the specific fake features of the hairline region.

[0086] (2) Training and optimization of gating mechanism: The gating mechanism can be used through a multi-layer perceptual neural network structure to determine the target expert model to be activated based on the input information. The training process includes discrete sampling and gradient backpropagation, and auxiliary loss function.

[0087] Discrete sampling and gradient backpropagation are used to address the problem that discrete sampling in the gating mechanism cannot be directly used for gradient descent. This can be solved by using the Top-k Gumbel-Softmax approach, where k is set to 2 to 3 by default. During the training of the gating network, Gumbel noise is introduced to perform a differentiable relaxation approximation on the output of the gating network, allowing the discrete operation of "selecting the target expert model" to be backpropagated. During inference, the top k experts with the highest probability (i.e., the target expert model) are directly selected, thus ensuring that only a few of the most relevant experts are activated each time, maintaining computational efficiency.

[0088] Meanwhile, an auxiliary loss function is used to encourage load balancing among expert models. Specifically, the expert utilization variance penalty loss is used to prevent the gating network from always activating a few popular expert models while ignoring others. The expert utilization variance penalty loss can be expressed by the following expression: aux_loss=Var(expert_activation_counts)×0.1; Here, Var represents the variance, expert_activation_counts represents the number of times each expert model is activated, and 0.1 is an adjustable coefficient constant used to control the weight of this loss term in the total loss. The purpose of the expression is to reduce the variance of the number of times expert models are activated, thereby ensuring that each expert model is used in a balanced way.

[0089] In addition, to ensure alignment between text embedding features and target expert models, a supervision signal can be constructed using a manually labeled "expert-forgery type" mapping table. For example, "hair artifact → Expert1" means that when the text embedding feature mentions hair artifact, the expert model named Expert1 should be activated.

[0090] The alignment loss function is used to supervise gated networks, ensuring that the output target expert distribution is as close as possible to the ideal target expert distribution. The alignment loss function can be expressed by the following expression: gate_loss=CE(gate_logits,expert_label); Where CE represents classification loss, namely the cross-entropy loss function, which measures the difference between two probability distributions; gate_logits represents the output value of the gating mechanism, which represents the probability of each expert model being activated; and expert_label represents the label value of the expert model, which is the target expert model to be activated by looking up the forgery type description in the current input text embedding features from the mapping table.

[0091] (3) The training of the overall hybrid expert model framework can adopt a phased strategy to ensure training stability and performance. The first phase is used to freeze the parameters of all pre-trained expert models and only optimize / update the parameters of the gating network so that the gating network can learn how to make the correct expert selection decision based on the input.

[0092] The second stage is used for fine-tuning all parameters. The parameters of the expert models are unfrozen, and the parameters of the gating network and all expert models are optimized together with a very small learning rate. This allows the overall hybrid expert model framework to be finely and collaboratively adjusted to achieve optimal recognition accuracy.

[0093] This embodiment achieves accurate and visual determination of whether an image to be identified is a forged image by comprehensively considering the outputs of multiple expert models and representing the decision logic through a directed acyclic graph, while also making the decision-making process interpretable.

[0094] In some embodiments, after generating an abnormal region heatmap of the suspected forged region corresponding to the image to be identified based on the forged region features, the method further includes: When the spatial distribution of the heatmap in the abnormal area is unreasonable, the image to be identified is determined to be an adversarial attack sample; Alarm information is generated based on adversarial attack samples and heatmaps of abnormal areas, and then sent to the terminal device.

[0095] Optionally, in this embodiment, if the spatial distribution of the heatmap of abnormal areas is unreasonable, the image to be identified will be determined as an adversarial attack sample. For example, if the heatmap of abnormal areas shows that the probability of forgery is distributed too evenly in the image to be identified, or if there are suspicious forgery areas that do not conform to the conventional forgery characteristics, it indicates that the image to be identified has been subjected to an adversarial attack, so as to mislead the forgery identification.

[0096] In addition, when the image to be identified is determined to be an adversarial attack sample, an alarm message is generated based on the adversarial attack sample and the heat map of the abnormal area, and the alarm message is sent to the terminal device so that security personnel or system administrators can be notified in a timely manner to take corresponding measures, such as further investigation or strengthening security protection.

[0097] The alarm information records in detail the location, forgery features, and corresponding confidence level of the suspicious forged area in the image to be identified.

[0098] Furthermore, adversarial attacks can be identified through an adversarial attack feature learning mechanism. This involves analyzing a large number of adversarial attack samples, learning the characteristic patterns of adversarial attacks, and continuously updating the adversarial attack feature database to improve the ability to identify new adversarial attacks and prevent potential security threats.

[0099] For example, it can be identified that adversarial attacks produce tiny but critical pixel changes in specific areas of the image to be identified, in order to mislead the forgery detection convolutional neural network in identifying forgeries in the image to be identified.

[0100] Furthermore, after detecting adversarial attack samples and generating alarm information, a real-time dynamic defense mechanism can be activated. This mechanism can adjust the parameters of the forgery detection convolutional neural network in real time to resist ongoing adversarial attacks.

[0101] For example, the detection sensitivity for specific regions can be temporarily increased, or the decision threshold of the convolutional neural network for forgery detection can be adjusted to make it more difficult for attackers to succeed.

[0102] This embodiment enables the timely detection and defense of potential security threats by identifying adversarial attack samples, thereby enhancing security. It also generates and sends alarm information so that relevant personnel can take swift action to reduce the occurrence of security incidents.

[0103] In some embodiments, the spatial distribution of the heat map of the abnormal area is reasonable and simultaneously meets the following conditions: The distribution of high-response areas in the abnormal area heatmap matches the distribution of high-response areas corresponding to the preset forgery pattern. The target expert model to be activated is the same as the target expert model corresponding to the suspicious forged area described in the text prompt; Frequency domain analysis determined that there were no abnormal high-frequency components in the image to be identified.

[0104] Optionally, for the spatial distribution of the heat map of the abnormal area in this embodiment to be reasonable, the following three conditions must be met simultaneously: (1) The distribution of high-response areas in the abnormal area heat map needs to conform to the distribution of high-response areas corresponding to the preset forgery mode. That is, in the process of generating the abnormal area heat map, the distribution of high-probability areas in the abnormal area heat map is evaluated according to the known forgery mode to ensure that the abnormal area heat map can accurately reflect the distribution of forgery areas in the image to be identified.

[0105] (2) The target expert model to be activated needs to be the same as the target expert model corresponding to the suspicious forged area described in the text prompt information in order to determine the accuracy of the activated target expert model. The text prompt information provides a detailed description of the suspicious forged area. The corresponding target expert model can be activated according to the description in the text prompt information to conduct in-depth analysis of the specific area. If the target expert model activated according to the text embedding features is inconsistent with the target expert model described in the text prompt information, it is easy to cause deviation in the analysis results and affect the accuracy of forgery identification.

[0106] (3) Determine that there are no abnormal high-frequency components in the image to be identified by frequency domain analysis. Frequency domain analysis is a method used to detect possible forgery traces in an image. In the frequency domain, forged images will produce abnormal high-frequency components. The high-frequency components may correspond to sharpening, edge enhancement or other forgery operations in the image. Abnormal components can be identified by frequency domain analysis and used as the basis for judging whether the image to be identified has been forged. That is, if there are abnormal high-frequency components in the image to be identified, it is considered that the image to be identified may be a forged image.

[0107] In addition, if the spatial distribution of the heatmap in the abnormal area meets any of the above conditions, it will be identified as a forged image.

[0108] This embodiment enables the determination of whether an image to be identified is a forged image by assessing the spatial distribution rationality of the heatmap of abnormal areas, thereby improving the accuracy and reliability of forged image identification.

[0109] To effectively address the shortcomings of traditional technologies, such as the easy omission of local forgery information, the untraceability of overall identification errors, and the low accuracy of forgery identification, and to significantly improve the accuracy of forgery identification and enhance the interpretability of the forgery identification results, this application provides an embodiment of a forgery identification device based on local image understanding for implementing all or part of the aforementioned forgery identification based on local image understanding. See [link to embodiment]. Figure 2 The forgery detection device based on local image understanding specifically includes the following components: The first processing module 10 is used to receive the image to be identified, extract the forgery region features of the image to be identified through a forgery recognition convolutional neural network and a multi-scale attention mechanism, generate an abnormal region heat map of the suspected forgery region corresponding to the image to be identified based on the forgery region features, and the tail of the forgery recognition convolutional neural network includes a deformable convolutional block with a preset number of layers. The second processing module 20 is used to generate structured text prompt information based on the abnormal area heat map and the image to be identified in a preset prompt information format when the spatial distribution of the abnormal area heat map is reasonable. The text prompt information and the image to be identified are encoded to obtain text embedding features. The text prompt information is used to describe the abnormal features of at least one suspected forgery area and the abnormal confidence level corresponding to the abnormal features. The third processing module 30 is used to activate the target expert model corresponding to the suspected forgery region based on the text embedding features, process the corresponding suspected forgery region through the target expert model, obtain the forgery confidence level corresponding to each suspected forgery region, and determine the forgery recognition result of the image to be recognized based on the forgery confidence level.

[0110] As described above, the forgery recognition device based on local image understanding provided in this application can innovatively receive an image to be recognized, extract forgery region features of the image to be recognized through a forgery recognition convolutional neural network and a multi-scale attention mechanism, generate an anomaly region heatmap of the suspected forgery region corresponding to the image to be recognized based on the forgery region features, wherein the tail of the forgery recognition convolutional neural network includes a preset number of deformable convolutions, and when the spatial distribution of the anomaly region heatmap is reasonable, generate structured text prompt information according to a preset prompt information format based on the anomaly region heatmap and the image to be recognized, encode the text prompt information and the image to be recognized to obtain text embedding features, wherein the text prompt information is used to describe the anomaly features of at least one suspected forgery region and the anomaly confidence level corresponding to the anomaly features, activate the target expert model corresponding to the suspected forgery region based on the text embedding features, process the corresponding suspected forgery region through the target expert model to obtain the forgery confidence level corresponding to each suspected forgery region, and determine the forgery recognition result of the image to be recognized based on the forgery confidence level. It can identify potential forgery regions in the image to be recognized, enhance the sensitivity to local forgery information, enrich the semantic information corresponding to the suspected forgery region through text prompt information, and improve the accuracy of forgery recognition. This method effectively addresses the shortcomings of traditional techniques, such as the easy omission of local forged information, the untraceability of overall identification errors, and the low accuracy of forgery identification. It significantly improves the accuracy of forgery identification and enhances the interpretability of forgery identification results.

[0111] From a hardware perspective, in order to effectively address the shortcomings of traditional technologies, such as the easy omission of local forged information, the untraceability of overall identification errors, and the low accuracy of forged identification, and to significantly improve the accuracy of forged identification and enhance the interpretability of the forged identification results, this application provides an embodiment of an electronic device for implementing all or part of the forged identification method based on local image understanding. The electronic device specifically includes the following components: The system comprises a processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to realize information transmission between the forgery recognition device based on local image understanding and core business systems, user terminals, and related databases and other related devices; the logic controller can be a desktop computer, tablet computer, or mobile terminal, etc., and this embodiment is not limited to these. In this embodiment, the logic controller can be implemented with reference to the embodiments of the forgery recognition method based on local image understanding and the embodiments of the forgery recognition device based on local image understanding in the embodiments, the contents of which are incorporated herein, and repeated details will not be described again.

[0112] It is understood that the user terminal may include smartphones, tablet computers, network set-top boxes, portable computers, desktop computers, personal digital assistants (PDAs), in-vehicle devices, smart wearable devices, etc. Among these, the smart wearable devices may include smart glasses, smartwatches, smart bracelets, etc.

[0113] In practical applications, parts of the forgery detection method based on local image understanding can be executed on the electronic device side as described above, or all operations can be completed in the client device. The specific choice depends on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations on this. If all operations are completed in the client device, the client device may further include a processor.

[0114] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.

[0115] Figure 3 This is a schematic block diagram illustrating the system configuration of the electronic device 9600 according to an embodiment of this application. Figure 3 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that... Figure 3 This is an example; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions.

[0116] In one embodiment, the forgery detection method based on local image understanding can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control: Step S101: Receive the image to be identified, extract the forgery region features of the image to be identified through a forgery recognition convolutional neural network and a multi-scale attention mechanism, generate an abnormal region heatmap of the suspected forgery region corresponding to the image to be identified based on the forgery region features, and the tail of the forgery recognition convolutional neural network includes deformable convolutional blocks with a preset number of layers. Step S102: When the spatial distribution of the heat map of the abnormal area is reasonable, a structured text prompt information is generated based on the heat map of the abnormal area and the image to be identified in accordance with the preset prompt information format. The text prompt information and the image to be identified are encoded to obtain text embedding features. The text prompt information is used to describe the abnormal features of at least one suspected forged area and the abnormal confidence level corresponding to the abnormal features. Step S103: Activate the target expert model corresponding to the suspected forgery region based on the text embedding features, process the corresponding suspected forgery region through the target expert model, obtain the forgery confidence of each suspected forgery region, and determine the forgery recognition result of the image to be recognized based on the forgery confidence. As described above, the electronic device provided in this application innovatively receives an image to be identified, extracts forgery region features from the image through a forgery recognition convolutional neural network and a multi-scale attention mechanism, and generates an anomaly region heatmap of the suspected forgery region corresponding to the image to be identified based on the forgery region features. The tail of the forgery recognition convolutional neural network includes a preset number of deformable convolutions. When the spatial distribution of the anomaly region heatmap is reasonable, structured text prompts are generated according to a preset prompt information format based on the anomaly region heatmap and the image to be identified. The text prompts and the image to be identified are encoded to obtain text embedding features. The text prompts describe the anomaly features of at least one suspected forgery region and the corresponding anomaly confidence level. A target expert model corresponding to the suspected forgery region is activated based on the text embedding features. The target expert model processes the corresponding suspected forgery region to obtain the forgery confidence level for each suspected forgery region. The forgery recognition result of the image to be identified is determined based on the forgery confidence level. This method can identify potential forgery regions in the image to be identified, enhances sensitivity to local forgery information, enriches the semantic information corresponding to the suspected forgery region through text prompts, and improves the accuracy of forgery recognition. This method effectively addresses the shortcomings of traditional techniques, such as the easy omission of local forged information, the untraceability of overall identification errors, and the low accuracy of forgery identification. It significantly improves the accuracy of forgery identification and enhances the interpretability of forgery identification results.

[0117] In another embodiment, the forgery recognition device based on local image understanding can be configured separately from the central processing unit 9100. For example, the forgery recognition device based on local image understanding can be configured as a chip connected to the central processing unit 9100, and the function of the forgery recognition method based on local image understanding can be realized through the control of the central processing unit.

[0118] like Figure 3 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily need to include these components. Figure 3 All components shown; in addition, the electronic device 9600 may also include Figure 3 For components not shown, please refer to existing technologies.

[0119] like Figure 3 As shown, the central processing unit 9100, sometimes also referred to as a controller or operating control, may include a microprocessor or other processor device and / or logic device, which receives inputs and controls the operation of various components of the electronic device 9600.

[0120] The memory 9140 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned failure-related information, and also store a program for executing that information. The central processing unit 9100 may execute the program stored in the memory 9140 to perform information storage or processing, etc.

[0121] Input unit 9120 provides input to central processing unit 9100. Input unit 9120 may be, for example, a keypad or touch input device. Power supply 9170 provides power to electronic device 9600. Display 9160 displays images and text. Display may be, for example, an LCD display, but is not limited thereto.

[0122] The memory 9140 can be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), a SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROMs. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 via the central processing unit 9100.

[0123] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for the electronic device for communication functions and / or for performing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0124] The communication module 9110 is a transmitter / receiver that sends and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processing unit 9100 to provide input signals and receive output signals, which is the same as in a conventional mobile communication terminal.

[0125] Based on different communication technologies, multiple communication modules 9110 can be configured in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless LAN modules. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby realizing typical telecommunications functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Additionally, the audio processor 9130 is coupled to a central processing unit 9100, enabling on-device recording via the microphone 9132 and on-device playback of stored audio via the speaker 9131.

[0126] Embodiments of this application also provide a computer-readable storage medium capable of implementing all steps of the forgery identification method based on local image understanding, where the execution subject is a server or client, as described in the above embodiments. The computer-readable storage medium stores a computer program that, when executed by a processor, implements all steps of the forgery identification method based on local image understanding, where the execution subject is a server or client, as described in the above embodiments. For example, when the processor executes the computer program, it implements the following steps: Step S101: Receive the image to be identified, extract the forgery region features of the image to be identified through a forgery recognition convolutional neural network and a multi-scale attention mechanism, generate an abnormal region heatmap of the suspected forgery region corresponding to the image to be identified based on the forgery region features, and the tail of the forgery recognition convolutional neural network includes deformable convolutional blocks with a preset number of layers. Step S102: When the spatial distribution of the heat map of the abnormal area is reasonable, a structured text prompt information is generated based on the heat map of the abnormal area and the image to be identified in accordance with the preset prompt information format. The text prompt information and the image to be identified are encoded to obtain text embedding features. The text prompt information is used to describe the abnormal features of at least one suspected forged area and the abnormal confidence level corresponding to the abnormal features. Step S103: Activate the target expert model corresponding to the suspected forgery region based on the text embedding features, process the corresponding suspected forgery region through the target expert model, obtain the forgery confidence of each suspected forgery region, and determine the forgery recognition result of the image to be recognized based on the forgery confidence.

[0127] As described above, the computer-readable storage medium provided in this application innovatively receives an image to be identified, extracts forgery region features from the image through a forgery recognition convolutional neural network and a multi-scale attention mechanism, and generates an anomaly region heatmap of the suspected forgery region corresponding to the image to be identified based on the forgery region features. The tail of the forgery recognition convolutional neural network includes a preset number of deformable convolutions. Given a reasonable spatial distribution of the anomaly region heatmap, structured text prompts are generated according to a preset prompt information format based on the anomaly region heatmap and the image to be identified. The text prompts and the image to be identified are encoded to obtain text embedding features. The text prompts describe the anomaly features of at least one suspected forgery region and the corresponding anomaly confidence level. A target expert model corresponding to the suspected forgery region is activated based on the text embedding features. The target expert model processes the corresponding suspected forgery region to obtain the forgery confidence level for each suspected forgery region. The forgery recognition result of the image to be identified is determined based on the forgery confidence level. This can identify potential forgery regions in the image to be identified, enhance sensitivity to local forgery information, enrich the semantic information corresponding to the suspected forgery region through text prompts, and improve the accuracy of forgery recognition. This method effectively addresses the shortcomings of traditional techniques, such as the easy omission of local forged information, the untraceability of overall identification errors, and the low accuracy of forgery identification. It significantly improves the accuracy of forgery identification and enhances the interpretability of forgery identification results.

[0128] Embodiments of this application also provide a computer program product capable of implementing all steps of the forgery recognition method based on local image understanding, where the execution subject is a server or client, as described in the above embodiments. When executed by a processor, this computer program / instruction implements the steps of the forgery recognition method based on local image understanding. For example, the computer program / instruction implements the following steps: Step S101: Receive the image to be identified, extract the forgery region features of the image to be identified through a forgery recognition convolutional neural network and a multi-scale attention mechanism, generate an abnormal region heatmap of the suspected forgery region corresponding to the image to be identified based on the forgery region features, and the tail of the forgery recognition convolutional neural network includes deformable convolutional blocks with a preset number of layers. Step S102: When the spatial distribution of the heat map of the abnormal area is reasonable, a structured text prompt information is generated based on the heat map of the abnormal area and the image to be identified in accordance with the preset prompt information format. The text prompt information and the image to be identified are encoded to obtain text embedding features. The text prompt information is used to describe the abnormal features of at least one suspected forged area and the abnormal confidence level corresponding to the abnormal features. Step S103: Activate the target expert model corresponding to the suspected forgery region based on the text embedding features, process the corresponding suspected forgery region through the target expert model, obtain the forgery confidence of each suspected forgery region, and determine the forgery recognition result of the image to be recognized based on the forgery confidence.

[0129] As described above, the computer program product provided in this application innovatively receives an image to be identified, extracts forgery region features from the image through a forgery recognition convolutional neural network and a multi-scale attention mechanism, and generates an anomaly region heatmap of the suspected forgery regions corresponding to the image to be identified based on the forgery region features. The tail of the forgery recognition convolutional neural network includes a preset number of deformable convolutions. When the spatial distribution of the anomaly region heatmap is reasonable, structured text prompts are generated according to a preset prompt information format based on the anomaly region heatmap and the image to be identified. The text prompts and the image to be identified are encoded to obtain text embedding features. The text prompts describe the anomaly features of at least one suspected forgery region and the corresponding anomaly confidence level. A target expert model corresponding to the suspected forgery region is activated based on the text embedding features. The target expert model processes the corresponding suspected forgery regions to obtain the forgery confidence level for each suspected forgery region. The forgery recognition result of the image to be identified is determined based on the forgery confidence level. This can identify potential forgery regions in the image to be identified, enhance sensitivity to local forgery information, enrich the semantic information corresponding to suspected forgery regions through text prompts, and improve the accuracy of forgery recognition. This method effectively addresses the shortcomings of traditional techniques, such as the easy omission of local forged information, the untraceability of overall identification errors, and the low accuracy of forgery identification. It significantly improves the accuracy of forgery identification and enhances the interpretability of forgery identification results.

[0130] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0131] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0132] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0133] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0134] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A forgery recognition method based on local image understanding, characterized in that, The method includes: The system receives an image to be identified, extracts the forgery region features of the image to be identified through a forgery recognition convolutional neural network and a multi-scale attention mechanism, and generates an abnormal region heatmap of the suspected forgery region corresponding to the image to be identified based on the forgery region features. The tail of the forgery recognition convolutional neural network includes deformable convolutional blocks with a preset number of layers. When the spatial distribution of the heatmap of the abnormal region is reasonable, a structured text prompt information is generated based on the heatmap of the abnormal region and the image to be identified in accordance with a preset prompt information format. The text prompt information and the image to be identified are encoded to obtain text embedding features. The text prompt information is used to describe the abnormal features of at least one of the suspected forgery regions and the abnormal confidence level corresponding to the abnormal features. The target expert model corresponding to the suspected forgery region is activated based on the text embedding features. The suspected forgery region is processed by the target expert model to obtain the forgery confidence score corresponding to each suspected forgery region. The forgery recognition result of the image to be identified is determined based on the forgery confidence score.

2. The method according to claim 1, characterized in that, The extraction of forgery region features from the image to be identified using a forgery detection convolutional neural network and a multi-scale attention mechanism includes: Based on the forgery detection convolutional neural network, multi-level features to be identified in the image to be identified are extracted. The features to be identified include low-level features and high-level features. The low-level features are used to obtain abnormal pixels in the image to be identified, and the high-level features are used to obtain semantic information of the image to be identified. By using a multi-scale attention mechanism, the features to be identified corresponding to a preset number of layers in the forgery recognition convolutional neural network are fused according to the feature pyramid network structure to obtain forgery region features.

3. The method according to claim 1, characterized in that, The step of generating an abnormal region heatmap of the suspected forged region corresponding to the image to be identified based on the forged region features includes: Pixel-level probability mapping is performed on the features of the forged region, and the probability value of each pixel in the image to be identified belonging to the forged region is mapped to a heatmap. The heatmap is colored by a preset mapping relationship between color and probability value to obtain an abnormal area heatmap.

4. The method according to claim 1, characterized in that, The preset prompt information format is a triple structure, including region type, anomaly description, and anomaly confidence level; The step of generating structured text prompt information based on the heatmap of the abnormal region and the image to be identified according to a preset prompt information format includes: Based on the heat map of the abnormal areas, the region type corresponding to the suspected forgery area and the anomaly confidence level corresponding to the suspected forgery area are determined. An anomaly description is generated for each suspected forgery region by using a preset anomaly description vocabulary. The region type, the anomaly description, and the anomaly confidence level are combined according to the preset prompt information format to obtain text prompt information.

5. The method according to claim 1, characterized in that, The step of determining the forgery identification result of the image to be identified based on the forgery confidence level includes: The forgery confidence scores output by each of the target expert models are received. Based on the target expert models and the forgery confidence scores, an identification decision logic is determined, wherein the identification decision logic is represented by a directed acyclic graph. The forgery confidence level is processed based on the aforementioned identification decision logic to obtain the forgery identification result.

6. The method according to claim 1, characterized in that, After generating the abnormal region heatmap of the suspected forged region corresponding to the image to be identified based on the forged region features, the method further includes: If the spatial distribution of the heatmap in the abnormal area is unreasonable, the image to be identified will be determined as an adversarial attack sample. Alarm information is generated based on the adversarial attack sample and the abnormal area heatmap, and the alarm information is sent to the terminal device.

7. The method according to claim 1, characterized in that, The spatial distribution of the heat map of the abnormal region is reasonable and simultaneously meets the following conditions: The distribution of high-response areas in the abnormal area heatmap conforms to the distribution of high-response areas corresponding to the preset forgery pattern. The target expert model to be activated is the same as the target expert model corresponding to the suspicious forged area described in the text prompt information; Frequency domain analysis determined that there were no abnormal high-frequency components in the image to be identified.

8. A forgery detection device based on local image understanding, characterized in that, The device includes: The first processing module is used to receive an image to be identified, extract the forgery region features of the image to be identified through a forgery recognition convolutional neural network and a multi-scale attention mechanism, and generate an abnormal region heatmap of the suspected forgery region corresponding to the image to be identified based on the forgery region features. The tail of the forgery recognition convolutional neural network includes a deformable convolutional block with a preset number of layers. The second processing module is used to generate structured text prompt information based on the abnormal area heat map and the image to be identified according to a preset prompt information format, provided that the spatial distribution of the abnormal area heat map is reasonable; and to encode the text prompt information and the image to be identified to obtain text embedding features. The text prompt information is used to describe the abnormal features of at least one of the suspected forgery areas and the abnormal confidence level corresponding to the abnormal features. The third processing module is used to activate the target expert model corresponding to the suspected forgery region based on the text embedding features, process the corresponding suspected forgery region through the target expert model, obtain the forgery confidence level corresponding to each suspected forgery region, and determine the forgery recognition result of the image to be identified based on the forgery confidence level.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the forgery recognition method based on local image understanding as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the forgery identification method based on local image understanding as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • False video detection method and system based on forensic embedding and context embedding

    CN116363764A

  • Multi-scale feature fusion depth forgery detection method based on reconstruction learning

    CN119068318A

  • Intelligent machine vision detection method and system based on image processing and storage medium

    CN119205719A

  • Image forgery detection method and system based on multi-expert model decision

    CN119762958A

  • Artificial intelligence forged content detection method based on multi-expert model joint decision

    CN119832552A

Cited By

  • Content counterfeiting detection and positioning method and electronic equipment

    CN121685527A

  • Image detection method and system based on artificial intelligence

    CN121810681A

  • Face deepfake detection method and system based on hybrid expert network

    CN122473829A