Method and apparatus for forgery recognition based on local image understanding
By employing a forgery recognition method based on local image understanding, this method utilizes a convolutional neural network for forgery recognition and a multi-scale attention mechanism to generate heatmaps of abnormal regions. Combined with a target expert model and textual prompts, it solves the problems of easily overlooked local forgery features and unexplainable decisions in existing technologies, achieving highly accurate and interpretable forgery recognition.
Patent Information
- Application Number
- CN202511485152.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Existing technologies tend to overlook local forgery features when identifying fake images, leading to decreased recognition accuracy and a lack of observability and interpretability in the decision-making process, thus reducing user experience.
A forgery recognition method based on local image understanding is adopted. Image features are extracted through a convolutional neural network for forgery recognition and a multi-scale attention mechanism to generate anomaly region heatmaps. Combined with target expert models and textual prompts, structured description and decision logic processing are performed to improve the accuracy and interpretability of forgery recognition.
It significantly improves the accuracy and interpretability of counterfeit identification, accurately locates suspicious counterfeit areas, and provides detailed semantic descriptions and reliable identification results.
Smart Images

Figure CN120953778B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, in particular to a forgery identification method and device based on local image understanding. BACKGROUND
[0002] With the development of digital media technology, deep forgery technology is increasingly realistic, and with it comes a number of security risks, such as the frequent occurrence of false information dissemination, privacy invasion, and identity theft. Deep forgery technology can generate highly realistic fake face images and videos through deep learning algorithms, making it difficult for ordinary users to distinguish between true and false.
[0003] In the prior art, the method of identifying forgery technology is easily limited to a specific data set, and local forgery features are easily ignored or not sensitive enough, resulting in a decrease in the accuracy of forgery identification. In addition, in the process of identifying forgery through the whole or a large model, the decision-making process lacks observability and explainability, making it difficult for users to understand the basis for judgment, reducing the trust in the forgery identification result, and reducing user experience. SUMMARY
[0004] To solve the problems in the prior art, the present application provides a forgery identification method and device based on local image understanding, which can effectively solve the problems of traditional technology in local forgery information being easily missed, whole identification errors being untraceable, and low accuracy of forgery identification, significantly improving the accuracy of forgery identification and enhancing the explainability of determining the forgery identification result.
[0005] To solve at least one of the above problems, the present application provides the following technical solutions:
[0006] In a first aspect, the present application provides a forgery identification method based on local image understanding, comprising:
[0007] Receiving an image to be identified, extracting forgery region features of the image to be identified through a forgery identification convolutional neural network and a multi-scale attention mechanism, generating an abnormal region heat map of a suspicious forgery region corresponding to the image to be identified based on the forgery region features, and the tail of the forgery identification convolutional neural network includes a deformable convolution block of a preset number of layers;
[0008] In the case that the spatial distribution of the abnormal region heat map is reasonable, generating a structured text prompt information based on the abnormal region heat map and the image to be identified according to a preset prompt information format, encoding the text prompt information and the image to be identified to obtain a text embedding feature, and the text prompt information is used to describe the abnormal features of at least one suspicious forgery region and the abnormal confidence corresponding to the abnormal features;
[0009] The suspiciously forged region corresponding to the target expert model is activated based on the text embedding feature, the suspiciously forged region is processed by the target expert model, the forgery confidence corresponding to each suspiciously forged region is obtained, and the forgery recognition result of the to-be-identified image is determined based on the forgery confidence.
[0010] Further, it also includes: based on the forgery recognition convolutional neural network, extracting the multi-level to-be-identified features of the to-be-identified image, the to-be-identified features including low-level to-be-identified features and high-level to-be-identified features, the low-level to-be-identified features being used to obtain abnormal pixels of the to-be-identified image, and the high-level to-be-identified features being used to obtain semantic information of the to-be-identified image;
[0011] Through the multi-scale attention mechanism, the to-be-identified features corresponding to the preset number of layers in the forgery recognition convolutional neural network are fused according to the feature pyramid network structure, and the forgery region feature is obtained.
[0012] Further, it also includes: performing pixel-level probability mapping on the forgery region feature, and mapping the probability value of each pixel point in the to-be-identified image belonging to the forgery region to a heat map;
[0013] The heat map is colored through a preset mapping relationship between the color and the probability value, and an abnormal region heat map is obtained.
[0014] Further, the preset prompt information format is a triple structure, including a region type, an abnormal description, and an abnormal confidence, and further including:
[0015] Based on the abnormal region heat map, the region type corresponding to the suspiciously forged region and the abnormal confidence corresponding to the suspiciously forged region are determined;
[0016] Each suspiciously forged region generates a corresponding abnormal description through a preset abnormal description vocabulary, the region type, the abnormal description, and the abnormal confidence are combined according to the preset prompt information format, and a text prompt information is obtained.
[0017] Further, it also includes: receiving the forgery confidence output by each target expert model, and determining an identification decision logic based on the target expert model and the forgery confidence, wherein the identification decision logic is represented by a directed acyclic graph;
[0018] The forgery recognition result is obtained by processing the forgery confidence based on the identification decision logic.
[0019] Further, after generating the abnormal region heat map of the to-be-identified image corresponding to the suspiciously forged region based on the forgery region feature, it also includes:
[0020] In the case that the spatial distribution of the abnormal region heat map is unreasonable, the to-be-identified image is determined as an adversarial attack sample;
[0021] The alarm information is generated based on the adversarial attack sample and the abnormal region heat map, and is sent to a terminal device.
[0022] Further, the spatial distribution of the abnormal region heat map satisfies the following conditions:
[0023] The distribution of the high response region in the abnormal region heat map conforms to the distribution of the high response region corresponding to the preset forgery mode;
[0024] The target expert model to be activated is the same as the target expert model corresponding to the suspicious forgery region described in the text prompt information;
[0025] The abnormal high frequency component does not exist in the to-be-identified image is determined through frequency domain analysis.
[0026] In a second aspect, the present application provides a forgery identification device based on local image understanding, comprising:
[0027] The first processing module is configured to receive a to-be-identified image, extract a forgery region feature of the to-be-identified image through a forgery identification convolutional neural network and a multi-scale attention mechanism, generate an abnormal region heat map of a suspicious forgery region corresponding to the to-be-identified image based on the forgery region feature, and the tail of the forgery identification convolutional neural network includes a deformable convolutional block with a preset number of layers;
[0028] The second processing module is configured to generate a structured text prompt information in a preset prompt information format based on the abnormal region heat map and the to-be-identified image in the case that the spatial distribution of the abnormal region heat map is reasonable, encode the text prompt information and the to-be-identified image to obtain a text embedding feature, and the text prompt information is used to describe the abnormal features of at least one suspicious forgery region and the abnormal confidence corresponding to the abnormal features;
[0029] The third processing module is configured to activate a target expert model corresponding to the suspicious forgery region based on the text embedding feature, process the corresponding suspicious forgery region through the target expert model, obtain a forgery confidence corresponding to each suspicious forgery region, and determine a forgery identification result of the to-be-identified image based on the forgery confidence.
[0030] In a third aspect, the present application provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to realize the steps of the forgery identification method based on local image understanding.
[0031] In a fourth aspect, the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to realize the steps of the forgery identification method based on local image understanding.
[0032] In a fifth aspect, the present application provides a computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of the method for forgery identification based on local image understanding.
[0033] According to the technical solution, the method and device for forgery identification based on local image understanding are provided. The image to be identified is received innovatively. The forgery area features of the image to be identified are extracted by the forgery identification convolutional neural network and the multi-scale attention mechanism. The abnormal area heat map of the suspicious forgery area corresponding to the image to be identified is generated according to the forgery area features. The tail of the forgery identification convolutional neural network comprises a deformable convolution of a preset number of layers. When the spatial distribution of the abnormal area heat map is reasonable, the structured text prompt information is generated according to the abnormal area heat map and the image to be identified in a preset prompt information format. The text prompt information is used to describe the abnormal features of at least one suspicious forgery area and the abnormal confidence corresponding to the abnormal features. The target expert model corresponding to the suspicious forgery area is activated according to the text embedding features. The forgery confidence corresponding to each suspicious forgery area is obtained by processing the corresponding suspicious forgery area through the target expert model. The forgery identification result of the image to be identified is determined according to the forgery confidence. The potential forgery area in the image to be identified can be identified. The sensitivity to local forgery information is enhanced. The semantic information corresponding to the suspicious forgery area is enriched through the text prompt information. The accuracy of the forgery identification is improved. The method effectively solves the problems of the traditional technology, such as easy omission of local forgery information, untraceable of overall identification error, and low accuracy of forgery identification. The accuracy of the forgery identification is significantly improved. The explainability of the determined forgery identification result is enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application. Those skilled in the art can also obtain other drawings according to these drawings without creating any inventive labor.
[0035] Figure 1 The flowchart of the method for forgery identification based on local image understanding in the embodiments of the present application;
[0036] Figure 2 The structural diagram of the device for forgery identification based on local image understanding in the embodiments of the present application;
[0037] Figure 3 The structural diagram of the electronic device in the embodiments of the present application.
[0038] Reference Signs:
[0039] The electronic device 9600, the central processor 9100, the memory 9140, the communication module 9110, the input unit 9120, the audio processor 9130, the display 9160, the power supply 9170, the buffer memory 9141, the application / function storage unit 9142, the data storage unit 9143, the driver program storage unit 9144, the antenna 9111, the speaker 9131, and the microphone 9132. DETAILED DESCRIPTION
[0040] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0041] The acquisition, storage, use, processing, and the like of data in the technical solutions of the present application comply with relevant provisions of national laws and regulations.
[0042] In the prior art, the current identification forgery algorithm can detect through prompt word text, however, the prompt word text is simple and fixed in structure, does not play a supervisory role, or misses local forgery information in the to-be-identified image, and there are problems of non-traceability and observation of identification errors in the use process of overall identification of the to-be-identified image and the preset model.
[0043] In view of the problems in the prior art, the present application provides a forgery identification method and device based on local image understanding, which can focus on the potential forgery area in the to-be-identified image through the expert model framework for local / global forgery identification, generate targeted analysis prompt word text, guide the activation of related expert networks of the expert model framework, and can realize accurate positioning and description of forgery traces through multi-granularity visual analysis and semantic reasoning, mainly focusing on face analysis, activating different target expert models through dynamically generated prompt word text, and predicting time cost according to the number of activated target expert models, so that the process of forgery identification has explainability.
[0044] In order to effectively solve the problems of easy omission of local forgery information, non-traceability of overall identification error, and low accuracy of forgery identification in the prior art, significantly improve the accuracy of forgery identification, and enhance the explainability of determining the forgery identification result, the present application provides an embodiment of a forgery identification method based on local image understanding, as shown in Figure 1 , which specifically includes the following contents:
[0045] Step S101: receiving an image to be identified, extracting a counterfeit area feature of the image to be identified through a counterfeit identification convolutional neural network and a multi-scale attention mechanism, and generating an abnormal area heat map of a suspicious counterfeit area corresponding to the image to be identified based on the counterfeit area feature.
[0046] The tail of the counterfeit identification convolutional neural network includes a deformable convolution block of a preset number of layers.
[0047] Optionally, the embodiment receives an image to be identified, extracts a counterfeit area feature from the image to be identified through a counterfeit identification convolutional neural network combined with a multi-scale attention mechanism, wherein the counterfeit area feature can represent macro counterfeit traces in the image to be identified, and can also represent micro-level abnormalities, that is, abnormal pixels and semantic information in the image to be identified are obtained through low-level features and high-level features. The abnormal pixels are pixel points with inconsistent textures and / or inconsistent gradients in the image to be identified.
[0048] The tail of the counterfeit identification convolutional neural network includes a deformable convolution block of a preset number of layers, for example, 3 deformable convolution blocks are included in the tails of the fifth and sixth layers of the counterfeit identification convolutional neural network, which are used to enhance the positioning ability of irregular suspicious counterfeit areas, so that the positioning of the suspicious counterfeit areas is more accurate. The deformable convolution block can be Deformable Conv.
[0049] In addition, an abnormal heat map of a suspicious counterfeit area of the image to be identified is generated according to the counterfeit area feature, that is, a suspicious counterfeit area of the image to be identified is generated according to the counterfeit area feature, and is converted into an abnormal heat map. The abnormal heat map is a pixel-level probability map, which is used to display the probability that a pixel point is a suspicious counterfeit area.
[0050] The local binary pattern branch can be inserted in the third layer of the counterfeit identification convolutional neural network, so as to enhance the detection of texture inconsistency of the image to be identified in the process of training the counterfeit identification convolutional neural network.
[0051] The embodiment realizes accurate positioning of a suspicious counterfeit area through a multi-scale attention mechanism and a deformable convolution block, and improves the accuracy of counterfeit identification.
[0052] Step S102: generating a structured text prompt information according to a preset prompt information format based on the abnormal area heat map and the image to be identified in the case that the spatial distribution of the abnormal area heat map is reasonable, and encoding the text prompt information and the image to be identified to obtain a text embedding feature.
[0053] The text prompt information is used to describe abnormal features of at least one suspicious counterfeit area and abnormal confidence corresponding to the abnormal features.
[0054] Optionally, in the case that the spatial distribution of the abnormal area heat map is reasonable, the abnormal heat map and the original image to be identified are used to generate a structured text prompt information according to a preset prompt information format, that is, the forged clues in the abnormal heat map are converted into a text form easy to understand and process, and the specific type of the suspicious forged area is determined in combination with the image to be identified, and the text prompt information is obtained.
[0055] The text prompt information is used to describe the abnormal features of the at least one suspicious forged area and the abnormal confidence corresponding to the abnormal features, for example, "hairline unnatural, abnormal confidence 0.7", wherein "hairline" represents the suspicious forged area, "unnatural" represents the abnormal feature, and "abnormal confidence 0.7" represents the abnormal confidence corresponding to the abnormal feature. The text prompt information can also be represented as "hairline unnatural, 0.7".
[0056] In addition, the text prompt information and the image to be identified are respectively encoded, and are fused into the same feature to obtain a text embedding feature, wherein the text embedding feature can provide rich context information for activating the target expert model.
[0057] The text prompt information can be encoded by a CLIP text encoder, the image to be identified can be encoded by a CLIP visual encoder, and the encoding corresponding to the text prompt information and the encoding corresponding to the image to be identified are fused to obtain the text embedding feature.
[0058] The embodiment realizes obtaining the text prompt information according to the abnormal area heat map, and thus obtains the text embedding feature according to the text prompt information and the image to be identified, which can provide detailed and accurate semantic description for the suspicious forged area and the image to be identified.
[0059] Step S103: activating the target expert model corresponding to the suspicious forged area based on the text embedding feature, processing the corresponding suspicious forged area by the target expert model, obtaining the forgery confidence corresponding to each suspicious forged area, and determining the forgery identification result of the image to be identified based on the forgery confidence.
[0060] Optionally, the target expert model corresponding to the suspicious forged area is activated according to the text embedding feature, wherein the target expert model is used for in-depth analysis focusing on a specific forged area, for example, a hairline area expert model focuses on identifying the forged features of the hairline area, and a lip area expert model focuses on identifying the forged features of the lip area.
[0061] In addition, the suspiciously forged region is processed by the target expert model to obtain a corresponding forgery confidence of each suspiciously forged region, and a forgery recognition result of the to-be-identified image is determined according to the forgery confidence. The determination of the forgery confidence includes comprehensive analysis of the output of the target expert model, the decision logic can be represented by a directed acyclic graph (DAG), and the output of each expert model is comprehensively analyzed according to a preset decision rule to determine the forgery recognition result.
[0062] The forgery recognition result includes that the to-be-identified image is a forged image or the to-be-identified image is not a forged image. In the case that the to-be-identified image is a forged image, the forgery recognition result can further include a forged region in the to-be-identified image.
[0063] Further, the target expert model that needs to be activated can be determined by a gating mechanism to process the suspiciously forged region.
[0064] Further, the target expert model can also be deployed on multiple computing sub-nodes for processing by using a distributed computing architecture, so that the target expert model can be processed in parallel, and the forgery confidence of each suspiciously forged region is obtained, and then the forgery recognition result is obtained, thereby improving the processing speed and efficiency of forgery recognition.
[0065] The embodiment realizes that the target expert model corresponding to the suspiciously forged region is activated by the text embedding feature, so that the target expert model processes the suspiciously forged region in the to-be-identified image to determine the forgery recognition result, which can focus on the local to-be-identified image and prevent missing details, thereby improving the accuracy of forgery recognition.
[0066] The embodiment realizes accurate positioning of the suspiciously forged region by using the multi-scale attention mechanism and the deformable convolution block, improves the accuracy of forgery recognition, and the structured text prompt information can describe the suspiciously forged region in detail, so that the user can intuitively understand the basis of the forgery recognition result. In addition, the target expert model can be used to identify the local to-be-identified image, prevent missing, and improve the accuracy of forgery recognition.
[0067] In some embodiments, the forgery region features of the to-be-identified image are extracted by using the forgery recognition convolutional neural network and the multi-scale attention mechanism, including:
[0068] The multi-level to-be-identified features of the to-be-identified image are extracted based on the forgery recognition convolutional neural network. The to-be-identified features include low-level to-be-identified features and high-level to-be-identified features. The low-level to-be-identified features are used to obtain abnormal pixels of the to-be-identified image, and the high-level to-be-identified features are used to obtain semantic information of the to-be-identified image.
[0069] The multi-scale attention mechanism is used to fuse the to-be-identified features corresponding to the preset layers in the forgery identification convolutional neural network according to a feature pyramid network structure, to obtain the forgery region features.
[0070] Optionally, the multi-level to-be-identified features are extracted by the forgery identification convolutional neural network from the to-be-identified image, wherein the to-be-identified features include but are not limited to low-level to-be-identified features and high-level to-be-identified features. The low-level to-be-identified features are mainly used to capture abnormal pixels in the to-be-identified image, such as discontinuous edges or abnormal textures. The high-level to-be-identified features are mainly used to obtain semantic information of the to-be-identified image, such as the layout of a face, irrelevant relative positions, etc. The multi-level to-be-identified features obtained through layered feature extraction can comprehensively understand and analyze the suspicious forgery region in the to-be-identified image from different levels.
[0071] In addition, the multi-scale attention mechanism is used to fuse the multi-level to-be-identified features according to a feature pyramid network (FPN). The feature pyramid network structure can comprehensively consider the to-be-identified features at multiple scales, so as to more accurately locate the suspicious forgery region, thereby fully utilizing the to-be-identified features at different levels and enhancing the identification ability of complex forgery methods.
[0072] For example, a forged image may have subtle pixel-level changes in a local region, and also have unreasonable semantic changes in the overall structure. The forgery region features fused by the multi-scale attention mechanism can represent abnormalities at different levels at the same time.
[0073] The multi-scale attention mechanism is used to fuse the to-be-identified features output by the third layer to the to-be-identified features output by the seventh layer in the forgery identification convolutional neural network. In the fusion process, the edge details of the low-level to-be-identified features are retained, and the semantic abnormality of the high-level to-be-identified features is captured, so as to obtain the forgery region features.
[0074] Further, in the process of extracting the low-level to-be-identified features, a local binary pattern (LBP) feature can be introduced. The local binary pattern is a texture description factor that can extract local texture information of the to-be-identified image. Therefore, by introducing the LBP feature in the low layer of the forgery identification convolutional neural network, the texture abnormalities in the image can be more sensitively captured.
[0075] For example, the forged image may have discontinuity or unnatural transition of textures in a local region. The LBP feature can effectively describe the abnormal textures, thereby improving the detection ability of local forgery traces.
[0076] Further, a dynamic adjustment mechanism can be added to the multi-scale attention mechanism, so that the weight of attention can be dynamically adjusted according to the content of the to-be-identified image and the to-be-identified feature.
[0077] For example, when processing a to-be-identified image mainly including local forgery, more attention is paid to low-level to-be-identified features, and when processing a to-be-identified image with the overall structure being tampered with, more attention is paid to high-level to-be-identified features, which improves the adaptability and can be optimized according to specific conditions, thereby improving the efficiency and accuracy of forgery identification.
[0078] The embodiment realizes extraction of the forgery region feature of the to-be-identified image through the forgery identification convolutional neural network and the multi-scale attention mechanism, can comprehensively understand and analyze the forgery information in the to-be-identified image from different levels, can more accurately locate the suspicious forgery region, and improves the accuracy of forgery identification.
[0079] In some embodiments, generating an abnormal region heat map of the suspicious forgery region corresponding to the to-be-identified image based on the forgery region feature comprises:
[0080] The pixel-level probability mapping is performed on the forgery region feature, and the probability value of each pixel point in the to-be-identified image belonging to the forgery region is mapped to the heat map.
[0081] The heat map is colored through a preset mapping relationship between color and probability value, and the abnormal region heat map is obtained.
[0082] Optionally, the pixel-level probability mapping is performed on the forgery region feature, that is, the probability value of each pixel point in the to-be-identified image belonging to the suspicious forgery region is mapped to the heat map, so as to intuitively display the probability of each pixel point in the to-be-identified image belonging to the suspicious forgery region.
[0083] For example, if the probability value of a pixel point belonging to the suspicious forgery region is high, the pixel point is assigned a high probability value in the heat map, so as to visually highlight the pixel point.
[0084] In addition, the heat map is colored through a preset mapping relationship between color and probability value, and the abnormal region heat map is obtained, wherein the coloring process enhances the visual effect of the heat map, so that the difference between the suspicious forgery region and the non-suspicious forgery region is more obvious, and the user can quickly identify and analyze.
[0085] For example, the pixel point with a high probability value can be colored red to represent a high probability of belonging to the suspicious forgery region, and the pixel point with a low probability value can be colored blue to represent a low probability of belonging to the suspicious forgery region, so that the user can intuitively see the suspicious forgery region in the to-be-identified image.
[0086] The anomaly region heatmap is the output of the forgery detection convolutional neural network processing the image to be identified. The process of obtaining the forgery detection convolutional neural network includes:
[0087] Training dataset construction: The training dataset mainly includes two types of samples: real samples and fake human samples. Real samples are derived from sources including but not limited to actual photographed portraits or videos, and public datasets (such as CelebA, FFHQ, etc.). Fake samples can be generated through various algorithms and APIs, including but not limited to local fake samples and global fake samples.
[0088] Local forged samples can be generated using algorithms such as mask fusion, T-shaped glasses fusion, mouth fusion, and nose fusion, and corresponding binary masks can be generated to annotate the local forged regions. Global forged samples can be generated using algorithms such as face swapping and face replay, and the corresponding mask regions can be extracted using pixel difference comparison algorithms. In addition, publicly available forged datasets, such as FaceForensics++, can be collected to enhance the data diversity of the training dataset.
[0089] Design of the hybrid loss function: The hybrid loss function is used to supervise the training of the forgery detection convolutional neural network. The hybrid loss function consists of three parts: focus loss, multi-scale result similarity loss, and gradient sensitive loss.
[0090] Focus loss is used to alleviate the problem of uneven distribution of positive and negative samples in classification tasks, and to improve the overall performance of convolutional neural networks for forgery detection. The expression for focus loss is: ;
[0091] in, Used to represent the predicted probability of a real sample, that is, the confidence level of a convolutional neural network for forgery detection in predicting that the current sample belongs to a real sample;
[0092] Used for class balancing weights, that is, adjusting the weight balance between positive and negative samples in the training data to prevent the convolutional neural network for forgery detection from being biased towards the majority class;
[0093] γ represents the adjustment factor, which can be set to γ≥0 to control the weight of easy and difficult samples. That is, when the value of γ is large, the convolutional neural network for forgery detection pays more attention to the training data that is difficult to classify.
[0094] Multi-scale result similarity loss (MS-SSIM Loss) is used to capture the differences in texture structure between fake and real regions, enabling the forgery detection convolutional neural network to better learn the subtle differences between fake and real samples. The expression for multi-scale result similarity loss is: L MS-SSIM =1−MS-SSIM(x,y);
[0095] wherein x represents a fake sample, y represents a real sample, MS-SSIM represents a multi-scale structural similarity measurement, and is used to measure the structural consistency of two training samples at multiple scales.
[0096] The gradient sensitive loss is used to strengthen the gradient difference of the fake boundary region, and the expression of the gradient sensitive loss is:
[0097] wherein, is used for the region weight mask, and is used to emphasize the importance of the fake region, represents element-wise multiplication, represents a gradient operator, represents a fake sample, represents a real sample, represents an L1 norm.
[0098] The expression of the hybrid loss function is: wherein a, b and c represent constant weight coefficients, and are used to balance the influence of different loss terms, and the model parameters with the best performance are selected according to the performance of the validation set in the training process of the fake identification convolutional neural network.
[0099] Further, in the process of processing dynamic image data, the abnormal region heat map is updated in real time to reflect the real-time fake feature changes in the image sequence to be identified, so that the fake behavior can be continuously monitored in the video stream, and new fake signs can be found in time. For example, in a real-time monitoring video, an abnormal region heat map can be generated and updated in real time to help monitoring personnel quickly identify and respond to fake behavior.
[0100] The embodiment realizes that the abnormal region heat map is generated based on the fake region features, the possibility of fake of each pixel point in the image to be identified can be intuitively displayed, and the abnormal region heat map can reflect the possibility of fake of each pixel point, so that the identification of suspicious fake regions is more accurate.
[0101] In some embodiments, the preset prompt information format is a three-tuple structure, including a region type, an abnormal description and an abnormal confidence;
[0102] Based on the abnormal region heat map and the image to be identified, a structured text prompt information is generated according to the preset prompt information format, including:
[0103] Based on the abnormal region heat map, a region type corresponding to the suspicious fake region and an abnormal confidence corresponding to the suspicious fake region are determined.
[0104] The corresponding abnormality description is generated for each suspicious counterfeit region by using a preset abnormality description vocabulary, and the region type, abnormality description, and abnormality confidence are combined according to a preset prompt information format to obtain text prompt information.
[0105] Optionally, the preset prompt information format of the embodiment adopts a triple structure including the region type, abnormality description, and abnormality confidence, wherein the structured prompt information format can accurately describe the characteristics of the suspicious counterfeit region in the image to be recognized.
[0106] For the probability distribution in the abnormal region heat map, a high-probability region is identified and determined as an abnormal region, and the region type of the abnormal region in the image to be recognized is determined, such as a hairline region, a lip region, an eye region, an eyebrow region, a nose region, a hair region, and a face region. Meanwhile, according to the probability value corresponding to the abnormal region, the corresponding abnormality confidence is determined, wherein the abnormality confidence represents the possibility of the region being counterfeit.
[0107] In addition, a corresponding abnormality description is generated for each suspicious counterfeit region by using a preset abnormality description vocabulary, wherein the abnormality description vocabulary includes but is not limited to various possible counterfeit feature descriptions, such as “blur”, “unnatural”, “texture inconsistency”, “shaking”, “disharmony”, and the like.
[0108] In addition, according to the specific characteristics of the abnormal region, a suitable description word is selected from the abnormality description vocabulary to generate the abnormality description corresponding to the abnormal region, and the region type, abnormality description, and abnormality confidence are combined according to the preset prompt information format to obtain complete text prompt information.
[0109] For example, the generated text prompt information can be “the hairline region is unnatural, with an abnormality confidence of 0.7; the lip region is blurred, with an abnormality confidence of 0.8”, and the structured text prompt information is easy to understand and provides guidance for expert model activation.
[0110] Further, the process of converting vision into natural language text prompt information can also be represented as: inputting the abnormal heat map and the corresponding image to be recognized, generating text prompt information according to the preset prompt information format, inputting the abnormal heat map, the image to be recognized, and the text prompt information into a multi-modal large model to obtain the text prompt information.
[0111] Further, through the context perception mechanism, more accurate and detailed abnormality descriptions can be generated according to the overall content and context information of the image to be recognized.
[0112] For example, a suspiciously forged region is located near the hairline region of the face, and the texture of the suspiciously forged region is obviously inconsistent with the surrounding region, which can be described as "texture inconsistency", and further described as "the texture of the hairline region is inconsistent with the texture of the surrounding hair, and there may be editing marks" in combination with the context information, so as to make the abnormal description more specific and accurate, and improve the reliability of the forgery identification.
[0113] The embodiment realizes the generation of structured text prompt information through the preset prompt information format, can accurately describe the region type, abnormal description and abnormal confidence of the suspiciously forged region in the to-be-identified image, clearly shows the related information of the suspiciously forged region to the user, and guides the activation of the target expert model.
[0114] In some embodiments, determining a forgery identification result of the to-be-identified image based on the forgery confidence includes:
[0115] Receiving the forgery confidence output by each target expert model, determining an identification decision logic based on the target expert model and the forgery confidence, wherein the identification decision logic is represented by a directed acyclic graph.
[0116] Processing the forgery confidence based on the identification decision logic to obtain the forgery identification result.
[0117] Optionally, the embodiment receives the forgery confidence output by each target expert model, wherein each target expert model focuses on the analysis of a specific suspiciously forged region and outputs the forgery confidence corresponding to the suspiciously forged region, and the forgery confidence represents the possibility of the suspiciously forged region being forged.
[0118] For example, the forgery confidence output by the target expert model corresponding to the hairline region is 0.9, indicating that the hairline region has a high possibility of being forged, and the forgery confidence output by the target expert model corresponding to the lip region is 0.4, indicating that the lip region has a low possibility of being forged.
[0119] In addition, the identification decision logic is determined based on the target expert model and the forgery confidence, wherein the identification decision logic can be represented by a directed acyclic graph (DAG), different nodes in the DAG represent different target expert models or intermediate decision nodes, and the edges represent the dependency relationship between the nodes.
[0120] Further, through the DAG, the outputs of multiple target expert models can be comprehensively considered to determine a forgery identification result according to a recognition decision logic, thereby ensuring the logicality and interpretability of the decision process. For example, if multiple target expert models all output high confidence, it can be determined that the to-be-identified image is forged, and if a few target expert models output high confidence, it can be determined that the to-be-identified image is not forged. In addition, according to the weights and rules in the DAG, a comprehensive confidence can be obtained, and it can be determined according to the comprehensive confidence whether the to-be-identified image is forged, and a corresponding forgery identification result can be output.
[0121] Further, the target expert models can be stored in the framework of the hybrid expert model, and the training process of the framework of the hybrid expert model includes:
[0122] (1) Before each expert model is put into the framework of the hybrid expert model, each expert model is individually pre-trained according to a specific data set. The specific data set is a region mask image, which can be used to strengthen the special ability of each expert model in a specific region. For example, if an expert model focuses on detecting the forgery of a hairline region, its training data set includes a large number of training samples labeled with the hairline region, so that the expert model can learn the specific forgery features of the hairline region.
[0123] (2) Training and optimization of the gating mechanism: The gating mechanism can be used to determine the target expert model to be activated according to the input information through a multi-layer perception neural network structure. The training process includes discrete sampling and gradient backpropagation, and an auxiliary loss function.
[0124] Wherein, the discrete sampling and gradient backpropagation are used to solve the problem that the discrete sampling in the gating mechanism cannot directly perform gradient descent. It can be solved by Top-k Gumbel-Softmax, where k is set to 2 to 3 by default. That is, during the training of the gating network of the gating mechanism, Gumbel noise is introduced to make the output of the gating network differentiable and relaxed approximation, so that the discrete operation of "selecting a target expert model" can be backpropagated. In the inference process, the top-k experts (i.e., target expert models) with the highest probability are directly selected, thereby ensuring that only a few most relevant experts are activated each time, and the calculation efficiency is maintained.
[0125] At the same time, the auxiliary loss function encourages load balancing among the expert models, that is, through an expert utilization rate variance penalty loss, it is prevented that the gating network always tends to activate a few popular expert models and ignore other expert models. The expert utilization rate variance penalty loss can be expressed by the following expression:
[0126] aux_loss = Var(expert_activation_counts) x 0.1;
[0127] wherein Var denotes variance, expert_activation_counts denotes the statistics of the number of times each expert model is activated, and 0.1 is a coefficient adjustable coefficient constant for controlling the weight of the loss term in the total loss. The purpose of the expression is to reduce the variance of the number of times the expert model is activated, so as to ensure that each expert model is used equally.
[0128] In addition, in order to ensure the alignment between the text embedding features and the target expert model, a supervised signal can be constructed through a manually annotated "expert-fake type" mapping table, for example, "hair artifact→Expert1" indicates that when the text embedding feature mentions a hair artifact, the expert model named Expert1 should be activated.
[0129] The alignment loss function is used to supervise the gating network to make the target expert distribution output by the gating network as close as possible to the ideal target expert distribution, and the alignment loss function can be expressed by the following expression:
[0130] gate_loss=CE(gate_logits,expert_label);
[0131] wherein CE denotes a classification loss, i.e., a cross-entropy loss function, for measuring the difference between two probability distributions, gate_logits denotes the output value of the gating mechanism, indicating the probability of each expert model being activated, and expert_label is used to represent the label value of the expert model, i.e., according to the description of the fake type in the current input text embedding feature, the target expert model to be activated is found from the mapping table.
[0132] (3) The training of the overall mixed expert model framework can adopt a phased strategy to ensure training stability and performance. In the first phase, the parameters of all pre-trained expert models are frozen, and only the parameters of the gating network are optimized / updated, so that the gating network learns how to make correct expert selection decisions according to the input.
[0133] In the second phase, all parameters are fine-tuned, the parameters of the expert models are unfrozen, and the parameters of the gating network and all expert models are optimized together using a very small learning rate, so that the overall mixed expert model framework is fine-tuned for optimal recognition accuracy.
[0134] The embodiment realizes accurate and visual determination of whether the to-be-recognized image is a fake image by comprehensively considering the outputs of multiple expert models and representing the decision logic through a directed acyclic graph, and at the same time makes the decision process interpretable.
[0135] In some embodiments, after generating the abnormal region heat map of the suspicious counterfeit region corresponding to the to-be-identified image based on the counterfeit region features, the method further comprises:
[0136] In the case that the spatial distribution of the abnormal region heat map is unreasonable, determining that the to-be-identified image is an adversarial attack sample;
[0137] Generating an alarm information based on the adversarial attack sample and the abnormal region heat map, and sending the alarm information to a terminal device.
[0138] Optionally, in the case that the spatial distribution of the abnormal region heat map is unreasonable, the embodiment determines that the to-be-identified image is an adversarial attack sample, for example, if the abnormal region heat map shows that the counterfeit probability is too uniform in the to-be-identified image, or there is a suspicious counterfeit region that does not conform to the conventional counterfeit features, indicating that the to-be-identified image is subjected to an adversarial attack to mislead the counterfeit identification.
[0139] In addition, when it is determined that the to-be-identified image is an adversarial attack sample, an alarm information is generated according to the adversarial attack sample and the abnormal region heat map, and the alarm information is sent to a terminal device, so that the security personnel or system administrator can be notified in time through the terminal device to take corresponding measures, such as further investigation or strengthening security protection.
[0140] The alarm information details the location of the suspicious counterfeit region, the counterfeit features, and the corresponding confidence in the to-be-identified image.
[0141] Further, the adversarial attack can be identified through an adversarial attack feature learning mechanism, that is, by analyzing a large number of adversarial attack samples, learning the feature patterns of the adversarial attack, and constantly updating the feature library of the adversarial attack, the identification ability of new adversarial attacks can be improved, and potential security threats can be prevented.
[0142] For example, it can be identified that the adversarial attack produces a small but critical pixel change in a specific region of the to-be-identified image to mislead the counterfeit identification convolutional neural network to perform counterfeit identification on the to-be-identified image.
[0143] Further, after detecting the adversarial attack sample and generating the alarm information, a real-time dynamic defense mechanism can also be started, which can adjust the parameters of the counterfeit identification convolutional neural network in real time to resist the ongoing adversarial attack.
[0144] For example, the detection sensitivity to a specific region can be temporarily increased, or the decision threshold of the counterfeit identification convolutional neural network can be adjusted, so that it is more difficult for the attacker to succeed.
[0145] The embodiment can identify an adversarial attack sample, discover and resist potential security threats in time, enhance security, generate and send alarm information, and enable relevant personnel to take measures quickly to reduce the occurrence of security incidents.
[0146] In some embodiments, the spatial distribution of the abnormal region heat map is reasonable and satisfies the following conditions:
[0147] The distribution of the high response region in the abnormal region heat map conforms to the distribution of the high response region corresponding to the preset forgery mode;
[0148] The target expert model to be activated is the same as the target expert model corresponding to the suspicious forgery region described in the text prompt information;
[0149] It is determined through frequency domain analysis that there is no abnormal high frequency component in the image to be identified.
[0150] Optionally, the spatial distribution of the abnormal region heat map in the embodiment needs to satisfy the following three conditions at the same time:
[0151] (1) The distribution of the high response region in the abnormal region heat map needs to conform to the distribution of the high response region corresponding to the preset forgery mode, that is, in the process of generating the abnormal region heat map, whether the distribution of the high probability region in the abnormal region heat map is reasonable is evaluated according to the known forgery mode, to ensure that the abnormal region heat map can accurately reflect the distribution of the forgery region in the image to be identified.
[0152] (2) The target expert model to be activated needs to be the same as the target expert model corresponding to the suspicious forgery region described in the text prompt information, to determine the accuracy of activating the target expert model. The text prompt information provides a detailed description of the suspicious forgery region, and the corresponding target expert model can be activated according to the description of the text prompt information to analyze the specific region in depth. If the target expert model activated according to the text embedding feature is inconsistent with the target expert model described in the text prompt information, it is easy to cause deviation of the analysis result and affect the accuracy of forgery identification.
[0153] (3) It is determined through frequency domain analysis that there is no abnormal high frequency component in the image to be identified. Frequency domain analysis is a way to detect possible forgery traces in an image. In the frequency domain, a forged image will produce an abnormal high frequency component. The high frequency component may correspond to sharpening, edge enhancement or other forgery operations in the image. Through frequency domain analysis, abnormal components can be identified and used as a basis for judging whether the image to be identified is a forged image. If there is an abnormal high frequency component in the image to be identified, it is considered that the image to be identified may be a forged image.
[0154] In addition, in the case where the spatial distribution of the abnormal region heat map satisfies any of the above conditions, it is determined as a forged image.
[0155] The embodiment realizes determination of whether the to-be-identified image is a fake image by determining the spatial distribution rationality of the abnormal area heat map, and improves the accuracy and reliability of fake image identification.
[0156] In order to effectively solve the problems of the conventional technology, such as easy omission of local fake information, untraceable of overall identification error, low accuracy of fake identification, and the like, significantly improve the accuracy of fake identification, and enhance the explainability of the determined fake identification result, an embodiment of a fake identification device based on local image understanding for implementing all or part of the content of the fake identification based on local image understanding is provided, which is described with reference to Figure 2 The fake identification device based on local image understanding specifically includes the following content:
[0157] The first processing module 10 is configured to receive a to-be-identified image, extract fake area features of the to-be-identified image through a fake identification convolutional neural network and a multi-scale attention mechanism, generate an abnormal area heat map of a suspicious fake area corresponding to the to-be-identified image based on the fake area features, and the tail of the fake identification convolutional neural network includes a deformable convolution block with a preset number of layers.
[0158] The second processing module 20 is configured to generate a structured text prompt information in a preset prompt information format based on the abnormal area heat map and the to-be-identified image in the case that the spatial distribution of the abnormal area heat map is reasonable, encode the text prompt information and the to-be-identified image to obtain a text embedding feature, and the text prompt information is used to describe abnormal features of at least one suspicious fake area and an abnormal confidence corresponding to the abnormal features.
[0159] The third processing module 30 is configured to activate a target expert model corresponding to the suspicious fake area based on the text embedding feature, process the corresponding suspicious fake area through the target expert model, obtain a fake confidence corresponding to each suspicious fake area, and determine a fake identification result of the to-be-identified image based on the fake confidence.
[0160] As can be known from the above description, the forgery identification device based on local image understanding provided by the embodiment of the application can receive an image to be identified innovatively, extract forgery region features of the image to be identified through a forgery identification convolutional neural network and a multi-scale attention mechanism, generate an abnormal region heat map of a suspicious forgery region corresponding to the image to be identified according to the forgery region features, wherein the tail of the forgery identification convolutional neural network includes a deformable convolution of a preset number of layers, and in the case that the spatial distribution of the abnormal region heat map is reasonable, generate a structured text prompt information according to the abnormal region heat map and the image to be identified in a preset prompt information format, encode the text prompt information and the image to be identified to obtain text embedding features, wherein the text prompt information is used to describe abnormal features of at least one suspicious forgery region and abnormal confidence corresponding to the abnormal features, activate a target expert model corresponding to the suspicious forgery region according to the text embedding features, process the corresponding suspicious forgery region through the target expert model to obtain a forgery confidence corresponding to each suspicious forgery region, and determine a forgery identification result of the image to be identified according to the forgery confidence. The potential forgery region in the image to be identified can be identified, the sensitivity to local forgery information is enhanced, the semantic information corresponding to the suspicious forgery region is enriched through the text prompt information, and the accuracy of forgery identification is improved. The method effectively solves the problems of the prior art, such as easy omission of local forgery information, untraceable errors in overall identification, and low accuracy of forgery identification, significantly improves the accuracy of forgery identification, and enhances the explainability of the determined forgery identification result.
[0161] From the hardware level, in order to effectively solve the problems of the prior art, such as easy omission of local forgery information, untraceable errors in overall identification, and low accuracy of forgery identification, significantly improve the accuracy of forgery identification, and enhance the explainability of the determined forgery identification result, the application provides an embodiment of an electronic device for implementing all or part of the contents of the forgery identification method based on local image understanding. The electronic device specifically includes the following contents:
[0162] A processor, a memory, a communications interface, and a bus; wherein the processor, the memory, and the communications interface complete mutual communication through the bus; the communications interface is used to realize information transmission between the forgery identification device based on local image understanding and related devices such as a core business system, a user terminal, and a related database; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, and the like, and the embodiment is not limited thereto. In the embodiment, the logic controller can be implemented by referring to the embodiment of the forgery identification method based on local image understanding and the embodiment of the forgery identification device based on local image understanding in the embodiment, the contents of which are incorporated herein, and the repeated parts will not be described herein.
[0163] It can be understood that the user terminal can include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. The smart wearable device can include smart glasses, a smart watch, a smart bracelet, etc.
[0164] In actual application, part of the forgery identification method based on local image understanding can be executed on the electronic device as described above, or all operations can be completed in the client device. Specifically, the selection can be made according to the processing capability of the client device and the restriction of the user's use scenario, etc. The present application does not limit this. If all operations are completed in the client device, the client device can further include a processor.
[0165] The client device described above can have a communication module (i.e., a communication unit) and can be communicatively connected with a remote server to realize data transmission with the server. The server can include a server of a task scheduling center side, and in other implementation scenarios, can also include a server of an intermediate platform, such as a server of a third-party server platform communicatively connected with the server of the task scheduling center. The server can include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.
[0166] Figure 3 A schematic block diagram of a system configuration of the electronic device 9600 of an embodiment of the present application is shown in FIG. 9. As shown in FIG. 9, the electronic device 9600 can include a central processor 9100 and a memory 9140; the memory 9140 is coupled to the central processor 9100. It is worth noting that the structure shown in FIG. 9 is exemplary; other types of structures can also be used to supplement or replace the structure to realize telecommunication functions or other functions. Figure 3 Figure 3 The structure shown in FIG. 9 is exemplary; other types of structures can also be used to supplement or replace the structure to realize telecommunication functions or other functions.
[0167] In an embodiment, the forgery identification method based on local image understanding can be integrated into the central processor 9100. The central processor 9100 can be configured to perform the following control:
[0168] Step S101: receiving an image to be identified, extracting a forgery region feature of the image to be identified through a forgery identification convolutional neural network and a multi-scale attention mechanism, generating an abnormal region heat map of a suspicious forgery region corresponding to the image to be identified based on the forgery region feature, and the tail of the forgery identification convolutional neural network includes a deformable convolution block of a preset number of layers;
[0169] Step S102: In the case that the spatial distribution of the abnormal region heat map is reasonable, generating a structured text prompt information according to the abnormal region heat map and the to-be-identified image based on a preset prompt information format, encoding the text prompt information and the to-be-identified image to obtain a text embedding feature, and the text prompt information is used to describe the abnormal features of at least one suspicious counterfeit region and the abnormal confidence corresponding to the abnormal features.
[0170] Step S103: Activating a target expert model corresponding to the suspicious counterfeit region based on the text embedding feature, processing the corresponding suspicious counterfeit region through the target expert model to obtain a counterfeit confidence corresponding to each suspicious counterfeit region, and determining a counterfeit identification result of the to-be-identified image based on the counterfeit confidence. As can be known from the above description, the electronic device provided in the embodiments of the present application receives a to-be-identified image innovatively, extracts counterfeit region features of the to-be-identified image through a counterfeit identification convolutional neural network and a multi-scale attention mechanism, generates an abnormal region heat map of suspicious counterfeit regions corresponding to the to-be-identified image according to the counterfeit region features, wherein the tail of the counterfeit identification convolutional neural network includes a deformable convolution of a preset number of layers, in the case that the spatial distribution of the abnormal region heat map is reasonable, generates a structured text prompt information according to the abnormal region heat map and the to-be-identified image based on a preset prompt information format, encodes the text prompt information and the to-be-identified image to obtain a text embedding feature, wherein the text prompt information is used to describe the abnormal features of at least one suspicious counterfeit region and the abnormal confidence corresponding to the abnormal features, activates a target expert model corresponding to the suspicious counterfeit region according to the text embedding feature, processes the corresponding suspicious counterfeit region through the target expert model to obtain a counterfeit confidence corresponding to each suspicious counterfeit region, and determines a counterfeit identification result of the to-be-identified image according to the counterfeit confidence. The potential counterfeit region in the to-be-identified image can be identified, the sensitivity to local counterfeit information is enhanced, the semantic information corresponding to the suspicious counterfeit region is enriched through the text prompt information, and the accuracy of counterfeit identification is improved. The method effectively solves the problems of the traditional technology, such as easy omission of local counterfeit information, untraceable of overall identification error, and low accuracy of counterfeit identification, significantly improves the accuracy of counterfeit identification, and enhances the explainability of the determined counterfeit identification result.
[0171] In another embodiment, the counterfeit identification device based on local image understanding can be configured separately from the central processor 9100, for example, the counterfeit identification device based on local image understanding can be configured as a chip connected with the central processor 9100, and the function of the counterfeit identification method based on local image understanding is realized through the control of the central processor.
[0172] As Figure 3As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily need to include these components. Figure 3 All components shown; in addition, the electronic device 9600 may also include Figure 3 For components not shown, please refer to existing technologies.
[0173] like Figure 3 As shown, the central processing unit 9100, sometimes also referred to as a controller or operating control, may include a microprocessor or other processor device and / or logic device, which receives inputs and controls the operation of various components of the electronic device 9600.
[0174] The memory 9140 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned failure-related information, and also store a program for executing that information. The central processing unit 9100 may execute the program stored in the memory 9140 to perform information storage or processing, etc.
[0175] Input unit 9120 provides input to central processing unit 9100. Input unit 9120 may be, for example, a keypad or touch input device. Power supply 9170 provides power to electronic device 9600. Display 9160 displays images and text. Display may be, for example, an LCD display, but is not limited thereto.
[0176] The memory 9140 can be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), a SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROMs. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 via the central processing unit 9100.
[0177] The memory 9140 can further include a data storage 9143 for storing data such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. A driver storage 9144 of the memory 9140 can include various drivers of the electronic device for communication functions and / or for performing other functions of the electronic device (e.g., a messaging application, a phonebook application, etc.).
[0178] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processor 9100 to provide input signals and receive output signals, as in the case of a conventional mobile communication terminal.
[0179] Based on different communication technologies, a plurality of communication modules 9110 can be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module, etc. The communication module 9110 (transmitter / receiver) is further coupled to the speaker 9131 and the microphone 9132 via the audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby implementing a conventional telecommunication function. The audio processor 9130 can include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 9130 is coupled to the central processor 9100, thereby enabling recording on the local device via the microphone 9132 and playing stored sound on the local device via the speaker 9131.
[0180] The embodiment of the present application further provides a computer readable storage medium capable of implementing all steps of the forgery identification method based on local image understanding in which the execution subject in the above-mentioned embodiment is a server or a client, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement all steps of the forgery identification method based on local image understanding in which the execution subject in the above-mentioned embodiment is a server or a client, for example, the processor implements the following steps when executing the computer program:
[0181] Step S101: receiving an image to be identified, extracting a forgery region feature of the image to be identified through a forgery identification convolutional neural network and a multi-scale attention mechanism, generating an abnormal region heat map of a suspicious forgery region corresponding to the image to be identified based on the forgery region feature, and the tail of the forgery identification convolutional neural network includes a deformable convolutional block with a preset number of layers;
[0182] Step S102: In the case that the spatial distribution of the abnormal region heat map is reasonable, generating a structured text prompt information according to the abnormal region heat map and the to-be-identified image based on a preset prompt information format, encoding the text prompt information and the to-be-identified image to obtain a text embedding feature, and the text prompt information is used to describe the abnormal features of at least one suspicious counterfeit region and the abnormal confidence corresponding to the abnormal features.
[0183] Step S103: Activating a target expert model corresponding to the suspicious counterfeit region based on the text embedding feature, processing the corresponding suspicious counterfeit region through the target expert model to obtain a counterfeit confidence corresponding to each suspicious counterfeit region, and determining a counterfeit identification result of the to-be-identified image based on the counterfeit confidence.
[0184] As can be seen from the above description, the computer readable storage medium provided by the embodiments of the present application receives the to-be-identified image innovatively, extracts the counterfeit region features of the to-be-identified image through the counterfeit identification convolutional neural network and the multi-scale attention mechanism, generates the abnormal region heat map of the suspicious counterfeit region corresponding to the to-be-identified image according to the counterfeit region features, wherein the tail of the counterfeit identification convolutional neural network includes a deformable convolution of a preset number of layers, in the case that the spatial distribution of the abnormal region heat map is reasonable, generates a structured text prompt information according to the abnormal region heat map and the to-be-identified image based on a preset prompt information format, encodes the text prompt information and the to-be-identified image to obtain a text embedding feature, wherein the text prompt information is used to describe the abnormal features of at least one suspicious counterfeit region and the abnormal confidence corresponding to the abnormal features, activates a target expert model corresponding to the suspicious counterfeit region according to the text embedding feature, processes the corresponding suspicious counterfeit region through the target expert model to obtain a counterfeit confidence corresponding to each suspicious counterfeit region, and determines a counterfeit identification result of the to-be-identified image according to the counterfeit confidence. The potential counterfeit region in the to-be-identified image can be identified, the sensitivity to local counterfeit information is enhanced, the semantic information corresponding to the suspicious counterfeit region is enriched through the text prompt information, and the accuracy of the counterfeit identification is improved. The method effectively solves the problems of the traditional technology, such as easy omission of local counterfeit information, untraceable of overall identification error, and low accuracy of counterfeit identification, significantly improves the accuracy of the counterfeit identification, and enhances the explainability of the determined counterfeit identification result.
[0185] The embodiments of the present application also provide a computer program product capable of implementing all steps of the local image understanding based counterfeit identification method in the above-mentioned embodiments, wherein the server or the client is the execution subject. The computer program / instruction is executed by the processor to implement the steps of the local image understanding based counterfeit identification method, for example, the computer program / instruction implements the following steps:
[0186] Step S101: receiving an image to be identified, extracting a counterfeit region feature of the image to be identified through a counterfeit identification convolutional neural network and a multi-scale attention mechanism, generating an abnormal region heat map of a suspicious counterfeit region corresponding to the image to be identified based on the counterfeit region feature, and the tail of the counterfeit identification convolutional neural network comprising a deformable convolutional block of a preset number of layers.
[0187] Step S102: in the case that the spatial distribution of the abnormal region heat map is reasonable, generating a structured text prompt information according to a preset prompt information format based on the abnormal region heat map and the image to be identified, encoding the text prompt information and the image to be identified to obtain a text embedding feature, and the text prompt information being used to describe an abnormal feature of at least one suspicious counterfeit region and an abnormal confidence corresponding to the abnormal feature.
[0188] Step S103: activating a target expert model corresponding to the suspicious counterfeit region based on the text embedding feature, processing the corresponding suspicious counterfeit region through the target expert model to obtain a counterfeit confidence corresponding to each suspicious counterfeit region, and determining a counterfeit identification result of the image to be identified based on the counterfeit confidence.
[0189] As can be seen from the above description, the computer program product provided by the embodiments of the present application receives an image to be identified, extracts a counterfeit region feature of the image to be identified through a counterfeit identification convolutional neural network and a multi-scale attention mechanism, generates an abnormal region heat map of a suspicious counterfeit region corresponding to the image to be identified according to the counterfeit region feature, wherein the tail of the counterfeit identification convolutional neural network comprises a deformable convolutional block of a preset number of layers, in the case that the spatial distribution of the abnormal region heat map is reasonable, generates a structured text prompt information according to a preset prompt information format based on the abnormal region heat map and the image to be identified, encodes the text prompt information and the image to be identified to obtain a text embedding feature, wherein the text prompt information is used to describe an abnormal feature of at least one suspicious counterfeit region and an abnormal confidence corresponding to the abnormal feature, activates a target expert model corresponding to the suspicious counterfeit region according to the text embedding feature, processes the corresponding suspicious counterfeit region through the target expert model to obtain a counterfeit confidence corresponding to each suspicious counterfeit region, and determines a counterfeit identification result of the image to be identified according to the counterfeit confidence. The potential counterfeit region in the image to be identified can be identified, the sensitivity to local counterfeit information is enhanced, the semantic information corresponding to the suspicious counterfeit region is enriched through the text prompt information, and the accuracy of counterfeit identification is improved. The method effectively solves the problems of the traditional technology, such as easy omission of local counterfeit information, untraceable of overall identification error, and low accuracy of counterfeit identification, significantly improves the accuracy of counterfeit identification, and enhances the explainability of determining the counterfeit identification result.
[0190] Those skilled in the art will appreciate that embodiments of the present application can be readily used as software, hardware, or a combination of software and hardware. In a software embodiment, the methods can be tangibly embodied in a machine-readable storage medium having stored thereon instructions that can be used to program a computer to perform any of the operations described herein. The software implementation can be for example, in the form of a computer program product which can include software agents or objects embedded in a computer readable storage medium. The computer readable storage medium can be a floppy disk, flexible disk, hard disk, USB (universal serial bus), RAM (random-access memory), flash memory, magnetic tape, or any other form of a computer readable storage medium.
[0191] The present application is described in relation to flowcharts and / or block diagrams of methods, apparatus (devices) and computer program products according to embodiments of the present application. It is understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowcharts and / or block diagrams block or blocks. Figure 1 one or more functions specified by one or more blocks Figure 1 one or more functions specified by one or more blocks.
[0192] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowcharts and / or block diagrams block or blocks. Figure 1 one or more functions specified by one or more blocks Figure 1 one or more functions specified by one or more blocks.
[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowcharts and / or block diagrams block or blocks. Figure 1 one or more functions specified by one or more blocks Figure 1 one or more functions specified by one or more blocks.
[0194] The principles and implementation modes of the present application are described in the specific embodiments. The above embodiments are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation modes and application ranges can be changed; in conclusion, the content of the present application should not be understood as limitation.
Claims
1. A forgery identification method based on local image understanding, characterized by, The method comprises: receiving an image to be identified, extracting a plurality of levels of features to be identified of the image to be identified based on a forgery identification convolutional neural network, the features to be identified including low-level features to be identified and high-level features to be identified, the low-level features to be identified being used to obtain abnormal pixels of the image to be identified, and the high-level features to be identified being used to obtain face layout and irrelevant relative position semantic information of the image to be identified; fusing the features to be identified corresponding to a preset number of layers in the forgery identification convolutional neural network according to a feature pyramid network structure through a multi-scale attention mechanism to obtain forgery region features, including: in the fusion process, retaining edge details for the low-level features to be identified and capturing semantic abnormal reactions for the high-level features to be identified, thereby obtaining forgery region features; generating an abnormal region heat map of a suspicious forgery region corresponding to the image to be identified based on the forgery region features, a tail part of the forgery identification convolutional neural network including deformable convolutional blocks of a preset number of layers for enhancing the positioning ability of irregular suspicious forgery regions, including: performing pixel-level probability mapping on the forgery region features, mapping a probability value of each pixel in the image to be identified belonging to a forgery region to a heat map, and coloring the heat map through a preset mapping relationship between colors and the probability value to obtain an abnormal region heat map; in the case that the spatial distribution of the abnormal region heat map is reasonable, generating a structured text prompt information according to a preset prompt information format based on the abnormal region heat map and the image to be identified, encoding the text prompt information and the image to be identified to obtain a text embedding feature, and the text prompt information being used to describe abnormal features of at least one suspicious forgery region and abnormal confidence corresponding to the abnormal features; activating a target expert model corresponding to the suspicious forgery region based on the text embedding feature, processing the corresponding suspicious forgery region through the target expert model to obtain a forgery confidence corresponding to each suspicious forgery region, and determining a forgery identification result of the image to be identified based on the forgery confidence.
2. The method of claim 1, wherein, The preset prompt information format is a triple structure, including a region type, an abnormal description, and an abnormal confidence; the generation of the structured text prompt information according to the preset prompt information format based on the abnormal region heat map and the image to be identified comprises: determining a region type corresponding to a suspicious forgery region and an abnormal confidence corresponding to the suspicious forgery region based on the abnormal region heat map; generating a corresponding abnormal description for each suspicious forgery region through a preset abnormal description vocabulary, and combining the region type, the abnormal description, and the abnormal confidence according to the preset prompt information format to obtain a text prompt information.
3. The method of claim 1, wherein, the determination of the forgery identification result of the image to be identified based on the forgery confidence comprises: receiving the forgery confidence output by each target expert model, and determining an identification decision logic based on the target expert model and the forgery confidence, wherein the identification decision logic is represented by a directed acyclic graph. The forgery confidence is processed based on the identification decision logic to obtain a forgery identification result.
4. The method of claim 1, wherein, After the abnormal region heat map of the suspicious forgery region corresponding to the image to be identified is generated based on the forgery region features, the method further includes: In a case where the spatial distribution of the abnormal region heat map is unreasonable, determining the image to be identified as an adversarial attack sample; Based on the adversarial attack sample and the abnormal region heat map, generating an alarm information, and sending the alarm information to a terminal device.
5. The method of claim 1, wherein, The spatial distribution of the abnormal region heat map is reasonable and meets the following conditions: The distribution of the high-response region in the abnormal region heat map meets the distribution of the high-response region corresponding to the preset forgery mode; The target expert model to be activated is the same as the target expert model corresponding to the suspicious forgery region described in the text prompt information; It is determined that there is no abnormal high-frequency component in the image to be identified through frequency domain analysis.
6. A forgery recognition apparatus based on local image understanding, characterized by, The device includes: A first processing module configured to receive an image to be identified, extract multi-level identification features of the image to be identified based on a forgery identification convolutional neural network, the identification features including low-level identification features and high-level identification features, the low-level identification features being used to obtain abnormal pixels of the image to be identified, and the high-level identification features being used to obtain face layout and irrelevant relative position semantic information of the image to be identified, and fuse the identification features corresponding to a preset number of layers in the forgery identification convolutional neural network according to a feature pyramid network structure through a multi-scale attention mechanism to obtain forgery region features, including: in the fusion process, retaining edge details for the low-level identification features and capturing semantic abnormal reactions for the high-level identification features, thereby obtaining forgery region features, generating an abnormal region heat map of a suspicious forgery region corresponding to the image to be identified based on the forgery region features, and the tail of the forgery identification convolutional neural network including deformable convolution blocks of a preset number of layers to enhance the positioning ability of irregular suspicious forgery regions, including: performing pixel-level probability mapping on the forgery region features, mapping a probability value of each pixel in the image to be identified belonging to a forgery region to a heat map, and coloring the heat map through a preset mapping relationship between colors and the probability value to obtain an abnormal region heat map; A second processing module configured to, in a case where the spatial distribution of the abnormal region heat map is reasonable, generate a structured text prompt information according to a preset prompt information format based on the abnormal region heat map and the image to be identified, encode the text prompt information and the image to be identified to obtain a text embedding feature, and the text prompt information being used to describe abnormal features of at least one suspicious forgery region and abnormal confidence corresponding to the abnormal features. A third processing module configured to activate a target expert model corresponding to the suspicious forgery region based on the text embedding feature, process the corresponding suspicious forgery region through the target expert model to obtain a forgery confidence of each suspicious forgery region, and determine a forgery identification result of the image to be identified based on the forgery confidence.
7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the local image understanding based forgery identification method of any one of claims 1 to 5 when executing the program.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the steps of the local image understanding based forgery identification method of any one of claims 1 to 5 when executed by the processor.
Citation Information
Patent Citations
Image forgery detection method and system based on multi-expert model decision
CN119762958A
CLIP-based double-prompt optimized few-sample industrial anomaly detection method
CN120672752A