Image forgery positioning method based on double-flow multi-scale feature fusion

Through the dual-stream multi-scale feature fusion method, RGB semantic features and noise features are extracted respectively, combined with global channel weighting and local space enhancement operations, and the coordinate attention fusion module encodes spatial information, the problems of insufficient adaptability of feature scales and global-local modeling imbalance in the existing technology are solved, and high-precision forged area positioning is achieved.

CN120580404AInactive Publication Date: 2025-09-02ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Patent Information

Application Number
CN202511088041.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient adaptability of feature scales, insufficient cross-domain feature fusion, and global-local modeling imbalance in image forgery positioning, resulting in insufficient accuracy in identifying and positioning forgery areas in complex scenarios.

Method used

The dual-stream multi-scale feature fusion method is adopted, and multi-scale features are extracted through RGB semantic branches and noise feature branches, combined with global channel weighting and local space enhancement operations, and the spatial information is encoded using the coordinate attention fusion module, and combined with Dice loss and Focal loss to optimize model parameters.

Benefits of technology

The accuracy of forgery positioning in image compression, blur and multifaceted scenes is improved, the model's robustness and adaptability to complex scenes is enhanced, and the local details and global structural contradictions of the forgery area are accurately positioned.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580404A_ABST
    Figure CN120580404A_ABST
Patent Text Reader

Abstract

The invention discloses an image forgery positioning method based on double-flow multi-scale feature fusion, and belongs to the technical field of image forgery positioning, and the method comprises the steps: respectively extracting multi-scale semantic features and noise features of a normalized RGB image through a double-branch structure; stacking the same-scale features to generate preliminary fusion features; multi-scale enhancement operation is executed, and enhancement features are generated by combining global channel weighting and local space enhancement; performing up-sampling splicing on the multi-scale enhanced features, encoding space information in horizontal and vertical directions through a coordinate attention fusion module, generating direction attention weights, and performing weighted fusion on the features; and generating a prediction mask based on the weighted fusion feature, and combining a Dice loss function and a Focal loss function to optimize the model. According to the method, semantic and noise information is comprehensively utilized, the multi-scale feature expression and spatial positioning capability is enhanced, and the counterfeit positioning precision and robustness in a complex scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image forgery positioning, and in particular relates to an image forgery positioning method based on dual-stream multi-scale feature fusion. Background Art

[0002] Image forgery detection technology, a core area of ​​digital content security, aims to accurately identify tampered or forged areas in an image through algorithms. It holds irreplaceable value in scenarios such as media content review and biometric security. With the advancement of generative technologies like generative adversarial networks (GANs) and diffusion models (DMs), forged images present two major challenges: First, the visual consistency between forged and authentic areas has significantly improved, making traditional detection methods based on single-domain features (such as the RGB spatial domain) difficult to detect subtle signs of tampering; second, post-processing operations (such as compression, blurring, and noise injection) further obscure forgery clues, resulting in insufficient generalization across scenarios.

[0003] While existing deep learning methods attempt to improve accuracy by incorporating cross-domain feature fusion, such as a dual-domain face forgery localization scheme that extracts frequency-domain features using a multi-band feature sensor and combines it with a spatial-frequency interaction module, they still suffer from fundamental limitations. First, feature scale adaptability is insufficient. A fixed-scale feature fusion strategy cannot simultaneously capture the multi-scale differences in traces of different forgery techniques (such as splicing and generative forgeries), resulting in failure to identify hybrid forgery attacks in biometric security scenarios. Second, there is an imbalance in global-local modeling. Relying on a purely convolutional architecture to process frequency-domain features lacks the ability to model long-range semantic relationships, making it difficult to distinguish global structural contradictions between forged and authentic regions. This is particularly true in multi-face scenarios, where the inability to establish cross-regional associations leads to localization bias. These shortcomings stem from the fact that existing technologies fail to fully exploit the complementary nature of multi-domain image information and fail to address the problem of co-optimization between multi-scale feature representation and spatial localization, becoming key bottlenecks hindering the practical application of forgery localization technology. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes an image forgery localization method based on dual-stream multi-scale feature fusion to solve the problems existing in the above-mentioned prior art.

[0005] In a first aspect, to achieve the above-mentioned objectives, the present invention provides an image forgery localization method based on dual-stream multi-scale feature fusion, comprising the following steps:

[0006] Input the RGB image to be detected and perform size normalization;

[0007] Extract multi-domain features separately through the dual-branch structure:

[0008] The RGB semantic branch uses the Transformer encoder to extract multi-scale semantic features;

[0009] The noise feature branch uses a convolutional network to extract multi-scale noise features;

[0010] The semantic features and noise features of the same scale are stacked and fused to generate preliminary fusion features;

[0011] Perform multi-scale enhancement operations on the preliminary fusion features, combining global channel weighting and local spatial enhancement to generate enhanced features;

[0012] The multi-scale enhanced features are upsampled to the highest resolution and spliced, and the spatial information in the horizontal and vertical directions is encoded through the coordinate attention fusion module to generate directional attention weights and weighted fusion features;

[0013] The tampered area prediction mask is generated based on the weighted fusion features, and the model parameters are optimized by combining the Dice loss function and the Focal loss function.

[0014] Optionally, the multi-domain feature extraction process includes:

[0015] The RGB semantic branch uses a Segformer encoder consisting of four Transformer Blocks, and outputs multi-scale semantic features with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input image respectively;

[0016] The noise feature branch extracts noise through a Bayer filter, and then outputs multi-scale noise features aligned with the resolution of the semantic branch through a four-layer convolutional network.

[0017] Optionally, the multi-scale enhancement operation includes:

[0018] Perform global average pooling on the preliminary fusion features, generate a channel weight vector through two 1×1 convolutions and ReLU activations, and multiply it element-wise with the original features;

[0019] Perform spatial enhancement operations on the channel enhancement results: fuse the maximum pooling and average pooling features, extract the spatial weights through 7×7 convolution, and generate spatial enhancement features through Sigmoid activation.

[0020] Optionally, the operation of the coordinate attention fusion module includes:

[0021] Concatenate multi-scale enhanced features into initial features;

[0022] Perform horizontal average pooling and vertical average pooling on the initial features to generate directional feature vectors;

[0023] The directional feature vectors are concatenated and then subjected to 1×1 convolution, batch normalization, and Sigmoid activation to generate the encoding tensor;

[0024] Separate the horizontal attention map and the vertical attention map from the encoding tensor and multiply them element-wise with the initial features to generate the final fused features.

[0025] Optionally, the process of generating the tampered area prediction mask includes:

[0026] Input the weighted fusion features output by the coordinate attention fusion module into the prediction network;

[0027] Output the tampering probability map of each pixel, and generate a positioning mask by threshold binarization of the probability map.

[0028] Optionally, the calculation process of the joint loss function includes:

[0029] The Dice loss function is used to optimize the intersection-over-union ratio between the predicted mask and the true mask;

[0030] The Focal loss function is used to adjust the weights of positive and negative samples and enhance the learning weights of pixels in the tampered area.

[0031] In a second aspect, the present invention further provides an image forgery location system based on dual-stream multi-scale feature fusion, which is used to implement an image forgery location method based on dual-stream multi-scale feature fusion. The system includes:

[0032] Image preprocessing module, used to input the RGB image to be detected and perform size normalization;

[0033] The dual-branch feature extraction module is used to extract multi-domain features separately, including:

[0034] RGB semantic branch unit, which uses Transformer encoder to extract multi-scale semantic features;

[0035] The noise feature branch unit uses a convolutional network to extract multi-scale noise features;

[0036] The feature stacking and fusion module is used to stack semantic features and noise features of the same scale to generate preliminary fusion features;

[0037] Multi-scale enhancement module, used to perform global channel weighting and local spatial enhancement operations on the preliminary fusion features to generate enhanced features;

[0038] The coordinate attention fusion module is used to upsample and splice the multi-scale enhanced features, encode the horizontal and vertical spatial information, generate directional attention weights and weighted fusion features;

[0039] A tampered region prediction module is used to generate a tampered region prediction mask based on weighted fusion features;

[0040] The loss optimization module is used to optimize model parameters by combining the Dice loss function and the Focal loss function.

[0041] Optionally, the RGB semantic branch unit includes a Segformer encoder, which includes four TransformerBlocks, and outputs multi-scale semantic features with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input image respectively;

[0042] The noise feature branch unit includes a Bayer filter unit and a four-layer convolutional network unit, wherein:

[0043] The Bayer filter unit extracts the noise characteristics of the input image;

[0044] A four-layer convolutional network unit processes the noise features and outputs multi-scale noise features aligned with the resolution of the semantic branch.

[0045] In a third aspect, the present invention further provides a computer terminal device, comprising:

[0046] one or more processors;

[0047] a memory, coupled to the processor, for storing one or more programs;

[0048] When the one or more programs are executed by the one or more processors, the one or more processors implement, for example, an image forgery localization method based on dual-stream multi-scale feature fusion.

[0049] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements an image forgery positioning method based on dual-stream multi-scale feature fusion.

[0050] Compared with the prior art, the present invention has the following advantages and technical effects:

[0051] This paper proposes a method for localizing forged images based on dual-stream multi-scale feature fusion. This method uses a dual-branch design to extract RGB semantic features and noise features, respectively, and combines multi-scale feature stacking and fusion to capture semantic inconsistencies and high-frequency traces in tampered areas. It also utilizes global channel weighting and local spatial enhancement operations to enhance multi-scale feature expression and improve feature robustness in complex scenarios. A coordinate attention fusion module encodes horizontal and vertical spatial information, achieving spatial alignment and weighting of cross-scale features to accurately locate local details and global structural inconsistencies in forged areas. Finally, the Dice loss and Focal loss are combined to optimize model parameters, effectively addressing the imbalance between positive and negative samples. Ultimately, the method achieves high-precision forgery localization in image compression, blurring, and multi-face scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings, which constitute part of the present invention, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0053] Figure 1 1. A schematic diagram of a flow chart of an image forgery localization method based on dual-stream multi-scale feature fusion according to an embodiment of the present invention;

[0054] Figure 2 This is a schematic diagram of a method flow of a multi-domain feature extraction module according to an embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram of a method flow of a multi-scale feature enhancement module according to an embodiment of the present invention;

[0056] Figure 4 Schematic diagram of the method flow of the coordinate attention fusion module of an embodiment of the present invention. DETAILED DESCRIPTION

[0057] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0058] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0059] Example 1

[0060] This embodiment provides an image forgery detection method based on dual-stream multi-scale feature fusion, including:

[0061] Input the RGB image to be detected and perform size normalization;

[0062] Extract multi-domain features separately through the dual-branch structure:

[0063] The RGB semantic branch uses the Transformer encoder to extract multi-scale semantic features;

[0064] The noise feature branch uses a convolutional network to extract multi-scale noise features;

[0065] The semantic features and noise features of the same scale are stacked and fused to generate preliminary fusion features;

[0066] Perform multi-scale enhancement operations on the preliminary fusion features, combining global channel weighting and local spatial enhancement to generate enhanced features;

[0067] The multi-scale enhanced features are upsampled to the highest resolution and spliced, and the spatial information in the horizontal and vertical directions is encoded through the coordinate attention fusion module to generate directional attention weights and weighted fusion features;

[0068] The tampered area prediction mask is generated based on the weighted fusion features, and the model parameters are optimized by combining the Dice loss function and the Focal loss function.

[0069] Specifically, Figure 1 The specific process of an image forgery localization method with dual-stream multi-scale feature fusion is demonstrated. Starting from the input RGB image and noise features, multi-scale features are extracted respectively through a dual-branch structure, and feature modeling is performed using Transformer and convolutional networks. Subsequently, features are gradually fused through multi-scale feature splicing, multi-view feature fusion module and coordinate attention fusion module, and finally a prediction mask of the tampered area is generated. The present invention proposes an image forgery localization method based on dual-branch feature extraction and multi-view fusion, aiming to solve the problems of insufficient adaptability of existing forged image localization technology to complex scenes and low feature fusion efficiency. The method captures the deep pattern of the tampered area and achieves efficient positioning through multi-domain feature extraction, hybrid modeling, multi-view fusion, coordinate attention fusion and precise prediction, combined with the semantic information of the RGB image and the detailed information of the noise features. The method processing process includes five core modules:

[0070] Multi-domain feature extraction module: extracts multi-scale features from RGB images and noise features, providing rich spatial and noise information for subsequent modeling.

[0071] Hybrid modeling module: Transformer and convolutional network are used to process RGB and noise features respectively to generate highly discriminative multi-scale feature representations.

[0072] Multi-view feature fusion module: By fusing the multi-scale features of RGB and noise, the complementarity of RGB and noise branches is enhanced.

[0073] Coordinate attention fusion module: further fuses multi-scale features based on the coordinate attention mechanism to highlight the spatial location of the tampered area.

[0074] Prediction module: Generates the prediction mask of the tampered area based on the fused features to achieve high-precision positioning.

[0075] As an implementation method in this embodiment, the multi-domain feature extraction process includes:

[0076] The RGB semantic branch uses a Segformer encoder consisting of four Transformer Blocks, and outputs multi-scale semantic features with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input image respectively;

[0077] The noise feature branch extracts noise through a Bayer filter, and then outputs multi-scale noise features aligned with the resolution of the semantic branch through a four-layer convolutional network.

[0078] Specifically, refer to Figure 2 ,The method flow of the multi-domain feature extraction module includes:

[0079] 1. RGB feature branch:

[0080] Input: Input image , resized to a uniform size and normalized.

[0081] Processing: Use pre-trained Segformer as encoder to extract RGB features, including four TransformerBlocks, to gradually generate multi-scale features , , , These features correspond to 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image resolution, respectively, and are used to capture high-level semantic information and other contextual features of the image.

[0082] Output: multi-scale RGB features in is the batch size, is the feature dimension, The resolution of each scale.

[0083] 2. Noise feature branch:

[0084] Input: Input image , resized to a uniform size and normalized.

[0085] Processing: Bayer filter (BayerConv) is used to extract noise features to capture forgery traces in the image. Subsequently, a convolutional network (containing four Conv Blocks) is used to process the noise features to generate multi-scale features. , , , , the resolution corresponds to the RGB features.

[0086] Output: multi-scale RGB features in is the batch size, is the feature dimension, The resolution of each scale.

[0087] Specifically, the hybrid modeling module:

[0088] Input: RGB features and noise characteristics , .

[0089] Processing: Stack the features after each block processing to obtain the fusion features of RGB and noise, stack the features of the corresponding scales, and obtain the multi-scale simple fusion features.

[0090] Output: The fused multi-scale features are , .

[0091] As an implementation manner in this embodiment, the multi-scale enhancement operation includes:

[0092] Perform global average pooling on the preliminary fusion features, generate a channel weight vector through two 1×1 convolutions and ReLU activations, and multiply it element-wise with the original features;

[0093] Perform spatial enhancement operations on the channel enhancement results: fuse the maximum pooling and average pooling features, extract the spatial weights through 7×7 convolution, and generate spatial enhancement features through Sigmoid activation.

[0094] Specifically, the multi-scale feature enhancement module includes:

[0095] High-level features have a larger receptive field and stronger ability to represent semantic information, but are weaker in representing geometric information (i.e., local texture details). On the contrary, low-level features have a smaller receptive field and are good at representing geometric details, but have limited ability to represent semantic information. In order to make full use of the information of features at different levels, a multi-scale feature enhancement module is designed. Figure 3 Schematic diagram of the method flow of the multi-scale feature enhancement module:

[0096] Input: multi-scale RGB features and multi-scale noise characteristics ,in Represents four scales. These features come from the hybrid modeling modules of the RGB branch and the noise branch respectively.

[0097] Processing: Feature enhancement and transformation: Apply global average pooling to generate channel description vectors:

[0098]

[0099] Will Channel transformation is performed through the 1x1 convolution layer, and the dimension is reduced to , and then nonlinearity is introduced through the ReLU activation function:

[0100]

[0101] Restore the number of channels to , and apply ReLU activation:

[0102]

[0103] For the other branch, channel fusion of local perspective is applied:

[0104]

[0105] Will With the original Perform element-wise multiplication to generate preliminary enhanced features:

[0106]

[0107] in, Represents the Sigmoid activation function.

[0108] Spatial enhancement: In order to enhance local details and suppress irrelevant areas, spatial enhancement based on CBAM blocks is adopted:

[0109]

[0110] in, Represents the Sigmoid activation function. Represent the features after maximum pooling and average pooling respectively. Inter-enhancement extracts local spatial information through pooling operation and convolution to highlight the detailed features of the tampered area.

[0111] Output: Enhanced multi-scale fusion features , used for subsequent coordinate attention enhancement, the fusion feature combines high-level semantic information and low-level geometric detail information, significantly improving the model's ability to detect tampered areas.

[0112] As an implementation method in this embodiment, the operation of the coordinate attention fusion module includes:

[0113] Concatenate multi-scale enhanced features into initial features;

[0114] Perform horizontal average pooling and vertical average pooling on the initial features to generate directional feature vectors;

[0115] The directional feature vectors are concatenated and then subjected to 1×1 convolution, batch normalization, and Sigmoid activation to generate the encoding tensor;

[0116] Separate the horizontal attention map and the vertical attention map from the encoding tensor and multiply them element-wise with the initial features to generate the final fused features.

[0117] Specifically, the coordinate attention fusion module includes:

[0118] The Coordinate Attention Fusion Module (CAF) aims to enhance the localization capability of multi-scale features by encoding spatial information in the horizontal and vertical directions. This module is particularly suitable for image forgery localization tasks. By introducing the coordinate attention mechanism, it can effectively capture the spatial distribution characteristics of the tampered area, making up for the shortcomings of the traditional attention mechanism in spatial localization, thereby improving the model's detection accuracy for subtle tampering in complex scenes. The method flow of the Coordinate Attention Fusion Module (CAF) is as follows: Figure 4 As shown, the process includes:

[0119] Input: Features after multi-scale enhancement ;

[0120] Processing: First, the features of the four scales , spliced ​​according to the channel dimension to generate initial features:

[0121]

[0122] The low-resolution features are upsampled to the highest resolution through upsampling operation.

[0123] Horizontal and vertical average pooling: Horizontal average pooling (HAP) and vertical average pooling (VAP) are applied to aggregate features respectively:

[0124]

[0125]

[0126] in and Represent the average pooling operations along the horizontal and vertical directions, respectively, for extracting spatial information.

[0127] Spatial information encoding: After concatenating the horizontal and vertical pooling results, the spatial information is encoded through 1x1 convolution and batch normalization (BN) to generate a tensor :

[0128]

[0129] in, Represents the concatenation operation and represents the Sigmoid activation function.

[0130] Attention weight generation:

[0131] Generating horizontal attention maps :

[0132]

[0133] Generating vertical attention maps :

[0134]

[0135] Final feature fusion:

[0136]

[0137] Output: strong fusion features Combining the spatial location information of multi-scale features significantly improves the model's ability to locate tampered areas, providing high-quality input for subsequent prediction tasks.

[0138] As an implementation method in this embodiment, the process of generating the tampered area prediction mask includes:

[0139] Input the weighted fusion features output by the coordinate attention fusion module into the prediction network;

[0140] Output the tampering probability map of each pixel, and generate a positioning mask by threshold binarization of the probability map.

[0141] As an implementation method in this embodiment, the calculation process of the joint loss function includes:

[0142] The Dice loss function is used to optimize the intersection-over-union ratio between the predicted mask and the true mask;

[0143] The Focal loss function is used to adjust the weights of positive and negative samples and enhance the learning weights of pixels in the tampered area.

[0144] Specifically, the prediction module includes:

[0145] Loss Function: By combining Dice loss and Focal loss, this approach addresses the imbalance between positive and negative samples in semantic segmentation tasks while improving the model's ability to learn difficult samples. Dice loss improves the intersection over union (IoU) of segmented regions, while Focal loss emphasizes the model's focus on tampered regions (positive samples), ensuring the model can accurately locate tampered areas even in complex scenarios.

[0146] Input: predicted label Represents the predicted probability of each pixel, the true label , represents the true label of each pixel of the sample image (0 represents the non-tampered area, 1 represents the tampered area), and the shape is same.

[0147] Processing: Dice loss is often used in semantic segmentation tasks to optimize the intersection-over-union ratio between the predicted mask and the true mask, and is defined as follows:

[0148]

[0149] Among them, are the predicted label and true label of each pixel in the sample image, Represents the sum of all pixels. Dice loss improves segmentation accuracy by minimizing the difference between the predicted and true masks.

[0150] Focal loss calculation: In order to solve the problem of unbalanced distribution of positive and negative samples, Focal loss is used to enhance the model's learning ability for difficult samples (such as tampered areas). It is defined as follows:

[0151]

[0152] in, and is a hyperparameter used to adjust the weight of the loss, Control the balance of positive and negative samples, Make the model pay more attention to difficult samples. Focal loss enhances the model's ability to detect tampered areas by reducing the weight of easy-to-classify samples. The overall loss function is:

[0153]

[0154] Based on this, the present invention provides an image forgery localization method based on dual-stream multi-scale feature fusion. This method uses a dual-branch design to extract RGB semantic features and noise features respectively, and combines multi-scale feature stacking and fusion to capture semantic inconsistencies and high-frequency traces in tampered areas. It also uses global channel weighting and local spatial enhancement operations to enhance multi-scale feature expression and improve feature robustness in complex scenarios. A coordinate attention fusion module is used to encode horizontal and vertical spatial information, achieving spatial alignment and weighting of cross-scale features, accurately locating local details and global structural contradictions in forged areas. Finally, the Dice loss and Focal loss are combined to optimize model parameters, effectively addressing the imbalance between positive and negative samples. Ultimately, high-precision forgery localization is achieved in image compression, blurring, and multi-face scenarios.

[0155] The beneficial effects of the present invention are mainly reflected in: through dual-branch design and hybrid modeling, the semantic information of RGB images and the forgery traces of noise features are comprehensively utilized to improve the feature representation ability; through the multi-view feature fusion module and the coordinate attention fusion module, the multi-scale feature interaction is optimized to significantly improve the positioning accuracy of the tampered area; it is suitable for complex scenes and distorted images, and shows strong robustness and adaptability.

[0156] Example 2

[0157] In this embodiment, a computer terminal device is provided, including:

[0158] one or more processors;

[0159] a memory, coupled to the processor, for storing one or more programs;

[0160] When the one or more programs are executed by the one or more processors, the one or more processors implement the methods in the above embodiments.

[0161] In this embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the method in the above embodiment is implemented.

[0162] In this embodiment, an electronic device is further provided, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute the method in the above embodiment.

[0163] The above program can be executed in a processor or stored in a memory (or computer-readable medium). Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0164] These computer programs can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one or more blocks can be implemented by different modules corresponding to different steps.

[0165] This embodiment provides such a device or system. The system is called an image forgery location system based on dual-stream multi-scale feature fusion, and includes:

[0166] Image preprocessing module, used to input the RGB image to be detected and perform size normalization;

[0167] The dual-branch feature extraction module is used to extract multi-domain features separately, including:

[0168] RGB semantic branch unit, which uses Transformer encoder to extract multi-scale semantic features;

[0169] The noise feature branch unit uses a convolutional network to extract multi-scale noise features;

[0170] The feature stacking and fusion module is used to stack semantic features and noise features of the same scale to generate preliminary fusion features;

[0171] Multi-scale enhancement module, used to perform global channel weighting and local spatial enhancement operations on the preliminary fusion features to generate enhanced features;

[0172] The coordinate attention fusion module is used to upsample and splice the multi-scale enhanced features, encode the horizontal and vertical spatial information, generate directional attention weights and weighted fusion features;

[0173] A tampered region prediction module is used to generate a tampered region prediction mask based on weighted fusion features;

[0174] The loss optimization module is used to optimize model parameters by combining the Dice loss function and the Focal loss function.

[0175] As an implementation method in this embodiment, the RGB semantic branch unit includes a Segformer encoder, which includes four Transformer Blocks, and outputs multi-scale semantic features with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input image respectively;

[0176] The noise feature branch unit includes a Bayer filter unit and a four-layer convolutional network unit, wherein:

[0177] The Bayer filter unit extracts the noise characteristics of the input image;

[0178] A four-layer convolutional network unit processes the noise features and outputs multi-scale noise features aligned with the resolution of the semantic branch.

[0179] As an implementation manner in this embodiment, the multi-scale enhancement module includes:

[0180] The channel weighting unit is used to perform global average pooling on the preliminary fusion features. After two 1×1 convolutions and ReLU activations, the channel weight vector is generated and multiplied element-wise with the original features.

[0181] The spatial enhancement unit is used to fuse the maximum pooling and average pooling features of the channel weighted results, extract the spatial weights through 7×7 convolution, and generate spatial enhancement features through Sigmoid activation.

[0182] As an implementation method in this embodiment, the coordinate attention fusion module includes:

[0183] Feature splicing unit, used to upsample multi-scale enhanced features to the highest resolution and splice them into initial features;

[0184] Directional pooling unit, used to perform horizontal average pooling and vertical average pooling on the initial features to generate a directional feature vector;

[0185] The spatial encoding unit is used to concatenate the directional feature vectors and generate the encoding tensor through 1×1 convolution, batch normalization and Sigmoid activation;

[0186] The attention weighting unit is used to separate the horizontal attention map and the vertical attention map from the encoding tensor, and multiply them element-wise with the initial features to generate the final fused features.

[0187] As an implementation method of this embodiment, the tampering area prediction module includes:

[0188] The probability map generation unit is used to input the weighted fusion features into the prediction network and output the tampering probability map of each pixel;

[0189] The mask generation unit is used to perform threshold binarization processing on the probability map to generate a positioning mask.

[0190] As an implementation method in this embodiment, the loss optimization module includes:

[0191] Dice loss calculation unit, used to optimize the intersection-over-union ratio between the predicted mask and the true mask;

[0192] Focal loss calculation unit, used to adjust the weights of positive and negative samples to enhance the learning weights of pixels in the tampered area.

[0193] The system or device is used to implement the functions of the method in the above-mentioned embodiment. Each module in the system or device corresponds to each step in the method, which has been explained in the method and will not be repeated here.

[0194] Through the above implementation, the problem of image forgery positioning based on dual-stream multi-scale feature fusion in the related art is solved, thereby ensuring that the problems existing in the existing technology are solved.

[0195] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for image forgery localization based on dual-stream multi-scale feature fusion, characterized in that: The following steps are involved: Input the RGB image to be detected and perform size normalization; Extract multi-domain features separately through the dual-branch structure: The RGB semantic branch uses the Transformer encoder to extract multi-scale semantic features; The noise feature branch uses a convolutional network to extract multi-scale noise features; The semantic features and noise features of the same scale are stacked and fused to generate preliminary fusion features; Perform multi-scale enhancement operations on the preliminary fusion features, combining global channel weighting and local spatial enhancement to generate enhanced features; The multi-scale enhanced features are upsampled to the highest resolution and spliced, and the spatial information in the horizontal and vertical directions is encoded through the coordinate attention fusion module to generate directional attention weights and weighted fusion features; The tampered area prediction mask is generated based on the weighted fusion features, and the model parameters are optimized by combining the Dice loss function and the Focal loss function.

2. The method according to claim 1, characterized in that The multi-domain feature extraction process includes: The RGB semantic branch uses a Segformer encoder consisting of four Transformer Blocks, and outputs multi-scale semantic features with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input image respectively; The noise feature branch extracts noise through a Bayer filter, and then outputs multi-scale noise features aligned with the resolution of the semantic branch through a four-layer convolutional network.

3. The method according to claim 1, characterized in that The multi-scale enhancement operation includes: Perform global average pooling on the preliminary fusion features, generate a channel weight vector through two 1×1 convolutions and ReLU activations, and multiply it element-wise with the original features; Perform spatial enhancement operations on the channel enhancement results: fuse the maximum pooling and average pooling features, extract the spatial weights through 7×7 convolution, and generate spatial enhancement features through Sigmoid activation.

4. The method according to claim 1, wherein The operations of the coordinate attention fusion module include: Concatenate multi-scale enhanced features into initial features; Perform horizontal average pooling and vertical average pooling on the initial features to generate directional feature vectors; The directional feature vectors are concatenated and then subjected to 1×1 convolution, batch normalization, and Sigmoid activation to generate the encoding tensor; Separate the horizontal attention map and the vertical attention map from the encoding tensor and multiply them element-wise with the initial features to generate the final fused features.

5. The method according to claim 1, wherein The process of generating the tampered area prediction mask includes: Input the weighted fusion features output by the coordinate attention fusion module into the prediction network; Output the tampering probability map of each pixel, and generate a positioning mask by threshold binarization of the probability map.

6. The method according to claim 1, characterized in that The calculation process of the joint loss function includes: The Dice loss function is used to optimize the intersection-over-union ratio between the predicted mask and the true mask; The Focal loss function is used to adjust the weights of positive and negative samples and enhance the learning weights of pixels in the tampered area.

7. An image forgery location system based on dual-stream multi-scale feature fusion, characterized by: The system comprises: Image preprocessing module, used to input the RGB image to be detected and perform size normalization; The dual-branch feature extraction module is used to extract multi-domain features separately, including: RGB semantic branch unit, which uses Transformer encoder to extract multi-scale semantic features; The noise feature branch unit uses a convolutional network to extract multi-scale noise features; The feature stacking and fusion module is used to stack semantic features and noise features of the same scale to generate preliminary fusion features; Multi-scale enhancement module, used to perform global channel weighting and local spatial enhancement operations on the preliminary fusion features to generate enhanced features; The coordinate attention fusion module is used to upsample and splice the multi-scale enhanced features, encode the horizontal and vertical spatial information, generate directional attention weights and weighted fusion features; A tampered region prediction module is used to generate a tampered region prediction mask based on weighted fusion features; The loss optimization module is used to optimize model parameters by combining the Dice loss function and the Focal loss function.

8. The system according to claim 7, characterized in that The RGB semantic branch unit includes a Segformer encoder, which contains four Transformer Blocks, and outputs multi-scale semantic features with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input image respectively; The noise feature branch unit includes a Bayer filter unit and a four-layer convolutional network unit, wherein: The Bayer filter unit extracts the noise characteristics of the input image; A four-layer convolutional network unit processes the noise features and outputs multi-scale noise features aligned with the resolution of the semantic branch.

9. A computer terminal device, characterized in that: include: one or more processors; a memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image forgery localization method based on dual-stream multi-scale feature fusion according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image forgery localization method based on dual-stream multi-scale feature fusion according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Double-flow U-Net image tampering detection network system and image tampering detection method thereof

    CN114998261A

  • Image tampering detection and positioning method based on multi-scale supervised contrast learning

    CN117541571A

Cited By

  • Multi-person three-dimensional human body posture and spatial positioning combined reconstruction method and system

    CN121527286A

  • Casting body sheet image segmentation method and system based on dynamic noise filtering and detail enhancement

    CN121564352A

  • A method and system for image segmentation of cast thin sections based on dynamic noise filtering and detail enhancement

    CN121564352B