A PRNU anonymization method based on multi-scale and hierarchical feature fusion

By combining a multi-scale dynamic Mamba U-shaped network with deep learning and learnable convolutional layers, the problem of insufficient visual quality and anonymization effect of existing PRNU anonymization methods in image processing is solved, and high-quality PRNU anonymized image generation is achieved.

CN120599058BActive Publication Date: 2025-10-28QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511086550.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-28
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing PRNU anonymization methods struggle to generate visually high-quality de-PRNU images when dealing with images that have many elements, high detail, and complex structures. Furthermore, traditional methods do not perform well in anonymizing images where the PRNU is unknown.

Method used

A PRNU anonymization method based on multi-scale and hierarchical feature fusion is adopted. It utilizes a multi-scale dynamic Mamba U-shaped network (MD-MUnet) combined with deep learning and learnable convolutional layers. Through multi-scale adaptive convolution, Mamba architecture and attention mechanism, the weights are dynamically adjusted to generate images with visual quality close to the target image and PRNU removed.

Benefits of technology

It improves image anonymization and visual quality, effectively processes multi-scale images, and generates high-quality anonymized images even when PRNU is unknown, thus overcoming the limitations of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599058B_ABST
    Figure CN120599058B_ABST
Patent Text Reader

Abstract

This invention discloses a PRNU anonymization method based on multi-scale and hierarchical feature fusion, relating to the field of information security technology. The method includes the following steps: S1: First, input a target image, then input the image into a deep learning-based PRNU extraction network to extract PRNU noise corresponding to the target image; S2: Randomly generate a noise tensor with the same width and height as the target image, and input this tensor into a multi-scale dynamic Mamba U-shaped network architecture, which mainly consists of a multi-scale dynamic Unet and Mamba modules. The technical problem this invention aims to solve is to provide a PRNU anonymization method based on multi-scale and hierarchical feature fusion that weakens the physical association between the image and the device, while preventing malicious identity tracking and ensuring the practical value of the image in multimedia, medical, and judicial scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information security technology, and more specifically, to a PRNU anonymization method based on multi-scale and hierarchical feature fusion. Background Art

[0002] Photo Response Non-Uniformity (PRNU) is generated based on the inherent noise characteristics of image sensors, much like a unique "biological signature" of a digital image. Extracting the PRNU can provide technical evidence for image attribution, aiding in tracing the physical source. However, this method of device source identification through PRNU extraction also raises the issue of digital image privacy protection. Therefore, methods for anonymizing PRNUs have attracted considerable attention.

[0003] Because PRNUs originate from the random noise distribution inherent in the physical characteristics of sensors, their generation mechanism is deeply coupled with image content, giving them extremely strong stability and resistance to tampering. Even with advanced image processing techniques, their features can only be weakened to a limited extent, and some damage to image details and quality is inevitable. Researching PRNU anonymization algorithms aims to find methods to remove PRNUs from images as much as possible while minimizing the harm to image quality.

[0004] The shortcomings of existing technology:

[0005] 1. Existing PRNU anonymization methods mostly use wavelet transform and other methods to directly scramble or modify the high-frequency and low-frequency information of the image. When dealing with images with many elements, high detail, and complex structure, they cannot obtain visually high-quality de-PRNU images.

[0006] 2. Current anonymization methods often use traditional signal processing techniques such as wavelet transform and Wiener filtering to extract PRNU features when processing images with unknown PRNUs. This makes it difficult to handle random noise or noise with complex patterns in the image, and the anonymization effect is often not very good. Summary of the Invention

[0007] The technical problem to be solved by this invention is to provide a PRNU anonymization method based on multi-scale and hierarchical feature fusion, which weakens the physical association between the image and the device, while preventing malicious identity tracking and ensuring the practical value of the image in multimedia, medical, judicial and other scenarios.

[0008] The present invention achieves its objective by employing the following technical solution:

[0009] A PRNU anonymization method based on multi-scale and hierarchical feature fusion is characterized by the following steps:

[0010] S1: First, input the target image, and then extract the PRNU noise corresponding to the target image through a deep learning-based PRNU extraction network;

[0011] S2: Randomly generate a noise tensor with the same width and height as the target image, and input the tensor into the network architecture of the multi-scale dynamic Mamba U-net, which mainly consists of the multi-scale dynamic Unet and the Mamba module;

[0012] S3: Utilize prior knowledge from a multi-scale dynamic Mamba U-shaped network to obtain an intermediate image that more closely resembles the image structure features of the target image. ;

[0013] S4: Transfer intermediate image The PRNU noise from the target image is input into the PRNU suppression module. The PRNU suppression module adds perturbation to the PRNU noise of the target image through a learnable convolutional layer controlled by parameter λ and an attention mechanism. The perturbed PRNU information is then injected into the intermediate image in a multiplicative manner. In the middle, making The correlation between PRNU noise in the image and PRNU noise in the target image decreases.

[0014] S5: Throughout the training process, an iterative optimization framework is used to continuously adjust the parameters according to our designed loss function. and The value of PRNU is used to obtain an image that is visually closer to the target image and has had PRNU removed. .

[0015] As a further limitation of this technical solution, the multi-scale dynamic Mamba U-shaped network includes a unique encoder-decoder structure.

[0016] As a further limitation of this technical solution, the encoder includes an initial convolutional layer, a hierarchical hybrid downsampling module, and an EMA attention mechanism at the end; the downsampling module is mainly divided into 4 layers, the first two layers are residual blocks based on multi-scale adaptive convolution, and the last two layers are residual blocks based on Mamba architecture and feature adaptive attention mechanism.

[0017] S21: For the input feature map , R The numerical type of the feature map is represented by C, which represents the number of channels, H, which represents the height of the feature map, and W, which represents the width of the feature map. It is first input into the initial convolutional layer - the full-dimensional dynamic convolutional layer. This initial convolutional layer not only needs to adaptively extract features, but also needs to adjust the number of channels of the input feature map to adapt the dimension to the entire network.

[0018] S22: The adjusted feature map is then input into the downsampling module. The first layer does not perform downsampling; it only extracts features and uses residual connections. The last three layers perform downsampling once each. The entire downsampling process is represented as follows:

[0019] (1);

[0020] in: The output feature map;

[0021] This indicates that the layer has 4 residual blocks based on the Mamba architecture and feature adaptive attention mechanism, but only the first residual block based on the Mamba architecture and feature adaptive attention mechanism is downsampled;

[0022] This indicates that the layer has two residual blocks based on the Mamba architecture and feature adaptive attention mechanism, but only the first residual block based on the Mamba architecture and feature adaptive attention mechanism is downsampled;

[0023] This indicates that the layer has two residual blocks based on multi-scale adaptive convolution, but only the first residual block based on multi-scale adaptive convolution is downsampled;

[0024] This indicates that the layer has only one residual block based on multi-scale adaptive convolution, but no downsampling is performed; only features are extracted and residual connections are used.

[0025] S23: The EMA attention mechanism processes feature maps in groups and extracts and fuses features at different scales, enabling the model to pay attention to feature information at different scales simultaneously.

[0026] As a further limitation of this technical solution, the multi-scale adaptive convolution module consists of three depthwise separable convolutions (K1, K2, K3) with different kernel sizes, a spatial attention mechanism, and a channel attention mechanism.

[0027] For the input feature map The three convolutional branches generate feature maps of different scales (F1, F2, F3), thereby capturing local details and global relationships in the feature maps. The formula is as follows:

[0028] (2);

[0029] in: For depthwise separable convolution;

[0030] It is the first Convolution weights at various scales;

[0031] Subsequently, the feature maps of three different sizes are preserved and integrated through feature stitching operations to preserve and integrate feature information of different scales;

[0032] Finally, a 1×1 convolution is used to fuse features, compressing the high-dimensional features back to the original number of channels and integrating cross-channel information, thus obtaining the output of the multi-scale adaptive convolution, as shown in the following formula:

[0033] (3);

[0034] in: It is the first Normalized weights for each scale;

[0035] It involves applying attention weights to the feature map based on both channel and spatial dimensions;

[0036] The three convolutional branches generate feature maps of different scales.

[0037] As a further limitation of this technical solution, for the input feature map First, group normalization and nonlinear transformation operations are performed. Then, the feature map is input into the Mamba layer module. For the input feature map, this module first performs a flattening operation on it, transforming the feature map... The dimension was changed Where L = H × W, L is the total length of each channel after its spatial dimension, height H and width W are flattened into a one-dimensional sequence;

[0038] The feature map will then be adjusted. The input is fed into two branches for capturing spatial long-range dependencies. In the first branch, full-dimensional dynamic convolution and state-space model modules are used to enhance local and global feature representations, respectively. In the second branch, linear layers are used to expand feature channels to provide the model with richer feature representation capabilities, and the SiLU activation function is used to enhance the model's ability to learn complex features.

[0039] Finally, the features from the two channels are aggregated, the feature maps are adjusted back to their original channels, and then output. The entire MambaLayer module can be formulated as follows:

[0040] (4);

[0041] (5);

[0042] (6);

[0043] in: Indicates a linear layer;

[0044] This represents a depthwise separable convolutional layer;

[0045] Indicates the SiLU activation function;

[0046] Represents the state-space model module;

[0047] Indicates pool normalization;

[0048] This indicates feature fusion.

[0049] To further define this technical solution, the decoder consists of an upsampling module and a feature enhancement module. The decoder performs three upsampling operations, and after each upsampling and skip connection, the feature enhancement module enhances and refines the features to improve the decoder's ability to recover detailed features. Finally, the feature map channels are adjusted and output through final convolution. The entire decoding process can be expressed by the following formula:

[0050] (7);

[0051] in: Represents upsampling;

[0052] Representative feature enhancement module.

[0053] As a further limitation of this technical solution, a loss function is constructed in S5;

[0054] An image taken by a digital device can be represented as:

[0055] (8);

[0056] in: It is an ideal image with no noise;

[0057] The representative PRNU and his weights;

[0058] This represents other types of noise;

[0059] When generating an image At that time, the anonymization of the target image was successfully achieved;

[0060] Meanwhile, to ensure that the generated image closely resembles the target image at the pixel level, a set of parameters was found that allows the generated image to maintain visual quality while suppressing PRNU artifacts. The loss function is designed as follows:

[0061] (9);

[0062] in: These are images generated by a multi-scale dynamic Mamba U-shaped network;

[0063] It is a device fingerprint;

[0064] These are the learnable weights of the PRNU suppression module;

[0065] These are the weights of a multi-scale dynamic Mamba U-shaped network;

[0066] This represents the Frobenius norm.

[0067] Compared with related technologies, the PRNU anonymization method based on multi-scale and hierarchical feature fusion provided by this invention has the following beneficial effects:

[0068] (1) A novel model architecture, MD-MUnet (Multi-scale Dynamic MambaUnet), was proposed. It innovatively combines multi-scale adaptive convolution, Mamba architecture, and full-dimensional dynamic convolution to handle the PRNU anonymization task. Experimental results show that the PRNU-de-PRNU images obtained by this model have good improvements in image quality and anonymization effect.

[0069] (2) When performing PRNU extraction on the target image, a deep learning-based PRNU extraction method is used. Compared with traditional extraction methods, this method not only extracts more accurate PRNU features, but also makes up for the limitations of traditional methods in extracting PRNU from a single image. Through this method, we can achieve better anonymization when anonymizing multiple images from the same camera with unknown PRNU or a single image with unknown PRNU.

[0070] (3) Unlike previous methods, this paper dynamically adjusts the weights of the extracted target image's PRNU by combining learnable convolution and attention mechanisms. This modifies the regions of the target image with stronger device features and regions that are easier for PRNU to identify, thereby achieving better anonymity while reducing the degradation of image quality. Attached Figure Description

[0071] Figure 1 This is an overall structural diagram of the present invention.

[0072] Figure 2 This is a schematic diagram of the encoder of the present invention.

[0073] Figure 3This is a schematic diagram of the residual block based on multi-scale adaptive convolution of the present invention.

[0074] Figure 4 This is a schematic diagram of the residual block based on the Mamba architecture and feature adaptive attention mechanism of the present invention.

[0075] Figure 5 This is a schematic diagram of the feature enhancement module of the present invention.

[0076] Figure 6 This is a schematic diagram of the PRNU suppression module of the present invention.

[0077] Figure 7 This is a comparison chart of the experimental results of the present invention; wherein, Figure 7 (a) is the original camera image. Figure 7 (b) is the image after anonymization. Detailed Implementation

[0078] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0079] A PRNU anonymization method based on multi-scale and hierarchical feature fusion is proposed. Utilizing a deep image prior algorithm and a constructed multi-scale dynamic MambaUnet (MD-MUet) network architecture, it effectively anonymizes PRNU information in images while preserving visual features to the greatest extent possible. Figure 1 As shown, the following steps are included:

[0080] S1: First, input a target image, and then input the image into a deep learning-based PRNU extraction network. This network consists of a feature extraction module, a feature enhancement module, and a compression module with a dual network structure. It has a high accuracy in PRNU extraction of a single image. The network extracts the PRNU noise corresponding to the target image.

[0081] S2: A noise tensor with the same width and height as the target image is randomly generated, and this tensor is input into the Multi-scale Dynamic Mamba U-Net (MD-MUnet) network architecture. This architecture mainly consists of the Multi-scale Dynamic Unet (MDUnet) and the Mamba module. The Multi-scale Dynamic Unet, with its unique encoder-decoder structure, can extract and represent features from the input image at different resolution levels, and is particularly adept at capturing rich local texture features and spatial details in images. The Mamba module, as a novel sequence modeling component, uses an efficient state-space model mechanism to accurately model long-distance dependencies in feature sequences while maintaining low computational complexity, thus effectively integrating the global contextual information of the image.

[0082] S3: Utilize the prior knowledge of MD-MUnet to obtain an intermediate image that more closely resembles the image structure features of the target image. Compared to traditional methods that utilize deep prior knowledge from CNNs to reconstruct target images, our MD-MUnet network offers stronger feature representation capabilities and task adaptability, enabling it to better reconstruct high-quality images from noise.

[0083] S4: Transfer intermediate image The PRNU noise from the target image is input into the PRNU suppression module. The PRNU suppression module adds perturbation to the PRNU noise of the target image through a learnable convolutional layer controlled by parameter λ and an attention mechanism. The perturbed PRNU information is then injected into the intermediate image in a multiplicative manner. In the middle, making The correlation between PRNU noise in the image and PRNU noise in the target image decreases.

[0084] like Figure 6As shown, the PRNU suppression module consists of a learnable 1×1 convolutional layer and a SimAM attention mechanism. The convolutional layer weights the PRNU information using a learnable weight λ. Different weights scale each element of the PRNU tensor to varying degrees, thus corrupting the original device fingerprint features. Since not all regions in the PRNU pattern are equally important for device identification—for example, high-variance regions (such as sensor edges and defect points) typically contain stronger device features, and PRNUs in textured regions (such as the sky and solid-color backgrounds) are more easily identified—the SimAM attention mechanism is used to automatically calculate the importance weight of each feature point by analyzing the activation patterns of neurons, and spatially adaptively enhances the weighted PRNU, focusing on perturbing key device feature regions. Subsequently, the weighted PRNU information is injected into the intermediate image in a multiplicative manner. In the middle, making The correlation between PRNU noise in the image and PRNU noise in the target image decreases.

[0085] S5: Throughout the training process, an iterative optimization framework is used to continuously adjust the parameters according to our designed loss function. and The value of PRNU is used to obtain an image that is visually closer to the target image and has had PRNU removed. .

[0086] Through continuous iterative optimization, the model will eventually learn an optimal PRNU suppression method (the most suitable weights λ and parameters α). Under the premise of ensuring that the visual quality of the generated image is similar to that of the target image, the PRNU information in the final output image is sufficiently different from the PRNU information in the original image, so that the model can both generate images with good visual quality and remove PRNU information from the image.

[0087] The multi-scale dynamic Mamba U-shaped network includes a unique encoder-decoder structure.

[0088] Traditional encoders struggle to capture complete and accurate image features from images with numerous target elements, high detail, and complex structures. To better capture these features, encoder designs such as... Figure 2 As shown, the encoder includes an initial convolutional layer, a hierarchical hybrid downsampling module, and an EMA attention mechanism at the end. The downsampling module is mainly divided into four layers: the first two layers are residual blocks based on multi-scale adaptive convolution (RMAConvBlock), and the last two layers are residual blocks based on Mamba architecture and feature adaptive attention mechanism (RMFBlock).

[0089] S21: For the input feature map R represents the numerical type of the feature map, C represents the number of channels, H represents the height of the feature map, and W represents the width of the feature map. It is first input into the initial convolutional layer - the full-dimensional dynamic convolutional layer. This initial convolutional layer not only needs to adaptively extract features, but also needs to adjust the number of channels of the input feature map to adapt the dimension to the entire network.

[0090] S22: The adjusted feature map is then input into the downsampling module. The first layer does not perform downsampling; it only extracts features and uses residual connections. The last three layers perform downsampling once each. The entire downsampling process can be represented as follows:

[0091] (1);

[0092] in: The output feature map;

[0093] This indicates that the layer has 4 residual blocks based on the Mamba architecture and feature adaptive attention mechanism, but only the first residual block based on the Mamba architecture and feature adaptive attention mechanism is downsampled;

[0094] This indicates that the layer has two residual blocks based on the Mamba architecture and feature adaptive attention mechanism, but only the first residual block based on the Mamba architecture and feature adaptive attention mechanism is downsampled;

[0095] This indicates that the layer has two residual blocks based on multi-scale adaptive convolution, but only the first residual block based on multi-scale adaptive convolution is downsampled;

[0096] This indicates that the layer has only one residual block based on multi-scale adaptive convolution, but no downsampling is performed; only features are extracted and residual connections are used.

[0097] S23: The EMA (Efficient Multi-Scale Attention) attention mechanism processes feature maps in groups and extracts and fuses features at different scales, enabling the model to pay attention to feature information at different scales simultaneously.

[0098] The main function of the encoder is to extract and compress features from the input image, while the decoder restores the image based on the feature map output by the encoder. Adding an EMA attention mechanism at the end of the encoder makes the output feature map more representative and discriminative, thus providing the decoder with richer and more accurate information, guiding it to perform more accurate image restoration.

[0099] Residual blocks based on multi-scale adaptive convolution (RMAConvBlock) are one of the key modules in the encoder for downsampling, such as... Figure 3 As shown, for the input feature map First, group normalization and nonlinear transformation are used to improve the stability and expressive power of the model. Then, the feature map is input into the multi-scale adaptive convolution (MAConv) module. In order to better capture the details and textures of complex target images and ensure that the generated images have better visual quality, the multi-scale adaptive convolution module (which is the core convolution calculation part of RMAConvBlock and is responsible for implementing multi-scale feature extraction and attention mechanism, and is also in the first two layers of the downsampling module) consists of three depthwise separable convolutions with different kernel sizes (K1, K2, K3), a spatial attention mechanism, and a channel attention mechanism.

[0100] For the input feature map The three convolutional branches generate feature maps of different scales (F1, F2, F3), thereby capturing local details and global relationships in the feature maps. The formula is as follows:

[0101] (2);

[0102] in: For depthwise separable convolution;

[0103] It is the first Convolution weights at various scales;

[0104] Subsequently, the feature maps of three different sizes were preserved and integrated through feature stitching. In order to enable the model to allocate attention resources more comprehensively and accurately and effectively improve the ability to capture target details in complex scenes, we performed channel attention weighting and spatial attention weighting operations on the stitched high-dimensional feature maps respectively.

[0105] Finally, a 1×1 convolution is used to fuse features, compressing the high-dimensional features back to the original number of channels and integrating cross-channel information, thus obtaining the output of the multi-scale adaptive convolution, as shown in the following formula:

[0106] (3);

[0107] in: It is the first Normalized weights for each scale;

[0108] It involves applying attention weights to the feature map based on both channel and spatial dimensions;

[0109] The three convolutional branches generate feature maps of different scales.

[0110] Residual blocks (RMFBlocks) based on the Mamba architecture and feature-adaptive attention mechanism are another important module for downsampling in the encoder. For example... Figure 4 As shown, for the input feature map First, group normalization and nonlinear transformation operations are performed. Then, the feature map is input into the MambaLayer module (MambaLayer is the core feature extraction component of RMFBlock, specifically belonging to the convolution calculation part of this module, also in the last two layers of the encoder). For the input feature map, this module first performs a flattening operation on it, making the feature map... The dimension was changed (Where, L = H × W, L is the total length of each channel's spatial dimensions (height H and width W) after being flattened into a one-dimensional sequence).

[0111] The feature map will then be adjusted. The input is fed into two branches for capturing spatial long-range dependencies. In the first branch, full-dimensional dynamic convolution (ODConv) and a state-space model (SSM) module are used to enhance local and global feature representations, respectively. ODConv, by dynamically adjusting the convolution kernel parameters, adaptively focuses on key local regions of the feature map, enhancing local feature representation. The SSM module, with its efficient structure, can quickly capture global features. The combination of the two overcomes the shortcomings of traditional methods in simultaneously capturing local details and global information. In the second branch, linear layers are used to expand the feature channels to provide the model with richer feature representation capabilities, and the SiLU activation function is used to enhance the model's ability to learn complex features.

[0112] Finally, the features from the two channels are aggregated, and the feature maps are adjusted back to the original channel output. The entire MambaLayer module can be formulated as follows:

[0113] (4);

[0114] (5);

[0115] (6);

[0116] in: Indicates a linear layer;

[0117] This represents a depthwise separable convolutional layer;

[0118] Indicates the SiLU activation function;

[0119] Represents the state-space model module;

[0120] Indicates pool normalization;

[0121] This indicates feature fusion.

[0122] To better capture dynamic changes in features, improve feature representation capabilities, and maintain a lightweight model, a feature adaptive attention mechanism (FAattention) is added after MambaLayer, such as... Figure 4 As shown, a global average pooling operation is performed on the input feature map, and the result is flattened into a two-dimensional tensor. The flattened feature vector y is then input into the fully connected layer sequence to obtain the attention weights. Subsequently, we perform a weighted operation with the feature map, and the weighted features are input into the next MambaLayer module.

[0123] In order to better reconstruct a generated image that is visually closer to the target image based on the feature vectors extracted by the encoder. The decoder structure is as follows Figure 2 As shown, the decoder consists of an upsampling module and a feature enhancement module (FRBlock). Corresponding to the encoder, the decoder performs three upsampling operations. After each upsampling and skip connection, the feature enhancement module enhances and refines the features, improving the decoder's ability to recover detailed features. Finally, through final convolution, the feature map channels are adjusted and output. The entire decoding process can be expressed by the following formula:

[0124] (7);

[0125] in: Represents upsampling;

[0126] Representative feature enhancement module.

[0127] The structure of the Feature Enhancement Module (FRBlock) is as follows: Figure 5As shown, in the convolutional layer, ODConv replaces the traditional convolution. ODConv, as a novel dynamic convolution algorithm, has the core advantage of dynamically adjusting the shape and size of the convolution kernel based on the features of the input data during convolution operations. This algorithm, by embedding learnable deformation modules, enables the convolutional neural network to have stronger adaptability to diverse input data, thereby improving performance. By using ODConv's multi-scale adaptability, not only can the decoder's detail preservation ability during upsampling be improved, but the visual quality of the reconstructed image can also be enhanced. Subsequently, the features processed by ODConv are subjected to residual fusion with the original input. Finally, the residual-fused features are subjected to group normalization and nonlinear transformation operations again, and feature maps are output.

[0128] In S5, construct the loss function;

[0129] An image taken by a digital device can be expressed as:

[0130] (8);

[0131] in: It is an ideal image with no noise;

[0132] The representative PRNU and his weights;

[0133] This represents other types of noise;

[0134] When generating an image At that time, the anonymization of the target image was successfully achieved;

[0135] At the same time, in order to ensure that the generated image is close to the target image at the pixel level, The Frobenius norm of the target image I needs to be as small as possible. Therefore, a set of parameters needs to be found that allows the generated image to maintain visual quality while suppressing PRNU artifacts. The loss function is designed as follows:

[0136] (9);

[0137] in: These are images generated by a multi-scale dynamic Mamba U-shaped network;

[0138] It is a device fingerprint;

[0139] These are the learnable weights of the PRNU suppression module;

[0140] These are the weights of a multi-scale dynamic Mamba U-shaped network;

[0141] This represents the Frobenius norm.

[0142] Correlation analysis:

[0143] The anonymization effect of the generated anonymized images is analyzed, and the PRNU of the generated images and the PRNU of the original images are extracted for correlation analysis.

[0144] Normalized Correlation (NCC) is used to analyze the anonymization effect. A higher NCC value (closer to 1) indicates a worse anonymization effect, while a value closer to 0 indicates a better anonymization effect. The formula is as follows:

[0145] (10);

[0146] in: This represents the image The noise residual extracted from it;

[0147] This represents the product of image I and the target device PRNU;

[0148] Represents noise residual With PRNU simulation matrix The sum of pixel-by-pixel products reflects the degree of linear correlation between the two. A larger value indicates a stronger linear correlation. and The more similar the distribution patterns of the images, the better. The more likely it is to originate from the target device.

[0149] yes and The Frobenius norm is used to normalize the numerator to the interval [-1,1], eliminating the influence of matrix scale on correlation calculation.

[0150] Experimental setup

[0151] Anonymous experiments were conducted on smooth and natural images taken by three cameras—Samsung Galaxy S3 Mini, Apple iPhone 6, and Samsung Galaxy S3—on the VISION dataset, with approximately 100 images in each group.

[0152] The maximum number of training rounds is 2000. Each iteration saves the generated image and the corresponding NCC and PSNR information. The iteration stops when the set PSNR_max is reached.

[0153] For natural images, images with a PSNR ≥ 37 are selected from the images saved in each iteration (if this condition is not met, the highest generated PSNR is used as the threshold). The selected images are segmented into 32×32 pixel blocks. Then, all segmented image blocks are sorted in ascending order of their NCC (Negative Classification of Images) based on their original positions. For each position, the top 10 image blocks with the highest NCC are selected and sequentially stitched together to form 10 images. Finally, these 10 reconstructed images are averaged pixel-by-pixel to generate the final anonymized image.

[0154] For smoothed images, the image with the highest PSNR after iteration is used as our final anonymized image.

[0155] Experimental results:

[0156] Figure 7 The images presented are four images each from three devices in the Vision dataset: Samsung Galaxy S3 (first row), Samsung Galaxy S3 Mini (second row), and Apple iPhone 6 (third row). As can be seen, after anonymization using our PRNU method, the images have high visual quality and do not exhibit artifacts or other issues.

[0157] Experimental results:

[0158] Table 1 Smoothed Images

[0159]

[0160] Table 2 Natural Images

[0161]

[0162] As shown in the table above, our method outperforms existing methods in terms of anonymity (closer to 0) and visual quality (higher PSNR) for both smooth and natural images.

[0163] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A PRNU anonymization method based on multi-scale and hierarchical feature fusion, characterized in that, Includes the following steps: S1: First, input the target image, and then extract the PRNU noise corresponding to the target image through a deep learning-based PRNU extraction network; S2: Randomly generate a noise tensor with the same width and height as the target image, and input the tensor into the network architecture of the multi-scale dynamic Mamba U-net, which consists of multi-scale dynamic Unet and Mamba modules; S3: Utilize prior knowledge from a multi-scale dynamic Mamba U-shaped network to obtain an intermediate image that more closely resembles the image structure features of the target image. ; S4: Transfer intermediate image The PRNU noise from the target image is input into the PRNU suppression module. The PRNU suppression module adds perturbation to the PRNU noise of the target image through a learnable convolutional layer controlled by parameter λ and an attention mechanism. The perturbed PRNU information is then injected into the intermediate image in a multiplicative manner. In the middle, making The correlation between PRNU noise in the image and PRNU noise in the target image decreases. S5: Throughout the training process, an iterative optimization framework is used to continuously adjust the parameters according to our designed loss function. and The value of PRNU is used to obtain an image that is visually closer to the target image and has had PRNU removed. ; The multi-scale dynamic Mamba U-shaped network includes an encoder-decoder structure; The encoder includes an initial convolutional layer, a hierarchical hybrid downsampling module, and an EMA attention mechanism at the end. The downsampling module consists of four layers: the first two layers are residual blocks based on multi-scale adaptive convolution, and the last two layers are residual blocks based on Mamba architecture and feature adaptive attention mechanism. S21: For the input feature map , R The numerical type of the feature map is represented by C, which represents the number of channels, H, which represents the height of the feature map, and W, which represents the width of the feature map. It is first input into the initial convolutional layer - the full-dimensional dynamic convolutional layer. This initial convolutional layer not only needs to adaptively extract features, but also needs to adjust the number of channels of the input feature map to adapt the dimension to the entire network. S22: The adjusted feature map is then input into the downsampling module. The first layer does not perform downsampling; it only extracts features and uses residual connections. The last three layers perform downsampling once each. The entire downsampling process is represented as follows: (1); in: The output feature map; This indicates that the layer has 4 residual blocks based on the Mamba architecture and feature adaptive attention mechanism, but only the first residual block based on the Mamba architecture and feature adaptive attention mechanism is downsampled; This indicates that the layer has two residual blocks based on the Mamba architecture and feature adaptive attention mechanism, but only the first residual block based on the Mamba architecture and feature adaptive attention mechanism is downsampled; This indicates that the layer has two residual blocks based on multi-scale adaptive convolution, but only the first residual block based on multi-scale adaptive convolution is downsampled; This indicates that the layer has only one residual block based on multi-scale adaptive convolution, but no downsampling is performed; only features are extracted and residual connections are used. S23: The attention mechanism processes feature maps in groups and extracts and fuses features at different scales, enabling the model to pay attention to feature information at different scales at the same time.

2. The PRNU anonymization method based on multi-scale and hierarchical feature fusion according to claim 1, characterized in that: The multi-scale adaptive convolution module consists of three depthwise separable convolutions with different kernel sizes (K1, K2, K3), a spatial attention mechanism, and a channel attention mechanism. For the input feature map The three convolutional branches generate feature maps of different scales (F1, F2, F3), thereby capturing local details and global relationships in the feature maps. The formula is as follows: (2); in: For depthwise separable convolution; It is Convolution weights at various scales; Subsequently, the feature maps of three different sizes are preserved and integrated through feature stitching operations to preserve and integrate feature information of different scales; Finally, a 1×1 convolution is used to fuse features, compressing the high-dimensional features back to the original number of channels and integrating cross-channel information, thus obtaining the output of the multi-scale adaptive convolution, as shown in the following formula: (3); in: It is Normalized weights for each scale; It involves applying attention weights to the feature map based on both channel and spatial dimensions; The three convolutional branches generate feature maps of different scales.

3. The PRNU anonymization method based on multi-scale and hierarchical feature fusion according to claim 2, characterized in that: For the input feature map First, group normalization and nonlinear transformation operations are performed. Then, the feature map is input into the Mamba layer module. For the input feature map, this module first performs a flattening operation on it, transforming the feature map... The dimension was changed Where L = H × W, L is the total length of each channel after its spatial dimension, height H and width W are flattened into a one-dimensional sequence; The feature map will then be adjusted. The input is fed into two branches for capturing spatial long-range dependencies. In the first branch, full-dimensional dynamic convolution and state-space model modules are used to enhance local and global feature representations, respectively. In the second branch, linear layers are used to expand feature channels to provide the model with richer feature representation capabilities, and the SiLU activation function is used to enhance the model's ability to learn complex features. Finally, the features from the two channels are aggregated, the feature maps are adjusted back to their original channels, and then output. The entire MambaLayer module can be formulated as follows: (4); (5); (6); in: Indicates a linear layer; This represents a depthwise separable convolutional layer; Indicates the SiLU activation function; Represents the state-space model module; Indicates pool normalization; This indicates feature fusion.

4. The PRNU anonymization method based on multi-scale and hierarchical feature fusion according to claim 3, characterized in that: The decoder consists of an upsampling module and a feature enhancement module. The decoder performs three upsampling operations, and after each upsampling and skip connection, the feature enhancement module enhances and refines the features, improving the decoder's ability to recover detailed features. Finally, a final convolution is performed to adjust the feature map channels and output the result. The entire decoding process can be expressed by the following formula: (7); in: Represents upsampling; Representative feature enhancement module.

5. The PRNU anonymization method based on multi-scale and hierarchical feature fusion according to claim 1, characterized in that: In S5, construct the loss function; An image taken by a digital device can be represented as: (8); in: It is an ideal image with no noise; The representative PRNU and his weights; This represents other types of noise; When generating an image At that time, the anonymization of the target image was successfully achieved; Meanwhile, to ensure that the generated image closely resembles the target image at the pixel level, a set of parameters was found that allows the generated image to maintain visual quality while suppressing PRNU artifacts. The loss function is designed as follows: (9); in: These are images generated by a multi-scale dynamic Mamba U-shaped network; It is a device fingerprint; These are the learnable weights of the PRNU suppression module; These are the weights of a multi-scale dynamic Mamba U-shaped network; This represents the Frobenius norm.

Citation Information

Patent Citations

  • Optical response non-uniform noise anonymization method and system based on DCT and Wiener filtering

    CN116232605A

  • Multi-feature fusion PRNU extraction method and device based on double-layer hybrid model

    CN116977657A