Safe and available visual anonymization pedestrian privacy protection method and system

By building an image reconstruction network, using conditional variational autoencoder and structural feature encoder to generate reversible anonymous images, the usability and privacy balance of anonymous images are solved, and high-quality pedestrian privacy protection is achieved.

CN120343168APending Publication Date: 2025-07-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510480315.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing pedestrian privacy protection methods are difficult to balance the availability and privacy of anonymous images. Traditional methods lead to limited image information destruction and privacy protection performance, and GAN-based methods ignore the visual effects of anonymous images.

Method used

A network of image reconstruction is constructed, and the appearance features and structural features are extracted from the original image using the conditional variational autoencoder and structural feature encoder, random appearance features are generated through the appearance feature decoder and steganography into an anonymous image, and reversible anonymous images are generated by combining style transfer technology.

Benefits of technology

It has achieved enhanced privacy protection capabilities with reduced accuracy of pedestrian identification models, high image quality, supports multi-scene applications, and has reversible recovery capabilities to adapt to the needs of electronic evidence forensic investigations and other needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343168A_ABST
    Figure CN120343168A_ABST
Patent Text Reader

Abstract

The invention belongs to the image processing technology, and particularly relates to a safe and available visual anonymization pedestrian privacy protection method and system, and the method comprises the steps: constructing an image reconstruction network which employs a conditional variation auto-encoder and a structural feature encoder to extract appearance features and structural features from an original image, the method comprises the following steps: extracting structural features in an image, enabling an appearance feature decoder to randomly generate appearance features according to the distribution of the appearance features in the image, inputting the randomly generated appearance features and the extracted structural features into a generator to obtain an anonymized image, and writing the extracted appearance features into the anonymized image based on an image steganography technology to obtain a reversible anonymized image. According to the method, the original pedestrian information is decoupled and then subjected to style migration, so that the feature change of the pedestrian is realized, the effect of changing the clothes by the pedestrian is visually presented, and the balance between privacy protection and availability is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, and particularly relates to a method and system for protecting pedestrian privacy with secure and usable visual anonymization. Background Art

[0002] With the popularization of monitoring devices, monitoring has started to enter thousands of households and is used by individuals. While these devices provide convenience for individuals, image capture devices such as monitoring devices will obtain a large number of original pedestrian images and videos, which may be stored in local storage or uploaded to a cloud server. On the one hand, the collection and storage of monitoring data play an important role in legitimate uses. On the other hand, this brings serious privacy problems to personal and public security. Because the original images or videos contain sensitive information about pedestrians, such as information that can reflect their true identities, such as the walking trajectories or home addresses of pedestrians. Especially for private users, due to the lack of strong security protection measures, the data in monitoring devices is more likely to become the target of attackers. For example, malicious attackers may identify personal identity information by stealing monitoring videos, and then monitor their life trajectories, or even use their facial and postural features for criminal activities, such as forging identities to carry out online fraud.

[0003] In the research field, the existing pedestrian privacy protection methods can be classified into three categories:

[0004] 1) Methods based on traditional image processing. In traditional methods, identity anonymization can be achieved by simple methods, such as blurring, pixelation, and covering. For example, by detecting sensitive attribute areas such as faces, bodies, and backpacks in the monitoring video, and using motion blur technology to destroy the contours of sensitive areas, identity anonymization and behavior preservation are achieved. Because of their easy application and low cost, these methods are still active in some practical application scenarios. Although traditional methods have played a role to a certain extent, there are still problems such as the destruction of useful image information, the degradation of image quality, and limited privacy protection performance that need to be further solved.

[0005] 2) Generative methods based on GAN. This type of method generates different postures through pose guidance, or generates anonymized images with stronger privacy protection capabilities based on traditional methods. The Lin Yutian team proposed a pedestrian privacy protection method that realizes the anonymization of pedestrian images by migrating features such as pedestrian postures, appearances, and lighting conditions to virtual samples through virtual sample generation and data augmentation technology, and balances the performance of privacy protection and re-identification tasks without revealing the appearance information of pedestrians.

[0006] However, none of the above methods consider the issue of balancing the availability and privacy of anonymized images. Although some studies mention adding a person re-identification model during privacy protection, the visual effect of anonymized images is ignored. Summary of the Invention

[0007] Aiming at the problem of imbalance between the availability and privacy of current pedestrian privacy protection technologies, the present invention proposes a pedestrian privacy protection method for secure reversible visual anonymization, including constructing an image reconstruction network. This network uses a conditional variational autoencoder and a structural feature encoder to extract appearance features and structural features from the original image, and makes the appearance feature decoder randomly generate appearance features according to the distribution of appearance features in the image. The randomly generated appearance features and the structural features extracted from the original image are input into the generator to obtain an anonymized image. Then, the appearance features extracted from the original image are hidden into the anonymized image through reversible image steganography technology to obtain a reversible anonymized image.

[0008] Furthermore, when recovering the anonymized image, the original image appearance features hidden in the reversible anonymized image are recovered. Since the structural features of the anonymized image are the same as those of the original image, only the structural feature encoder is needed to extract the structural features from the reversible anonymized image. The generator is made to generate a restored image according to the original image appearance features and structural features obtained from the reversible anonymized image.

[0009] Furthermore, during the training of the image reconstruction network, two images with different appearance features are used as an image pair. The conditional variational autoencoder is used to obtain appearance features from the images and at the same time map the appearance features to the appearance feature latent space. Then, the appearance feature decoder is used to reconstruct the appearance features from the appearance feature latent space. The reconstructed appearance features and the structural features extracted from the other image in the image pair by the structural feature encoder are input into the generator to obtain the generated anonymized image.

[0010] The present invention also provides a pedestrian privacy protection system for secure and available visual anonymization, which is used for a pedestrian privacy protection method for secure and available visual anonymization, including:

[0011] A pedestrian image acquisition module, which is used to extract pedestrian images from video frames;

[0012] An appearance feature extraction module, which is used to obtain appearance features from the acquired pedestrian images;

[0013] A structural feature extraction module, which is used to obtain structural features from the acquired pedestrian images;

[0014] An appearance feature generation module, which is used to generate random appearance features according to the distribution of appearance features in the pedestrian images;

[0015] A generator for generating images based on structural features and appearance features;

[0016] A steganography module for writing the extracted appearance features into an anonymized image based on steganography technology.

[0017] In the present invention, by decoupling the original pedestrian information and performing style transfer, the features of the pedestrian are changed, presenting the effect of the pedestrian changing clothes visually, effectively solving the balance between privacy protection and usability. The specific beneficial effects of the present invention include:

[0018] 1), The present invention has enhanced privacy protection ability based on feature confusion. The reason is that the pedestrian recognition model relies on the clothing color information of the pedestrian. The present invention effectively decouples and reconstructs the semantic feature space of the pedestrian image, causing the accuracy rate of the pedestrian recognition model to decline. Through experiments, the privacy protection method has a Rank-1 recognition rate of 23% in the pedestrian recognition model AGW;

[0019] 2), The present invention has reversibility and can better adapt to some scenarios where the anonymized image needs to be restored, such as in electronic forensics investigations. Therefore, the application scope of the present invention is wider;

[0020] 3), The present invention has strong usability. The present invention can support machine recognition applications such as pedestrian detection and pose detection in different scenarios, and the anonymized image of the present invention has high imaging quality and is more easily accepted by people for the anonymized effect, taking into account both privacy protection and machine vision requirements. Description of the Drawings

[0021] Figure 1 It is a flowchart of a method for secure and usable visual anonymization of pedestrian privacy protection according to the present invention;

[0022] Figure 2 It is a schematic diagram of the training process of the privacy protection network in the present invention;

[0023] Figure 3 It is a schematic diagram of the inference process of the privacy protection network in the present invention;

[0024] Figure 4 It is a comparison schematic diagram of pedestrian anonymization in the present invention. Detailed Embodiments

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] The present invention proposes a pedestrian privacy protection method for secure reversible visual anonymization, which constructs an image reconstruction network. This network uses a conditional variational autoencoder and a structural feature encoder to extract appearance features and structural features from the original image, and enables the appearance feature decoder to randomly generate appearance features according to the distribution of appearance features in the image. The randomly generated appearance features and the extracted structural features are input into the generator to obtain an anonymized image. Then, the extracted appearance features are written into the anonymized image based on image steganography technology to obtain a reversible anonymized image.

[0027] In this embodiment, as Figure 1 , the present invention is a pedestrian privacy method for secure reversible visual anonymization, which specifically includes the following steps:

[0028] For an original pedestrian image to be processed, randomly select another image, and simultaneously perform deep decoupling of appearance features and structural features. The appearance features are learned and generated by a conditional variational autoencoder (CVAE), and the structural features are extracted by a structural feature extractor.

[0029] Exchange the appearance features of the two images while keeping the structural features unchanged, and deeply fuse them through a style transfer network. Extract the appearance and structural features of the two generated anonymized images again and exchange them again, that is, reconstruct the original image from the anonymized image.

[0030] Hide the appearance features of the original image into the finally generated anonymized image to make it reversible.

[0031] The construction process of the pedestrian anonymization network adopted in this embodiment includes:

[0032] Step 1) Retain the backbone network of the ResNet50 network, replace the final spatial average pooling layer with an adaptive pooling layer (AdaptiveAvgPool2d), and connect two parallel fully connected layers to generate the mean and variance of the latent space, which serve as the encoder of the conditional variational autoencoder (CVAE). The decoder learns the latent space features through conditional encoding and embeds them into a multi-level residual network to reconstruct the appearance features.

[0033] Step 2) In this embodiment, a progressive downsampling architecture is constructed in cooperation with deep residual connections to achieve decoupling of image structural features; furthermore, an adaptive instance normalization (AdaIN) module based on dynamic parameter injection is adopted to perform cross-modal fusion of the semantic features extracted from the identity features and the structural features, and finally generate an anonymized image with strong visual consistency.

[0034] Step 3) In this embodiment, the learned CVAE is used to generate the appearance features of random pedestrians, and then the original pedestrian features are generated. The random appearance features and the original pedestrian structure features are combined to generate an anonymized image. Using an advanced steganographic image method, the appearance code of the original image is hidden in the anonymized image.

[0035] The depth image fusion network is constrained by three loss functions, namely the conditional variational autoencoder (CVAE) loss function, the reconstruction loss function, and the adversarial loss function. Among them, the conditional variational autoencoder (CVAE) loss function is used to ensure that appearance features can be generated from the latent space, the reconstruction loss function is used to ensure that the generated image has the effect of style transfer, and the adversarial loss function is used to ensure that the quality of the generated image can be close to the original image, making the anonymized image more realistic and more usable. In this embodiment, the total loss of the image reconstruction network is expressed as:

[0036]

[0037] Among them, represents the total loss function for generating the anonymized image; is the KL divergence of the conditional variational autoencoder, which is used to constrain the statistical characteristics of the latent space and balance the controllability of the generation process; is the image reconstruction information loss between the anonymized image and the original image;

[0038] is the adversarial loss; λ1, λ2, and λ3 are the weights of respectively.

[0039] In this example, the KL divergence constraint of the conditional variational autoencoder (CVAE) is expressed as:

[0040]

[0041] Among them, β represents the weight hyperparameter of the KL divergence term, p(z) is the standard normal prior distribution, and q φ (z|x) is the posterior distribution parameterized by the encoder; D KL (q φ (z|x)||p(z)) represents calculating the divergence, which is expressed as is the mean of the i-th dimension output by the encoder, is the variance of the i-th dimension output by the encoder, d is the dimension of the encoder output features, z is the latent representation, which is an abstract feature extracted from the input original image by the conditional variational autoencoder in the present invention, usually located in a low-dimensional space, and the decoder of the present invention will use this latent representation to generate the feature code.

[0042] In this example, the image reconstruction information loss between the anonymized image and the original image Expressed as:

[0043]

[0044] Wherein, is the reconstruction loss of appearance features and structural features; a i is the appearance feature of the original image i, and s j is the structural feature of the original image j; i and j represent original images of different identities; G(a i , s j ) represents the image generated by the generator G according to the appearance feature a i , and the structural feature s j ; G(a i , s i ) represents the image generated by the generator G according to the appearance feature a i , and the structural feature s i ; x i represents the original image; represents the expectation value, and the present invention represents the average loss of the data distribution; E a (·) represents the appearance code encoder; E s (·) represents the structure code encoder; ‖·‖1 represents the L1 norm; is the reconstruction loss of the image; λ r1 , λ r2 are the weights of the feature reconstruction loss and the image reconstruction loss respectively; through the constraint of the reconstruction loss, the structural feature extractor and the conditional variational autoencoder (CVAE) can accurately extract the corresponding codes, and the style transferer can generate realistic and high-quality images.

[0045] In this example, the adversarial loss function is expressed as:

[0046]

[0047] Wherein, D(x i ) represents the result that the discriminator D judges that the original image x i is a real image; D(G(a i , s j ) represents the result that the discriminator D judges that the image generated by the generator G according to the appearance feature a i , and the structural feature s j is a real image; in this embodiment, the adversarial loss function hopes to minimize the recognition ability of the discriminator for the generated samples, that is, to make the discriminator misjudge the generated samples as real samples, and ensure the authenticity of the target domain of the image after style transfer.

[0048] In this example, the appearance features generated by a conditional variational autoencoder (CVAE) are fused with the extracted structural features through a style transfer network to generate anonymized images. When generating anonymized images, there is no need for paired corresponding images to train the style transfer network. Only two unrelated images need to be randomly selected. An anonymized image pair is generated by confusing the appearance features of the image pair, and then the appearance features of the anonymized image pair are swapped and confused to generate reconstructed images. By comparing the differences between the reconstructed images and the original images, backpropagation is carried out to the conditional variational autoencoder (CVAE), the structural code extractor, and the anonymized image generator, forcing them to learn appearance and structural features. Through such a training method, this method better adapts to the dataset used in this example.

[0049] A pedestrian privacy protection system for secure and usable visual anonymization according to the present invention is used for a method for secure and usable visual anonymization of pedestrian privacy protection, including:

[0050] A pedestrian image acquisition module for extracting pedestrian images from video frames;

[0051] An appearance feature extraction module for obtaining appearance features from the acquired pedestrian images;

[0052] A structural feature extraction module for obtaining structural features from the acquired pedestrian images;

[0053] An appearance feature generation module for generating random appearance features according to the distribution of appearance features in pedestrian images;

[0054] A generator for generating images according to structural features and appearance features;

[0055] A steganography module for writing the extracted appearance features into anonymized images based on steganography techniques.

[0056] The method of the present invention can be used in scenarios such as front-end hardware such as surveillance cameras and in-vehicle cameras. First, a conditional variational autoencoder (CVAE), a structural extractor, and an anonymized image generator are trained using template pedestrian images, enabling the method to have the ability to extract features and generate images. At this time, the method is deployed on the front-end server. After the front-end camera starts working, the detection module can be used to detect pedestrians. The detected pedestrian images are transmitted into the system. First, the conditional variational autoencoder (CVAE) is used to generate the original pedestrian appearance features, and then the conditional variational autoencoder (CVAE) is used to randomly generate the appearance features of other pedestrians. The original pedestrian structural features are extracted by the structural extractor, and finally, anonymized images are generated from the appearance features generated by the decoder and the original structural features, and the original image features are hidden in the anonymized images and output to the back-end monitoring video storage device.

[0057] This embodiment also presents the specific training process of the anonymization network, which specifically includes the following steps:

[0058] 1) Dataset

[0059] Market-1501 dataset: This dataset was collected on the campus of Tsinghua University, taken in summer, and constructed and made public in 2015. It includes 1,501 pedestrians captured by 6 cameras (5 high-definition cameras and 1 low-definition camera), 12,936 training images, 3,368 query images, and 19,732 gallery images.

[0060] 2) Training of the network

[0061] The proposed deep model is trained using the Market-1501 dataset. The model is saved every 10,000 iterations during training, and a total of 100,000 iterations are trained. The original dataset is optimized during the training process using the Adam optimizer. Every 2,000 iterations, 8 images are saved to observe the anonymization effect and reconstruction effect during the iteration, so as to observe the model training situation at any time.

[0062] As Figure 2 , when the image reconstruction network is trained, two images with different appearance features are used as an image pair. After obtaining the appearance features from the image using the conditional variational autoencoder Ce and mapping them to the latent space, the appearance feature decoder Cd is used to reconstruct the features. The reconstructed appearance feature Ia and the structural feature Is extracted from the other image in the image pair by the structural feature encoder Es are input into the generator G to obtain the generated anonymized image. After training, only the appearance feature decoder Cd and the conditional variational autoencoder Ce need to be deployed on the server. The appearance feature decoder Cd directly samples the appearance features from the latent space, and the conditional variational autoencoder Ce is used to extract the appearance features from the input image and implicitly write the extracted appearance features into the anonymized image.

[0063] As Figure 3 , the image reconstruction network uses the conditional variational autoencoder Ce and the structural feature encoder Es to extract the original appearance feature Ia’ and the structural feature Is from the original image. The appearance feature decoder Cd is made to randomly generate the appearance feature Ia according to the distribution of the appearance features in the image. The randomly generated appearance feature and the extracted structural feature Is are input into the generator G to obtain the anonymized image. Then, the original appearance feature Ia’ of the original image is based on the image steganography technique to implicitly write the appearance feature code into the anonymized image to obtain a reversible anonymized image. When restoring the anonymized image, the steganographed original image appearance feature is restored from the reversible anonymized image, and the structural feature encoder Es is used to generate the structural feature Is from the reversible anonymized image. The generator is made to generate the restored image according to the appearance feature and the structural feature obtained from the reversible anonymized image.

[0064] This embodiment also quantitatively verifies the proposed anonymization model through simulation experiments, the purpose of which is to test the privacy protection performance of the generated images. The experiment compares it with the traditional privacy protection method and effectively proves the superiority of the method of the present invention.

[0065] Table 1 shows the objective indicators of the privacy protection performance of the method of the present invention on the Market-1501 dataset. The results show that the privacy protection performance of the method of the present invention after the anonymization operation reaches the effect of the traditional privacy protection method, and the quality of the anonymized image is much higher than that of the traditional anonymization method.

[0066] Table 1 Anonymization performance evaluation indicators

[0067]

[0068] The AGW pedestrian identifier is used as the baseline for the recognition rate test. Traditional privacy protection methods use three independent methods to mask sensitive information in images:

[0069] Gaussian blur: Use a 12×12 convolution kernel to perform Gaussian blur on the image, remove high-frequency detail information through low-pass filtering, and reduce the recognizability of local features;

[0070] Block pixelation: The image is divided into 24×24 non-overlapping regions, and the pixel value in each region is replaced by the mean value of the region, achieving coarse-grained expression by reducing the spatial resolution;

[0071] Noise injection: Gaussian noise with a mean of 0 and a standard deviation of 0.5 is superimposed in the pixel domain to destroy the original texture features through random perturbations.

[0072] like Figure 4 This embodiment also provides three groups of comparative examples, namely phase 1 to phase 3. Each group of comparisons respectively displays the original image, the anonymized image, the restored image, the stego image and the stego restored image. The feasibility of the scheme in this embodiment is verified from the comparative examples. The example of the present invention provides a universal pedestrian privacy protection method, which effectively solves the problems of privacy protection and usability of pedestrian privacy images.

[0073] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A pedestrian privacy protection method for secure and usable visual anonymization, characterized in that Construct an image reconstruction network. This network uses a conditional variational autoencoder and a structural feature encoder to extract appearance features and structural features from the original image. Then, let the appearance feature decoder randomly generate appearance features according to the distribution of appearance features in the image, input the randomly generated appearance features and the structural features extracted from the original image into the generator to obtain an anonymized image. Next, write the appearance features extracted from the original image into the anonymized image through reversible image steganography technology to obtain a reversible anonymized image.

2. The method according to claim 1, wherein When restoring the anonymized image, restore the steganographed original image appearance features from the reversible anonymized image, use the structural feature encoder to extract structural features from the reversible anonymized image, and let the generator generate a restored image according to the original image appearance features and structural features obtained from the reversible anonymized image.

3. The method according to claim 1, characterized in that During the training of the image reconstruction network, let two images with different appearance features be used as an image pair. Use the conditional variational autoencoder to obtain appearance features from the images and at the same time map the appearance features to the appearance feature latent space. Then, use the appearance feature decoder to reconstruct the appearance features from the appearance feature latent space, and input the reconstructed appearance features and the structural features extracted from the other image in the image pair by the structural feature encoder into the generator to obtain the generated anonymized image.

4. The method according to claim 3, wherein Use the ResNet50 network to construct an encoder, extract the high-level appearance features of the original image through the deep residual structure in the ResNet50 network, and map them to a low-dimensional latent space to obtain a latent representation; the decoder performs reconstruction based on the latent representation.

5. The method according to claim 3, characterized in that During the training of the image reconstruction network, train with the KL divergence of the conditional variational autoencoder, the image reconstruction information loss between the anonymized image and the original image, and the adversarial loss of the generated image as the loss function.

6. The method according to claim 5, characterized in that, The total loss of the image reconstruction network is expressed as: Among them, represents the total loss function for generating anonymized images; is the KL divergence of the conditional variational autoencoder; is the image reconstruction information loss between the anonymized image and the original image; is the adversarial loss; λ1, λ2, and λ3 are the weights of respectively.

7. The method according to claim 6, characterized in that, KL divergence of conditional variational autoencoder It is expressed as: Among them, β represents the weight hyperparameter of the KL divergence term, p(z) is the standard normal prior distribution, and q φ (z|x) is the posterior distribution parameterized by the encoder; D KL (q φ (z|x)||p(z)) represents calculating the divergence, expressed as is the mean of the i-th dimension output by the encoder, is the variance of the i-th dimension output by the encoder, d is the dimension of the encoder output features, and z is the latent representation.

8. The method according to claim 6, wherein Image reconstruction information loss between the anonymized image and the original image Expressed as: Among them, is the reconstruction loss of appearance features and structural features; a i is the appearance feature of the original image i, s j is the structural feature of the original image j; G(a i , s j ) represents the image generated by the generator G according to the appearance feature a i and the structural feature s j ; G(a i , s i ) represents the image generated by the generator G according to the appearance feature a i and the structural feature s i ; x i represents the original image; represents the expected value; E a (·) represents the appearance code encoder; E s (·) represents the structure code encoder; ‖·‖1 represents the L1 norm; is the reconstruction loss of the image; λ r1 , λ r2 are the weights of the feature reconstruction loss and the image reconstruction loss respectively.

9. The method according to claim 6, wherein Adversarial loss Expressed as: Among them, denotes the expectation value; D(x i ) represents the result that the discriminator D determines that the original image x i is a real image; D(G(a i , s j ) represents the result that the discriminator D determines that the image generated by the generator G according to the appearance feature a i , the structural feature s j is a real image.

10. A pedestrian privacy protection system for secure and usable visual anonymization, characterized in that, A pedestrian privacy protection method for realizing secure and usable visual anonymization described in claim 1, including: A pedestrian image acquisition module, used to extract pedestrian images from video frames; An appearance feature extraction module, used to obtain appearance features from the acquired pedestrian images; A structural feature extraction module, used to obtain structural features from the acquired pedestrian images; An appearance feature generation module, used to generate random appearance features according to the distribution of appearance features in the pedestrian image; A generator, used to generate images according to structural features and appearance features; A steganography module, used to write the extracted appearance features into the anonymized image based on steganography technology.