Position privacy protection method and system for road building image
By extracting and combining building edge and non-building image content, combining style learning networks and adversarial training techniques, false road building images with different levels of style differences are solved, and the problems of insufficient utility retention, low image quality and insufficient privacy protection intensity control in the existing technology are solved, and efficient location privacy protection and image quality improvement are achieved.
Patent Information
- Application Number
- CN202510232783.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
When the prior art protects the privacy of locations in the images of road buildings, the utility remains insufficient, the image quality is low, and the intensity of privacy protection is lacking, which limits the applicable scenarios of the model.
By extracting the combination of building edges and non-building image contents, input them into the generator's encoder to generate realistic images that differ from the appearance of the original building. Use style learning network and adaptive instance normalization technology to generate false images of different styles, and conduct adversarial training through basic discriminators and pixel-level semantic discriminators to achieve control of specific privacy protection intensity.
The generated false road building images significantly change the building image content while retaining non-building area content, effectively protecting the privacy of the building location, improving the authenticity of the image, and achieving flexible control of the intensity of privacy protection.
Smart Images

Figure CN120145448A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of visual image privacy protection, and relates to a method and system for protecting the location privacy of road building images. Background Art
[0002] With the rapid development of the Internet and the wide deployment of deep learning models in daily life, image sharing and image collection have become very common. People share various image contents they have taken through social media and network applications, and a large number of images are also collected and labeled for model training. The resulting risk of privacy leakage is worrying. Privacy information such as users' personal identities, social relationships, and locations is likely to be leaked along with the image contents. Among them, location privacy is crucial. The leakage of personal location information may pose a threat to personal safety, and some lawbreakers may use the obtained location information for criminal activities. In addition, the leakage of location information may also lead to the infringement of other personal privacy rights, such as personal living habits, social circles, and other private information.
[0003] In the research on location privacy, traditional location-based services (LBS) usually require users to explicitly share their location information to obtain relevant services (such as navigation, social media check-ins, etc.), and have always been the focus of location privacy protection research. There have been many efforts to protect the location privacy of users in LBS. However, compared with LBS, the potential risk of location privacy leakage in image contents has often been overlooked or underestimated. Although images do not contain explicit geographical coordinates like LBS, the image contents themselves may still contain a large amount of locatable geographical information. Among them, buildings play an important role in the leakage of image location privacy, especially in the urban street view environment. There are many buildings and landmarks in the city, and these buildings contain a large amount of location information. Without any protection measures, attackers can easily infer the location where the photographer is located through manual analysis or geolocation technology. If the attackers have more images, they may even infer more sensitive information, such as the user's home address, trajectory, behavior habits, etc. With the rapid development of machine learning technology, the means of location inference are more diverse and the accuracy has been significantly improved, bringing new challenges to location privacy protection.
[0004] In this context, the importance of building location privacy protection begins to emerge. However, there is often a certain contradiction between privacy protection and the effective utilization of data. Strict privacy protection will inevitably reduce the quality of data services. Therefore, in the location privacy protection of buildings, how to conduct sufficient location privacy protection while retaining the utility information in the image (so as not to affect the normal use of the image) is a prerequisite that must be considered when designing location privacy protection methods. Currently, there have been some studies attempting to protect the location privacy of buildings on the premise of retaining the utility objects in the image. Xiong et al. were the first to focus on the location privacy protection of images collected by the cameras of autonomous vehicles. They used a generative adversarial network (GAN) to generate images to eliminate the privacy building objects in the images while maintaining the utility of other valuable objects, mainly using the structural similarity index measure (SSIM) and L 1 distance as the utility and privacy measures and using them as loss functions to guide the training of GAN. However, such measures are rough, and the quality of the generated buildings is not high. So they further proposed the ADGAN I and ADGAN-II models, no longer using SSIM and L 1 distance, but using semantic accuracy as the utility and privacy measures and setting a pre-trained FCN semantic segmentation model as the guiding model. ADGAN-I uses a generator built with a U-net network to generate noise and add it to the original image to generate a fake image. This method retains a relatively high quality, but the visual change of the buildings for human eyes is very small, and it can only resist the location inference attacks of machine models; ADGAN-II combines a generative adversarial network with a variational autoencoder (VAE) to directly generate fake street view building images. The images generated in this way have a more powerful protection effect at the cost of a greater utility loss.
[0005] Most existing methods reduce the risk of privacy leakage by deleting location-related information from the images. Specifically, a generative model is used to eliminate the privacy objects in the images while maintaining the practicality of the utility objects, so that the processed images can still be used for practical applications. But there are still some problems to be solved here:
[0006] 1. Insufficient utility retention. When processing privacy objects, other utility objects are modified, resulting in a lower overall usability of the images;
[0007] 2. The quality of the images generated by the models is relatively low. Although existing methods usually use high-quality generative models such as GAN and VAE, due to the interference of privacy processing, the authenticity of the generated images has decreased significantly compared to when privacy protection is not considered;
[0008] 3. The current location privacy protection methods lack consideration for privacy protection intensity control. When using these models for privacy processing, one can only choose to protect or not protect, and users have no room to customize the privacy protection intensity, which severely limits the applicable scenarios of the models. Summary of the Invention
[0009] To more effectively solve the problem of building location privacy leakage in road building images, the present invention proposes a location privacy protection method and system for road building images. First, extract the edges from the original building and combine the remaining non-building image content, then inject it into the encoder of the generator to prompt the generator to generate a realistic image with a different appearance from the original building. There is a skip connection between the encoder and decoder of the generator. Second, input the original road building image into the style learning network, perform adaptive instance normalization together with the intermediate features output by the encoder, and then input it into the decoder to obtain false road building images with different degrees of style differences. Finally, use the basic discriminator and pixel-level semantic discriminator to discriminate between the original real image and the false image, and perform adversarial training with the generator.
[0010] To achieve the above object, the present invention provides the following technical solutions:
[0011] On the one hand, a location privacy protection method for road building images is proposed. The method includes the following steps:
[0012] S1. Extract the building edge image from the original road building image I real with the help of the building mask M b and splice it with the remaining non-building part of the image to synthesize the edge composite image I B ;
[0013] S2. Input the edge composite image I B and the original image I real into the encoder G down of the generator respectively to obtain the intermediate features F B and F real ;
[0014] S3. Use the mean and standard deviation of the intermediate features F real after encoding the original image I real as the minimum difference style representation; input the original road building image I real into the style learning network to generate the standard deviation and mean, and use the generated standard deviation and mean as the maximum difference style representation;
[0015] S4. Align the mean and standard deviation of the intermediate features F B of the edge composite image to the minimum difference style and the maximum difference style respectively by using adaptive instance normalization to obtain the aligned intermediate features Fmin , F max ;
[0016] S5. Upsample the aligned intermediate feature F of the edge composite map in the decoder of the generator to generate false road and building images I with the maximum and minimum privacy protection levels respectively min , I max ; and calculate the style difference loss max , I min ;
[0017] S6. Use the multi-scale Patchgan discriminator D b as the basic discriminator to distinguish between true and false for the false road and building image I max , I min and the original real road and building image I real and calculate the adversarial loss
[0018] S7. Reflect the relative position distribution of the building area and the non-building area by calculating the correlation between the building mask elements, and calculate the losses of the basic discriminator and the generator according to the mask guidance loss function
[0019] S8. Use the pixel-level semantic discriminator D p to distinguish between true and false for the false road and building image I max , I min and the original real road and building image I real and calculate its semantic discrimination loss
[0020] S9. Backpropagate according to the calculated generator-related losses, iteratively train and update the parameters of the generator; for the trained generator, given the privacy protection intensity α, generate the required false road and building image I α .
[0021] Furthermore, in step S1, the following sub-steps are included
[0022] S11. Use the Canny algorithm to extract the edge image of the building from the original road and building image. The expression of the Canny algorithm is as follows
[0023] B α = canny(I real , α) ⊙ M b
[0024] where ⊙ represents the Hadamard product Hadamard, α represents the privacy protection intensity, and the value of α controls the completeness of the extracted building edge B α ; when α takes the value of 0, the edge extracted at this time is the most complete and is used as the ground truth edge, marked as Bgt , B gt = canny(I real , 0) ⊙ M b ;
[0025] S12. Extract the real edge B gt and the real street view image I real to jointly construct an edge composite map I B under the indication of the building mask:
[0026] I B = B gt ⊙ M b + I real ⊙ (1 - M b ).
[0027] Furthermore, in step S3, the following sub - steps are included:
[0028] Minimum style difference representation: Calculate the mean μ real and standard deviation σ real of the encoded intermediate feature F min of the original image I min along the channel and batch dimensions, which are respectively:
[0029]
[0030] In the formula, H and W respectively represent the height and width of the intermediate feature; h and w respectively represent the serial numbers of the elements in F real in the H and W dimensions;
[0031] Minimum style difference representation: Input the original image I real into the style learning network to generate a set of standard deviation and mean (μ max , σ max ) as the maximum difference style representation of the image, where the style learning network includes multiple convolutional modules and a final pooling layer, and each convolutional module includes a convolutional layer, an activation function, and a batch normalization layer.
[0032] Furthermore, in step S4, the following sub - steps are included:
[0033] S41. Calculate the mean and standard deviation of the encoded intermediate feature F B of the edge composite map I B along the batch and channel dimensions:
[0034]
[0035] S42. Perform style alignment to obtain the aligned intermediate feature F min , F max :
[0036]
[0037] In the formula, (μ min , σ min ), (μ max , σ max ) are the mean and standard deviation representing the minimum style difference and the maximum style difference respectively.
[0038] Furthermore, in step S5, the following sub-steps are included:
[0039] S51. Input the intermediate features F min , F max into the decoder G up of the generator, and upsample the intermediate features F min , F max to obtain the privacy-protected false road building images I min , I max respectively:
[0040] I min = G up (F min ), I max = G up (F max )
[0041] S52. Use the L 1 distance of the Gram matrices of the intermediate features of the two images in the pre-trained neural network to measure the style difference between the two images. With the goal that I min has a smaller style difference from the original image and I max has a larger style difference from the original image, the style difference loss function is:
[0042]
[0043] where is the style difference loss, and φ i (I) represents the feature map calculated after the activation function of the i-th layer after inputting I into the pre-trained classical classification neural network VGG19, and its shape is C i × H i × W i ;
[0044] The way of the Gram matrix is:
[0045] Gram(φ i (I)) = φ i (I)(φ i (I)) T
[0046] It is indicated that φ i (I) After vectorizing the features in the H i ,W i dimension, calculate its Gram matrix.
[0047] Furthermore, in step S6, the following sub-steps are included:
[0048] S61. Input I min , I max , I real to the basic discriminator for discrimination, calculate the adversarial loss on the basic discriminator, perform backpropagation, and update the parameters of the basic discriminator. Among them, the adversarial loss is:
[0049]
[0050] Among them, I fake represents the set of I min and I max , represents the adversarial loss on the basic discriminator;
[0051] S62. Calculate the adversarial loss corresponding to the generator:
[0052]
[0053] Among them, represents the adversarial loss of the basic discriminator on the generator.
[0054] Furthermore, in step S7, the following sub-steps are included:
[0055] S71. Given the image I fake The discrimination result obtained after inputting it into the basic discriminator is denoted as In f I Select N s query elements. For each query element p i Randomly sample N k elements and arrange them into a vector Calculate the correlation between the query element and these multiple elements as:
[0056]
[0057] Arrange the N s correlation queries into a correlation query set
[0058] S72. Given the mask M fake corresponding to the image I I , scale it to the same size as I fakeFor the same size, calculate the mask M I The correlation query set
[0059] S73. According to two correlation query sets Calculate the mask guidance loss:
[0060]
[0061]
[0062] where S c is the matrix obtained in the intermediate step, and p ij represents the element in the i-th row and j-th column of the matrix. That is, the average value of the elements of the matrix obtained by taking the difference between the two correlation query sets represents the mask guidance loss;
[0063] According to the mask guidance loss and the basic discriminator adversarial loss, perform backpropagation and iteratively train to update the parameters of the basic discriminator.
[0064] Furthermore, in step S8, it includes the following sub-steps:
[0065] S81. Given the fake road building image I fake , that is, I min , I max two types of images, input I fake into the pixel-level semantic discriminator for true / false discrimination;
[0066] The pixel-level semantic discriminator adopts a U-net architecture to perform a multi-classification task on each pixel of the input image. Its classification labels include: real building, fake building, real non-building, fake non-building; the final output has the same size as the input image; use the cross-entropy between the discriminator prediction result and the ground truth label as the loss function, which is called the semantic discrimination loss and is expressed as:
[0067]
[0068] where represents the semantic discrimination loss on the pixel-level semantic discriminator. The size of the discrimination result of the pixel-level semantic discriminator is 4×H×W, where H and W are the image size of I real , and t c,i,j represents the element value at the coordinate (c, i, j) in the discrimination result of the discrimination image, that is, the proportion of the pixel at the coordinate (i, j) in the discrimination image being classified as c;
[0069] According to the calculated loss, perform backpropagation and iteratively train to update the parameters of the pixel-level semantic discriminator;
[0070] S82. Calculate the semantic discrimination loss on the generator:
[0071]
[0072] where represents the semantic discrimination loss on the generator;
[0073] Backpropagate according to the calculated losses related to the generator, and iteratively train and update the parameters of the generator.
[0074] Furthermore, in step S9, given the privacy protection strength α through the trained generator, for F min , F max Perform feature interpolation using α in the feature space, and then upsample the interpolated features to obtain the target building image I with a specific privacy protection level α :
[0075] I α = G up (F max *α+(1-α)*F min )
[0076] In the formula, G up (·) represents the upsampling operation.
[0077] On the other hand, a system for implementing the aforementioned method for location privacy protection of road building images is also proposed. The system includes:
[0078] An edge composite module that extracts building edges and stitches together non-building images to form an edge composite map;
[0079] A generator module for generating a fake road building image after location privacy protection;
[0080] A mask guidance module for calculating the correlation between building mask elements to reflect the relative position distribution between the building area and the non-building area;
[0081] A basic discriminator module for discriminating the authenticity of the input image and performing adversarial training with the generator;
[0082] A pixel-level semantic discriminator module for discriminating the authenticity of each pixel of the input image and simultaneously performing building semantic classification, and performing adversarial training with the generator;
[0083] Among them, the generator module includes the following sub-modules:
[0084] An encoding module that performs multi-layer convolution on the input image to learn image features;
[0085] A decoding module that decodes intermediate features to generate a fake building image;
[0086] A style learning module, which includes multiple convolutional layers and one pooling layer. It inputs the original road building image and generates a set of mean and standard deviation (μ max , σ max ) as the style representation of the maximum image difference;
[0087] A style difference loss calculation module that calculates the style difference of the image using the Gram matrix, prompting I min to have a small style difference from the original image, and I max to have a large style difference from the original image.
[0088] The beneficial effects of the present invention are as follows:
[0089] (1) The fake road building image generated by the present invention significantly changes the building image content while retaining the non-building areas of the image, which can effectively protect the location privacy of buildings and resist location inference attacks.
[0090] (2) By additionally designing a pixel-level semantic discriminator and a mask-guided loss function, the present invention effectively improves the authenticity of the generated building image.
[0091] (3) The present invention realizes the control of the privacy protection intensity. By setting different privacy protection intensities, it can flexibly interpolate between the intermediate features of the maximum and minimum style differences, generating building styles with different degrees of difference from the original building style, achieving a better balance of privacy utility.
[0092] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings
[0093] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0094] Figure 1 is the overall flowchart of the method for protecting the location privacy of road building images provided by the embodiment of the present invention.
[0095] Figure 2 is the detailed flowchart of the method for protecting the location privacy of road building images provided by the embodiment of the present invention.
[0096] Figure 3Schematic diagram of the modules of the location privacy protection system for road building images provided by the embodiments of the present invention.
[0097] Figure 4 Schematic working principle diagram of the location privacy protection system for road building images provided by the embodiments of the present invention. Specific implementation manners
[0098] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0099] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be construed as a limitation on the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0100] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be construed as a limitation on the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0101] Please refer to Figures 1 to 4 , a method and system for protecting the location privacy of road building images.
[0102] Embodiment 1
[0103] This embodiment provides the specific steps of a method for protecting the location privacy of road building images. As Figure 1 shows the overall steps of the method of the present invention. Figure 2The specific steps of the method of the present invention are shown. Taking the training of a generative adversarial network model with a street view image dataset containing 3,000 training images as an example, the specific implementation steps of the present invention will be elaborated below. The goal is to change the building images in the input road street view images, protect the location privacy contained therein, and at the same time keep the image content of the non-building parts unchanged to generate false road building images. The steps include respectively:
[0104] Step S1, synthesize an edge composite map. Extract the building edge image from the original road building image I real and splice it with the remaining non-building part image by means of the building mask M b to obtain the edge composite map I B . It includes the following sub-steps:
[0105] Step S1-1, use the Canny algorithm to extract the building edge image from the original road building image. The expression of the Canny algorithm is as follows:
[0106] B α = canny(I real ,α)⊙M b
[0107] where ⊙ represents the Hadamard product, α represents the privacy protection intensity, and the value of α controls the integrity of the extracted building edge B α . When α takes the value of 0, the edge extracted at this time is the most complete, which is used as the ground truth edge and marked as B gt , B gt = canny(I real ,0)⊙M b .
[0108] In this embodiment, assume that I real has a batch size of 4, a channel number of 3, and the height and width dimensions of the image are 512 and 512 respectively; then the building mask M b is:
[0109]
[0110] Further calculate the ground truth edge B gt = canny(I real ,0)⊙M b , and get:
[0111]
[0112] Step S1-2, put the true edge B gtConstruct an edge composite graph I jointly with the real street view image under the indication of the building mask B :
[0113] I B = B gt ⊙ M b + I real ⊙ (1 - M b )
[0114] In this embodiment, the constructed edge composite graph is as follows:
[0115]
[0116] Step S2, downsample to calculate the minimum-difference style representation of the original building image. Input the edge composite graph I B obtained in step S1 and the original image I real into the encoder G down of the generator to obtain the mean and standard deviation of the encoded intermediate features. Use the encoder to downsample the edge composite graph and I B the original image I real, respectively to obtain the intermediate features F B , F real .
[0117] In this embodiment, let:
[0118]
[0119] The encoder G down consists of 7 convolutional layers. The convolutional kernel size of each convolutional layer is (4, 4), the stride is 2, and the padding is 1. The intermediate features after convolution
[0120] Step S3, minimum-difference style calculation and learning the maximum-difference style representation.
[0121] Calculate the mean μ real and standard deviation σ min of F min along the channel and batch dimensions, which is called the minimum-difference style representation and is respectively:
[0122]
[0123] In the formula, H and W respectively represent the height and width of the intermediate feature; h and w respectively represent the serial numbers of the elements in F real in the H and W dimensions.
[0124] In this embodiment, there is:
[0125]
[0126] The calculated σ min =([0.6421, 0.0658,..., 0.1819, 0.9962],...);
[0127]
[0128] The calculated μ min =([[0.1167, 0.8905,..., 0.3577, 0.7530],...]),
[0129] Input the original road building image I real into the style learning network to generate a set of standard deviation and mean (μ max , σ max ) as the maximum difference style representation of the image. The style learning network includes multiple convolutional modules and a final pooling layer. Each convolutional module includes a convolutional layer, an activation function, and a batch normalization layer.
[0130] In this embodiment, the size of the input image of the style learning network is 4×3×512×512. After two layers of convolution, the size of each convolutional kernel is (4, 4), the stride is 2, and the padding is 1. The size of the convolutional feature is 4×512×128×128. Then, the features are respectively input into two separate convolutional layers to obtain two output features with a size of 4×512×64×64. Average pooling is performed in the W and H dimensions to obtain the mean and standard deviation of the maximum difference style representation:
[0131] μ max =([[0.8664, 0.6567,…, 0.7066, 0.6699],…])
[0132] σ max =([0.8860, 0.2503,…, 0.2748, 0.6551],…])
[0133] Step S4, style alignment. Use adaptive instance normalization to align the mean and standard deviation of F B to (μ min , σ min ), (μ max , σ max ) respectively to obtain the aligned intermediate features F min , F max .
[0134] Step S4-1, calculate the mean and standard deviation of F B along the batch and channel dimensions.
[0135]
[0136] In this embodiment, μ(F B ) = ([[0.3821, 0.3304,..., 0.3751, 0.7162],…]), σ(F B ) = ([[0.7427, 0.6046,..., 0.8594, 0.2843],...]) is calculated according to the formula.
[0137] Step S4-2: Perform style alignment to obtain the aligned intermediate feature F min , F max .
[0138]
[0139] In the formula, (μ min , σ min ), (μ max , σ max ) are the mean and standard deviation representing the minimum style difference and the maximum style difference respectively.
[0140] In this embodiment, it is calculated according to the formula that:
[0141]
[0142] Step S5: Generate a fake road building image and calculate the style difference loss. Input the obtained intermediate features F min , F max into the upsampling in the generator decoder to generate fake road building images with the maximum privacy protection level and the minimum privacy protection level, and calculate the loss using the image according to the style difference loss function. It includes the following sub-steps:
[0143] Step S5-1: Input the intermediate features F min , F max into the decoder G up of the generator, and upsample the intermediate features F min , F max to obtain the fake road building images I min , I max
[0144] I min = G up (F min ), I max = G up (F max )
[0145] In this embodiment, it is calculated that:
[0146]
[0147] Step S5-2: Use the L of the Gram matrix of the intermediate features of two images in the pre-trained neural network 1 distance to measure the style difference between the two images, such that I min has a smaller style difference from the original image, and I max has a larger style difference from the original image. The style difference loss function formula is as follows:
[0148]
[0149] Among them, That is, the style difference loss. φ i (I) represents the feature map after calculating the activation function of the i-th layer after inputting I in the trained VGG19 network, and its shape is C i ×H i ×W i , and this network takes the relu_1, relu2_1, relu3_1, relu4_1 layers of VGG19. Indicates that after vectorizing the features of φ i (I) in the H i , W i dimensions, calculate its Gram matrix, and the specific calculation formula is as follows.
[0150] Gram(φ i (I)) = φ i (I)(φ i (I)) T
[0151] In this embodiment, it is calculated according to the style difference loss function formula
[0152] Step S6: Basic discriminator discrimination. Use the multi-scale Patchgan discriminator D b to distinguish between true and false for the fake road building images and the original real road building images obtained in Step S5, and calculate the adversarial loss. It includes the following sub-steps:
[0153] Step S6-1: Input I min , I max , I real to the basic discriminator for discrimination, calculate the adversarial loss on the basic discriminator, and perform backpropagation to update the parameters of the basic discriminator.
[0154]
[0155] Among them, I fake represents Imin and I max of the set represents the adversarial loss on the basic discriminator.
[0156] In this embodiment, calculated according to the above formula
[0157] Step S6-2, calculate the adversarial loss corresponding to the generator.
[0158]
[0159] Among them, represents the adversarial loss of the basic discriminator on the generator.
[0160] In this embodiment, calculated according to the above formula
[0161] Step S7, calculate the mask guidance loss. By calculating the correlation between the building mask elements to reflect the relative position distribution of the building area and the non-building area, so as to improve the discrimination ability of the basic discriminator, and calculate the losses of the basic discriminator and the generator according to the mask guidance loss function. It includes the following sub-steps:
[0162] Step S7-1, given the image I fake , that is, I min , I max For the two types of images, the discrimination results obtained after inputting into the basic discriminator are expressed as In f I Select N s query elements. For each query element p i Randomly sample N k elements and arrange them into a vector Calculate the correlation between the query element and these multiple elements as:
[0163]
[0164] Arrange the N s correlation queries into a set
[0165] In this embodiment, take N s = 250, N k = 100, and calculate to get:
[0166]
[0167] Step S7-2, given the mask M fake corresponding to the image I I , scale it to the same size as I fakeThe same size. Implement according to the steps of S7-1, and calculate the mask M I The relevance query set of
[0168] In this embodiment, it is calculated that:
[0169]
[0170] Step S7-3, use the two relevance query sets obtained in steps S7-1 and S7-2 Calculate the mask guidance loss:
[0171]
[0172]
[0173] Among them, S c is the matrix obtained in the intermediate step, and p ij represents the element in the i-th row and j-th column of the matrix. That is, the average value of the elements of the matrix obtained by taking the difference between the two relevance query sets represents the mask guidance loss.
[0174] According to the mask guidance loss and the basic discriminator adversarial loss, perform backpropagation and iteratively train and update the parameters of the basic discriminator.
[0175] In this embodiment, it is calculated that:
[0176]
[0177] Step S8, pixel-level semantic discriminator discrimination. Input the fake road building image and the original real road building image obtained in step S5 into the pixel-level semantic discriminator D p for true / false discrimination and calculate the relevant adversarial loss. It includes the following sub-steps:
[0178] Step S8-1, given the fake road building image I fake , that is, I min , I max two types of images, input I fake into the pixel-level semantic discriminator for true / false discrimination. The pixel-level semantic discriminator adopts a U-net architecture and performs a multi-classification task on each pixel of the input image. It has four types of labels: real building, fake building, real non-building, and fake non-building. The final output is the same size as the input image. The form of its loss function is the cross-entropy between the discriminator prediction result and the ground truth label, which is called the semantic discrimination loss. Calculate the adversarial loss as follows:
[0179]
[0180] Among them, represents the semantic discrimination loss on the pixel-level semantic discriminator. The size of the discrimination result of the pixel-level semantic discriminator is 4×H×W, where H and W are the real image size of I, c,i,j represents the element value at the coordinate (c, i, j) in the discrimination result of the discriminated image, that is, the proportion of the pixel at the coordinate (i, j) in the discriminated image being classified as c.
[0181] Backpropagate according to the calculated loss, and iteratively train and update the parameters of the pixel-level semantic discriminator.
[0182] In this embodiment, the pixel-level semantic discriminator is composed of 7 downsampling convolutional layers and 7 transposed convolutional layers. According to the formula, the pixel-level semantic discriminator loss is calculated as
[0183] Step S8-2, correspondingly, calculate the semantic discrimination loss on the generator.
[0184]
[0185] Among them, represents the semantic discrimination loss on the generator.
[0186] Backpropagate according to the calculated generator-related loss, and iteratively train and update the parameters of the generator.
[0187] In this embodiment, according to the formula, the pixel-level semantic discrimination loss corresponding to the generator is calculated as
[0188] Step S9, for the trained generator, given the privacy protection strength α, generate the required false road building images. For F min , F max Perform feature interpolation using α in the feature space, and then upsample the interpolated features. Obtain the target building image I with a specific privacy protection degree α :
[0189] I α = G up (F max *α+(1-α)*F min )
[0190] In this embodiment, let α = 0.5, generate the false road building images under this privacy protection strength, and calculate to obtain:
[0191]
[0192] Embodiment 2
[0193] This embodiment provides a location privacy protection system for road building images for implementing the aforementioned location privacy protection method for road building images, as Figure 3 shown, including:
[0194] An edge composite module that extracts building edges and stitches non-building images to form an edge composite map.
[0195] A generator module for generating a false road building image after location privacy protection. It includes the following sub-modules,
[0196] An encoding module that performs multi-layer convolution on the input image to learn image features;
[0197] A decoding module that decodes intermediate features to generate a false building image;
[0198] A style learning module that includes multi-layer convolutional layers and one pooling layer. It inputs the original road building image and generates a set of mean and standard deviation (μ max , σ max ) as the style representation of the maximum image difference;
[0199] A style difference loss calculation module that calculates the style difference of the image using the Gram matrix, prompting I min to have a smaller style difference from the original image, and I max to have a larger style difference from the original image;
[0200] A mask guidance module for calculating the correlation between building mask elements to reflect the relative position distribution of the building area and the non-building area.
[0201] A basic discriminator module for discriminating the authenticity of the input image and performing adversarial training with the generator.
[0202] A pixel-level semantic discriminator module for discriminating the authenticity of each pixel of the input image and simultaneously performing building semantic classification, and performing adversarial training with the generator.
[0203] As Figure 4 shown in the schematic diagram of the principle structure, the present invention designs a location privacy protection method and system for road building images by using building semantics to improve the discriminator of the generative adversarial network and combining building edge information. By extracting the building edges of the image and inputting them into the generator, then calculating the style features of the image, setting up a maximum style learning network, generating two false building images with different degrees of style difference, and then performing adversarial training through the mask-guided basic discriminator and pixel-level semantic discriminator, and finally outputting a false road building image with a specific privacy protection level.
[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for protecting location privacy of road and building images, characterized by: The method comprises the following steps: S1. From the original road building image I real The building edge image is extracted from the b Combined with the remaining non-building images to synthesize edge composite image I B ; S2, composite edge graph I B and the original image I real Input to the encoder G of the generator down The intermediate features F are obtained respectively B and F real ; S3, the original image I real The encoded intermediate feature F real The mean and standard deviation of real Input into the style learning network to generate standard deviation and mean, and use the generated standard deviation and mean as the maximum difference style representation; S4, using adaptive instance normalization to normalize the intermediate features F of the edge composite graph B The mean and standard deviation of are aligned to the minimum difference style and the maximum difference style respectively, and the aligned intermediate feature F is obtained. min ,F max ; S5. Align the intermediate features F of the edge composite graph min ,F max The decoder of the input generator is upsampled to generate false road building images I with maximum privacy protection and minimum privacy protection respectively. max ,I min ; and calculate the style difference loss; S6. Use multi-scale PatchGAN discriminator D b As the basic discriminator for the false road building image I max ,I min and the original real road building image I real Perform true and false discrimination and calculate adversarial loss; S7, by calculating the correlation between the elements of the building mask to reflect the relative position distribution of the building area and the non-building area, which calculates the loss of the basic discriminator and the generator according to the mask guidance loss function; S8, through pixel-level semantic discriminator D p For the false road building image I max ,I min and the original real road building image I real Perform true and false discrimination and calculate its semantic discrimination loss; S9. According to the calculated generator-related loss back propagation, iterative training updates the parameters of the generator; for the trained generator, given the privacy protection strength α, generate the required false road building image I α .
2. The method for protecting location privacy of road and building images according to claim 1, characterized in that: In step S1, the following sub-steps are included: S11. Use the Canny algorithm to extract the edge image of the building from the original road building image. The expression of the Canny algorithm is as follows: B α =canny(I real ,α)⊙M b Where ⊙ represents the Hadamard product, α represents the privacy protection strength, and the value of α controls the extracted building edge B. α When α is 0, the extracted edge is the most complete and is marked as the ground truth edge. gt , B gt = canny(I real ,0)⊙M b ; S12, extract the real edge B gt Compared with real street view images real Jointly construct edge composite graph I under the guidance of building mask B : I B =B gt ⊙M b +I real ⊙(1-M b )。 3. The method for protecting location privacy of road and building images according to claim 1, characterized in that: In step S3, the following sub-steps are included: Minimum style difference representation: Calculate the original image I along the channel and batch dimensions real The encoded intermediate feature F real The mean μ min and standard deviation σ min , which are: In the formula, H, W represent the height and width of the intermediate feature, h, w represent F real The sequence number of the elements in the H and W dimensions; Minimum style difference representation: the original image I real Input into the style learning network to generate a set of standard deviation and mean (μ max ,σ max ) is used as the maximum difference style representation of the image, where the style learning network contains multiple convolution modules and a final pooling layer. Each convolution module includes a convolution layer, an activation function and a batch normalization layer.
4. The method for protecting location privacy of road and building images according to claim 1, characterized in that: In step S4, the following sub-steps are included: S41, calculate edge composite graph I along batch and channel dimensions B The encoded intermediate feature F B The mean and standard deviation of : S42: Perform style alignment to obtain the aligned intermediate feature F min ,F max : In the formula, (μ min ,σ min )、(μ max ,σ max ) are the mean and standard deviation of the minimum style difference and the maximum style difference, respectively.
5. The method for protecting location privacy of road and building images according to claim 1, characterized in that: In step S5, the following sub-steps are included: S51, the intermediate feature F min ,F max Input to the decoder G of the generator up In the middle, upsample the intermediate feature F min ,F max , and obtain the false road building images I after privacy protection min ,I max : I min =G up (F min ),I max =G up (F max ) S52, use the L1 distance of the Gram matrix of the intermediate features of the two images in the pre-trained neural network to measure the style difference between the two images, with I min The style of the original image is smaller, max The goal is to have a greater style difference with the original image, then the style difference loss function is: in, That is, the style difference loss, φ i (I) represents the feature map after input I is calculated to the i-th layer activation function in the pre-trained classic classification neural network VGG19, and its shape is C i ×H i ×W i ; The Gram matrix format is: Gram(φ i (I))=φ i (I)(φ i (AND)) T Indicates that φ i (I) In H i ,W i After the features on the dimension are vectorized, their Gram matrix is calculated.
6. The method for protecting location privacy of road and building images according to claim 1, characterized in that: In step S6, the following sub-steps are included: S61, Input I min , I max , I real Go to the base discriminator for identification, calculate the adversarial loss on the base discriminator, perform back propagation, and update the parameters of the base discriminator, where the adversarial loss is: Among them, I fake I min and I max A collection of represents the adversarial loss on the base discriminator; S62. Calculate the adversarial loss corresponding to the generator: in, represents the adversarial loss of the base discriminator on the generator.
7. The method for protecting location privacy of road and building images according to claim 1, characterized in that: In step S7, the following sub-steps are included: S71, given image I fake The identification result obtained after inputting into the basic discriminator is expressed as In f I Select N s query elements, for each query element p i Random sampling N k Elements are arranged into vectors The correlation between the query element and these multiple elements is calculated as: N s The relevant queries are arranged into a relevant query set S72, given correspondence and image I fake The mask M I , scaled to the same fake The same size, the mask M is calculated I A collection of related queries S73. Querying a set based on two correlations Compute the mask-guided loss: Among them, S c is the matrix obtained in the intermediate step, p ij represents the element in the i-th row and j-th column of the matrix, That is, two sets of related queries The element average of the matrix obtained by difference represents the mask guidance loss; According to the mask guidance loss and the base discriminator adversarial loss, back propagation and iterative training are performed to update the parameters of the base discriminator.
8. The method for protecting location privacy of road and building images according to claim 1, characterized in that: In step S8, the following sub-steps are included: S81: Given a false road building image I fake , that is I min ,I max Two types of images, I fake Input to the pixel-level semantic discriminator for true or false discrimination; The pixel-level semantic discriminator uses the U-net architecture to perform multi-classification tasks on each pixel of the input image. Its classification labels include: real building, fake building, real non-building, fake non-building; the final output has the same size as the input image; the cross entropy between the discriminator prediction result and the ground truth label is used as the loss function, which is called the semantic discrimination loss, which is expressed as: in, represents the semantic discrimination loss on the pixel-level semantic discriminator. The size of the pixel-level semantic discriminator discrimination result is 4×H×W, where H and W are I real The image size, t c,i,j Represents the element value of the coordinates (c, i, j) in the identification result of the identification image, that is, the proportion of pixels with coordinates (i, j) in the identification image that are classified as c; According to the calculated loss back propagation, iterative training updates the parameters of the pixel-level semantic discriminator; S82. Calculate the semantic discrimination loss on the generator: in, represents the semantic discrimination loss on the generator; According to the calculated generator-related loss backpropagation, the parameters of the generator are updated through iterative training.
9. The method for protecting location privacy of road and building images according to claim 1, characterized in that: In step S9, through the trained generator, given the privacy protection strength α, min ,F max In the feature space, α is used to perform feature interpolation, and then the interpolated features are upsampled to obtain the target building image I with a specific privacy protection degree. α : I α =G up (F max *a+(1-a)*F min ) In the formula, G up (·) indicates an upsampling operation.
10. A system for executing the method for protecting the location privacy of road and building images as claimed in any one of claims 1 to 9, characterized in that: The system comprises: The edge composite module extracts the building edges and combines the non-building images to form an edge composite image; A generator module, used to generate fake road and building images after location privacy protection; The mask guidance module is used to calculate the correlation between the building mask elements to reflect the relative position distribution of the building area and the non-building area; The basic discriminator module is used to identify the true and false input images and conduct adversarial training with the generator; The pixel-level semantic discriminator module is used to identify the authenticity of each pixel in the input image, classify the building semantics, and conduct adversarial training with the generator; Among them, the generator module contains the following submodules, The encoding module performs multi-layer convolution on the input image to learn image features; Decoding module, decoding intermediate features and generating fake building images; The style learning module, which contains multiple convolutional layers and one pooling layer, inputs the original road building image and generates a set of mean and standard deviation (μ max ,σ max ) as the maximum difference style representation of the image; The style difference loss calculation module uses the Gram matrix to calculate the style difference of the image, which promotes I min The style of the original image is slightly different. max There is a big difference in style from the original picture.
Citation Information
Patent Citations
Information hiding method, device and system based on cavity space pyramid
CN113254891A
Image privacy protection method and system based on generative adversarial network
CN114329549A
Multi-modal private data generation model training method, data generation method and system
CN118862144A
CycleGAN-based sunny and snowy weather image data style migration system and method
CN118864630A
Medical image segmentation method based on u-net
US20220309674A1