Spatial data information desensitization processing method based on generative adversarial model

Through the digital watermark added by the spatial data information desensitization processing method based on the generative adversarial model and the fast Fourier transform method, the shortcomings of concealment and traceability in the existing GIS spatial data desensitization methods are solved, and more concealed desensitization effects and fast traceability are achieved.

CN120070141APending Publication Date: 2025-05-30HANGZHOU SHUIWU KONGGU GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411856376.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing GIS spatial data desensitization method is difficult to achieve more concealed desensitization effects, and it can quickly trace the source of the leakage after data leakage.

Method used

The spatial data information desensitization processing method based on the generative adversarial model is adopted to desensitize the spatial data images by constructing a generative adversarial neural network, and digital watermarks are added using the fast Fourier transform method to achieve data concealment and traceability.

Benefits of technology

The generated desensitized spatial data images can achieve more real-like effects than real, completely desensitizing spatial information, and at the same time, the digital watermark is more concealed and cannot be cracked, and the source of the leakage can be quickly located after the data is leaked.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070141A_ABST
    Figure CN120070141A_ABST
Patent Text Reader

Abstract

The invention provides a spatial data information desensitization processing method based on a generative adversarial model. The method comprises the following steps: constructing a generative adversarial neural network, and training the generative adversarial neural network; using the trained generative adversarial neural network to perform desensitization processing on a to-be-processed spatial data image to obtain a desensitized spatial data image; adding a digital watermark to the desensitized spatial data image by using a fast Fourier transform method; when the desensitized spatial data image is leaked, extracting a first digital watermark from the leaked desensitized spatial data image by adopting a watermark extraction algorithm, and carrying out similarity matching on the extracted first digital watermark and all second digital watermarks in the original watermark data set; and determining the desensitized spatial data image corresponding to the second digital watermark with the highest matching degree as a leakage source. According to the method, the spatial information can be desensitized thoroughly, and a leakage source can be quickly positioned after the desensitized spatial data image is leaked.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing technology, and specifically relates to a spatial data information desensitization processing method based on a generative adversarial model. Background Art

[0002] GIS spatial data desensitization is a method of transforming sensitive data according to rules and policies. Its purpose is to remove sensitive information while maintaining the original characteristics of the data to ensure the security and validity of the data. Data desensitization technology can process sensitive data so that it cannot be identified when it is accessed and obtained without authorization. By using data desensitization technology, personal privacy data and privacy data of social institutions can be used in non-secure environments without being exposed to potential risks.

[0003] The main task of the existing common GIS spatial data desensitization methods is to shield the sensitive information in the spatial structure data, such as Figure 1 As shown in the figure, a piece of content is missing in the middle of the image, which is likely to produce the effect of covering one's ears and stealing the bell. Some other existing desensitization methods also use the projection coordinate conversion method, using some GIS processing software, such as ArcGIS, QGIS, etc., using the projection conversion tool attached to the software, and offsetting the GIS spatial information by customizing the projection conversion coordinate system.

[0004] In addition, to ensure that the source of GIS spatial data can be traced after it is leaked, the existing technology uses digital watermark technology. Digital watermark is a kind of meta-information hidden in the data, which is commonly used to verify the integrity and authenticity of the data. However, in the scenario of GIS data, digital watermark is usually used to hide important information such as ownership and usage rights of geographic information, and is mainly used for tracing the source after the GIS data is leaked. The existing common GIS spatial data digital watermark encryption method is the LSB replacement algorithm, which embeds the watermark information by modifying the least significant bit of the pixel. This encryption method is too simple and easy to be cracked.

[0005] It can be seen that how to desensitize GIS spatial data more covertly and quickly trace the source of leakage after GIS spatial data leakage are technical problems that need to be solved at present. Summary of the invention

[0006] In order to solve the above technical problems, the present invention provides a spatial data information desensitization processing method based on a generative adversarial model, the method comprising the following steps: Construct a generative adversarial neural network and train the generative adversarial neural network; wherein, the generative adversarial neural network includes a generator model and a discriminator model, the generator model is composed of a convolutional layer, a dilated convolutional layer, and a transposed convolutional layer, the discriminator model is composed of an overall image discriminator model and a local image discriminator model, both the overall image discriminator model and the local image discriminator model are implemented using CNN, and the local image discriminator model has one less convolutional layer than the overall image discriminator model; Use the trained generative adversarial neural network to desensitize the spatial data image to be processed to obtain a desensitized spatial data image; Use the fast Fourier transform method to add a digital watermark to the desensitized spatial data image; When the desensitized spatial data image is leaked, use a watermark extraction algorithm to extract the first digital watermark from the leaked desensitized spatial data image, perform similarity matching between the extracted first digital watermark and all the second digital watermarks in the original watermark dataset, and determine the desensitized spatial data image corresponding to the second digital watermark with the highest matching degree as the leakage source.

[0007] Optionally, the convolutional layer in the generator model is used to reduce the resolution of the spatial data image to be completed, the dilated convolutional layer is used to perform image completion processing on the spatial data image with reduced resolution, and the transposed convolutional layer is used to restore the completed spatial data image to the original resolution.

[0008] Optionally, the overall image discriminator in the image discriminator model consists of 6 convolutional layers and 1 fully connected layer, which compresses the entire spatial data image to 256×256 as the input, and the output is a 1024-dimensional vector. All convolutional layers use a 5×5 convolutional kernel and use a 2×2 stride to reduce the resolution of the image; the input of the local image discriminator is a 128×128 image patch centered on the completed area, and it consists of 5 convolutional layers and 1 fully connected layer; And, connect the outputs of the overall image discriminator and the local image discriminator together to generate a 2048-dimensional vector, and then after being processed by the sigmoid function, obtain a value in the range of 0 to 1, and this value is the probability that the corresponding spatial data image is the original image.

[0009] Optionally, the training of the generative adversarial neural network includes: Represent the generator model as G(z,θg); where z is the spatial data image to be completed, θg is the parameter in the generator model, and G(z) is the output of the generator model; The discriminator model is denoted as D(x,θd); where x is the spatial data image after completion by the generator model, θd is the parameter in the discriminator model, and D(x) is the output of the discriminator model. The generator model is trained using a first loss function, which is a loss function that combines the introduction of MRF and MSE losses. Specifically: L ( z , x , θg ) = L ( z , θg ) + E ( x ) In the formula, L ( z , θg ) is the MSE loss function, E ( x ) is the MRF loss function; The MSE loss function is specifically: L ( z , θg ) = || θg × ( G ( z , θg ) - z ) ||2 In the formula, × represents the convolution operation, and ||∙|| represents the Euclidean norm; The MRF loss function is specifically: ; In the formula, is the pixel point in the completed area, is the pixel point in the intact area, is the weight parameter that makes the completed area and the intact area of the spatial data image consistent; The discriminator model is trained using the GAN loss function.

[0010] Optionally, the convolution calculation formula of the convolution layer is expressed as: ; In the formula, (i,j) represents the position of the pixel on the image, and h(k,l) represents the size of the convolution kernel; The dilated convolution calculation formula of the dilated convolution layer is expressed as: ; Where Xp,q represents the pixel component of the input layer, Yp,q represents the pixel component of the output layer, Wi,j represents the size of the convolution kernel, B represents the bias term, × represents the convolution operation, η represents the dilation coefficient, and f() represents the activation function; The input of the fully connected layer is the two-dimensional image features extracted by the convolutional layer. Convolution is performed using a convolutional kernel of the same size as the extracted two-dimensional image features. The output of the fully connected layer is a one-dimensional vector composed of individual nodes.

[0011] Optionally, adding a digital watermark to the desensitized spatial data image using the fast Fourier transform method includes: Preprocessing the digital watermark to be embedded, including at least scrambling the watermark information, radix conversion, and fast Fourier transform; Selecting the embedding position using the method of non-repeating pseudo-random sequences: Using the key key obtained after scrambling in the preprocessing part of the digital watermark as a seed to generate non-repeating pseudo-random numbers, the number of which is the number of data in the complex sequence, and the range of the random numbers is the number of points in the desensitized spatial data image used. After obtaining the non-repeating random number sequence, find the abscissas of the points corresponding to each element in the random number sequence in the desensitized spatial data image and record them in a one-dimensional array. Each element in the array is the embedding position of the watermark; Embedding the watermark information using a non-blind watermark algorithm: Performing a fast Fourier transform on the elements in the obtained one-dimensional array recording the embedding positions to obtain frequency-domain data; Using the additive criterion, performing corresponding operations on the watermark data obtained after the fast Fourier transform in the watermark preprocessing and the frequency-domain data to obtain a new complex sequence; Then performing an inverse fast Fourier transform to obtain data containing the digital watermark; Replacing the abscissas according to the corresponding positions of the non-repeating random sequence, obtaining data containing watermark information, and using the real part to replace the abscissas of the corresponding position points in the desensitized spatial data image to obtain a desensitized spatial data image of a vector containing the digital watermark.

[0012] Optionally, extracting the first digital watermark from the leaked desensitized spatial data image using a watermark extraction algorithm includes: Using the key key obtained during watermark scrambling in the watermark preprocessing as a seed to generate a non-repeating pseudo-random number sequence, finding the corresponding points in the desensitized spatial data image according to the data in the pseudo-random number sequence, and recording the relevant information in a one-dimensional array to obtain a coordinate sequence that should theoretically contain the watermark; After modifying the repeated ordinates in the same way as when embedding the digital watermark, traversing all the points in the watermark-containing data using the ordinates of the elements in the coordinate sequence, and extracting the points with the same ordinate and a small error between the abscissa of the watermark-containing data and the abscissa of the data in the coordinate sequence, and recording them in the extraction sequence; Multiply the obtained watermarked data by the sufficiently large number used in the watermark preprocessing process and round it, then convert the obtained decimal number to obtain a binary sequence. Encode it into a binary image of 128 pixels * 32 pixels according to this sequence, and then divide the binary image into four parts and perform Arnold scrambling respectively, and then splice them together to obtain the final first digital watermark.

[0013] Optionally, the step of performing similarity matching between the extracted first digital watermark and all the second digital watermarks in the original watermark dataset, and determining the desensitized spatial data image corresponding to the second digital watermark with the highest matching degree as the leakage source includes: Perform similarity matching between the extracted first digital watermark and all the second digital watermarks in the original watermark dataset to obtain the corresponding similarity; Set a threshold p. If the similarity ≥ p, it is considered that the target data contains digital watermark information, that is, the first digital watermark matches the corresponding second digital watermark. If the similarity < p, it is considered that the target data does not contain digital watermark information, that is, the first digital watermark does not match the corresponding second digital watermark; Determine the desensitized spatial data image corresponding to the second digital watermark with the highest matching degree as the leakage source.

[0014] The beneficial effects of the present invention are as follows: The present invention uses the generative adversarial neural network completion algorithm to desensitize the spatial data image. The generated desensitized spatial data image can achieve the effect of being more like real but not real, and completely desensitize the spatial information. At the same time, the present invention also uses the digital watermark made by the fast Fourier transform method, which has better concealment, cannot be cracked, and can quickly locate the leakage source after the desensitized spatial data image is leaked. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 is an example of desensitization of a spatial data image; Figure 2 is a schematic flowchart of a method for desensitizing spatial data information based on a generative adversarial model disclosed in an embodiment of the present invention; Figure 3 is a schematic diagram of a convolution operation; Figure 4 is a schematic diagram of the input area of a feature map; Figure 5 It is a schematic diagram of the model training process. Specific implementation manners

[0017] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0018] In addition, the technical features involved in different implementation manners of the present application described below can be combined with each other as long as they do not conflict with each other.

[0019] As Figure 2 shown, an embodiment of the present invention discloses a method for desensitizing spatial data information based on a generative adversarial model. The method includes the following steps: Construct a generative adversarial neural network and train the generative adversarial neural network; wherein, the generative adversarial neural network includes a generator model and a discriminator model. The generator model is composed of a convolutional layer, a dilated convolutional layer, and a transposed convolutional layer. The discriminator model is composed of an overall image discriminator model and a local image discriminator model. Both the overall image discriminator model and the local image discriminator model are implemented using CNN, and the local image discriminator model has one less convolutional layer than the overall image discriminator model; Use the trained generative adversarial neural network to desensitize the spatial data image to be processed to obtain a desensitized spatial data image; Use the fast Fourier transform method to add a digital watermark to the desensitized spatial data image; When the desensitized spatial data image is leaked, use a watermark extraction algorithm to extract the first digital watermark from the leaked desensitized spatial data image, perform similarity matching between the extracted first digital watermark and all the second digital watermarks in the original watermark dataset, and determine the desensitized spatial data image corresponding to the second digital watermark with the highest matching degree as the leakage source.

[0020] The present invention uses a generative adversarial neural network completion algorithm to desensitize the spatial data image. The generated desensitized spatial data image can achieve a more realistic but not real effect, completely desensitizing the spatial information. At the same time, the present invention also uses a digital watermark made by the fast Fourier transform method, which has better concealment, cannot be cracked, and can quickly locate the leakage source after the desensitized spatial data image is leaked.

[0021] Optionally, the convolutional layer in the generator model is used to reduce the resolution of the spatial data image to be completed, the dilated convolutional layer is used to perform image completion processing on the spatial data image with reduced resolution, and the transposed convolutional layer is used to restore the spatial data image after completion processing to the original resolution.

[0022] In this embodiment, the generator model is composed of three parts: a convolutional layer, a dilated convolutional layer, and a transposed convolutional layer. Before image completion, the convolutional layer can be used to reduce the resolution of the image to be completed: the convolution stride of a part of the convolutional layers is 2, aiming to reduce the size of the input image to be completed to half of the original, so as to reduce the storage space and calculation time of the image. At the same time, a part of the convolutional layers with a convolution stride of 1 are also added to the generator model, aiming to extract more feature maps at some feature levels of the image to be completed. After reducing the resolution of the image to be completed, the dilated convolutional layer can be used for image completion. The difference in the dilation coefficient η is exactly the advantage of the dilated convolutional layer. Without increasing the computational consumption, it multiplies the surrounding area of the missing area that can be covered, and this is also the key for this method to complete high-resolution images. After the completion work is done, the transposed convolutional layer can be used to restore the completed image to the original resolution. It should be noted that considering the rationality of the overall texture structure of the image, the resolution of the image can be reduced to 1 / 4 of the original.

[0023] Optionally, the overall image discriminator in the image discriminator model consists of 6 convolutional layers and 1 fully connected layer. It compresses the entire spatial data image to 256×256 as the input, and the output is a 1024-dimensional vector. All convolutional layers use a 5×5 convolutional kernel and a 2×2 stride to reduce the resolution of the image; the input of the local image discriminator is a 128×128 image patch centered on the completion area, and it consists of 5 convolutional layers and 1 fully connected layer; Moreover, the outputs of the overall image discriminator and the local image discriminator are connected together to generate a 2048-dimensional vector, and then through the processing of the sigmoid function, a value in the range of 0 to 1 is obtained. This value is the probability that the corresponding spatial data image is the original image.

[0024] In this embodiment, the image discriminator model is divided into two parts, namely the global image discriminator model and the local image discriminator model. Using these two discriminator models, it is possible to determine whether an image is an original image or a completed image. The discriminator model is implemented using CNN. First, the convolutional layer is used to continuously compress the image, and then the fully connected layer is used to classify the image to determine the authenticity of the image. The global image discriminator consists of 6 convolutional layers and 1 fully connected layer. It compresses the entire image to 256×256 as the input, and the output is a 1024-dimensional vector. All convolutional layers use a 5×5 convolutional kernel and a 2×2 stride to reduce the resolution of the image.

[0025] The local image discriminator follows the same pattern as the global image discriminator, but the input of the local image discriminator is a 128×128 image patch centered on the completed region. Since the resolution of the input image is half that of the global image discriminator, in terms of the composition structure, the local image discriminator can reduce one convolutional layer.

[0026] Finally, the outputs of the global and local image discriminators are concatenated to generate a 2048-dimensional vector, which is then processed by the sigmoid function to obtain a value in the range of 0 to 1. This value is the probability that the image is an original image.

[0027] Optionally, training the generative adversarial neural network includes: Represent the generator model as G(z,θg); where z is the spatial data image to be completed, θg is the parameter in the generator model, and G(z) is the output of the generator model; Represent the discriminator model as D(x,θd); where x is the spatial data image completed by the generator model, θd is the parameter in the discriminator model, and D(x) is the output of the discriminator model; Train the generator model using the first loss function, and the first loss function is a loss function that combines the introduction of MRF and MSE losses. Specifically: L ( z , x , θg ) = L ( z , θg ) + E ( x ) In the formula, L ( z , θg ) is the MSE loss function, E ( x ) is the MRF loss function; The specific MSE loss function is as follows: L ( z , θg ) = || θg ×( G ( z , θg ) - z )||2 In the formula, × represents the convolution operation, and ||∙|| represents the Euclidean norm; The specific MRF loss function is as follows: ; In the formula, is the pixel point in the completed area, is the pixel point in the intact area, is the weight parameter that makes the completed area and the intact area of the spatial data image consistent; The discriminator model is trained using the GAN loss function.

[0028] In this embodiment, in order to make the reconstructed and supplemented spatial data image after desensitization more like authenticity rather than real, on the basis of using the GAN loss function to train the model, a loss function combining MRF and MSE losses is introduced to train the generator model. The advantages of the two loss functions are combined with each other, and a more stable high-performance network model can be trained.

[0029] The loss function for training the generator model includes two parts: MRF and MSE losses.

[0030] On the basis of training the generator model using the MSE loss function, in order to further improve the ability to complete the image, the energy function of MRF is added to the loss function. The addition of this energy function keeps the pixels in the completed area continuous with their surrounding pixels in terms of image features with the highest probability. The coherence of the image features between the completed area and the intact area is continuously strengthened, making the loss function decline rapidly and further improving the training effect of the generator model.

[0031] The training of the discriminator model is completed through the GAN loss function. The GAN loss function directly transforms the adversarial process between the generator model and the discriminator model into a min-max problem, so that in each adversarial process, the generator model and the discriminator model are jointly updated.

[0032] During the process of model training optimization, the states of the generator model and the discriminator model are constantly changing, which actually means that the parameters of the convolutional kernels in the model are constantly changing until the model reaches the optimal state. The model training optimization process can be divided into two stages: in the first stage, the discriminator model is updated using the GAN loss function; in the second stage, the generator model is updated using a loss function that combines MRF and MSE. The model training optimization process is as Figure 3 shown.

[0033] Optionally, the convolution calculation formula of the convolutional layer is expressed as: ; where (i,j) represents the position of the pixel on the image, and h(k,l) represents the size of the convolutional kernel; The dilated convolution calculation formula of the dilated convolutional layer is expressed as: ; where Xp,q represents the pixel component of the input layer, Yp,q represents the pixel component of the output layer, Wi,j represents the size of the convolutional kernel, B represents the bias term, × represents the convolution operation, η represents the dilation coefficient, and f() represents the activation function; The input of the fully connected layer is the two-dimensional image features extracted by the convolutional layer. A convolutional kernel of the same size as the extracted two-dimensional image features is used for convolution, and the output of the fully connected layer is a one-dimensional vector composed of individual nodes.

[0034] In this embodiment, the Convolutional Neural Network (CNN), as a type of deep neural network, has outstanding advantages in image processing. The basic functions of the CNN are divided into two parts: the feature extraction layer and the classifier. The feature extraction layer is mainly used to extract image features layer by layer, and the main task of the classifier is to summarize and classify the extracted image features. For the reconstruction of the spatial data desensitization region information, the feature extraction layer consists of a convolutional layer and a dilated convolutional layer, and the classifier consists of a fully connected layer.

[0035] The convolutional layer is an important part of the convolutional neural network and an important means for image feature extraction. The local connectivity and weight sharing of the convolutional layer can well help the convolutional neural network process large-size images. The operation process of convolution is as Figure 4 shown.

[0036] The dilated convolutional layer is a variant of the convolutional layer. This convolutional layer increases the input area of the feature map while keeping the number of weights unchanged, and at the same time ensures that the size of the output feature map remains unchanged. When the dilation coefficient η is different, the input area of the feature map that the dilated convolutional layer can read is different. The input areas of the feature map under different dilation coefficients are as Figure 5 shown.

[0037] On the basis of using a convolutional layer to extract image features, a fully connected layer is adopted to classify the image features extracted by the convolutional layer. The input of the fully connected layer is the two-dimensional image features extracted by the convolutional layer. A convolutional kernel of the same size as the extracted two-dimensional image features is used for convolution. The output of the fully connected layer is a one-dimensional vector composed of individual nodes.

[0038] Optionally, the method of adding a digital watermark to the desensitized spatial data image using the fast Fourier transform method includes: Preprocess the digital watermark to be embedded, including at least scrambling the watermark information, base conversion, and fast Fourier transform; Adopt the method of non-repeating pseudo-random sequences to select the embedding positions: use the key key obtained after scrambling in the preprocessing part of the digital watermark as a seed to generate non-repeating pseudo-random numbers, the number of which is the number of data in the complex sequence, and the range of the random numbers is the number of points in the used desensitized spatial data image. After obtaining the non-repeating random number sequence, find the abscissas of the points corresponding to each element in the random number sequence in the desensitized spatial data image and record them in a one-dimensional array. Each element in the array is the embedding position of the watermark; Adopt a non-blind watermarking algorithm to embed the watermark information: Perform a fast Fourier transform on the elements in the obtained one-dimensional array recording the embedding positions to obtain frequency-domain data; use the additive criterion to perform corresponding operations on the watermark data obtained after the fast Fourier transform in the watermark preprocessing and the frequency-domain data to obtain a new complex sequence; then perform an inverse fast Fourier transform to obtain the data containing the digital watermark; Replace the abscissas according to the corresponding positions of the non-repeating random sequence, obtain the data containing the watermark information, and use the real part to replace the abscissas of the corresponding positions of the points in the desensitized spatial data image to obtain the desensitized spatial data image of the vector containing the digital watermark.

[0039] In this embodiment, during the process of embedding the digital watermark, if Fourier transforms are performed on all desensitized spatial data images, the speed of embedding the digital watermark will be very slow, reducing work efficiency. Therefore, the present invention adopts the method of non-repeating pseudo-random sequences to select the embedding positions.

[0040] At the same time, the present invention adopts a non-blind watermarking algorithm, which requires the participation of the desensitized spatial data image during detection, and uses the ordinate of the point as the detection identifier. In order to prevent the situation where the ordinates are the same in the same group of data, the repeated ordinates need to be modified. Therefore, when the ordinate sequence of the selected embedding points itself appears repeated or the ordinate of the selected embedding points appears repeated with the ordinates of the non-watermarked part, make slight changes to the later-appearing ordinates to minimize the impact on the original data.

[0041] It should be noted that by performing the above preprocessing on the digital watermark, the watermark information can be encoded to adapt to the algorithm. Assume that the original watermark image is a 128-pixel * 32-pixel binary image, which can satisfy the information encoding combination of about 15 characters.

[0042] The Arnold scrambling algorithm is used for watermark information scrambling, also known as cat face scrambling. The main process is to multiply the horizontal and vertical coordinates of each pixel point of the original image by a matrix to obtain the new coordinates of each pixel point. The original 128 * 32-pixel image is divided into four parts, each part being 32 pixels * 2 pixels. Arnold scrambling is performed on each part, and then the scrambled watermark information of the four parts is combined together to obtain the scrambled watermark information, and the number of scrambling times is recorded as the key key. The Arnold scrambling algorithm is as follows: ; The purpose of watermark information base conversion is to reduce the data volume of the watermark information and reduce the impact of the watermark information on the desensitized spatial data image. The specific method is to convert the scrambled watermark information into binary representation in row-column order and encode it into a one-dimensional sequence, and then combine every eight binary bits in the sequence into a decimal number. Divide the obtained decimal number by a sufficiently large number to make it less than 0.01 to achieve the effect of reducing the impact.

[0043] Perform a fast Fourier transform on the watermark information after base conversion. The specific method is to perform a fast Fourier transform on the obtained decimal sequence to obtain a complex number sequence.

[0044] ; Among them, 0 indicates that the imaginary part in the complex number is 0, DArray represents the newly obtained complex number sequence, k and l are the sequence numbers after the fast Fourier transform, and i represents the imaginary unit.

[0045] Optionally, the step of extracting the first digital watermark from the leaked desensitized spatial data image by using the watermark extraction algorithm includes: Using the key key obtained during watermark scrambling in watermark preprocessing as a seed to generate a non-repeating pseudo-random number sequence, finding the corresponding points in the desensitized spatial data image according to the data in the pseudo-random number sequence, and recording the relevant information in a one-dimensional array to obtain a coordinate sequence that should theoretically contain the watermark; After modifying the repeated vertical coordinates by the same method as when embedding the digital watermark, traverse all the points in the watermark-containing data using the vertical coordinates of the elements in the coordinate sequence, and extract the points with the same vertical coordinates and the horizontal coordinates of the watermark-containing data having a small error from the horizontal coordinates of the data in the coordinate sequence, and record them in the extraction sequence; Multiply the obtained watermarked data by a sufficiently large number (i.e., a number greater than the preset value) used in the watermark preprocessing process and round it, then convert the obtained decimal number to obtain a binary sequence, encode it into a binary image of 128 pixels * 32 pixels according to this sequence, and then divide the binary image into four parts and perform Arnold scrambling on each part respectively, and then splice them together to obtain the final first digital watermark.

[0046] In this embodiment, the process of extracting watermark information is the reverse process of embedding watermark information. Since the desensitized spatial data image may encounter operations such as addition, deletion, and movement during use, some watermark information may not be extracted. Therefore, the present invention adopts a method of information filling to fill the extraction sequence completely. For a sequence with null values, traverse it from front to back. If a null value is encountered, make the value at this position the same as the previous value until the traversal ends. Then, traverse the sequence from back to front again. If there are still null values, make the value at this position the same as the next value until the traversal ends. In this way, a sequence without null values and with a numerical curve approximately similar to the original numerical curve can be obtained. The purpose of the second traversal is to prevent the situation where the first few values are null values.

[0047] Optionally, the process of performing similarity matching on the extracted first digital watermark with all the second digital watermarks in the original watermark dataset and determining the desensitized spatial data image corresponding to the second digital watermark with the highest matching degree as the leakage source includes: Perform similarity matching on the extracted first digital watermark with all the second digital watermarks in the original watermark dataset to obtain the corresponding similarity; Set a threshold p. If the similarity ≥ p, it is considered that the target data contains digital watermark information, that is, the first digital watermark matches the corresponding second digital watermark. If the similarity < p, it is considered that the target data does not contain digital watermark information, that is, the first digital watermark does not match the corresponding second digital watermark; Determine the desensitized spatial data image corresponding to the second digital watermark with the highest matching degree as the leakage source.

[0048] In this embodiment, after obtaining the final watermark image, it is necessary to match it with all the images in the watermark set. Set a threshold p. If the similarity ≥ p, it is considered that the target data contains digital watermark information. If the similarity < p, it is considered that the target data does not contain digital watermark information. Output the digital watermark with the highest matching degree as the final result.

[0049] It should be understood that various forms of the above - shown processes can be used, re - ordering, adding, or deleting steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0050] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present disclosure is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples may be omitted or replaced by their equivalent elements. In addition, the steps may be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein may be replaced by equivalent elements that emerge after the present disclosure.

Claims

1. A spatial data information desensitization processing method based on a generative adversarial model, characterized in that: The method comprises the following steps: Constructing a generative adversarial neural network and training the generative adversarial neural network; wherein the generative adversarial neural network includes a generator model and a discriminator model, the generator model is composed of a convolution layer, a dilated convolution layer, and a deconvolution layer, the discriminator model is composed of an overall image discriminator model and a local image discriminator model, the overall image discriminator model and the local image discriminator model are both implemented using CNN, and the local image discriminator model has one less convolution layer than the overall image discriminator model; Using the trained generative adversarial neural network to perform desensitization processing on the spatial data image to be processed, to obtain a desensitized spatial data image; Use fast Fourier transform method to add digital watermark to desensitized spatial data image; When a desensitized spatial data image is leaked, a watermark extraction algorithm is used to extract the first digital watermark from the leaked desensitized spatial data image. The extracted first digital watermark is similarly matched with all the second digital watermarks in the original watermark data set, and the desensitized spatial data image corresponding to the second digital watermark with the highest matching degree is determined as the source of the leak.

2. According to claim 1, a spatial data information desensitization processing method based on a generative adversarial model is characterized by: The convolution layer in the generator model is used to reduce the resolution of the spatial data image to be completed, the dilated convolution layer is used to perform image completion processing on the spatial data image after the resolution is reduced, and the deconvolution layer is used to restore the spatial data image after the completion processing to the original resolution.

3. The method for desensitizing spatial data information based on a generative adversarial model according to claim 2, characterized in that: The overall image discriminator in the image discriminator model consists of 6 convolutional layers and 1 fully connected layer, which compresses the entire spatial data image to 256×256 as input, and outputs a 1024-dimensional vector. All convolutional layers use 5×5 convolution kernels and use a 2×2 stride to reduce the resolution of the image; the input of the local image discriminator is a 128×128 image block centered on the completion area, which consists of 5 convolutional layers and 1 fully connected layer; Furthermore, the outputs of the overall image discriminator and the local image discriminator are connected together to generate a 2048-dimensional vector, which is then processed by a sigmoid function to obtain a value in the range of 0 to 1, which is the probability that the corresponding spatial data image is the original image.

4. The method for desensitizing spatial data information based on a generative adversarial model according to claim 3, characterized in that: Training the generative adversarial neural network includes: The generator model is represented by G(z,θg); wherein z is the spatial data image to be completed, θg is a parameter in the generator model, and G(z) is the output of the generator model; The discriminator model is represented by D(x,θd); wherein x is the spatial data image completed by the generator model, θd is a parameter in the discriminator model, and D(x) is the output of the discriminator model; The generator model is trained using a first loss function, where the first loss function is a loss function combining MRF and MSE losses, specifically: L ( z , x , θg ) = L ( z , θg )+ E ( x ) In the formula, L ( z , θg ) is the MSE loss function, E ( x ) is the MRF loss function; The MSE loss function is specifically: L ( z , θg ) = || θg ×( G ( z , θg ) - z )||2 In the formula, × represents the convolution operation, ||∙|| represents the Euclidean normal form; The MRF loss function is specifically: ; In the formula, is the pixel in the completed area, is the pixel in the intact area, It is a weight parameter that makes the completed area and the intact area of ​​the spatial data image consistent; The discriminator model is trained using the GAN loss function.

5. The method for desensitizing spatial data information based on a generative adversarial model according to claim 4, characterized in that: The convolution calculation formula of the convolution layer is expressed as: ; Where (i, j) represents the position of the pixel on the image, and h(k, l) represents the size of the convolution kernel; The calculation formula of the dilated convolution layer is expressed as: Where Xp,q represents the pixel component of the input layer, Yp,q represents the pixel component of the output layer, Wi,j represents the size of the convolution kernel, B represents the bias term, × represents the convolution operation, η represents the expansion coefficient, and f() represents the activation function; The input of the fully connected layer is the two-dimensional image features extracted by the convolutional layer. Convolution is performed using a convolution kernel of the same size as the extracted two-dimensional image features. The output of the fully connected layer is a one-dimensional vector composed of nodes.

6. The method for desensitizing spatial data information based on a generative adversarial model according to claim 5, characterized in that: Use the fast Fourier transform method to add digital watermarks to desensitized spatial data images, including: Preprocessing the digital watermark to be embedded, including at least watermark information scrambling, base conversion, and fast Fourier transform; The embedding position is selected by using a non-repeating pseudo-random sequence method: the key obtained after scrambling in the digital watermark preprocessing part is used as a seed to generate a non-repeating pseudo-random number, the number of which is the number of data in the complex sequence, and the range of the random number is the number of points in the desensitized spatial data image used. After obtaining a non-repeating random number sequence, the horizontal coordinates of the points corresponding to each element in the random number sequence are found in the desensitized spatial data image and recorded in a one-dimensional array. Each element in the array is the embedding position of the watermark; Use non-blind watermark algorithm to embed watermark information: Perform fast Fourier transform on the elements in the one-dimensional array of the obtained record embedding position to obtain frequency domain data; use the additive criterion to perform corresponding operations on the watermark data obtained after fast Fourier transform in watermark preprocessing and the frequency domain data to obtain a new complex number sequence; then perform inverse fast Fourier transform to obtain data containing digital watermark; The horizontal coordinate is replaced according to the corresponding position of the non-repeating random sequence to obtain data containing watermark information, and the horizontal coordinate of the corresponding position point in the desensitized spatial data image is replaced with the real part to obtain a desensitized spatial data image of the vector containing the digital watermark.

7. The method for desensitizing spatial data information based on a generative adversarial model according to claim 6, characterized in that: The first digital watermark is extracted from the leaked desensitized spatial data image using a watermark extraction algorithm, including: The key obtained during the watermark scrambling in the watermark preprocessing is used as a seed to generate a non-repeating pseudo-random number sequence. According to the data in the pseudo-random number sequence, the corresponding points are found in the desensitized spatial data image, and the relevant information is recorded in a one-dimensional array to obtain the coordinate sequence that should theoretically contain the watermark. After modifying the repeated ordinates in the same way as when embedding a digital watermark, all points in the watermarked data are traversed using the ordinates of the elements in the coordinate sequence, and points with the same ordinates and a small error between the abscissas of the watermarked data and the abscissas of the data in the coordinate sequence are extracted and recorded in the extraction sequence; The obtained watermarked data is multiplied by a sufficiently large number used in the watermark preprocessing process and rounded, and then the obtained decimal number is converted to a binary sequence, which is encoded into a binary image of 128 pixels * 32 pixels according to the sequence. The binary image is then divided into four parts and Arnold scrambled separately, and then spliced ​​together to obtain the final first digital watermark.

8. The method for desensitizing spatial data information based on a generative adversarial model according to claim 7, characterized in that: The extracted first digital watermark is similarly matched with all the second digital watermarks in the original watermark data set, and the desensitized spatial data image corresponding to the second digital watermark with the highest matching degree is determined as the source of the leakage, including: Performing similarity matching between the extracted first digital watermark and all the second digital watermarks in the original watermark data set to obtain corresponding similarities; Set a threshold p. If the similarity is ≥p, it is considered that the target data contains digital watermark information, that is, the first digital watermark matches the corresponding second digital watermark. If the similarity <p, it is considered that the target data does not contain digital watermark information, that is, the first digital watermark does not match the corresponding second digital watermark; Determine the desensitized spatial data image corresponding to the second digital watermark with the highest matching degree as the leakage source.