Real-scene three-dimensional model pattern texture desensitization method and system and medium
By combining the technical means of Mask R-CNN and U-NetGAN, automatic detection and texture repair of sensitive content in the three-dimensional model texture image is achieved, and the problems of 3D model data desensitization efficiency and accuracy in the existing technology are solved, and efficient and automated data desensitization and high-quality texture repair are achieved.
Patent Information
- Application Number
- CN202510423050.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to efficiently and automatically desensitize textured three-dimensional models in three-dimensional geographic information systems, especially to ensure data privacy protection while maintaining data accuracy and availability.
Deep learning-based instance segmentation networks (such as Mask R-CNN) are used to detect and segment sensitive content in three-dimensional model texture images, and texture repair is used to generate desensitized texture images that seamlessly fuse with surrounding non-sensitive areas.
The automation, high precision and high fidelity desensitization of three-dimensional model data is realized, ensuring the security and availability of desensitized data, promoting the secure sharing of real-life three-dimensional model data, and providing important potential for digital twins and smart city applications.
Smart Images

Figure CN119941583A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern texture desensitization, and in particular relates to a method, system and medium for desensitizing pattern texture of a real-scene three-dimensional model. Background Art
[0002] With the development of smart cities, 3D GIS plays an increasingly important role in the visualization and analysis of urban infrastructure and services. As a basic data resource, textured 3D models are essential for the construction of 3D GIS and smart city applications. However, these models may contain sensitive or private content such as personal information, commercial logos, or security-related signs. Therefore, it is of great significance to effectively process the sensitive information in the texture of real-life 3D models before data distribution to reduce the risk of information leakage.
[0003] Desensitization is a cutting-edge data protection method that achieves a balance between data privacy protection and availability while ensuring data integrity by modifying or eliminating sensitive content. Currently, existing data desensitization methods can be mainly divided into three categories: dynamic data desensitization, obfuscation-based desensitization, and image restoration-based desensitization.
[0004] Dynamic data desensitization achieves real-time data protection by dynamically adjusting data visibility based on user roles and access rights. The core of the method lies in access control and encryption algorithms. At present, although the existing technology has strong resistance to statistical attacks and brute force cracking. However, although this type of method can effectively restrict unauthorized access, sensitive textures can still be decrypted with the correct key, thereby exposing sensitive content during rendering or interaction.
[0005] Desensitization methods based on obfuscation usually smooth or blur texture images, or slightly offset or simplify coordinate points, to retain data availability to a certain extent to support data analysis. Currently, there are existing technologies that use neighboring locations with similar geographical features to replace the original location to achieve location protection. However, the disturbance of such methods often reduces spatial and texture accuracy, which may affect the availability of data. Summary of the invention
[0006] In order to solve the above technical problems, the present invention proposes a method, system and medium for desensitizing the pattern texture of a real-life three-dimensional model, which realizes the automation, high precision and high fidelity of three-dimensional model data desensitization, fills the gap in the existing technology, and provides a reliable solution for the data security of smart cities.
[0007] To achieve the above-mentioned object, the present invention provides a method for desensitizing a pattern texture of a real-scene three-dimensional model, comprising: mapping texture information of the real-scene three-dimensional model to a two-dimensional image to generate a corresponding texture image; Detecting sensitive content in the texture image through a deep learning-based instance segmentation network, and generating a pixel-level segmentation mask to mark sensitive areas; Using a generative adversarial network to perform texture restoration on the sensitive area, and generate a desensitized texture image that is seamlessly integrated with the surrounding non-sensitive areas; The repaired texture image is remapped to the real-scene three-dimensional model to obtain desensitized three-dimensional model data; wherein the instance segmentation network adopts a two-stage detection framework, combining the region proposal network and the feature pyramid network to achieve multi-scale sensitive target recognition; the generative adversarial network includes a generator of an encoder-decoder structure, which fuses multi-level features through jump connections, and a discriminator constructed by a convolutional neural network, which is used for adversarial training to optimize the output quality of the generator.
[0008] Optionally, the process of mapping the texture information of the real-scene three-dimensional model to the two-dimensional image includes: Traversing the tree structure nodes of the three-dimensional model to extract triangular facets containing texture information; According to the texture coordinates of the triangular patch, the pixel values of the surface of the three-dimensional model are mapped to a blank image; The tree structure is traversed hierarchically, and only branches whose parent nodes contain sensitive content are deeply traversed.
[0009] Optionally, the instance segmentation network is Mask R-CNN, and the backbone network is ResNet50; The detection process includes: Generate candidate regions through the region proposal network and perform bilinear interpolation alignment on the candidate regions; Predict bounding box parameters, category confidence, and pixel-level segmentation mask of candidate regions; Sensitive targets are screened according to the preset sensitive pattern library and confidence threshold to generate a mask image.
[0010] Optionally, the generator of the generative adversarial network is a U-Net structure, including: The encoder gradually extracts features through multi-layer convolution and maximum pooling; The decoder restores the image resolution by upsampling and performs feature fusion with the corresponding layer of the encoder through skip connections; The authenticity of the repaired image output by the generator is determined by the discriminator to drive adversarial training.
[0011] Optionally, the process of generating a mask image further includes: A dilation operation is performed on the pixel-level segmentation mask so that the boundary of the sensitive area extends outward by a preset pixel width.
[0012] Optionally, the adversarial training process of the generative adversarial network includes: Inputting the mask image and the original texture image into the generator to generate a repaired image; The discriminator receives the repaired image and the real non-sensitive image, and outputs a probability of authenticity; The parameters of the generator are adjusted according to the feedback of the discriminator until the difference between the repaired image and the real image is lower than a preset threshold.
[0013] A system comprises: a memory, a processor, and a real-scene 3D model pattern texture desensitization program stored in the memory and executable on the processor, wherein the real-scene 3D model pattern texture desensitization program implements the steps of the real-scene 3D model pattern texture desensitization method when executed by the processor.
[0014] A computer-readable storage medium stores a real-scene three-dimensional model pattern texture desensitization program, and when the real-scene three-dimensional model pattern texture desensitization program is executed by a processor, the steps of the real-scene three-dimensional model pattern texture desensitization method are implemented.
[0015] Technical effect of the invention: The invention discloses a method, system and medium for desensitizing the texture of a real-life 3D model pattern, which combines Mask R-CNN to realize automatic detection and segmentation of sensitive content, and uses a generative adversarial network (GAN) based on U-Net to perform realistic texture restoration, thereby realizing automated detection and high-quality texture restoration. The invention can effectively perform automated and accurate data desensitization while achieving high-quality texture restoration. This ensures the security and availability of desensitized data, promotes the secure sharing of real-life 3D model data, and provides important potential for digital twin and smart city applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A schematic diagram of a process of a method for desensitizing a pattern texture of a real-scene three-dimensional model according to an embodiment of the present invention; Figure 2 Schematic diagram of the structure of Mask R-CNN according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the structure of U-NetGAN according to an embodiment of the present invention; Figure 4 A schematic diagram of further selective traversal of real-scene three-dimensional data according to an embodiment of the present invention; Figure 5Schematic diagram of street scene images under extreme conditions according to an embodiment of the present invention, where (a) is noise, (b) is rotation, (c) is long distance, (d) is partial occlusion, (e) is dim light, and (f) is complex background; Figure 6 Schematic diagram of the detection results of sensitive shapes in the test set of an embodiment of the present invention, where (a) is MobileNetv2+Mask R-CNN, (b) is the YOLO+Patch Match algorithm, and (c) is the algorithm proposed by the present invention; Figure 7 The figure is a schematic diagram of the repair result of an example of repair through qualitative evaluation according to an embodiment of the present invention. DETAILED DESCRIPTION
[0017] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0018] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0019] Image restoration-based desensitization is a static desensitization technology that conceals sensitive areas through realistic reconstruction, taking into account both privacy protection and visual integrity. Compared with the desensitization method based on obfuscation, this method can ensure that sensitive content is not leaked while maintaining the realism of the data. Therefore, it has a high application value in the field of GIS data desensitization. For example, the existing technology uses convolutional neural networks (CNN) to automatically erase sensitive content in high-quality images, effectively preventing privacy leakage. However, the existing desensitization methods based on image restoration are mainly aimed at two-dimensional data, and it is difficult to handle the structural complexity of textured three-dimensional models.
[0020] In summary, dynamic data desensitization can limit data access through permission control, but there is still a risk of sensitive information leakage in some cases. Desensitization methods based on obfuscation reduce security risks by directly modifying data, but usually sacrifice data accuracy and availability. Therefore, these two types of methods are less applicable in data privacy protection. In contrast, desensitization algorithms based on image restoration achieve a better balance between data openness and data accuracy. However, most of the existing image restoration-based methods are targeted at two-dimensional data and are not suitable for three-dimensional models with textures. In addition, current desensitization methods for three-dimensional model textures still rely on a lot of manual processing, which limits their application in urban scene data distribution. How to design an automated and efficient three-dimensional model data desensitization algorithm based on image restoration has become an urgent problem to be solved.
[0021] In recent years, deep learning has demonstrated excellent performance in computer vision tasks such as object detection and image restoration. Some existing technologies have attempted to use deep learning to improve the automation and quality of desensitization algorithms based on image restoration. For example, You Only Look Once version 8 (YOLOv8) is used for object detection, and texture restoration is performed manually. However, this semi-automatic method is still inefficient in large-scale application scenarios. Similarly, You Only Look Once version 5 Small (YOLOv5s) is used to automatically detect sensitive targets, and multi-scale PatchMatch is used for texture restoration. Although the algorithm is optimized in terms of automation, the Patch Match method is prone to block artifacts, which affects the texture quality. In general, although deep learning has brought new opportunities for the study of desensitization algorithms based on image restoration for three-dimensional data, existing algorithms still have limitations in terms of automation and restoration quality.
[0022] like Figure 1 As shown, this embodiment provides a method for desensitizing a real-scene 3D model pattern texture, comprising: Mapping the texture information of the real-scene three-dimensional model to the two-dimensional image to generate a corresponding texture image; Detecting sensitive content in the texture image through a deep learning-based instance segmentation network, and generating a pixel-level segmentation mask to mark sensitive areas; Using a generative adversarial network to perform texture restoration on the sensitive area, and generate a desensitized texture image that is seamlessly integrated with the surrounding non-sensitive areas; The repaired texture image is remapped to the real-scene three-dimensional model to obtain desensitized three-dimensional model data; wherein the instance segmentation network adopts a two-stage detection framework, combining the region proposal network and the feature pyramid network to achieve multi-scale sensitive target recognition; the generative adversarial network includes a generator of an encoder-decoder structure, which fuses multi-level features through jump connections, and a discriminator constructed by a convolutional neural network, which is used for adversarial training to optimize the output quality of the generator.
[0023] Furthermore, the process of mapping the texture information of the real-scene three-dimensional model to the two-dimensional image includes: Traversing the tree structure nodes of the three-dimensional model to extract triangular facets containing texture information; According to the texture coordinates of the triangular patch, the pixel values of the surface of the three-dimensional model are mapped to a blank image; The tree structure is traversed hierarchically, and only branches whose parent nodes contain sensitive content are deeply traversed.
[0024] Furthermore, the instance segmentation network is Mask R-CNN, and the backbone network is ResNet50; The detection process includes: Generate candidate regions through the region proposal network and perform bilinear interpolation alignment on the candidate regions; Predict bounding box parameters, category confidence, and pixel-level segmentation mask of candidate regions; Sensitive targets are screened according to the preset sensitive pattern library and confidence threshold to generate a mask image.
[0025] Furthermore, the generator of the generative adversarial network is a U-Net structure, including: The encoder gradually extracts features through multi-layer convolution and maximum pooling; The decoder restores the image resolution by upsampling and performs feature fusion with the corresponding layer of the encoder through skip connections; The authenticity of the repaired image output by the generator is determined by the discriminator to drive adversarial training.
[0026] Furthermore, the process of generating the mask image also includes: A dilation operation is performed on the pixel-level segmentation mask so that the boundary of the sensitive area extends outward by a preset pixel width.
[0027] Furthermore, the adversarial training process of the generative adversarial network includes: Inputting the mask image and the original texture image into the generator to generate a repaired image; The discriminator receives the repaired image and the real non-sensitive image, and outputs a probability of authenticity; The parameters of the generator are adjusted according to the feedback of the discriminator until the difference between the repaired image and the real image is lower than a preset threshold.
[0028] A system comprises: a memory, a processor, and a real-scene 3D model pattern texture desensitization program stored in the memory and executable on the processor, wherein the real-scene 3D model pattern texture desensitization program implements the steps of the real-scene 3D model pattern texture desensitization method when executed by the processor.
[0029] A computer-readable storage medium stores a real-scene three-dimensional model pattern texture desensitization program, and when the real-scene three-dimensional model pattern texture desensitization program is executed by a processor, the steps of the real-scene three-dimensional model pattern texture desensitization method are implemented.
[0030] Data desensitization methods based on image restoration involve detecting and replacing sensitive content in three-dimensional models. The key to designing deep learning-driven desensitization methods is to achieve automatic detection and realistic restoration of targets. To this end, the algorithm uses Mask R-CNN for automatic target detection, and uses its high-precision pixel-level segmentation capabilities to ensure accurate identification and isolation of sensitive areas. Subsequently, a U-Net-enhanced generative adversarial network (U-NetGAN) was designed for high-quality texture restoration to ensure that the desensitized content can be seamlessly restored to a realistic appearance. Ultimately, the algorithm is able to generate non-sensitive real-life three-dimensional model data to ensure the secure distribution of data. Figure 1 The overall framework of the proposed algorithm is presented, which includes three key steps: data preprocessing based on texture mapping, sensitive content detection, and sensitive content regeneration.
[0031] The texture of 3D models is highly complex, and the scale of sensitive targets varies significantly. This complexity requires the detection algorithm to maintain high accuracy on multi-scale targets. Compared with single-stage detection algorithms, Mask R-CNN, as a two-stage target detection method, shows higher detection accuracy and is therefore more suitable for multi-scale target detection in real-life 3D models.
[0032] The structure of Mask R-CNN is as follows Figure 2As shown in Figure 1, it mainly consists of two stages. The first stage uses the region proposal network (RPN) to identify potential areas in the image that may contain objects and generate region proposals. The second stage finely classifies these proposed regions and calculates the mask of each object region. By adding a branch for predicting object masks on the basis of Faster R-CNN, Mask R-CNN can not only detect objects, but also generate pixel-level segmentation masks for each detected object. Therefore, Mask R-CNN can provide high-precision pixel-level segmentation results for automatic detection of sensitive objects.
[0033] Generative Adversarial Networks (GANs) are a type of generative model based on a game theory framework, consisting of two adversarial neural networks: the generator and the discriminator. The two compete with each other during the training process to achieve high-quality generation results. The generator extracts multi-scale features through the encoder and decoder, and regenerates the repaired image to generate a pseudo image similar to the real image. The discriminator consists of multiple convolutional layers and is responsible for distinguishing the generated image from the real image. In this adversarial training process, the two models are continuously optimized, and ultimately generate high-resolution images with realistic visual effects.
[0034] Based on this feature of GAN, it can be used for high-quality restoration of textures of real-life 3D models. Although GAN has made significant progress in image processing tasks, there are still problems such as unstable training, potential mode collapse, and poor performance on complex backgrounds and textures. These problems may lead to blurred restoration results and unclear textures, thus limiting its application in complex urban 3D scenes. Therefore, improving the ability of GAN in the task of high-fidelity restoration of textures of real-life 3D models is an urgent problem to be solved.
[0035] U-Net is a convolutional neural network based on an encoder-decoder architecture, and performs feature fusion through skip connections. The encoder extracts hierarchical features by gradually downsampling, while the decoder restores the image resolution layer by layer during upsampling. Skip connections effectively establish a connection between the encoder and decoder, combining local and global context information to achieve high-precision reconstruction and restore finer texture details. This method can reduce texture inconsistencies during the restoration process and improve the quality of the restored image.
[0036] Based on this feature of U-Net, the present invention combines it with GAN to enhance the feature extraction and context understanding capabilities of the generator network. The resulting architecture is called U-NetGAN. Figure 3 As shown in Figure 1, U-Net is used as a generator, while the discriminator works under the adversarial training framework of GAN. This architecture significantly improves the quality of texture restoration and ensures the usability of post-processed real-life 3D models.
[0037] Repairing the network Data preprocessing: In order to perform texture mapping, traverse the input real scene 3D data nodes and retrieve the nodes containing texture information , and traverse all triangles to get the texture coordinate set , is the number of triangles, Represents the horizontal and vertical coordinates of the first index vertex of the triangle face.
[0038] Create a blank image The image length and width are , as shown in formula (1), the mapping relationship between texture coordinates and blank image is established: ; On this basis, the pixel value mapping of the triangular surface texture is performed to transfer the pixel RGB value within the triangular surface to the image In the texture image after mapping .
[0039] like Figure 4 As shown, in order to reduce the desensitization calculation pressure, the real-scene three-dimensional data is further selectively traversed.
[0040] Assuming that the real-life 3D data has an n-layer tree structure, each layer of data indexes several sub-data of the next layer. First, the sensitivity of the texture image of the first n-1 layers of data is detected, and the last layer of data is selectively traversed. Only when the parent node of the data has sensitive content, the node data is traversed.
[0041] Texture Mapping and Selective Traversal Sensitive pattern detection: The input image Resized to a fixed size suitable for Mask R-CNN model input.
[0042] According to the pre-trained weight model, as shown in formula (2), the adjusted Extract feature maps through the backbone network ResNet50: ; In the formula, The convolution operation representing the backbone network contains image information at different levels, which is passed to the Feature Pyramid Network (FPN) to generate multi-scale feature maps.
[0043] Then, the feature map is fed into the region proposal network (RPN) to generate regions of interest (RoIs) and Then Align and apply bilinear interpolation to maintain spatial resolution to obtain the aligned feature vector .
[0044] As shown in formula (3), the bounding box adjustment parameters of each detected potential pattern are predicted: ; In the formula, and Represents the pre-trained weights and aligned feature vectors for bounding box regression; and is the predicted offset of the x- and y-coordinates of the center of the bounding box, and are the predicted scale factors for width and height.
[0045] According to the above adjustment parameters, the pattern bounding box is adjusted according to formula (4) to formula (7): ; In the formula, and are the adjusted center coordinates of the predicted bounding box, and is the center coordinate of the original bounding box, and is the adjusted bounding box width and height, and are the width and height of the original bounding box.
[0046] Calculate the probability that the bounding box belongs to different categories, as shown in formula (8): ; In the formula is the confidence that the bounding box belongs to category k, is the classification weight, is the total number of categories.
[0047] Compare the confidence levels of different categories and get the maximum confidence value With category .
[0048] The mask prediction branch enables the model to generate a segmentation mask for each detected object, allowing it to both detect and segment instances. The segmentation mask is obtained by pixel classification within each region. .
[0049] The detection result of each potential target in the input image is ,in represents the bounding box, Indicates the category to which the target belongs. and are the confidence and pixel mask of pattern recognition, respectively.
[0050] Pre-set sensitive pattern library ,when It is a preset sensitive pattern rule library A subset of , and the confidence When it is greater than the set threshold, the target is considered sensitive. Added to the sensitive content collection.
[0051] Draw each mask range that meets the conditions on the mask image. The pixel value of the sensitive area in the image is 255, which is the sensitive area to be blanked, and the pixel value of the black part is 0, which is the non-sensitive area.
[0052] In order to reduce the impact of shadows on pattern segmentation accuracy, the mask dilation operation is shown in formula (9): ; In the formula is the pixel value, and the original mask after expansion extends 10 pixels outward to be identified as the sensitive area.
[0053] Sensitive texture regeneration: Use U-NetGAN to regenerate the texture image to obtain non-sensitive desensitized real-scene 3D model data. The specific steps are as follows: The input image is resized to a fixed size suitable for the inpainting model U-NetGAN.
[0054] The size of the convolution kernel of the restoration model is set to 3x3, and the texture image and mask image are input into the U-Net network. After multi-layer convolution operations, maximum pooling is performed to obtain the feature map.
[0055] The size of the feature map is reduced, the depth is increased, and features such as edges and textures are gradually extracted. Then it enters the image decoder, restores the spatial resolution of the image by upsampling, and makes a jump connection with the corresponding layer of the encoder at each layer to utilize the encoder features at each scale.
[0056] Generate a texture image containing the restored texture image using a 1x1 convolution.
[0057] The desensitized texture image is texture mapped to obtain a desensitized real-scene three-dimensional model.
[0058] This paper verifies the effectiveness of the proposed algorithm through real data sets and compares it with the classic method and the desensitization algorithm YOLO+Patch Match. This experiment mainly evaluates two key aspects: (1) algorithm security: evaluating the algorithm's detection accuracy for sensitive targets; (2) algorithm usability: evaluating the quality of texture reconstruction.
[0059] Experimental Dataset Real street scene images are an important texture basis for 3D model reconstruction, which contain multi-scale and multi-category urban sensitive targets, so they are suitable for evaluating the effectiveness of target detection and restoration. In order to construct the experimental dataset, this paper uses the open source datasets TT-100K, CCTSDB and VisDrone. The specific information is shown in Table 1.
[0060] To address the challenges of texture inversion and complex backgrounds commonly found in urban 3D GIS data, the dataset was enhanced to improve the model’s ability to detect objects at multiple scales, resolutions, and angles. The preprocessing step included rotating the image clockwise by 60°, 120°, 180°, 240°, and 300° to enhance the model’s generalization capabilities. Figure 5 As shown in (a)-(f) of the dataset, the dataset contains street view images with added noise, as well as rotated targets, distant targets, partial texture loss, dark lighting, and complex background.
[0061] In order to verify the detection accuracy of the algorithm under extreme conditions, the final dataset contains 1,500 annotated street view images, which are divided into training set and test set in a ratio of 8:2. The data is annotated using the Labelme tool, and the final sample contains street view images in .jpg format and corresponding .json format annotation files.
[0062] Table 1
[0063] Evaluation indicators Object detection accuracy: Object detection accuracy is a key indicator for evaluating the ability of the desensitization model to identify, locate and process specific sensitive objects in a three-dimensional scene. Quantitative evaluation indicators include recall, precision and F1 value.
[0064] As shown in formula (10), Recall represents the proportion of correctly detected targets among all sensitive targets. A higher recall rate means that the algorithm has a stronger ability to correctly classify samples.
[0065] ; In the formula and They represent the number of correctly detected sensitive targets and the total number of sensitive targets respectively.
[0066] Precision refers to the proportion of correctly classified samples among all samples predicted to be sensitive. A higher precision means that the detection algorithm has a stronger ability to classify samples. The precision calculation formula is shown in formula (11): ; In the formula Indicates the number of non-sensitive targets that were incorrectly detected as sensitive.
[0067] F1 is the harmonic mean of precision and recall, and is used to evaluate the performance of the target detection algorithm as a whole. Its calculation formula is shown in formula (12). The F1 score ranges from 0 to 1. The closer the value is to 1, the higher the overall classification accuracy of the algorithm.
[0068] ; In the formula Indicates the number of sensitive targets that are incorrectly detected as non-sensitive.
[0069] Texture Restoration Quality: The quality of texture after desensitization will affect the availability of data and is an important indicator for evaluating the effectiveness of the algorithm. The evaluation includes qualitative and quantitative aspects. Qualitative evaluation uses subjective observation methods to focus on the clarity of the texture image after desensitization and its transition with the surrounding texture to ensure that the image has no obvious distortion, blurred edges or abrupt changes. Quantitative evaluation objectively evaluates the ability of the desensitization model to generate high-quality textures through measurable indicators. The present invention uses peak signal-to-noise ratio (PSNR), root mean square error (RMSE) and structural similarity index (SSIM) as quantitative evaluation indicators.
[0070] PSNR is used to evaluate the texture difference between the original image and the repaired image, and quantify the degree of noise in the image repair process in decibels (dB). The calculation formula of PSNR is as follows: ; in, and represent the original image and the repaired image respectively, Represents the pixel coordinates in the image, and the image size is The higher the PSNR value, the lower the image distortion and the higher the image quality. Generally speaking, a PSNR higher than 40 dB indicates good image quality, while a PSNR lower than 20 dB indicates poor image quality.
[0071] RMSE is an indicator that measures the average difference between the predicted value and the true value. Generally, the smaller the RMSE value, the smaller the difference between the predicted value and the true value, and the better the prediction performance of the model. The calculation formula of RMSE is as follows: ; in, represents the total number of pixels in the image, and Represent the pixel values of the original texture image and the repaired image respectively.
[0072] The SSIM comprehensive image considers brightness, contrast and structural information and is used to evaluate the structural similarity of two images. The SSIM value range is -1 to 1. The closer the value is to 1, the higher the similarity between the two images and the better the restoration quality. The calculation method of SSIM is as follows: ; in, and represent the original image and the repaired image respectively, and Respectively represent images The pixel mean and variance of and Respectively represent images The pixel mean and variance of represents the covariance between the two images, and To ensure the stability of the calculation, represents the dynamic range of pixel values. In the proposed algorithm, .
[0073] Experimental results analysis: The experimental dataset consists of 300 street view images containing sensitive objects and 200 street view images without sensitive objects. The experimental hardware environment includes Windows 11 operating system, 13th generation Intel® Core™ i7-13700 2.10 GHz processor, 64GB memory, and NVIDIA GeForce RTX 3090 GPU. The experimental process uses GPU acceleration. Training and testing are based on Python language compilation and use the PyTorch framework. The Python version is 3.7.
[0074] Sensitive target detection performance analysis: In order to verify the performance of the sensitive object detection algorithm, detection tests are performed on sensitive shapes and sensitive texts. Table 2 shows the datasets used for object detection. The test dataset is not repeated with the model training dataset. Based on experience and the high security requirements of data desensitization, the detection threshold is set to 0.75 in the experiment.
[0075] Table 2
[0076] Height limit, width limit, axle load limit, and weight limit signs are considered as sensitive targets, while other non-sensitive traffic signs are considered as interference items. The specific number is shown in Table 3. In the experiment, MobileNetv2 is used as the feature extraction network of Mask R-CNN to conduct pattern target detection experiments, and the detection results are compared with the desensitization algorithm YOLO+Patch Match algorithm.
[0077] Table 3
[0078] The experimental results of this algorithm and the comparison algorithm were statistically analyzed. The detection results of sensitive shapes in the test set are shown in Tables 4 and Figure 6 Table 4 counts the number of correctly detected sensitive targets, excluding false detections.
[0079] Table 4
[0080] Confusion matrix of pattern detection results According to Table 4 and Figure 6 , when MobileNetv2 is used as the backbone network of Mask R-CNN, the algorithm has poor detection effect on the data set, and there are many false detections and missed detections of the targets. For example, for the height limit signs, there are 11 missed detections and 14 false detections, and the missed detection rate and false detection rate are 6.32% and 8.05% respectively, indicating that its detection results are unstable and far from meeting the requirements of data security sharing. The desensitization algorithm YOLO+Patch Match uses YOLOv5 for detection. Although the detection accuracy has been improved, there are still many false detections. In the algorithm proposed in the present invention, ResNet50 is used as the backbone network of Mask R-CNN. Its deep network structure and residual connection make its feature extraction ability stronger than the comparison algorithm. The number of errors in the detection results is only 2, which is lower than 0.5% of the total number of targets. Therefore, it can be considered that the detection results are reliable.
[0081] Based on the above statistical results, the recall, precision and F1 scores of each category of targets in the test data set are calculated, and the arithmetic mean of the four categories of targets is taken as the indicator evaluation result of the algorithm, as shown in Table 5.
[0082] Table 5
[0083] As shown in Table 5, when MobileNetv2 is used, the recall rate, precision rate and F1 score are 0.9331, 0.9151 and 0.9222 respectively, which are significantly lower than the YOLO+Patch Match algorithm and the algorithm of the present invention, indicating that the detection accuracy of pattern targets of MobileNetv2 is low. The detection result of the YOLO+Patch Match algorithm is close to 0.95, but still does not reach the higher standard. The algorithm proposed in the present invention has all indicators close to 1, indicating its superiority in the task of sensitive target detection of real-life 3D model data.
[0084] Image restoration performance analysis: The quality of texture restoration in sensitive areas determines the usability of data after desensitization. To verify the performance of the U-NetGAN texture restoration algorithm, the restoration results of the algorithm proposed in this paper are compared with those of the YOLO+Patch Match algorithm and the standard GAN method. The experiment is conducted on a test dataset containing 300 street view images containing sensitive targets.
[0085] (1) Qualitative evaluation Three test images were selected, in which sensitive targets were located in different scenes - close up, far away, and against complex backgrounds, to observe the texture details after desensitization. The quality of texture restoration was judged by qualitatively evaluating the degree of integration of the restored area with the surrounding texture, focusing on whether there were artifacts such as distortion, blurred edges, or unnatural transitions. Figure 7 The repair results of these examples are shown.
[0086] After observing the texture details, the images repaired using Patch Match in the prior art show varying degrees of edge texture distortion and abrupt transitions. Although the GAN algorithm is superior to Patch Match in restoration quality and has less texture distortion, the resolution drops significantly in long-distance images. In contrast, U-NetGAN shows significant advantages in local texture restoration of close-up, distant and complex background images. This method can effectively retain detailed textures and avoid obvious distortion, blur or abrupt transitions. In summary, the U-NetGAN scheme proposed in the present invention performs well in visual effects and ensures high usability, making it very suitable for the desensitization task of real-life 3D model data.
[0087] (2) Quantitative evaluation The RMSE, PSNR and SSIM values of the original image and the repaired image are calculated and the arithmetic mean is taken. The results are shown in Table 6.
[0088] Table 6
[0089] As shown in Table 6, the RMSE of the algorithm proposed in the present invention is 1.00 and 1.03 lower than that of the Patch Match and GAN algorithms, respectively. This shows that the algorithm has a smaller difference between the original texture and the repaired texture, and therefore has a lower noise repair effect. The PSNR of the Patch Match algorithm is the lowest, only 38.43dB, reflecting that its texture regeneration quality is the worst. The GAN algorithm has improved on this basis, with a PSNR of more than 40dB, indicating that the texture quality it generates is good. In addition, the PSNR of the algorithm of the present invention is further improved by 0.89dB, indicating that it has the lowest noise and the highest quality of repaired texture. The SSIM index of the Patch Match algorithm is the lowest, only 0.98, while the GAN algorithm and the algorithm proposed in the present invention both reach 1, indicating that the repaired image is almost completely consistent with the original texture. In summary, the Patch Match algorithm is difficult to generate high-quality, detailed images, and the GAN algorithm significantly improves the repair quality. The U-NetGAN algorithm proposed in the present invention performs best on the test data set, not only effectively retaining texture and structural features, but also showing excellent ability in texture repair.
[0090] The present invention discloses a method, system and medium for desensitizing patterns and textures of real-life 3D models. The method combines Mask R-CNN to realize automatic detection and segmentation of sensitive content, and uses a generative adversarial network (GAN) based on U-Net to perform realistic texture restoration, thereby realizing automatic detection and high-quality texture restoration. The present invention can effectively perform automatic and accurate data desensitization, while achieving high-quality texture restoration. This ensures the security and availability of desensitized data, promotes the secure sharing of real-life 3D model data, and provides important potential for digital twin and smart city applications.
[0091] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for desensitizing a pattern texture of a real-life three-dimensional model, characterized in that: include: Mapping the texture information of the real-scene three-dimensional model to the two-dimensional image to generate a corresponding texture image; Detecting sensitive content in the texture image through a deep learning-based instance segmentation network, and generating a pixel-level segmentation mask to mark sensitive areas; Using a generative adversarial network to perform texture restoration on the sensitive area, and generate a desensitized texture image that is seamlessly integrated with the surrounding non-sensitive areas; The repaired texture image is remapped to the real-scene three-dimensional model to obtain desensitized three-dimensional model data; wherein the instance segmentation network adopts a two-stage detection framework, combining the region proposal network and the feature pyramid network to achieve multi-scale sensitive target recognition; the generative adversarial network includes a generator of an encoder-decoder structure, which fuses multi-level features through jump connections, and a discriminator constructed by a convolutional neural network, which is used for adversarial training to optimize the output quality of the generator.
2. The method for desensitizing a real-life 3D model pattern texture according to claim 1, characterized in that: The process of mapping the texture information of the real-scene three-dimensional model to the two-dimensional image includes: Traversing the tree structure nodes of the three-dimensional model to extract triangular facets containing texture information; According to the texture coordinates of the triangular patch, the pixel values of the surface of the three-dimensional model are mapped to a blank image; The tree structure is traversed hierarchically, and only branches whose parent nodes contain sensitive content are deeply traversed.
3. The method for desensitizing a real-life 3D model pattern texture according to claim 1, characterized in that: The instance segmentation network is Mask R-CNN, and the backbone network is ResNet50; The detection process includes: Generate candidate regions through the region proposal network and perform bilinear interpolation alignment on the candidate regions; Predict bounding box parameters, category confidence, and pixel-level segmentation mask of candidate regions; Sensitive targets are screened according to the preset sensitive pattern library and confidence threshold to generate a mask image.
4. The method for desensitizing a real-life 3D model pattern texture according to claim 1, characterized in that: The generator of the generative adversarial network is a U-Net structure, including: The encoder gradually extracts features through multi-layer convolution and maximum pooling; The decoder restores the image resolution by upsampling and performs feature fusion with the corresponding layer of the encoder through skip connections; The authenticity of the repaired image output by the generator is determined by the discriminator to drive adversarial training.
5. The method for desensitizing a real-life 3D model pattern texture according to claim 1, characterized in that: The process of generating the mask image further includes: A dilation operation is performed on the pixel-level segmentation mask so that the boundary of the sensitive area extends outward by a preset pixel width.
6. The method for desensitizing a real-life 3D model pattern texture according to claim 1, characterized in that: The adversarial training process of the generative adversarial network includes: Inputting the mask image and the original texture image into the generator to generate a repaired image; The discriminator receives the repaired image and the real non-sensitive image, and outputs a probability of authenticity; The parameters of the generator are adjusted according to the feedback of the discriminator until the difference between the repaired image and the real image is lower than a preset threshold.
7. A system, characterized in that: The system includes: a memory, a processor, and a real-scene three-dimensional model pattern texture desensitization program stored in the memory and executable on the processor. When the real-scene three-dimensional model pattern texture desensitization program is executed by the processor, the steps of the real-scene three-dimensional model pattern texture desensitization method as described in any one of claims 1 to 6 are implemented.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a real-scene 3D model pattern texture desensitization program, and when the real-scene 3D model pattern texture desensitization program is executed by the processor, the steps of the real-scene 3D model pattern texture desensitization method as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Automatic generation method of color LOD model based on block decomposition
CN109118588A
Geographic grid intelligent local desensitization method based on generative adversarial network
CN113066094A
Method and device for automatically replacing tree model in live-action three-dimensional data
CN115861549A
Interactive three-dimensional model texture editing method and system and electronic equipment
CN118379470A
Deep learning-based crowdsourcing acquisition high-precision map desensitization method and system
CN119293835A