Image processing method and device

Through the privacy protection network, individual information in images or video data is identified and positioned, and privacy protection is used using preset masks, which solves the problem of inefficient processing in the prior art, and achieves efficient privacy protection and image readability.

CN115690639BActive Publication Date: 2025-08-15HISENSE GRP HLDG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110878160.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-30
Publication Date
2025-08-15
Estimated Expiration
2041-07-30

AI Technical Summary

Technical Problem

When the prior art protects privacy of images or video data with large data volumes, the processing efficiency is low and cannot effectively improve the efficiency of privacy protection without affecting the readability of images.

Method used

The individual positioning subnet in the privacy protection network is used to identify and locate individual information in the image, generate boundary positioning information, and use the individual protection subnet to perform privacy protection processing based on the preset mask to block the designated area.

Benefits of technology

Improve the privacy protection and processing efficiency of image or video data, ensuring data privacy and security without affecting the readability of images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690639B_ABST
    Figure CN115690639B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of computer technology, and in particular to a method and apparatus for image processing. The method is used to address the problem of low processing efficiency for privacy protection of large-volume image or video data under existing technologies. The method comprises the following steps: obtaining an image to be processed containing at least one individual information requiring privacy protection, inputting the image to be processed into a privacy protection network, employing an individual positioning sub-network within the privacy protection network to identify and locate each individual information contained in the image to be processed, and obtaining boundary positioning information corresponding to each individual information; then, employing an individual protection sub-network within the privacy protection network to perform privacy protection processing on the positioning areas associated with each obtained boundary positioning information, and obtaining a target image corresponding to the image to be processed; thereby, the processing efficiency of privacy protection is improved, and the security of data privacy is ensured without affecting the readability of the image to be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to an image processing method and device. Background Art

[0002] With the rapid development of computer technology, users can access information more conveniently and efficiently through products such as smart terminals and cloud processors. In the era of big data, an increasing number of services and products are built around user data. While this provides personalized services and improves service quality and accuracy, it also inevitably exposes user privacy to risks during data collection, use, and disclosure.

[0003] Under existing technologies, the following two methods are usually used to achieve privacy protection of image or video data:

[0004] The first method adopts the most direct protection method, for example, protecting the privacy of image or video data by blocking sensitive information, or preventing third parties from accessing image or video data through encryption and other methods.

[0005] However, with the development of multimedia technology, the amount of image or video data is getting larger and larger. When using the above method to encrypt or mask image or video data with a large amount of data, it will bring higher operation complexity, thereby reducing processing efficiency.

[0006] The second method is to extract and identify sensitive information in image or video data based on the traditional image recognition algorithm network, and then hide and replace the identified sensitive information to achieve the corresponding privacy protection effect.

[0007] However, since traditional image recognition algorithms have the disadvantages of large number of parameters and are not easy to be transplanted and widely used, the above method of hiding and replacing sensitive information in large amounts of image or video data also has the problem of low processing efficiency.

[0008] Therefore, a new method needs to be designed to solve the above problems. Summary of the Invention

[0009] The purpose of the present disclosure is to provide an image processing method and apparatus for solving the problem of low processing efficiency when performing privacy protection on large amounts of image or video data under existing technologies.

[0010] The specific technical solutions provided by the embodiments of the present disclosure are as follows:

[0011] In a first aspect, an image processing method includes:

[0012] Acquiring an image to be processed, wherein the image to be processed contains at least one individual information requiring privacy protection;

[0013] Inputting the image to be processed into a privacy protection network, and using the individual positioning sub-network in the privacy protection network to respectively identify and locate each individual information contained in the image to be processed, and obtaining boundary positioning information corresponding to each individual information;

[0014] The individual protection subnetwork in the privacy protection network is used to perform privacy protection processing on the positioning areas associated with each piece of obtained boundary positioning information to obtain a target image corresponding to the image to be processed; wherein the individual protection subnetwork performs privacy protection processing on the image to be processed based on a preset mask having the same appearance features as the image to be processed; the preset mask is generated by the individual positioning subnetwork and is used to shield a specified area in the specified individual information.

[0015] Optionally, obtaining the image to be processed includes:

[0016] Select any candidate image from an image set and use the any candidate image as the image to be processed, wherein the image set includes candidate images that require privacy protection; or

[0017] Video data is acquired, and frames of images are extracted from the video data, and the frames of images are used as the images to be processed, wherein the frames of images contain at least one individual information that needs to be privacy protected.

[0018] Optionally, the individual positioning sub-network comprises at least a feature extraction module, a feature pyramid fusion module, a prototype mask generation module and a boundary positioning information output module; wherein,

[0019] The feature extraction module is used to extract each first feature information contained in the image to be processed and generate a first feature map corresponding to the image to be processed; each first feature information represents a one-dimensional feature information contained in the image to be processed;

[0020] The feature pyramid fusion module is configured to obtain second feature information corresponding to each piece of individual information contained in the image to be processed, and fuse the first feature map with each piece of second feature information to generate a second feature map corresponding to the image to be processed, wherein each piece of second feature information contains two-dimensional feature information corresponding to individual information associated with one piece of second feature information;

[0021] The prototype mask generation module is configured to generate the preset mask, wherein the preset mask is obtained by fusing the first feature map and each piece of third feature information obtained, each piece of third feature information is obtained by the feature pyramid fusion module and is one piece of first feature information among each piece of first feature information except each piece of second feature information obtained;

[0022] The boundary positioning information output module is used to output the boundary positioning information corresponding to each piece of individual information.

[0023] Optionally, inputting the image to be processed into a privacy-preserving network, and using an individual positioning subnetwork in the privacy-preserving network to respectively identify and locate individual information contained in the image to be processed, and obtaining boundary positioning information corresponding to each individual information, includes:

[0024] Inputting the image to be processed into the feature extraction module in the privacy protection network, extracting each first feature information contained in the image to be processed, and generating a first feature map corresponding to the image to be processed based on the each first feature information;

[0025] Inputting the first feature map into the feature pyramid fusion module, obtaining second feature information corresponding to each individual information contained in the image to be processed, and fusing the first feature map with each obtained second feature information to obtain a second feature map corresponding to the image to be processed;

[0026] Inputting the second feature map into the prototype mask generation module to generate a preset mask corresponding to the image to be processed;

[0027] The preset mask is input into the boundary positioning information output module. In the image to be processed, based on the preset mask, the boundary positioning information corresponding to each individual information contained in the preset mask is drawn, and the drawn image is used as a positioning image, and the positioning image is output.

[0028] Optionally, the step of using an individual protection subnetwork in the privacy protection network to perform privacy protection processing on positioning areas associated with each obtained boundary positioning information to obtain a target image corresponding to the image to be processed includes:

[0029] Obtaining the preset mask, and using the individual protection sub-network to superimpose the preset mask and the positioning image to obtain a superimposed image;

[0030] Obtaining the correspondence between preset category attribute information and foreground color;

[0031] Filling the positioning areas corresponding to the individual pieces of information with corresponding foreground colors based on the category attribute information contained in the obtained boundary positioning information and the preset correspondence between the category attribute information and the foreground color;

[0032] The category attribute information corresponding to each of the individual information is labeled, and the target image is obtained and output.

[0033] Optionally, the individual positioning sub-network is trained in the following manner:

[0034] The individual positioning sub-network to be trained is trained for multiple rounds of iterative training based on the training sample set, and when the preset convergence conditions are met, the trained individual positioning sub-network is output. In one cycle of iterative training, the following operations are performed:

[0035] Inputting each training sample obtained from the training sample set into the feature extraction module of the individual positioning subnetwork to be trained, extracting each first feature information corresponding to each training sample, and generating a first feature map corresponding to each training sample; wherein, each training sample includes the actual boundary positioning information corresponding to the designated area marked with the designated individual information that needs to be privacy protected;

[0036] Inputting each generated first feature map into the feature pyramid fusion model in the individual positioning subnetwork to be trained, obtaining second feature information corresponding to each individual information contained in each training sample, and fusing the first feature map corresponding to each training sample with the corresponding second feature information obtained to obtain each second feature map corresponding to each training sample;

[0037] Inputting each second feature map obtained into the prototype mask generation module in the individual positioning sub-network to be trained, respectively, to generate a preset mask corresponding to each training sample;

[0038] The obtained preset masks are respectively input into the boundary positioning information output module in the individual positioning sub-network to be trained, and based on the respective preset masks, the predicted boundary positioning information corresponding to each individual information contained in the corresponding training samples is drawn; and based on the comparison results of each predicted boundary positioning information obtained and the corresponding true boundary positioning information, the corresponding total loss value is determined, and based on the obtained total loss values, the model parameters of the individual positioning sub-network to be trained are adjusted.

[0039] Optionally, the total loss value includes a first loss value, a second loss value and a third loss value;

[0040] The determining of the respective corresponding total loss values based on the comparison results of the obtained respective predicted boundary positioning information and the respective corresponding real boundary positioning information includes:

[0041] Comparing the predicted contour information corresponding to each of the obtained training samples with the corresponding true contour information to obtain a corresponding first loss value;

[0042] Comparing the predicted category attribute information corresponding to each of the obtained training samples with the corresponding true category attribute information to obtain a corresponding second loss value;

[0043] Comparing the predicted labeling information corresponding to each of the obtained training samples with the corresponding true labeling information to obtain a corresponding third loss value;

[0044] According to the preset relationship, the obtained first loss values, the obtained second loss values, and the obtained third loss values are respectively integrated to obtain the corresponding total loss value.

[0045] In a second aspect, an image processing apparatus includes:

[0046] An acquisition module, configured to acquire an image to be processed, wherein the image to be processed contains at least one individual information requiring privacy protection;

[0047] an acquisition module, configured to input the image to be processed into a privacy-preserving network, and employ an individual positioning subnetwork in the privacy-preserving network to identify and locate each individual information contained in the image to be processed, thereby obtaining boundary positioning information corresponding to each individual information;

[0048] An output module is configured to perform privacy protection processing on the positioning areas associated with each piece of obtained boundary positioning information using an individual protection subnetwork in the privacy protection network, thereby obtaining a target image corresponding to the image to be processed; wherein the individual protection subnetwork performs privacy protection processing on the image to be processed based on a preset mask having the same appearance features as the image to be processed; the preset mask is generated by the individual positioning subnetwork and is used to shield a specified area in the specified individual information.

[0049] Optionally, the acquiring module for acquiring the image to be processed is configured to:

[0050] Select any candidate image from an image set and use the any candidate image as the image to be processed, wherein the image set includes candidate images that require privacy protection; or

[0051] Video data is acquired, and frames of images are extracted from the video data, and the frames of images are used as the images to be processed, wherein the frames of images contain at least one individual information that needs to be privacy protected.

[0052] Optionally, the individual positioning sub-network comprises at least a feature extraction module, a feature pyramid fusion module, a prototype mask generation module and a boundary positioning information output module; wherein,

[0053] The feature extraction module is used to extract each first feature information contained in the image to be processed and generate a first feature map corresponding to the image to be processed; each first feature information represents a one-dimensional feature information contained in the image to be processed;

[0054] The feature pyramid fusion module is configured to obtain second feature information corresponding to each piece of individual information contained in the image to be processed, and fuse the first feature map with each piece of second feature information to generate a second feature map corresponding to the image to be processed, wherein each piece of second feature information contains two-dimensional feature information corresponding to individual information associated with one piece of second feature information;

[0055] The prototype mask generation module is configured to generate the preset mask, wherein the preset mask is obtained by fusing the first feature map and each piece of third feature information obtained, each piece of third feature information is obtained by the feature pyramid fusion module and is one piece of first feature information among each piece of first feature information except each piece of second feature information obtained;

[0056] The boundary positioning information output module is used to output the boundary positioning information corresponding to each piece of individual information.

[0057] Optionally, the image to be processed is input into a privacy protection network, and an individual positioning subnetwork in the privacy protection network is used to identify and locate each individual information contained in the image to be processed, thereby obtaining boundary positioning information corresponding to each individual information. The obtaining module is configured to:

[0058] Inputting the image to be processed into the feature extraction module in the privacy protection network, extracting each first feature information contained in the image to be processed, and generating a first feature map corresponding to the image to be processed based on the each first feature information;

[0059] Inputting the first feature map into the feature pyramid fusion module, obtaining second feature information corresponding to each individual information contained in the image to be processed, and fusing the first feature map with each obtained second feature information to obtain a second feature map corresponding to the image to be processed;

[0060] Inputting the second feature map into the prototype mask generation module to generate a preset mask corresponding to the image to be processed;

[0061] The preset mask is input into the boundary positioning information output module. In the image to be processed, based on the preset mask, the boundary positioning information corresponding to each individual information contained in the preset mask is drawn, and the drawn image is used as a positioning image, and the positioning image is output.

[0062] Optionally, the individual protection subnetwork in the privacy protection network is used to perform privacy protection processing on the positioning areas associated with each obtained boundary positioning information to obtain a target image corresponding to the image to be processed, and the output module is used to:

[0063] Obtaining the preset mask, and using the individual protection sub-network to superimpose the preset mask and the positioning image to obtain a superimposed image;

[0064] Obtaining the correspondence between preset category attribute information and foreground color;

[0065] Filling the positioning areas corresponding to the individual pieces of information with corresponding foreground colors based on the category attribute information contained in the obtained boundary positioning information and the preset correspondence between the category attribute information and the foreground color;

[0066] The category attribute information corresponding to each of the individual information is labeled, and the target image is obtained and output.

[0067] Optionally, the individual positioning sub-network is trained in the following manner:

[0068] The individual positioning sub-network to be trained is trained for multiple rounds of iterative training based on the training sample set, and when the preset convergence conditions are met, the trained individual positioning sub-network is output. In one cycle of iterative training, the following operations are performed:

[0069] Inputting each training sample obtained from the training sample set into the feature extraction module of the individual positioning subnetwork to be trained, extracting each first feature information corresponding to each training sample, and generating a first feature map corresponding to each training sample; wherein, each training sample includes the actual boundary positioning information corresponding to the designated area marked with the designated individual information that needs to be privacy protected;

[0070] Inputting each generated first feature map into the feature pyramid fusion model in the individual positioning subnetwork to be trained, obtaining second feature information corresponding to each individual information contained in each training sample, and fusing the first feature map corresponding to each training sample with the corresponding second feature information obtained to obtain each second feature map corresponding to each training sample;

[0071] Inputting each second feature map obtained into the prototype mask generation module in the individual positioning sub-network to be trained, respectively, to generate a preset mask corresponding to each training sample;

[0072] The obtained preset masks are respectively input into the boundary positioning information output module in the individual positioning sub-network to be trained, and based on the respective preset masks, the predicted boundary positioning information corresponding to each individual information contained in the corresponding training samples is drawn; and based on the comparison results of each predicted boundary positioning information obtained and the corresponding true boundary positioning information, the corresponding total loss value is determined, and based on the obtained total loss values, the model parameters of the individual positioning sub-network to be trained are adjusted.

[0073] Optionally, the total loss value includes a first loss value, a second loss value and a third loss value;

[0074] The output module is used to determine the total loss value corresponding to each of the predicted boundary positioning information and the corresponding real boundary positioning information based on the comparison result obtained.

[0075] Comparing the predicted contour information corresponding to each of the obtained training samples with the corresponding true contour information to obtain a corresponding first loss value;

[0076] Comparing the predicted category attribute information corresponding to each of the obtained training samples with the corresponding true category attribute information to obtain a corresponding second loss value;

[0077] Comparing the predicted labeling information corresponding to each of the obtained training samples with the corresponding true labeling information to obtain a corresponding third loss value;

[0078] According to the preset relationship, the obtained first loss values, the obtained second loss values, and the obtained third loss values are respectively integrated to obtain the corresponding total loss value.

[0079] According to a third aspect, a computer device includes:

[0080] a memory for storing a computer program executable by the controller;

[0081] The controller is connected to the memory and is configured to execute any one of the methods of the first aspect above.

[0082] In a fourth aspect, a computer-readable storage medium is provided. When a computer program in the computer-readable storage medium is executed by a processor, the processor is enabled to perform any one of the methods described in the first aspect.

[0083] In an embodiment of the present disclosure, an image to be processed is obtained, wherein the image to be processed contains at least one individual information requiring privacy protection. The image to be processed is then input into a privacy protection network, and an individual positioning sub-network in the privacy protection network is used to identify and locate each individual information contained in the image to be processed, thereby obtaining boundary positioning information corresponding to each individual information. Finally, the individual protection sub-network in the privacy protection network is used to perform privacy protection processing on the positioning areas associated with each obtained boundary positioning information, thereby obtaining a target image corresponding to the image to be processed. The individual protection sub-network performs privacy protection processing on the image to be processed based on a preset mask having the same external features as the image to be processed. The preset mask is generated by the individual positioning sub-network and is used to shield a specified area in the specified individual information. In this way, since the privacy protection network is trained based on each training sample and their corresponding real boundary positioning information, when the image to be processed is input into the trained privacy protection network, the specified area in the specified individual information that needs to be privacy protected can be accurately located, and a preset mask for shielding the specified area in the specified individual information can be obtained. Then, based on the obtained preset mask, the specified area in the specified individual information in the image to be processed can be better privacy-protected, thereby improving the processing efficiency of privacy protection. In addition, on the basis of achieving privacy protection for the specified area in the specified individual information, other individual information that does not need to be privacy protected can also be clearly displayed, thereby ensuring the security of data privacy without affecting the readability of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 Schematic diagram of the architecture of a privacy protection network in an embodiment of the present disclosure;

[0085] Figure 2 Schematic diagram of the network structure of the feature pyramid submodule in the embodiment of the present disclosure;

[0086] Figure 3 Schematic diagram of the network structure of a submodule in the boundary positioning information output module in an embodiment of the present disclosure;

[0087] Figure 4 Schematic diagram of the training process of the individual positioning sub-network in an embodiment of the present disclosure;

[0088] Figure 5 This is a schematic diagram of an image processing process in an embodiment of the present disclosure;

[0089] Figure 6 This is a schematic diagram of a privacy protection process in an embodiment of the present disclosure;

[0090] Figure 7 A schematic diagram of an application scenario in an embodiment of the present disclosure;

[0091] Figure 8 Schematic diagram of the logical architecture of the image processing device in an embodiment of the present disclosure;

[0092] Figure 9 Schematic diagram of the physical architecture of a computer device in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0093] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure and not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0094] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0095] To address the problem of low processing efficiency when performing privacy protection on large-volume image or video data under existing technologies, in an embodiment of the present disclosure, an image to be processed is obtained, wherein the image to be processed contains at least one individual information requiring privacy protection. The image to be processed is then input into a privacy protection network, and an individual positioning sub-network in the privacy protection network is used to identify and locate each individual information contained in the image to be processed, thereby obtaining boundary positioning information corresponding to each individual information. Finally, the individual protection sub-network in the privacy protection network is used to perform privacy protection processing on the positioning areas associated with each obtained boundary positioning information, thereby obtaining a target image corresponding to the image to be processed. The individual protection sub-network performs privacy protection processing on the image to be processed based on a preset mask having the same external features as the image to be processed. The preset mask is generated by the individual positioning sub-network and is used to shield a specified area in the specified individual information, thereby improving the processing efficiency of privacy protection for large-volume image or video data.

[0096] In the embodiments of the present disclosure, the execution subject can be a terminal device or a server, which can be determined according to actual needs and is not specifically limited here. Among them, the terminal device can be a smart mobile terminal, a tablet computer, a laptop computer, a smart handheld device, a personal computer (PC), a computer, a smart screen, various wearable devices, a personal digital assistant (PDA), etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms. It is not limited in the present disclosure.

[0097] The preferred embodiments of the present disclosure are further described in detail below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present disclosure. In addition, the embodiments of the present disclosure and the features in the embodiments can be combined with each other if there is no conflict.

[0098] In an embodiment of the present disclosure, a method for image processing is provided. The method can be divided into a training part of an individual positioning subnetwork in a privacy protection network, and an application part of the individual positioning subnetwork and the individual protection subnetwork.

[0099] Optionally, in an embodiment of the present disclosure, an individual positioning sub-network to be trained can be selectively built based on the improved real-time instance segmentation open source library (YouOnly Look At CoefficienTs, YOLACT) network.

[0100] In the embodiment of the present disclosure, Figure 1 Figure 2 shows a schematic diagram of the network architecture of a privacy-preserving network. Figure 1 As shown, the privacy protection network 100 includes an individual positioning subnetwork 110 and an individual protection subnetwork 120. The individual positioning subnetwork 110 is used to extract and identify features of the input original image, output boundary positioning information corresponding to each individual information that needs to be privacy protected according to user needs, and generate a preset mask for the individual protection subnetwork 120 to perform privacy protection processing; the individual protection subnetwork 120 is used to implement privacy protection for a specified area of specified individual information in the original image based on the preset mask, and output the final target image after privacy protection.

[0101] In the disclosed embodiment, the individual positioning sub-network 110 comprises at least a feature extraction module 1101 , a feature pyramid fusion module 1102 , a prototype mask generation module 1103 , and a boundary positioning information output module 1104 ;

[0102] The feature extraction module 1101 is used to extract each first feature information corresponding to the input original image and generate a corresponding first feature map; each first feature information represents a one-dimensional feature information contained in the original image;

[0103] The feature pyramid fusion module 1102 is configured to obtain second feature information corresponding to each individual information contained in the original image, and fuse the first feature map with each second feature information to generate a second feature map corresponding to the original image, wherein each second feature information includes two-dimensional feature information corresponding to the individual information associated with the second feature information;

[0104] The prototype mask generation module 1103 is configured to generate a preset mask, wherein the preset mask is obtained by fusing the first feature map and each piece of third feature information obtained. Each piece of third feature information is obtained by the feature pyramid fusion module 1102 and is one piece of first feature information among each piece of first feature information excluding each piece of second feature information obtained.

[0105] The boundary positioning information output module 1104 is configured to output boundary positioning information corresponding to each piece of individual information.

[0106] Therefore, in the embodiment of the present disclosure, the individual positioning sub-network should be trained first to obtain a trained individual positioning sub-network; then, the privacy-protected target image is output based on the trained individual positioning sub-network.

[0107] (1) Training of individual positioning sub-network:

[0108] In the embodiment of the present disclosure, first, a sensitive data set needs to be established in advance, where the sensitive data set contains a large amount of training data, which are recorded as training samples. Each training sample is marked with a designated area of designated individual information that needs to be privacy protected, as well as corresponding category attribute information.

[0109] Optionally, in an embodiment of the present disclosure, each training sample in the sensitive data set is divided into several training sample sets according to the different designated areas of the designated individual information that need to be privacy protected as mentioned above, wherein each training sample set contains each training sample for training a designated area of a designated individual information that needs to be privacy protected.

[0110] Optionally, in an embodiment of the present disclosure, each training sample contained in each training set is divided into two parts, respectively recorded as positive samples and negative samples, each positive sample represents a training sample labeled with true category attribute information, and each negative sample represents a training sample labeled with false category attribute information; optionally, the ratio of positive samples to negative samples in each training sample set is 3:1.

[0111] In the embodiment of the present disclosure, the specific training process of the individual positioning sub-network is as follows:

[0112] The individual positioning sub-network to be trained is trained for multiple rounds of iterative training based on the training sample set, and when the preset convergence conditions are met, the trained individual positioning sub-network is output. In one cycle of iterative training, the following operations are performed:

[0113] In the embodiment of the present disclosure, in order to facilitate the description of the training process, the following description is given of the training of an individual positioning sub-network using only one training sample from each training sample in the training sample set.

[0114] 1. Feature extraction module:

[0115] In an embodiment of the present disclosure, one training sample is selected from each training sample obtained from a training sample set, and the training sample is input into a feature extraction module to extract the corresponding first feature information, and generate a first feature map corresponding to the training sample, wherein each first feature information represents a one-dimensional feature information contained in a training sample associated with a feature information.

[0116] Optionally, in an embodiment of the present disclosure, the feature extraction module may use a Bottleneck Attention Module (BAM) to extract first feature information corresponding to a training sample.

[0117] Optionally, in the disclosed embodiment, the BAM module is divided into a channel attention mechanism submodule and a spatial attention mechanism submodule. The channel attention mechanism submodule is constructed using global average pooling and two fully connected layers to increase cross-channel feature fusion; the spatial attention mechanism submodule uses a convolutional network for feature extraction to improve the feature correlation of spatial pixels. Specifically, the channel attention mechanism submodule and the spatial attention mechanism submodule adopt a parallel structure.

[0118] 2. Feature pyramid fusion module:

[0119] In the disclosed embodiment, a first feature map corresponding to a training sample output by the feature extraction module is input into a feature pyramid fusion module. The feature pyramid fusion module performs dimensionality reduction processing on the first feature map, thereby mining more important features and realizing the adjustment and fusion of feature channels of different scales.

[0120] Specifically, the feature pyramid fusion template further processes the input first feature map, obtains the second feature information corresponding to each individual information contained in a training sample, and fuses the first feature map and each second feature information to generate a second feature map corresponding to the training sample, wherein each second feature information contains two-dimensional feature information of the individual information contained in a training sample associated with one second feature information.

[0121] like Figure 2 As shown, Figure 2 Figure 2 shows the network structure diagram of the feature pyramid submodule in the feature pyramid fusion module. Figure 2 As shown, in the embodiment of the present disclosure, the input of the 1*1 convolution operation is the first feature map obtained by the feature extraction module; after the 1*1 convolution operation, each third feature map with the same number of channels is obtained. Then, after passing through three feature filters of different scales and multiple upsampling, downsampling and convolution operations, each first feature information in the first feature map is classified to obtain the second feature information and third feature information corresponding to a training sample, wherein each third feature information is one first feature information in each first feature information, excluding each second feature information obtained. Finally, the feature fusion submodule in the feature pyramid fusion module fuses the obtained second feature information with the first feature map to generate a second feature map corresponding to a training sample.

[0122] Optionally, in the embodiment of the present disclosure, in the feature pyramid sub-module, five feature filters with scales of 69*69, 35*35, 18*18, 9*9 and 5*5 can be used to divide the first feature information corresponding to a training sample obtained into candidate feature maps of different scales.

[0123] 3. Prototype mask generation module:

[0124] In the disclosed embodiment, the prototype mask generation module is used to generate a preset mask to be fused. During the model training phase, a training sample is selected from each training sample in the training sample set. The second feature information and third feature information of the training sample are obtained through the feature pyramid fusion module. The third feature information is then fused with the first feature map in the prototype mask generation module to obtain a preset mask for the training sample.

[0125] The specific training process is as follows:

[0126] First, the prototype mask generation module is trained using 32 basic prototypes as fused prototype masks. Prototypes with large contributions are automatically selected from the individual information of different instances, while prototypes with small contributions are reduced.

[0127] Secondly, the training coefficients are linearly combined with the prototype mask to effectively increase the inference speed of the individual localization sub-network and thus reduce the time complexity.

[0128] Finally, in the embodiment of the present disclosure, a weight feature selection submodule is added to the preset mask generation process, so that the generated intermediate mask can assist in adjusting the correlation between the various intermediate masks through global average pooling and two fully connected layers.

[0129] In the embodiment of the present disclosure, during the specific training process, the first feature map corresponding to a training sample output by the feature extraction module is fused with each third feature information corresponding to a training sample to obtain a third feature map, and then the third feature map is subjected to three convolution operations, as well as global average pooling and a fully connected layer, to obtain the final preset mask.

[0130] 4. Boundary positioning information output module:

[0131] The boundary positioning information output module is used to output the boundary positioning information of each individual information corresponding to each training sample and the corresponding category attribute information.

[0132] In the specific implementation, in the feature pyramid fusion module, for one training sample in each training sample set, candidate feature maps with scales of 69*69, 35*35, 18*18, 9*9 and 5*5 are obtained respectively. Correspondingly, in the boundary positioning information output module, a sub-module is matched for each of the above candidate feature maps.

[0133] Specifically, first, the boundary positioning information output module performs convolution processing on each input candidate feature map, adjusts the size of the feature channel, and performs regression adjustment on each pixel point and predicted boundary positioning information contained in a training sample. The individual information of a training sample that needs to be privacy protected, as well as other individual information other than the above-mentioned individual information that needs to be privacy protected, are constrained by their respective corresponding classification tasks, thereby obtaining the boundary positioning information corresponding to the individual information of a training sample that needs to be privacy protected.

[0134] Optionally, for ease of description, in the embodiment of the present disclosure, individual information in a training sample that requires privacy protection is recorded as foreground information; and other individual information except the individual information that requires privacy protection is recorded as background information.

[0135] Optionally, in the embodiment of the present disclosure, refer to Figure 3 As shown, a schematic diagram of the network structure of a submodule of a boundary positioning information output module is provided.

[0136] In the embodiment of the present disclosure, the total loss function of the individual positioning sub-network can be expressed by the following formula:

[0137] loss(x,y) total =α1loss(x,y) class +α2loss(x,y) boundingbox +α3loss(x,y) mask Formula 1

[0138] Among them, α1loss(x, y) class is the category attribute information constraint loss function, α2loss(x, y) boundingbox is the target prediction box constraint loss function, α3loss(x, y) mask is the mask-constrained loss function, and α1, α2, and α3 correspond to the weights of the three loss functions in the total loss function. In the disclosed embodiment, the target prediction box-constrained loss function, the category attribute information-constrained loss function, and the mask-constrained loss function described above enable targeted prediction of individual information in the individual positioning subnetwork, thereby achieving better convergence of the individual positioning subnetwork.

[0139] The following is a brief introduction to the target prediction box constraint loss function, category attribute information constraint loss function, and mask constraint loss function.

[0140] 1. Constraining loss function for target prediction box

[0141] In the embodiment of the present disclosure, in the boundary positioning information output module, the real boundary information can be captured by selecting sub-modules of different scales and different pixel shapes.

[0142] During training, the target prediction box loss function is used to match all possible predicted contour information for an individual with the true contour information, obtaining the corresponding intersection-over-union ratio. Based on this intersection-over-union ratio, candidate boxes are determined that meet a preset threshold with the true boundary information. The obtained candidate boxes are then calculated with the true contour information to obtain the smoothed L1 loss.

[0143] Optionally, in the embodiment of the present disclosure, the preset threshold value can be arbitrarily set according to actual usage, which will not be described in detail here.

[0144] Optionally, in the embodiment of the present disclosure, the target prediction box loss function may be expressed by the following formula:

[0145]

[0146] Among them, i is the serial number, y i Represents the true profile information of an individual whose information needs to be protected in a training sample, numbered i, f(x i ) represents the predicted contour information corresponding to the serial number i obtained by the boundary positioning information output module, and n represents the total number of contour information preset for the above-mentioned individual information that needs to be privacy protected.

[0147] In the embodiment of the present disclosure, based on the above formula, for an individual information of a training sample that needs to be privacy protected, the obtained predicted profile information can be compared with the true profile information to obtain a corresponding first loss value.

[0148] 2. Constraining loss function for category attribute information

[0149] In the embodiment of the present disclosure, the category attribute information constraint loss function can be expressed by using the following formula:

[0150] loss(x,y) class =-∑ i y i logf(x i ) Formula 3

[0151] Among them, i is the serial number, y i Represents the true category attribute information of an individual whose information needs to be protected in a training sample, numbered i, f(x i ) represents the predicted category attribute information corresponding to the serial number i obtained through the boundary positioning information output module.

[0152] Optionally, in the embodiment of the present disclosure, the category attribute information constraint loss function adopts a multi-classification loss function. i and f(x i ) uses cross entropy to measure whether the distribution of predicted category attribute information and true category attribute information is consistent.

[0153] In the embodiment of the present disclosure, based on the above formula, for an individual information of a training sample that needs to be privacy protected, the predicted category attribute information can be compared with the actual category attribute information to obtain a corresponding second loss value.

[0154] 3. Mask Constrained Loss Function

[0155] In the disclosed embodiment, the mask constraint loss function is used to refine the boundary positioning information corresponding to the individual information, thereby constraining the prototype mask of the individual information.

[0156] It should be noted that, based on the above-mentioned target prediction box constraint loss function and category attribute information constraint loss function, the coordinates and category attribute information corresponding to the boundary positioning information of a training sample have been predicted respectively. Then, based on the mask constraint loss function, the coefficients of each positive sample are obtained from the branch of the boundary positioning information output module, and the obtained coefficients are linearly combined to obtain the total coefficient after the linear combination; then, the prototype mask is obtained, and the obtained total coefficient is multiplied with the matrix corresponding to the prototype mask, and the obtained product is matched with the label.

[0157] In the embodiment of the present disclosure, the mask constraint loss function can be expressed by the following formula:

[0158] loss(x,y) mask =-[ylogf(x)+(1-y)log(1-f(x))] Formula 4

[0159] Among them, y represents the true labeling information of an individual information that needs to be privacy protected contained in a training sample, and f(x) represents the predicted labeling information of the foreground information and background information of an individual information that needs to be privacy protected contained in a training sample.

[0160] In the embodiment of the present disclosure, based on the above formula, the predicted labeling information can be compared with the actual labeling information for an individual information of a training sample that needs to be privacy protected to obtain a corresponding third loss value.

[0161] In this way, in the embodiment of the present disclosure, a binary classification loss function is used to accurately characterize the boundary positioning information of the specified individual information by judging whether it is foreground information or background information, thereby greatly reducing the constraint difficulty of the loss function.

[0162] Optionally, during training, the coefficients and prototype masks corresponding to each positive sample can be selected to effectively distinguish the intra-class differences between individuals of the same category.

[0163] In the embodiment of the present disclosure, based on the above formula 1 and the obtained first loss value, second loss value, and third loss value, the total loss value corresponding to each training sample in a training sample set is determined; then, based on the obtained total loss values, the model parameters of the individual positioning sub-network to be trained are adjusted.

[0164] See Figure 4 As shown, in the embodiment of the present disclosure, the specific process of training the individual positioning sub-network in the privacy protection network is as follows:

[0165] Step 400: Obtain a training sample set.

[0166] In the disclosed embodiment, before and after training the individual positioning sub-network in the privacy protection network, it is necessary to first obtain a corresponding training sample set.

[0167] Step 410: Perform multiple rounds of iterative training on the individual positioning sub-network to be trained based on the training sample set, and output the trained individual positioning sub-network when the preset convergence condition is met. In one round of iterative training, the following operations are performed:

[0168] It should be noted that, in the embodiment of the present disclosure, when training the individual positioning sub-network to be trained, all training samples in the training sample set are used to train the individual positioning sub-network to be trained, which is called one round of iterative training.

[0169] When executing step 410, relevant functions are implemented by executing steps 4101 to 4106.

[0170] Step 4101: Input each training sample obtained from the training sample set into the feature extraction module in the individual positioning subnetwork to be trained, extract each first feature information corresponding to each training sample, and generate a first feature map corresponding to each training sample; wherein, each training sample is marked with the real boundary positioning information corresponding to the designated area of the designated individual information that needs to be privacy protected.

[0171] Step 4102: Input each generated first feature map into the feature pyramid fusion model in the individual positioning sub-network to be trained, obtain the second feature information corresponding to each individual information contained in each training sample, and fuse the first feature map corresponding to each training sample with the corresponding second feature information obtained to obtain each second feature map corresponding to each training sample.

[0172] Step 4103: Input each of the obtained second feature maps into the prototype mask generation module in the individual positioning sub-network to be trained to generate a preset mask corresponding to each training sample.

[0173] Step 4104: Input each of the obtained preset masks into the boundary positioning information output module in the individual positioning sub-network to be trained, and based on each preset mask, draw the predicted boundary positioning information corresponding to each individual information contained in the corresponding training sample.

[0174] Step 4105: Based on the obtained predicted boundary positioning information and the comparison results with the corresponding real boundary positioning information, the corresponding total loss value is determined.

[0175] In the embodiment of the present disclosure, since each training sample input into the individual positioning subnetwork to be trained carries its own true boundary positioning information, the corresponding total loss value can be determined by comparing each predicted boundary positioning information with its corresponding true boundary positioning information.

[0176] During the specific training process, the total loss value includes the first loss value, the second loss value, and the third loss value; when executing step 4105, the corresponding function can be achieved by performing the following operations.

[0177] Operation 1: Compare the predicted contour information corresponding to each training sample with the corresponding true contour information to obtain a corresponding first loss value;

[0178] Operation 2: Compare the predicted category attribute information corresponding to each training sample with the corresponding true category attribute information to obtain a corresponding second loss value;

[0179] Operation three: Compare the predicted labeling information corresponding to each training sample with the corresponding true labeling information to obtain the corresponding third loss value;

[0180] Operation 4: Integrate the first loss values, the second loss values, and the third loss values obtained according to a preset relationship to obtain a corresponding total loss value.

[0181] Step 4106: Based on the obtained total loss values, adjust the model parameters of the individual positioning sub-network to be trained.

[0182] In the embodiment of the present disclosure, after executing step 4105, a corresponding total loss value is determined. Based on the total loss value, the model parameters of the individual positioning sub-network to be trained are adjusted to prepare the individual positioning sub-network to be trained for the next round of iterative training.

[0183] Step 420: When the preset convergence condition is met, the trained individual positioning sub-network is output.

[0184] In the embodiment of the present disclosure, multiple rounds of iterative training are performed according to the iterative training method described in step 410 until a preset convergence condition is met, thereby obtaining a trained individual positioning sub-network.

[0185] See Figure 5 As shown, in the embodiment of the present disclosure, the specific process of image processing is as follows:

[0186] Step 500: Acquire an image to be processed, wherein the image to be processed contains at least one individual information requiring privacy protection.

[0187] In the embodiment of the present disclosure, the methods for obtaining the image to be processed include but are not limited to the following two methods:

[0188] Method 1: Select any candidate image from an image set and use any candidate image as an image to be processed, wherein the image set includes candidate images that need to be privacy protected.

[0189] Method 2 is to obtain video data, extract frames of images from the obtained video data, and use the frames of images as images to be processed, wherein the frames of images contain at least one individual information that needs to be protected for privacy.

[0190] Step 510: Input the image to be processed into the privacy protection network, and use the individual positioning sub-network in the privacy protection network to identify and locate each individual information contained in the image to be processed, and obtain the boundary positioning information corresponding to each individual information.

[0191] In the embodiment of the present disclosure, the individual positioning sub-network at least includes a feature extraction module, a feature pyramid fusion module, a prototype mask generation module and a boundary positioning information output module; wherein,

[0192] Module 1, a feature extraction module, is used to extract each first feature information contained in the image to be processed and generate a first feature map corresponding to the image to be processed; each first feature information represents a one-dimensional feature information contained in the image to be processed.

[0193] Module 2, a feature pyramid fusion module, is used to obtain second feature information corresponding to each individual information contained in the image to be processed, and fuse the first feature map with each second feature information to generate a second feature map corresponding to the image to be processed, wherein each second feature information contains two-dimensional feature information corresponding to the individual information associated with one second feature information;

[0194] Module 3, prototype mask generation module, is used to generate a preset mask, wherein the preset mask is obtained by fusing the first feature map and each third feature information obtained. Each third feature information is obtained through the feature pyramid fusion module and is one of the first feature information among the first feature information except for the second feature information obtained.

[0195] Module 4, a boundary positioning information output module, is used to output the boundary positioning information corresponding to each individual information.

[0196] In the specific implementation, after the image to be processed is input into the privacy protection network, it enters the feature extraction module, feature pyramid fusion module, prototype mask generation module and boundary positioning information output module in sequence.

[0197] In the specific implementation, first, the feature extraction module receives the image to be processed, extracts the first feature information contained in the image to be processed, and generates a first feature map corresponding to the image to be processed based on the first feature information; each first feature information represents a one-dimensional feature information contained in the image to be processed.

[0198] Secondly, the feature pyramid fusion module obtains the second feature information corresponding to each individual information contained in the image to be processed, and fuses the first feature map with each second feature information to generate a second feature map corresponding to the image to be processed, wherein each second feature information contains two-dimensional feature information corresponding to the individual information associated with one second feature information.

[0199] Again, the prototype mask generation module fuses the first feature map and the obtained third feature information to generate a preset mask, wherein each third feature information is obtained through the feature pyramid fusion module and is one of the first feature information among the first feature information except for the obtained second feature information.

[0200] Finally, the boundary positioning information output module outputs the boundary positioning information corresponding to each individual information; specifically, in the image to be processed, based on the preset mask, the boundary positioning information corresponding to each individual information contained in the preset mask is drawn, and the drawn image is used as the positioning image to output the positioning image.

[0201] Step 520: Using the individual protection subnetwork in the privacy protection network, perform privacy protection processing on the positioning areas associated with each obtained boundary positioning information to obtain a target image corresponding to the image to be processed; wherein the individual protection subnetwork performs privacy protection processing on the image to be processed based on a preset mask having the same appearance characteristics as the image to be processed; the preset mask is generated by the individual positioning subnetwork and is used to shield a specified area in the specified individual information.

[0202] In the embodiments of the present disclosure, see Figure 6 As shown, when executing step 520, the target image corresponding to the image to be processed can be obtained by executing the following steps:

[0203] Step 5201: Obtain a preset mask, and use the individual protection sub-network to superimpose the preset mask and the positioning image to obtain a superimposed image.

[0204] In the disclosed embodiment, the individual protection subnetwork in the privacy protection network obtains a preset mask corresponding to the image to be processed from the individual positioning subnetwork, and then superimposes the preset mask and the positioning image output by the individual positioning subnetwork to obtain a superimposed image.

[0205] Step 5202: Obtain the correspondence between preset category attribute information and foreground color.

[0206] In the embodiment of the present disclosure, in the superimposed image, the positioning areas associated with the boundary positioning information corresponding to each individual information that needs to be protected for privacy have been preliminarily hidden. Then, the correspondence between the preset category attribute information and the foreground color is obtained. In this way, the corresponding foreground color can be filled for each obtained positioning area according to user needs.

[0207] Step 5203: Based on the category attribute information contained in each obtained boundary positioning information and the preset correspondence between the category attribute information and the foreground color, the positioning area corresponding to each individual information is filled with the corresponding foreground color.

[0208] In the embodiment of the present disclosure, each piece of boundary positioning information includes corresponding category attribute information. Therefore, according to the preset correspondence between the category attribute information and the foreground color, the corresponding foreground color is filled into the positioning area associated with each piece of boundary positioning information.

[0209] Step 5204: Label the category attribute information corresponding to each individual information, obtain and output the target image.

[0210] In the embodiment of the present disclosure, after executing step 5203, the corresponding foreground color has been filled in the positioning area associated with each boundary positioning information. When executing step 5204, the category attribute information corresponding to each individual information is marked, the target image corresponding to the image to be processed is obtained, and the target image is output.

[0211] Optionally, for the input video data, the output target images need to be spliced in time sequence to obtain the corresponding privacy-protected video file.

[0212] The above embodiment is further described in detail below using specific examples.

[0213] See Figure 7 As shown, assume that user A inputs the original image A into the privacy-preserving network.

[0214] First, the feature extraction module in the individual positioning sub-network receives the input original image A, performs feature extraction on the original image A, obtains the first feature information contained in the original image A, and inputs the generated first feature map corresponding to the original image A into the feature pyramid fusion module;

[0215] After receiving the first feature map, the feature pyramid fusion module obtains the second feature information corresponding to each individual information, fuses the first feature map with each second feature information, generates the second feature map corresponding to the original image A, and inputs the second feature map into the prototype mask generation module;

[0216] After receiving the second feature map, the prototype mask generation module generates a preset mask corresponding to the specified area of the specified individual information, and inputs the preset mask and the second feature map into the boundary positioning information output module;

[0217] After receiving the preset mask and the second feature map, the boundary positioning information output module draws the boundary positioning information corresponding to each individual information contained in the preset mask in the original image A based on the preset mask, and uses the drawn original image A as the positioning image A and outputs the positioning image A.

[0218] Then, after the individual positioning sub-network outputs the positioning image A, and after determining that the boundary positioning information in the positioning image A meets the preset conditions, the individual positioning sub-network inputs the positioning image A and the preset mask into the individual protection sub-network. After receiving the positioning image A and the preset mask, the individual protection sub-network superimposes the preset mask and the positioning image A to obtain the superimposed image A.

[0219] Optionally, in the embodiment of the present disclosure, the individual protection subnetwork can fill the corresponding positioning areas of each individual information in the superimposed image A with corresponding foreground colors according to the preset correspondence between category attribute information and foreground color, as well as the category attribute information corresponding to each individual information contained in the superimposed image A, and mark the category attribute information corresponding to each individual information to obtain and output the target image A.

[0220] Based on the same inventive concept, see Figure 8 As shown, an embodiment of the present disclosure provides an image processing device, including:

[0221] An acquisition module 810 is configured to acquire an image to be processed, wherein the image to be processed contains at least one individual information requiring privacy protection;

[0222] An acquisition module 820 is configured to input the image to be processed into a privacy-preserving network, and employ an individual positioning subnetwork within the privacy-preserving network to identify and locate each individual information contained in the image to be processed, thereby obtaining boundary positioning information corresponding to each individual information;

[0223] Output module 830 is used to use the individual protection sub-network in the privacy protection network to perform privacy protection processing on the positioning areas associated with each piece of boundary positioning information obtained, so as to obtain a target image corresponding to the image to be processed; wherein, the individual protection sub-network performs privacy protection processing on the image to be processed based on a preset mask having the same appearance features as the image to be processed; the preset mask is generated by the individual positioning sub-network and is used to shield a specified area in the specified individual information.

[0224] Optionally, the acquiring module 810 is configured to:

[0225] Select any candidate image from an image set and use the any candidate image as the image to be processed, wherein the image set includes candidate images that require privacy protection; or

[0226] Video data is acquired, and frames of images are extracted from the video data, and the frames of images are used as the images to be processed, wherein the frames of images contain at least one individual information that needs to be privacy protected.

[0227] Optionally, the individual positioning sub-network comprises at least a feature extraction module, a feature pyramid fusion module, a prototype mask generation module and a boundary positioning information output module; wherein,

[0228] The feature extraction module is used to extract each first feature information contained in the image to be processed and generate a first feature map corresponding to the image to be processed; each first feature information represents a one-dimensional feature information contained in the image to be processed;

[0229] The feature pyramid fusion module is configured to obtain second feature information corresponding to each piece of individual information contained in the image to be processed, and fuse the first feature map with each piece of second feature information to generate a second feature map corresponding to the image to be processed, wherein each piece of second feature information contains two-dimensional feature information corresponding to individual information associated with one piece of second feature information;

[0230] The prototype mask generation module is configured to generate the preset mask, wherein the preset mask is obtained by fusing the first feature map and each piece of third feature information obtained, each piece of third feature information is obtained by the feature pyramid fusion module and is one piece of first feature information among each piece of first feature information except each piece of second feature information obtained;

[0231] The boundary positioning information output module is used to output the boundary positioning information corresponding to each piece of individual information.

[0232] Optionally, the image to be processed is input into a privacy-preserving network, and an individual positioning subnetwork in the privacy-preserving network is used to identify and locate each individual information contained in the image to be processed, thereby obtaining boundary positioning information corresponding to each individual information. The obtaining module 820 is configured to:

[0233] Inputting the image to be processed into the feature extraction module in the privacy protection network, extracting each first feature information contained in the image to be processed, and generating a first feature map corresponding to the image to be processed based on the each first feature information;

[0234] Inputting the first feature map into the feature pyramid fusion module, obtaining second feature information corresponding to each individual information contained in the image to be processed, and fusing the first feature map with each obtained second feature information to obtain a second feature map corresponding to the image to be processed;

[0235] Inputting the second feature map into the prototype mask generation module to generate a preset mask corresponding to the image to be processed;

[0236] The preset mask is input into the boundary positioning information output module. In the image to be processed, based on the preset mask, the boundary positioning information corresponding to each individual information contained in the preset mask is drawn, and the drawn image is used as a positioning image, and the positioning image is output.

[0237] Optionally, the individual protection subnetwork in the privacy protection network is used to perform privacy protection processing on the positioning areas associated with each obtained boundary positioning information to obtain a target image corresponding to the image to be processed. The output module 830 is used to:

[0238] Obtaining the preset mask, and using the individual protection sub-network to superimpose the preset mask and the positioning image to obtain a superimposed image;

[0239] Obtaining the correspondence between preset category attribute information and foreground color;

[0240] Filling the positioning areas corresponding to the individual pieces of information with corresponding foreground colors based on the category attribute information contained in the obtained boundary positioning information and the preset correspondence between the category attribute information and the foreground color;

[0241] The category attribute information corresponding to each of the individual information is labeled, and the target image is obtained and output.

[0242] Optionally, the individual positioning sub-network is trained in the following manner:

[0243] The individual positioning sub-network to be trained is trained for multiple rounds of iterative training based on the training sample set, and when the preset convergence conditions are met, the trained individual positioning sub-network is output. In one cycle of iterative training, the following operations are performed:

[0244] Inputting each training sample obtained from the training sample set into the feature extraction module of the individual positioning subnetwork to be trained, extracting each first feature information corresponding to each training sample, and generating a first feature map corresponding to each training sample; wherein, each training sample includes the actual boundary positioning information corresponding to the designated area marked with the designated individual information that needs to be privacy protected;

[0245] Inputting each generated first feature map into the feature pyramid fusion model in the individual positioning subnetwork to be trained, obtaining second feature information corresponding to each individual information contained in each training sample, and fusing the first feature map corresponding to each training sample with the corresponding second feature information obtained to obtain each second feature map corresponding to each training sample;

[0246] Inputting each second feature map obtained into the prototype mask generation module in the individual positioning sub-network to be trained, respectively, to generate a preset mask corresponding to each training sample;

[0247] The obtained preset masks are respectively input into the boundary positioning information output module in the individual positioning sub-network to be trained, and based on the respective preset masks, the predicted boundary positioning information corresponding to each individual information contained in the corresponding training samples is drawn; and based on the comparison results of each predicted boundary positioning information obtained and the corresponding true boundary positioning information, the corresponding total loss value is determined, and based on the obtained total loss values, the model parameters of the individual positioning sub-network to be trained are adjusted.

[0248] Optionally, the total loss value includes a first loss value, a second loss value and a third loss value;

[0249] Based on the comparison results of each predicted boundary positioning information and the corresponding real boundary positioning information, the corresponding total loss value is determined, and the output module 830 is used to:

[0250] Comparing the predicted contour information corresponding to each of the obtained training samples with the corresponding true contour information to obtain a corresponding first loss value;

[0251] Comparing the predicted category attribute information corresponding to each of the obtained training samples with the corresponding true category attribute information to obtain a corresponding second loss value;

[0252] Comparing the predicted labeling information corresponding to each of the obtained training samples with the corresponding true labeling information to obtain a corresponding third loss value;

[0253] According to the preset relationship, the obtained first loss values, the obtained second loss values, and the obtained third loss values are respectively integrated to obtain the corresponding total loss value.

[0254] See Figure 9 As shown, the computer device includes a memory 901 and a controller 902. Specifically:

[0255] The memory 901 is used to store computer programs that can be executed by the controller 902 .

[0256] The controller 902 is connected to the memory and is configured to execute any one of the methods executed by the data transmission device in the above embodiments.

[0257] Based on the same inventive concept, an embodiment of the present disclosure provides a computer-readable storage medium. When a computer program in the computer-readable storage medium is executed by a processor, any one of the methods performed by the data transmission device in the above embodiments can be executed.

[0258] In summary, in the embodiment of the present disclosure, an image to be processed is obtained, wherein the image to be processed contains at least one individual information that needs to be privacy protected; then the image to be processed is input into a privacy protection network, and the individual positioning sub-network in the privacy protection network is used to identify and locate each individual information contained in the image to be processed, and obtain the boundary positioning information corresponding to each individual information; finally, the individual protection sub-network in the privacy protection network is used to perform privacy protection processing on the positioning area associated with each obtained boundary positioning information, and obtain the target image corresponding to the image to be processed; wherein the individual protection sub-network performs privacy protection processing on the image to be processed based on a preset mask having the same appearance features as the image to be processed; the preset mask is generated by the individual positioning sub-network and is used to shield a specified area in the specified individual information. domain; thus, since the privacy protection network is trained based on each training sample and the corresponding real boundary positioning information, when the image to be processed is input into the trained privacy protection network, the specified area in the specified individual information that needs to be privacy protected can be accurately located, and a preset mask for shielding the specified area in the specified individual information can be obtained. Then, based on the obtained preset mask, the specified area in the specified individual information in the image to be processed can be better privacy-protected, thereby improving the processing efficiency of privacy protection. Moreover, on the basis of realizing the privacy protection of the specified area in the specified individual information, other individual information that does not need to be privacy protected can also be clearly displayed, thereby ensuring the security of data privacy without affecting the readability of the image.

[0259] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0260] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0261] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0262] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0263] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. An image processing method, characterized in that: include: Acquiring an image to be processed, wherein the image to be processed contains at least one individual information requiring privacy protection; Inputting the image to be processed into a privacy protection network, and using the individual positioning sub-network in the privacy protection network to respectively identify and locate each individual information contained in the image to be processed, and obtaining boundary positioning information corresponding to each individual information; Using the individual protection subnetwork in the privacy protection network, privacy protection processing is performed on the positioning areas associated with each piece of obtained boundary positioning information to obtain a target image corresponding to the image to be processed; wherein the individual protection subnetwork performs privacy protection processing on the image to be processed based on a preset mask having the same appearance characteristics as the image to be processed; the preset mask is generated by the individual positioning subnetwork and is used to shield a specified area in the specified individual information; The individual positioning subnetwork includes at least a feature extraction module, a feature pyramid fusion module, a prototype mask generation module, and a boundary positioning information output module. The feature extraction module is used to extract each first feature information contained in the image to be processed and generate a first feature map corresponding to the image to be processed. Each first feature information represents a one-dimensional feature information contained in the image to be processed. The feature pyramid fusion module is used to obtain second feature information corresponding to each individual information contained in the image to be processed, and fuse the first feature map with each second feature information to generate a second feature map corresponding to the image to be processed, wherein each second feature information contains two-dimensional feature information corresponding to the individual information associated with the second feature information. The prototype mask generation module is used to generate the preset mask, wherein the preset mask is obtained by fusing the first feature map and each third feature information obtained. Each third feature information is obtained by the feature pyramid fusion module and is one first feature information among each first feature information, excluding each second feature information obtained. The boundary positioning information output module is used to output the boundary positioning information corresponding to each individual information.

2. The method according to claim 1, wherein The step of obtaining an image to be processed includes: Select any candidate image from an image set and use the any candidate image as the image to be processed, wherein the image set includes candidate images that require privacy protection; or Video data is acquired, and frames of images are extracted from the video data, and the frames of images are used as the images to be processed, wherein the frames of images contain at least one individual information that needs to be privacy protected.

3. The method according to claim 1, wherein Inputting the image to be processed into the privacy protection network, using the individual positioning sub-network in the privacy protection network to respectively identify and locate each individual information contained in the image to be processed, and obtaining the boundary positioning information corresponding to each individual information, including: Inputting the image to be processed into the feature extraction module in the privacy protection network, extracting each first feature information contained in the image to be processed, and generating a first feature map corresponding to the image to be processed based on the each first feature information; Inputting the first feature map into the feature pyramid fusion module, obtaining second feature information corresponding to each individual information contained in the image to be processed, and fusing the first feature map with each obtained second feature information to obtain a second feature map corresponding to the image to be processed; Inputting the second feature map into the prototype mask generation module to generate a preset mask corresponding to the image to be processed; The preset mask is input into the boundary positioning information output module. In the image to be processed, based on the preset mask, the boundary positioning information corresponding to each individual information contained in the preset mask is drawn, and the drawn image is used as a positioning image, and the positioning image is output.

4. The method according to claim 3, wherein The individual protection sub-network in the privacy protection network is used to perform privacy protection processing on the positioning areas associated with each obtained boundary positioning information to obtain a target image corresponding to the image to be processed, including: Obtaining the preset mask, and using the individual protection sub-network to superimpose the preset mask and the positioning image to obtain a superimposed image; Obtaining the correspondence between preset category attribute information and foreground color; Filling the positioning areas corresponding to the individual pieces of information with corresponding foreground colors based on the category attribute information contained in the obtained boundary positioning information and the preset correspondence between the category attribute information and the foreground color; The category attribute information corresponding to each of the individual information is labeled, and the target image is obtained and output.

5. The method according to any one of claims 1 to 4, characterized in that: The individual positioning sub-network is trained in the following way: The individual positioning sub-network to be trained is trained for multiple rounds of iterative training based on the training sample set, and when the preset convergence conditions are met, the trained individual positioning sub-network is output. In one cycle of iterative training, the following operations are performed: Inputting each training sample obtained from the training sample set into the feature extraction module of the individual positioning subnetwork to be trained, extracting each first feature information corresponding to each training sample, and generating a first feature map corresponding to each training sample; wherein, each training sample includes the actual boundary positioning information corresponding to the designated area marked with the designated individual information that needs to be privacy protected; Inputting each generated first feature map into the feature pyramid fusion model in the individual positioning subnetwork to be trained, obtaining second feature information corresponding to each individual information contained in each training sample, and fusing the first feature map corresponding to each training sample with the corresponding second feature information obtained to obtain each second feature map corresponding to each training sample; Inputting each second feature map obtained into the prototype mask generation module in the individual positioning sub-network to be trained, respectively, to generate a preset mask corresponding to each training sample; The obtained preset masks are respectively input into the boundary positioning information output module in the individual positioning sub-network to be trained, and based on the respective preset masks, the predicted boundary positioning information corresponding to each individual information contained in the corresponding training samples is drawn; and based on the comparison results of each predicted boundary positioning information obtained and the corresponding true boundary positioning information, the corresponding total loss value is determined, and based on the obtained total loss values, the model parameters of the individual positioning sub-network to be trained are adjusted.

6. The method according to claim 5, wherein The total loss value includes a first loss value, a second loss value and a third loss value; The determining of the respective corresponding total loss values based on the comparison results of the obtained respective predicted boundary positioning information and the respective corresponding real boundary positioning information includes: Comparing the predicted contour information corresponding to each of the obtained training samples with the corresponding true contour information to obtain a corresponding first loss value; Comparing the predicted category attribute information corresponding to each of the obtained training samples with the corresponding true category attribute information to obtain a corresponding second loss value; Comparing the predicted labeling information corresponding to each of the obtained training samples with the corresponding true labeling information to obtain a corresponding third loss value; According to the preset relationship, the obtained first loss values, the obtained second loss values, and the obtained third loss values are respectively integrated to obtain the corresponding total loss value.

7. An image processing device, characterized in that: include: An acquisition module, configured to acquire an image to be processed, wherein the image to be processed contains at least one individual information requiring privacy protection; an acquisition module, configured to input the image to be processed into a privacy-preserving network, and employ an individual positioning subnetwork in the privacy-preserving network to identify and locate each individual information contained in the image to be processed, thereby obtaining boundary positioning information corresponding to each individual information; an output module, configured to perform privacy protection processing on the positioning areas associated with each piece of obtained boundary positioning information using the individual protection subnetwork in the privacy protection network, thereby obtaining a target image corresponding to the image to be processed; wherein the individual protection subnetwork performs privacy protection processing on the image to be processed based on a preset mask having the same appearance features as the image to be processed; the preset mask is generated by the individual positioning subnetwork and is used to shield a specified area in the specified individual information; The individual positioning subnetwork includes at least a feature extraction module, a feature pyramid fusion module, a prototype mask generation module, and a boundary positioning information output module. The feature extraction module is used to extract each first feature information contained in the image to be processed and generate a first feature map corresponding to the image to be processed. Each first feature information represents a one-dimensional feature information contained in the image to be processed. The feature pyramid fusion module is used to obtain second feature information corresponding to each individual information contained in the image to be processed, and fuse the first feature map with each second feature information to generate a second feature map corresponding to the image to be processed, wherein each second feature information contains two-dimensional feature information corresponding to the individual information associated with the second feature information. The prototype mask generation module is used to generate the preset mask, wherein the preset mask is obtained by fusing the first feature map and each third feature information obtained. Each third feature information is obtained by the feature pyramid fusion module and is one first feature information among each first feature information, excluding each second feature information obtained. The boundary positioning information output module is used to output the boundary positioning information corresponding to each individual information.

8. A computer device, characterized in that: include: a memory for storing a computer program executable by the controller; The controller is connected to the memory and is configured to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor, the processor is enabled to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image processing method, device and apparatus and storage medium

    CN112347512A

  • Object attribute detection method and device, neural network training method and device, and regional detection method and device

    WO2018121690A1