Multi-select full-body anonymization method, device and apparatus for character image data
Through multi-select whole-body anonymization method, combined with the graphic and text model, anonymization prompt word engineering, palette algorithm and Controlnet, the problems of limited anonymization selection and insufficient diversity in the existing technology are solved, and high-quality and high-diversity anonymization effect is achieved.
Patent Information
- Application Number
- CN202510270146.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Existing full-body anonymization technology is difficult to provide different degrees of anonymization choices while ensuring data quality and fidelity. At the same time, it cannot effectively ensure the diversity of anonymous data, which increases the risk of being traced by data.
The multi-selective whole-body anonymization method is adopted to obtain the human body part to be anonymized through the target segmentation algorithm, and combined with the graphical text model, anonymization prompt word engineering, color palette algorithm and Controlnet, the diffusion model is guided to perform the Inpainting task to generate the anonymized result map.
It realizes that while ensuring high anonymous data quality and fidelity, it provides different degrees of anonymous choices, improves the diversity of anonymous data and reduces the risk of being traced by data.
Smart Images

Figure CN119808160B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of artificial intelligence security technology, and in particular to a multi-select full-body anonymization method, device and equipment for character image data. Background Art
[0002] In the data age, the importance of data is self-evident. However, with the surge in online social networking and ubiquitous data collection, people are also deeply concerned about personal privacy and information security. The emergence of anonymization technology is precisely to protect users' personal privacy while effectively exploring and utilizing data. Portrait image data contains a large amount of personal privacy information. If it is not anonymized, direct use will infringe on personal privacy, threaten personal information security, and violate relevant data protection regulations. Anonymization technology is needed to protect personal privacy information in the data while retaining important information in the data. The privacy information of the human body does not only exist on the face. Parts outside the face still contain a lot of personal identification information. The anonymization of the human body should be extended from the face to the whole body.
[0003] Traditional anonymization techniques include blurring, pixelation, etc. However, these methods seriously damage the data and cause the effective information in the data to be greatly destroyed. In order to protect people's privacy while retaining the important features in the data and preserving the utility of the data, the generative model is introduced into the anonymization technology. Based on the generation ability of the deep generative model, anonymization is achieved by performing the Inpainting task on the image or transforming the image while retaining the data distribution. Existing anonymization technologies widely focus on face anonymization, but personal sensitive information does not only exist in the face. The parts outside the face still contain a lot of personal identification information. If only the face part is anonymized, the data will still leak privacy. Therefore, the new anonymization technology is extended from the face to the whole body. Through the whole body anonymization technology, the privacy of the original data is fully protected while retaining the characteristic information of the data. The whole body anonymization method based on the adversarial generative network anonymizes the sensitive information of the whole body, but the data quality and fidelity are not good, especially in the case of high resolution. With the development of the diffusion model, the diffusion model has shown a strong generation ability. The whole body anonymization method based on the diffusion model has excellent data quality and fidelity after anonymization, but the diversity of anonymous data is poor. The current full-body anonymization method can only perform relatively single anonymization and cannot retain different degrees of data attributes to achieve different degrees of anonymization. At the same time, it cannot guarantee the diversity of anonymous data. For full-body anonymization, including image diversity and face diversity, if the diversity of anonymous data cannot be guaranteed, the original data cannot be further effectively utilized, and the risk of data traceability is increased. Summary of the invention
[0004] In view of the technical problems existing in the prior art, the present invention proposes a multi-select full-body anonymization method, device and equipment for character image data. Under the premise of ensuring image quality and anonymity effect, different degrees of anonymization options can be provided at the macro-semantic level, while ensuring the diversity of full-body anonymized images and faces.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] On the one hand, a multi-select full-body anonymization method for person image data is provided, comprising:
[0007] Input character image data;
[0008] Obtain a mask image of the human body part to be anonymized in the human image data;
[0009] Use the image-to-text model to obtain the text description of the mask image of the human body part to be anonymized;
[0010] The text description is passed through an anonymous prompt word project to obtain an anonymous prompt word;
[0011] A palette image representing the color feature distribution of the person image data is obtained by using a palette algorithm;
[0012] Obtain a false and random face identity feature based on the face database;
[0013] Controlnet is introduced to control the non-sensitive attribute information of the character image data as image prompts to guide the diffusion model image generation process;
[0014] According to the obtained mask image of the human body part to be anonymized, the anonymization prompt word, the palette image and the false and random face identity features, the Inpainting task based on the diffusion model is guided to generate a new character image to replace the human body part to be anonymized in the character image data, and the anonymized result image is obtained.
[0015] On the other hand, a multi-select full-body anonymization device for character image data is provided, comprising:
[0016] The first module is used to input character image data;
[0017] The second module is used to obtain a mask image of the human body part to be anonymized in the human image data;
[0018] The third module is used to obtain the text description of the mask image of the human body part to be anonymized by using the image-to-text model;
[0019] The fourth module is used to obtain anonymized prompt words by passing the text description through an anonymized prompt word project;
[0020] A fifth module is used to obtain a palette image representing the color feature distribution of the character image data using a palette algorithm;
[0021] The sixth module is used to obtain a false and random face identity feature based on the face database;
[0022] The seventh module is used to introduce the non-sensitive attribute information of the ControlNet controlled person image data as an image prompt to guide the diffusion model image generation process;
[0023] The eighth module is used to guide the Inpainting task based on the diffusion model according to the obtained mask image of the human body part to be anonymized, the anonymization prompt word, the palette image and the false and random face identity features, generate a new character image to replace the human body part to be anonymized in the character image data, and obtain the anonymized result image.
[0024] On the other hand, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned multi-selection full-body anonymization method for character image data are implemented.
[0025] On the other hand, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned multi-selection full-body anonymization method for character image data when the computer program is executed by a processor.
[0026] On the other hand, the present invention provides a computer program product, which is stored on a computer-readable storage medium and includes computer instructions, which, when executed by a processor, enable a computer device to implement the steps of the above-mentioned multi-select full-body anonymization method for character image data.
[0027] Compared with the prior art, the present invention proposes a multi-select full-body anonymization method for character image data. First, a target segmentation algorithm is used to obtain the human body part to be anonymized; the human body part to be anonymized is passed through a picture-to-text model to obtain a text description of the character, and the text description is passed through an anonymization prompt word project to obtain anonymization prompt words; at the same time, a palette algorithm is used to obtain the color feature distribution of the image; Controlnet is introduced to control the non-sensitive attribute information of the character image data; a face guidance module is used to obtain false and random face identity features and input them into the anonymization generation process, and according to the different image information obtained, the Inpainting task based on the diffusion model is guided to replace the sensitive character image in the original image to obtain the anonymized result image. The present invention has the following advantages:
[0028] The present invention introduces multiple ControlNets in the full-body anonymization based on the diffusion model, including prompt word engineering, face guidance and palette algorithm, which realizes more detailed anonymous attribute control while ensuring high anonymous data quality, and greatly improves the diversity of anonymous data.
[0029] The present invention proposes a new multi-select full-body anonymization method, which provides different degrees of anonymization selection at the semantic level under the premise of ensuring data quality, fidelity and anonymity effect. Specifically, it can dynamically control the palette algorithm and the anonymization prompt word engineering according to the different anonymization degrees selected by the user, wherein the palette algorithm is used to control the color distribution, and the prompt word engineering provides random colors and random details to obtain different anonymization selections and retain different degrees of data attributes to achieve different anonymization degree effects, while ensuring the diversity of full-body anonymization pictures and the diversity of faces. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0031] Figure 1 It is a flowchart of a multi-select full-body anonymization method for character image data provided by an embodiment. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] Reference Figure 1 In one embodiment, a method for anonymizing a person's image data using multiple selections is provided, comprising:
[0034] Input character image data;
[0035] Obtain a mask image of the human body part to be anonymized in the human image data;
[0036] Use the image-to-text model to obtain the text description of the mask image of the human body part to be anonymized;
[0037] The text description is passed through an anonymous prompt word project to obtain an anonymous prompt word;
[0038] A palette image representing the color feature distribution of the person image data is obtained by using a palette algorithm;
[0039] Obtain a false and random face identity feature based on the face database;
[0040] Controlnet is introduced to control the non-sensitive attribute information of the character image data as image prompts to guide the diffusion model image generation process;
[0041] According to the obtained mask image of the human body part to be anonymized, the anonymization prompt word, the palette image and the false and random face identity features, the Inpainting task based on the diffusion model is guided to generate a new character image to replace the human body part to be anonymized in the character image data, and the anonymized result image is obtained.
[0042] Furthermore, a mask image of the human body part to be anonymized in the human image data is obtained, and the specific method is as follows:
[0043] For the input character image data , use the yolo algorithm to obtain character image data The classification of each pixel in the Person class is set to 255, and the pixel values of all other pixels are set to 0. In this way, the mask image of the human body part to be anonymized is obtained. .
[0044] Next, the image-to-text model is used to obtain the text description of the human body part image to be anonymized, including:
[0045] Character image data Mask image of the human body part to be anonymized Overlap and keep only the mask image The part with pixel value of 255 in the middle gets the overlapping picture ;
[0046] ;
[0047] Overlapping pictures Remove irrelevant information from the picture and keep only the person in the picture, that is, the picture Focus the image on the person.
[0048] Use the graph-to-text model to obtain pictures The character text description is used as the text description of the human body part image to be anonymized, including the description of the clothing features worn by the character. The image-to-text model is not limited. For example, LLava is used as the image-to-text model, and "describe this character clothing, gender, and status" is used as the query. Obtain the character text description containing the desired content.
[0049] In one embodiment, the text description is subjected to an anonymous prompt word engineering to obtain an anonymous prompt word, including:
[0050] Prompt word engineering is performed on the acquired text description: first, data cleaning is performed on the text description to remove irrelevant descriptions, and then clothing keywords are extracted. The acquired clothing keywords are randomly anonymized using the pre-built color text library and clothing detail text library to obtain anonymous prompt words with different anonymization levels. Depending on the degree of anonymization selected, you can freely decide whether to use random colors or random details to obtain anonymous prompt words with different anonymization levels.
[0051] Generate n different anonymous prompt words for the text description , use the clip encoder to combine n anonymous prompt words with the character image data They are all mapped to a unified latent space of text and image. By calculating the cosine similarity between the text and image vectors, the alignment relationship between the image and text is obtained. The anonymous prompt word with the largest difference from the image is selected from n different anonymous prompt words. As the final anonymization prompt :
[0052] ;
[0053] in, For the i An anonymous reminder word, , It is a clip encoder, which is used to map both text and images into a unified text and image latent space.
[0054] The present invention uses a palette algorithm to obtain a palette image representing the color feature distribution of the character image data, as follows:
[0055] For character image data Gaussian blur is performed to remove image details.
[0056] The mathematical formula of the Gaussian convolution kernel is:
[0057] ;
[0058] in, is the standard deviation of the Gaussian distribution and controls the degree of blur. are the coordinates of each element in the convolution kernel.
[0059] For character image data The mathematical formula for Gaussian blur processing can be expressed as:
[0060] ;
[0061] in, is the blurred image at coordinates The pixel value at It is the character image data In coordinates The pixel value at is the Gaussian convolution kernel at coordinates The value at is the size of the convolution kernel.
[0062] Use the K-means algorithm to represent the image after Gaussian blur processing with a certain number of colors, and then perform image enhancement operations to improve the color contrast of the processed image, thereby obtaining a palette image that represents the color feature distribution of the character image data .
[0063] The mathematical formula of the K-means algorithm is as follows:
[0064] ;
[0065] in, It is categories, It is The center of each category (i.e., the representative color), is the pixel value With Category Center The Euclidean distance between .
[0066] The palette image obtained from this , which preserves the color distribution information in the character image data while removing the detail features.
[0067] In the present invention, Controlnet is introduced to control the non-sensitive attribute information of the character image data, including:
[0068] Get person image data The corresponding three types of non-sensitive partial attribute information are soft edge map, posture map, and facial information map. The soft edge map describes the edge contour of the human body in the image. The posture map is used to capture and represent the posture information of the human body in the image. The facial information map contains the key feature points and facial position information of the face.
[0069] The soft edge map, posture map, and facial information map are input into ControlNet as image prompts to guide the diffusion model image generation process.
[0070] In the present invention, a false and random face identity feature is obtained based on the face database, and the specific method is as follows:
[0071] Randomly select n face images from the face database and extract the identity feature vectors corresponding to the n face images , and generate random weights corresponding to n face images , weighted sum these n facial identity features to obtain a false and random facial identity feature as shown in the formula ;
[0072] ;
[0073] According to the obtained mask image of the human body part to be anonymized, the anonymization prompt word, the palette image and the false and random face identity features, the Inpainting task based on the diffusion model is guided, including:
[0074] Use the mask image of the human body part to be anonymized The masking guidance of the Inpainting task of the diffusion model is carried out. According to the different anonymization levels selected by the user, the palette algorithm and the anonymization prompt word engineering are dynamically controlled. The palette algorithm is used to control the color distribution, and the prompt word engineering provides random colors and random details to obtain different anonymization options. Without loss of generality, three macro-semantic level anonymous options can be provided for users to choose. In the highest degree of anonymization, the colors and styles are inconsistent and vary greatly; in the medium degree of anonymization, the colors are consistent, and the style details are inconsistent; in the lowest degree of anonymization, the colors and style details are mostly consistent. The palette algorithm is used to control the color distribution, and the prompt word engineering provides random colors and random details. According to the anonymization level selected by the user, the palette algorithm and the anonymization prompt word engineering modules are dynamically controlled to obtain different anonymity results.
[0075] Combining the idea of img2img, when performing the Inpainting task of the diffusion model, according to the mask image of the human body part to be anonymized For the image part that needs to be reconstructed, that is, the human body part to be anonymized in the character image data, the pure Gaussian noise is replaced by the noisy palette image. That is, the color distribution map obtained by the palette algorithm is reconstructed by adding noise to obtain a noisy palette map, and the noisy palette map is used to replace the pure Gaussian noise.
[0076] Anonymization prompt word Serves as a guide text to guide the anonymization generation process. As a separate facial information guide. In the present invention, the facial identity feature A lightweight adaptive module with decoupled cross attention is used to support identity feature prompts. At the same time, the facial information map obtained by Annotators is used to control the position space of the face through ControlNet to guide the generation of facial parts.
[0077] A new character image will be generated to replace the human body part to be anonymized in the character image data, and an anonymized result image will be obtained.
[0078] To prove the effectiveness of the multi-select full-body anonymization method for person image data provided by the present invention, the data quality and authenticity of the full-body anonymization data generated by the multi-select full-body anonymization method (referred to as MCFA) for person image data provided by the present invention are verified by the following experiments. The evaluation indicators are FID, IS, and LPIPS, which are commonly used to test the quality of generated images. FID evaluates the similarity by calculating the Wasserstein distance (also known as Earth-Mover's distance) between the feature vector distributions of the generated image and the real image. The lower the FID value, the closer the generated image is to the real image in distribution, and the better the quality. IS calculates the index of the KL divergence of the predicted probability that the generated image is correctly classified into each category. A high IS value usually indicates that the generated image is both diverse and realistic. A lower FID corresponds to better data quality, a higher IS indicates better generation quality, and a lower LPIPS value indicates that the perceptual difference between the two images is smaller. FID (Frechet Inception Distance) is used to measure the distribution distance between the generated image and the real image. IS (Inception Score) is used to evaluate the clarity and diversity of the generated image based on the classification confidence. LPIPS (Learned Perceptual Image Patch Similarity) is a perceptual similarity metric based on deep learning, which measures the perceptual similarity of images through pre-trained deep networks (such as AlexNet and VGG). LPIPS can capture the subtle differences in image quality perceived by the human visual system, so it is often used to evaluate the performance of image generation models.
[0079] First, we select a dataset. We select two datasets: VION-HD and COCO-Body. VION-HD is a high-resolution human image dataset with a resolution of 768*1024. COCO-Body is a derived dataset of COCO-dataset for full-body anonymization with a resolution of 288*160.
[0080] In order to conveniently reflect the experimental effect, three SOTA full-body anonymization methods, SG-GAN, DeepPrivacy2, and FADM, are selected for comparison with the multi-select full-body anonymization method for character image data provided by the present invention. At the same time, three multi-select full-body anonymization methods for character image data corresponding to three different anonymization degree selections are provided, which are referred to as MCFA(a), MCFA(b), and MCFA(c), respectively. MCFA(a), MCFA(b), and MCFA(c) correspond to three different anonymization degrees selected in the multi-select full-body anonymization method for character image data, namely, the highest anonymization degree, the medium anonymization degree, and the lowest anonymization degree, respectively. The color styles of the highest anonymization degree are inconsistent and vary greatly (anonymization prompt words (random colors + random details)). The colors of the medium anonymization degree remain consistent, and the style details are inconsistent (palette algorithm + anonymization prompt words (random details)). Most of the lowest anonymization degrees remain consistent. (Palette algorithm + original prompt words).
[0081] Table 1. Results of different methods on VION-HD and COCO-Body image data quality indicators
[0082]
[0083] As shown in Table 1, the data quality effects of different anonymity methods were tested on the VION-HD and COCO-Body datasets using the three indicators of FID, IS, and LPIPS. The GAN-based SG-GAN and DeepPrivacy2 have good generation performance on the COCO-Body dataset with lower resolution, but the generation effect on the high-resolution VION-HD is far inferior to the method based on the diffusion model. The method of the present invention shows good data quality on both high-resolution and low-resolution datasets. At the same time, by comparing the image data quality index results corresponding to the three different anonymization degree selections (MCFA(a), MCFA(b), MCFA(c)) provided by the present invention, it can be seen that as the anonymity level increases, the larger the LPIPS index, that is, the greater the perceived difference between the two images.
[0084] In order to verify the performance of the multi-select full-body anonymization method for character image data provided by the present invention in the diversity of full-body anonymized images, the discussion is divided into two parts, namely, image diversity and face diversity. Five anonymous images are generated for each verification sample to test the diversity. Image diversity represents the overall perceptual diversity of the image, and is judged using LPIPSdiversity. The larger the index, the higher the diversity of the generated images. This embodiment proposes Facediversity, which uses the corresponding face identity features to calculate the Euclidean distance between face features. The larger the index, the higher the face diversity.
[0085] Table 2 Results of image diversity and face diversity
[0086]
[0087] Table 2 contains the results of the indicators on diversity. It is found that the diversity of the whole-body anonymization based on GAN is better than that of FADM based on the diffusion model. The whole-body anonymization framework MCFA proposed in the present invention ensures the quality and authenticity of anonymous data based on the diffusion model, and at the same time achieves good results in image diversity (MCFA(a) is selected here). For face diversity, it can be seen that the MCFA without the face injection module (Face Guide module) (i.e., MCFA(withoutFace Guide)) has a performance of only 0.982 in face diversity, indicating that the diversity of faces is poor. However, the MCFA with the face injection module (Face Guide module) (i.e., MCFA(with Face Guide)) has a significant improvement in the face diversity index, and can obtain anonymized data with high face diversity.
[0088] In order to verify the anonymity effect of the multi-select full-body anonymization method for person image data provided by the present invention, that is, to anonymize sensitive information in the image so that it cannot be recognized. The risk of re-identification is measured by comparing each anonymous image with the real data set through nearest neighbor retrieval, while considering re-identification at the face level and image level.
[0089] For face-level re-identification, the face identity features are obtained, and then the Euclidean distance is calculated to perform face identity comparison. For image-level re-identification, the image is passed through the clip image encoder and a K-NN search is performed based on the obtained image features.
[0090] Table 3 Face-level re-recognition rate
[0091]
[0092] Table 4 Image-level re-identification rate
[0093]
[0094] The indicators in Table 3 and Table 4 illustrate the re-identification effects at the face level and the image level. As can be seen from Table 3, the anonymity effect of the face is relatively outstanding. The face contains the most privacy information and is the top priority for human anonymity. The method of the present invention achieves an excellent face anonymity effect, and the re-identification rate remains at a very low level. For the re-identification rate at the picture level, since the GAN-based SG-GAN and DeepPrivacy2 methods perform poorly in high-resolution anonymous generation, the fuzzy and chaotic generation results will also lead to a low re-identification rate. It can be seen that the multi-select full-body anonymization method for character image data provided by the present invention has different re-identification rates with different anonymity choices, which also represents the retention of different attribute features in the clip semantic space.
[0095] Table 5 Re-identification rate of OSNet on Market1501 dataset under different anonymization methods
[0096]
[0097] OSNet is an algorithm for person re-identification. It uses different full-body anonymization methods to anonymize the Market1501 dataset, and then uses OSNet to re-identify the acquired anonymous data. Table 5 shows the re-identification rate results of different anonymization methods. It is found that since the anonymous selection of MCFA(a) (without palette) of the present invention changes the color distribution of the anonymous person, the re-identification rate is significantly reduced compared to the anonymous selection MCFA(c) (with palette) that keeps its color, indicating that the color distribution of the image is an important basis for the OSNet pedestrian re-identification algorithm to make judgments. By changing the color distribution, the anonymity effect is significantly improved and the risk of being re-identified is reduced.
[0098] Finally, for the anonymization task, it is expected that the anonymized data needs to retain the usefulness of the original data as much as possible. For human body images, the usefulness of human body images is evaluated using two tasks: human body key point estimation and human body semantic segmentation. The human body key points are obtained, and KPd and OKS are used to judge the similarity between the original data and the anonymized data in terms of human body key points. For the human body semantic segmentation task, the human body semantic segmentation map of the original data and the anonymized data is obtained, and the accuracy metric in the semantic segmentation is used to judge the difference between the two human body semantic segmentation maps. The usefulness of anonymous data is judged in two types of human-related tasks: human body key point estimation and human body semantic segmentation.
[0099] Table 6 Performance of different anonymous methods in human key point estimation and human semantic segmentation
[0100]
[0101] In the task of estimating the key points of the human body, the key points of the human body of the original data and the anonymized data are obtained through the key point prediction algorithm. KPd is the Euclidean distance of the gap between the key points of the human body, and OKS is the similarity of the key points of the human body calculated between the two. The smaller the KPd, the closer the OKS is to 1, indicating that the utility of anonymous data in estimating the key points of the human body is better. The results of the key point estimation of the human body in Table 6 show the advantage of the method of the present invention in ensuring the utility of this task. The test of the utility of anonymous data in the task of human semantic segmentation is calculated using three indicators: Pixel Accuracy (PA), Mean Pixel Accuracy (MPA), and Mean Intersection over Union (MIoU). The larger the three indicators, the higher the accuracy of semantic segmentation. In Table 6, the relevant indicators of human semantic segmentation show that MCFA has achieved good results, indicating that the method of the present invention has better guaranteed the utility of data in the two tasks of human semantic segmentation.
[0102] In general, the present invention proposes a multi-select full-body anonymization method for human image data. It introduces multiple ControlNets, prompt word engineering, face guidance and palette algorithms in the full-body anonymization based on the diffusion model, achieves more detailed anonymous attribute control while ensuring high anonymous data quality, and greatly improves the diversity of anonymous data. Compared with other methods, it provides different degrees of anonymization options at the semantic level.
[0103] In one embodiment, a multi-select full-body anonymization device for character image data is provided, comprising:
[0104] The first module is used to input character image data;
[0105] The second module is used to obtain a mask image of the human body part to be anonymized in the human image data;
[0106] The third module is used to obtain the text description of the mask image of the human body part to be anonymized by using the image-to-text model;
[0107] The fourth module is used to obtain anonymized prompt words by passing the text description through an anonymized prompt word project;
[0108] A fifth module is used to obtain a palette image representing the color feature distribution of the character image data using a palette algorithm;
[0109] The sixth module is used to obtain a false and random face identity feature based on the face database;
[0110] The seventh module is used to introduce the non-sensitive attribute information of the ControlNet controlled person image data as an image prompt to guide the diffusion model image generation process;
[0111] The eighth module is used to guide the Inpainting task based on the diffusion model according to the obtained mask image of the human body part to be anonymized, the anonymization prompt word, the palette image and the false and random face identity features, so as to generate a new character image to replace the human body part to be anonymized in the character image data, and obtain the anonymized result image.
[0112] The implementation method of each module in the above-mentioned device and the construction of the model can adopt the method described in any of the above-mentioned embodiments, which will not be repeated here.
[0113] On the other hand, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the multi-select full-body anonymization method for character image data provided in any of the above embodiments are implemented. The computer device may be a server. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store sample data. The network interface of the computer device is used to communicate with an external terminal via a network connection.
[0114] On the other hand, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-selection full-body anonymization method for character image data provided in any of the above embodiments.
[0115] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0116] Matters not covered by the present invention are known technologies.
[0117] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0118] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
[0119] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-select full-body anonymization method for person image data, characterized in that: include: Input character image data; Obtain a mask image of the human body part to be anonymized in the human image data; Use the image-to-text model to obtain the text description of the mask image of the human body part to be anonymized; The text description is passed through an anonymous prompt word project to obtain an anonymous prompt word; A palette image representing the color feature distribution of the person image data is obtained by using a palette algorithm; Obtain a false and random face identity feature based on the face database; Controlnet is introduced to control the non-sensitive attribute information of the character image data as image prompts to guide the diffusion model image generation process; According to the obtained mask image of the human body part to be anonymized, the anonymization prompt word, the palette image and the false and random face identity features, the Inpainting task based on the diffusion model is guided to generate a new character image to replace the human body part to be anonymized in the character image data, and the anonymized result image is obtained.
2. The multi-select full-body anonymization method for character image data according to claim 1, characterized in that: The method for obtaining the mask image of the human body part to be anonymized in the character image data is: For the input character image data , use the yolo algorithm to obtain character image data The classification of each pixel in the Person class is set to 255, and the pixel values of all other pixels are set to 0. In this way, the mask image of the human body part to be anonymized is obtained. .
3. The multi-select full-body anonymization method for character image data according to claim 1, characterized in that: The image-to-text model is used to obtain the text description of the human body part image to be anonymized, including: Character image data Mask image of the human body part to be anonymized Overlap and keep only the mask image The part with pixel value of 255 in the middle gets the overlapping picture : Use the graph-to-text model to obtain pictures The character text description is used as the text description of the human body part image to be anonymized, including a description of the characteristics of the clothes worn by the person.
4. The multi-select full-body anonymization method for character image data according to claim 3, characterized in that: The text description is passed through the anonymization prompt word project to obtain anonymized prompt words, including: Prompt word engineering is performed on the acquired text descriptions: first, data cleaning is performed on the text descriptions to remove irrelevant descriptions, and then clothing keywords are extracted. The acquired clothing keywords are randomly anonymized using the pre-built color text library and clothing detail text library to obtain anonymous prompt words with different anonymization degrees; Generate n different anonymous prompt words for the text description , use the clip encoder to combine n anonymous prompt words with the character image data They are all mapped to a unified latent space of text and image. By calculating the cosine similarity between the text and image vectors, the alignment relationship between the image and text is obtained. The anonymous prompt word with the largest difference from the image is selected from n different anonymous prompt words. As a final anonymization hint: in, For the i An anonymous reminder word, , For the clip encoder.
5. The multi-select full-body anonymization method for person image data according to claim 1, characterized in that: Get the color palette image that represents the color feature distribution of the person image data, as follows: For character image data Perform Gaussian blur processing; Use the K-means algorithm to represent the image after Gaussian blur processing with a certain number of colors, and then perform image enhancement operations to improve the color contrast of the processed image, thereby obtaining a palette image that represents the color feature distribution of the character image data .
6. The multi-select full-body anonymization method for person image data according to claim 1, characterized in that: Introduce Controlnet to control the non-sensitive attribute information of character image data, including: Get person image data The corresponding three types of non-sensitive partial attribute information are soft edge map, posture map, and facial information map. The soft edge map describes the edge contour of the human body in the image. The posture map is used to capture and represent the posture information of the human body in the image. The facial information map contains the key feature points and facial position information of the face.
7. The multi-select full-body anonymization method for person image data according to claim 6, characterized in that: Obtaining a false and random face identity feature based on a face database, including: randomly selecting n face images from the face database, extracting identity feature vectors corresponding to the n face images, generating random weights for the n face images, and performing weighted summation on the identity feature vectors corresponding to the n face images to obtain a false and random face identity feature. .
8. The multi-select full-body anonymization method for person image data according to claim 1, characterized in that: According to the obtained mask image of the human body part to be anonymized, the anonymization prompt word, the palette image and the false and random face identity features, the Inpainting task based on the diffusion model is guided, including: Use the mask image of the human body part to be anonymized The masking guidance of the Inpainting task of the diffusion model is carried out. According to the different choices of anonymization degree, the palette algorithm and the prompt word engineering are dynamically controlled. The palette algorithm is used to control the color distribution, and the prompt word engineering provides random colors and random details. When performing the Inpainting task of the diffusion model, according to the mask map of the human body part to be anonymized For the human body part in the image to be anonymized, the pure Gaussian noise is replaced by the noisy palette image. Reconstructed based on the anonymized prompt word Serves as a guide text to guide the anonymization generation process; As a separate facial information guide, the facial identity features A lightweight adaptive module with decoupled cross attention is used to support identity feature prompts. At the same time, the facial information map obtained by Annotators is used to control the position space of the face through ControlNet to guide the generation of facial parts.
9. A multi-select full-body anonymization device for person image data, characterized in that: include: The first module is used to input character image data; The second module is used to obtain a mask image of the human body part to be anonymized in the human image data; The third module is used to obtain the text description of the mask image of the human body part to be anonymized by using the image-to-text model; The fourth module is used to obtain anonymized prompt words by passing the text description through an anonymized prompt word project; A fifth module is used to obtain a palette image representing the color feature distribution of the character image data using a palette algorithm; The sixth module is used to obtain a false and random face identity feature based on the face database; The seventh module is used to introduce the non-sensitive attribute information of the ControlNet controlled person image data as an image prompt to guide the diffusion model image generation process; The eighth module is used to guide the Inpainting task based on the diffusion model according to the obtained mask image of the human body part to be anonymized, the anonymization prompt word, the palette image and the false and random face identity features, generate a new character image to replace the human body part to be anonymized in the character image data, and obtain the anonymized result image.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the multi-select full-body anonymization method for person image data as claimed in claim 1 are implemented.
Citation Information
Patent Citations
Key point differential privacy-driven face image privacy protection method
CN114169002A
Face anonymization method based on multi-condition diffusion model, storage medium and equipment
CN118658188A