Facial feature desensitization methods, devices, electronic equipment and storage media

By extracting features from the target facial image and generating a virtual face, the problem of coarse and inefficient desensitization in existing technologies is solved, achieving high-precision facial desensitization and meeting the needs of large-scale data processing.

CN119625808BActive Publication Date: 2025-11-14SUZHOU AUTOMOBILE RES INST OF TSINGHUA UNIV (WUJIANG) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411781899.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-11-14
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

Existing facial desensitization technologies result in rough facial surfaces and noticeable desensitization marks after desensitization, and are inefficient, making them unsuitable for large-scale data processing.

Method used

By identifying the target facial image, feature extraction is performed to obtain the target text features and draft image. A virtual face is generated using a virtual face generation model and replaces the target face in the original vehicle data. The generation process is optimized by combining the YOLOv5-face model and the StableDiffusion model with the ControlNet network.

Benefits of technology

It improves the accuracy of virtual face generation, reduces desensitization traces, minimizes the loss of original vehicle data, and meets the needs of large-scale face desensitization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625808B_ABST
    Figure CN119625808B_ABST
Patent Text Reader

Abstract

This invention discloses a facial feature desensitization method, apparatus, electronic device, and storage medium. The method includes: determining a target facial image; extracting features from the target facial image to obtain target text features and a target draft image; generating a virtual face based on the target text features and the target draft image; and replacing the target face in the original vehicle data with the virtual face. This method generates a virtual face based on the target text features and the target draft image, eliminating the need for manual editing. It meets the facial desensitization requirements of large-scale vehicle data collection while reducing desensitization traces and minimizing the loss of desensitization data in the original image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, electronic device, and storage medium for facial feature desensitization. Background Technology

[0002] In recent years, the intelligent development of automobiles has flourished. Among the key aspects, collecting massive amounts of data to train and optimize various AI models has become a core path for companies to continuously improve their intelligent capabilities. However, at the same time, the government's compliance requirements for data collection are becoming increasingly stringent. Facial data, as a matter of personal privacy, typically requires anonymization during the raw data (video or image) collection process to meet compliance requirements.

[0003] Currently, most facial desensitization techniques only employ methods such as facial bounding box detection, facial key point detection, and facial region image segmentation to directly erase or cover the identified face using color blocks, mosaics, or Gaussian blur. However, this approach often results in a rough appearance of the desensitized face, with obvious desensitization traces, thus reducing the image quality of the original data. Some facial desensitization techniques use virtual faces to replace real faces, but the generation process requires manual editing in an editing interface, and the model used needs multiple adjustments to generation parameters during facial recognition and virtual face generation to achieve satisfactory results. This leads to low processing efficiency and makes it unsuitable for large-scale facial desensitization. Summary of the Invention

[0004] This invention provides a facial feature desensitization method, device, electronic device, and storage medium to solve the problem of rough facial areas and obvious desensitization marks after desensitization.

[0005] According to one aspect of the present invention, a method for desensitizing facial features is provided, comprising:

[0006] The target facial image is determined; the target facial image is an image obtained by cropping the target face corresponding to the target control object in the original vehicle data; the target control object is the object that controls the vehicle.

[0007] Feature extraction is performed on the target facial image to obtain target text features and a target draft image; the target text features are used to characterize the attribute information of the target control object; the target draft image is used to characterize the facial contour of the target control object;

[0008] A virtual face is generated based on the target text features and the target draft image;

[0009] The virtual face replaces the target face in the original vehicle data.

[0010] According to another aspect of the present invention, a facial feature desensitization device is provided, comprising:

[0011] A facial image determination module is used to determine a target facial image; the target facial image is an image obtained by cropping the target face corresponding to the target control object in the original vehicle data; the target control object is the object that controls the vehicle;

[0012] The feature extraction module is used to extract features from the target facial image to obtain target text features and a target draft image; the target text features are used to characterize the attribute information of the target control object; the target draft image is used to characterize the facial contour of the target control object.

[0013] A virtual face generation module is used to generate a virtual face based on the target text features and the target draft image;

[0014] The replacement module is used to replace the target face in the original vehicle data using the virtual face.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the facial feature desensitization method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the facial feature desensitization method according to any embodiment of the present invention.

[0020] The technical solution of this invention involves determining a target facial image; extracting features from the target facial image to obtain target text features and a target draft image, which provides a basis for generating a virtual face and improves the accuracy of virtual face generation; generating a virtual face based on the target text features and the target draft image ensures that the generated virtual face meets the facial features of the target controlled object without exposing the target controlled object's biological information; replacing the target face in the original vehicle data with the virtual face reduces desensitization traces and minimizes the loss of original vehicle data. This method generates a virtual face based on target text features and the target draft image without manual editing, meeting the facial desensitization requirements of large-scale vehicle data collection while reducing desensitization traces and minimizing the loss of original image desensitization.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 A flowchart of a facial feature desensitization method provided in an embodiment of the present invention;

[0024] Figure 2 A flowchart of another facial feature desensitization method provided in an embodiment of the present invention;

[0025] Figure 3 A flowchart for generating a virtual face is provided for an embodiment of the present invention;

[0026] Figure 4 This is a schematic diagram of a facial feature desensitization device provided in an embodiment of the present invention;

[0027] Figure 5 A schematic diagram of the structure of an electronic device for implementing the facial feature desensitization method of this invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Figure 1 This is a flowchart illustrating a facial feature desensitization method provided in an embodiment of the present invention. This embodiment is applicable to desensitizing facial information contained in acquired vehicle images. The method can be executed by a facial feature desensitization device, which can be implemented in hardware and / or software. This facial feature desensitization device can be configured in any electronic device with network communication capabilities. Figure 1 As shown, the method includes:

[0031] S110. Determine the target facial image.

[0032] The target facial image is obtained by cropping the face of the target control object from the original vehicle data. The target control object is the object used to control the vehicle.

[0033] Specifically, raw vehicle data is acquired from an image acquisition device, facial feature detection is performed on the raw vehicle data using a facial detection model, and regions with facial features are cropped to obtain the target facial image.

[0034] Furthermore, if the original vehicle data is image data, facial feature detection is performed directly; if the original vehicle data is video data, the video data is first cropped according to the frame information, and then facial feature detection is performed on the cropped image data.

[0035] For example, the face detection model can use the YOLOv5-Face model. The image output by the YOLOv5-Face model contains 21 data dimensions. The data in dimensions 1 to 4 (denoted as Y1-Y4) represent the position and width and height information of the target detection box. Dimensions 5 and 6 represent the detection category and confidence score, respectively (Y5-Y6). The remaining 15 dimensions, from 7 to 21, give the position information of 5 facial key points (i.e., left eye, right eye, nose, right corner of mouth, and left corner of mouth) (Y7-Y21), which are denoted as K5.

[0036] Furthermore, for facial feature images with a confidence level Y6 greater than a set threshold, the regions containing the facial features are cropped according to the position and width / height information of the target detection boxes represented by Y1-Y4, to obtain the target facial image.

[0037] Furthermore, since an image data set containing facial features contains at least one facial feature, the obtained target facial image is also at least one.

[0038] S120. Extract features from the target facial image to obtain target text features and target draft image.

[0039] Among them, the target text features are used to represent the attribute information of the target controlled object. The target draft image is used to represent the facial contour of the target controlled object.

[0040] Specifically, detail features and edge features are extracted from the target facial image, and the target text features are determined based on the detail features; a target draft image is constructed based on the edge features.

[0041] S130. Generate a virtual face based on the target text features and the target draft image.

[0042] Specifically, a virtual face is generated by combining the attribute information of the target control object given by the target text features with the target sketch.

[0043] For example, assuming the target controlled object is wearing glasses, the target text features will contain glasses cue words. Based on the glasses cue words, glasses elements will be added to the target sketch, thereby generating a virtual face.

[0044] Furthermore, the generated virtual face can represent the attribute information of the target control object, but the biometric features that can be used to locate the target control object cannot be obtained from the virtual face.

[0045] The above steps, through the generation of virtual faces, can reduce the leakage of information of the target controlled object and ensure the privacy of the target controlled object.

[0046] S140. Replace the target face in the original vehicle data with a virtual face.

[0047] Specifically, the virtual face is used to replace the corresponding target face according to its position in the original vehicle data.

[0048] Furthermore, if the original vehicle data is image data, then a virtual face is directly used for overlay; if the original vehicle data is video data, then the target face in the corresponding image is first replaced by a virtual face. After the replacement is completed, the replaced image replaces the image of the corresponding frame in the original vehicle data according to the frame information.

[0049] For example, assuming the virtual face is generated from the target face in the 20th frame image, the target face in the 20th frame image is first replaced with the virtual face, and then the original 20th frame image in the video is replaced with the replaced 20th frame image.

[0050] For example, such as Figure 2 As shown, text feature extraction and draft generation are performed on the target facial image to obtain the attribute information and target draft image of the target control object. The attribute information and target draft image of the target control object are input into the prediction model to generate a virtual face. The virtual face replaces the corresponding target face in the original vehicle data to obtain the desensitized synthetic image.

[0051] The prediction model uses a joint model of the Stable Diffusion (SD) model and the ControlNet plugin.

[0052] The SD model serves as the base model, and its parameters are optimized in conjunction with the ControlNet network architecture. While keeping the SD model's backbone unchanged and its trained parameters locked, ControlNet adds only a small number of additional convolutional layers at specific locations to control the image generation results of the SD model, making it more accurate.

[0053] The above steps employ a combined SD model and ControlNet model because, in practical applications, directly using the SD model cannot guarantee that the generated results will match the real-world requirements in terms of content and style. Therefore, to generate more detailed virtual faces, the ControlNet network architecture is used to optimize the parameters of the SD model.

[0054] Optionally, the target facial image is determined, including steps A1-A5:

[0055] Step A1: Determine the original vehicle data.

[0056] The raw vehicle data includes facial information of the target controlled object.

[0057] Specifically, raw vehicle data is acquired from the image acquisition device.

[0058] Furthermore, the original vehicle data includes at least one target control object and the target face corresponding to the target control object.

[0059] Step A2: Perform facial feature detection on the original vehicle data to obtain the location information of the facial features.

[0060] Specifically, facial feature detection is performed on the original vehicle data using a facial detection model, and the location information of the facial features is determined.

[0061] Step A3: Determine the pixel information within the preset bounding box based on the location information of facial features.

[0062] Specifically, a preset bounding box is generated based on the location information of facial features, and the pixel information within the preset bounding box is determined.

[0063] For example, assuming the location information of the facial features is (5,5), (5,7), (6,5), (6,7), then the preset bounding box can be (4,4), (4,8), (7,4), (7,8).

[0064] The above steps involve selecting a preset bounding box to prevent the acquired location information from not including all facial features.

[0065] Step A4: Determine the size and position coordinates of the target detection box based on the pixel information.

[0066] The position coordinates can be the coordinates of the center point of the target face, or the coordinates of the four corners of a rectangle that can cover all the features of the target face.

[0067] Specifically, the size and position coordinates of the target detection box are determined based on the color changes reflected by the pixel information within the preset bounding box.

[0068] Step A5: Crop the original vehicle data according to the size and position coordinates of the target detection box to obtain the target face image.

[0069] Specifically, the original vehicle data is located based on the position coordinates of the target detection box, and the original vehicle data is cropped according to the size of the target detection box to obtain the target facial image.

[0070] For example, such as Figure 2 As shown, target face detection is performed on the original image, i.e., the original vehicle data, to obtain target detection boxes. Based on the target detection boxes, the original vehicle data is cropped to obtain n target face images.

[0071] Optionally, feature extraction is performed on the target facial image to obtain target text features and a target draft image, including steps B1-B3:

[0072] Step B1: Extract facial features from the target facial image to obtain facial feature information.

[0073] Among them, facial feature information is used to characterize the facial contour information and facial structure information of the target control object.

[0074] Among them, facial contour information is used to characterize the face shape and approximate location of the facial features of the target control object.

[0075] Among them, facial structure information is used to characterize the shape and size of facial features, as well as the degree of prominence and depression of facial bones.

[0076] Specifically, facial features are extracted from the target facial image to obtain the facial contour and facial structure information of the target control object.

[0077] Furthermore, facial contour information is extracted using the canny function configured in the OpenCV image processing library.

[0078] Step B2: Generate a target draft image based on facial feature information.

[0079] Specifically, a target draft image is generated based on facial contour information and the position and size of facial features.

[0080] Step B3: Perform attribute feature analysis on facial feature information to obtain target text features.

[0081] Among them, attribute features are used to characterize the type and aging degree of the target controlled object.

[0082] Specifically, attribute feature analysis is performed based on facial structure and facial contour information, and the analysis results are processed into text to obtain target text features.

[0083] For example, the type of target subject can be determined based on facial structure information, such as the height of the nose, the shape and size of the eyes. The degree of aging of the target subject can be determined based on facial contour information, changes in facial bones, or facial size.

[0084] Furthermore, the target text features are obtained using the DeepFace framework. The DeepFace framework obtains the target text features by performing attribute analysis on facial feature information.

[0085] Optionally, a target draft image is generated based on facial feature information, including steps C1-C3:

[0086] Step C1: Determine facial key point information and facial edge lines based on facial feature information.

[0087] Among them, facial key point information is used to represent the location information and contour of the facial features.

[0088] Furthermore, when determining the target facial image, facial key point information can be obtained, as described above with K5. However, in certain special circumstances, such as rain or fog, the obtained target facial image may not be clear, and the key point information obtained when determining the target facial image may not meet the requirements. Therefore, it is necessary to supplement the information with more accurate and denser facial key point information.

[0089] Furthermore, supplementary methods can be achieved by using the Dlib tool library, which provides 68 facial key points (denoted as K68), or the Openpose tool library, which provides 70 facial key points (denoted as K70).

[0090] Specifically, facial key point information is determined based on facial feature information using a facial key point determination method; and facial edge lines are extracted based on facial feature information using an edge extraction method.

[0091] Among them, the edge extraction method can adopt the Canny edge detection algorithm.

[0092] Step C2: Generate the edge contour of the target face based on the facial edge lines.

[0093] Specifically, the broken lines of the facial edges are filled and the contours are regularized to obtain the edge contour of the target face.

[0094] Step C3: Generate a target draft based on the facial key point information and the edge contour of the target face.

[0095] Specifically, the target draft image is obtained by synthesizing the facial features and their contours with the edge contours of the target face according to their corresponding positions.

[0096] Optionally, attribute feature analysis is performed on the facial feature information to obtain the target text features, including steps D1-D4:

[0097] Step D1: Perform 3D transformation on the target facial image based on facial feature information to obtain a 3D facial image.

[0098] Specifically, the DeepFace framework determines facial contour information and facial key point information based on facial feature information, aligns the target facial image and annotates anchor points based on the facial contour information and facial key point information, and performs three-dimensional transformation on the target facial image based on the annotation results to obtain a three-dimensional facial image.

[0099] Step D2: Add the first feature anchor point to the 3D facial image to obtain the first anchor point image.

[0100] The first feature anchor point is used to characterize the structural changes of the target face.

[0101] Specifically, a first feature anchor point is added based on the distribution of facial structures in the 3D facial image to obtain a first anchor point image.

[0102] Step D3: Determine the attribute information of the target control object corresponding to the target face based on the first anchor point image.

[0103] Specifically, feature detection is performed on the first anchor point image, and the attribute information of the target control object corresponding to the target face is determined based on the detection results.

[0104] Step D4: Generate target text features based on the attribute information of the target control object.

[0105] Specifically, the obtained attribute information of the target controlled object is processed through text sorting and lexical recombination to obtain target text features.

[0106] For example, assuming the attribute information of the target controlled object is moderate aging and wearing glasses, the generated target text feature is a middle-aged object wearing glasses.

[0107] Optionally, the target facial image is transformed into a three-dimensional image based on facial feature information to obtain a three-dimensional facial image, including steps E1-E3:

[0108] Step E1: Rotate, translate, and scale the target facial image based on facial feature information to obtain a two-dimensional corrected image.

[0109] Among them, the two-dimensional corrected image is the image after positive correction of the target face in the target face image.

[0110] "Positive" here refers to being able to see the entire target's face.

[0111] Specifically, based on facial structure information, namely the position of the facial features, the target facial image is rotated, translated, and scaled, and the processed image is used as a two-dimensional corrected image.

[0112] The above steps involve processing the target facial image because the acquired target face may not be in a forward orientation due to the actions of the target control object. Therefore, forward orientation correction is required to generate a more accurate virtual face.

[0113] Step E2: Add a second feature anchor point to the two-dimensional corrected image to obtain the second anchor point image.

[0114] The second feature anchor point is used to characterize the detailed location of the facial features within the target face.

[0115] Furthermore, the first feature anchor point and the second feature anchor point can represent the same facial structure or different facial structures.

[0116] Specifically, after acquiring the two-dimensional corrected image, a second feature anchor point is added to the two-dimensional corrected image to obtain the second anchor point image.

[0117] Step E3: Generate a three-dimensional facial model based on the second anchor point image, and fuse the three-dimensional facial model with the target facial image to obtain a three-dimensional facial image.

[0118] Specifically, a three-dimensional facial model is generated based on the changes in the facial structure represented in the second anchor point image. The target facial image is then fitted into the three-dimensional facial model according to the position of the facial features to obtain a three-dimensional facial image.

[0119] Optionally, a virtual face is generated based on the target text features and the target draft image, including steps F1-F4:

[0120] Step F1: Encode the target text features and the target facial image to obtain the first encoding vector.

[0121] Specifically, textual semantic information is determined based on the target text features; image semantic information is determined based on the target facial image; the textual semantic information and image semantic information are then fused to obtain fused semantic information. The fused semantic information is then encoded to obtain the first encoding vector.

[0122] Image semantic information is used to describe the content contained in the target facial image through text.

[0123] Furthermore, the first encoding vector can also be determined by determining the textual semantic information of the target text features; determining the facial feature information such as texture features and color features of the target facial image; and fusing the facial feature information with the semantic text information to obtain the first encoding vector.

[0124] Step F2: Encode the target draft image to obtain the second encoding vector.

[0125] The second encoding vector has the same size as the first encoding vector, but is generated using a different encoder.

[0126] Specifically, the image information contained in the target draft image is encoded to obtain a two-dimensional encoded vector.

[0127] Step F3: Input the first and second encoding vectors into the preset model to obtain the third encoding vector.

[0128] The third encoding vector is used to characterize the features of the virtual face.

[0129] The preset model is a pre-established model used to generate a virtual face based on the first encoding vector and the second encoding vector.

[0130] Specifically, the first and second encoding vectors are input into a preset model, and the prediction model analyzes the target facial features represented by the first and second encoding vectors to obtain the third encoding vector.

[0131] Step F4: Decode the third encoding vector to obtain the virtual face.

[0132] Specifically, the virtual face is obtained by decoding the third encoded vector using the decoder.

[0133] For example, such as Figure 3 As shown, the attribute information of the target control object and a reference image or random noise image are passed through a first encoder to obtain a 64x64 fixed-size SD hidden code. The reference image is a target face image obtained by directly cropping the original vehicle data; the random noise image is a target face image with added random noise. The target draft image is encoded by a second encoder to obtain a feature encoding vector of the same size as the SD code. Using a ControlNet structure, the backbone network module of the SD is copied as a trainable copy, and connected before and after it by 1x1 convolutional layers with weights and biases initialized to zero.

[0134] Furthermore, in the image generation process (reverse diffusion process), a 64x64 hidden code is first obtained through an encoder, and then the sampling process is repeatedly performed, i.e., Z... T Z T Z T-1 Z T-1 …, Z0Z0 (sampled T times), the generation quality usually improves gradually, and the number of samples can be adjusted appropriately according to specific time requirements. Finally, after sampling Z0Z0, it goes through a decoder to obtain the virtual face.

[0135] The technical solution of this embodiment determines the target facial image; extracts features from the target facial image to obtain target text features and a target draft image, which can provide a basis for the generation of virtual faces and improve the accuracy of virtual face generation; generating virtual faces based on target text features and target draft images ensures that the generated virtual faces meet the facial features of the target controlled object while not exposing the target controlled object's biological information; replacing the target face in the original vehicle data with virtual faces can reduce desensitization traces and reduce the loss of original vehicle data. This method generates virtual faces based on target text features and target draft images without manual editing, meeting the facial desensitization requirements of large-scale vehicle data collection while reducing desensitization traces and minimizing the loss of original image desensitization.

[0136] Figure 4 This is a schematic diagram of a facial feature desensitization device provided in an embodiment of the present invention. This embodiment is applicable to situations where facial information contained in acquired vehicle images is desensitized. The facial feature desensitization device can be implemented in hardware and / or software, and can be configured in any electronic device with network communication capabilities. Figure 4 As shown, the device includes: a facial image determination module 210, a feature extraction module 220, a virtual face generation module 230, and a replacement module 240, wherein:

[0137] Facial image determination module 210: used to determine the target facial image; the target facial image is an image obtained by cropping the target face corresponding to the target control object in the original vehicle data; the target control object is the object that controls the vehicle;

[0138] Feature extraction module 220: used to extract features from the target facial image to obtain target text features and target draft image; target text features are used to characterize the attribute information of the target control object; target draft image is used to characterize the facial contour of the target control object;

[0139] Virtual face generation module 230: used to generate a virtual face based on the target text features and the target draft image;

[0140] Replacement module 240: Used to replace the target face in the original vehicle data with a virtual face.

[0141] Optionally, the facial image determination module 210 includes:

[0142] Original vehicle data determination unit: used to determine the original vehicle data; the original vehicle data contains facial information of the target control object;

[0143] Location information determination unit: used to perform facial feature detection on the original vehicle data to obtain the location information of the facial features;

[0144] Pixel information determination unit: used to determine pixel information within a preset bounding box based on the positional information of facial features;

[0145] Target detection bounding box determination unit: used to determine the size and position coordinates of the target detection bounding box based on pixel information;

[0146] Facial image determination unit: used to crop the original vehicle data according to the size and position coordinates of the target detection box to obtain the target facial image.

[0147] Optionally, the feature extraction module 220 includes:

[0148] Facial Feature Information Determination Unit: Used to extract facial features from the target facial image to obtain facial feature information; facial feature information is used to characterize the facial contour information and facial structure information of the target control object;

[0149] Target draft image generation unit: used to generate target draft images based on facial feature information;

[0150] Target text feature determination unit: used to perform attribute feature analysis on facial feature information to obtain target text features;

[0151] Optional, the target draft diagram generation unit includes:

[0152] Edge line determination subunit: used to determine facial key point information and facial edge lines based on facial feature information; facial key point information is used to represent the location information and contour of the facial features.

[0153] Edge contour determination subunit: used to generate the edge contour of the target face based on the facial edge lines;

[0154] Target draft image determination sub-unit: used to generate a target draft image based on facial key point information and the edge contour of the target face.

[0155] Optionally, the target text feature determination unit includes:

[0156] Three-dimensional facial image determination subunit: used to perform three-dimensional transformation on the target facial image based on facial feature information to obtain a three-dimensional facial image;

[0157] First Feature Anchor Point Determination Subunit: Used to add first feature anchor points to a 3D facial image to obtain a first anchor point image; the first feature anchor points are used to characterize the structural changes of the target face.

[0158] Attribute information determination subunit: used to determine the attribute information of the target control object corresponding to the target face based on the first anchor point image;

[0159] Target text feature determination subunit: used to generate target text features based on the attribute information of the target controlled object.

[0160] Optionally, a 3D facial image determination subunit is used specifically for:

[0161] The target face image is rotated, translated, and scaled based on facial feature information to obtain a two-dimensional corrected image; the two-dimensional corrected image is the image after positive correction of the target face in the target face image.

[0162] A second feature anchor point is added to the two-dimensional corrected image to obtain a second anchor point image; the second feature anchor point is used to characterize the detailed position of the facial features within the target face.

[0163] A 3D facial model is generated based on the second anchor point image, and the 3D facial model and the target facial image are fused to obtain a 3D facial image.

[0164] Optional, the virtual face generation module includes:

[0165] First encoding vector determination unit: used to encode the target text features and the target facial image to obtain the first encoding vector;

[0166] The second encoding vector determination unit is used to encode the target draft image to obtain a second encoding vector; the second encoding vector has the same size as the first encoding vector and is generated using a different encoder.

[0167] The third encoding vector determination unit is used to input the first encoding vector and the second encoding vector into a preset model to obtain the third encoding vector; the third encoding vector is used to characterize the features of the virtual face; the preset model is a pre-established model used to generate the virtual face based on the first encoding vector and the second encoding vector.

[0168] Virtual face generation unit: used to decode the third encoding vector to obtain a virtual face.

[0169] The facial feature desensitization device provided in the embodiments of the present invention can perform the facial feature desensitization method provided in any of the embodiments of the present invention, and has the corresponding functions and beneficial effects of performing the facial feature desensitization method. For details, please refer to the relevant operations of the facial feature desensitization method in the foregoing embodiments.

[0170] Figure 5This is a schematic diagram of the structure of an electronic device for implementing the facial feature desensitization method of this invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0171] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0172] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0173] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as facial feature desensitization methods.

[0174] In some embodiments, the facial feature desensitization method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the facial feature desensitization method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the facial feature desensitization method by any other suitable means (e.g., by means of firmware).

[0175] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0176] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0177] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0178] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0179] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0180] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0181] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0182] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for desensitizing facial features, characterized in that, include: The target facial image is determined; the target facial image is an image obtained by cropping the target face corresponding to the target control object in the original vehicle data; the target control object is the object that controls the vehicle. Feature extraction is performed on the target facial image to obtain target text features and a target draft image; the target text features are used to characterize the attribute information of the target control object; the target draft image is used to characterize the facial contour of the target control object; A virtual face is generated based on the target text features and the target draft image; The target face in the original vehicle data is replaced by the virtual face; The step of extracting features from the target facial image to obtain target text features and target draft image includes: Facial features are extracted from the target facial image to obtain facial feature information; the facial feature information is used to characterize the facial contour information and facial structure information of the target control object. Generate a target draft image based on the facial feature information; Attribute feature analysis is performed on the facial feature information to obtain the target text features; The step of performing attribute feature analysis on the facial feature information to obtain target text features includes: The target facial image is transformed into a three-dimensional image based on the facial feature information. A first feature anchor point is added to the three-dimensional facial image to obtain a first anchor point image; the first feature anchor point is used to characterize the structural changes of the target face. Determine the attribute information of the target control object corresponding to the target face based on the first anchor point image; Generate target text features based on the attribute information of the target control object; The step of performing three-dimensional transformation on the target facial image based on the facial feature information to obtain a three-dimensional facial image includes: The target facial image is rotated, translated, and scaled based on facial feature information to obtain a two-dimensional corrected image; the two-dimensional corrected image is the image after positive correction of the target face in the target facial image. A second feature anchor point is added to the two-dimensional corrected image to obtain a second anchor point image; the second feature anchor point is used to characterize the detailed position of the facial features within the target face. A three-dimensional facial model is generated based on the second anchor point image, and the three-dimensional facial model and the target facial image are fused to obtain a three-dimensional facial image.

2. The method according to claim 1, characterized in that, The determination of the target facial image includes: Determine the original vehicle data; the original vehicle data includes facial information of the target controlled object; Facial feature detection is performed on the original vehicle data to obtain the location information of the facial features; The pixel information within the preset bounding box is determined based on the location information of the facial features; The size and position coordinates of the target detection box are determined based on the pixel information; The original vehicle data is cropped based on the size and position coordinates of the target detection box to obtain the target facial image.

3. The method according to claim 1, characterized in that, The step of generating a target draft image based on the facial feature information includes: Facial key point information and facial edge lines are determined based on facial feature information; the facial key point information is used to represent the location information and outline of the facial features. Generate the edge contour of the target face based on the facial edge lines; A target draft image is generated based on the facial key point information and the edge contour of the target face.

4. The method according to claim 1, characterized in that, The step of generating a virtual face based on the target text features and the target draft image includes: The target text features and the target facial image are encoded to obtain a first encoding vector; The target draft image is encoded to obtain a second encoded vector; the second encoded vector has the same size as the first encoded vector, but is generated using a different encoder; The first encoding vector and the second encoding vector are input into a preset model to obtain a third encoding vector; the third encoding vector is used to characterize the features of the virtual face; the preset model is a pre-established model for generating a virtual face based on the first encoding vector and the second encoding vector. Decoding the third encoded vector yields the virtual face.

5. A facial feature desensitization device, characterized in that, include: A facial image determination module is used to determine a target facial image; the target facial image is an image obtained by cropping the target face corresponding to the target control object in the original vehicle data; the target control object is the object that controls the vehicle; The feature extraction module is used to extract features from the target facial image to obtain target text features and a target draft image; the target text features are used to characterize the attribute information of the target control object; the target draft image is used to characterize the facial contour of the target control object. A virtual face generation module is used to generate a virtual face based on the target text features and the target draft image; The replacement module is used to replace the target face in the original vehicle data using the virtual face; The feature extraction module includes: Facial feature information determination unit: used to extract facial features from the target facial image to obtain facial feature information; the facial feature information is used to characterize the facial contour information and facial structure information of the target control object; Target draft image generation unit: used to generate a target draft image based on the facial feature information; Target text feature determination unit: used to perform attribute feature analysis on the facial feature information to obtain target text features; The target text feature determination unit includes: Three-dimensional facial image determination subunit: used to perform three-dimensional transformation on the target facial image based on the facial feature information to obtain a three-dimensional facial image; First Feature Anchor Point Determination Subunit: Used to add a first feature anchor point to the three-dimensional facial image to obtain a first anchor point image; the first feature anchor point is used to characterize the structural changes of the target face. Attribute information determination subunit: used to determine the attribute information of the target control object corresponding to the target face based on the first anchor point image; Target text feature determination subunit: used to generate target text features based on the attribute information of the target control object; The three-dimensional facial image determination subunit is specifically used for: The target facial image is rotated, translated, and scaled based on facial feature information to obtain a two-dimensional corrected image; the two-dimensional corrected image is the image after positive correction of the target face in the target facial image. A second feature anchor point is added to the two-dimensional corrected image to obtain a second anchor point image; the second feature anchor point is used to characterize the detailed position of the facial features within the target face. A three-dimensional facial model is generated based on the second anchor point image, and the three-dimensional facial model and the target facial image are fused to obtain a three-dimensional facial image.

6. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the facial feature desensitization method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the facial feature desensitization method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Image desensitization method and device, electronic equipment and storage medium

    CN116977484A

  • Virtual image display method and device, electronic equipment, medium and chip

    CN117876635A