A Method for Generating Heterogeneous Facial Mask Images Based on Progressive Adversarial Generation Architecture
Through the progressive adversarial generation architecture, combining the key points of heterogeneous face images and the pre-trained generation module, a mask image with consistent style is generated, which solves the problem of modal inconsistency in heterogeneous face images synthesis and improves image quality and the performance of the recognition system.
Patent Information
- Application Number
- CN202310397214.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-04-13
AI Technical Summary
In the prior art, when generating heterogeneous face images, especially near-infrared or sketched images, the inconsistency between mask modes and face modes leads to poor quality of synthesis results, lack of authenticity and visual perception, and lack of image synthesis requirements in heterogeneous scenes.
Using a progressive confrontation generation architecture, by obtaining the key points of heterogeneous face images, selecting the target mask template and deformation and synthesis, combining the pre-trained generation module to generate mask images with the same style, and finally fusing with the heterogeneous face images to crop the mask area to ensure the consistent image style.
The generation quality of heterogeneous face mask images is improved, so that the mask area is consistent with the image style of the background part is consistent, the authenticity and quality of the generated images are enhanced, and the performance of the heterogeneous face recognition system is improved.
Smart Images

Figure CN116631024B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for generating heterogeneous face mask images based on a progressive adversarial generation architecture. Background Art
[0002] The purpose of heterogeneous face recognition is to recognize and match face images under cross-modal conditions, and it plays an important role in public security fields such as security monitoring. With the continuous promotion of in-depth learning research, many effective heterogeneous face recognition methods have been proposed. The partial loss of facial features caused by the normalization of wearing masks has brought a huge impact to traditional face recognition and heterogeneous face recognition systems. For example, in the near-infrared face recognition scenario, people may wear masks and pass through the access control system or near-infrared video surveillance based on near-infrared cameras at night, and the captured near-infrared masked faces must be compared and recognized with the color ID photos in the pre-registered database. In the sketch face recognition scenario, a suspect at the crime scene may wear a mask to hide his face, and at this time, it is necessary to draw a portrait of the masked face according to the description of the witness to identify the suspect. Therefore, it is crucial to promote the research on heterogeneous masked face recognition.
[0003] At present, most existing heterogeneous face recognition methods are designed for traditional recognition scenarios without masks, which will encounter some problems in heterogeneous face recognition scenarios with masks. First, mask occlusion widens the domain gap between heterogeneous face images, making it challenging to effectively learn face representations using traditional heterogeneous face recognition. In addition, due to the lack of available heterogeneous face data with masks, recognition models based on deep neural networks cannot work well. In response to these problems, there are some mask-wearing face synthesis methods in the related art, but the mask-wearing face synthesis methods in the related art usually pay more attention to the position issues such as mask size and angle. For example, Wang et al. proposed MaskTool with two-dimensional image processing technology, which includes operations such as landmark detection, affine transformation, blending, brightness adjustment, and boundary blurring. Similarly, MaskTheFace proposed by Anwar et al. also relies on 2D image technology, which adjusts the mask to directly cover the face after calculating facial key points and tilt angles. Wang et al.'s 3D-based Face Mask Adding (FMA-3D) maps masks and faces into UV space through 3D face reconstruction technology, and then mixes and renders them into a face with a mask. Du et al. also use 3D face reconstruction technology to first add a mask template to the UV texture map of the face without a mask, and then restore the 2D face with a mask from the UV texture map of the face with a mask. The above methods have achieved satisfactory synthesis effects on ordinary face images, but they ignore the image synthesis requirements in heterogeneous scenes such as near-infrared and sketches. For heterogeneous images, due to the inconsistency of the modalities of the color mask template and the heterogeneous face, using existing methods to directly synthesize the two may result in poor quality and lack of authenticity in the generated image.
[0004] That is to say, the methods for generating face mask images in related technologies are all studies on face mask synthesis for visible light images, and only consider positional issues such as mask size and face angle, without paying attention to the overall color of the synthesized image. Although the synthesis effect is excellent when dealing with color photos with the same mask modality, once faced with images such as near-infrared images or sketches, the color difference caused by modality inconsistency (inconsistent image style) will cause the synthesis result to be seriously unrealistic and the visual perception to be poor. In addition, the image quality will also be affected by the existing synthesis method and degraded. Summary of the invention
[0005] In order to solve the above problems existing in the related art, the present invention provides a method for generating heterogeneous face mask images based on a progressive adversarial generation architecture. The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0006] The present invention provides a method for generating heterogeneous face mask images based on a progressive adversarial generation architecture, comprising:
[0007] Obtain heterogeneous face images;
[0008] Extract the face key points of the heterogeneous face image, select a target mask template from multiple preset mask templates according to the face key points, and determine a set of target key points according to the face key points; each preset mask template corresponds to a set of preset key points;
[0009] Determine a mask image according to a set of preset key points corresponding to the target mask template and the set of target key points;
[0010] Synthesize the mask image and the heterogeneous face image to obtain an initial heterogeneous face mask image, and generate a mask image of the initial heterogeneous face image;
[0011] Input the feature vector of the initial heterogeneous face image into a pre-trained generation module to obtain a generated image; the image styles of the mask part and the background part in the generated image are consistent; wherein, the pre-trained generation module includes n pre-trained generation networks, the output of the i-th generation network is used as the input of the (i + 1)-th generation network, the first n - 1 generation networks are all used to generate feature maps according to the input images and perform upsampling processing, and the n-th generation network is used to obtain the generated image according to the input feature map; the n pre-trained generation networks are obtained by training a progressive adversarial generation network; the progressive adversarial generation network includes n generation networks and a discriminator network; i is an integer from 1 to n - 1; n is an integer greater than or equal to 2;
[0012] Crop out the mask area from the generated image according to the mask image;
[0013] Fuse the mask area and the heterogeneous face image to obtain a target heterogeneous face mask image.
[0014] The present invention has the following beneficial technical effects:
[0015] Select a mask template based on the facial key points on the obtained heterogeneous face image. Based on the key points corresponding to the selected mask template and a set of target key points obtained from the facial key points, obtain a mask image. Synthesize an initial heterogeneous face mask image based on the mask image and the heterogeneous face image, which can determine the position of the mask to be synthesized on the heterogeneous face image in combination with the heterogeneous face image, making the position of the mask on the face image more accurate. Moreover, through a pre-trained generation module, generate a generated image with the same image style for the mask part and the background part based on the initial heterogeneous face mask image, crop the mask area from the generated image, and fuse the mask area with the heterogeneous face image to obtain a target heterogeneous face mask image. In this way, the style (modality) of the mask area used for fusing with the heterogeneous face image is consistent with the heterogeneous face image, thereby improving the quality of the finally generated heterogeneous face mask image.
[0016] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0017] Figure 1 It is a flowchart of a method for generating a heterogeneous face mask image based on a progressive adversarial generation architecture provided by an embodiment of the present invention;
[0018] Figure 2 It is a flowchart of an exemplary method for generating a heterogeneous face mask image based on a progressive adversarial generation architecture provided by an embodiment of the present invention;
[0019] Figure 3 It is a comparison result of images generated by exemplary different face mask synthesis methods provided by an embodiment of the present invention. Detailed Embodiments
[0020] The present invention will be further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0021] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0022] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0023] Although the present invention has been described in connection with various embodiments herein, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0024] Figure 1 is a flowchart of a method for generating heterogeneous face mask images based on a progressive adversarial generation architecture provided by an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:
[0025] S101. Obtain a heterogeneous face image.
[0026] In the embodiment of the present invention, a heterogeneous face image can be obtained by photographing with devices such as an infrared camera, or can also be obtained by receiving an image from an external device. The embodiment of the present invention does not make any limitations in this regard.
[0027] In the embodiment of the present invention, the heterogeneous face image can be an image taken under non-visible light. For example, it can be an image taken with an infrared camera, or a sketch image, etc. The present invention will explain the provided method through the processing process of a single heterogeneous face image.
[0028] S102. Extract the face key points of the heterogeneous face image, select a target mask template from multiple preset mask templates according to the face key points, and determine a set of target key points according to the face key points; each preset mask template corresponds to a set of preset key points.
[0029] In the embodiments of the present invention, a key point detector (for example, a face key point detector based on dlib) can be used to detect the face of a heterogeneous face image, and key points of the eyebrows, eyes, nose, mouth, and chin contour (the coordinates of each key point on the heterogeneous face image) can be obtained respectively. Exemplarily, 68 facial key points can be obtained in total. For example, 5 key points correspond to the left eyebrow, 6 key points represent the left eye, and so on.
[0030] In some embodiments, the face deflection angle can be determined according to the face key points, the face angle state can be determined according to the face deflection angle and a preset angle value, and a target mask template can be selected from multiple preset mask templates according to the face angle state.
[0031] Exemplarily, the mask template can be a colored mask template image, and can include a left-face mask template, a right-face mask template, and a front-face mask template; moreover, the face angle state corresponding to the left-face mask template is the left face, the face angle state corresponding to the right-face mask template is the right face, and the face angle state corresponding to the front-face mask template is the front face. When the determined face angle state is the left face, the target mask template is the left-face mask template, and when the determined face angle state is the right face, the target mask template is the right-face mask template.
[0032] Here, the face deflection angle can be calculated according to the obtained face key points. When the face deflection angle is greater than the preset angle value th, the face angle state can be determined to be the left face. When the face deflection angle is less than -th, the face angle state can be determined to be the right face; when the face deflection angle is in other situations, the face angle state can be determined to be the front face.
[0033] Here, the method for calculating the face deflection angle according to the obtained face key points is as follows:
[0034] The coordinates of the left and right eyes are averaged respectively, denoted as the midpoints of the left and right eyes; then the midpoint values of the left and right eyes are averaged, denoted as the eye midpoint;
[0035] 1. The midpoint of the lower lip = the coordinates of the lower lip are averaged to be used as the "right point" and the "midpoint"; the first point in the nasal bridge point set is taken as the "left point"; a facial symmetry center line (using existing methods such as np.polyfit, fitting a function through two points, and then taking 50 equally spaced values on this function to form a line) is determined through the "left point" and the "right point", denoted as perp_line (perpendicular line);
[0036] 2. left_point = the first point in the nasal bridge point set; right_point = the last point in the nasal bridge point set; a vertical line is determined through the left and right points, denoted as the nose center line nose_mid_line;
[0037] 3. Calculate the radian values a of the perp and nose lines respectively through tan a = Δy / Δx, and then convert them into angular values. The difference between the two angular values is denoted as the face deflection angle.
[0038] Here, the preset angular value th can be set according to actual needs.
[0039] In the embodiment of the present invention, after extracting the face key points, six coordinate points including the middle of the nose bridge, the left and right cheekbones, the left and right mandibles, and the center of the chin can also be obtained through the calculation of the extracted face key points. These six coordinate points are used as a set of target key points.
[0040] In the embodiment of the present invention, each mask template also corresponds to a set of preset key points, and this set of preset key points is also the six coordinate points of the preset middle of the nose bridge, the left and right cheekbones, the left and right mandibles, and the center of the chin.
[0041] S103. Determine the mask image according to a set of preset key points and a set of target key points corresponding to the target mask template.
[0042] In the embodiment of the present invention, a transformation matrix and a set of binary masks can be obtained by calculating according to a set of preset key points and a set of target key points corresponding to the target mask template. Then, the target mask template is deformed according to the transformation matrix to obtain a deformed image, and the deformed image is the determined mask image.
[0043] Here, when deforming the target mask template according to the transformation matrix, the cv2.warpPerspective function can be used to implement perspective transformation.
[0044] Here, through the calculation of the face deflection angle and the calculation of the target key points, the mask position can be made more accurate.
[0045] Here, the principle of obtaining a transformation matrix and a set of binary masks according to a set of preset key points and a set of target key points corresponding to the target mask template is: using the existing cv2.findHomography function to obtain the transformation matrix and the binary masks.
[0046] S104. Synthesize the mask image with the heterogeneous face image to obtain an initial heterogeneous face mask image, and generate a mask image of the initial heterogeneous face image.
[0047] S105. Input the feature vector of the initial heterogeneous face image into the pre-trained generation module to obtain a generated image; the image styles of the mask part and the background part in the generated image are consistent; wherein, the pre-trained generation module includes n pre-trained generation networks, the output of the i-th generation network is used as the input of the (i + 1)-th generation network, the first n - 1 generation networks are all used to generate a feature map according to the input image and perform upsampling processing, and the n-th generation network is used to obtain the generated image according to the input feature map; the n pre-trained generation networks are obtained by training a progressive adversarial generation network; the progressive adversarial generation network includes n generation networks and a discriminator network; i is an integer from 1 to n - 1; n is an integer greater than or equal to 2.
[0048] In the embodiment of the present invention, n can be set according to actual needs. For example, it can be 3.
[0049] In the embodiment of the present invention, each pre-trained generation network can adopt a fully convolutional neural network. In this way, it can accept images of any size as input, and while paying more attention to image details, it can avoid losing relevant features.
[0050] Here, the convolution kernel of the convolutional layer constituting the generation network is 3, there are 64 filters in each convolutional layer, and except for the last convolutional layer (behind which there is only a Tanh function), there is a batch normalization (BN) and a Leaky-ReLU activation function behind each convolutional layer.
[0051] Here, the convolution kernel of the convolutional layer constituting the discriminator network is also 3, there are also 64 filters in each convolutional layer, and except for the last convolutional layer (behind which there is only a Tanh function), there is also a Leaky-ReLU activation function behind each convolutional layer.
[0052] Exemplarily, the generation network can be a generator or other networks with generator functions.
[0053] Exemplarily, the discriminator network can be a Markov discriminator. In this way, only one discriminator is sufficient to judge the modal consistency of images at different scales, and it pays more attention to image texture.
[0054] In some embodiments, before S105, it may further include: S201 - S206:
[0055] S201. Obtain the heterogeneous face image to be trained.
[0056] Here, one or more heterogeneous face images to be trained can be obtained. The following will describe a single heterogeneous face image to be trained.
[0057] S202. Downsample the heterogeneous face images to be trained according to n preset sizes respectively, and correspondingly obtain n downsampled images; the sizes of the n downsampled images gradually increase.
[0058] Exemplarily, when n is 3, the sizes of the first to the third downsampled images gradually increase, and the size of the nth downsampled image is less than or equal to the size of the heterogeneous face image to be trained. For example, the size of the first downsampled image is 120*120, the size of the second downsampled image is 180*180, and the size of the third downsampled image is 240*240.
[0059] S203. Use the ith downsampled image to train the ith generation network and discriminator network respectively, and obtain the ith generation network in the ith stage and the discriminator network in the ith stage; i is an integer from 1 to n-1; n is an integer greater than or equal to 2.
[0060] Here, i is successively 1,..., n-1, n. In each stage, the generation network can be trained m times and the discriminator network can be trained k times, where m is an integer greater than or equal to 3, k is an integer greater than or equal to 2, and m is greater than k; for example, m can be 1000 and k can be 3.
[0061] Here, each time the generation network or the discriminator network is trained, the learning rate of the generation network or the learning rate of the discriminator network can be adjusted. Exemplarily, in the ith stage, the initial value of the learning rate of the ith generation network can be 0.0005, in the (i+1)th stage, the initial value of the learning rate of the (i+1)th generation network can be 0.0005, and in the first stage, the initial value of the learning rate of the discriminator network can also be 0.0005.
[0062] For example, when there are 3 stages, in the first stage, use the first downsampled image to train the first generation network and the discriminator network; then, enter the second stage, in the second stage, use the second downsampled image to train the first generation network, the second generation network and the discriminator network; then, enter the third stage, in the third stage, use the third downsampled image to train the first generation network, the second generation network, the third generation network and the discriminator network.
[0063] In some embodiments, the above S203 can be implemented through S2031 to S2037:
[0064] S2031. Input the feature vector of the ith downsampled image into the discriminator network to obtain the first discriminant matrix; the first discriminant matrix represents the authenticity of the ith downsampled image.
[0065] Here, the discrimination matrix generated by the discrimination network can be an N×N matrix and is used to represent whether the image input to the discrimination network is a real face image or a synthetic face image.
[0066] S2032. During the j-th training, perform data augmentation on the feature vector of the i-th downsampled image, and input the j-th augmented vector obtained into the i-th generation network to obtain the feature vector of the j-th fake image; j is an integer from 1 to m, and m is an integer greater than or equal to 3.
[0067] Here, j is successively 1, 2, …, m.
[0068] Here, the data augmentation process can specifically be any one of adding Gaussian noise processing, color-changing processing for changing the overall color of the image, and local cropping and color-changing processing.
[0069] S2033. Input the feature vector of the j-th fake image into the discrimination network to obtain the second discrimination matrix of the j-th time; the second discrimination matrix of the j-th time represents the authenticity of the j-th fake image.
[0070] S2034. When j is less than or equal to k, determine the j-th reconstruction loss and the j-th style consistency loss according to the feature vector of the i-th downsampled image, the feature vector of the j-th fake image, the first discrimination matrix, the second discrimination matrix of the j-th time, the first preset coefficient, and the second preset coefficient.
[0071] In the embodiment of the present invention, the result value can be obtained according to the feature vector of the j-th fake image and the feature vector of the i-th downsampled image, and the j-th reconstruction loss is determined according to the result value and the first preset coefficient; and, the difference between the mean value of the second discrimination matrix of the j-th time and the mean value of the first discrimination matrix is determined, and the j-th style consistency loss is determined according to the difference, the first discrimination matrix, and the second preset coefficient.
[0072] Here, the reconstruction loss for each training can be calculated using a reconstruction loss function. Exemplarily, the reconstruction loss function is shown in the following formula (1):
[0073]
[0074] Among them, when is the j-th reconstruction loss: G(f) is the feature vector of the j-th fake image, f is the feature vector of the i-th downsampled image, ||.||2 is the Euclidean norm, is the square of ||G(f)-f||2, which is the above result value, and α is the first preset coefficient.
[0075] Here, the style consistency loss for each training can be calculated using a style consistency loss function. Exemplarily, the style consistency loss function is shown in the following formula (2):
[0076]
[0077] Wherein, is the j-th style consistency loss, f represents an image, and f ∼ P g represents the data distribution for generating a pseudo-image, is the average value of the second discriminant matrix for the j-th time, and f ∼ P r represents the data distribution of the real downsampled image, is the average value of the first discriminant matrix, and f ∼ P f represents the sampled data distribution, with the sampling range between the generated pseudo-image and the real image, and λ is the second preset parameter, is the derivative symbol, ||.||2 is the Euclidean norm, represents the mean value.
[0078] In the embodiments of the present invention, the purpose of training the progressive adversarial generation network is to enable the generation network to reduce the style difference between the color mask template and the heterogeneous face image and deceive the discriminant network with the same color style. To this end, through the above-mentioned reconstruction loss and style consistency loss, an adversarial loss function can be constructed for training the generator and the discriminator, so as to achieve a balance; thus, compared with the traditional generative adversarial network using information entropy as the loss function, the above-mentioned loss function we used can train more robustly and produce quite reliable image quality.
[0079] S2035. Obtain the i-th generation network for the j-th time according to the j-th reconstruction loss and the i-th generation network, and obtain the discriminant network for the j-th time according to the j-th style consistency loss and the discriminant network.
[0080] Here, according to the j-th reconstruction loss, the parameters (e.g., learning rate) of the i-th generation network can be adjusted, and the i-th generation network with adjusted parameters is used as the i-th generation network for the j-th time; and according to the j-th style consistency loss, the parameters (e.g., learning rate) of the discriminant network can be adjusted, and the discriminant network with adjusted parameters is used as the discriminant network for the j-th time.
[0081] S2036. During the (j + 1)-th training, perform data augmentation on the feature vector of the i-th downsampled image, and input the augmented vector obtained for the (j + 1)-th time into the i-th generation network for the j-th time to obtain the feature vector of the pseudo-image for the (j + 1)-th time.
[0082] Here, during each training, the data augmentation process for the downsampled image processing is different, so the pseudo-images generated each time are also different. For example, when training a total of 1000 times, 1000 completely different pseudo-images can be generated through different data augmentation processes.
[0083] S2037. Input the feature vector of the (j + 1)-th pseudo-image into the j-th discrimination network to obtain the second discrimination matrix of the (j + 1)-th time.
[0084] S2038. When j + 1 is less than or equal to k, determine the (j + 1)-th reconstruction loss and the (j + 1)-th style consistency loss according to the first discrimination matrix, the second discrimination matrix of the (j + 1)-th time, the first preset coefficient, and the second preset coefficient. Obtain the i-th generation network of the (j + 1)-th time according to the i-th generation network of the j-th time and the (j + 1)-th reconstruction loss, and obtain the discrimination network of the (j + 1)-th time according to the discrimination network of the j-th time and the (j + 1)-th style consistency loss. When j + 1 is greater than k, determine the (j + 1)-th reconstruction loss according to the first discrimination matrix, the second discrimination matrix of the (j + 1)-th time, and the first preset coefficient, and obtain the i-th generation network of the (j + 1)-th time according to the i-th generation network of the j-th time and the (j + 1)-th reconstruction loss.
[0085] Here, the reconstruction loss can be generated using the above formula (1) during each training, and the style consistency loss can be generated using the above formula (2) during each training.
[0086] S2039. Until the i-th generation network of the m-th time and the discrimination network of the k-th time are obtained, take the i-th generation network of the m-th time as the i-th generation network of the i-th stage, and take the discrimination network of the m-th time as the discrimination network of the i-th stage; both m and k are integers greater than or equal to 2, and m is greater than k.
[0087] Here, when the (j + 1)-th training ends, continue the (j + 2)-th training using the same principle as in S2032 - S2035 above until the i-th generation network of the m-th time and the discrimination network of the k-th time are obtained.
[0088] Exemplarily, when k is 3, the discrimination network of this stage can be obtained after training the discrimination network 3 times in each stage; when m = 1000 times, the generation network of this stage can be obtained after training the generation network to be trained in this stage 1000 times in each stage.
[0089] S204. Obtain the i-th generation network to be trained in the (i + 1)-th stage according to the i-th generation network of the i-th stage, and obtain the discrimination network to be trained in the (i + 1)-th stage according to the discrimination network of the i-th stage.
[0090] Here, the learning rate of the i-th generation network finally obtained in the i-th stage can be reduced by a preset ratio to obtain the i-th generation network to be trained in the (i + 1)-th stage, and the learning rate of the discriminator network finally obtained in the i-th stage can be reduced by a preset ratio to obtain the discriminator network to be trained in the (i + 1)-th stage.
[0091] Here, the preset ratio can be set according to actual needs. For example, it can be nine-tenths. Thus, after reducing the learning rate of the first generation network finally obtained in the first stage to one-tenth of itself, it can be used as the first generation network to be trained in the second stage to continue training the first generation network to be trained in the second stage; and after reducing the learning rate of the discriminator network finally obtained in the first stage to one-tenth of itself, it can be used as the discriminator network to be trained in the second stage to continue training the discriminator network to be trained in the second stage; the same applies to other stages.
[0092] S205. Use the (i + 1)-th downsampled image to train the (i + 1)-th generation network, the i-th generation network to be trained in the (i + 1)-th stage, and the discriminator network to be trained in the (i + 1)-th stage respectively, to obtain the (i + 1)-th generation network in the (i + 1)-th stage, the i-th generation network in the (i + 1)-th stage, and the discriminator network in the (i + 1)-th stage, until n generation networks in the n-th stage and the discriminator network in the n-th stage are obtained.
[0093] In some embodiments, the above S205 can be implemented through S2051 to S2058:
[0094] S2051. Input the feature vector of the (i + 1)-th downsampled image into the discriminator network to be trained in the (i + 1)-th stage to obtain a third discrimination matrix; the third discrimination matrix represents the authenticity of the (i + 1)-th downsampled image.
[0095] S2052. At the z-th training, perform data augmentation on the feature vector of the (i + 1)-th downsampled image, and input the obtained z-th augmented vector into the i-th generation network to be trained in the (i + 1)-th stage to obtain an upsampled vector, and input the upsampled vector into the (i + 1)-th generation network to obtain the feature vector of the z-th fake image; z is an integer from 1 to m.
[0096] Here, z is successively 1, 2,..., m.
[0097] S2053. Input the feature vector of the z-th fake image into the discriminator network to be trained in the (i + 1)-th stage to obtain the z-th fourth discrimination matrix; the z-th fourth discrimination matrix represents the authenticity of the z-th fake image.
[0098] S2054. When z is less than or equal to k, according to the feature vector of the (i + 1)-th downsampled image, the feature vector of the z-th pseudo-image, the third discrimination matrix, the fourth discrimination matrix of the z-th time, the first preset coefficient, and the second preset coefficient, determine the z-th reconstruction loss and the z-th style consistency loss. According to the z-th reconstruction loss, the (i + 1)-th generation network, and the i-th generation network to be trained in the (i + 1)-th stage, obtain the (i + 1)-th generation network of the z-th time and the i-th generation network of the z-th time. And according to the z-th style consistency loss and the discrimination network to be trained in the (i + 1)-th stage, obtain the discrimination network of the z-th time.
[0099] Here, when z is 1, the parameters of the (i + 1)-th generation network can be adjusted according to the z-th reconstruction loss, and the (i + 1)-th generation network with adjusted parameters is used as the (i + 1)-th generation network of the z-th time; according to the z-th reconstruction loss, the parameters of the i-th generation network to be trained in the (i + 1)-th stage are adjusted, and the i-th generation network to be trained in the (i + 1)-th stage with adjusted parameters is used as the i-th generation network of the z-th time; according to the z-th style consistency loss, the parameters of the discrimination network to be trained in the (i + 1)-th stage are adjusted to obtain the discrimination network of the z-th time, and then the next iterative training is performed until the training ends.
[0100] Here, in each training in the (i + 1)-th stage, the adjustment method of the parameters of the (i + 1)-th generation network is the same as that of the parameters of the i-th generation network in each training in the i-th stage, while in each training in the (i + 1)-th stage, the adjustment method of the parameters of the i-th generation network is different from that of the parameters of the (i + 1)-th generation network. Exemplarily, in each training in the (i + 1)-th stage, the adjustment amplitude of the parameters of the (i + 1)-th generation network is greater than that of the parameters of the i-th generation network in each training in the (i + 1)-th stage. For example, in each training in the (i + 1)-th stage, the parameters of the (i + 1)-th generation network are increased by 0.0005, while in each training in the (i + 1)-th stage, the parameters of the i-th generation network are increased by 0.0001.
[0101] Here, in the training of each stage, when calculating the reconstruction loss, the above reconstruction loss function can be used for calculation, and when calculating the style consistency loss, the above style consistency loss function can be used for calculation.
[0102] S2055. In the (z + 1)-th training, perform data augmentation on the feature vector of the (i + 1)-th downsampled image, and input the obtained (z + 1)-th augmented vector into the i-th generation network of the z-th time to obtain an upsampled vector, and input the upsampled vector into the (i + 1)-th generation network of the z-th time to obtain the feature vector of the (z + 1)-th pseudo-image.
[0103] S2056. Input the feature vector of the (z + 1)-th pseudo-image into the z-th discrimination network to obtain the (z + 1)-th fourth discrimination matrix.
[0104] S2057. When z + 1 is less than or equal to k, determine the (z + 1)-th reconstruction loss and the (z + 1)-th style consistency loss according to the feature vector of the (i + 1)-th downsampled image, the feature vector of the (z + 1)-th pseudo-image, the third discrimination matrix, the (z + 1)-th fourth discrimination matrix, the first preset coefficient, and the second preset coefficient. According to the (z + 1)-th reconstruction loss, the (z)-th (i + 1)-th generation network, and the (z)-th i-th generation network, obtain the (z + 1)-th (i + 1)-th generation network and the (z + 1)-th i-th generation network. And according to the (z + 1)-th style consistency loss and the (z)-th discrimination network, obtain the (z + 1)-th discrimination network. When z + 1 is greater than k, determine the (z + 1)-th reconstruction loss according to the feature vector of the (i + 1)-th downsampled image, the feature vector of the (z + 1)-th pseudo-image, and the first preset coefficient. According to the (z + 1)-th reconstruction loss, the (z)-th (i + 1)-th generation network, and the (z)-th i-th generation network, obtain the (z + 1)-th (i + 1)-th generation network and the (z + 1)-th i-th generation network.
[0105] S2058. Until the (m)-th (i + 1)-th generation network, the (m)-th i-th generation network, and the k-th discrimination network are obtained, take the (m)-th (i + 1)-th generation network as the (i + 1)-th generation network in the (i + 1)-th stage, take the (m)-th i-th generation network as the i-th generation network in the (i + 1)-th stage, and take the k-th discrimination network as the discrimination network in the (i + 1)-th stage. Both m and k are integers greater than or equal to 2, and m is greater than k.
[0106] S206. Take the n generation networks in the n-th stage as n pre-trained generation networks.
[0107] For example, when n is 3, the first, second, and third generation networks obtained in the third stage can be taken as 3 pre-trained generation networks.
[0108] Here, the above training process of optimizing the generation network in the current stage using the generation network in the previous stage can significantly improve the training accuracy.
[0109] S106. Crop the mask area from the generated image according to the mask image.
[0110] Here, the method of cropping the mask area according to the mask image is as follows: The part of the mask in the mask image is all 1, and the rest is all 0. Therefore, the mask image can be AND-operated with the generated image to obtain only the mask image with the style changed.
[0111] S107. Merge the mask region with the heterogeneous face image to obtain a target heterogeneous face mask image.
[0112] Here, the principle of merging is as follows: Invert the mask image, perform an AND operation on the inverted mask image and the heterogeneous face image obtained in step S101 to obtain a background image, and add the mask image obtained in step S106 to the background image to obtain a target heterogeneous face mask image.
[0113] In the embodiment of the present invention, the position of the mask on the heterogeneous face image can be made more accurate, and the style (modality) of the mask region fused with the heterogeneous face image can be made consistent with the heterogeneous face image, thereby improving the quality of the finally generated heterogeneous face mask image.
[0114] Exemplarily, Figure 2 is a flowchart of an exemplary heterogeneous face mask image generation method based on a progressive adversarial generation architecture provided by the present invention. As Figure 2 shown, when training the progressive adversarial generation network, the original image is downsampled 3 times using 3 different preset sizes to obtain 3 downsampled images, namely the smallest size image f1, the intermediate size image f2, and the largest size image f3. In the first training stage, the image f1 is used to iteratively train the first generator G1 m times and the discriminator D k times. Specifically, in the first training stage, first input the feature vector of the image f1 into the discriminator D to obtain the discrimination matrix D(f1). Then, in each training, perform data augmentation on the feature vector of the image f1 and input the augmented vector into the generator G1. The generator G1 generates a fake image G1(f1). Then, input the fake image G1(f1) into the discriminator D to obtain the discrimination matrix D(G1(f1)), and calculate the corresponding reconstruction loss and style consistency loss based on D(f1) and D(G1(f1)) obtained this time respectively to adjust the generator G1 and the discriminator D respectively, so as to obtain the generator G1 and the discriminator D for this time. Iteratively train in this way until the generator G1 for the mth time and the discriminator D for the kth time are obtained, and the first stage is completed.
[0115] Based on the generator G1 finally obtained in the first stage, the generator G1 to be trained in the second stage is generated, and based on the discriminator D finally obtained in the first stage, the discriminator D to be trained in the second stage is generated. In the second training stage, the image f2 is used to iteratively train the generator G1 to be trained in the second stage and the second generator G2 for m times, and the discriminator D to be trained in the second stage is iteratively trained for k times. Specifically, in the second training stage, first, the feature vector of the image f2 is input into the discriminator D to be trained in the second stage to obtain the discrimination matrix D(f2). Then, during each training, the feature vector of the image f2 is subjected to data augmentation processing, and the augmented vector is input into the generator G1. The generator G1 upsamples the input features, and the upsampled vector is input into the generator G2. The generator G2 generates a fake image G2(f2). Then, the fake image G2(f2) is input into the discriminator D to obtain the discrimination matrix D(G2(f2)). Based on the D(f2) and D(G2(f2)) obtained this time, the corresponding reconstruction loss and style consistency loss are calculated respectively to adjust the generator G1, the generator G2, and the discriminator D respectively, to obtain the generator G1, the generator G2, and the discriminator D for this time. Such iterative training is performed until the generator G1 for the mth time, the generator G2 for the mth time, and the discriminator D for the kth time are obtained, completing the second stage.
[0116] The training principle of the third stage is the same as that of the second stage. The difference is that the image used is f3, and the fake image generated during each training is G3(f3). Moreover, during each training, the feature vector of the image f3 is subjected to data augmentation processing, and the augmented vector is input into the generator G1. The generator G1 upsamples the input features, and the upsampled vector is input into the generator G2. The generator G2 upsamples the input features and then inputs them into the generator G3. The generator G3 generates a fake image G3(f3). Then, the fake image G3(f3) is input into the discriminator D to obtain the discrimination matrix D(G3(f3)). Based on the D(f3) and D(G3(f3)) obtained this time, the corresponding reconstruction loss and style consistency loss are calculated respectively to adjust the generator G1, the generator G2, the generator G3, and the discriminator D respectively, to obtain the generator G1, the generator G2, the generator G3, and the discriminator D for this time. After the third stage is completed, the finally obtained G1, G2, G3, and D in the third stage are correspondingly used as the pre-trained G1, the pre-trained G2, the pre-trained G3, and the pre-trained D.
[0117] In the actual usage stage, key point detection is performed on the original image of the heterogeneous face image for which a mask needs to be generated. Based on the detected key points and the mask template, image initialization is carried out to obtain an initialization result (the above-mentioned initial heterogeneous face image), and a mask image of the initial heterogeneous face image is generated. After that, the feature vector of the initial heterogeneous face image is determined, and the feature vector is input into the pre-trained G1. After the pre-trained G1 upsamples the input feature vector, the upsampled vector is input into the pre-trained G2. After the pre-trained G2 upsamples the input feature vector, the upsampled vector is input into the pre-trained G3, and the pre-trained G3 generates an image to obtain a generation result (the above-mentioned generated image). After that, according to the mask image, the mask area is cropped from the generated image, and the mask area is fused with the heterogeneous face image to obtain the target heterogeneous face mask image.
[0118] The following further illustrates the effects that can be obtained by the method provided by the present invention through experimental data.
[0119] Images are generated using DirectADD, MaskTheFace, FMA-3D and the method of the present invention (Ours), and 1500 images are randomly selected from each of them and mixed to form a test set. Then, the above four methods and CASIA NIR-VIS 2.0 are used as training sets to train a face recognition model respectively. The experimental results are shown in Table 1, and Table 1 is the recognition results of masked faces under different training models. It can be observed that when these models are trained on masked faces, the recognition accuracy is higher than that of the models trained on unmasked images. Moreover, the method of the present invention performs the best in all recognition systems, which indicates that the method of the present invention is effective and can improve the performance of the face recognition system on masked faces.
[0120]
[0121] Table 1
[0122] Next, the image quality of DirectADD, MaskTheFace, FMA-3D and the method of the present invention is evaluated from both qualitative and quantitative aspects.
[0123] 1) Qualitative evaluation
[0124] Figure 3 Shows the comparison results of the images generated by different face mask synthesis methods. The (a) column is the original unmasked image. The (b) column is the image generated by the method of directly adding a colored mask (DirectADD). The (c) column is the image generated by adding a gray mask using MaskTheFace. The (d) column is the image generated by adding a gray mask using FMA-3D. The (e) column is the image generated by adding a colored mask using the method of the present invention. As Figure 3As shown, in column (b), the modal difference between the colored mask and the heterogeneous face results in a lack of realism in the image. Although the visual perception of other methods has been improved compared to column (b), the feeling of forgery is still inevitable. For example, as shown in the fourth row of column (b), the third row of column (c), and the second and third rows of column (d), when the background brightness is too dark, the gray mask causes an obvious contrast, making the style between the mask and the face image inconsistent. In addition, in terms of stability, the FMA-3D method is not as stable as our method, and there is a problem of interfering with the original image tone on some images, specifically resulting in the result shown in the fourth row of column (d). Obviously, in contrast, the images generated by the method of the present invention have sufficient realism and reliable quality.
[0125] 2) Quantitative evaluation
[0126] Six quality assessment metrics are used. Among them, the larger the values of SSIM, MSSSIM, PSNR, and FQA, the better the image quality, while the smaller the values of LPIPS and FID, the better the quality. The quality assessment results are shown in Table 2. In SSIM, the methods of the present invention and DirectADD scored 0.84, higher than FMA-3D and MaskTheFace, which scored 0.81. Under LPIPS and FID, the image quality comparison is significant. In LPIPS, the scores of DirectADD, MaskTheFace, and FMA-3D are 0.27, 0.25, and 0.23 respectively, at least 0.07 points higher than the method of the present invention. In FID, the score of the method of the present invention is 170.21, 100 points lower than the scores of DirectADD (284.1), MaskTheFace (285.6), and FMA-3D (270.24). Obviously, under all objective quality assessment criteria, the method of the present invention has the best score, which means that the images generated by the method of the present invention have better quality than existing synthesis methods.
[0127] Metric DirectADD MaskTheFace FMA-3D Ours SSIM↑ 0.84 0.81 0.81 0.84 MSSSIM↑ 0.85 0.79 0.76 0.89 PSNR↑ 21.14 19.10 18.65 24.63 LPIPS↓ 0.27 0.25 0.23 0.16 FID↓ 284.10 285.60 270.24 170.21 FQA↑ 0.19 0.21 0.32 0.34
[0128] Table 2
[0129] In summary, the method of the present invention synthesizes heterogeneous faces and colored mask templates with the help of a progressive generation architecture, which can ensure the quality and visual consistency of the generated images. The generated images contribute to improving the performance of the heterogeneous masked face recognition system. Compared with previous face mask synthesis methods, the method of the present invention can synthesize realistic masked results for faces in various scenarios, and the quality of the generated images is also higher.
[0130] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for generating heterogeneous face mask images based on a progressive adversarial generation architecture, characterized in that Including: Obtain heterogeneous face images; Extract face key points of the heterogeneous face images, select a target mask template from multiple preset mask templates according to the face key points, and determine a set of target key points according to the face key points; Each preset mask template corresponds to a set of preset key points; Determine a mask image according to the set of preset key points corresponding to the target mask template and the set of target key points; Synthesize the mask image and the heterogeneous face image to obtain an initial heterogeneous face mask image, and generate a mask image of the initial heterogeneous face image; Input the feature vector of the initial heterogeneous face image into a pre-trained generation module to obtain a generated image; the image styles of the mask part and the background part in the generated image are consistent; wherein, the pre-trained generation module includes n pre-trained generation networks, the output of the i-th generation network is used as the input of the (i + 1)-th generation network, the first n - 1 generation networks are all used to generate feature maps according to the input images and perform upsampling processing, and the n-th generation network is used to obtain the generated image according to the input feature map; the n pre-trained generation networks are obtained by training a progressive adversarial generation network; the progressive adversarial generation network includes n generation networks and a discriminator network; i is an integer from 1 to n - 1; n is an integer greater than or equal to 2; Crop the mask area from the generated image according to the mask image; Fuse the mask area and the heterogeneous face image to obtain a target heterogeneous face mask image.
2. The method for generating heterogeneous face mask images based on a progressive adversarial generation architecture according to claim 1, wherein Each preset mask template also corresponds to a face angle state; the step of selecting a target mask template from multiple preset mask templates according to the face key points includes: Determine the face deflection angle according to the face key points; Determine the face angle state according to the face deflection angle and a preset angle value; Select the target mask template from the multiple preset mask templates according to the face angle state.
3. The method for generating heterogeneous face mask images based on a progressive adversarial generation architecture according to claim 1, wherein Before inputting the feature vector of the initial heterogeneous face image into the pre-trained generation module to obtain a generated image, the method further includes: Obtain heterogeneous face images to be trained; Perform downsampling processing on the heterogeneous face images to be trained respectively according to n preset sizes, and correspondingly obtain n downsampled images; the sizes of the n downsampled images gradually increase; Use the i-th downsampled image to train the i-th generation network and the discriminator network respectively to obtain the i-th generation network and the i-th discriminator network in the i-th stage; Obtain the i-th generation network to be trained in the (i + 1)-th stage according to the i-th generation network in the i-th stage, and obtain the i-th discriminator network to be trained in the (i + 1)-th stage according to the discriminator network in the i-th stage; Use the (i + 1)-th downsampled image to train the (i + 1)-th generation network, the i-th generation network to be trained in the (i + 1)-th stage and the i-th discriminator network to be trained in the (i + 1)-th stage respectively, to obtain the (i + 1)-th generation network, the i-th generation network in the (i + 1)-th stage and the (i + 1)-th discriminator network in the (i + 1)-th stage, until obtaining n generation networks and the n-th discriminator network in the n-th stage; Take the n generation networks in the nth stage as n pre-trained generation networks.
4. The method for generating heterogeneous face mask images based on a progressive adversarial generation architecture according to claim 3, wherein The step of training the ith generation network and the discriminator network respectively using the ith downsampled image to obtain the ith generation network in the ith stage and the discriminator network in the ith stage includes: Input the feature vector of the ith downsampled image into the discriminator network to obtain a first discrimination matrix; the first discrimination matrix represents the authenticity of the ith downsampled image; During the jth training, perform data augmentation on the feature vector of the ith downsampled image, and input the obtained jth augmented vector into the ith generation network to obtain the feature vector of the jth fake image; j is an integer from 1 to m; m is an integer greater than or equal to 3; Input the feature vector of the jth fake image into the discriminator network to obtain the second discrimination matrix of the jth time; the second discrimination matrix of the jth time represents the authenticity of the jth fake image; When j is less than or equal to k, determine the jth reconstruction loss and the jth style consistency loss according to the feature vector of the ith downsampled image, the feature vector of the jth fake image, the first discrimination matrix, the second discrimination matrix of the jth time, the first preset coefficient, and the second preset coefficient; Obtain the ith generation network of the jth time according to the jth reconstruction loss and the ith generation network, and obtain the discriminator network of the jth time according to the jth style consistency loss and the discriminator network; During the (j + 1)th training, perform data augmentation on the feature vector of the ith downsampled image, and input the obtained (j + 1)th augmented vector into the ith generation network of the jth time to obtain the feature vector of the (j + 1)th fake image; Input the feature vector of the (j + 1)th fake image into the discriminator network of the jth time to obtain the second discrimination matrix of the (j + 1)th time; When j + 1 is less than or equal to k, determine the (j + 1)th reconstruction loss and the (j + 1)th style consistency loss according to the first discrimination matrix, the second discrimination matrix of the (j + 1)th time, the first preset coefficient, and the second preset coefficient, obtain the ith generation network of the (j + 1)th time according to the ith generation network of the jth time and the (j + 1)th reconstruction loss, and obtain the discriminator network of the (j + 1)th time according to the discriminator network of the jth time and the (j + 1)th style consistency loss; when j + 1 is greater than k, determine the (j + 1)th reconstruction loss according to the first discrimination matrix, the second discrimination matrix of the (j + 1)th time, and the first preset coefficient, and obtain the ith generation network of the (j + 1)th time according to the ith generation network of the jth time and the (j + 1)th reconstruction loss; Until the ith generation network of the mth time and the discriminator network of the kth time are obtained, take the ith generation network of the mth time as the ith generation network in the ith stage, and take the discriminator network of the mth time as the discriminator network in the ith stage; k is an integer greater than or equal to 2, and m is greater than k.
5. The heterogeneous face mask image generation method based on the progressive adversarial generation architecture according to claim 4, wherein The step of determining the jth reconstruction loss and the jth style consistency loss according to the feature vector of the ith downsampled image, the feature vector of the jth fake image, the first discrimination matrix, the second discrimination matrix of the jth time, the first preset coefficient, and the second preset coefficient includes: Obtain a result value based on the feature vector of the j-th pseudo-image and the feature vector of the i-th downsampled image; Determine the j-th reconstruction loss according to the result value and the first preset coefficient; Determine the difference between the mean of the second discriminant matrix of the j-th time and the mean of the first discriminant matrix; Determine the j-th style consistency loss according to the difference, the first discriminant matrix, and the second preset coefficient.
6. The method for generating heterogeneous face mask images based on a progressive adversarial generation architecture according to claim 3, wherein The step of using the (i + 1)-th downsampled image to train the (i + 1)-th generation network, the i-th generation network to be trained in the (i + 1)-th stage, and the discriminant network to be trained in the (i + 1)-th stage respectively, so as to obtain the (i + 1)-th generation network in the (i + 1)-th stage, the i-th generation network in the (i + 1)-th stage, and the discriminant network in the (i + 1)-th stage, until obtaining n generation networks in the n-th stage and the discriminant network in the n-th stage includes: Use the (i + 1)-th downsampled image to train the (i + 1)-th generation network, the i-th generation network to be trained in the (i + 1)-th stage, and the discriminant network to be trained in the (i + 1)-th stage respectively, so as to obtain the (i + 1)-th generation network in the (i + 1)-th stage, the i-th generation network in the (i + 1)-th stage, and the discriminant network in the (i + 1)-th stage; Obtain the (i + 1)-th generation network to be trained in the (i + 2)-th stage according to the (i + 1)-th generation network in the (i + 1)-th stage, obtain the i-th generation network to be trained in the (i + 2)-th stage according to the i-th generation network in the (i + 1)-th stage, and obtain the discriminant network to be trained in the (i + 2)-th stage according to the discriminant network in the (i + 1)-th stage; Use the (i + 2)-th downsampled image to train the (i + 2)-th generation network, the (i + 1)-th generation network to be trained in the (i + 2)-th stage, the i-th generation network to be trained in the (i + 2)-th stage, and the discriminant network to be trained in the (i + 2)-th stage respectively, so as to obtain the (i + 2)-th generation network in the (i + 2)-th stage, the (i + 1)-th generation network in the (i + 2)-th stage, the i-th generation network in the (i + 2)-th stage, and the discriminant network in the (i + 2)-th stage, until obtaining n generation networks in the n-th stage and the discriminant network in the n-th stage.
7. The method for generating heterogeneous face mask images based on a progressive adversarial generation architecture according to claim 6, wherein The step of using the (i + 1)-th downsampled image to train the (i + 1)-th generation network, the i-th generation network to be trained in the (i + 1)-th stage, and the discriminant network to be trained in the (i + 1)-th stage respectively, so as to obtain the (i + 1)-th generation network in the (i + 1)-th stage, the i-th generation network in the (i + 1)-th stage, and the discriminant network in the (i + 1)-th stage includes: Input the feature vector of the (i + 1)-th downsampled image into the discriminant network to be trained in the (i + 1)-th stage to obtain a third discriminant matrix; the third discriminant matrix represents the authenticity of the (i + 1)-th downsampled image; During the z-th training, perform data augmentation on the feature vector of the (i + 1)-th downsampled image, input the obtained z-th augmented vector into the i-th generation network to be trained in the (i + 1)-th stage to obtain an upsampled vector, and input the upsampled vector into the (i + 1)-th generation network to obtain the feature vector of the z-th pseudo-image; z is an integer from 1 to m; m is an integer greater than or equal to 3; Input the feature vector of the z-th fake image into the discriminative network to be trained in the (i + 1)-th stage, and obtain the fourth discriminant matrix of the z-th time; the fourth discriminant matrix of the z-th time represents the authenticity of the z-th fake image. When z is less than or equal to k, determine the z-th reconstruction loss and the z-th style consistency loss according to the feature vector of the (i + 1)-th downsampled image, the feature vector of the z-th fake image, the third discriminant matrix, the fourth discriminant matrix of the z-th time, the first preset coefficient, and the second preset coefficient. According to the z-th reconstruction loss, the (i + 1)-th generation network, and the i-th generation network to be trained in the (i + 1)-th stage, obtain the (i + 1)-th generation network of the z-th time and the i-th generation network of the z-th time, and according to the z-th style consistency loss and the discriminative network to be trained in the (i + 1)-th stage, obtain the discriminative network of the z-th time. During the (z + 1)-th training, perform data augmentation on the feature vector of the (i + 1)-th downsampled image, and input the obtained (z + 1)-th augmented vector into the i-th generation network of the z-th time to obtain an upsampled vector, and input the upsampled vector into the (i + 1)-th generation network of the z-th time to obtain the feature vector of the (z + 1)-th fake image. Input the feature vector of the (z + 1)-th fake image into the discriminative network of the z-th time to obtain the fourth discriminant matrix of the (z + 1)-th time. When z + 1 is less than or equal to k, determine the (z + 1)-th reconstruction loss and the (z + 1)-th style consistency loss according to the feature vector of the (i + 1)-th downsampled image, the feature vector of the (z + 1)-th fake image, the third discriminant matrix, the fourth discriminant matrix of the (z + 1)-th time, the first preset coefficient, and the second preset coefficient. According to the (z + 1)-th reconstruction loss, the (i + 1)-th generation network of the z-th time, and the i-th generation network of the z-th time, obtain the (i + 1)-th generation network of the (z + 1)-th time and the i-th generation network of the (z + 1)-th time, and according to the (z + 1)-th style consistency loss and the discriminative network of the z-th time, obtain the discriminative network of the (z + 1)-th time; when z + 1 is greater than k, determine the (z + 1)-th reconstruction loss according to the feature vector of the (i + 1)-th downsampled image, the feature vector of the (z + 1)-th fake image, and the first preset coefficient, and according to the (z + 1)-th reconstruction loss, the (i + 1)-th generation network of the z-th time, and the i-th generation network of the z-th time, obtain the (i + 1)-th generation network of the (z + 1)-th time and the i-th generation network of the (z + 1)-th time. Until the (i + 1)-th generation network of the m-th time, the i-th generation network of the m-th time, and the discriminative network of the k-th time are obtained, take the (i + 1)-th generation network of the m-th time as the (i + 1)-th generation network in the (i + 1)-th stage, take the i-th generation network of the m-th time as the i-th generation network in the (i + 1)-th stage, and take the discriminative network of the k-th time as the discriminative network in the (i + 1)-th stage; k is an integer greater than or equal to 2, and m is greater than k.
8. The method for generating heterogeneous face mask images based on a progressive adversarial generation architecture according to claim 3, wherein The obtaining of the i-th generation network to be trained in the (i + 1)-th stage according to the i-th generation network in the i-th stage and the obtaining of the discriminative network to be trained in the (i + 1)-th stage according to the discriminative network in the i-th stage include: Reduce the learning rate of the $i$-th generation network in the $i$-th stage by a preset ratio to obtain the $i$-th generation network to be trained in the $(i + 1)$-th stage; Reduce the learning rate of the discriminator network in the $i$-th stage by a preset ratio to obtain the discriminator network to be trained in the $(i + 1)$-th stage.
9. The method for generating heterogeneous face mask images based on a progressive adversarial generation architecture according to claim 5, wherein, The $j$-th style consistency loss is calculated by the style consistency loss function; the calculation formula of the style consistency loss function is as follows: Among them, is the j-th style consistency loss, f represents the image, and f ∼ P g represents the data distribution of the generated pseudo-image, is the average value of the second discriminant matrix at the j-th time, and f ∼ P r represents the data distribution of the real downsampled image, is the average value of the first discriminant matrix, and f ∼ P f represents the sampled data distribution, and the sampling range is between the generated pseudo-image and the real image. λ is the second preset parameter, is the derivative symbol, and ||.||2 is the Euclidean norm, represents the mean value.
10. The heterogeneous face mask image generation method based on the progressive adversarial generation architecture according to claim 1, wherein each pre-trained generation network adopts a fully convolutional neural network; the discriminator network is a Markov discriminator.
Citation Information
Patent Citations
Model training method and device, electronic equipment and computer readable storage medium
CN115376182A
Oct image-based image recognition method and apparatus, and device and storage medium
WO2021151276A1