Model training method, image processing method and device, electronic device, medium
By combining the joint training of the decoder reference vector space and the encoder feature vector space, the problem of low image attribute editing accuracy in the prior art is solved, and higher image attribute editing accuracy and stability are achieved.
Patent Information
- Application Number
- CN202111538449.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2041-12-15
AI Technical Summary
In the prior art, model training methods only consider similarity, resulting in low quality and accuracy after image attribute editing.
The encoder is jointly trained by the spatial distribution of the reference vector of the trained decoder and the spatial distribution of the eigenvector of the encoder, combining the training objectives of multiple dimensions, including minimizing the consistency of the spatial distribution of the eigenvector, and improving the training process of the encoder.
It improves the accuracy and operating range of face editing, and enhances the accuracy and stability of image attribute adjustment model.
Smart Images

Figure CN114239717B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to a model training method, an image processing method, a model training device, an image processing device, an electronic device, and a computer-readable storage medium. Background Art
[0002] During the image processing process, facial images can be edited and processed to meet the needs of various application scenarios.
[0003] In related technologies, the similarity between the reconstructed face and the input face can be used to train an encoder, thereby training a model for image editing. However, this approach only considers similarity when training the model, resulting in low model accuracy. This results in poor quality and low accuracy of the resulting image after attribute adjustment.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention
[0005] The purpose of the present disclosure is to provide a model training method, an image processing method and device, an electronic device, and a storage medium, thereby overcoming, at least to a certain extent, the problem of low model accuracy caused by the limitations and defects of related technologies.
[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.
[0007] According to one aspect of the present disclosure, a model training method is provided, including: training a decoder based on a sample image to obtain a trained decoder; extracting multiple feature vectors of the sample image through an encoder, and training the encoder in combination with the spatial distribution of the reference vector of the trained decoder and the spatial distribution of the multiple feature vectors to obtain a trained encoder; and obtaining an image attribute adjustment model for performing attribute editing on the image based on the trained decoder, the trained encoder, and the attribute editing model.
[0008] According to one aspect of the present disclosure, there is provided an image processing method, comprising: obtaining an image to be processed; extracting a feature vector of the image to be processed according to an image attribute adjustment model, and performing an editing operation on the feature vector to obtain an editing vector to generate an attribute image corresponding to the image to be processed; the image attribute adjustment model is trained according to any one of the model training methods described above.
[0009] According to one aspect of the present disclosure, a model training device is provided, including: a decoder training module, used to train a decoder according to a sample image to obtain a trained decoder; an encoder training module, used to extract multiple feature vectors of the sample image through an encoder, and train the encoder in combination with the spatial distribution of the reference vector of the trained decoder and the spatial distribution of the multiple feature vectors to obtain a trained encoder; a model acquisition module, used to obtain an image attribute adjustment model for performing attribute editing on an image based on the trained decoder, the trained encoder and the attribute editing model.
[0010] According to one aspect of the present disclosure, an image processing device is provided, comprising: an image acquisition module for acquiring an image to be processed; an image generation module for extracting a feature vector of the image to be processed according to an image attribute adjustment model, and performing an editing operation on the feature vector to obtain an editing vector to generate an attribute image corresponding to the image to be processed; the image attribute adjustment model is trained according to the model training method described in any one of the above items.
[0011] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned model training methods or any one of the above-mentioned image processing methods by executing the executable instructions.
[0012] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above-mentioned model training methods or any one of the above-mentioned image processing methods.
[0013] In the model training method, model training device, image processing method, image processing device, electronic device and computer-readable storage medium provided in the embodiments of the present disclosure, on the one hand, the encoder is trained in combination with the spatial distribution of the reference vector of the trained decoder and the spatial distribution of the feature vector output by the encoder, thereby improving the training process of the encoder, wherein the balance between reconstruction error and editability is taken into account, so that the feature vector obtained by the encoder can be within the spatial distribution corresponding to the decoder. Since face editing can only be performed if the obtained feature vector is in the spatial distribution of the decoder, the editability of face attributes is enhanced, and the accuracy and operating range of face editing are improved. On the other hand, the encoder can be trained in combination with multiple dimensions, avoiding the limitation of model training based on only one target, improving the accuracy of the encoder, and thereby improving the accuracy and stability of the image attribute adjustment model.
[0014] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0016] Figure 1 A schematic diagram of a system architecture to which the model training method or image processing method of an embodiment of the present disclosure can be applied is shown.
[0017] Figure 2 A schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown.
[0018] Figure 3 A schematic diagram schematically illustrates a model training method in an embodiment of the present disclosure.
[0019] Figure 4 A schematic diagram of training an encoder in an embodiment of the present disclosure is schematically shown.
[0020] Figure 5 A schematic diagram schematically illustrates the overall process of encoder training in an embodiment of the present disclosure.
[0021] Figure 6 A schematic diagram schematically illustrates an image attribute adjustment model in an embodiment of the present disclosure.
[0022] Figure 7 The following is a flow chart schematically illustrating an image processing method in an embodiment of the present disclosure.
[0023] Figure 8 The following schematically illustrates a flow chart of generating an attribute image in an embodiment of the present disclosure.
[0024] Figure 9 The following schematically illustrates a flow chart of determining a feature vector in an embodiment of the present disclosure.
[0025] Figure 10 A schematic diagram schematically illustrates image attribute editing in an embodiment of the present disclosure.
[0026] Figure 11 A block diagram of a model training device in an embodiment of the present disclosure is schematically shown.
[0027] Figure 12 The following schematically shows a block diagram of an image processing device in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0029] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0030] The present disclosure provides a model training method that can be applied to application scenarios where attributes of facial images are partially edited and adjusted.
[0031] Figure 1 A schematic diagram shows a system architecture to which the model training method and apparatus or image processing method and apparatus of the embodiments of the present disclosure can be applied.
[0032] like Figure 1As shown, the system architecture 100 may include a client 101, a network 102, and a server 103. The client may be a client, for example, a smartphone, a computer, a tablet computer, a smart speaker or other terminal. The network 102 is used to provide a medium for a communication link between the client 101 and the server 103. The network 102 may include various connection types, such as a wired communication link, a wireless communication link, and the like. In the embodiment of the present disclosure, the network 102 between the client 101 and the server 103 may be a wired communication link, for example, a communication link may be provided through a serial port connection line; or it may be a wireless communication link, providing a communication link through a wireless network. The server 103 may be a server or a client with computing functions, such as a portable computer, a desktop computer, a smartphone or other terminal device with computing functions, which is used to perform model training on the images sent by the client and perform image processing according to the trained model.
[0033] This model training method can be applied to the application scenario of the model for image editing. Figure 1 As shown in , this method can be specifically applied to a process in which a client 101 sends a sample image 104 to a server 103, and the server 103 extracts features from the sample image obtained from the client to train a model. The client can be any type of computing device, such as a smartphone, tablet computer, desktop computer, in-vehicle device, wearable device, etc. The sample image can be any type of image, such as a facial image.
[0034] The server 103 can use the sample images sent by the client 101 to train the decoder, further train the encoder in combination with the reference vector space distribution of the trained decoder, and form an image attribute adjustment model based on the trained encoder, the trained decoder and the attribute editing model.
[0035] Furthermore, when the server 103 receives the image to be processed sent by the client 101, it can use the encoder, attribute editing model, and decoder to perform image processing on the image to be processed to obtain a corresponding attributed image. The server 103 can also send the attributed image to the client 101 for display and other image processing operations.
[0036] It should be noted that the model training method and image processing method provided in the embodiments of the present disclosure can be completely executed by a server. Accordingly, the model training device and image processing device can be set in the server.
[0037] Figure 2 Schematic diagram of an electronic device suitable for implementing an exemplary embodiment of the present disclosure is shown. The terminal of the present disclosure can be configured as follows Figure 2The form of the electronic device shown, however, needs to be explained. Figure 2 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0038] The electronic device of the present disclosure includes at least a processor and a memory, where the memory is used to store one or more programs. When the one or more programs are executed by the processor, the processor can implement the method of the exemplary embodiment of the present disclosure.
[0039] Specifically, such as Figure 2 As shown, the electronic device 200 may include: a processor 210, an internal memory 221, an external memory interface 222, a Universal Serial Bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 271, a receiver 272, a microphone 273, an earphone interface 274, a sensor module 280, a display 290, a camera module 291, an indicator 292, a motor 293, a button 294, and a Subscriber Identification Module (SIM) card interface 295. The sensor module 280 may include a depth sensor, a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, and a bone conduction sensor.
[0040] It is understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0041] The processor 210 may include one or more processing units, for example: the processor 210 may include an application processor, a modem processor, a graphics processor, an image signal processor, a controller, a video codec, a digital signal processor, a baseband processor and / or a neural network processor (Neural-etwork Processing Unit, NPU), etc. Among them, different processing units can be independent devices or integrated into one or more processors. In addition, a memory can be provided in the processor 210 for storing instructions and data. The model training method in this exemplary embodiment can be executed by an application processor, a graphics processor or an image signal processor. When the method involves processing related to a neural network, it can be executed by an NPU.
[0042] The internal memory 221 can be used to store computer executable program code, which includes instructions. The internal memory 221 can include a program storage area and a data storage area. The external memory interface 222 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200.
[0043] The communication functions of mobile terminal 200 are implemented through a mobile communication module, antenna 1, a wireless communication module, antenna 2, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module can provide 2G, 3G, 4G, and 5G mobile communication solutions for mobile terminal 200. The wireless communication module can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 200.
[0044] The display module is used to implement display functions, such as displaying user interfaces, images, and videos. The camera module is used to implement shooting functions, such as capturing images and videos. The audio module is used to implement audio functions, such as playing audio and capturing voice. The power module is used to implement power management functions, such as charging the battery, powering the device, and monitoring the battery status.
[0045] The present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device.
[0046] Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device.
[0047] Computer-readable storage media can transmit, propagate, or transfer programs for use by or in conjunction with an instruction execution system, apparatus, or device. Program code contained on a computer-readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, or any suitable combination thereof.
[0048] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments.
[0049] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0050] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, and the units described may also be provided in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0051] In related technologies, the specific process of model training includes: using the original face as a sample, training a face reconstructor for face restoration, and training latent vectors representing facial attributes on the original face; finding a normal vector in the space of the latent vectors that indicates the direction of editing each of the facial attributes; and adjusting the latent vectors along the normal vectors to generate target face data in the face reconstructor. The model's single goal is to reconstruct a face that represents the similarity between the original face and the original face, resulting in poor quality faces after attribute editing.
[0052] In order to solve the problems in the related art, the training process of the encoder is adjusted in the embodiment of the present disclosure. Figure 3 The model training method in the embodiment of the present disclosure is described in detail.
[0053] In step S310, the decoder is trained according to the sample image to obtain a trained decoder.
[0054] In the embodiments of the present disclosure, the sample image may be a facial image or any type of image such as a landscape. The sample image herein is taken as an example for explanation. The facial image may be any type of facial image, such as a color facial image or a black and white facial image. The facial image may be a facial image in a variety of application scenarios. The various application scenarios may include, but are not limited to, different groups of people, expressions, postures, lighting, and environments. Furthermore, the facial image refers to a real facial image, i.e., a facial image directly captured by a camera or obtained from a network, a storage device, or a terminal's photo album without any image processing operations.
[0055] The decoder can be a GAN (Generative Adversarial Network), a generative model capable of generating new content. It can be used for applications such as generating synthetic training data, creating art, style transfer, and image-to-image translation. A generative adversarial network consists of two networks: a generator for generating samples and a discriminator. The generator attempts to generate fake samples and trick the discriminator into believing they are real. The discriminator is responsible for distinguishing between generated samples and fake ones.
[0056] In the disclosed embodiments, a decoder can be used to represent the mapping from a latent vector (feature vector) to a facial image. Specifically, the decoder can be a GAN (Generative Adversarial Network), such as the styleGAN2 model, which includes a generator and a discriminator. The decoder can include an input layer, a convolutional layer, a pooling layer, a fully connected layer, an output layer, and the like.
[0057] To obtain an accurate decoder, the decoder can be trained based on sample images to obtain a trained decoder. A generative adversarial approach can be used to train the decoder. During decoder training, feature extraction is performed on the sample image to determine the sample image's feature vector. The sample image's feature vector is mapped using a generator in a generative adversarial network to obtain a predicted sample image. A discriminant operation is performed on the sample image and the predicted sample image to train the decoder, thereby obtaining a trained decoder. Specifically, the discriminator in the generative adversarial network performs a discriminative operation to determine whether the predicted sample image is a real image, that is, to determine whether the predicted sample image is true or false. If the input is a real sample, the discriminator output is close to 1; if the input is a false sample, the discriminator output is close to 0. Furthermore, with the training goal of minimizing the difference between the sample image and the predicted sample image, the generator and discriminator are used to adjust the model parameters until the discriminator cannot distinguish whether the predicted sample image is a real image or an image generated by the generator, completing the entire decoder training process.
[0058] Based on this, the styleGAN2 model can be trained on real face datasets, or open-source pre-trained styleGAN2 models can be used. Training styleGAN2 requires a large amount of real face image data. The decoder is trained using an adversarial generation approach, enabling the decoder to generate realistic images of faces, improving the accuracy of the decoder and the precision of the images generated by the decoder.
[0059] In step S320, a plurality of feature vectors of the sample image are extracted by an encoder, and the encoder is trained in combination with the spatial distribution of the reference vector of the trained decoder and the spatial distribution of the plurality of feature vectors to obtain a trained encoder.
[0060] In the disclosed embodiments, the sample image may be a real face image, such as face data from a real face dataset. The encoder may be a residual network structure. The residual network structure is composed of multiple residual blocks, each of which is divided into a direct mapping part and a residual part. The residual part is generally composed of multiple convolution operations.
[0061] When training an encoder, the encoder can be trained in conjunction with a trained decoder. When training an encoder in conjunction with a trained decoder, the parameters of the trained decoder can be used as a reference for training. Referring to the parameters of a trained decoder means fixing the parameters of the trained decoder so that they remain unchanged during the encoder training process. Specifically, during the encoder training process, the encoder is connected to the trained decoder, and the parameters of the trained decoder are fixed to train the encoder. Figure 4As shown in , the parameters of the trained decoder 403 are fixed, and the encoder 401 is trained according to the sample image 402 to obtain the trained encoder 404. The decoder can be located downstream of the encoder, that is, the encoder is followed by the trained decoder.
[0062] Based on this, the encoder can be trained based on the training objectives and the trained decoder. To address the technical issues in related technologies, the encoder training objectives can be jointly determined based on the trained decoder and encoder. Specifically, the training objectives are jointly constructed based on the consistency between the spatial distribution of the reference vectors of the trained decoder and the spatial distribution of the feature vectors output by the encoder, and minimizing the reconstruction error.
[0063] Among them, minimizing the reconstruction error is minimizing the face reconstruction error, which is used to indicate minimizing the error between the input image and the output image. The input image refers to the real face image represented by the sample image input to the encoder. Since the encoder is connected to the trained decoder, and the decoder is used to generate images, the output image refers to the face image corresponding to the sample image generated by the trained decoder. The face image can be the same as or different from the sample image, which is not limited here. The specific processing process includes: extracting features from the sample image through the encoder to obtain a feature vector of the sample image; decoding the feature vector of the sample image through the trained decoder to achieve a mapping of the feature vector of the sample image to the face image, thereby generating a face image corresponding to the sample image. Based on this, the error between the output image and the input image can be calculated and minimized. Specifically, the similarity between the output image and the input image can be calculated to determine the error between the two.
[0064] After obtaining the error, the training target can be determined based on minimizing the face reconstruction error and the spatial distribution of the feature vector, and the encoder can be trained based on the training target to obtain a trained encoder.
[0065] In the embodiment of the present disclosure, the number of feature vectors may be N. When obtaining the feature vector of the sample image through the encoder, a main sample feature vector may be extracted first. At the same time, multiple offset sample feature vectors may be obtained based on the main sample feature vector. The number of multiple offset sample feature vectors may be N-1. Both the main sample feature vector and the multiple offset sample feature vectors may be 512-dimensional feature vectors. Specifically, the value of the main sample feature vector may be increased or decreased to adjust the main sample feature vector, thereby obtaining N-1 offset sample feature vectors. The N-1 offset sample feature vectors may be obtained in different ways, and the multiple offset sample feature vectors may be the same or different. Furthermore, the main sample feature vector and the offset sample feature vector may be fused to obtain multiple feature vectors corresponding to the sample image. The fusion operation may be an addition operation, which is not specifically limited here.
[0066] In the related art, when training the encoder, the training objective can be represented by the reconstruction loss. The reconstruction loss is used to enable the feature vector obtained by the encoder to represent the features of the input sample image, that is, the feature vector can restore the input sample image after passing through the decoder. However, if the encoder is trained only based on the reconstruction loss, the spatial distribution of the feature vector output by it is quite different from the spatial distribution of the reference vector of the decoder, thereby affecting the editability of the feature vector. In order to solve the above problems in the embodiment of the present disclosure, the spatial distribution of the reference vector can be consistent with the spatial distribution of the multiple feature vectors, and the minimization of the reconstruction error can be determined as the training objective, so as to construct the training objective from multiple dimensions.
[0067] In the process of determining the training target based on the vector space distribution and minimizing the reconstruction error, in order to accurately determine the training target, the spatial distribution of the feature vector can be adjusted. Based on the above-mentioned feature extraction method of obtaining a main sample feature vector and an offset sample feature vector, the spatial distribution of the feature vector can be adjusted by minimizing the variance. Specifically, minimizing the variance adjusts the spatial distribution of multiple feature vectors corresponding to the sample image by constraining the norm of the offset term (offset sample feature vector), so that the spatial distribution of multiple feature vectors is more compact. The norm here can be the L2 norm. The L2 norm refers to the square root of the sum of the squares of the elements of the offset sample feature vector. In addition, the norm can also be the L0 norm or the L1 norm, etc., which is not limited here. Minimizing the variance is used to represent the variance corresponding to the offset sample feature vector, and can be calculated according to formula (1):
[0068]
[0069] Where △ represents N-1 offset sample feature vectors.
[0070] In addition, a discriminator for the feature vector can also be used to constrain the encoder through adversarial learning to adjust the spatial distribution of the feature vector so that the spatial distribution of the feature vector is consistent with the spatial distribution of the reference vector. The discriminator is equivalent to a classifier, which is used to perform a discriminant operation on the feature vector obtained by the encoder and the input vector required by the decoder to determine whether the vector comes from the decoder or the encoder. In the process of constraining the encoder with the discriminator corresponding to the feature vector, the training is stopped when the discriminator cannot determine the difference between the feature vector output by the encoder and the input vector of the decoder to obtain a trained encoder, and the distribution space of the feature vector output by the encoder is consistent with the spatial distribution of the reference vector of the decoder through the trained encoder. The input vector can be sampled from the actual vector obtained by mapping the actual distribution. In the embodiment of the present disclosure, based on the discriminator for the feature vector, the encoder is constrained and trained through adversarial learning so that the spatial distribution of the feature vector is consistent with the spatial distribution of the reference vector of the decoder, so that the feature vector obtained by the encoder is within the spatial distribution of the reference vector corresponding to the decoder.
[0071] In order to solve the technical problems in the related art, the loss function corresponding to the training target can be determined. Specifically, the loss function can be determined by minimizing the variance and minimizing the offset loss function and the reconstruction loss function. Among them, minimizing the variance and minimizing the offset loss function can be used to constrain the consistency of the spatial distribution of the feature vector. The minimizing the offset loss function can be calculated by formula (2):
[0072]
[0073] Among them, D w is the discriminator, and γ is used to indicate the importance of the corresponding item.
[0074] In addition, the reconstruction loss (L2,L LPIPS ,L sim ), L2 represents the norm; L LPIPS It is the feature of the image obtained by the weighted convolutional neural network, which is used to constrain the features so that the features are similar; L sim It is a structural similarity loss that makes two images similar at the structural level.
[0075] Furthermore, by adding the minimization of variance and minimization of offset loss functions as loss functions, the encoder can be trained to constrain the feature vectors output by the encoder, so that the spatial distribution of multiple feature vectors is consistent with the spatial distribution of the reference vector, that is, the same. Specifically, the loss function can be obtained by performing a weighted sum operation on the minimization of variance and minimization of offset loss functions and the reconstruction loss function, that is, determining the weight corresponding to each item and the product of each item, and then adding all the products.
[0076] After determining the loss function corresponding to the training objective, the encoder can be trained based on the loss function, with the trained decoder parameters fixed. Specifically, the encoder model parameters can be adjusted until the loss function is minimized, terminating the model training process to obtain the trained encoder. The trained encoder maps images to vectors, generating latent vectors (feature vectors) that represent facial features. This also considers the balance between reconstruction error and editability.
[0077] Figure 5 The overall flow chart of encoder training is shown schematically in Figure 5 As shown in , a sample image 501 is input to an encoder 502 to obtain a main sample feature vector 5021 and N-1 offset sample feature vectors 5022. The main sample feature vector 5021 and the N-1 offset sample feature vectors 5022 are further added to obtain N-1 feature vectors 5023. The main feature vector 5021 and the N-1 feature vectors 5023 are combined to obtain N feature vectors 5024. The minimum variance 5025 is determined based on the offset sample feature vector, and the discriminator 504 is added to the feature vector 5024 to obtain the minimum offset loss function 5026 to adjust the spatial distribution of the feature vector. The feature vector 5024 is input to the decoder 503 for decoding to obtain an output image 505 corresponding to the sample image, and the reconstruction loss is determined based on the output image 505 to train the encoder.
[0078] It should be added that transformers can also be used to construct encoders and decoders, as long as they can achieve the corresponding functions, and there is no limitation here.
[0079] Continue to refer Figure 3 As shown in , in step S330, an image attribute adjustment model for performing attribute editing on an image is obtained based on the trained decoder, the trained encoder, and the attribute editing model.
[0080] In the disclosed embodiments, an attribute editing model is used to adjust the target attributes of an image. The attribute editing model aims to obtain the target facial image by manipulating the feature vector (moving it in a specific direction) and then decoding it through a decoder. This means finding the mapping relationship between the direction vector in the feature vector space and the attributes of the generated image. To this end, a batch of sampled feature vectors can be randomly sampled, and then the decoder decodes the sampled feature vectors to generate an intermediate image. The facial images represented by these intermediate images are then attribute-labeled to obtain various attribute values, such as gender, age, and expression. For each attribute, the attribute is binarized to obtain a classification (e.g., male / female, old / young, smiling / not smiling) and a label is obtained. A linear binary classifier is used on the data pairs consisting of the sampled feature vectors and the labels (latents labels) to determine the classification hyperplane for the binary attribute, thereby completing the training of the attribute editing model. That is, the classification hyperplane for each attribute can be determined based on the sampled feature vectors and the intermediate image output by the decoder to train the attribute editing model. Alternatively, vector extraction can be performed on the sample image to obtain the sampled feature vectors, or a batch of feature vectors can be directly obtained as the sampled feature vectors, which is not limited here.
[0081] On this basis, the trained decoder, the trained encoder, and the attribute editing model can be combined to obtain an image attribute adjustment model for image attribute editing. Figure 6 As shown in , the trained decoder 601 , the attribute editing model 602 , and the trained decoder 603 are combined to obtain the image attribute adjustment model 600 .
[0082] In the disclosed embodiment, the structure of the encoder is adjusted by improving the training objectives and loss function of the encoder module, and a discriminator for the feature vector is added to constrain the training process of the encoder to obtain a trained encoder. Since the balance between reconstruction error and editability is taken into account, the consistency of the spatial distribution of the feature vectors of the encoder and the decoder is added to the training objective of the encoder, which can make the feature vector obtained by the encoder fall within the spatial distribution corresponding to the decoder. Since face editing can only be performed if the obtained feature vector is in the spatial distribution of the decoder, the editability of face attributes is enhanced, and the accuracy and operating range of face editing are improved. In addition, the encoder can be trained in combination with multiple dimensions, avoiding the limitation of model training based on only one objective, improving the accuracy of the encoder, and thereby improving the accuracy and stability of the image attribute adjustment model.
[0083] In the embodiment of the present disclosure, an image processing method is also provided. Figure 7 As shown in , it mainly includes the following steps:
[0084] In step S710, an image to be processed is obtained;
[0085] In step S720, a feature vector of the image to be processed is extracted according to the image attribute adjustment model, and an editing operation is performed on the feature vector to obtain an editing vector, so as to generate an attribute image corresponding to the image to be processed.
[0086] In the embodiments of the present disclosure, the image to be processed may be any type of image, specifically a facial image or other type of image, etc. The facial image may be any type of facial image, such as a color facial image or a black and white facial image. The facial image may be a facial image in a variety of application scenarios. Based on this, the image to be processed may be a real facial image, that is, a face directly captured by a camera or a facial image obtained from a network, a storage device, or a terminal album, etc., without undergoing any image processing operation. The image to be processed may be one image or a batch of images, which is not limited here. When the image to be processed is a batch of images, batch processing may be performed based on an image attribute adjustment model to improve image processing efficiency.
[0087] Next, the image attribute adjustment model can be used to adjust the attributes of the image to be processed, thereby obtaining an attribute image corresponding to the image to be processed. Specifically, feature extraction can be performed on the image to obtain a feature vector, which can then be attribute-adjusted. Image generation can then be performed based on the attribute-adjusted edit vector to obtain the attribute image. A trained encoder can be used to extract features from the image to obtain a feature vector, which can then be edited using the attribute editing model to obtain an edit vector. The edit vector can then be generated using a trained decoder to map the feature vector to the facial image, thereby obtaining the attribute image.
[0088] refer to Figure 8 As shown in the flowchart of generating attribute images, the image to be processed 801, such as a face image, is subjected to feature extraction by the encoder 802 to obtain a feature vector; the attribute editing module 803 obtains the edited vector after attribute editing by editing the feature vector (moving it in a specific direction), and then the edited vector is input into the decoder 804 to obtain the face image after the target attribute is changed, that is, the attribute image 805.
[0089] The trained encoder is a mapping from image to vector, which can obtain the latent vector representing the facial features, that is, the feature vector. Figure 9 The flow chart for determining the feature vector is shown schematically in FIG. Figure 9 As shown in , it mainly includes the following steps:
[0090] In step S910, feature extraction is performed on the image to be processed to obtain a main feature vector and multiple offset feature vectors;
[0091] In step S920, the main eigenvector and the multiple offset eigenvectors are fused to obtain the eigenvector.
[0092] In the embodiment of the present disclosure, the number of feature vectors may be N, and each feature vector is a 512-dimensional vector. In the process of extracting N feature vectors, a 512-dimensional main feature vector may be extracted first. At the same time, multiple offset feature vectors may be obtained based on the main feature vector. Specifically, the main feature vector may be adjusted to obtain multiple offset feature vectors. The adjustment method may be to increase or decrease the value of the main feature vector, and the number of offset feature vectors may be the number of feature vectors minus the number of main feature vectors, for example, N-1.
[0093] Furthermore, the main eigenvector and the offset eigenvector can be fused to obtain multiple eigenvectors. Specifically, the main eigenvector and the offset eigenvector can be added to obtain eigenvectors corresponding to N-1 offset eigenvectors, and the eigenvectors corresponding to the main eigenvector and the offset eigenvector can be combined to obtain multiple eigenvectors, i.e., N eigenvectors.
[0094] In the embodiment of the present disclosure, a main feature vector is obtained through feature extraction, and the main feature vector is adjusted to obtain multiple offset feature vectors, and then the feature vector of the image to be processed is obtained by fusing the main feature vector and the offset feature vector. This can improve the accuracy and comprehensiveness of the feature vector, and also improve the efficiency of obtaining the feature vector.
[0095] Specifically, the target attribute among multiple attributes can be adjusted in response to the movement operation of the feature vector along the normal vector, and the attribute image can be generated according to the target attribute. The multiple attributes may include gender, age, expression, and the like. For each attribute, the attribute is binarized (such as male or female, old or young, smiling / not smiling) to obtain labels, and a linear binary classifier (such as linear SVM) is used on the data pairs (latents labels) consisting of the feature vector and the label to find the classification hyperplane and normal vector of the binary attribute (such as male or female). The target attribute can be any one of the multiple attributes, such as age. If it is detected that the feature vector moves along the normal vector, the target attribute of the image to be processed can be adjusted according to the moving direction of the feature vector, and the specific adjustment value can be determined according to the moving operation. For example, the age attribute of the generated face is changed (such as getting older).
[0096] refer to Figure 10As shown in , the image to be processed 1001 is input to the encoder 1002 to obtain a main feature vector 1021 and N-1 offset feature vectors 1022. The main feature vector 1021 and the N-1 offset feature vectors 1022 are further added to obtain N-1 feature vectors 1023. The main feature vector 1021 and the N-1 feature vectors 1023 are combined to obtain N feature vectors 1024. After obtaining the feature vector 1024, the feature vector can be attribute edited by the attribute editing model 1004 to obtain an edit vector 1025. The edit vector 1025 is input to the decoder 1003 so that the decoder 1003 realizes the mapping of the latent vector to the attribute image, thereby obtaining the attribute image 1005 corresponding to the image to be processed. For example Figure 10 As shown in , the age attribute of the image to be processed is adjusted to obtain an aged edited image.
[0097] The technical solution in the embodiments of the present disclosure, by adjusting the structure of the encoder, takes into account the editability of the feature vector output by the encoder and the consistency of the spatial distribution of the feature vectors of the encoder and decoder during the model training process, thereby improving the authenticity and naturalness of the generated attribute image, and can improve the face quality and image accuracy after editing the attributes, and also improves the editing intensity, and increases the flexibility and operability of image editing.
[0098] The present disclosure provides a model training device, referring to Figure 11 As shown in , the model training device 1100 may include:
[0099] The decoder training module 1101 is used to train the decoder according to the sample image to obtain a trained decoder;
[0100] An encoder training module 1102 is configured to extract multiple feature vectors of the sample image through an encoder, and train the encoder based on the spatial distribution of the reference vector of the trained decoder and the spatial distribution of the multiple feature vectors to obtain a trained encoder;
[0101] The model acquisition module 1103 is used to acquire an image attribute adjustment model for performing attribute editing on an image based on the trained decoder, the trained encoder, and the attribute editing model.
[0102] In an exemplary embodiment of the present disclosure, the decoder training module includes: an adversarial training module, configured to train the decoder based on the sample image in an adversarial generation manner to obtain a trained decoder.
[0103] In an exemplary embodiment of the present disclosure, the encoder training module includes: a joint training module, which is used to make the reference vector space part consistent with the spatial distribution of the multiple feature vectors, and minimize the reconstruction error as a training goal, and train the encoder in combination with the trained decoder to obtain a trained encoder.
[0104] In an exemplary embodiment of the present disclosure, the joint training module includes: a loss function determination module, which is used to determine the loss function according to the training objective; a parameter fixing module, which is used to fix the parameters of the trained decoder, train the encoder according to the loss function corresponding to the training objective, and obtain a trained encoder.
[0105] In an exemplary embodiment of the present disclosure, the loss function determination module includes: a determination control module, configured to determine the loss function according to minimizing the variance and minimizing the offset loss function and the reconstruction loss function.
[0106] In an exemplary embodiment of the present disclosure, the encoder training module includes: a feature vector acquisition module, used to obtain a main sample feature vector and multiple offset sample feature vectors corresponding to the sample image, and fuse the main sample feature vector and the multiple offset sample feature vectors to obtain multiple feature vectors.
[0107] In an exemplary embodiment of the present disclosure, the apparatus further includes: a spatial distribution adjustment module, configured to constrain the norm of the offset sample feature vector by minimizing variance, so as to adjust the spatial distribution of the multiple sample feature vectors.
[0108] In an exemplary embodiment of the present disclosure, the device further includes: a training constraint module, configured to employ a discriminator for feature vectors and perform constraint training on the encoder through adversarial learning so that the spatial distribution of the feature vectors is consistent with the spatial distribution of the reference vectors.
[0109] The present disclosure provides an image processing device, referring to Figure 12 As shown in , the image processing apparatus 1200 may include:
[0110] The image acquisition module 1201 is used to acquire the image to be processed;
[0111] The image generation module 1202 is used to extract the feature vector of the image to be processed according to the image attribute adjustment model, and edit the feature vector to obtain an editing vector to generate an attribute image corresponding to the image to be processed; the image attribute adjustment model is trained according to any one of the model training methods described above.
[0112] In an exemplary embodiment of the present disclosure, the image generation module includes: an attribute adjustment module for adjusting a target attribute among multiple attributes in response to a movement operation of the feature vector along a normal vector corresponding to the target attribute; and a generation control module for generating the attribute image according to the target attribute.
[0113] It should be noted that the specific details of each module in the above-mentioned model training device and the above-mentioned image processing device have been described in detail in the corresponding model training method and image processing method, so they will not be repeated here.
[0114] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0115] Furthermore, the figures above are merely illustrative of the processes included in the methods according to exemplary embodiments of the present disclosure and are not intended to be limiting. It is readily understood that the processes illustrated in the figures above do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0116] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0117] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing what is disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims. It should be understood that the present disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and that various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A model training method, characterized in that: include: Train the decoder according to the sample image to obtain a trained decoder; Extracting multiple feature vectors of the sample image through an encoder, and training the encoder in combination with a reference vector spatial distribution of a trained decoder and the spatial distribution of the multiple feature vectors to obtain a trained encoder, including: aligning a portion of the reference vector space with the spatial distribution of the multiple feature vectors and minimizing a reconstruction error as a training objective, and training the encoder in combination with the trained decoder to obtain a trained encoder; Obtaining an image attribute adjustment model for performing attribute editing on an image according to the trained decoder, the trained encoder, and the attribute editing model; The step of training the encoder in combination with the trained decoder to obtain the trained encoder includes: Determining a loss function according to the training objective; wherein the loss function is determined according to minimizing the variance and minimizing the offset loss function and the reconstruction loss function; Fixing the parameters of the trained decoder, training the encoder according to the loss function corresponding to the training objective, and obtaining a trained encoder; The method further comprises: A discriminator for the feature vector is used to constrain the encoder training through adversarial learning so that the spatial distribution of the feature vector is consistent with the spatial distribution of the reference vector.
2. The model training method according to claim 1, characterized in that The step of training the decoder according to the sample image to obtain the trained decoder includes: The decoder is trained based on the sample image using an adversarial generation method to obtain a trained decoder.
3. The model training method according to claim 1, characterized in that The extracting a plurality of feature vectors of the sample image by an encoder includes: A main sample feature vector and a plurality of offset sample feature vectors corresponding to the sample image are obtained, and the main sample feature vector and the plurality of offset sample feature vectors are fused to obtain a plurality of feature vectors.
4. The model training method according to claim 3, characterized in that The method further comprises: The norm of the offset sample feature vector is constrained by minimizing the variance, so as to adjust the spatial distribution of the multiple sample feature vectors.
5. An image processing method, characterized in that: include: Get the image to be processed; The feature vector of the image to be processed is extracted according to the image attribute adjustment model, and the feature vector is edited to obtain an editing vector to generate an attribute image corresponding to the image to be processed; the image attribute adjustment model is trained according to the model training method according to any one of claims 1-4.
6. The image processing method according to claim 5, characterized in that The editing operation on the feature vector to obtain an editing vector to generate an attribute image corresponding to the image to be processed includes: In response to a movement operation on the feature vector along a normal vector corresponding to the target attribute, adjusting a target attribute among the plurality of attributes; The attribute image is generated according to the target attribute.
7. A model training device, characterized in that: include: The decoder training module is used to train the decoder according to the sample image and obtain the trained decoder; an encoder training module, configured to extract multiple feature vectors of the sample image through an encoder, and train the encoder in combination with the spatial distribution of the reference vector of the trained decoder and the spatial distribution of the multiple feature vectors to obtain a trained encoder, including: aligning a portion of the reference vector space with the spatial distribution of the multiple feature vectors and minimizing reconstruction error as a training objective, and training the encoder in combination with the trained decoder to obtain a trained encoder; A model acquisition module, configured to acquire an image attribute adjustment model for performing attribute editing on an image based on the trained decoder, the trained encoder, and the attribute editing model; The step of training the encoder in combination with the trained decoder to obtain the trained encoder includes: determining a loss function according to the training objective; wherein the loss function is determined according to minimizing the variance and minimizing the offset loss function and the reconstruction loss function; fixing the parameters of the trained decoder, and training the encoder according to the loss function corresponding to the training objective to obtain the trained encoder; The device is further configured to employ a discriminator for feature vectors to perform constraint training on an encoder in an adversarial learning manner, so that the spatial distribution of the feature vectors is consistent with the spatial distribution of the reference vectors.
8. An image processing device, characterized in that: include: An image acquisition module, used for acquiring an image to be processed; An image generation module is used to extract the feature vector of the image to be processed according to the image attribute adjustment model, and edit the feature vector to obtain an editing vector to generate an attribute image corresponding to the image to be processed; the image attribute adjustment model is trained according to the model training method according to any one of claims 1-4.
9. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the model training method described in any one of claims 1-4 or the image processing method described in any one of claims 5-6 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the model training method described in any one of claims 1 to 4 or the image processing method described in any one of claims 5 to 6.
Citation Information
Patent Citations
Pedestrian re-identification method capable of defending against adversarial attack
CN113205030A
Face editor training, face editing and live broadcasting methods and related devices
CN113255551A