Image processing method and device, electronic equipment and storage medium
By determining the model to process image data based on the target facial attributes, the problem of limb and facial distortion caused by the addition of special effects in existing technologies is solved, and more realistic and interesting special effects image generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2021-12-29
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies often cause distortion of the subject's limbs and face when adding effects to images, resulting in poor effect and a poor user experience.
By acquiring Gaussian noise or the image to be converted, the model is determined using the target facial attributes to generate a target facial image corresponding to the data to be processed, and corresponding facial features are added.
It improves the realism of special effects images and user experience, and increases the richness and interest of video images.
Smart Images

Figure CN114387373B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to an image processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of technology, more and more applications have entered users' lives, gradually enriching their leisure time, such as short video applications. Users can record their lives using videos and photos and upload them to short video applications. To further enhance the fun of video and photo content, corresponding special effects are usually added to the objects in the images.
[0003] Currently, existing techniques for adding special effects typically use software to edit images and add effects to objects, such as making an object turn its head. However, this method can easily cause distortion of the object's limbs or face, resulting in poor effect. Summary of the Invention
[0004] This disclosure provides an image processing method, apparatus, electronic device, and storage medium to add facial effects to an object, so that the generated image best matches the effect of facial changes in reality, thereby improving the accuracy of adding facial effects.
[0005] In a first aspect, embodiments of this disclosure provide an image processing method, the method comprising:
[0006] Acquire the data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted;
[0007] The data to be processed is processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed; wherein at least one target feature in the target facial image matches at least one corresponding preset facial feature.
[0008] Secondly, embodiments of this disclosure also provide an image processing apparatus, the apparatus comprising:
[0009] A data acquisition module is used to acquire data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted;
[0010] The target image determination module is used to process the data to be processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed; wherein, at least one target feature in the target facial image matches at least one corresponding preset facial feature.
[0011] Thirdly, embodiments of this disclosure also provide an electronic device, the device comprising:
[0012] One or more processors;
[0013] Storage device for storing one or more programs.
[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method as described in any of the embodiments of this disclosure.
[0015] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the image processing method as described in any of the embodiments of this disclosure.
[0016] The technical solution of this disclosure acquires Gaussian noise or image data to be converted, and then determines a model based on the target facial attributes to process the data, thereby obtaining a target facial image with added effects corresponding to the data to be processed. This solves the problem in the prior art where the special effects images generated by using image editing technology have low realism, resulting in a poor user experience. It adds corresponding facial features to the faces of each target object in the image to be processed, thereby increasing the realism of the special effects and enhancing the richness and interest of the video image content, further improving the user experience. Attached Figure Description
[0017] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0018] Figure 1 This is a schematic flowchart of an image processing method provided in Embodiment 1 of the present disclosure;
[0019] Figure 2 This is a schematic diagram of a target facial image that matches preset facial features, provided in Embodiment 1 of this disclosure;
[0020] Figure 3 This is a schematic flowchart of an image processing method provided in Embodiment 2 of this disclosure;
[0021] Figure 4 This is a schematic diagram of the target facial attribute determination model provided in Embodiment 2 of this disclosure;
[0022] Figure 5 This is a schematic flowchart of an image processing method provided in Embodiment 3 of this disclosure;
[0023] Figure 6This is a schematic diagram of the structure of the facial attribute determination model to be trained provided in Embodiment 3 of this disclosure;
[0024] Figure 7 This is a schematic flowchart of an image processing method provided in Embodiment 4 of this disclosure;
[0025] Figure 8 This is a structural block diagram of an image processing apparatus provided in Embodiment 5 of this disclosure;
[0026] Figure 9 This is a schematic diagram of the structure of an electronic device provided in Embodiment Six of this disclosure. Detailed Implementation
[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0028] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should also be noted that the modifications of "a" and "a plurality of" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0031] Before introducing this technical solution, the application scenarios can be illustrated by examples. This disclosed technical solution can be applied to any scenario that requires special effects display. For example, special effects can be displayed during video calls; or, in live streaming scenarios, special effects can be displayed for the broadcaster; of course, it can also be applied during video shooting, where special effects can be displayed on the image corresponding to the subject being filmed. For example, in short video shooting scenarios, the captured image can be processed into a special effects image, and then the processed special effects image can be displayed; it can also be applied during still image shooting, for example, after capturing an image using the camera built into a terminal device, the captured image can be processed into a special effects image for display.
[0032] Example 1
[0033] Figure 1 This is a schematic flowchart of an image processing method provided in Embodiment 1 of this disclosure. This embodiment is applicable to any image display scenario supported by the Internet, used to process the facial image of a target object into a special effects image and display it. The method can be executed by an image processing device, which can be implemented in software and / or hardware, optionally through an electronic device, such as a mobile terminal, PC, or server. Scenarios of arbitrary image display are typically implemented through cooperation between a client and a server. The method provided in this embodiment can be executed by the server, the client, or through cooperation between the client and the server.
[0034] like Figure 1 The method in this embodiment includes:
[0035] S110. Obtain the data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted.
[0036] It should be noted that the various applicable scenarios have been briefly described above, and will not be elaborated further here. The apparatus for executing the image processing method provided in this embodiment can be integrated into application software that supports image processing functions, and this software can be installed on an electronic device, optionally a mobile terminal or a PC. The application software can be a type of software for image / video processing; specific application software will not be described in detail here, as long as it can achieve image / video processing.
[0037] In this context, the data to be processed can be understood as the data that needs to be processed, which can be Gaussian noise or an image. Gaussian noise can be random sampling noise, and can include at least one type of noise such as fluctuation noise, cosmic noise, thermal noise, and shot noise. The image to be converted can be an image acquired by the application software or an image pre-stored in the storage space by the application software. In specific application scenarios, the image to be converted can be acquired in real time or periodically. For example, in live streaming or video shooting scenarios, the camera device acquires images of the target scene, including the target, in real time. In this case, the images acquired by the camera device can be used as the image to be converted. Correspondingly, the image to be converted can include the target subject, which can be a user, a pet, flowers, trees, etc. It should be noted that video frames corresponding to the captured video can also be processed. For example, a target subject corresponding to the captured video can be pre-set. When the target subject is detected in the image corresponding to the video frame, the image corresponding to that video frame can be used as the image to be converted, so that facial feature processing can be performed on the target subject in each video frame image in the video.
[0038] It should be noted that the number of target subjects in the same shooting scene can be one or more. Regardless of whether it is one or more, the technical solution provided in this disclosure can be used to determine the special effects display image.
[0039] Specifically, in any video shooting or live streaming scenario, images including the target subject can be acquired in real-time or intermittently as the images to be converted. Simultaneously, the device can randomly collect noise to obtain Gaussian noise. This Gaussian noise and the images to be converted can then be used as input to the model for training.
[0040] In this embodiment, the original facial data when it is necessary to add corresponding features to the facial image is used as the data to be processed.
[0041] S120. The data to be processed is processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed.
[0042] The target facial attribute determination model can be pre-trained. This model processes input Gaussian noise or the image to be converted, resulting in an image with added facial features. It's important to note that before training the model, the facial features to be added by the model can be pre-determined and used as preset facial features. The image output by the target facial attribute determination model can then be used as the target facial image, in which the preset facial features have been added.
[0043] In this embodiment, the preset facial features include at least one of the following: facial features with various accessories, facial features of different age groups, facial features from different angles, facial features with different hairstyles, facial features with different hairstyle colors, and facial features with different facial expressions.
[0044] The accessories worn can be glasses, sunglasses, face masks, etc., and the features corresponding to these accessories can be used as preset facial features. Facial features for different age groups can be the facial features corresponding to different age groups of the subject. Different age groups can be defined by adding or subtracting age from the original captured image as a reference, such as youth, middle-aged, and elderly age groups. Facial features at different angles can be the features corresponding to different facial orientations of the subject. Optionally, different angles can also be based on the facial angle of the subject in the original captured image, and can be rotated 20 degrees to the left, 20 degrees to the right, 20 degrees downward, etc., and the facial features corresponding to these angles can also be used as preset facial features. For example, if you want the target facial attribute determination model output to add effects such as wearing glasses, increasing age by 20 years, and rotating 20 degrees to the left, these facial features can be used as preset facial features. Preset facial features can also be facial features of different hairstyles, such as long hair, short hair, curly hair, straight hair, etc. Facial features with different hairstyles and colors can be various colors such as purple, white, and black, or combinations of different hairstyles and colors. Facial features with different expressions can include facial expressions such as smiling, anger, and rage.
[0045] Specifically, in this embodiment, the data to be processed can be input into a pre-trained target facial attribute determination model, which can add corresponding features to the image corresponding to the data to be processed, and then output a facial image with added features corresponding to the data to be processed, i.e., the target facial image.
[0046] For example, the preset facial features may include at least one of the facial features corresponding to images with the face turned 20° to the right (yaw+20°), age increased by 20 years (age+20), wearing sunglasses, smiling, the face turned 30° to the right (yaw+30°), and age decreased by 20 years (age-20). After inputting the original image (Origin) into the model, a corresponding schematic diagram can be obtained, which can be found in the following diagram. Figure 2 .
[0047] The technical solution of this disclosure acquires the data to be processed, and then determines the model based on the target facial attributes to process the data to be processed, thereby obtaining the target facial image with added effects corresponding to the data to be processed. This achieves the technical effect of adding corresponding facial features to the faces of each target object in the data to be processed, so that the obtained effects have a high degree of realism, and increases the richness and interest of the video image content, thereby further improving the user experience.
[0048] Example 2
[0049] Figure 3 This is a schematic flowchart of an image processing method provided in Embodiment 2 of this disclosure. Based on the foregoing embodiments, S120 is further refined, and the specific implementation method can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0050] like Figure 3 As shown, the method specifically includes the following steps:
[0051] S210. Obtain the data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted.
[0052] S220. Determine the feature vector to be spliced corresponding to the data to be processed.
[0053] It should be noted that the target facial attribute determination model includes a feature preprocessing sub-model, an attribute editing sub-model, and an image generation sub-model. These three sub-models process the data to be processed to obtain the target facial image corresponding to the data to be processed.
[0054] The feature preprocessing sub-model is used to extract relevant features. The attribute editing sub-model adds preset facial features to the extracted features. The image generation sub-model generates an image based on the features output by the attribute editing sub-model. The feature vector to be concatenated is the vector output by the feature preprocessing sub-model, which can be used as the feature vector to be concatenated.
[0055] Specifically, the data to be processed can be used as input to the feature preprocessing sub-model. After the data undergoes feature extraction processing by the model, the feature vector corresponding to the data to be processed can be obtained, and this feature vector can be used as the feature vector to be concatenated.
[0056] For example, a structural diagram of the target facial attribute determination model can be found here. Figure 4 The model may include a feature preprocessing sub-model, an attribute editing sub-model, and an image generation sub-model. The feature preprocessing sub-model includes a first feature extraction module and a second feature extraction module.
[0057] It should be noted that the data to be processed may include Gaussian noise or images to be converted. For the sake of accuracy, different processing methods can be used depending on the data, with each method employing its own algorithm module to target different data types. Specific processing methods are described below:
[0058] Optionally, the feature preprocessing sub-model includes a first feature extraction module and a second feature extraction module. The step of determining the feature vector to be concatenated corresponding to the data to be processed based on the feature preprocessing sub-model includes: if the data to be processed is Gaussian noise, then determining the feature vector to be concatenated corresponding to the Gaussian noise based on the first feature extraction module; if the data to be processed is the image to be converted, then determining the feature vector to be concatenated corresponding to the image to be converted based on the second feature extraction module.
[0059] The first feature extraction module is used to extract the feature vector corresponding to Gaussian noise. The second feature extraction module is used to extract the feature vector corresponding to facial attributes in the image to be converted.
[0060] It should be noted that, in this embodiment, in order to process Gaussian noise and the image to be converted separately, two feature extraction modules can be preset in the feature preprocessing sub-model to process the two types of data respectively. Accordingly, the data to be processed that is Gaussian noise can be input into the first feature extraction module, which can process the Gaussian noise and obtain the feature vector to be concatenated corresponding to the Gaussian noise; the data to be processed that is the image to be converted can be input into the second feature extraction module, which can process the image to be converted and obtain the feature vector to be concatenated corresponding to the image to be converted. Accordingly, the feature vectors to be concatenated corresponding to all the data to be processed can be obtained.
[0061] For example, see [link to example]. Figure 4 The Gaussian noise can be processed using the first feature extraction module, outputting a feature vector to be concatenated corresponding to the Gaussian noise. Optionally, the first feature extraction module can be a Mapping Network model. The image to be converted can be processed using the second feature extraction module, outputting a feature vector to be concatenated corresponding to the image to be converted. Optionally, the second feature extraction module can be an Encoder model; for example, by fixing the generator parameters of a pre-trained StyleGan model, an Encoder model can be trained. After training, a facial image can be input, encoded by the Encoder, and then passed through the StyleGan generator to reconstruct the facial image. The output feature vector to be concatenated can be used as W+ for input to subsequent models.
[0062] In practical applications, after inputting data into the model, the type of input data can be determined based on the data interface, and then the appropriate module can be selected to process it based on the data type. Specifically, if the data to be processed is Gaussian noise, the Gaussian noise can be used as input to the first feature extraction module, which can output the corresponding feature vector to be concatenated. If the data to be processed is an image to be converted, the image to be converted can be used as input to the second feature extraction module, which can output the corresponding feature vector to be concatenated.
[0063] S230. Concatenate the feature vector to be concatenated with a preset feature vector corresponding to the at least one preset facial feature to obtain a target feature vector corresponding to the target facial image.
[0064] Among them, the preset feature vector refers to the vector corresponding to the preset facial features, and the target feature vector can be understood as the feature vector after the feature vector to be concatenated is concatenated with the preset feature vector.
[0065] It should be noted that, in this embodiment, in order to add corresponding special effects to the data to be processed, the feature vector to be concatenated can be concatenated with the feature vector corresponding to the preset facial features. This yields the target feature vector after adding facial features to the data, enabling the generation of a target facial image with special effects based on the target feature vector. For example, the feature vector to be concatenated can be input into a pre-trained attribute editing sub-model. The sub-model can then concatenate the feature vector to be concatenated with the preset feature vector corresponding to the preset facial features. For instance, feature vector A and feature vector B can be concatenated to form AB. Correspondingly, the concatenated feature vector can be obtained, which can then be used as the target feature vector corresponding to the target facial image.
[0066] For example, see [link to example]. Figure 4 The attribute editing sub-model can be a Dynamic Network model. For example, a Dynamic Network model can be used to concatenate the feature vector to be concatenated with a preset feature vector, and the output W++ is the target feature vector corresponding to the target facial image. For instance, the input preset feature vector is encoded by a multilayer perceptron (MLP), then passed through two fully connected layers (FC) and an activation function sigmoid, and multiplied and added with the input feature vector to be concatenated to obtain the target feature vector. The identifier corresponding to the preset feature vector can be set in the Dynamic Network model.
[0067] Specifically, the feature vector to be spliced can be used as input to the attribute editing sub-model. The sub-model can splice the feature vector to be spliced with at least one preset feature vector corresponding to a preset facial feature, thereby obtaining the target feature vector corresponding to the target facial image with the preset facial features added, so that a target facial image with special effects can be generated based on the target feature vector in the future.
[0068] S240. Process the target feature vector to obtain the target facial image.
[0069] In this embodiment, in order to generate an image with corresponding effects added to the data to be processed, the target feature vector corresponding to the spliced feature vector and the preset feature vector can be input into a pre-trained image generation sub-model. The model can reconstruct the target feature vector and output the target facial image corresponding to the target feature vector.
[0070] For example, see [link to example]. Figure 4 The image generation sub-model can be a Generator model. For example, the Generator model can be used to process the target feature vector and output the target facial image corresponding to the target feature vector.
[0071] It should be noted that after obtaining the target facial image, in order to further optimize the parameters of the target facial attribute determination model based on the original image input to the model and the target facial image output by the model, an optional pre-trained model attribute classifier can be added. This attribute classifier can be a model used to extract and classify image attribute features. Furthermore, this attribute classifier can be used to extract feature data from the target facial image to determine the facial features in the target facial image. This allows for further verification of whether the model output image has added preset facial features. If not, the parameters in the model can be further corrected to improve the model's accuracy.
[0072] Optionally, after obtaining the target facial image, the method further includes: determining at least one target attribute corresponding to the target facial image based on a pre-trained attribute classifier, so as to correct the model parameters in the target facial attribute determination model based on the at least one target attribute; wherein the at least one target attribute matches the attribute identifier of the at least one preset facial feature.
[0073] Among them, the attribute identifier can refer to the identifier corresponding to the preset facial features. That is to say, the preset facial features can be represented by the corresponding identifiers. For example, age feature is represented by A1, angle feature by A2, and wearing glasses feature by A3. Then, identifiers such as A1, A2, and A3 can be used as attribute identifiers of the corresponding features. The target attribute can also be the identifier information corresponding to the preset facial features.
[0074] It should be noted that, in this embodiment, in order to further optimize the parameters in the target facial attribute determination model, the output target facial image can be compared with the corresponding theoretical facial image with added effects. Accordingly, the attribute features corresponding to the output target facial image can also be compared with the added effect features, so as to correct the model parameters in the classification model to be trained based on whether the attribute features corresponding to the target facial image contain the added effect features.
[0075] It should also be noted that the target facial image can be input into a pre-trained attribute classifier, which can then perform feature extraction on the target facial image and output the corresponding feature attributes, i.e., the target attributes. Further, the currently output feature attributes are compared with the attribute labels of preset facial features, and the comparison error value is calculated. The model parameters in the model can then be adjusted based on the comparison error value.
[0076] For example, see [link to example]. Figure 4 The attribute classifier can be a ResNet model. For example, the target facial image can be input into the attribute classifier. After the image is processed by the feature extraction of the classifier, the target attributes corresponding to the target facial image can be output.
[0077] Specifically, the target facial image can be used as input to an attribute classifier, which can then output the corresponding target attributes. Furthermore, an algorithm can be used to perform error processing between the target attributes corresponding to the image and the attribute labels of preset facial features. The model parameters of the target facial attribute determination model can then be corrected based on the obtained error results, further improving the training accuracy of the target facial attribute determination model.
[0078] The technical solution of this disclosure acquires Gaussian noise or images to be converted and processed data, and then determines a model based on the target facial attributes to process the data, thereby obtaining target facial images with added effects corresponding to different types of data. After obtaining the target facial images, the parameters in the model are further corrected based on the target facial images, which improves the accuracy of the model, thereby improving the realism of the added effects, as well as increasing the richness and interest of the video image content, and further improving the technical effect of the user experience.
[0079] Example 3
[0080] Figure 5 This is a flowchart illustrating an image processing method provided in Embodiment 3 of this disclosure. Based on the foregoing embodiments, a facial attribute determination model to be trained can be pre-constructed, and training processing can be performed on this model to obtain a target facial attribute determination model. Specific implementation details can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0081] like Figure 5 As shown, the method specifically includes the following steps:
[0082] S310. Construct a facial attribute determination model to be trained.
[0083] Among them, the facial attribute determination model to be trained refers to the model whose model parameters are set to default values, and the model that needs to be trained is the model obtained from this model.
[0084] In this example, in order to obtain a high-precision target facial attribute determination model that determines facial attribute features, it can be trained based on a pre-built training facial attribute determination model. After the training of the training facial attribute determination model is completed, the final applicable facial attribute determination model, i.e., the target facial attribute determination model, can be obtained.
[0085] In this embodiment, constructing the facial attribute determination model to be trained includes: constructing the facial attribute determination model to be trained based on the attribute editing sub-model to be trained, a pre-trained target adversarial model, a target attribute classification model, and a facial matching model; wherein, the target attribute classification model is used to determine the facial features of the image output by the target adversarial model; the facial matching model is used to determine the matching degree of the facial image output by the target adversarial model; the target adversarial model is used to output two facial images, one of which is consistent with the preset facial features set in the attribute editing sub-model to be trained.
[0086] The target adversarial model includes a feature preprocessing sub-model and an image generation sub-model; the feature preprocessing sub-model includes a first feature extraction module; the image generation sub-model includes a first image generation sub-module and a second image generation sub-module; and the feature preprocessing sub-model further includes a second feature extraction module.
[0087] The target adversarial model can be a StyleGAN model. The target attribute classification model (ResNet model) is used to determine the facial features of the image output by the second image generation submodule. The face matching model (face recognition model) is used to determine the matching degree between the facial images output by the first and second image generation submodules. The first feature extraction module (Mapping Network model) is used to extract the feature vector corresponding to Gaussian noise. The second feature extraction module (Encoder model) is used to extract the feature vector corresponding to facial attributes in the image. The first image generation submodule (Generator model) is used to determine the image corresponding to the feature vector output by the feature preprocessing submodel. The second image generation submodule (Generator model) is used to determine the image corresponding to the feature vector output by the attribute editing submodel to be trained. The attribute editing submodel to be trained (Dynamic Network model) is used to concatenate the feature vector output by the feature preprocessing submodel with a preset feature vector. This can also be understood as the attribute editing sub-model that needs to be trained. At this stage, the output of the attribute editing sub-model may not yet meet the expected results, so it needs to be trained to ensure that the output of the trained model matches the expected results. After training, an applicable attribute editing sub-model can be obtained. For example, an image code (w code) can be obtained through random noise or an image encoder. Two StyleGAN generator branches are designed. One branch directly generates a face image img1 through a pre-trained StyleGAN generator; the other branch passes through an attribute editing sub-model to be trained (the input to the editing module is a specific attribute value, such as 1 for wearing glasses, 0 for not wearing glasses, 1 for smiling, and 0 for not smiling), and then through the pre-trained StyleGAN generator to generate an attribute-edited face image img2. img1 and img2 are then fed into a pre-trained target attribute classification model and a face matching model, so that the difference between the attributes and IDs of the two images can be used as the loss to train the attribute editing sub-model to be trained.
[0088] It should be noted that before constructing the facial attribute determination model to be trained, a large number of facial images can be collected and their attributes labeled, such as whether glasses are worn. A ResNet model can then be used to train an attribute classifier to obtain a target attribute classification model. Additionally, a pre-trained facial recognition model can be developed to obtain a facial matching model.
[0089] Specifically, in combination Figure 6To explain, training data can be input into the feature preprocessing sub-model (either the first or second feature extraction module) to obtain the corresponding feature vector. This feature vector is then used as input to the attribute editing sub-model to obtain a concatenated feature vector with a preset feature vector. This processed feature vector is then used as input to the second image generation sub-module, outputting an effect face image. This effect face image may deviate slightly from the preset facial features. The feature vector is then used as input to the first image generation sub-module to obtain the corresponding face image. Finally, the effect face image and the face image are used as input to the target attribute classification model and face matching model, respectively, to obtain the feature differences and face matching degree between the two images.
[0090] S320. By training the facial attribute determination model to be trained, a facial attribute determination model to be used is obtained.
[0091] Among them, the facial attribute determination model to be used refers to the model that has been trained to determine the facial attributes of the model to be trained.
[0092] Specifically, after obtaining the sample data, the facial attribute determination model to be trained is trained using the sample data to obtain the facial attribute determination model to be used.
[0093] S330. The target facial attribute determination model is obtained by cropping the facial attribute determination model to be used.
[0094] In this embodiment, in order to reduce the redundancy of the model, some modules that are not actually used in the facial attribute determination model can be pruned, and the model obtained after pruning can be used as the target facial attribute determination model.
[0095] Specifically, the target attribute classification model, the face matching model, and the first image generation sub-model are removed from the facial attribute determination model to be used, resulting in the target facial attribute determination model.
[0096] S340. Obtain the data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted.
[0097] S350. The data to be processed is processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed.
[0098] The technical solution of this disclosure involves acquiring Gaussian noise or image data to be converted, then processing the data based on a target facial attribute determination model to obtain a target facial image with added effects corresponding to the data to be processed. Simultaneously, a training facial attribute determination model is constructed based on different types of models, and a high-precision model is obtained by training the constructed model. The model is then cropped to obtain the target facial attribute determination model, thereby improving the accuracy of the target facial attribute determination model and the efficiency of adding effects, ultimately resulting in higher realism of the added effects to the image.
[0099] Example 4
[0100] Figure 7 This is a flowchart illustrating an image processing method provided in Embodiment 4 of this disclosure. Based on the foregoing embodiments, the step of training the face attribute determination model to obtain the face attribute determination model to be used is further refined. Specific implementation details can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0101] like Figure 7 As shown, the method specifically includes the following steps:
[0102] S410: Obtain multiple training samples.
[0103] It should be noted that before training the facial attribute determination model, training samples need to be obtained first to train the model. To improve the model's accuracy, as many and varied training samples as possible should be obtained.
[0104] The training samples include training data. Training samples can be used to train the model, allowing the model's parameters to be adjusted during training to match the expected output. Training data can be data used to train the model, such as Gaussian noise (e.g., Gaussian sampled random noise). It can also be images; for example, images of the user subject taken from different viewing angles can generate facial images corresponding to those angles, which can be used as training data. Thus, multiple training samples can be obtained. For instance, in practical applications, a large number of facial images can be collected as training samples to train a StyleGan model. After training, the StyleGan generator can generate different types of facial images by inputting Gaussian sampled random noise z ~ N(0,1).
[0105] Specifically, each training sample can be stored in a pre-defined database, and then the training samples in the database can be extracted using an interface.
[0106] S420. For each training sample, input the training data in the current training sample into the first feature extraction module or the second feature extraction module to obtain the first feature vector corresponding to the current training sample.
[0107] When it is necessary to determine the feature vector corresponding to each training sample, the feature vector of any training sample can be used as the feature vector of the current training sample for processing, thus illustrating that one of the training samples is the current training sample. The first feature vector refers to the feature vector output by the first feature extraction module or the second feature extraction module. For example, after the current training sample is input into the first feature extraction module or the second feature extraction module, the feature vector extracted by the module can be used as the first feature vector.
[0108] It should be noted that since the training data in the training samples may vary (it could be Gaussian noise or an image), the training data can be processed differently based on the first and second feature extraction modules. For example, the training data in the current training samples can be input into the feature preprocessing sub-model. If the training data is Gaussian noise, it can be processed using the first feature extraction module to obtain the corresponding feature vector. If the training data is an image, it can be processed using the second feature extraction module to obtain the corresponding feature vector. The feature vectors output by both the first and second feature extraction modules can be used as the first feature vector corresponding to the current training sample.
[0109] S430. Based on the attribute editing sub-model to be trained, the first attribute feature vector is concatenated with the first feature vector and attribute feature vectors corresponding to at least one preset facial feature to obtain the first attribute feature vector.
[0110] The attribute feature vector can be a vector representation of a preset facial feature, or any feature vector corresponding to a preset facial feature can be used as the attribute feature vector. The first attribute feature vector can be understood as the feature vector obtained by concatenating the first feature vector and the attribute feature vector, and the vector output by the attribute editing sub-model to be trained can be used as the first attribute feature vector.
[0111] Specifically, the first feature vector corresponding to the current training sample can be used as the input to the attribute editing sub-model to be trained. The sub-model can concatenate the first feature vector with the attribute feature vector corresponding to the preset facial features, and then output the first attribute feature vector after adding the preset facial features, so that the model parameters can be adjusted after generating facial images with special effects based on the first attribute feature vector.
[0112] S440. Input the first feature vector into the first image generation submodule to obtain an image without attribute features; and input the first attribute feature vector into the second image generation submodule to obtain an image with attribute features.
[0113] In this context, the image without attribute features can be understood as a facial image without special effects, and the image output by the first image generation submodule can be used as the image without attribute features. The image with attribute features can be understood as a facial image with special effects, and the image output by the second image generation submodule can be used as the image with attribute features.
[0114] Specifically, the first feature vector corresponding to the current training sample can be used as input to the first image generation submodule. The model can reconstruct the image from the first feature vector. At this time, the first feature vector has not been processed by the attribute editing model to be trained, and no features have been added. The model can output an image without attribute features. At the same time, the first attribute feature vector corresponding to the current training sample can also be input into the second image generation submodule. Because the first attribute feature vector is a vector after adding effects, the model can output an image with attribute features.
[0115] S450. Input the attached attribute feature image and the unattached attribute feature image into the target attribute classification model to obtain the attribute information to be compared; and input the attached attribute feature image and the unattached attribute feature image into the face matching model to obtain face matching information.
[0116] The attribute information to be compared can be understood as the facial features of the image output by the target attribute classification model. The facial matching information can be understood as the facial image matching degree output by the facial matching model.
[0117] Specifically, the image with and without attribute features corresponding to the current training sample can be used as input to the target attribute classification model, which can output the facial features corresponding to the two images, i.e., the attribute information to be compared. Alternatively, the image with and without attribute features can be used as input to the face matching model, which can output the matching degree between the two images, i.e., the face matching information.
[0118] S460. Based on the loss function in the attribute editing sub-model to be trained, the attribute information to be compared, the facial matching information, and the preset facial features in the attribute editing sub-model to be trained are processed to obtain the target loss value.
[0119] The target loss value can be used to characterize the loss between the attribute information to be compared and the preset facial features, as well as the loss between the facial matching information.
[0120] Specifically, the loss function in the attribute editing sub-model can be used to process the loss between the comparison attribute information and the preset facial features, thereby calculating the loss value between the two. The loss function can also be used to process the loss of facial matching information, thereby calculating the corresponding loss value. Correspondingly, all calculated loss values can be fused to obtain a fused loss value, which can be used as the target loss value to adjust the model parameters.
[0121] S470. Based on the target loss value, the model parameters in the attribute editing sub-model to be trained are corrected, and the convergence of the loss function is taken as the training objective to train and obtain the facial attribute determination model to be used.
[0122] Among them, the convergence of the preset loss function can be used as the training objective. When it is determined that the preset loss function of the sub-model for editing attributes to be trained has converged, it indicates that the adjustment result meets the requirements of the scheme and the trained model has been obtained, thus obtaining the facial attribute determination model to be used.
[0123] In this embodiment, the first feature vector corresponding to the current training sample can be processed using attribute feature addition technology. The attribute editing sub-model to be trained can output the first attribute feature vector corresponding to the current training sample, so as to generate the attached attribute feature image corresponding to the current training sample based on the first attribute feature vector. Since the model parameters in the attribute editing sub-model to be trained are uncorrected, the resulting attached attribute feature image will also have corresponding differences from the image without attached attribute features corresponding to the current training sample after the actual addition of features. The error value can be determined by processing the comparison attribute information, facial matching information, and preset facial features corresponding to the two types of images corresponding to the current training sample. Then, the model parameters in the attribute editing sub-model to be trained can be corrected based on the error value.
[0124] Specifically, the loss function in the attribute editing sub-model to be trained can be used to compare the attribute information to be compared with the preset attribute values of the current training sample, and the loss value can be calculated. The similarity error value corresponding to the facial matching information can also be calculated. Then, based on the loss value and the similarity error value, the target loss value can be calculated to correct the model parameters of the attribute editing sub-model to be trained. Furthermore, the training error of the loss function, i.e., the loss parameter, can be used as a condition to detect whether the loss function has reached convergence. For example, whether the training error is less than the preset error or whether the error change trend is stable, or whether the current iteration number is equal to the preset number. If the convergence condition is met, such as the training error of the loss function being less than the preset error or the error change trend being stable, it indicates that the attribute editing sub-model to be trained has completed training, and iterative training can be stopped. If the convergence condition is not met, further training sample data can be obtained to continue training the attribute editing sub-model to be trained until the training error of the loss function is within the preset range. When the training error of the loss function converges, the attribute editing sub-model to be trained can be considered to be trained well. When the first feature vector is input into the trained attribute editing sub-model, the model can accurately concatenate the attribute feature vector into the first feature vector, so as to generate an image with facial features.
[0125] S480. Remove the target attribute classification model, the face matching model, and the first image generation sub-model from the face attribute determination model to be used to obtain the target face attribute determination model.
[0126] In this embodiment, the determined facial attribute determination model includes not only a feature preprocessing sub-model and a second image generation sub-module, but also a first image generation sub-module, a target attribute classification model, and a face matching model. To achieve fast data processing and generate effective facial images with low computational requirements on terminal devices, the first image generation sub-module, the target attribute classification model, and the face matching model in the facial attribute determination model can be removed. That is, models that will not be used in the application are removed. For example, after training the facial attribute determination model, only the attribute editing sub-model branch is retained, and the edited facial image is generated by inputting attribute values.
[0127] S490. Obtain the data to be processed; wherein the data to be processed includes Gaussian noise or an image to be converted.
[0128] S4100. The data to be processed is processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed.
[0129] The technical solution of this disclosure involves acquiring Gaussian noise or image data to be converted, then processing the data based on a target facial attribute determination model to obtain a target facial image with added effects corresponding to the data to be processed. Simultaneously, the model to be trained is trained using training samples to continuously optimize the model parameters in the attribute editing sub-model to be trained, thereby obtaining a facial attribute determination model to be used. The facial attribute determination model to be used is then cropped to obtain the target facial attribute determination model, thereby improving the accuracy of the target facial attribute determination model and the efficiency of adding effects, and thus making the effects added to the image more realistic.
[0130] Example 5
[0131] Figure 8 This is a structural block diagram of an image processing apparatus provided in Embodiment 5 of this disclosure. It can execute the image processing method provided in any embodiment of this disclosure and possesses the corresponding functional modules and beneficial effects for executing the method. For example... Figure 8 As shown, the device specifically includes a data acquisition module 510 and a target image determination module 520.
[0132] The data acquisition module 510 is used to acquire data to be processed, including Gaussian noise or an image to be converted. The target image determination module 520 is used to process the data to be processed based on a target facial attribute determination model to obtain a target facial image corresponding to the data to be processed. At least one target feature in the target facial image matches at least one corresponding preset facial feature.
[0133] Based on the above technical solutions, the target image determination module 520 includes a feature vector determination unit to be stitched, a target feature vector acquisition unit, and a target facial image acquisition unit.
[0134] The feature vector to be spliced unit is used to determine the feature vector to be spliced corresponding to the data to be processed;
[0135] The target feature vector acquisition unit is used to concatenate the feature vector to be concatenated with a preset feature vector corresponding to the at least one preset facial feature, so as to obtain a target feature vector corresponding to the target facial image.
[0136] The target facial image acquisition unit is used to process the target feature vector to obtain the target facial image.
[0137] Based on the above technical solutions, the feature preprocessing sub-model includes a first feature extraction module and a second feature extraction module, and a feature vector determination unit to be spliced, including a first unit to determine the feature vector to be spliced and a second unit to determine the feature vector to be spliced.
[0138] The first unit for determining the feature vector to be spliced is used to determine the feature vector to be spliced corresponding to the Gaussian noise based on the first feature extraction module if the data to be processed is the Gaussian noise.
[0139] The second unit for determining the feature vector to be spliced is used to determine the feature vector to be spliced corresponding to the image to be converted based on the second feature extraction module if the data to be processed is the image to be converted.
[0140] Based on the above technical solutions, the device further includes: a target facial attribute determination model parameter correction module.
[0141] The target facial attribute determination model parameter correction module is used to determine at least one target attribute corresponding to the target facial image based on a pre-trained attribute classifier, so as to correct the model parameters in the target facial attribute determination model based on the at least one target attribute; wherein, the at least one target attribute matches the attribute identifier of the at least one preset facial feature.
[0142] Based on the above technical solutions, the device further includes: a module for constructing a facial attribute determination model to be trained.
[0143] The training facial attribute determination model construction module is used to edit sub-models based on the training attributes, and to construct the training facial attribute determination model by pre-training the target adversarial model, target attribute classification model and face matching model.
[0144] The target attribute classification model is used to determine the facial features of the image output by the target adversarial model; the face matching model is used to determine the matching degree of the face image output by the target adversarial model; the target adversarial model is used to output two face images, one of which is consistent with the preset face features set in the attribute editing sub-model to be trained.
[0145] Based on the above technical solutions, the target adversarial model includes: a feature preprocessing sub-model and an image generation sub-model; the feature preprocessing sub-model includes a first feature extraction module; the image generation sub-model includes a first image generation sub-module and a second image generation sub-module; the feature preprocessing sub-model also includes a second feature extraction module.
[0146] Based on the above technical solutions, the training facial attribute determination model construction module further includes a training facial attribute determination model construction unit.
[0147] The face attribute determination model construction unit is used to take the output of the first feature extraction module or the second feature extraction module as the input of the attribute editing sub-model to be trained and the first image generation sub-module, respectively; take the output of the attribute editing sub-model to be trained as the input of the second image generation sub-module; and take the output of the first image generation sub-module and the output of the second image generation sub-module as the input of the target attribute classification model and the face matching model, so as to construct the face attribute determination model to be trained.
[0148] Based on the above technical solutions, the facial attribute determination model acquisition unit further includes a training sample acquisition subunit, a first feature vector acquisition subunit, a first attribute feature vector acquisition subunit, an attribute feature image acquisition subunit, an information acquisition subunit, a target loss value determination subunit, and a facial attribute determination model acquisition subunit.
[0149] The training sample acquisition subunit is used to acquire multiple training samples, which include the data to be trained.
[0150] The first feature vector acquisition subunit is used to input the training data in the current training sample into the first feature extraction module or the second feature extraction module to obtain the first feature vector corresponding to the current training sample for each training sample.
[0151] The first attribute feature vector acquisition sub-unit is used to edit the sub-model based on the attribute to be trained, and then concatenate the first feature vector with an attribute feature vector corresponding to at least one preset facial feature to obtain the first attribute feature vector.
[0152] The attribute feature image acquisition subunit is used to input the first feature vector into the first image generation submodule to obtain an image without attribute features; and to input the first attribute feature vector into the second image generation submodule to obtain an image with attribute features.
[0153] The information acquisition subunit is used to input the attached attribute feature image and the unattached attribute feature image into the target attribute classification model to obtain the attribute information to be compared; and to input the attached attribute feature image and the unattached attribute feature image into the face matching model to obtain face matching information.
[0154] The target loss value determination subunit is used to process the comparison attribute information, facial matching information, and preset attribute values in the attribute editing submodel based on the loss function in the attribute editing submodel to be trained, in order to obtain the target loss value.
[0155] The facial attribute determination model obtains a sub-unit, which is used to correct the model parameters in the attribute editing sub-model to be trained based on the target loss value, and uses the convergence of the loss function as the training objective to train the facial attribute determination model to be used.
[0156] Based on the above technical solutions, the target facial attribute determination model acquisition unit includes a target facial attribute determination model acquisition subunit.
[0157] The target facial attribute determination model acquisition subunit is used to remove the target attribute classification model, the face matching model, and the second image generation submodel from the facial attribute determination model to be used, so as to obtain the target facial attribute determination model.
[0158] Based on the above technical solutions, the preset facial features include at least one of the following: facial features with various accessories, facial features of different age groups, facial features from different angles, facial features with different hairstyles, facial features with different hairstyle colors, and facial features with different expressions.
[0159] The technical solution of this disclosure acquires Gaussian noise or image data to be converted, and then determines a model based on the target facial attributes to process the data, thereby obtaining a target facial image with added effects corresponding to the data to be processed. This solves the problem in the prior art where the special effects images generated by using image editing technology have low realism, resulting in a poor user experience. It adds corresponding facial features to the faces of each target object in the image to be processed, thereby increasing the realism of the special effects and enhancing the richness and interest of the video image content, further improving the user experience.
[0160] The image processing apparatus provided in this disclosure can execute the image processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0161] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0162] Example 6
[0163] Figure 9 This is a schematic diagram of the structure of an electronic device provided in Embodiment Six of this disclosure. Refer to the following... Figure 9 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 9 The diagram below shows the structure of the terminal device or server 600. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0164] like Figure 9 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0165] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0166] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0167] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0168] The electronic device provided in this embodiment and the image processing method provided in the above embodiments belong to the same disclosed concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0169] Example 7
[0170] Embodiment 7 of this disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the image processing method provided in the above embodiments.
[0171] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0172] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0173] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0174] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0175] Acquire the data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted;
[0176] The data to be processed is processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed; wherein at least one target feature in the target facial image matches at least one corresponding preset facial feature.
[0177] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0178] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0179] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0180] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0181] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0182] According to one or more embodiments of this disclosure, [Example 1] provides an image processing method, the method comprising:
[0183] Acquire the data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted;
[0184] The data to be processed is processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed; wherein at least one target feature in the target facial image matches at least one corresponding preset facial feature.
[0185] According to one or more embodiments of this disclosure, [Example 2] provides an image processing method, further comprising:
[0186] Optionally, the step of processing the data to be processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed includes:
[0187] Determine the feature vector to be concatenated corresponding to the data to be processed;
[0188] The feature vector to be spliced is spliced with a preset feature vector corresponding to the at least one preset facial feature to obtain a target feature vector corresponding to the target facial image;
[0189] The target feature vector is processed to obtain the target facial image.
[0190] According to one or more embodiments of this disclosure, [Example 3] provides an image processing method, further comprising:
[0191] Optionally, determining the feature vector to be concatenated corresponding to the data to be processed includes:
[0192] If the data to be processed is Gaussian noise, then the feature vector to be spliced corresponding to the Gaussian noise is determined based on the first feature extraction module;
[0193] If the data to be processed is the image to be converted, then based on the second feature extraction module, the feature vector to be spliced corresponding to the image to be converted is determined.
[0194] According to one or more embodiments of this disclosure, [Example 4] provides an image processing method, further comprising:
[0195] Optionally, based on a pre-trained attribute classifier, at least one target attribute corresponding to the target facial image is determined, so as to correct the model parameters in the target facial attribute determination model based on the at least one target attribute;
[0196] Wherein, the at least one target attribute matches the attribute identifier of the at least one preset facial feature.
[0197] According to one or more embodiments of this disclosure, [Example 5] provides an image processing method, further comprising:
[0198] Optionally, construct a model to determine the facial attributes to be trained;
[0199] By training the facial attribute determination model to be trained, a facial attribute determination model to be used is obtained.
[0200] The target facial attribute determination model is obtained by cropping the facial attribute determination model to be used.
[0201] According to one or more embodiments of this disclosure, [Example Six] provides an image processing method, further comprising:
[0202] Optionally, constructing the facial attribute determination model to be trained includes:
[0203] Based on the sub-model edited according to the attributes to be trained, the pre-trained target adversarial model, target attribute classification model and face matching model are used to construct the face attribute determination model to be trained.
[0204] The target attribute classification model is used to determine the facial features of the image output by the target adversarial model; the face matching model is used to determine the matching degree of the face image output by the target adversarial model; the target adversarial model is used to output two face images, one of which is consistent with the preset face features set in the attribute editing sub-model to be trained.
[0205] According to one or more embodiments of this disclosure, [Example Seven] provides an image processing method, further comprising:
[0206] Optionally, the target adversarial model includes: a feature preprocessing sub-model and an image generation sub-model; the feature preprocessing sub-model includes a first feature extraction module; the image generation sub-model includes a first image generation sub-module and a second image generation sub-module; the feature preprocessing sub-model further includes a second feature extraction module.
[0207] According to one or more embodiments of this disclosure, [Example Eight] provides an image processing method, further comprising:
[0208] Optionally, constructing the facial attribute determination model to be trained includes: using the output of the first feature extraction module or the second feature extraction module as the input of the attribute editing sub-model to be trained and the first image generation sub-module, respectively; using the output of the attribute editing sub-model to be trained as the input of the second image generation sub-module; and using the output of the first image generation sub-module and the output of the second image generation sub-module as the input of the target attribute classification model and the face matching model, so as to construct the facial attribute determination model to be trained.
[0209] According to one or more embodiments of this disclosure, [Example Nine] provides an image processing method, further comprising:
[0210] Optionally, the step of training the facial attribute determination model to obtain the facial attribute determination model to be used includes:
[0211] Obtain multiple training samples, where the training samples include the data to be trained;
[0212] For each training sample, the training data in the current training sample is input into the first feature extraction module or the second feature extraction module to obtain the first feature vector corresponding to the current training sample;
[0213] Based on the attribute editing sub-model to be trained, the first attribute feature vector is obtained by concatenating the first feature vector with an attribute feature vector corresponding to at least one preset facial feature.
[0214] The first feature vector is input into the first image generation submodule to obtain an image without attribute features; and the first attribute feature vector is input into the second image generation submodule to obtain an image with attribute features.
[0215] The attached attribute feature image and the unattached attribute feature image are input into the target attribute classification model to obtain the attribute information to be compared; and the attached attribute feature image and the unattached attribute feature image are input into the face matching model to obtain face matching information.
[0216] Based on the loss function in the sub-model for editing the attribute to be trained, the information of the attribute to be compared, the facial matching information, and the preset facial features in the sub-model for editing the attribute to be trained are processed to obtain the target loss value.
[0217] Based on the target loss value, the model parameters in the attribute editing sub-model to be trained are corrected, and the convergence of the loss function is taken as the training objective to train the facial attribute determination model to be used.
[0218] According to one or more embodiments of this disclosure, [Example 10] provides an image processing method, further comprising:
[0219] Optionally, obtaining the target facial attribute determination model by cropping the facial attribute determination model to be used includes:
[0220] The target facial attribute determination model is obtained by removing the target attribute classification model, the facial matching model, and the first image generation sub-model from the facial attribute determination model to be used.
[0221] According to one or more embodiments of this disclosure, [Example Eleven] provides an image processing method, further comprising:
[0222] Optionally, the preset facial features include at least one of the following: facial features with various accessories, facial features of different age groups, facial features from different angles, facial features with different hairstyles, facial features with different hairstyle colors, and facial features with different expressions.
[0223] According to one or more embodiments of this disclosure, [Example Twelve] provides an image processing apparatus, the apparatus comprising:
[0224] A data acquisition module is used to acquire data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted;
[0225] The target image determination module is used to process the data to be processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed; wherein, at least one target feature in the target facial image matches at least one corresponding preset facial feature.
[0226] Optionally, the above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0227] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0228] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An image processing method, characterized in that, include: Acquire the data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted; The data to be processed is processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed; wherein at least one target feature in the target facial image matches at least one corresponding preset facial feature; The target facial attribute determination model is obtained based on the following method: Based on the attribute editing sub-model to be trained, a pre-trained target adversarial model, target attribute classification model, and face matching model are used to construct a face attribute determination model to be trained. The target attribute classification model is used to determine the facial features of the image output by the target adversarial model. The face matching model is used to determine the matching degree of the face image output by the target adversarial model. The target adversarial model is used to output two face images, one of which matches the preset face features set in the attribute editing sub-model to be trained. Obtain multiple training samples, where the training samples include the data to be trained; For each training sample, the training data in the current training sample is input into the first feature extraction module or the second feature extraction module to obtain the first feature vector corresponding to the current training sample; Based on the attribute editing sub-model to be trained, the first attribute feature vector is obtained by concatenating the first feature vector with an attribute feature vector corresponding to at least one preset facial feature. The first feature vector is input into the first image generation submodule to obtain an image without attribute features; and the first attribute feature vector is input into the second image generation submodule to obtain an image with attribute features. The attached attribute feature image and the unattached attribute feature image are input into the target attribute classification model to obtain the attribute information to be compared; and the attached attribute feature image and the unattached attribute feature image are input into the face matching model to obtain face matching information. The target loss value is obtained by processing the attribute information to be compared, the facial matching information, and the preset facial features in the attribute editing sub-model to be trained based on the loss function in the attribute editing sub-model to be trained. Based on the target loss value, the model parameters in the attribute editing sub-model to be trained are corrected, and the convergence of the loss function is taken as the training objective to train and obtain the facial attribute determination model to be used. The target facial attribute determination model is obtained by cropping the facial attribute determination model to be used.
2. The method according to claim 1, characterized in that, The process of processing the data to be processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed includes: Determine the feature vector to be concatenated corresponding to the data to be processed; The feature vector to be spliced is spliced with a preset feature vector corresponding to the at least one preset facial feature to obtain a target feature vector corresponding to the target facial image; The target feature vector is processed to obtain the target facial image.
3. The method according to claim 2, characterized in that, The step of determining the feature vector to be concatenated corresponding to the data to be processed includes: If the data to be processed is Gaussian noise, then the feature vector to be spliced corresponding to the Gaussian noise is determined based on the first feature extraction module; If the data to be processed is the image to be converted, then based on the second feature extraction module, the feature vector to be spliced corresponding to the image to be converted is determined.
4. The method according to claim 1, characterized in that, After obtaining the target facial image, the process also includes: Based on a pre-trained attribute classifier, at least one target attribute corresponding to the target facial image is determined, and the model parameters in the target facial attribute determination model are corrected based on the at least one target attribute. Wherein, the at least one target attribute matches the attribute identifier of the at least one preset facial feature.
5. The method according to claim 1, characterized in that, The target adversarial model includes: a feature preprocessing sub-model and an image generation sub-model; the feature preprocessing sub-model includes a first feature extraction module; the image generation sub-model includes a first image generation sub-module and a second image generation sub-module; the feature preprocessing sub-model also includes a second feature extraction module.
6. The method according to claim 5, characterized in that, The step of constructing the facial attribute determination model to be trained includes: using the output of the first feature extraction module or the second feature extraction module as the input of the attribute editing sub-model to be trained and the first image generation sub-module, respectively; using the output of the attribute editing sub-model to be trained as the input of the second image generation sub-module; and using the output of the first image generation sub-module and the output of the second image generation sub-module as the input of the target attribute classification model and the face matching model, so as to construct the facial attribute determination model to be trained.
7. The method according to claim 1, characterized in that, The step of obtaining the target facial attribute determination model by cropping the facial attribute determination model to be used includes: The target facial attribute determination model is obtained by removing the target attribute classification model, the facial matching model, and the first image generation sub-model from the facial attribute determination model to be used.
8. The method according to any one of claims 1-7, characterized in that, The preset facial features include at least one of the following: facial features with various accessories, facial features of different age groups, facial features from different angles, facial features with different hairstyles, facial features with different hairstyle colors, and facial features with different expressions.
9. An image processing apparatus, characterized in that, include: A data acquisition module is used to acquire data to be processed; wherein, the data to be processed includes Gaussian noise or an image to be converted; The target image determination module is used to process the data to be processed based on the target facial attribute determination model to obtain a target facial image corresponding to the data to be processed; wherein, at least one target feature in the target facial image matches at least one corresponding preset facial feature, and the at least one preset facial feature is a facial feature added to the user by the target facial attribute determination model before training the target facial attribute determination model; The image processing device further includes: a model construction module for determining facial attributes to be trained, used to construct a model for determining facial attributes to be trained based on a sub-model for editing the attributes to be trained, a pre-trained target adversarial model, a target attribute classification model, and a face matching model; the target attribute classification model is used to determine the facial features of the image output by the target adversarial model; the face matching model is used to determine the matching degree of the facial image output by the target adversarial model, the target adversarial model is used to output two facial images, one of which is consistent with the preset facial features set in the sub-model for editing the attributes to be trained; acquiring multiple training samples, wherein the training samples include data to be trained; for each training sample, inputting the data to be trained in the current training sample into a first feature extraction module or a second feature extraction module to obtain a first feature vector corresponding to the current training sample; concatenating the first feature vector with an attribute feature vector corresponding to at least one preset facial feature based on the sub-model for editing the attributes to be trained to obtain a first attribute feature vector; inputting the first feature vector into a first image generation sub-module to obtain an image without attribute features; and inputting the first attribute feature vector into a second image generation sub-module to obtain an image with attribute features. The attached attribute feature image and the unattached attribute feature image are input into a target attribute classification model to obtain the attribute information to be compared; and the attached attribute feature image and the unattached attribute feature image are input into a face matching model to obtain face matching information; the attribute information to be compared, the face matching information, and the preset face features in the attribute editing sub-model to be trained are processed based on the loss function in the attribute editing sub-model to be trained to obtain a target loss value; the model parameters in the attribute editing sub-model to be trained are corrected based on the target loss value, and the convergence of the loss function is used as the training objective to train a face attribute determination model to be used; the target face attribute determination model is obtained by pruning the face attribute determination model to be used.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method as described in any one of claims 1-8.
11. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the image processing method as described in any one of claims 1-8.
Citation Information
Patent Citations
Face image age conversion method based on gradient adversarial attack and generative adversarial model
CN113569780A