Model training method and device based on small data generation network, equipment and medium

By iteratively training and configuring parameters, the image generation network is trained using face and cartoon image sets, which solves the problem of poor performance of the generation network with limited training data and achieves efficient generation of high-quality cartoon portraits.

CN115223013BActive Publication Date: 2026-04-28SHENZHEN WONDERSHARE SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN WONDERSHARE SOFTWARE CO LTD
Filing Date
2022-07-04
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently train image generation networks with limited training data, resulting in poor quality cartoon portraits.

Method used

The initial image generation network is trained iteratively, the target encoder is trained, and the second image generation network with fixed weights is trained iteratively using paired sets of face and cartoon images to ensure the parameter configuration and weight adjustment of the generation network.

Benefits of technology

This method enables efficient training of an image generation network with limited training data, generating high-quality cartoon portraits and improving the performance of the image generation network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223013B_ABST
    Figure CN115223013B_ABST
Patent Text Reader

Abstract

The application discloses a model training method and device based on a small data generation network, equipment and a medium. The method comprises the following steps: iteratively training an initial image generation network according to a face image set to obtain a trained first image generation network; training an encoder according to the face image set to obtain a target encoder and performing parameter configuration on the first image generation network to obtain a second image generation network; fixing the weight of the second image generation network; and iteratively training the second image generation network with fixed weight according to a cartoon image set to obtain a target image generation network. The application belongs to the technical field of model training, and the generation network can be trained based on a pair of face images and cartoon images. After the initial image generation network is trained by the face images, the second image generation network with fixed weight is iteratively trained by the cartoon images. The initial image generation network can be efficiently trained based on less training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of model training technology, and in particular to a model training method, apparatus, device, and medium based on a small data generation network. Background Technology

[0002] Chat and dating apps often require portraits to represent users' personal characteristics. Existing technologies can intelligently generate cartoon avatars based on user images for use in these apps. However, training image generation networks for cartoon avatars is often difficult due to the scarcity of training data. Since cartoon avatars are the target information for the model's output, the amount of cartoon face data that can be created or collected is very limited. This lack of training data hinders effective training of the image generation network, resulting in poor quality cartoon images generated based on real-life input images.

[0003] Therefore, existing techniques have the problem of being unable to efficiently train image generation networks with limited training data. Summary of the Invention

[0004] This invention provides a model training method, apparatus, device, and medium based on a small data generation network, aiming to solve the problem in existing methods that cannot efficiently train image generation networks with limited training data.

[0005] In a first aspect, embodiments of the present invention provide a model training method based on a small data generation network, the method comprising:

[0006] The initial image generation network is iteratively trained based on a pre-stored set of face images to obtain the first trained image generation network.

[0007] A preset encoder is trained based on the multi-layer analysis network in the first image generation network and the face image set to obtain the corresponding target encoder.

[0008] The first image generation network is configured with parameters according to the target encoder to obtain the corresponding second image generation network;

[0009] The second image generation network with fixed weights is iteratively trained according to the preset training conditions and the pre-stored cartoon image set to obtain the target image generation network that satisfies the training conditions. Each cartoon image in the cartoon image set corresponds to a face image in the face image set.

[0010] Secondly, embodiments of the present invention provide a model training apparatus based on a small data generation network, comprising:

[0011] The initial image generation network training unit is used to iteratively train the initial image generation network based on a pre-stored set of face images to obtain the first trained image generation network.

[0012] The encoder training unit is used to train a preset encoder based on the multi-layer analysis network in the first image generation network and the face image set, so as to train the corresponding target encoder.

[0013] The parameter configuration unit is used to configure the parameters of the first image generation network according to the target encoder to obtain the corresponding second image generation network.

[0014] The iterative training unit is used to iteratively train the second image generation network with fixed weights according to preset training conditions and a pre-stored cartoon image set to obtain a target image generation network that satisfies the training conditions. Each cartoon image in the cartoon image set corresponds to a face image in the face image set.

[0015] Thirdly, embodiments of the present invention provide a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the model training method based on a small data generation network described in the first aspect.

[0016] Fourthly, embodiments of the present invention also provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the model training method based on a small data generation network described in the first aspect.

[0017] This invention provides a model training method, apparatus, device, and medium for generating networks based on small data. An initial image generation network is iteratively trained using a set of face images to obtain a trained first image generation network. An encoder is then trained using the same set of face images to obtain a target encoder. The parameters of the first image generation network are configured to obtain a second image generation network. The weights of the second image generation network are fixed, and the second image generation network with fixed weights is iteratively trained using a set of cartoon images to obtain a target image generation network. This method allows for efficient training of the initial image generation network using relatively little training data, resulting in a target image generation network that meets practical needs. High-quality cartoon images can be generated based on this target image generation network. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic flowchart of a model training method based on a small data generation network provided in an embodiment of the present invention.

[0020] Figure 2 A schematic diagram of a sub-process of the model training method based on a small data generation network provided in an embodiment of the present invention;

[0021] Figure 3 This is a schematic diagram of another sub-process of the model training method based on a small data generation network provided in an embodiment of the present invention;

[0022] Figure 4 This is a schematic diagram of another sub-process of the model training method based on a small data generation network provided in an embodiment of the present invention;

[0023] Figure 5 This is a schematic diagram of the latter sub-process of the model training method based on a small data generation network provided in an embodiment of the present invention;

[0024] Figure 6 This is another sub-process diagram of the model training method based on a small data generation network provided in an embodiment of the present invention;

[0025] Figure 7 This is a schematic diagram of a subsequent sub-process of the model training method based on a small data generation network provided in an embodiment of the present invention.

[0026] Figure 8 A schematic block diagram of a model training device based on a small data generation network provided in an embodiment of the present invention;

[0027] Figure 9 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0030] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0031] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0032] Please see Figure 1 , Figure 1 This is a flowchart illustrating the model training method based on a small data generation network provided in an embodiment of the present invention. This method is applied to a user terminal or management server and is executed by application software installed on the user terminal or management server. The user terminal is a terminal device that can be used to execute the model training method based on a small data generation network to receive input sets of face images and cartoon images, and train the image generation network to obtain the target image generation network. The user terminal can be a desktop computer, laptop computer, tablet computer, or mobile phone, etc. The management server is a server-side device used to execute the model training method based on a small data generation network to obtain the set of face images and cartoon images uploaded by the user terminal, and train the image generation network to obtain the target image generation network, such as a server-side device built within an enterprise or government department. Figure 1 As shown, the method includes steps S110 to S140.

[0033] S110. The initial image generation network is iteratively trained based on the pre-stored set of face images to obtain the first trained image generation network.

[0034] The initial image generation network is iteratively trained using a pre-stored set of face images to obtain the first trained image generation network. Users can input a set of face images into the management server or user terminal, and the initial image generation network can be iteratively trained based on this set. Specifically, the face image set contains a certain number of real face images, all of the same size (e.g., 1024×1024). The initial image generation network can be an intelligent image generation model built based on Style GAN (Style-Generative Adversarial Networks). Since the initial image generation network contains multiple parameter values, iterative training can be performed using the face images in the face image set, that is, iteratively adjusting the parameter values ​​in the model to obtain the model with adjusted parameters, which serves as the first trained image generation network. Because the number of cartoon images obtained is relatively small, and face images in the face image set are used in pairs with cartoon images in the cartoon image set, the number of corresponding face images in the face image set is also relatively small.

[0035] In one embodiment, such as Figure 2 As shown, step S110 includes sub-steps S111, S112, S113, S114, S115, S116, S117 and S118.

[0036] S111. Obtain a face image from the face image set as the current face image; S112. Extract the input features corresponding to the current face image based on the convolutional input layer of the initial image generation network, wherein the input features include hidden variable features and noise features.

[0037] It can acquire any face image from a set of face images as the current face image, and train an initial image convolutional model using this current face image. Specifically, the initial image convolutional model includes a convolutional input layer, which performs convolution processing on the current face image to obtain corresponding input features. The input features include hidden variable features and noise features. The hidden variable feature z can be represented as a feature vector to reflect the hidden feature information of the current face image, and the noise feature B can be represented as a feature vector to reflect the noise feature information of the current face image. The noise feature B is used to enrich the detailed feature information of the generated image.

[0038] S113. Perform feature analysis on the input features according to the multi-layer analysis network in the initial image generation network to output an output image corresponding to the input features.

[0039] The input features can be analyzed by the multi-layer analysis layers in the initial image generation network. Specifically, multi-layer convolution processing based on the hidden variable z can further obtain intermediate hidden variables W. The obtained intermediate hidden variables W are then subjected to affine transformation to obtain feature A. Feature A is the feature vector used to control the overall style of the generated image. The intermediate hidden variable W can have multiple variables w (w∈W), and each variable w in the intermediate hidden variables corresponds to a feature A. Feature A and noise feature B are respectively input into each analysis layer of the multi-layer analysis network. The output image size of each analysis layer is different, but the image size corresponding to each analysis layer is a multiple of 2, such as 4×4, 8×8, 16×16, etc. After feature analysis is performed layer by layer by the multi-layer analysis network, the multi-layer basic images obtained are integrated to obtain the corresponding output image. The size of the output image is the same as the size of the current face image (e.g., both are 1024×1024).

[0040] S114. Calculate the loss value between the output image and the current face image according to the preset loss function.

[0041] The loss value between the output image and the current face image can be calculated based on a preset loss function. Specifically, the difference value of each pixel between the output image and the current face image can be obtained. Generally speaking, each pixel in the output image and the current face image contains the pixel value corresponding to three color channels (based on RGB, it contains red, green and blue color channels). The loss value between the two images can be calculated based on the difference value of each pixel. The closer the output image is to the current face image, the smaller the corresponding image loss value is, and vice versa.

[0042] For example, the loss function can be expressed using formula (1):

[0043]

[0044] Where N is the number of rows of pixels in the image, M is the number of columns of pixels in the image, and P ij To output the pixel values ​​of the three color channels corresponding to the i-th row and j-th column of the image, R ij These are the pixel values ​​for the three color channels of the current face image.

[0045] S115. Adjust the parameter values ​​contained in the initial image generation network according to the preset parameter adjustment rules and the loss value.

[0046] The parameter values ​​contained in the initial image generation network can be adjusted based on parameter adjustment rules and loss values. Specifically, the parameter adjustment rules can be gradient descent rules, which are configured with a learning rate for gradient descent. The updated parameter values ​​can be calculated based on the gradient descent rules, loss values, and the calculated values ​​of each calculation process in the model. The original parameter values ​​of each parameter are updated based on the updated parameter values. The parameter values ​​in the model include the parameter values ​​configured in the analysis layer, convolutional layer, and convolutional input layer.

[0047] S116. Determine whether the face image set contains other untrained face images; S117. If the face image set contains other untrained face images, obtain the next face image as the current face image and return to execute the step of extracting the input features corresponding to the current face image from the convolutional input layer of the initial image generation network; S118. If the face image set does not contain other untrained face images, determine the currently obtained initial image generation network as the first image generation network after training.

[0048] The system can determine whether the face image set contains other unused face images. If so, it continues to acquire the next face image as the current face image for model training. If not, it uses the currently trained first image generation network as the next trained first image generation network. A single face image can be used to train the initial image generation network once. Another face image can be used to train the previously trained initial image generation network again. Multiple face images can be used to iteratively train the initial image generation network.

[0049] In one embodiment, such as Figure 3 As shown, step S115 is followed by steps S1151, S1152 and S1153.

[0050] S1151. Determine whether the loss value is not greater than a preset loss threshold; S1152. If the loss value is not greater than the loss threshold, determine the currently obtained initial image generation network as the first image generation network after training; S1153. If the loss value is greater than the loss threshold, randomly select a face image from the face image set as the current face image and return to execute the step of extracting the input features corresponding to the current face image based on the convolutional input layer of the initial image generation network.

[0051] Specifically, after calculating the loss value during model training, it can be determined whether the loss value is not greater than a preset loss threshold. If the calculated loss value is not greater than the loss threshold, it indicates that the currently trained first image generation network can generate an output image close to the original face image, and the training process can be terminated, with the currently obtained first image generation network used as the trained first image generation network. If the loss value is greater than the loss threshold, it indicates that the model training is insufficient, and the next face image can be obtained for further training of the model.

[0052] S120. Train a preset encoder based on the multi-layer analysis network in the first image generation network and the face image set to obtain the corresponding target encoder.

[0053] A pre-set encoder is trained using the multi-layer analysis network in the first image generation network and the face image set to obtain the corresponding target encoder. Specifically, the encoder is a virtual structure used for encoding image processing. It can be trained in conjunction with the multi-layer analysis network in the already trained first image generation network to obtain the trained target encoder.

[0054] In one embodiment, such as Figure 4 As shown, step S120 includes sub-steps S121, S122, S123, S124, S125, S126, S127 and S128.

[0055] S121. Obtain a face image from the face image set as the current face image; S122. Perform convolution processing on the current face image according to the encoder to obtain the convolution feature vectors corresponding to the multiple convolution layers in the encoder.

[0056] A single face image from a set of face images can be acquired as the current face image. An encoder performs convolution processing on this face image, comprising multiple concatenated convolutional layers. Each convolutional layer processes the pixel values ​​of the face image. The convolution result from the previous layer is used as input to the next layer. Specifically, the convolutional feature vector obtained from each convolutional layer's processing of the face image can be acquired. This feature vector is the convolution result of the layer and consists of a three-dimensional array containing multiple convolutional values. Thus, each convolutional layer corresponds to one feature vector.

[0057] For example, the encoder contains nine convolutional layers arranged in series. After the initial image is processed by the convolutional processing model, the convolutional feature vectors corresponding to the nine convolutional layers can be obtained in sequence.

[0058] S123. Input the convolutional feature vectors of the multiple convolutional layers into the multiple analysis layers of the multi-layer analysis network to analyze and obtain the corresponding training face images.

[0059] Specifically, the convolutional feature vectors in the convolutional layer can be used as intermediate hidden variables W. A set of convolutional feature vectors can be decomposed into multiple variables w, and these multiple variables w together form a set of intermediate hidden variables W, such as intermediate hidden variables W = {w1, w2, ... w...} 18 The obtained convolutional feature vectors are subjected to an affine transformation to obtain feature A. Each intermediate hidden variable corresponds to one feature A, resulting in 18 features A. Feature A is the feature vector used to control the overall style of the generated image. The multiple features A are then input into the analysis layers of a multi-layer analysis network. For example, the first two features A are input into the first analysis layer, resulting in an output image size of 4×4; the third and fourth features A are input into the second analysis layer, resulting in an output image size of 8×8; and so on, with the last two features A being input into the last analysis layer, resulting in an output image size of 1024×1024. The images output from each analysis layer are combined to obtain the corresponding training face image.

[0060] S124. Calculate the loss value between the training face image and the current face image according to the loss function.

[0061] The loss value between the training face image and the current face image can be calculated based on the loss function. The process of calculating the loss value is the same as the process of obtaining the loss value between the output image and the current face image, and will not be described in detail here.

[0062] S125. Adjust the parameter values ​​contained in the encoder according to the parameter adjustment rules and the loss value.

[0063] The encoder's functional components, such as convolutional layers and affine transformation layers, are all configured with parameter values. The parameter values ​​contained in the encoder can be adjusted according to the parameter adjustment rules and the calculated loss values. The process of adjusting the parameter values ​​is the same as the process of adjusting the parameter values ​​in the initial image generation network, and will not be elaborated here.

[0064] S126. Determine whether the face image set contains other untrained face images; S127. If the face image set contains other untrained face images, obtain the next face image as the current face image and return to execute the convolution processing of the current face image according to the encoder to obtain the convolution feature vectors corresponding to the multiple convolution layers in the encoder; S128. If the face image set does not contain other untrained face images, use the currently obtained encoder as the trained target encoder.

[0065] Since it is necessary to obtain the latent features of all face images in the face image set for training, each face image in the face image set can be used to train the encoder sequentially until each face image in the face image set has been used to train the encoder, and finally the target encoder is obtained.

[0066] S130. Configure the parameters of the first image generation network according to the target encoder to obtain the corresponding second image generation network.

[0067] The first image generation network is configured with parameters based on the target encoder to obtain the corresponding second image generation network. To ensure that the real-life photo and the subsequently generated cartoon photo have the same distance and orientation in the latent space, the latent space is composed of the hidden variable z, i.e., part of the weight layers in the first image generation network. By configuring the parameters of the trained first image generation network through the target encoder, the cartoon photo generated using the second image generation network can be made to have the same distance and orientation as the real-life photo in the latent space.

[0068] In one embodiment, such as Figure 5 As shown, step S130 includes sub-steps S131 and S132.

[0069] S131. Extract the attribute encoding vector corresponding to the face image set from the target encoder.

[0070] Specifically, each face image in the face image set can be encoded separately according to the target encoder, thereby extracting the convolutional feature vector corresponding to each face image. The convolutional feature vectors of all face images are combined and then split into multiple corresponding variables w. The multiple variables w can be combined into an attribute encoding vector corresponding to the face image set.

[0071] S132. The attribute encoding vector is transferred to the first image generation network to configure the intermediate hidden feature parameters in the first image generation network, thereby obtaining the second image generation network after parameter configuration.

[0072] The obtained attribute encoding vector is transferred to the first image generation network, thus configuring the attribute encoding vector as an intermediate hidden feature parameter, which also fixes the corresponding intermediate hidden feature weight parameters in the first image generation network. The attribute encoding vector contains multiple variables w, each corresponding to a feature A that needs to be input to the multilayer analysis network in the first image generation network. Therefore, the attribute encoding vector can be directly configured as the upstream input information of feature A and fixed, and the upstream input information of feature A remains unchanged during the subsequent use of the second image generation network.

[0073] S140. The second image generation network with fixed weights is iteratively trained according to the preset training conditions and the pre-stored cartoon image set to obtain the target image generation network that satisfies the training conditions.

[0074] The second image generation network with fixed weights is iteratively trained based on preset training conditions and a pre-stored cartoon image set to obtain a target image generation network that meets the training conditions. To ensure that the second image generation network can generate cartoon-style face images, iterative training can be performed on the configured second image generation network according to the training conditions and the cartoon image set. Each cartoon image in the cartoon image set corresponds to a face image in the face image set; for example, the size of each cartoon image in the cartoon image set is the same as the size of the face image (e.g., both are 1024×1024). The weights of the convolutional layers in the second image generation network are fixed. During this training, other non-weight parameter values ​​in the second image generation network can be adjusted. After iterative training of the second image generation network, a target image generation network containing cartoon styles can be obtained. To ensure that the cartoon images generated by the target image generation network are more similar to the corresponding face images, thereby avoiding excessive differences between the generated cartoon images and the input face images, in addition to fixing the weight parameters in the second image generation network, iterative training is also performed using a set of cartoon images corresponding to the face image set used in the initial training. That is, in this embodiment of the application, paired face images and cartoon images are used for training when training the image generation network; wherein, in addition to the aforementioned intermediate hidden feature weight parameters, the weight parameters also include other weight parameters in the image generation network.

[0075] In one embodiment, such as Figure 6 As shown, step S140 includes sub-steps S141, S142, S143, S144, S145, S146 and S147.

[0076] S141. Obtain a cartoon image from the cartoon image set as the current cartoon image; S142. Extract cartoon noise features corresponding to the current cartoon image based on the convolutional input layer in the second image generation network.

[0077] A cartoon image can be obtained as the current cartoon image, and the cartoon noise features corresponding to the current cartoon image can be extracted. Since this training does not involve the hidden variable part in the second image generation network, it is only necessary to obtain the cartoon noise features of the cartoon image.

[0078] S143. Perform feature analysis on the cartoon noise features according to the multilayer analysis network in the second image generation network, so as to obtain the real-person output image corresponding to the current cartoon image from the latent vector layer of the multilayer analysis network.

[0079] Based on the multi-layer analysis network and fixed intermediate hidden feature weight parameters in the second image generation network, the cartoon noise features are analyzed, so that the corresponding real-life output image can be obtained from the latent vector layer of the multi-layer analysis network. The size of the real-life output image is the same as the size of the current cartoon image.

[0080] S144. Calculate the contrast loss value between the face image corresponding to the current cartoon image and the real-person output image according to the loss calculation formula in the training conditions; S145. Adjust the non-weight parameter values ​​in the multilayer analysis network according to the contrast loss value.

[0081] Specifically, the contrast loss value can be calculated based on the pixel content of the image. To calculate the contrast loss value, it is necessary to use the face image corresponding to the current cartoon image and the real-life output image output by the image generation network.

[0082] In one embodiment, such as Figure 7 As shown, step S144 includes sub-steps S1441, S1442, S1443, S1444 and S1445.

[0083] S1441. Calculate the mean square error between the pixel values ​​of the face image corresponding to the current cartoon image and the real-person output image; S1442. Calculate the mean absolute error between the pixel values ​​of the face image corresponding to the current cartoon image and the real-person output image; S1443. Obtain the similarity value between the face image corresponding to the current cartoon image and the real-person output image according to a preset face recognition model; S1444. Obtain the perceptual error value between the face image corresponding to the current cartoon image and the real-person output image according to a preset classification deep learning model; S1445. Combine and calculate the mean square error, the mean absolute error, the similarity value, and the perceptual error value according to a preset combined calculation formula to obtain the contrast loss value between the current cartoon image and the real-person output image.

[0084] Specifically, the pixel values ​​of the real-life output image generated by the second generator network can be compared and analyzed with the pixel values ​​of the face image corresponding to the current cartoon image. Specifically, the mean-square error (MSE) between the pixel values ​​of the images can be calculated. MSE is a measure reflecting the degree of difference between the estimator and the estimated quantity. The MSE is obtained by calculating the pixel values ​​of each pixel in the real-life output image across the three color channels (e.g., RBG represents red, blue, and green) and the pixel values ​​of each pixel in the face image corresponding to the current cartoon image across the three color channels.

[0085] The absolute error between two images can be obtained by subtracting the pixel value of each pixel in a certain color channel (e.g., RGB represents the red, blue, and green color channels) in the real-person output image from the pixel value of the same pixel in the same color channel in the face image corresponding to the current cartoon image, and then averaging the absolute values ​​of all pixels in the three color channels.

[0086] The obtained real-life output image and the corresponding face image of the current cartoon image can be simultaneously input into the face recognition model. The pre-set face recognition model performs similarity analysis on the two images to obtain the corresponding similarity value. Alternatively, two images of the same person or two images of two different people can be simultaneously input into the face recognition model for training, resulting in a trained face recognition model that yields a similarity value. The similarity value ranges from 0 to 1.

[0087] The obtained real-life output image and the corresponding face image of the current cartoon image can be simultaneously input into a pre-set classification deep learning model. The classification deep learning model performs perceptual analysis on the two images to obtain the corresponding perceptual error value. This perceptual error value reflects the actual perceptual error of the human body between the two images. Alternatively, similarity or dissimilarity labels can be manually added between the two images, and the two labeled images can be simultaneously input into the face recognition model for training. The trained face recognition model then yields the perceptual error value. The perceptual error value ranges from 0 to 1.

[0088] The comparison loss value can be obtained by combining the obtained mean squared error, mean absolute error, similarity value, and perceptual error value. Specifically, the multiple values ​​obtained above can be combined using a combination calculation formula. For example, the combination calculation formula can be expressed as formula (2):

[0089]

[0090] Where G is the calculated contrast loss value, l1 is the mean squared error, l2 is the mean absolute error, l3 is the similarity value, and l4 is the perceptual error value.

[0091] The specific process of adjusting the parameter values ​​in the corresponding convolutional layer based on the contrast loss value is the same as the process of adjusting the parameter values ​​in the initial image generation network, and will not be elaborated here.

[0092] S146. Determine whether the number of training iterations exceeds the preset number of training conditions, and determine whether the number of times the contrast loss value is less than the contrast loss threshold in the training conditions exceeds the preset number threshold; S147. If the number of training iterations exceeds the preset number, or the number of times the contrast loss value is less than the contrast loss threshold exceeds the preset number threshold, determine the currently obtained second image generation network as the trained target image generation network.

[0093] It can determine whether the number of training iterations for the second image generation network exceeds the preset number of training conditions, and at the same time determine whether the contrast loss value is less than the contrast loss threshold in the training conditions. If it is less, it can further determine whether the number of times the contrast loss value is less than the contrast loss threshold exceeds the preset number threshold.

[0094] In the above judgment results, if the number of training attempts exceeds the preset number, or the number of times the contrast loss value is less than the contrast loss threshold exceeds the preset threshold, then the training of the second image generation network ends, and the currently obtained second image generation network is determined as the target image generation network.

[0095] If the number of training iterations does not exceed the preset number of iterations and the number of times the contrast loss value is less than the contrast loss threshold does not exceed the preset threshold, then continue to obtain cartoon images from the cartoon image set to train the second image generation network until the above training conditions are met.

[0096] The above training method ensures that during the training of the second generative network based on cartoon images, the latent vectors in the latent vector layer corresponding to the cartoon images are as close as possible to the latent vectors of the corresponding face images. This ensures that the overall style and appearance of the generated cartoon images are closer to the face images, avoiding excessive differences between the generated cartoon images and the input face images, and significantly improving the quality of the cartoon images generated by the target image-based generative network.

[0097] Subsequently, if the input image to be converted is received, the style conversion of the image to be converted can be performed according to the trained target image generation network to obtain the target cartoon image corresponding to the image to be converted.

[0098] For example, the specific implementation process is as follows: image noise features corresponding to the image to be converted are extracted from the convolutional input layer in the target image generation network; feature analysis is performed on the image noise features using the multi-layer analysis network in the target image generation network to convert and obtain a target cartoon image corresponding to the image to be converted. The size of the obtained target cartoon image is the same as the size of the image to be converted.

[0099] In the model training method for generating networks based on small data provided in this embodiment of the invention, an initial image generation network is iteratively trained using a set of face images to obtain a trained first image generation network. An encoder is then trained using the face image set to obtain a target encoder, and parameters of the first image generation network are configured to obtain a second image generation network. The weights of the second image generation network are fixed, and the second image generation network with fixed weights is iteratively trained using a set of cartoon images to obtain a target image generation network. Through this method, the generation network can be trained based on pairs of face images and cartoon images. After training the initial image generation network using face images to obtain the second image generation network, iterative training of the second image generation network with fixed weights using a set of cartoon images allows for efficient training of the initial image generation network with relatively little training data, resulting in a target image generation network that meets practical needs. High-quality cartoon images can be generated based on this target image generation network.

[0100] This invention also provides a model training apparatus based on a small data generation network. This apparatus can be configured in a user terminal or a management server and is used to execute any of the aforementioned embodiments of the model training method based on a small data generation network. Specifically, please refer to... Figure 8 , Figure 8 This is a schematic block diagram of a model training device based on a small data generation network provided in an embodiment of the present invention.

[0101] like Figure 8 As shown, the model training device 100 based on the small data generation network includes an initial image generation network training unit 110, an encoder training unit 120, a parameter configuration unit 130, and an iterative training unit 140.

[0102] The initial image generation network training unit 110 is used to iteratively train the initial image generation network based on a pre-stored set of face images to obtain the trained first image generation network.

[0103] The encoder training unit 120 is used to train a preset encoder based on the multi-layer analysis network in the first image generation network and the face image set, so as to train the corresponding target encoder.

[0104] The parameter configuration unit 130 is used to configure the parameters of the first image generation network according to the target encoder to obtain the corresponding second image generation network.

[0105] The iterative training unit 140 is used to iteratively train the second image generation network with fixed weights according to preset training conditions and a pre-stored cartoon image set to obtain a target image generation network that satisfies the training conditions. Each cartoon image in the cartoon image set corresponds to a face image in the face image set.

[0106] The model training device based on a small-data generated network provided in this embodiment of the invention applies the aforementioned model training method based on a small-data generated network. Iterative training of an initial image generation network using a set of face images yields a trained first image generation network. An encoder is trained using the face image set to obtain a target encoder, and parameters of the first image generation network are configured to obtain a second image generation network. The weights of the second image generation network are fixed, and iterative training of the second image generation network with fixed weights using a set of cartoon images yields a target image generation network. Through this method, the generation network can be trained based on paired face images and cartoon images. After training the initial image generation network using face images to obtain the second image generation network, iterative training of the second image generation network with fixed weights using a set of cartoon images allows for efficient training of the initial image generation network with relatively little training data, resulting in a target image generation network that meets practical needs. High-quality cartoon images can be generated based on this target image generation network.

[0107] The aforementioned model training device based on small data generation networks can be implemented as a computer program, which can be used in, for example... Figure 9 It runs on the computer device shown.

[0108] Please see Figure 9 , Figure 9 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. The computer device can be a user terminal or management server used to execute a model training method based on a small data generation network to receive input sets of face images and cartoon images, and to train the image generation network to obtain a target image generation network.

[0109] See Figure 9 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a storage medium 503 and internal memory 504.

[0110] The storage medium 503 may store the operating system 5031 and the computer program 5032. When the computer program 5032 is executed, it enables the processor 502 to execute a model training method based on a small data generation network. The storage medium 503 may be a volatile storage medium or a non-volatile storage medium.

[0111] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.

[0112] The internal memory 504 provides an environment for the operation of the computer program 5032 in the storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a model training method based on a small data generation network.

[0113] This network interface 505 is used for network communication, such as providing data transmission. Those skilled in the art will understand that... Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device 500 to which the present invention is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0114] The processor 502 is used to run the computer program 5032 stored in the memory to implement the corresponding functions in the above-described model training method based on small data generation network.

[0115] Those skilled in the art will understand that Figure 9 The embodiments of the computer device shown do not constitute a limitation on the specific configuration of the computer device. In other embodiments, the computer device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. For example, in some embodiments, the computer device may include only memory and a processor. In such embodiments, the structure and function of the memory and processor are different from those shown. Figure 9 The embodiments shown are consistent and will not be repeated here.

[0116] It should be understood that, in this embodiment of the invention, the processor 502 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0117] In another embodiment of the invention, a computer-readable storage medium is provided. This computer-readable storage medium may be volatile or non-volatile. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps included in the above-described model training method based on a small data generation network.

[0118] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.

[0119] In the embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Units with the same function may be grouped into one unit. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, or it may be an electrical, mechanical, or other form of connection.

[0120] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.

[0121] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0122] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned computer-readable storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks.

[0123] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A model training method based on a small data generation network, characterized in that, The method includes: The initial image generation network is iteratively trained based on a pre-stored set of face images to obtain the first trained image generation network. A preset encoder is trained based on the multi-layer analysis network in the first image generation network and the face image set to obtain the corresponding target encoder. The first image generation network is configured with parameters according to the target encoder to obtain the corresponding second image generation network; The second image generation network with fixed weights is iteratively trained according to the preset training conditions and the pre-stored cartoon image set to obtain the target image generation network that satisfies the training conditions. Each cartoon image in the cartoon image set corresponds to a face image in the face image set. The step of training a pre-set encoder based on the multi-layer analysis network in the first image generation network and the face image set to train a corresponding target encoder includes: Obtain one face image from the face image set as the current face image; The current face image is convolved according to the encoder to obtain the convolutional feature vectors corresponding to the multiple convolutional layers in the encoder. The feature maps of the multiple convolutional layers are respectively input into the multiple analysis layers of the multi-layer analysis network to analyze and obtain the corresponding training face images; The loss value between the training face image and the current face image is calculated based on a preset loss function. The parameter values ​​contained in the encoder are adjusted according to the preset parameter adjustment rules and the loss value; Determine whether the face image set contains other face images that have not been trained; If the face image set contains other untrained face images, obtain the next face image as the current face image and return to perform the convolution processing on the current face image according to the encoder to obtain the convolution feature vectors corresponding to the multiple convolution layers in the encoder. If the face image set does not contain other untrained face images, the currently obtained encoder is used as the target encoder after training. The step of iteratively training a first image generation network with fixed weights according to preset training conditions and a pre-stored set of cartoon images to obtain a target image generation network that satisfies the training conditions includes: Obtain one cartoon image from the cartoon image set as the current cartoon image; Cartoon noise features corresponding to the current cartoon image are extracted from the convolutional input layer in the second image generation network. The cartoon noise features are analyzed using the multi-layer analysis network in the second image generation network to obtain a real-person output image corresponding to the current cartoon image from the latent vector layer of the multi-layer analysis network. Calculate the contrast loss value between the face image corresponding to the current cartoon image and the real-person output image according to the loss calculation formula in the training conditions; The non-weight parameter values ​​in the multilayer analysis network are adjusted based on the contrastive loss value; Determine whether the number of training iterations exceeds the preset number of training conditions, and determine whether the number of times the contrast loss value is less than the contrast loss threshold in the training conditions exceeds the preset number threshold. If the number of training iterations exceeds the preset number of iterations, or the number of times the contrast loss value is less than the contrast loss threshold exceeds the preset number of iterations threshold, the currently obtained second image generation network is determined as the trained target image generation network.

2. The model training method based on a small data generation network according to claim 1, characterized in that, The step of iteratively training the initial image generation network based on a pre-stored set of face images to obtain the trained first image generation network includes: Obtain one face image from the face image set as the current face image; The input features corresponding to the current face image are extracted from the convolutional input layer of the initial image generation network. The input features include hidden variable features and noise features. The input features are analyzed using a multi-layer analysis network in the initial image generation network to output an output image corresponding to the input features; The loss value between the output image and the current face image is calculated based on a preset loss function; The parameter values ​​contained in the initial image generation network are adjusted according to the preset parameter adjustment rules and the loss value; Determine whether the face image set contains other face images that have not been trained; If the face image set contains other untrained face images, obtain the next face image as the current face image and return to execute the convolutional input layer of the network generated from the initial image to extract the input features corresponding to the current face image; If the face image set does not contain any other untrained face images, the currently obtained initial image generation network is determined as the first trained image generation network.

3. The model training method based on a small data generation network according to claim 2, characterized in that, After adjusting the parameter values ​​contained in the initial image generation network according to the preset parameter adjustment rules and the loss value, the method further includes: Determine whether the loss value is not greater than a preset loss threshold; If the loss value is not greater than the loss threshold, the currently obtained initial image generation network is determined as the first image generation network after training; If the loss value is greater than the loss threshold, a face image is randomly selected from the face image set as the current face image, and the process returns to execute the convolutional input layer of the network that generates the image from the initial image to extract the input features corresponding to the current face image.

4. The model training method based on a small data generation network according to claim 1, characterized in that, The step of configuring the parameters of the first image generation network according to the target encoder to obtain the corresponding second image generation network includes: The attribute encoding vector corresponding to the face image set is extracted from the target encoder; The attribute encoding vector is transferred to the first image generation network to configure the intermediate hidden feature parameters in the first image generation network, resulting in a second image generation network with configured parameters.

5. The model training method based on a small data generation network according to claim 1, characterized in that, The step of calculating the contrast loss value between the current cartoon image and the live-action output image according to the loss calculation formula in the training conditions includes: Calculate the mean square error between the pixel values ​​of the face image corresponding to the current cartoon image and the real-life output image; Calculate the mean absolute error between the pixel values ​​of the face image corresponding to the current cartoon image and the real-person output image; The similarity value between the face image corresponding to the current cartoon image and the real-person output image is obtained according to the preset face recognition model; The perceptual error value between the face image corresponding to the current cartoon image and the real-person output image is obtained according to the preset classification deep learning model; The mean squared error, the mean absolute error, the similarity value, and the perceptual error value are combined and calculated according to a preset combined calculation formula to obtain the contrast loss value between the current cartoon image and the real-person output image.

6. A model training apparatus based on a small data generation network, the apparatus being used to execute the model training method based on a small data generation network as described in any one of claims 1-5, characterized in that, The device includes: The initial image generation network training unit is used to iteratively train the initial image generation network based on a pre-stored set of face images to obtain the first trained image generation network. The encoder training unit is used to train a preset encoder based on the multi-layer analysis network in the first image generation network and the face image set, so as to train the corresponding target encoder. The parameter configuration unit is used to configure the parameters of the first image generation network according to the target encoder to obtain the corresponding second image generation network. The iterative training unit is used to iteratively train the second image generation network with fixed weights according to preset training conditions and a pre-stored cartoon image set to obtain a target image generation network that satisfies the training conditions. Each cartoon image in the cartoon image set corresponds to a face image in the face image set.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the model training method based on a small data generation network as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the model training method based on a small data generation network as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and device for generating a cartoon head portrait generation model

    CN109800732A

  • Image generation model training method and device, electronic equipment and storage medium

    CN113096055A