Model training method, method for completing a person image, device, and electronic device

By generating multiple rounds of adversarial training for neural networks and discriminative neural networks, combining preset random templates and dense coordinate images, the problem of surface texture completion difficulties in character image completion is solved, and a more efficient image completion effect is achieved.

CN112488284BActive Publication Date: 2025-05-27BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910860914.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-09-11
Publication Date
2025-05-27
Estimated Expiration
2039-09-11

AI Technical Summary

Technical Problem

When the prior art completes the image of the character, it is difficult to effectively complete the surface texture of different characters, resulting in a large gap between the completed image and the original image.

Method used

A model training method is adopted to generate multiple rounds of adversarial training of neural networks and discriminative neural networks, combined with preset random templates and dense coordinate images, to generate complete character images and complete edge images, thereby reducing the gap between the complete image and the original image.

Benefits of technology

By introducing dense coordinate images, the generation neural network can complete the surface texture according to the posture of different characters, significantly reducing the gap between the complete image and the original image, and improving the completion effect of the character image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112488284B_ABST
    Figure CN112488284B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of image processing, and specifically relates to a model training method, a human image completion method, a model training device, a human image completion device, a medium, and an electronic device. The model training method includes: preprocessing a human sample image according to a preset algorithm, a preset random template, and a preset model respectively to obtain a target image and a to-be-processed image with a pixel loss region; inputting the preset random template and the to-be-processed image into a generation neural network to generate a completed human image and a completed edge image; inputting the completed human image, the completed edge image, the human sample image, and the target image into a discrimination neural network to generate a discrimination result; and performing multiple rounds of adversarial training on the generation neural network and the discrimination neural network according to the discrimination results corresponding to multiple human sample images. The technical solution of the embodiments of the present disclosure can complete the corresponding surface texture according to the postures of different people, thereby reducing the gap between the completed human image and the human sample image.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] During the processes of storing, transcoding, and transmitting images, it often occurs that pixel loss areas appear in the images. To ensure the quality of the images, developers often complete the images in the following two ways: one is to fill the pixel loss areas by matching and copying background blocks; the other is to train a generative adversarial network so that the generator in it can generate a complete and consistent completion image with the original missing image to complete the original image.

[0003] However, when encountering complex human images, the first method mentioned above will have the problem that it cannot capture the high-dimensional features of the images and thus cannot complete the filling; although the second method can capture high-dimensional features, it is difficult to complete the surface textures of different people. Therefore, when completing human images, there is still a large gap between the completed image and the original image before the loss.

[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a model training method, a human image completion method, a model training device, a human image completion device, a computer-readable storage medium, and an electronic device, so as to at least to some extent overcome the problem that when completing a human image, there is a large gap between the completed image and the original image before the loss due to the inability to complete the surface texture of the human.

[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.

[0007] According to a first aspect of the present disclosure, there is provided a model training method, including:

[0008] Preprocessing a human sample image according to a preset algorithm to obtain a target image, and preprocessing the human sample image according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss area; wherein, the to-be-processed image includes a to-be-processed human image and a to-be-processed dense coordinate image;

[0009] Inputting the preset random template and the to-be-processed image into a generative neural network to generate a completed human image and a completed edge image;

[0010] Inputting the completed human image, the completed edge image, the human sample image, and the target image into a discriminant neural network to generate a discriminant result;

[0011] Perform multiple rounds of adversarial training on the generation neural network and the discriminant neural network according to the discrimination results corresponding to multiple said character sample images.

[0012] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the target image includes a target edge image;

[0013] The preprocessing of the character sample image according to a preset algorithm to obtain a target image includes:

[0014] Extract the edge information in the character sample image according to a preset algorithm to obtain a target edge image.

[0015] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the preprocessing of the character sample image according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss area includes:

[0016] Occlude the character sample image according to a preset random template to obtain a to-be-processed character image with a pixel loss area;

[0017] Perform pose analysis on the to-be-processed character image according to a preset model to obtain the to-be-processed dense coordinate image.

[0018] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the to-be-processed image further includes a to-be-processed edge image;

[0019] The method further includes:

[0020] Extract edge information in the character sample image according to the preset algorithm;

[0021] Extract the edge information occluded by the preset random template from the edge information to obtain a to-be-processed edge image.

[0022] In an exemplary embodiment of the present disclosure, based on the foregoing solution, after extracting the edge information occluded by the preset random template from the edge information, the method further includes:

[0023] Randomly erase the edge information occluded by the preset random template.

[0024] In an exemplary embodiment of the present disclosure, based on the foregoing solution, before performing multiple rounds of adversarial training on the generation neural network and the discriminant neural network according to the discrimination results corresponding to multiple said character sample images, the method further includes:

[0025] Calculate an image loss function based on the completed character image, the completed edge image, the character sample image, and the target image.

[0026] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the multi-round adversarial training of the generation neural network and the discriminant neural network according to the discrimination results corresponding to multiple pieces of the human sample images includes:

[0027] Alternately performing the following two training processes:

[0028] Training the generation neural network according to the image loss function and the discrimination results corresponding to multiple pieces of the human sample images;

[0029] Training the discriminant neural network according to the discrimination results corresponding to multiple pieces of the human sample images.

[0030] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the image loss function includes at least one or a combination of multiple ones of a reconstruction loss function, a content loss function, and a style loss function.

[0031] According to a second aspect of the present disclosure, there is provided a method for completing a human image, including:

[0032] Obtaining a to-be-processed human image having a pixel loss area, and performing pose analysis on the to-be-processed human image according to a preset model to obtain a to-be-processed dense coordinate image;

[0033] Inputting preset marking information, the to-be-processed human image, and the to-be-processed dense coordinate image into a trained generation neural network to generate a completed human image;

[0034] Wherein, the preset marking information includes area marking information for the area with pixel loss on the to-be-processed human image; the trained generation neural network is obtained by training according to the model training method described in the first aspect.

[0035] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the preset marking information includes edge marking information for the area with pixel loss in the to-be-processed human image.

[0036] According to a third aspect of the present disclosure, there is provided a model training device, including:

[0037] A first processing module, configured to preprocess a human sample image according to a preset algorithm to obtain a target image, and preprocess the human sample image according to a preset random template and a preset model to obtain a to-be-processed image having a pixel loss area; wherein, the to-be-processed image includes a to-be-processed human image and a to-be-processed dense coordinate image;

[0038] An image generation module, configured to input the preset random template and the image to be processed into a generation neural network to generate a completed human image and a completed edge image;

[0039] A result discrimination module, configured to input the completed human image, the completed edge image, the human sample image, and the target image into a discrimination neural network to generate a discrimination result;

[0040] A model training module, configured to perform multiple rounds of adversarial training on the generation neural network and the discrimination neural network according to the discrimination results corresponding to multiple human sample images.

[0041] According to a fourth aspect of the present disclosure, there is provided a human image completion device, including:

[0042] An image processing module, configured to obtain an image of a human to be processed having a pixel loss region, and perform pose analysis on the image of the human to be processed according to a preset model to obtain a dense coordinate image to be processed;

[0043] An image completion module, configured to input preset marker information, the image of the human to be processed, and the dense coordinate image to be processed into a trained generation neural network to generate a completed human image;

[0044] Wherein, the preset marker information includes region marker information for the region with pixel loss on the image of the human to be processed; the trained generation neural network is obtained by training according to the model training method described in the first aspect.

[0045] According to a fifth aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the model training method described in the first aspect of the above embodiments or the human image completion method described in the second aspect of the above embodiments.

[0046] According to a sixth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0047] A processor; and

[0048] A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the model training method described in the first aspect of the above embodiments or the human image completion method described in the second aspect of the above embodiments.

[0049] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0050] In the technical solution provided by an embodiment of the present disclosure, the person sample image is preprocessed according to a preset algorithm to obtain a target image, and the person sample image is preprocessed according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss area. Then, the preset random template, the to-be-processed person image, and the to-be-processed dense coordinate image are input into a generation neural network to generate a completed person image and a completed edge image. Subsequently, a discriminant neural network discriminates the generated completed person image and completed edge image according to the person sample image and the target image to obtain a discrimination result. Finally, multiple rounds of adversarial training are performed on the generation neural network and the discriminant neural network according to the discrimination results corresponding to multiple person sample images. Since the to-be-processed dense coordinate image is introduced when training the generation neural network and the discriminant neural network, and the generation neural network generates a completed picture according to the preset random template, the to-be-processed person image, and the to-be-processed dense coordinate image, the corresponding surface texture can be completed according to the postures of different persons, thereby reducing the gap between the completed person image and the person sample image.

[0051] When the generation neural network obtained by using the above model training method completes a to-be-processed person image with a missing pixel area, it can determine different person postures according to the to-be-processed dense coordinate image. Therefore, it can complete the surface texture of the to-be-processed person image according to different person postures, reduce the gap between the completed person image and the original image before the missing part, and improve the completion effect of the person image.

[0052] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0054] Figure 1 Schematically showing a flowchart of a model training method in an exemplary embodiment of the present disclosure;

[0055] Figure 2 Schematically showing a flowchart of a method for preprocessing a person sample image according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss area in an exemplary embodiment of the present disclosure;

[0056] Figure 3 Schematically showing a flowchart of a method for obtaining a to-be-processed edge image in an exemplary embodiment of the present disclosure;

[0057] Figure 4 A flowchart schematically showing a method for multi-round adversarial training of a generation neural network and a discriminant neural network according to discrimination results corresponding to multiple said person sample images in an exemplary embodiment of the present disclosure;

[0058] Figure 5 A flowchart schematically showing a method for completing a person image in an exemplary embodiment of the present disclosure;

[0059] Figure 6 A schematic diagram schematically showing an adversarial neural network including a generation neural network and a discriminant neural network in an exemplary embodiment of the present disclosure;

[0060] Figure 7 Showing a specific person sample image in an exemplary embodiment of the present disclosure;

[0061] Figure 8 Showing a target edge image obtained by extracting edge information of a person sample image in an exemplary embodiment of the present disclosure;

[0062] Figure 9 Showing a dense coordinate image obtained by analyzing the person pose of a person sample image in an exemplary embodiment of the present disclosure;

[0063] Figure 10 A schematic diagram schematically showing region marking information made in a region with pixel loss in a person image to be processed in an exemplary embodiment of the present disclosure;

[0064] Figure 11 A schematic diagram schematically showing edge marking information made in a region with pixel loss in a person image to be processed in an exemplary embodiment of the present disclosure;

[0065] Figure 12 Showing a completed person image generated according to region marking information and edge marking information in an exemplary embodiment of the present disclosure;

[0066] Figure 13 A schematic diagram schematically showing region marking information and edge marking information made in another region with pixel loss in a person image to be processed in an exemplary embodiment of the present disclosure;

[0067] Figure 14 Showing a completed person image generated according to another region marking information and edge marking information in an exemplary embodiment of the present disclosure;

[0068] Figure 15 A schematic diagram schematically showing the composition of a model training device in an exemplary embodiment of the present disclosure;

[0069] Figure 16Schematically show a schematic diagram of the composition of a human image completion device in an exemplary embodiment of the present disclosure;

[0070] Figure 17 Schematically show a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the exemplary embodiments of the present disclosure;

[0071] Figure 18 Schematically show a schematic diagram of a computer-readable storage medium according to some embodiments of the present disclosure. Detailed implementation manners

[0072] Now, the exemplary embodiments will be described more comprehensively with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0073] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0074] The model training method and / or the human image completion method of the exemplary embodiments of the present disclosure can be implemented by a server, that is, the server can execute each step of the following model training method and / or human image completion method. In this case, the corresponding device and module of the model training method and / or the human image completion method can be configured in the server. Additionally, the model training method can be implemented on one server, and the human image completion method can be implemented on another server, that is, model training and model application (human image completion) can be two different servers. However, it is easily understood that model training and model application can be implemented based on the same server, and no special limitation is made in this exemplary embodiment.

[0075] In addition, it should be understood that a terminal device (such as a mobile phone, a tablet, etc.) can also implement each step of the following method, and the corresponding device and module can be configured in the terminal device. In this case, for example, a human image with a pixel loss area can be complemented and processed through the terminal device.

[0076] Figure 1 The flowchart of a model training method in an exemplary embodiment of the present disclosure is schematically shown. Among them, the model refers to an adversarial neural network including a generative neural network and a discriminative neural network. Specifically, it can be an adversarial neural network as shown in Figure 6 the following figure, but it can also be other adversarial neural networks, and the present disclosure does not make special limitations on this.

[0077] Referring to Figure 1 the following figure, the model training method may include the following steps:

[0078] In step S110, the human sample image is preprocessed according to a preset algorithm to obtain a target image, and the human sample image is preprocessed according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss region.

[0079] In an exemplary embodiment of the present disclosure, the target image includes a target edge image. The preprocessing of the human sample image according to the preset algorithm to obtain the target image includes: extracting the edge information in the human sample image according to the preset algorithm to obtain the target edge image. Among them, the preset algorithm can be algorithms such as the canny algorithm, Sobel algorithm, Laplacian algorithm, Marr-Hildreth algorithm, etc. that can be used to extract the edges of images. Through the above algorithms, the edge information in the human sample can be extracted to obtain the target edge image. For example, extracting the edge information of the human image shown in Figure 7 the following figure can obtain the target edge image as shown in Figure 8 the following figure.

[0080] In an exemplary embodiment of the present disclosure, the to-be-processed image includes a to-be-processed human image and a to-be-processed dense coordinate image. The preprocessing of the human sample image according to the preset random template and the preset model to obtain the to-be-processed image with a pixel loss region, referring to Figure 2 the following figure, includes the following steps S210 to S220:

[0081] Step S210, occluding the human sample image according to the preset random template to obtain a to-be-processed human image with a pixel loss region.

[0082] In an exemplary embodiment of the present disclosure, the human sample image is occluded according to a preset random template to simulate the situation where pixel loss regions appear in the image during the processes of image storage, transcoding, and transmission. Among them, the preset random template can be automatically generated according to a random algorithm set by the user or randomly generated by the system, and it is a preset random template with high randomness. By setting the preset random template to occlude the human sample image, it is possible to simulate the situation where uncertain pixel loss regions appear during the processes of image storage, transcoding, and transmission, so as to increase the randomness of training samples and improve the training effect.

[0083] Step S220, perform pose analysis on the to-be-processed human image according to a preset model to obtain the to-be-processed dense coordinate image.

[0084] In an exemplary embodiment of the present disclosure, the preset model can be a DensePose model, an OpenPose model, an AlphaPose model, etc., which are models that can be used to analyze human poses. Through the preset model, the human image can be processed to extract the human pose corresponding to the human image to obtain the to-be-processed dense coordinate image. For example, analyzing Figure 7 the shown human image can obtain Figure 9 the shown dense coordinate image.

[0085] In an exemplary embodiment of the present disclosure, when the to-be-processed image further includes a to-be-processed edge image, referring to Figure 3 shown, the method further includes the following steps S310 to S320:

[0086] Step S310, extract edge information from the human sample image according to a preset algorithm;

[0087] Step S320, extract the edge information occluded by the preset random template from the edge information to obtain the to-be-processed edge image.

[0088] In an exemplary embodiment of the present disclosure, in order to enable the generated neural network to also complete the pixel loss region in the to-be-processed human image based on the edge information, when training the model, the edge information of the pixel loss region corresponding to the to-be-processed human image can be used as the input. In order to obtain the edge information of the pixel loss region corresponding to the to-be-processed human image, the edge information of the human sample image can be extracted according to the preset algorithm proposed in step S110, and then, in the extracted edge information, the edge information of the occluded region can be further extracted according to the preset random template to obtain the edge information of the region occluded by the preset random template, and thus the to-be-processed edge image can be obtained.

[0089] Further, in order to simulate that when actually completing a person image, it is not necessarily possible to obtain the edge information of the pixel loss area corresponding to the complete person image. Therefore, after extracting the edge information occluded by the preset random template from the edge information, the edge information occluded by the preset random template can also be randomly erased, so as to obtain the edge image to be processed.

[0090] Step S120: Input the preset random template and the image to be processed into a generation neural network to generate a completed person image and a completed edge image.

[0091] In an exemplary embodiment of the present disclosure, after inputting the preset random template and the image to be processed into the generation neural network, the generation neural network can determine the area with pixel loss in the image to be processed according to the preset random template, and then complete the area with pixel loss in the person image to be processed based on the dense coordinate image to be processed, generating a completed person image and a completed edge image. Since the input of the generation neural network includes the dense coordinate image to be processed, when the generation neural network generates a completed person image, it will consider the person pose represented by the dense coordinate image to be processed. Therefore, when the generation neural network generates a completed person image, it will complete the surface texture of the person according to the person pose, achieving a better completion effect.

[0092] In an exemplary embodiment of the present disclosure, when the image to be processed further includes an edge image to be processed, after inputting the preset random template and the image to be processed into the generation neural network, the generation neural network can determine the area with pixel loss in the image to be processed according to the random template, and then complete the area with pixel loss in the person image to be processed based on the dense coordinate image to be processed and the edge image to be processed, generating a completed person image and a completed edge image. By using the edge image to be processed as the input of the generation neural network, the generation neural network can generate a completed person image and a completed edge image based on the guidance of the edge image to be processed, forming a completed person image and a completed edge image corresponding to the edge image to be processed.

[0093] Step S130: Input the completed person image, the completed edge image, the person sample image, and the target image into a discrimination neural network to generate a discrimination result.

[0094] In an exemplary embodiment of the present disclosure, the generated completion person image and completion edge image generated by the generation neural network, as well as the person sample image and the target image are input into the discrimination neural network, so that the discrimination neural network discriminates the completion person image and the person sample image, as well as the completion edge image and the target image respectively, to obtain discrimination results. By adding the process of discriminating the completion edge image and the target image, the generation neural network and the discrimination neural network can be trained according to the discrimination results of the completion edge image and the target image, so that the generation neural network can generate a completion edge image consistent with the target edge image.

[0095] Step S140, perform multiple rounds of adversarial training on the generation neural network and the discrimination neural network according to the discrimination results corresponding to multiple said person sample images.

[0096] In an exemplary embodiment of the present disclosure, before performing multiple rounds of adversarial training on the generation neural network and the discrimination neural network according to the discrimination results corresponding to multiple said person sample images, the method further includes: calculating an image loss function based on the completion person image, the completion edge image, the person sample image, and the target image. Wherein, the image loss function refers to the image loss function between the completion person image and the person sample image, and between the completion edge image and the target edge image in the target image.

[0097] In an exemplary embodiment of the present disclosure, the image loss function includes at least one or a combination of multiple of a reconstruction loss function, a content loss function, and a style loss function. Wherein, the reconstruction loss function can be used to represent the difference between the completion person image, the completion edge image and the person sample image, the target image; the content loss function can be used to represent the difference in content between the completion person image, the completion edge image and the person sample image, the target image; the style loss function can be used to represent the difference in style between the completion person image, the completion edge image and the person sample image, the target image.

[0098] For example, if P is used to represent the person sample image or the target image, and P′ correspondingly represents the completion person image and the completion edge image, the reconstruction loss function can be calculated according to the following formula:

[0099] L R (P,P′)=‖P - P′‖ 1 (1);

[0100] The content loss function can be calculated according to the following formula:

[0101]

[0102] Wherein, n refers to the number of layers for extracting features, Refers to the features extracted at different scales, where i is the number of different feature layers; the style loss can be calculated according to the following formula:

[0103]

[0104]

[0105] Among them, G i (x) refers to the Gram matrix of each feature layer, where w and h respectively refer to the horizontal and vertical scales of a certain feature layer. Refers to the features extracted at different scales, where c and c′ are the specific position coordinates on a certain feature layer, i is the number of different feature layers, and n refers to the number of feature extraction layers. It should be noted that the reconstruction loss function, content loss function, and style loss function can also be calculated according to other calculation formulas, and the present disclosure does not make special restrictions on this.

[0106] In an exemplary embodiment of the present disclosure, when performing multiple rounds of adversarial training on the generation neural network and the discriminative neural network based on the calculated image loss function of the complemented person image, the complemented edge image, the person sample image, and the target image, referring to Figure 4 As shown, two training processes can be alternately performed. The specific training process includes the following steps S410 to S420:

[0107] Step S410, training the generation neural network according to the image loss function corresponding to multiple person sample images and the discrimination result.

[0108] Step S420, training the discriminative neural network according to the discrimination result corresponding to multiple person sample images.

[0109] In an exemplary embodiment of the present disclosure, the generation neural network and the discriminative neural network can be alternately trained when performing multiple rounds of adversarial training. It should be noted that when training the generation neural network, it is based on the image loss function corresponding to multiple person sample images and the discrimination result; when training the discriminative neural network, it is based on the discrimination result corresponding to multiple person sample images. On the basis of the existing training of the generation neural network according to the discrimination result of the discriminative neural network, adding the training of the generation neural network based on the image loss function calculated from the complemented person image, the complemented edge image, the person sample image, and the target image can constrain the complemented person image and the complemented edge image generated by the generation neural network, and improve the complementing effect of the generation neural network.

[0110] In addition, in an exemplary embodiment of the present disclosure, a method for complementing a person image is also provided. Referring toFigure 5 As shown in the figure, the method for completing a human image includes the following steps S510 to S520:

[0111] Step S510: Obtain a to-be-processed human image with a pixel loss area, and perform pose analysis on the to-be-processed human image according to a preset model to obtain a to-be-processed dense coordinate image.

[0112] In an exemplary embodiment of the present disclosure, when it is necessary to complete a to-be-processed human image with a pixel loss area, first, it is necessary to perform pose analysis on the to-be-processed human image according to a preset model to obtain a to-be-processed dense coordinate image.

[0113] Step S520: Input the preset marker information, the to-be-processed human image, and the to-be-processed dense coordinate image into a trained generation neural network to generate a completed human image.

[0114] In an exemplary embodiment of the present disclosure, the generation neural network in step S520 refers to the generation neural network trained by the above model training method. When the generation neural network completes the to-be-processed human image, it is necessary to input the preset marker information, the to-be-processed human image, and the to-be-processed dense coordinate image into the trained generation neural network. Among them, the preset marker information may include region marker information for the area with pixel loss on the to-be-processed human image, which is equivalent to the preset random template input into the generation neural network during the model training process. For example, in Figure 7 the human image shown in, mark it to determine the area with pixel loss in the human image, such as Figure 10 shown.

[0115] In an exemplary embodiment of the present disclosure, the preset marker information may further include edge marker information for the area with pixel loss in the to-be-processed human image, which is equivalent to the to-be-processed edge image input into the generation neural network during the model training process. For example, on the basis of the area with pixel loss marked in Figure 10 if the edge marker information is as shown in Figure 11 then a completed human image as shown in Figure 12 can be output; again, if the region marker information and the edge marker information are as shown in Figure 13 then a completed human image as shown in Figure 14The completed human image shown. By inputting the edge marking information on the human image to be processed into the generative neural network, the generative neural network can generate the corresponding completed human image according to the edge marking information. Since the corresponding completed human images generated are different when the edge marking information input by the user is different, the user can change the generated completed human image by inputting different edge marking information, realizing the function of allowing the user to restore or modify the completed human image according to the edge marking information.

[0116] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0117] In addition, in the exemplary embodiment of the present disclosure, a model training device is also provided. Referring to Figure 15 As shown, the model training device 1500 includes: a first processing module 1510, an image generation module 1520, a result discrimination module 1530, and a model training module 1540.

[0118] Among them, the first processing module 1510 can preprocess the human sample image according to a preset algorithm to obtain a target image, and preprocess the human sample image according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss region; wherein, the to-be-processed image includes a to-be-processed human image and a to-be-processed dense coordinate image;

[0119] The image generation module 1520 can be used to input the preset random template and the to-be-processed image into the generative neural network to generate a completed human image and a completed edge image;

[0120] The result discrimination module 1530 can be used to input the completed human image, the completed edge image, the human sample image, and the target image into the discrimination neural network to generate a discrimination result;

[0121] The model training module 1540 can be used to perform multiple rounds of adversarial training on the generative neural network and the discrimination neural network according to the discrimination results corresponding to multiple human sample images.

[0122] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the first processing module 1510 can be used to extract the edge information in the human sample image according to a preset algorithm to obtain a target edge image.

[0123] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the first processing module 1510 may be configured to occlude the person sample image according to a preset random template to obtain a to-be-processed person image with a pixel loss region; perform pose analysis on the to-be-processed person image according to a preset model to obtain the to-be-processed dense coordinate image.

[0124] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the first processing module 1510 may be configured to extract edge information in the person sample image according to the preset algorithm; extract the edge information occluded by the preset random template from the edge information to obtain a to-be-processed edge image.

[0125] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the first processing module 1510 may be configured to randomly erase the edge information occluded by the preset random template.

[0126] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the model training device 1500 further includes a loss calculation module 1550, which may be configured to calculate an image loss function based on the completed person image, the completed edge image, the person sample image, and the target image.

[0127] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the image loss function includes at least one or a combination of a reconstruction loss function, a content loss function, and a style loss function.

[0128] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the model training module 1540 may be configured to train the generative neural network according to the image loss function and the discrimination result corresponding to multiple person sample images; train the discriminative neural network according to the discrimination result corresponding to multiple person sample images.

[0129] In addition, in an exemplary embodiment of the present disclosure, a person image completion device is further provided. Refer to Figure 16 As shown, the person image completion device 1600 includes: an image processing module 1610 and an image completion module 1620.

[0130] Among them, the image processing module 1610 may be configured to obtain a to-be-processed person image with a pixel loss region, and perform pose analysis on the to-be-processed person image according to a preset model to obtain a to-be-processed dense coordinate image;

[0131] The image completion module 1620 can be used to input the preset marker information, the to-be-processed person image, and the to-be-processed dense coordinate image into the trained generation neural network to generate a completed person image; wherein, the preset marker information includes region marker information for regions with pixel loss on the to-be-processed person image; the trained generation neural network is obtained by training according to the model training method.

[0132] In an exemplary embodiment of the present disclosure, based on the foregoing solution, the preset marker information includes edge marker information for regions with pixel loss in the to-be-processed person image.

[0133] Since each functional module of the model training device and the person image completion device in the exemplary embodiments of the present disclosure corresponds to the steps of the exemplary embodiments of the above model training method and person image completion method, for details not disclosed in the device embodiments of the present disclosure, please refer to the embodiments of the above model training method and person image completion method of the present disclosure.

[0134] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0135] In addition, in the exemplary embodiments of the present disclosure, an electronic device capable of implementing the above model training method and person image completion method is also provided.

[0136] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0137] The following refers to Figure 17 to describe the electronic device 1700 according to this embodiment of the present disclosure. Figure 17 The shown electronic device 1700 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0138] As Figure 17As shown, the electronic device 1700 is presented in the form of a general-purpose computing device. The components of the electronic device 1700 may include, but are not limited to: at least one of the above-mentioned processing units 1710, at least one of the above-mentioned storage units 1720, a bus 1730 connecting different system components (including the storage unit 1720 and the processing unit 1710), and a display unit 1740.

[0139] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1710, so that the processing unit 1710 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of the present specification above. For example, the processing unit 1710 can execute steps such as Figure 1 shown in S110: preprocess the person sample image according to a preset algorithm to obtain a target image, and preprocess the person sample image according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss area; wherein, the to-be-processed image includes a to-be-processed person image and a to-be-processed dense coordinate image; S120: input the preset random template and the to-be-processed image into a generation neural network to generate a completed person image and a completed edge image; S130: input the completed person image, the completed edge image, the person sample image, and the target image into a discriminant neural network to generate a discriminant result; S140: perform multiple rounds of adversarial training on the generation neural network and the discriminant neural network according to the discriminant results corresponding to multiple person sample images.

[0140] Again, the above-mentioned electronic device can implement each step as shown in Figures 2 to 5 shown.

[0141] The storage unit 1720 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 1721 and / or a cache storage unit 1722, and may further include a read-only storage unit (ROM) 1723.

[0142] The storage unit 1720 may further include a program / utilities 1724 having a set (at least one) of program modules 1725. Such program modules 1725 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0143] The bus 1730 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.

[0144] The electronic device 1700 can also communicate with one or more external devices 1770 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1700, and / or communicate with any device that enables the electronic device 1700 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 1750. Moreover, the electronic device 1700 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1760. As shown in the figure, the network adapter 1760 communicates with other modules of the electronic device 1700 through the bus 1730. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0145] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software in combination with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0146] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0147] Referring to Figure 18 , a program product 1800 for implementing the above method according to an embodiment of the present disclosure is described. It can adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0148] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0149] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0150] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0151] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0152] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed, for example, synchronously or asynchronously in multiple modules.

[0153] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known or customary technical means in the art not disclosed herein. The specification and examples are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0154] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A model training method, characterized in that, comprising: Preprocessing a character sample image according to a preset algorithm to obtain a target image, and preprocessing the character sample image according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss region; wherein, the target image includes a target edge image; the to-be-processed image includes a to-be-processed character image, a to-be-processed edge image, and a to-be-processed dense coordinate image; Inputting the preset random template and the to-be-processed image into a generation neural network to generate a completed character image and a completed edge image; Inputting the completed character image, the completed edge image, the character sample image, and the target image into a discrimination neural network to generate a discrimination result; Performing multiple rounds of adversarial training on the generation neural network and the discrimination neural network according to the discrimination results corresponding to multiple character sample images; Wherein, the preprocessing the character sample image according to a preset algorithm to obtain a target image includes: Extracting edge information in the character sample image according to a preset algorithm to obtain a target edge image; Wherein, the preprocessing the character sample image according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss region includes: Occluding the character sample image according to a preset random template to obtain a to-be-processed character image with a pixel loss region; Performing pose analysis on the to-be-processed character image according to a preset model to obtain the to-be-processed dense coordinate image; Extracting the edge information occluded by the preset random template from the edge information to obtain a to-be-processed edge image.

2. The method according to claim 1, characterized in that, After extracting the edge information occluded by the preset random template from the edge information, the method further includes: Randomly erasing the edge information occluded by the preset random template.

3. The method according to claim 1, characterized in that, Before performing multiple rounds of adversarial training on the generation neural network and the discrimination neural network according to the discrimination results corresponding to multiple character sample images, the method further includes: Calculating an image loss function based on the completed character image, the completed edge image, the character sample image, and the target image.

4. The method according to claim 3, characterized in that, The performing multiple rounds of adversarial training on the generation neural network and the discrimination neural network according to the discrimination results corresponding to multiple character sample images includes: Alternately performing the following two training processes: Training the generation neural network according to the image loss function and the discrimination results corresponding to multiple character sample images; Training the discrimination neural network according to the discrimination results corresponding to multiple character sample images.

5. The method according to claim 3, characterized in that, The image loss function includes at least one or a combination of multiple of a reconstruction loss function, a content loss function, and a style loss function.

6. A method for completing a character image, characterized in that, comprising: Obtain a to-be-processed human image with a pixel loss area, and perform pose analysis on the to-be-processed human image according to a preset model to obtain a to-be-processed dense coordinate image; Input the preset marking information, the to-be-processed human image, and the to-be-processed dense coordinate image into a trained generation neural network to generate a completed human image; Wherein, the preset marking information includes area marking information for the area with pixel loss on the to-be-processed human image; the trained generation neural network is obtained by training according to the model training method described in any one of claims 1 to 5.

7. According to the method described in claim 6, characterized in that, the preset marking information includes edge marking information for the area with pixel loss in the to-be-processed human image.

8. A model training device, characterized in that, comprises: A first processing module, configured to preprocess a human sample image according to a preset algorithm to obtain a target image, and preprocess the human sample image according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss area; wherein, the target image includes a target edge image; the to-be-processed image includes a to-be-processed human image, a to-be-processed edge image, and a to-be-processed dense coordinate image; An image generation module, configured to input the preset random template and the to-be-processed image into a generation neural network to generate a completed human image and a completed edge image; A result discrimination module, configured to input the completed human image, the completed edge image, the human sample image, and the target image into a discrimination neural network to generate a discrimination result; A model training module, configured to perform multiple rounds of adversarial training on the generation neural network and the discrimination neural network according to the discrimination results corresponding to multiple human sample images; Wherein, the preprocessing of the human sample image according to a preset algorithm to obtain a target image includes: Extracting edge information in the human sample image according to a preset algorithm to obtain a target edge image; Wherein, the preprocessing of the human sample image according to a preset random template and a preset model to obtain a to-be-processed image with a pixel loss area includes: Occluding the human sample image according to a preset random template to obtain a to-be-processed human image with a pixel loss area; Performing pose analysis on the to-be-processed human image according to a preset model to obtain the to-be-processed dense coordinate image; Extracting the edge information occluded by the preset random template from the edge information to obtain a to-be-processed edge image.

9. A human image completion device, characterized in that, comprises: An image processing module, configured to obtain a to-be-processed human image with a pixel loss area, and perform pose analysis on the to-be-processed human image according to a preset model to obtain a to-be-processed dense coordinate image; An image completion module, configured to input the preset marking information, the to-be-processed human image, and the to-be-processed dense coordinate image into a trained generation neural network to generate a completed human image; Among them, the preset marker information includes region marker information for regions with pixel loss on the to-be-processed human image; the trained generation neural network is obtained by training according to the model training method described in any one of claims 1 to 5.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the model training method described in any one of claims 1 to 5 or the human image completion method described in any one of claims 6 to 7.

11. An electronic device, characterized in that it includes: a processor; and a memory for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the model training method described in any one of claims 1 to 5 or the human image completion method described in any one of claims 6 to 7.

Citation Information

Patent Citations

  • Gait recognition method based on generative adversarial image completion network

    CN109753935A

  • Face restoration method based on generative adversarial network

    CN110222628A