Image Augmentation and Neural Network Training Method, Apparatus, Device, and Storage Medium

By acquiring and repairing the defective parts of the three-dimensional image and generating an augmented image set, the problems of insufficient data and inaccurate manual labeling in deep learning model training are solved, and efficient data expansion and labeling accuracy are achieved.

CN111797264BActive Publication Date: 2025-06-17BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201910282291.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-04-09
Publication Date
2025-06-17
Estimated Expiration
2039-04-09

AI Technical Summary

Technical Problem

In the prior art, deep learning face recognition model training requires a large amount of precisely marked data, but the amount of manual labeling is limited. Especially for face images with large postures, manual labeling is difficult and subjective, making it difficult to guarantee the accuracy of labeling.

Method used

By obtaining three-dimensional images carrying the setting key point annotation of the target object, using neural networks for feature extraction and repair, and generating an augmented image set, solving the problems of insufficient training data and inaccurate manual annotation.

Benefits of technology

It realizes automatic and rapid acquisition of augmented images carrying the key point information of the specified target object, expands the target object data from different angles, improves the accuracy of manual annotation under large postures, and thus improves the effect of neural network model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111797264B_ABST
    Figure CN111797264B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides an image augmentation and neural network training method, apparatus, device, and storage medium. A three-dimensional image with set key point annotations of a target object is obtained, where the three-dimensional image is reconstructed from a two-dimensional image of the target object; a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle is obtained, and the defective two-dimensional image includes transformed coordinates corresponding to the set angle of the set key points of the target object; feature extraction is performed on the defective two-dimensional image based on a trained neural network, and the defective two-dimensional image is repaired based on the correspondence between the transformed coordinates of the set key points and the pose to obtain a repaired image corresponding to the defective two-dimensional image; an augmented image set of the target object is obtained based on the repaired image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to a method and device for image augmentation, a method and device for neural network training, a computer device, and a storage medium. Background Art

[0002] Training of deep learning face recognition models requires a large amount of accurately labeled data, but the amount of manually labeled data is very limited. In addition, it is very difficult to manually label large pose face images. Due to self-occlusion and large poses, during the manual labeling process, for the labeling of invisible positions, it is often necessary to guess the positions of key points, which has a certain degree of subjectivity. For example, to determine the position of a person's left mouth corner, the labeling results of different people may have some deviations, and it is difficult to grasp the accuracy of the labeling. In existing face key point databases, there is also relatively little accurate large pose face key point data. Summary of the Invention

[0003] In view of this, the main purpose of the embodiments of the present invention is to provide a method and device for image augmentation, a method and device for neural network training, a computer device, and a storage medium, which can automatically and quickly obtain augmented images carrying key point information of a specified target object.

[0004] To achieve the above object, the technical solution of the embodiments of the present invention is implemented as follows:

[0005] In the first aspect of the embodiments of the present invention, an image augmentation method is provided, including: obtaining a three-dimensional image carrying set key point annotations of a target object, where the three-dimensional image is reconstructed from a two-dimensional image of the target object; obtaining a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle, where the defective two-dimensional image includes conversion coordinates corresponding to the set angle of the set key points of the target object; performing feature extraction on the defective two-dimensional image based on a trained neural network, and repairing the defective two-dimensional image based on the correspondence between the conversion coordinates of the set key points and the pose to obtain a repaired image corresponding to the defective two-dimensional image; and obtaining an augmented image set of the target object based on the repaired image.

[0006] In the second aspect of the embodiments of the present invention, a neural network training method is provided, including: obtaining an augmented image set of a target object by using the image augmentation method provided in any embodiment of the present invention; forming a training sample set according to the two-dimensional image of the target object and the augmented image set; and inputting the training sample set into a neural network model for training until the neural network model converges to obtain the trained neural network model.

[0007] In a third aspect of the embodiments of the present invention, an image augmentation device is provided. The device includes: an acquisition module configured to acquire a three-dimensional image with set key point annotations of a target object, where the three-dimensional image is reconstructed from a two-dimensional image of the target object; a projection module configured to acquire a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle, where the defective two-dimensional image includes conversion coordinates corresponding to the set angle of the set key points of the target object; a first processing module configured to perform feature extraction on the defective two-dimensional image based on a trained neural network, and repair the defective two-dimensional image based on the correspondence between the conversion coordinates of the set key points and the pose to obtain a repaired image corresponding to the defective two-dimensional image; a second processing module configured to obtain an augmented image set of the target object based on the repaired image.

[0008] In a fourth aspect of the embodiments of the present invention, a neural network training device is provided. The device includes: a sample generation module configured to obtain an augmented image set of a target object by using the image augmentation method provided in any embodiment of the present invention, and form a training sample set according to the two-dimensional image of the target object and the augmented image set; a training module configured to input the training sample set into a neural network model for training until the neural network model converges, to obtain the trained neural network model.

[0009] In a fifth aspect of the embodiments of the present invention, a computer device is provided, including: a processor and a memory for storing a computer program that can run on the processor;

[0010] wherein, when the processor is used to run the computer program, it implements the image augmentation method provided in any embodiment of the present invention, or implements the neural network training method provided in any embodiment of the present invention.

[0011] In a sixth aspect of the embodiments of the present invention, a computer storage medium is provided. The computer storage medium stores a computer program, and when the computer program is executed by a processor, it implements the image augmentation method provided in any embodiment of the present invention, or implements the neural network training method provided in any embodiment of the present invention.

[0012] In the above embodiments of the present invention, a three-dimensional image with set key point annotations of a target object is obtained, where the three-dimensional image is reconstructed from a two-dimensional image of the target object; a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle is obtained, and the defective two-dimensional image includes the converted coordinates corresponding to the set angle of the set key points of the target object. Thus, a large number of defective two-dimensional images can be obtained from one original two-dimensional image of the target object; feature extraction is performed on the defective two-dimensional images based on a trained neural network, and the defective two-dimensional images are repaired based on the correspondence between the converted coordinates of the set key points and the pose to obtain a repaired image corresponding to the defective two-dimensional image, so that an augmented image set containing the key point information of the target object can be accurately and efficiently obtained. Thus, an augmented image set of the target object is obtained based on the repaired image, effectively solving the problems of less training data and inaccurate manual annotation in the neural network, expanding the data of the target object at different angles, and improving the accuracy of manual annotation in large poses, thereby further improving the training effect of the neural network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 FIG. is a schematic flow chart of a known process for obtaining 106 key point coordinates;

[0014] Figure 2 FIG. is a schematic flow chart of a known process for generating 106 key point coordinates based on a three-dimensional model;

[0015] Figure 3 FIG. is a schematic flow chart of an image augmentation method provided by an embodiment of the present invention;

[0016] Figure 4 FIG. is a schematic diagram of 106 key points of a human face provided by an embodiment of the present invention;

[0017] Figure 5 FIG. is an example diagram of a three-dimensional standard model provided by an embodiment of the present invention;

[0018] Figure 6 FIG. is a schematic flow chart of a neural network training method provided by an embodiment of the present invention

[0019] Figure 7 FIG. is a schematic structural diagram of an image augmentation device provided by an embodiment of the present invention;

[0020] Figure 8 FIG. is a schematic structural diagram of a neural network training device provided by an embodiment of the present invention;

[0021] Figure 9 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present invention;

[0022] Figure 10 Schematic flowchart of the image augmentation method provided by another embodiment of the present invention. Detailed implementation manners

[0023] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0025] Before further describing the present invention in detail, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations.

[0026] 1) Target object refers to the object contained in the image to be recognized during neural network training, which refers to a human face herein.

[0027] 2) Defective two-dimensional image refers to the two-dimensional image obtained by projection after the three-dimensional image is rotated. Since it is obtained by rotating and projecting the three-dimensional image, it may cause texture holes in the image, and it is called a defective two-dimensional image.

[0028] 3) Three-dimensional standard model refers to a three-dimensional model established from multiple pixel points at multiple specific positions of the target object. Taking the target object as a human face as an example, it refers to a model composed of all pixel points representing specific positions of the human face. For example, if a three-dimensional standard model has 30,000 pixel points, then these 30,000 pixel points are arranged in an orderly manner and each pixel point can represent a specific position on the human face, such as eyes, mouth or nose, etc.

[0029] 4) Two-dimensional training image is a sample image for image training.

[0030] 5) Loss function, also called cost function, is the objective function for neural network optimization;

[0031] 6) Neural Networks (NN) is a complex network system formed by a large number of simple processing units (called neurons) widely interconnected. It reflects many basic characteristics of the human brain function and is a highly complex non-linear dynamic learning system.

[0032] Please refer to Figure 1 , which is a currently known method for generating facial key points. By establishing a three-dimensional facial model, rotating the three-dimensional facial model by a certain angle and performing projection after sampling 68 key points, two-dimensional images containing the coordinates of 68 key points at different angles are obtained. Then, the key points are supplemented again using the interpolation method to obtain 106 key points. This solution will cause a certain degree of loss of key point accuracy during the process of projecting the three-dimensional facial model to obtain two-dimensional images. Secondly, using the interpolation method to supplement the key points to obtain 106 key points will cause a secondary loss of accuracy.

[0033] Therefore, based on the above solution, another solution that reduces the loss of key point accuracy caused by one-dimensional conversion has been proposed. Please refer to Figure 2 , which is another known method for generating facial key points. By establishing a three-dimensional facial model, rotating the three-dimensional facial model by a certain angle and performing projection after sampling 106 key points, two-dimensional images containing the coordinates of 106 key points at different angles are obtained. However, there will be defects in the images generated by this method, which will also cause a certain degree of loss of key point accuracy. Based on the problems existing in the above-known solutions, please refer to Figure 3 , an embodiment of the present invention provides an image augmentation method, which includes the following steps:

[0034] Step 101: Obtain a three-dimensional image carrying the set key point annotation of the target object, where the three-dimensional image is reconstructed from the two-dimensional image of the target object;

[0035] The three-dimensional image is reconstructed from the two-dimensional image of the target object, which means that for the input two-dimensional image containing the target object, by adjusting the combined parameters of the three-dimensional standard model, specifically, it can include a shape expression model and a texture model, to obtain the three-dimensional image of the target object with the highest similarity to the input two-dimensional image; here, the target object refers to a human face, and the three-dimensional image can be a facial model. For other objects corresponding to the target object, the three-dimensional image can also be other object models accordingly.

[0036] A three-dimensional image carrying the set key point annotation of the target object means that the corresponding vertices of the facial key points are included in the three-dimensional image. Here, the key points can be 106 facial key points. Please refer to Figure 4 , the 106 facial key points of the human face are respectively used to represent a certain specific position on the human face, mainly including eyebrows, eyes, mouth or nose, and facial contours, etc. Compared with 68 key points, they can more completely outline the upper and lower edges, contour information of the eyebrows, and information at the nasal wings, so they can more completely describe the contour of the human face and its facial features.

[0037] Step 102: Obtain the defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle. The defective two-dimensional image includes the conversion coordinates corresponding to the set key points of the target object at the set angle.

[0038] Obtaining the defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle means rotating the obtained three-dimensional image in three-dimensional space. For example, placing the three-dimensional image at the origin of the three-dimensional coordinate system, projecting from the front of the three-dimensional image, and rotating the three-dimensional image on the three-dimensional coordinate system, rotating around the coordinate axes X, Y, and Z on the three-dimensional coordinate system respectively, so as to obtain the defective two-dimensional images corresponding to the projections after different set angles.

[0039] The conversion coordinates corresponding to the set key points of the target object at the set angle refer to establishing the correspondence between the two-dimensional defective images and the corresponding three-dimensional images at different angles. Here, for example, the coordinate of a certain key point A on the three-dimensional image at the initial position is A(x1, y1, z1), and after rotating by an angle α, the three-dimensional coordinate of the key point A is obtained as A'(x2, y2, z2), so as to determine the conversion coordinates corresponding to the α angle.

[0040] Step 103: Extract features from the defective two-dimensional image based on the trained neural network, and repair the defective two-dimensional image based on the correspondence between the conversion coordinates of the set key points and the pose to obtain a repaired image corresponding to the defective two-dimensional image.

[0041] Extracting features from the defective two-dimensional image based on the trained neural network means extracting features from the defective two-dimensional image and obtaining the conversion coordinates of the set key points corresponding to the angle of the pose corresponding to each defective two-dimensional image.

[0042] The defective two-dimensional image refers to a two-dimensional image with texture holes caused by the projection after the three-dimensional image is rotated; the repaired image corresponding to the defective two-dimensional image refers to an image obtained by repairing the key point coordinates at the texture hole based on the trained neural network.

[0043] Step 104: Obtain the augmented image set of the target object based on the repaired image.

[0044] Obtaining the augmented image set of the target object based on the repaired image means an image set composed of the repaired images corresponding to each defective two-dimensional image repaired according to each defective two-dimensional image, that is, the augmented image set.

[0045] In the above embodiments of the present invention, by obtaining a three-dimensional image with set key point annotations of a target object, where the three-dimensional image is obtained by reconstructing a two-dimensional image of the target object; obtaining a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle, the defective two-dimensional image includes the conversion coordinates corresponding to the set angle of the set key points of the target object. In this way, a large number of defective two-dimensional images can be obtained through one original two-dimensional image of the target object; based on the trained neural network, feature extraction is performed on the defective two-dimensional images, and the defective two-dimensional images are repaired based on the corresponding relationship between the conversion coordinates of the set key points and the pose, to obtain a repaired image corresponding to the defective two-dimensional image, so that an augmented image set containing the key point information of the target object can be obtained accurately and efficiently. In this way, based on the repaired image, an augmented image set of the target object is obtained, effectively solving the problems of less training data and inaccurate manual annotation in the neural network, expanding the data of the target object at different angles, and improving the accuracy of manual annotation in large poses, thereby further improving the training effect of the neural network model.

[0046] In one embodiment, the obtaining a three-dimensional image with set key point annotations of a target object includes:

[0047] Obtaining a two-dimensional image of the target object, and based on the two-dimensional image and the key point mapping relationship included in the three-dimensional standard model, determining the three-dimensional image corresponding to the two-dimensional image and the three-dimensional coordinates of the set key points.

[0048] A two-dimensional image refers to an original picture containing a target object taken or drawn for reconstructing and obtaining a three-dimensional image. Here, the target object generally refers to a human face, and can also be other objects.

[0049] The three-dimensional standard model refers to a model composed of multiple pixel points representing specific positions on a human face. For example, see Figure 5 , if a three-dimensional standard model has 30,000 pixel points, then the 30,000 pixel points are arranged in an orderly manner and each pixel point can represent a specific position on a human face, such as eyes, mouth or nose, etc.; among them, the key points can be 106-point human face key points, and the three-dimensional standard model can be a three-dimensional human face model constructed using a three-dimensional variable human face model (3DMM).

[0050] The three-dimensional variable human face model (3DMM) is established on the basis of a three-dimensional human face database, with human face shape and human face texture statistics as constraints, and at the same time taking into account the influence of human face pose and illumination factors, and can generate a high-precision three-dimensional human face model. A linear combination of the face data objects in the 3DMM model database. On the basis of the above 3D face representation, assuming that we establish a 3D deformed human face model consisting of m human face models, where each human face model contains the corresponding Si and T i These are two vectors. When representing a new 3D face model, refer to Formulas (1) and (2).

[0051]

[0052]

[0053] where represents the average facial shape model, S i and e i respectively represent the Principal Component Analysis (PCA) parts of shape and expression, and α i and β i respectively represent the corresponding coefficients of shape and expression; the texture model is the same. In this way, a new face model can be linearly combined from the existing facial models. That is to say, three-dimensional images can be generated based on the existing standard face models by changing the coefficients.

[0054] Here, the three-dimensional standard model pre-sets the coordinates corresponding to the key points and the corresponding indexes. Based on the mapping relationship between the key points included in the two-dimensional image and the three-dimensional standard model, it means determining the index value corresponding to each key point in the three-dimensional image in the three-dimensional standard model based on the key points included in the three-dimensional standard model and the corresponding pixel points in the two-dimensional image. Further, determine the three-dimensional coordinates corresponding to the key points included in the three-dimensional image corresponding to the input two-dimensional image based on the three-dimensional standard model.

[0055] In the above implementation manner of the present application, the three-dimensional coordinates of the key points are set for the three-dimensional image set corresponding to the two-dimensional image based on the two-dimensional image and the three-dimensional standard model. In this way, it is ensured that the two-dimensional image obtained by rotating and projecting the three-dimensional image contains the conversion coordinates corresponding to the set key points of the target object at the set angle.

[0056] In one implementation manner, the obtaining of the defective two-dimensional image corresponding to the three-dimensional image after rotating by a set angle includes:

[0057] Based on the conversion coordinates corresponding to the set key points of the target object at the set angle, determine the posture of the target object corresponding to the three-dimensional image after rotating by the set angle, and determine the corresponding projection matrix based on the posture;

[0058] Determine the defective two-dimensional image corresponding to the posture based on the projection matrix.

[0059] Determining the posture of the target object corresponding to the rotated three-dimensional image by a set angle based on the conversion coordinates corresponding to the set key points of the target object and the set angle means determining the corresponding coordinate offset value based on the rotation angle of the three-dimensional image. For example, the coordinate of a key point A on the three-dimensional image at the initial position is A(x1, y1, z1), and after rotating by an angle α, the three-dimensional coordinate of key point A is obtained as A'(x2, y2, z2), that is, determining the posture corresponding to the first posture under the corresponding conversion coordinates.

[0060] The projection matrix refers to the matrix that converts the coordinates in three-dimensional coordinates to the corresponding two-dimensional coordinates. Here, the corresponding defective two-dimensional image and the conversion coordinates corresponding to the set key points of the target object and the set angle can be determined by determining the projection matrix corresponding to the first posture.

[0061] In the above embodiment, different projection matrices are obtained based on the determined posture of the target object corresponding to the rotated three-dimensional image by a set angle, and then the defective two-dimensional images corresponding to the projection matrices are obtained, thus realizing the augmentation processing of a two-dimensional image to obtain a set of defective two-dimensional images.

[0062] In one embodiment, before obtaining the three-dimensional image with the set key points of the target object marked, it further includes:

[0063] Obtaining the original two-dimensional image of the target object as the image to be augmented, and processing the image to be augmented to obtain the processed two-dimensional image of the target object; where the processing includes scaling processing and / or normalization processing.

[0064] Here, the original two-dimensional image of the target object obtained is used as the image to be augmented, and the image to be augmented is scaled, for example, scaled to a fixed size (such as 128*128). The scaled image to be augmented is normalized, such as subtracting the mean or dividing by the variance, to obtain the processed two-dimensional image of the target object. In this way, the influence of the external environment on the image, such as light, noise, rotation, etc., is reduced.

[0065] In one embodiment, the neural network is a generative adversarial network, and the generative adversarial network includes a generator network and a discriminator network; the feature extraction of the defective two-dimensional image based on the trained neural network, and the repair of the defective two-dimensional image based on the corresponding relationship between the conversion coordinates and the posture of the set key points to obtain the repaired image corresponding to the defective two-dimensional image includes:

[0066] Inputting the defective two-dimensional image into the trained generative adversarial network, and obtaining the generated two-dimensional repaired image through the generator network based on the corresponding relationship between the conversion coordinates and the posture of the set key points;

[0067] Input the generated two-dimensional repaired image and the two-dimensional image into the adversarial network to determine the discrimination result of the generated two-dimensional repaired image and the two-dimensional image, and determine the repaired image corresponding to the defective two-dimensional image based on the discrimination result.

[0068] Generative Adversarial Networks (GANs) is a deep learning model. The model generates outputs through the mutual game learning of (at least) two modules in the framework: the generative model and the discriminative model. Here, the generative network corresponds to the generative model in the generative adversarial network, and the adversarial network corresponds to the adversarial model in the generative adversarial network.

[0069] Inputting the defective two-dimensional image into the trained generative adversarial network, and obtaining the generated two-dimensional repaired image through the generative network based on the conversion coordinate and pose correspondence relationship of the set key points means generating an image with repaired texture through the generative network based on the conversion coordinate and pose correspondence relationship of the set key points, that is, the generated two-dimensional repaired image.

[0070] Inputting the generated two-dimensional repaired image and the two-dimensional image into the adversarial network to determine the discrimination result of the generated two-dimensional repaired image and the two-dimensional image means judging whether the key points corresponding to the area to be filled in the two-dimensional repaired image are within the set range, that is, basically consistent, with the key points of the two-dimensional image based on the input of the generated two-dimensional repaired image and the two-dimensional image into the adversarial network. If so, determine that the generated two-dimensional repaired image is the repaired image corresponding to the defective two-dimensional image.

[0071] In the above embodiment, the generated two-dimensional repaired image corresponding to the defective two-dimensional image is generated based on the generative network, and judgment is performed based on the adversarial network, so as to obtain the repaired image corresponding to the defective two-dimensional image. In this way, the repair of the texture hole caused by the rotation projection of the three-dimensional image is realized.

[0072] In one embodiment, before obtaining the three-dimensional image with the set key point annotation of the target object, it includes:

[0073] Reconstruct a three-dimensional training image based on the two-dimensional image containing the target object, and the three-dimensional training image carries the set key point annotation label of the target object;

[0074] Obtain a set of two-dimensional training images corresponding to the projections of the three-dimensional training image after rotating different set angles respectively.

[0075] Here, the two-dimensional image can be a two-dimensional face image, and a corresponding three-dimensional training image with a set key point annotation label of the target object is obtained based on the two-dimensional image.

[0076] Obtaining a two-dimensional training image set of multiple two-dimensional training images respectively projected after rotating the three-dimensional training image by different set angles means obtaining a two-dimensional training image set composed of multiple two-dimensional training images by rotating and projecting the three-dimensional training image according to different set angles. For example, placing the three-dimensional training image at the origin of the three-dimensional coordinate system, projecting from the front of the three-dimensional training image, rotating the three-dimensional training image on the three-dimensional coordinate, and rotating around the coordinate axes X, Y, and Z on the three-dimensional coordinate respectively, so as to obtain a two-dimensional training image set corresponding to the projections after different set angles.

[0077] In the above embodiment, the three-dimensional training image reconstructed from the two-dimensional image, and a two-dimensional image training set of multiple two-dimensional training images respectively projected after rotating the three-dimensional training image by different set angles. In this way, the heavy dependence of the face key point positioning task on a large amount of labeled data and the data preparation time before training are greatly reduced, and a two-dimensional image training set can be automatically obtained through one two-dimensional image.

[0078] In one embodiment, before obtaining the three-dimensional image with the set key point annotation of the target object, it further includes:

[0079] Inputting the two-dimensional training image into an initial generative adversarial network, and obtaining a corresponding generated training repair image through the generative network based on the corresponding relationship between the transformed coordinates and postures of the set key points;

[0080] Inputting the two-dimensional training image and the generated training repair image into the adversarial network, and determining the discrimination result of the two-dimensional image and the generated training repair image;

[0081] Based on the discrimination result, performing separate alternating iterations on the generative adversarial network until the set loss function satisfies the convergence condition, and obtaining the trained generative adversarial network.

[0082] Here, the loss function is also called the cost function, which is the objective function for neural network optimization. The process of neural network training or optimization is the process of minimizing the loss function. The smaller the loss function value, the closer the corresponding predicted result is to the true result. In the present invention, the loss function may include an adversarial loss function and a reconstruction loss function.

[0083] Determine the discrimination result of the two-dimensional training image and the generated training repaired image. If the set conditions are not met, perform separate alternating iterations on the generative adversarial network until the set loss function meets the convergence condition, and obtain the trained generative adversarial network.

[0084] Here, performing separate alternating iterations on the generative adversarial network until the set loss function meets the convergence condition means updating the parameters of the generative network with the two-dimensional training image obtained by resampling, re-obtaining the generated training repaired image, and then inputting the two-dimensional training image and the generated training repaired image into the adversarial network until the set loss function meets the convergence condition, and obtaining the trained generative network and the trained adversarial network. Specifically, using the neural network backpropagation algorithm, iteratively update the values of the parameters of the generative network and the adversarial network. First, update the parameters of the adversarial network, and then update the parameters of the generative network with the training color patches obtained by resampling until the set loss function meets the convergence condition, and obtain the trained generative adversarial network. In this way, the trained generative network and the trained adversarial network are obtained through alternating iterative training, and the trained generative adversarial network is used to repair the texture holes caused by rotation, reducing human error and high labor costs.

[0085] In one embodiment, before performing separate alternating iterations on the generative adversarial network based on the discrimination result until the set loss function meets the convergence condition, it further includes:

[0086] According to the combination of the adversarial loss function and the reconstruction loss function, obtain the loss function corresponding to the generative adversarial network.

[0087] Here, the loss function includes two parts, the adversarial loss function and the reconstruction loss function. Refer to formulas (3) and (4) respectively;

[0088] L adv =E x [logD(x)]+E x [log(1-(D(G(x))))] (3)

[0089] L rec =E x [w⊙(x-G(x))1] (4)

[0090] Among them, L adv is the adversarial loss function, L recLet \(L_{rec}\) be the reconstruction loss function, \(G\) be the generator network, \(D\) be the discriminator network, \(x\) be the input sample, i.e., the image with texture holes after rotation, and \(G(x)\) be the generated image based on the input image \(x\), i.e., the image after texture repair. The purpose of the adversarial loss is to make the generated image more realistic and natural. A weight coefficient \(w\) is introduced into the reconstruction loss function, and its purpose is to make the areas of the repaired image other than the filling area as consistent with the original image as possible, so as to ensure the correctness of the key point coordinates.

[0091] In another embodiment, as Figure 6 shown, a neural network training method is also provided, including:

[0092] Step 201: Obtain an augmented image set of the target object; form a training sample set according to the two-dimensional image of the target object and the augmented image set; wherein, the augmented image set of the target object can be obtained by using the image augmentation method provided in any embodiment of the present invention.

[0093] Step 202: Input the training sample set into the neural network model for training until the neural network model converges, and obtain the trained neural network model.

[0094] Here, a corresponding augmented image set is formed based on the two-dimensional image of the target object as the training sample of the neural network, so as to ensure that more effective training samples can be obtained more quickly to realize the training of the neural network and improve the classification accuracy of the trained neural network. Here, since the augmented image includes the accurate positioning of the set key points, the augmented image set can be used to train the neural network to obtain the trained neural network model, and this neural network model can be used in application scenarios such as expression recognition, animation synthesis, live broadcast, beauty, and special effects camera.

[0095] In another embodiment, as Figure 7 shown, an image augmentation device is also provided, and the device includes:

[0096] An acquisition module 31, configured to acquire a three-dimensional image carrying the set key point annotation of the target object, where the three-dimensional image is reconstructed from the two-dimensional image of the target object;

[0097] A projection module 32, configured to acquire a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle, where the defective two-dimensional image includes the conversion coordinates corresponding to the set angle of the set key points of the target object;

[0098] The first processing module 33 is configured to extract features from the defective two-dimensional image based on the trained neural network, and repair the defective two-dimensional image based on the conversion coordinate and pose correspondence relationship of the set key points, so as to obtain a repaired image corresponding to the defective two-dimensional image;

[0099] The second processing module 34 is configured to obtain an augmented image set of the target object based on the repaired image.

[0100] In the above embodiments of the present application, by obtaining a three-dimensional image with set key point annotations of a target object, where the three-dimensional image is reconstructed from a two-dimensional image of the target object; obtaining a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle, the defective two-dimensional image includes the conversion coordinates corresponding to the set angle of the set key points of the target object; thus, a large number of defective two-dimensional images can be automatically and quickly obtained through a single two-dimensional image; extracting features from the defective two-dimensional image based on the trained neural network, and repairing the defective two-dimensional image based on the conversion coordinate and pose correspondence relationship of the set key points to obtain a repaired image corresponding to the defective two-dimensional image; obtaining an augmented image set of the target object based on the repaired image. Thus, obtaining the augmented image set of the target object based on the repaired image effectively solves the problems of few training data and inaccurate manual annotation in the neural network, can expand face data at different angles, and improves the accuracy of manual annotation in large poses, thereby further improving the training effect of the neural network model.

[0101] Optionally, the acquisition module 31 is further configured to obtain a two-dimensional image of the target object, and determine a three-dimensional image corresponding to the two-dimensional image and the three-dimensional coordinates of the set key points based on the key point mapping relationship included in the two-dimensional image and the three-dimensional standard model.

[0102] Optionally, the projection module 32 is further configured to determine the pose corresponding to the target object after the three-dimensional image is rotated by the set angle based on the conversion coordinates corresponding to the set angle of the set key points of the target object, and determine a corresponding projection matrix based on the pose; determine a defective two-dimensional image corresponding to the pose based on the projection matrix.

[0103] Optionally, the acquisition module 31 is further configured to obtain the original two-dimensional image of the target object as the image to be augmented, and process the image to be augmented to obtain the processed two-dimensional image of the target object; where the processing includes scaling processing and / or normalization processing.

[0104] Optionally, the first processing module 33 is further configured to input the defective two-dimensional image into the trained generative adversarial network, and obtain the generated two-dimensional repaired image through the generative network based on the conversion coordinate and pose correspondence relationship of the set key points; input the generated two-dimensional repaired image and the two-dimensional image into the adversarial network, determine the discrimination result of the generated two-dimensional repaired image and the two-dimensional image, and determine the repaired image corresponding to the defective two-dimensional image based on the discrimination result.

[0105] Optionally, the obtaining module 31 is further configured to reconstruct a three-dimensional training image based on a two-dimensional image including a target object, where the three-dimensional training image carries a set key point annotation label of the target object; obtain a two-dimensional training image set corresponding to projections of the three-dimensional training image after rotating different set angles.

[0106] Optionally, the first processing module 33 is further configured to input the two-dimensional training image into the initial generative adversarial network, and obtain the corresponding generated two-dimensional training image through the generative network based on the conversion coordinate and pose correspondence relationship of the set key points; input the two-dimensional image and the generated two-dimensional training image into the adversarial network, and determine the discrimination result of the two-dimensional image and the generated two-dimensional training image; perform separate alternating iterations on the generative adversarial network based on the discrimination result until the set loss function meets the convergence condition, and obtain the trained generative adversarial network.

[0107] Optionally, the first processing module 33 is further configured to obtain the loss function corresponding to the generative adversarial network according to the combination of the adversarial loss function and the reconstruction loss function.

[0108] In another embodiment, as Figure 8 shown, there is also provided a neural network training device, where the device includes:

[0109] A sample generation module 41, configured to obtain an augmented image set of a target object by using the image augmentation method provided in any embodiment of the present invention, and form a training sample set according to the two-dimensional image of the target object and the augmented image set;

[0110] A training module 42, configured to input the training sample set into a neural network model for training until the neural network model converges, and obtain the trained neural network model.

[0111] In another embodiment, as Figure 9 shown, there is also provided a computer device, including: at least one processor 210 and a memory 211 for storing a computer program that can run on the processor 210; where Figure 9The processor 210 shown in the figure does not refer to the number of processors being one, but only refers to the positional relationship of the processor relative to other devices. In actual applications, the number of processors can be one or more; similarly, Figure 9 The memory 211 shown in the figure has the same meaning, that is, it only refers to the positional relationship of the memory relative to other devices. In actual applications, the number of memories can be one or more.

[0112] Wherein, when the processor 210 is used to run the computer program, the following steps are executed:

[0113] Obtain a three-dimensional image carrying the set key point annotations of the target object, where the three-dimensional image is reconstructed from a two-dimensional image of the target object; obtain a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle, and the defective two-dimensional image includes the converted coordinates corresponding to the set key points of the target object and the set angle; perform feature extraction on the defective two-dimensional image based on the trained neural network, and repair the defective two-dimensional image based on the correspondence between the converted coordinates of the set key points and the pose to obtain a repaired image corresponding to the defective two-dimensional image; obtain an augmented image set of the target object based on the repaired image.

[0114] In an alternative embodiment, when the processor 210 is further used to run the computer program, the following steps are executed:

[0115] Obtain a two-dimensional image of the target object, and determine the three-dimensional image corresponding to the two-dimensional image and the three-dimensional coordinates of the set key points based on the mapping relationship between the two-dimensional image and the key points included in the three-dimensional standard model.

[0116] In an alternative embodiment, when the processor 210 is further used to run the computer program, the following steps are executed:

[0117] Determine the pose corresponding to the target object after the three-dimensional image is rotated by the set angle based on the converted coordinates corresponding to the set key points of the target object and the set angle, and determine the corresponding projection matrix based on the pose; determine the defective two-dimensional image corresponding to the pose based on the projection matrix.

[0118] In an alternative embodiment, when the processor 210 is further used to run the computer program, the following steps are executed:

[0119] Obtain the original two-dimensional image of the target object as the image to be augmented, and process the image to be augmented to obtain the processed two-dimensional image of the target object; wherein the processing includes scaling processing and / or normalization processing.

[0120] In an alternative embodiment, when the processor 210 is further configured to run the computer program, the following steps are performed:

[0121] Input the defective two-dimensional image into the trained generative adversarial network. Through the generative network, obtain the generated two-dimensional repaired image based on the correspondence between the transformed coordinates of the set key points and the pose. Input the generated two-dimensional repaired image and the two-dimensional image into the adversarial network, determine the discrimination result of the generated two-dimensional repaired image and the two-dimensional image, and determine the repaired image corresponding to the defective two-dimensional image based on the discrimination result.

[0122] In an alternative embodiment, when the processor 210 is further configured to run the computer program, the following steps are performed:

[0123] Reconstruct a three-dimensional training image based on a two-dimensional image containing the target object. The three-dimensional training image carries the set key point annotation labels of the target object. Obtain a set of two-dimensional training images corresponding to the projections after rotating the three-dimensional training image by different set angles.

[0124] In an alternative embodiment, when the processor 210 is further configured to run the computer program, the following steps are performed:

[0125] Input the two-dimensional training image into the initial generative adversarial network. Through the generative network, obtain the corresponding generated training repaired image based on the correspondence between the transformed coordinates of the set key points and the pose. Input the two-dimensional training image and the generated training repaired image into the adversarial network, determine the discrimination result of the two-dimensional training image and the generated training repaired image. Based on the discrimination result, perform separate alternating iterations on the generative adversarial network until the set loss function meets the convergence condition, and obtain the trained generative adversarial network.

[0126] In an alternative embodiment, when the processor 210 is further configured to run the computer program, the following steps are performed:

[0127] Obtain the loss function corresponding to the generative adversarial network according to the combination of the adversarial loss function and the reconstruction loss function.

[0128] In an alternative embodiment, when the processor 210 is further configured to run the computer program, the following steps are performed:

[0129] Obtain an augmented image set of the target object, and form a training sample set according to the two-dimensional image of the target object and the augmented image set;

[0130] Input the training sample set into a neural network model for training until the neural network model converges, and obtain the trained neural network model.

[0131] The computer device may further include: at least one network interface 212. Each component in the sending end is coupled together through a bus system 213. It can be understood that the bus system 213 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 213 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, Figure 5 all kinds of buses are labeled as the bus system 213 in

[0132] Among them, the memory 211 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM, Static Random Access Memory), synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), dynamic random access memory (DRAM, Dynamic Random Access Memory), synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 211 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memories.

[0133] The memory 211 in the embodiments of the present invention is used to store various types of data to support the operations of the sending end. Examples of such data include: any computer programs for operating on the sending end, such as operating systems and application programs. Among them, the operating system contains various system programs, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs can include various application programs for implementing various application services. Here, the program for implementing the method of the embodiments of the present invention can be included in the application programs.

[0134] This embodiment also provides a computer storage medium, for example, including a memory 211 storing a computer program, and the above computer program can be executed by a processor 210 in the sending end to complete the steps described in the foregoing method. The computer storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM; it can also be various devices including one or any combination of the above memories, such as smart phones, tablet computers, laptop computers, etc. A computer storage medium stores a computer program, and when the computer program is run by a processor, the following steps are executed:

[0135] Among them, when the processor 210 is used to run the computer program, the following steps are executed:

[0136] Obtain a three-dimensional image carrying the set key point annotations of the target object, where the three-dimensional image is obtained by reconstructing a two-dimensional image of the target object; obtain a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle, and the defective two-dimensional image includes the converted coordinates corresponding to the set angle of the set key points of the target object; perform feature extraction on the defective two-dimensional image based on a trained neural network, and repair the defective two-dimensional image based on the correspondence between the converted coordinates of the set key points and the pose to obtain a repaired image corresponding to the defective two-dimensional image; obtain an augmented image set of the target object based on the repaired image.

[0137] In an optional embodiment, when the computer program is run by a processor, the following steps are further executed:

[0138] Obtain a two-dimensional image of the target object, and determine a three-dimensional image corresponding to the two-dimensional image and the three-dimensional coordinates of the set key points based on the key point mapping relationship included in the three-dimensional standard model.

[0139] In an optional embodiment, when the computer program is run by a processor, the following steps are further performed:

[0140] Based on the conversion coordinates corresponding to the set angle of the set key points of the target object, determine the posture of the target object corresponding to the three-dimensional image rotated by the set angle, and determine the corresponding projection matrix based on the posture; determine the defective two-dimensional image corresponding to the posture based on the projection matrix.

[0141] In an optional embodiment, when the computer program is run by a processor, the following steps are further performed:

[0142] Obtain the original two-dimensional image of the target object as the image to be augmented, and process the image to be augmented to obtain the processed two-dimensional image of the target object; wherein the processing includes scaling processing and / or normalization processing.

[0143] In an optional embodiment, when the computer program is run by a processor, the following steps are further performed:

[0144] Input the defective two-dimensional image into the trained generative adversarial network, and obtain the generated two-dimensional repaired image through the generative network based on the correspondence between the conversion coordinates of the set key points and the posture; input the generated two-dimensional repaired image and the two-dimensional image into the adversarial network, determine the discrimination result of the generated two-dimensional repaired image and the two-dimensional image, and determine the repaired image corresponding to the defective two-dimensional image based on the discrimination result.

[0145] In an optional embodiment, when the computer program is run by a processor, the following steps are further performed:

[0146] Reconstruct a three-dimensional training image based on a two-dimensional image containing the target object, and the three-dimensional training image carries a set key point annotation label of the target object; obtain a two-dimensional training image set corresponding to the projections respectively after rotating the three-dimensional training image by different set angles.

[0147] In an optional embodiment, when the computer program is run by a processor, the following steps are further performed:

[0148] Input the two-dimensional training image into the initial generative adversarial network, and obtain the corresponding generated two-dimensional training image through the generative network based on the correspondence between the conversion coordinates of the set key points and the posture; input the two-dimensional training image and the generated training repaired image into the adversarial network, determine the discrimination result of the two-dimensional training image and the generated training repaired image; perform separate alternating iterations on the generative adversarial network based on the discrimination result until the set loss function meets the convergence condition, and obtain the trained generative adversarial network.

[0149] In an optional embodiment, when the computer program is run by a processor, the following steps are further performed:

[0150] According to the combination of the adversarial loss function and the reconstruction loss function, the loss function corresponding to the generative adversarial network is obtained.

[0151] In an optional embodiment, when the computer program is run by a processor, the following steps are further performed:

[0152] An augmented image set of the target object is obtained, and a training sample set is formed according to the two-dimensional image of the target object and the augmented image set;

[0153] The training sample set is input into a neural network model for training until the neural network model converges, and the trained neural network model is obtained.

[0154] Please refer to Figure 6 , taking 3DMM as the face reconstruction method and the neural network as the conditional adversarial generative network as an example, a more detailed example is used to further elaborate on the image augmentation method of the embodiments of the present application. The image augmentation method includes the following steps:

[0155] S11: Obtain a two-dimensional image;

[0156] Here, obtaining a two-dimensional image means obtaining a two-dimensional image containing a face image;

[0157] S12: Generate a three-dimensional image;

[0158] Here, generating a three-dimensional image means determining a three-dimensional image corresponding to the two-dimensional image and setting the three-dimensional coordinates of key points based on the key point mapping relationship included in the two-dimensional image and the three-dimensional standard model.

[0159] S13: 106-point key point sampling;

[0160] Here, 106-point key point sampling is based on the three-dimensional image to obtain the three-dimensional coordinates of the set key points.

[0161] Here, after obtaining the 106-point key points of the three-dimensional image, steps S14 and S16 are respectively performed;

[0162] S14: Project the three-dimensional image;

[0163] Here, projecting the three-dimensional image means obtaining the corresponding two-dimensional image based on the three-dimensional image;

[0164] S15: Obtain the 106-point key point coordinates of the two-dimensional image;

[0165] Here, obtaining the key point coordinates of 106 points of the two-dimensional image means determining the corresponding key point coordinates of 106 points of the two-dimensional image based on the set three-dimensional coordinates of the key points and the corresponding projection matrix; the two-dimensional image here is the key point coordinates of 106 points corresponding to the original image.

[0166] S16: Rotate the three-dimensional image.

[0167] Here, rotating the three-dimensional image means rotating the three-dimensional image by a set angle to obtain corresponding different poses.

[0168] S17: Project the three-dimensional image after rotating by the set angle.

[0169] Here, projecting the three-dimensional image after rotating by the set angle means obtaining the corresponding two-dimensional missing image based on the different poses of the three-dimensional image rotated by the set angle; for example, placing the three-dimensional image at the origin of the three-dimensional coordinate system and projecting it from the front of the three-dimensional image, rotating the three-dimensional image on the three-dimensional coordinate system, and rotating around the coordinate axes X, Y, and Z on the three-dimensional coordinate system respectively, so as to obtain the corresponding defective two-dimensional images after different set angles of projection.

[0170] Here, after projecting the three-dimensional image after rotating by a set angle, steps S18 and S20 are respectively executed.

[0171] S18: Repair the image based on the trained generative adversarial network.

[0172] Here, repairing the image based on the trained generative adversarial network means extracting features from the defective two-dimensional image based on the trained generative adversarial network, and repairing the defective two-dimensional image based on the conversion coordinate and pose correspondence relationship of the set key points to obtain a repaired image corresponding to the defective two-dimensional image.

[0173] S19: Obtain the augmented image set after rotation and repair.

[0174] Here, obtaining the augmented image set after rotation and repair means that the augmented image set of the target object obtained based on the repaired image means an image set composed of a large number of repaired images corresponding to each defective two-dimensional image obtained by repairing each defective two-dimensional image, that is, the augmented image set.

[0175] S20: Rotate the 106 key point coordinates of the augmented image.

[0176] Here, rotating the 106 key point coordinates of the augmented image means obtaining the corresponding 106 key point coordinates respectively based on the defective two-dimensional images obtained in step S17.

[0177] Compared with the prior art, the above embodiments of the present application solve at least the following problems: on the one hand, it can reduce the heavy dependence on labeled data and the data preparation time before training in face key point localization, and collect a large amount of training data for model training; on the other hand, it effectively solves the problems of less training data and inaccurate manual annotation in large poses, can expand face data at different angles, and improves the accuracy of manual annotation in large poses, thereby further improving the effect of model training.

[0178] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.

Claims

1. An image augmentation method, characterized in that, Including: Obtaining a three-dimensional image carrying the set key point annotations of the target object, where the three-dimensional image is obtained by reconstructing a two-dimensional image of the target object; Obtaining a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle, where the defective two-dimensional image includes the converted coordinates corresponding to the set angle of the set key points of the target object; Performing feature extraction on the defective two-dimensional image based on a trained neural network, and repairing the defective two-dimensional image based on the correspondence between the converted coordinates of the set key points and the pose, to obtain a repaired image corresponding to the defective two-dimensional image; Obtaining an augmented image set of the target object based on the repaired image; Wherein, the neural network is a generative adversarial network, and the generative adversarial network includes a generator network and a discriminator network; the performing feature extraction on the defective two-dimensional image based on a trained neural network, and repairing the defective two-dimensional image based on the correspondence between the converted coordinates of the set key points and the pose, to obtain a repaired image corresponding to the defective two-dimensional image, includes: Inputting the defective two-dimensional image into the trained generative adversarial network, and obtaining a generated two-dimensional repaired image through the generator network based on the correspondence between the converted coordinates of the set key points and the pose; Inputting the generated two-dimensional repaired image and the two-dimensional image into the discriminator network, determining the discrimination result of the generated two-dimensional repaired image and the two-dimensional image, and determining the repaired image corresponding to the defective two-dimensional image based on the discrimination result.

2. The image augmentation method according to claim 1, characterized in that, The obtaining a three-dimensional image carrying the set key point annotations of the target object includes: Obtaining a two-dimensional image of the target object, and determining the three-dimensional image corresponding to the two-dimensional image and the three-dimensional coordinates of the set key points based on the two-dimensional image and the key point mapping relationship included in the three-dimensional standard model.

3. The image augmentation method according to claim 1, characterized in that, The obtaining a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle includes: Based on the converted coordinates corresponding to the set angle of the set key points of the target object, determining the pose corresponding to the target object after the three-dimensional image is rotated by the set angle, and determining the corresponding projection matrix based on the pose; Determining the defective two-dimensional image corresponding to the pose based on the projection matrix.

4. The image augmentation method according to claim 1, characterized in that, Before the obtaining a three-dimensional image carrying the set key point annotations of the target object, it further includes: Obtaining the original two-dimensional image of the target object as the image to be augmented, and processing the image to be augmented to obtain the processed two-dimensional image of the target object; where the processing includes scaling processing and / or normalization processing.

5. The image augmentation method according to claim 1, characterized in that, Before the obtaining a three-dimensional image carrying the set key point annotations of the target object, it includes: Reconstructing a three-dimensional training image from a two-dimensional image including the target object, where the three-dimensional training image carries the set key point annotation labels of the target object; Obtaining a two-dimensional training image set of a plurality of two-dimensional training images respectively corresponding to the projections after the three-dimensional training image is rotated by different set angles.

6. The image augmentation method according to claim 5, characterized in that, Before the obtaining a three-dimensional image carrying the set key point annotations of the target object, it further includes: Input the two-dimensional training image into an initial generative adversarial network, and obtain a corresponding generated training repaired image through the generative network based on the conversion coordinate and pose correspondence relationship of the set key points; Input the two-dimensional training image and the training repaired image into the adversarial network, and determine the discrimination results of the two-dimensional training image and the training repaired image; Based on the discrimination results, perform separate alternating iterations on the generative adversarial network until the set loss function meets the convergence condition, and obtain the trained generative adversarial network.

7. The image augmentation method according to claim 6, characterized in that, Before performing separate alternating iterations on the generative adversarial network based on the discrimination results until the set loss function meets the convergence condition, it further includes: According to the combination of the adversarial loss function and the reconstruction loss function, obtain the loss function corresponding to the generative adversarial network.

8. A neural network training method, characterized in that, It includes: Obtain an augmented image set of the target object by using the image augmentation method described in any one of claims 1 to 7, and form a training sample set according to the two-dimensional image of the target object and the augmented image set; Input the training sample set into a neural network model for training until the neural network model converges, and obtain the trained neural network model.

9. An image augmentation device, characterized in that, The device includes: An acquisition module, configured to acquire a three-dimensional image with set key point annotations of the target object, where the three-dimensional image is reconstructed from the two-dimensional image of the target object; A projection module, configured to acquire a defective two-dimensional image corresponding to the projection after the three-dimensional image is rotated by a set angle, where the defective two-dimensional image includes the conversion coordinates corresponding to the set angle of the set key points of the target object; A first processing module, configured to perform feature extraction on the defective two-dimensional image based on the trained neural network, and repair the defective two-dimensional image based on the conversion coordinate and pose correspondence relationship of the set key points to obtain a repaired image corresponding to the defective two-dimensional image; A second processing module, configured to obtain an augmented image set of the target object based on the repaired image; Wherein, the neural network is a generative adversarial network, and the generative adversarial network includes a generative network and an adversarial network; the first processing module is specifically configured to: Input the defective two-dimensional image into the trained generative adversarial network, and obtain a generated two-dimensional repaired image through the generative network based on the conversion coordinate and pose correspondence relationship of the set key points; Input the generated two-dimensional repaired image and the two-dimensional image into the adversarial network, determine the discrimination results of the generated two-dimensional repaired image and the two-dimensional image, and determine the repaired image corresponding to the defective two-dimensional image based on the discrimination results.

10. A neural network training device, characterized in that, It includes: A sample generation module, configured to obtain an augmented image set of the target object by using the image augmentation method described in any one of claims 1 to 7, and form a training sample set according to the two-dimensional image of the target object and the augmented image set; A training module, configured to input the training sample set into a neural network model for training until the neural network model converges, and obtain the trained neural network model.

11. A computer device, characterized in that, It includes: A processor and a memory for storing a computer program that can run on the processor; Wherein, when the processor is used to run the computer program, the image augmentation method described in any one of claims 1 to 7 is implemented, or the neural network training method described in claim 8 is implemented.

12. A computer storage medium, characterized in that, A computer program is stored in the computer storage medium, wherein the computer program, when executed by a processor, implements the image augmentation method described in any one of claims 1 to 7, or implements the neural network training method described in claim 8.

Citation Information

Patent Citations

  • Gesture image generating method and device and storage medium

    CN108346168A