Image generation method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202211730112.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-12-30
AI Technical Summary
[0003]本发明实施例提供一种图像生成方法,旨在解决现有困难样本的获取难度较高,使得训练过程中可用的困难样本数量较少,影响深度神经网络训练效率的问题
[0036]本发明实施例中,获取原始图像的前景分布特征以及与所述前景分布特征对应的前景图像;将所述原始图像的前景分布特征输入到预设的图像生成网络中进行背景图像生成,得到所述原始图像对应的背景生成图像;将所述前景图像与所述原始图像对应的背景生成图像进行图像融合,得到目标图像。通过提取原始图像的前景分布特征来进行背景图像生成,得到对应的背景生成图像,将前景分布特征对应的前景图像与背景生成图像进行融合,得到对应的目标图像,由于背景生成图像是根据前景分布特征进行生成,使得目标图像中背景包含有前景图像的部分隐含特征,从而增加目标图像中前景与背景的识别难度,使得目标图像可以作为困难样本,进而降低了困难样本的获取难度,提高深度神经网络的训练效率。
Smart Images

Figure CN116091766B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and more particularly to an image generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] In image recognition tasks based on deep neural networks, image samples are needed to train the deep neural network. Image samples that are difficult to recognize are called hard samples, while those that are easy to recognize are called easy samples. Hard samples generate significant error loss during training, resulting in more pronounced gradient descent and thus increasing the training speed. The more difficult the sample, the more significant the improvement in the recognition performance of the deep neural network, and the faster the training speed. However, the difficulty in obtaining hard samples limits the number of usable hard samples during training, affecting the training efficiency of the deep neural network. Summary of the Invention
[0003] This invention provides an image generation method aimed at addressing the problem that the difficulty in obtaining difficult samples in existing methods results in a limited number of usable difficult samples during training, thus affecting the training efficiency of deep neural networks. The method generates a background image by extracting the foreground distribution features of the original image. The foreground image corresponding to the foreground distribution features is then fused with the generated background image to obtain the corresponding target image. Since the generated background image is based on the foreground distribution features, the background in the target image contains some of the implicit features of the foreground image, thereby increasing the difficulty of distinguishing between the foreground and background in the target image. This allows the target image to be used as a difficult sample, reducing the difficulty of obtaining difficult samples and improving the training efficiency of deep neural networks.
[0004] In a first aspect, embodiments of the present invention provide an image generation method, the method comprising:
[0005] Obtain the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features;
[0006] The foreground distribution features of the original image are input into a preset image generation network to generate a background image, thereby obtaining a background generated image corresponding to the original image;
[0007] The foreground image is fused with the background image corresponding to the original image to obtain the target image.
[0008] Optionally, obtaining the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features includes:
[0009] Obtain the original image;
[0010] Foreground features are extracted from the original image using a preset feature extraction network to obtain the foreground distribution features of the original image;
[0011] Extract the foreground image corresponding to the foreground distribution features from the original image.
[0012] Optionally, before inputting the foreground distribution features of the original image into a preset image generation network to generate a background image and obtain the background generated image corresponding to the original image, the method further includes:
[0013] Acquire a dataset and a model to be trained. The dataset includes sample images, and the model to be trained includes a generator network, a discriminator network, and a feature extraction network. The output of the feature extraction network is connected to the input of the generator network, and the output of the generator network is connected to the input of the discriminator network.
[0014] The trained generative network is obtained by training the model to be trained using the dataset.
[0015] The image generation network is determined based on the trained generation network.
[0016] Optionally, training the model to be trained using the dataset to obtain the trained generative network includes:
[0017] The sample image is input into the feature extraction network to perform foreground feature extraction, thereby obtaining the foreground distribution features of the sample image;
[0018] The foreground distribution features of the sample image are input into the generation network to generate a background image, thereby obtaining the background generated image of the sample image;
[0019] The background image of the sample image is input into the discrimination network for discrimination to obtain the discrimination result;
[0020] Calculate the error loss of the discrimination result, and adjust the parameters of the generator network and the discrimination network based on the error loss of the discrimination result;
[0021] The model to be trained is iteratively trained using the dataset, and a trained generative network is obtained after the training termination condition is met.
[0022] Optionally, the discrimination network includes a first discrimination branch and a second discrimination branch, wherein inputting the background image of the sample image into the discrimination network for discrimination to obtain a discrimination result includes:
[0023] The background image of the sample image is input into the first discrimination branch to identify the authenticity of the image, and the authenticity identification result is obtained.
[0024] The background image of the sample image is input into the second discrimination branch to identify image differences and obtain the difference identification result.
[0025] Optionally, calculating the error loss of the discrimination result and adjusting the parameters of the generator network and the discrimination network based on the error loss of the discrimination result includes:
[0026] A first error calculation is performed on the authenticity identification result to obtain a first error loss, and a second error calculation is performed on the difference identification result to obtain a second error loss;
[0027] The parameters of the generator network and the discrimination network are adjusted based on the first error loss and the second error loss.
[0028] Optionally, the step of fusing the foreground image with the background image corresponding to the original image to obtain the target image includes:
[0029] The foreground image is added to the background image corresponding to the original image, and the foreground image is filtered to obtain the target image.
[0030] In a second aspect, embodiments of the present invention provide an image generation apparatus, the apparatus comprising:
[0031] The first acquisition module is used to acquire the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features;
[0032] The generation module is used to input the foreground distribution features of the original image into a preset image generation network to generate a background image, thereby obtaining a background generated image corresponding to the original image.
[0033] The fusion module is used to fuse the foreground image with the background image corresponding to the original image to obtain the target image.
[0034] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the image generation method provided in embodiments of the present invention.
[0035] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the image generation method provided in the embodiments of the invention.
[0036] In this embodiment of the invention, the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features are obtained; the foreground distribution features of the original image are input into a preset image generation network to generate a background image, resulting in a background generated image corresponding to the original image; the foreground image and the background generated image corresponding to the original image are fused to obtain a target image. By extracting the foreground distribution features of the original image to generate the background image, a corresponding background generated image is obtained. The foreground image corresponding to the foreground distribution features is fused with the background generated image to obtain the corresponding target image. Since the background generated image is generated based on the foreground distribution features, the background in the target image contains some implicit features of the foreground image, thereby increasing the difficulty of recognizing the foreground and background in the target image. This allows the target image to be used as a difficult sample, thereby reducing the difficulty of obtaining difficult samples and improving the training efficiency of the deep neural network. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 This is a flowchart of an image generation method provided in an embodiment of the present invention;
[0039] Figure 2 This is a schematic diagram of the structure of an image generation device provided in an embodiment of the present invention;
[0040] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0042] Please see Figure 1 , Figure 1 This is a flowchart of an image generation method provided in an embodiment of the present invention, such as... Figure 1 As shown, the image generation method includes the following steps:
[0043] 101. Obtain the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features.
[0044] In this embodiment of the invention, the image generation method described above can be used to generate images of difficult samples, which can be used as data for training a deep neural network. The original image can be an image containing a target object, where the target object serves as the foreground and non-target objects serve as the background. Before training a deep neural network, image samples required for training can be collected to construct an initial dataset, and image samples can be extracted from this initial dataset to serve as the original images.
[0045] Feature extraction algorithms can be used to extract features from the original image to obtain the foreground distribution features. These foreground distribution features can include size, texture, and positional features. After obtaining the foreground distribution features, a foreground image corresponding to these features can be cropped from the original image. The feature extraction algorithms mentioned above can be based on SIFT (Scale-invariant feature transform), SURF (Speeded UpRobust Features), ORB (Oriented Fast and Rotated BRIEF), HOG (Histogram of Oriented Gradient), LBP (Local Binary Patterns), and Haar, among others.
[0046] 102. Input the foreground distribution features of the original image into a preset image generation network to generate a background image, and obtain the background generated image corresponding to the original image.
[0047] In this embodiment of the invention, after obtaining the foreground distribution features of the original image, the foreground distribution features are input as input data into a preset image generation network. The foreground distribution features are upsampled or deconvolved by the generation network to reconstruct the foreground distribution features and obtain the background generated image corresponding to the original image.
[0048] The image generation network described above can be a network built based on a deep neural network, and it is a pre-trained image generation network. Through this network, the foreground distribution features of the original image are used to generate a background image. During background image generation, the foreground distribution features are utilized for reconstruction, resulting in a certain implicit relationship between the generated background image and the foreground distribution features. For example, if the target object in the original image is a fish, the corresponding foreground distribution features are the distribution features of the fish in the original image. Based on these fish distribution features, a corresponding background image is generated, which might contain background elements such as a lake shaped like a fish's body or clouds shaped like a fish's tail.
[0049] 103. The foreground image and the background image corresponding to the original image are merged to obtain the target image.
[0050] In this embodiment of the invention, the background image generated corresponding to the original image does not contain a foreground, and the target image includes both a foreground and a background, wherein the foreground in the target image is the aforementioned foreground image. Specifically, after obtaining the background image generated corresponding to the original image, the foreground image can be added to the background image generated corresponding to the original image, thereby completing the addition of the foreground.
[0051] The aforementioned target images can be used as hard samples for training deep neural networks, thereby increasing the amount of hard sample data for training deep neural networks. Specifically, the aforementioned target images can be added to the initial dataset, thus increasing the amount of hard sample data in the initial dataset.
[0052] In this embodiment of the invention, the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features are obtained; the foreground distribution features of the original image are input into a preset image generation network to generate a background image, resulting in a background generated image corresponding to the original image; the foreground image and the background generated image corresponding to the original image are fused to obtain a target image. By extracting the foreground distribution features of the original image to generate the background image, a corresponding background generated image is obtained. The foreground image corresponding to the foreground distribution features is fused with the background generated image to obtain the corresponding target image. Since the background generated image is generated based on the foreground distribution features, the background in the target image contains some implicit features of the foreground image, thereby increasing the difficulty of recognizing the foreground and background in the target image. This allows the target image to be used as a difficult sample, thereby reducing the difficulty of obtaining difficult samples and improving the training efficiency of the deep neural network.
[0053] Optionally, in the step of obtaining the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features, the original image can be obtained; the foreground features of the original image can be extracted through a preset feature extraction network to obtain the foreground distribution features of the original image; and the foreground image corresponding to the foreground distribution features can be cropped from the original image.
[0054] In this embodiment of the invention, before training a deep neural network, image samples required for training can be collected to construct an initial dataset, and several image samples can be randomly selected from the initial dataset as original images.
[0055] The aforementioned feature extraction network can be an existing feature extraction network, such as the feature extraction part of the YOLO-V series, CenterNet, etc. It should be noted that the YOLO-V series, CenterNet, etc., include both feature extraction and feature prediction parts; this embodiment of the invention only uses the feature extraction part. The aforementioned feature extraction network can also be a feature extraction network built based on feature extraction algorithms such as SIFT, SURF, ORB, HOG, LBP, and Haar.
[0056] Foreground features are extracted from the original image using a feature extraction network, resulting in foreground distribution features. These foreground distribution features can include size, texture, and position features. The size feature describes the size of the foreground, the texture feature describes the texture distribution of the foreground, and the position feature describes the location of the foreground in the original image. These foreground distribution features are represented as vectors, also known as foreground feature vectors.
[0057] After obtaining the foreground distribution features, a screenshot can be taken from the original image based on these features to obtain the corresponding foreground image. Specifically, a screenshot can be taken from the original image based on the size and position features of the foreground to obtain the corresponding foreground image.
[0058] Optionally, before inputting the foreground distribution features of the original image into a preset image generation network to generate a background image and obtain the background image corresponding to the original image, a dataset and a model to be trained can be acquired. The dataset includes sample images, and the model to be trained includes a generation network, a discrimination network, and a feature extraction network. The output of the feature extraction network is connected to the input of the generation network, and the output of the generation network is connected to the input of the discrimination network. The model to be trained is trained using the dataset to obtain a trained generation network. Based on the trained generation network, the image generation network is determined.
[0059] In this embodiment of the invention, the aforementioned dataset is the dataset for training the model. Different sample images within the dataset may contain different target objects. The aforementioned feature extraction network is used to extract foreground distribution features from the sample images, the aforementioned generation network is used to generate a background image based on the foreground distribution features of the sample images, and the aforementioned discrimination network is used to distinguish whether the generated background image of the sample image is a generated image or a real image.
[0060] The training model is trained using a dataset. During training, the error loss is calculated based on the discrimination network's results, and the parameters of the generator and discrimination networks are adjusted to minimize this loss. When an image generated by the generator is identified as fake, its parameters are adjusted to produce more realistic images. Conversely, when an image generated by the generator is identified as real, the parameters of the discrimination network are adjusted to improve its discrimination ability. This training process is iterated until the generated images become increasingly closer to real images. Training ends when a preset number of iterations are reached, yielding a trained model. This trained model includes both a trained generator and a trained discrimination network. In this embodiment, only the trained generator network is used as the image generation network.
[0061] Optionally, in the step of training the model to be trained using the dataset to obtain a trained generative network, the sample image can be input into the feature extraction network to perform foreground feature extraction, thereby obtaining the foreground distribution features of the sample image; the foreground distribution features of the sample image can be input into the generative network to generate the background image, thereby obtaining the background generated image of the sample image; the background generated image of the sample image can be input into the discrimination network for discrimination, thereby obtaining the discrimination result; the error loss of the discrimination result can be calculated, and the parameters of the generative network and the discrimination network can be adjusted according to the error loss of the discrimination result; the model to be trained using the dataset can be iteratively trained, and the trained generative network can be obtained after the training termination condition is met.
[0062] In this embodiment of the invention, the foreground distribution features of the sample image may include size features, texture features, and position features, etc. The size features are used to describe the size of the foreground, the texture features are used to describe the texture distribution of the foreground, and the position features are used to describe the position of the foreground in the sample image.
[0063] The input data for the aforementioned generative network is a vector. Specifically, the input data is the foreground distribution features corresponding to the sample image. These foreground distribution features are represented by a vector, also known as a foreground feature vector. After obtaining the foreground distribution features of the sample image, these features are input into the generative network to generate the background image. The resulting background image of the sample image is obtained by the generative network performing multiple upsampling and pixel reconstructions on the foreground distribution features corresponding to the sample image, resulting in a background image of the same size as the sample image. More specifically, during the upsampling process, upsampling and pixel reconstruction are performed through a transposed convolution operation with a span of n until the size of the sampling result matches a preset size, thus obtaining the background image of the sample image. The preset size is the size of the sample image.
[0064] After obtaining the generated background image of the sample image, the generated background image is input into the discrimination network. The discrimination network then determines the authenticity of the generated background image, obtaining a discrimination result. The discrimination result can be true or false, and an error loss is calculated based on the discrimination result. During the discrimination process, it is desirable for the discrimination network to output a probability close to 1 for real images and a probability close to 0 for generated images. The discrimination network can calculate two error losses, and the total error loss of the discrimination network is the sum of these two error losses. One error loss is used to maximize the probability of the real image, and the other error loss is used to minimize the probability of the generated image. Specifically, the calculation of the above total error loss is shown in the following formula:
[0065]
[0066] Where lossA represents the total error loss, D(x) represents the identification result, x represents the foreground distribution feature corresponding to the sample image, Ex~pr[logD(x)] represents the identification of the real image as close to 1 as possible, and Ex~pr[log(1-D(x))] represents the identification result of the generated image as close to 0 as possible.
[0067] Based on the total error loss, the parameters of the generator network and the discriminator network are adjusted using the error backpropagation method.
[0068] Optionally, the discrimination network includes a first discrimination branch and a second discrimination branch. In the step of inputting the background generated image of the sample image into the discrimination network for discrimination and obtaining the discrimination result, the background generated image of the sample image can be input into the first discrimination branch for image authenticity discrimination to obtain authenticity discrimination result; the background generated image of the sample image can be input into the second discrimination branch for image difference discrimination to obtain difference discrimination result.
[0069] In this embodiment of the invention, the first discrimination branch is used to discriminate the authenticity of the background generated image of the sample image, and the second discrimination branch is used to discriminate the difference between the background generated image of the sample image and the sample image.
[0070] Considering that the final target image cannot be identical to the original image, a second discrimination branch is used to identify image differences, specifically the differences between the generated background image and the sample image. The greater the difference between the generated background image and the sample image, the greater the difference between the final target image and the original image. Adding the target image to the initial dataset increases the sample breadth of the initial dataset. This difference can be reflected in the similarity between the generated background image and the sample image; a higher similarity indicates a smaller difference, and vice versa. This similarity can be cosine similarity.
[0071] Optionally, in the step of calculating the error loss of the identification result and adjusting the parameters of the generator network and the identification network according to the error loss of the identification result, a first error calculation can be performed on the authenticity identification result to obtain the first error loss, and a second error calculation can be performed on the difference identification result to obtain the second error loss; the parameters of the generator network and the identification network can be adjusted according to the first error loss and the second error loss.
[0072] In this embodiment of the invention, the calculation of the first error can be referred to the calculation of lossA, where the first error loss is lossA, and the calculation of the second error can be shown in the following formula:
[0073] lossB = mins(G(x), Y)
[0074] Wherein, lossB represents the second error loss, s() represents the similarity calculation, G(x) represents the background generated image of the sample image, Y represents the sample image, and x represents the foreground distribution features corresponding to the sample image.
[0075] The sum of the first error loss and the second error loss can be used as the final error loss to adjust the parameters of the generator network and the discriminator network.
[0076] Optionally, in the step of fusing the foreground image with the background generated image corresponding to the original image to obtain the target image, the foreground image can be added to the background generated image corresponding to the original image, and the foreground image can be filtered to obtain the target image.
[0077] In this embodiment of the invention, the above-mentioned filtering process can be performed using a preset filter. Specifically, parameters can be collected from the background image corresponding to the original image to obtain filtering parameters. A corresponding filter can be constructed based on the filtering parameters, which may include contrast, white balance, brightness, color saturation, etc. By filtering the foreground image, the fusion between the foreground image and the background image corresponding to the original image becomes more natural, thereby improving the image quality of the target image.
[0078] It should be noted that the image generation method provided in this embodiment of the invention can be applied to devices such as smart cameras, smartphones, computers, and servers that are capable of image generation.
[0079] Optional, please see Figure 2 , Figure 2 This is a schematic diagram of the structure of an image generation device provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the device includes:
[0080] The first acquisition module 201 is used to acquire the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features;
[0081] The generation module 202 is used to input the foreground distribution features of the original image into a preset image generation network to generate a background image, thereby obtaining a background generated image corresponding to the original image;
[0082] The fusion module 203 is used to fuse the foreground image with the background image corresponding to the original image to obtain the target image.
[0083] Optionally, the acquisition module 201 includes:
[0084] The `get` submodule is used to acquire the original image;
[0085] The first extraction submodule is used to extract foreground features from the original image through a preset feature extraction network to obtain the foreground distribution features of the original image;
[0086] The cropping submodule is used to crop the foreground image corresponding to the foreground distribution features from the original image.
[0087] Optionally, the device further includes:
[0088] The second acquisition module is used to acquire a dataset and a model to be trained. The dataset includes sample images, and the model to be trained includes a generator network, a discriminator network, and a feature extraction network. The output of the feature extraction network is connected to the input of the generator network, and the output of the generator network is connected to the input of the discriminator network.
[0089] The training module is used to train the model to be trained using the dataset to obtain a trained generative network.
[0090] The determination module is used to determine the image generation network based on the trained generation network.
[0091] Optionally, the training module includes:
[0092] The second extraction submodule is used to input the sample image into the feature extraction network for foreground feature extraction to obtain the foreground distribution features of the sample image;
[0093] The generation submodule is used to input the foreground distribution features of the sample image into the generation network to generate a background image, thereby obtaining a background generated image of the sample image;
[0094] The discrimination submodule is used to input the background image of the sample image into the discrimination network for discrimination and obtain the discrimination result.
[0095] An adjustment submodule is used to calculate the error loss of the discrimination result and adjust the parameters of the generation network and the discrimination network based on the error loss of the discrimination result.
[0096] The training submodule is used to iteratively train the model to be trained using the dataset, and obtain the trained generative network after the training termination condition is met.
[0097] Optionally, the identification submodule includes:
[0098] The first identification unit is used to input the background generated image of the sample image into the first identification branch to identify the authenticity of the image and obtain the authenticity identification result.
[0099] The second identification unit is used to input the background image of the sample image into the second identification branch to identify the image differences and obtain the difference identification result.
[0100] Optionally, the adjustment submodule includes:
[0101] The calculation unit is used to perform a first error calculation on the authenticity identification result to obtain a first error loss, and to perform a second error calculation on the difference identification result to obtain a second error loss;
[0102] The adjustment unit is used to adjust the parameters of the generator network and the discrimination network based on the first error loss and the second error loss.
[0103] Optionally, the fusion module 203 includes:
[0104] The processing submodule is used to add the foreground image to the background generated image corresponding to the original image, and to filter the foreground image to obtain the target image.
[0105] It should be noted that the image generation device provided in this embodiment of the invention can be applied to devices such as smart cameras, smartphones, computers, and servers that can perform image generation methods.
[0106] The image generation apparatus provided in this embodiment of the invention can implement all the processes implemented by the image generation method in the above-described method embodiments, and can achieve the same beneficial effects. To avoid repetition, further details are omitted here.
[0107] See Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 3 As shown, it includes: a memory 302, a processor 301, and a computer program for an image generation method stored in the memory 302 and executable on the processor 301, wherein:
[0108] The processor 301 is used to call the computer program stored in the memory 302 and perform the following steps:
[0109] Obtain the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features;
[0110] The foreground distribution features of the original image are input into a preset image generation network to generate a background image, thereby obtaining a background generated image corresponding to the original image;
[0111] The foreground image is fused with the background image corresponding to the original image to obtain the target image.
[0112] Optionally, the process of acquiring the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features, performed by processor 301, includes:
[0113] Obtain the original image;
[0114] Foreground features are extracted from the original image using a preset feature extraction network to obtain the foreground distribution features of the original image;
[0115] Extract the foreground image corresponding to the foreground distribution features from the original image.
[0116] Optionally, before inputting the foreground distribution features of the original image into a preset image generation network to generate a background image and obtain the background generated image corresponding to the original image, the method executed by the processor 301 further includes:
[0117] Acquire a dataset and a model to be trained. The dataset includes sample images, and the model to be trained includes a generator network, a discriminator network, and a feature extraction network. The output of the feature extraction network is connected to the input of the generator network, and the output of the generator network is connected to the input of the discriminator network.
[0118] The trained generative network is obtained by training the model to be trained using the dataset.
[0119] The image generation network is determined based on the trained generation network.
[0120] Optionally, the process executed by processor 301 to train the model to be trained using the dataset to obtain a trained generative network includes:
[0121] The sample image is input into the feature extraction network to perform foreground feature extraction, thereby obtaining the foreground distribution features of the sample image;
[0122] The foreground distribution features of the sample image are input into the generation network to generate a background image, thereby obtaining the background generated image of the sample image;
[0123] The background image of the sample image is input into the discrimination network for discrimination to obtain the discrimination result;
[0124] Calculate the error loss of the discrimination result, and adjust the parameters of the generator network and the discrimination network based on the error loss of the discrimination result;
[0125] The model to be trained is iteratively trained using the dataset, and a trained generative network is obtained after the training termination condition is met.
[0126] Optionally, the discrimination network includes a first discrimination branch and a second discrimination branch. The step of processor 301 inputting the background image of the sample image into the discrimination network for discrimination to obtain a discrimination result includes:
[0127] The background image of the sample image is input into the first discrimination branch to identify the authenticity of the image, and the authenticity identification result is obtained.
[0128] The background image of the sample image is input into the second discrimination branch to identify image differences and obtain the difference identification result.
[0129] Optionally, the processor 301 performs the calculation of the error loss of the discrimination result and adjusts the parameters of the generator network and the discrimination network based on the error loss of the discrimination result, including:
[0130] A first error calculation is performed on the authenticity identification result to obtain a first error loss, and a second error calculation is performed on the difference identification result to obtain a second error loss;
[0131] The parameters of the generator network and the discrimination network are adjusted based on the first error loss and the second error loss.
[0132] Optionally, the step of processor 301 performing image fusion between the foreground image and the background image corresponding to the original image to obtain the target image includes:
[0133] The foreground image is added to the background image corresponding to the original image, and the foreground image is filtered to obtain the target image.
[0134] The electronic device provided in this embodiment of the invention can implement all the processes of the image generation method in the above-described method embodiments, and can achieve the same beneficial effects. To avoid repetition, it will not be described again here.
[0135] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the image generation method provided in this invention and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0136] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (RON), or random access memory (RAN), etc.
[0137] The above description discloses only preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.
Claims
1. An image generation method, characterized in that, Includes the following steps: Obtain the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features. The foreground distribution features include the size features, texture features, and position features of the foreground. The size features are used to describe the size of the foreground, the texture features are used to describe the texture distribution of the foreground, and the position features are used to describe the position of the foreground in the original image. The foreground distribution features of the original image are input into a preset image generation network to generate a background image, thereby obtaining a background generated image corresponding to the original image; During the training process, the background image of the sample image output by the image generation network is identified by the discrimination network, which includes a first discrimination branch and a second discrimination branch. The first discrimination branch is used to discriminate the authenticity of the background generated image of the sample image, and the second discrimination branch is used to discriminate the difference between the background generated image of the sample image and the sample image. The difference is reflected in the similarity between the background generated image of the sample image and the sample image. By maximizing the probability of real images and minimizing similarity, the parameters of the generator network and the discriminator network are adjusted to obtain the trained generator network as the image generation network. Specifically, the background image of the sample image is generated and input into the first identification branch to identify the authenticity of the image, and an authenticity identification result is obtained; The background image of the sample image is input into the second discrimination branch to identify image differences and obtain a difference discrimination result; a first error calculation is performed on the authenticity discrimination result to obtain a first error loss, and a second error calculation is performed on the difference discrimination result to obtain a second error loss; The sum of the first error loss and the second error loss is used as the final error loss to adjust the parameters of the generator network and the discriminator network. The foreground image and the background image corresponding to the original image are fused to obtain the target image, which serves as a hard sample for training the deep neural network.
2. The image generation method as described in claim 1, characterized in that, The step of obtaining the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features includes: Obtain the original image; Foreground features are extracted from the original image using a preset feature extraction network to obtain the foreground distribution features of the original image; Extract the foreground image corresponding to the foreground distribution features from the original image.
3. The image generation method as described in claim 2, characterized in that, Before inputting the foreground distribution features of the original image into a preset image generation network to generate a background image and obtain the background generated image corresponding to the original image, the method further includes: Acquire a dataset and a model to be trained. The dataset includes sample images, and the model to be trained includes a generator network, a discriminator network, and a feature extraction network. The output of the feature extraction network is connected to the input of the generator network, and the output of the generator network is connected to the input of the discriminator network. The trained generative network is obtained by training the model to be trained using the dataset. The image generation network is determined based on the trained generation network.
4. The image generation method as described in claim 3, characterized in that, The step of training the model to be trained using the dataset to obtain the trained generative network includes: The sample image is input into the feature extraction network to perform foreground feature extraction, thereby obtaining the foreground distribution features of the sample image; The foreground distribution features of the sample image are input into the generation network to generate a background image, thereby obtaining the background generated image of the sample image; The background image of the sample image is input into the discrimination network for discrimination to obtain the discrimination result; Calculate the error loss of the discrimination result, and adjust the parameters of the generator network and the discrimination network based on the error loss of the discrimination result; The model to be trained is iteratively trained using the dataset, and a trained generative network is obtained after the training termination condition is met.
5. The image generation method as described in claim 4, characterized in that, The step of fusing the foreground image with the background image corresponding to the original image to obtain the target image includes: The foreground image is added to the background image corresponding to the original image, and the foreground image is filtered to obtain the target image.
6. An image generation apparatus, characterized in that, The device includes: The first acquisition module is used to acquire the foreground distribution features of the original image and the foreground image corresponding to the foreground distribution features. The foreground distribution features include the size features, texture features and position features of the foreground. The size features are used to describe the size of the foreground, the texture features are used to describe the texture distribution of the foreground, and the position features are used to describe the position of the foreground in the original image. A generation module is used to input the foreground distribution features of the original image into a preset image generation network to generate a background image, thereby obtaining a background generated image corresponding to the original image. During the training process of the image generation network, the background generated images of the sample images output by the generation network are identified by a discrimination network. The discrimination network includes a first discrimination branch and a second discrimination branch. The first discrimination branch is used to identify the authenticity of the background generated image of the sample image, and the second discrimination branch is used to identify the difference between the background generated image of the sample image and the sample image. The difference is reflected in the similarity between the background generated image of the sample image and the sample image. By maximizing the probability of the real image and minimizing... With similarity as the target, the parameters of the generator network and the discriminator network are adjusted to obtain a trained generator network as the image generation network. Specifically, the background generated image of the sample image is input into the first discriminator branch to identify the authenticity of the image, and an authenticity identification result is obtained. The background generated image of the sample image is input into the second discriminator branch to identify the difference between the images, and a difference identification result is obtained. A first error is calculated on the authenticity identification result to obtain a first error loss, and a second error is calculated on the difference identification result to obtain a second error loss. The sum of the first error loss and the second error loss is used as the final error loss to adjust the parameters of the generator network and the discriminator network. The fusion module is used to fuse the foreground image with the background image corresponding to the original image to obtain a target image, which serves as a hard sample for training a deep neural network.
7. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the image generation method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the image generation method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Target identification method for complex environment under small sample
CN111582345A
Method for improving pedestrian attribute recognition accuracy, terminal and medium
CN113221757A
Target recognition model training method and device and target recognition method and device
CN114358249A
Cover design model training method and device, medium and computing equipment
CN114926705A