Image generation method and apparatus
By determining the first semantic target image and the first prior distribution of the source image, a diversity distribution target image that conforms to the full-image texture is generated, which solves the problems of high labor costs and poor image quality when generating complex scene pictures in the prior art, and achieves efficient and diverse image generation.
Patent Information
- Application Number
- CN202111006421.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-08-30
AI Technical Summary
When generating complex scene pictures, the prior art requires a large number of manual annotation and preparation of samples, and can only generate pictures with a single semantic target, which is of poor quality, resulting in image distortion or missing part of semantic targets.
By determining the first semantic target image and the first prior distribution of the source image, a target image with diversity distribution is generated to ensure that the image conforms to the full-picture texture and avoid distortion. The method includes determining a background image from the first semantic target image, generating a priori distribution, and generating a target image by random sampling.
It realizes the automatic generation of stable and diverse target image sets in a small number of sample images, avoiding the workload of manual annotation, and improving the generation efficiency and image quality.
Smart Images

Figure CN114037640B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to an image generation method and apparatus. Background Art
[0002] Currently, image generation technology is increasingly widely used in the field of computer vision and can be widely applied to fields such as industrial automation, biomedicine, and automobiles. For example, autonomous driving, intelligent detection, and video surveillance.
[0003] In the prior art, the powerful computing ability of the deep neural network model can effectively process image data. For example, through a discriminative model, the unknown attributes of a given known input can be predicted, that is, the type of the given picture can be identified and all semantic targets existing in the picture can be detected; through a generative model, a model of the data distribution can be constructed to describe the observable data set, so as to generate data with the same distribution. Combining the current discriminative model and generative model can realize the mutual conversion of different types of images.
[0004] The current generative model has greatly inspired the research on generating large and complex scene pictures. For example, conditional generative adversarial nets (cGAN). However, on the one hand, the labor cost required before the implementation of the current generative model cannot be ignored. For example, a sufficient number of original picture sets and target picture sets need to be prepared, and the semantic targets therein need to be manually labeled; on the other hand, only pictures for a single semantic target can be generated in a complex scene, and the generated pictures have poor quality. For example, the details of the semantic targets and the image texture are missing, resulting in serious distortion of the pictures and even the absence of some semantic targets.
[0005] How to automatically generate a target picture set with stable quality and diversity for multi-semantic target pictures contained in a small sample picture set has become an urgent problem to be solved in the industry. Summary of the Invention
[0006] This application provides an image generation method and apparatus, which can automatically generate a target picture set with stable quality and diversity for multi-semantic target images contained in a small sample picture set.
[0007] In a first aspect, an image generation method is provided. The method includes: determining a first semantic target image of a source image, where the first semantic target image is at least one of at least one semantic target image included in the source image; determining a target image according to the first semantic target image and a first prior distribution, where the first prior distribution is a prior distribution of noise obtained according to a first background image and information of a semantic target to be generated, the target image includes multiple images including the information of the semantic target to be generated, the information of the semantic target to be generated includes a label of the semantic target image to be generated and indication information of the semantic target image to be generated, and the first background image is the background image of the above source image.
[0008] According to the above solution provided by the present application, a prior distribution is generated based on the background image of the source image and the information of the semantic target image to be generated. Random sampling according to the prior distribution can generate a diversity distribution of the semantic target image to be generated in the first background image and a target image that conforms to the texture of the whole image.
[0009] In combination with the first aspect, in some implementation manners of the first aspect, the method further includes: the determining the target image according to the first semantic target image and the first prior distribution includes: determining the first background image according to the first semantic target image; generating the first prior distribution according to the first background image and the information of the semantic target to be generated; generating the target image according to the first prior distribution and the first background image.
[0010] According to the above technical solution, the generation module determines the region of the first background image according to the first semantic target image, further determines the features of the first background image, and generates the first prior distribution according to the features of the first background image and the semantic target to be generated, so as to ensure that the generated image can conform to the texture of the whole image, avoid image distortion, and the prior distribution of the semantic target to be generated based on the features of the first background image can make the subsequent sampling have diversity, thus ensuring the diversity of the target image.
[0011] In combination with the first aspect, in some implementation manners of the first aspect, the method further includes: the generating the target image according to the first prior distribution and the first background image includes: generating noise of the semantic target image to be generated according to the first prior distribution; generating the target image according to the noise of the semantic target image to be generated and the first background image.
[0012] According to the above technical solution, noise sampling is performed from the first prior distribution, and further combined with the first background image to generate a target image. The noise sampling is random, thus ensuring the diversity of the target image.
[0013] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: determining a first background image according to the first semantic target image, including: performing smoothing processing on the first semantic target image; determining the first background image according to the smoothed first semantic target image.
[0014] According to the above technical solution, smoothing the semantic target image to be replaced ensures that the area of the semantic target image to be generated can completely cover the replaced semantic target image, thereby ensuring that the generated target image conforms to the texture of the whole image and avoiding distortion of the semantic target image.
[0015] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: removing the smoothed first semantic target image from the source image.
[0016] According to the above technical solution, deleting the first semantic target image from the source image can enable the semantic target image to be generated to better fuse with the first background image, which is beneficial to the consistency of the texture of the whole image.
[0017] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: using the target image as an input image and using a to-be-trained image discriminator to identify the authenticity of the target image;
[0018] According to the output result of the to-be-trained image discriminator and the input image, adjusting the network parameter values of the image generator, where the image generator is used to generate the target image;
[0019] Using the target image generated by the image generator with adjusted network parameter values as the input image, repeating the identification action of the to-be-trained image discriminator until the training process converges.
[0020] According to this technical solution, training the generator and the discriminator makes the image generated by the generator tend to be real, ensuring that the generated target image is more vivid and natural.
[0021] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: using a to-be-trained image discriminator to identify the authenticity of the target image, including:
[0022] The to-be-trained image discriminator discriminates the images in different regions included in the target image, including:
[0023] The to-be-trained image discriminator weights the different regions when calculating the loss function in combination with the to-be-generated semantic target information.
[0024] According to this technical solution, the discriminator performs regional recognition on the target image, and uses a convolutional network to obtain a two-dimensional regional discrimination result, which is used to represent the probability that the image in each region is real. In particular, in this application, the result of regional discrimination is weighted to obtain a loss function, so that the model pays more attention to the result of the generated region, that is, focuses on discriminating the semantic target image region to be generated, thereby improving the discrimination ability of the discriminator and making the target image generated by the generator more realistic and natural.
[0025] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the image discriminator to be trained performs regional division on the target image.
[0026] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the image discriminator to be trained performs smoothing processing on the semantic target image to be generated included in the target image; the image discriminator to be trained performs authenticity discrimination on the target image including the smoothed semantic target image to be generated.
[0027] According to this technical solution, the discriminator performs smoothing processing on the replaced semantic target image to ensure that the region of the semantic target image can completely cover the replaced semantic target image, thereby ensuring that the generated target image conforms to the full-image texture and avoiding distortion of the semantic target image.
[0028] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: completing the annotation of the semantic target of the target image according to the information of the semantic target to be generated and the information of the first semantic target image.
[0029] According to this technical solution, based on the semantic target label in the information of the semantic target to be generated, the generated target image can be automatically annotated, avoiding the workload of manual annotation and effectively improving the efficiency of preparing the training data set.
[0030] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the image discriminator to be trained includes an image detector, and the image detector is used to extract features of the first semantic target image included in the source image.
[0031] According to this technical solution, the recognizer and the discriminator are coupled. In this case, the discriminator and the detector can share feature extraction, thereby greatly reducing the computational amount of the model and effectively improving the image generation efficiency.
[0032] In a second aspect, an image generation device is provided, and the image generation device executes the units of the method in the first aspect or its various implementation manners.
[0033] In a third aspect, an image generation device is provided, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the communication device executes the image generation method in the first aspect and its various possible implementation manners.
[0034] Optionally, there is one or more processors and one or more memories.
[0035] Optionally, the memory may be integrated with the processor, or the memory is separately provided from the processor.
[0036] In a fourth aspect, a computer-readable storage medium is provided, characterized in that the computer-readable medium stores program code for a device to execute, and the program code includes a method for executing the first aspect or the second aspect.
[0037] In a fifth aspect, a computer program product containing instructions is provided. When the computer program product runs on a computer, the computer is caused to execute the method in any one of the above aspects.
[0038] In a sixth aspect, a chip is provided. The chip includes a processor and a data interface. The processor reads instructions stored on a memory through the data interface and executes the method in any one of the above aspects.
[0039] Optionally, as an implementation manner, the chip may further include a memory. Instructions are stored in the memory, and the processor is used to execute the instructions stored on the memory. When the instructions are executed, the processor is used to execute the method in any one of the above aspects.
[0040] The above-mentioned chip may specifically be a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). Description of the Drawings
[0041] Figure 1 A schematic structural diagram of the system architecture provided by an embodiment of the present application is shown;
[0042] Figure 2 A schematic structural diagram of a convolutional neural network provided by an embodiment of the present application;
[0043] Figure 3 A schematic diagram of a product implementation form provided by an embodiment of the present application;
[0044] Figure 4Schematic flowchart of an image generation method provided by an embodiment of the present application;
[0045] Figure 5 Schematic diagrams of a discriminator structure and a detector structure for shared feature extraction provided by an embodiment of the present application;
[0046] Figure 6 Schematic block diagram of an image generation device provided by an embodiment of the present application. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.
[0048] It should be understood that the names of all nodes and devices in the present application are only set for the convenience of description in the present application, and the names in actual applications may be different. It should not be understood that the present application limits the names of various nodes and devices. On the contrary, any name having the same or similar functions as the nodes or devices used in the present application is regarded as a method or equivalent replacement of the present application and is within the protection scope of the present application. This will not be elaborated below.
[0049] Since the embodiments of the present application involve a large number of applications of neural networks, for the convenience of understanding, the relevant terms and concepts of the neural networks that may be involved in the embodiments of the present application will be introduced below.
[0050] (1) Neural network
[0051] A neural network can be composed of neural units. A neural unit can refer to an operation unit with x s and intercept 1 as inputs. The output of this operation unit can be:
[0052]
[0053] where s = 1, 2,..., n, and n is a natural number greater than 1. W s is the weight of x s , b is the bias of the neural unit. f is the activation function of the neural unit, which is used to perform a non-linear transformation on the features obtained in the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting multiple such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0054] (2) Deep Neural Network
[0055] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple hidden layers. When dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is to say, any neuron in the i-th layer must be connected to any neuron in the (i + 1)-th layer.
[0056] Although the DNN seems very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: Among them, is the input vector, is the output vector, is the bias vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer simply performs such a simple operation on the input vector to obtain the output vector Due to the large number of layers in the DNN, the number of coefficients W and bias vectors is also relatively large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscripts correspond to the index 2 of the output third layer and the index 4 of the input second layer.
[0057] To sum up, the coefficient from the k-th neuron in the (L - 1)-th layer to the j-th neuron in the L-th layer is defined as
[0058] It should be noted that the input layer does not have the W parameter. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically speaking, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is also the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0059] (3) Convolutional Neural Network
[0060] The convolutional neural network (CNN) is a deep neural network with a convolutional structure. The convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers, and this feature extractor can be regarded as a filter. The convolutional layer refers to the neuron layer in the convolutional neural network that performs convolutional processing on the input signal. In the convolutional layer of the convolutional neural network, a neuron can only be connected to some neighboring layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some neurons arranged in a rectangle. The neurons in the same feature plane share weights, and the shared weights here are the convolutional kernels. The shared weights can be understood as a way of extracting image information that is independent of position. The convolutional kernel can be formalized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolutional kernel can obtain reasonable weights through learning. In addition, the direct benefit brought by the shared weights is to reduce the connections between the layers of the convolutional neural network and at the same time reduce the risk of overfitting.
[0061] (4) Loss function
[0062] During the process of training a deep neural network, since it is hoped that the output of the deep neural network is as close as possible to the value that is truly desired to be predicted, the predicted value of the current network can be compared with the truly desired target value, and then the weight vector of each layer of the neural network can be updated according to the difference between the two (of course, there is usually a process of initialization before the first update, that is, configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is too high, the weight vector is adjusted to make it predict lower, and continuous adjustment is made until the deep neural network can predict the truly desired target value or a value very close to the truly desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, and then the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0063] (5) Backbone network
[0064] In neural networks, especially in the field of computer vision (CV), generally, feature extraction is first performed on the image. This part is the foundation of the entire CV task because subsequent tasks all depend on the extracted image features. Therefore, this part of the network structure is called the backbone network.
[0065] (6) U-Net
[0066] It is one of the earlier algorithms for semantic segmentation using fully convolutional networks. Its network structure is divided into a downsampling stage and an upsampling stage. There are only convolutional layers and pooling layers in the network structure, without fully connected layers. The shallower high-resolution layers in the network are used to solve the problem of pixel localization, and the deeper layers are used to solve the problem of pixel classification, so as to achieve semantic-level segmentation of images. In the structure of U-Net, it includes a contracting path that captures context information and a symmetric expanding path that allows for precise localization. In this application, this method can be understood as completing end-to-end training with very little data, that is, the input is an image and the output is also an image, and obtaining the best results.
[0067] (7) Image features
[0068] Image features can be understood as numerical features transformed from raw data through feature extraction operations, which are convenient for algorithms to understand and process. In the embodiments of this application, image features specifically refer to the image features extracted using the backbone network model.
[0069] (8) Discriminant model
[0070] In the embodiments of this application, the discriminant model can predict the unknown attributes of a given known input, that is, identify the category to which a given picture belongs, and detect all semantic targets existing in the picture. Such discriminant models include, for example, object classification models and object detection models.
[0071] In the discriminant model, the first thing to be realized is the classification of single semantic target images, that is, the target classification model. This model uses convolutional layers to extract features of local regions of the image, and inputs the obtained features into a fully connected layer for classification. On this basis, the target detection model uses a similar backbone network to extract features from images containing multiple semantic targets, and then locates and identifies the semantic targets in the original image according to the feature maps output by the backbone network. According to the different model structures, the target detection model can be further divided into single-step models and two-step models. Among them, the single-step methods represented by the YOLO neural network (you only look once neural network, YOLO) and the SSD detector (single shot multibox detector, SSD) use the regression method to calculate the probability of each grid point containing semantic targets with the overall feature map as the input; the two-step methods represented by the Faster Region Based Convolutional Neural Networks (Faster RCNN) and the Mask Region Based Convolutional Neural Networks (Mask RCNN) introduce a candidate region network. First, it discriminates whether the preset candidate boxes contain semantic targets based on feature information, and then inputs the regions with a higher probability of containing semantic targets into the classifier for target detection.
[0072] (9) Generative model
[0073] In the embodiments of this application, the generative model takes the generative adversarial model as an example. This model consists of a generator and a discriminator. The generator aims to learn the mapping from the noise distribution to the target data distribution, while the discriminator is used to discriminate whether the received data is real. During training, first, the discriminator is trained with the fake data generated by the generator and the real data respectively. Then, the generator and the discriminator are connected in series, and the parameters of the discriminator are fixed. With the goal of "generating real data", the generator is trained. During the training process, the ability of the discriminator to distinguish between true and false is improved, so that the pictures generated by the generator are also closer to the real distribution, and finally reach an equilibrium. The training of the generative adversarial network can be understood as the following optimization problem:
[0074] The generative adversarial model can perform well in generating small single-class images, but it is difficult to process high-definition images and cannot generate images of different classes based on a single model. On this basis, the conditional generative adversarial model (cGAN) introduces control conditions on the basis of the original model, achieving the effect of controllable generation of semantic target categories, so that different classes of images can be generated using a single model. Specifically, the model represents the category of the semantic target to be generated as a one-dimensional vector as the conditional vector. In the generator, an embedding layer is used to encode the noise and the condition, and then it is transmitted into the network for generation; in the discriminator, the features obtained from the backbone network and the noise are encoded to determine whether the input image is a real image under the target category.
[0075] To facilitate the understanding of the embodiments of the present application, first, in combination with Figure 1 Briefly describe the structural schematic diagram of a system architecture 100 according to an embodiment of the present application. As Figure 1 shown, the system architecture 100 includes a dataset module 110 and a model training module 120. The dataset module 110 is one of the bases for model training. A high-quality and large-scale dataset is the key to training a high-quality model.
[0076] Among them, in the dataset module 110, as Figure 1 shown, the data acquisition device 111 is used to acquire raw data; the data preprocessing device 112 can be used for screening, filtering, and data annotation of the raw data. Among them, the data annotation can be manual annotation or automatic annotation, and the embodiments of the present application do not make any limitations in this regard. The data generation device 113 is used to generate new data according to the annotated data. Among them, data generation is an important means to solve the problem of insufficient dataset scale.
[0077] In the dataset module 110, there is also a data repository 114 for storing the automatically generated dataset. This dataset can be used by the training module for model training.
[0078] As Figure 1 shown, the model training module 120 includes a training device 121, and the training device 121 can train a target model 122 based on the training data maintained in the database 130.
[0079] Next, describe how the training device 121 obtains the target model 122 based on the training data. The training device 120 processes the input data and compares the output data with the input raw data until the difference between the output data of the training device 120 and the input raw data is less than a certain threshold, thereby completing the training of the target model 122.
[0080] The above target model 122 can be used to implement the image generation method of the embodiments of the present application, that is, by inputting the picture set automatically generated by the data generation device 113 into the target model 122, it is possible to determine the authenticity of the pictures included in the picture set. The target model 122 in the embodiments of the present application may specifically be a cGAN network.
[0081] It should be noted that, in actual applications, the training data maintained in the database 114 may not necessarily all come from the collection of the data collection device 111, and it is also possible to receive it from other devices. Additionally, it should be noted that the training device 121 does not necessarily completely train the target model based on the training data maintained in the database 114, and it is also possible to obtain training data from the cloud or other places for model training. The above description should not be regarded as a limitation to the embodiments of the present application.
[0082] The target model 122 trained by the training device 121 can be applied to different systems or devices, such as an execution device. The execution device can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR), a vehicle-mounted terminal, etc., or it can also be a server or the cloud, etc.
[0083] It is worth noting that the above training device 121 can generate corresponding target models 122 based on different targets or different tasks and based on different training data. The corresponding target model 122 can be used to achieve the above targets or complete the above tasks, so as to provide the required results for users.
[0084] It should be noted that Figure 1 This is only a schematic diagram of a system architecture provided by the embodiments of the present application. The positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation.
[0085] Since CNN is a very common neural network, the following will combine Figure 2 to introduce the structure of CNN in detail. As described in the basic concept introduction above, a convolutional neural network is a deep neural network with a convolutional structure and is a deep learning architecture. A deep learning architecture refers to an algorithm for updating a neural network model to perform multiple levels of learning at different abstraction levels. As a deep learning architecture, CNN is a feed-forward artificial neural network, and each neuron in the feed-forward artificial neural network can respond to the input image.
[0086] In Figure 2In this case, the convolutional neural network (CNN) 200 may include an input layer 210, a convolutional layer / pooling layer 220 (where the pooling layer is optional), and a fully connected layer 230. Figure 2 The convolutional neural network in this case can be applied to the image classification model structure. The following will Figure 2 introduce in detail the internal layer structure of the CNN 200 in this case.
[0087] Convolutional layer / pooling layer 220:
[0088] Convolutional layer:
[0089] As Figure 2 shown, the convolutional layer / pooling layer 220 may include layers such as examples 221 - 226. For example: in one implementation, layer 221 is a convolutional layer, layer 222 is a pooling layer, layer 223 is a convolutional layer, layer 224 is a pooling layer, 225 is a convolutional layer, and 226 is a pooling layer; in another implementation, 221 and 222 are convolutional layers, 223 is a pooling layer, 224 and 225 are convolutional layers, and 226 is a pooling layer. That is, the output of the convolutional layer can be used as the input of the subsequent pooling layer or as the input of another convolutional layer to continue the convolutional operation.
[0090] The following will take the convolutional layer 221 as an example to introduce the internal working principle of one convolutional layer.
[0091] The convolutional layer 221 can include a number of convolutional operators, also known as kernels, which act as filters for extracting specific information from the input image matrix in image processing. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the convolution operation on the image, the weight matrix is typically processed pixel by pixel (or two pixels by two pixels... depending on the value of the stride) along the horizontal direction of the input image to complete the task of extracting specific features from the image. The size of this weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as that of the input image, and during the convolution operation, the weight matrix extends to the entire depth of the input image. Therefore, convolving with a single weight matrix produces a convolved output with a single depth dimension. However, in most cases, instead of using a single weight matrix, multiple weight matrices of the same size (rows × columns), i.e., multiple matrices of the same type, are applied. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image, where the dimension can be understood as being determined by the "multiple" mentioned above. Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract edge information of the image, another weight matrix is used to extract specific colors of the image, and yet another weight matrix is used to blur the unwanted noise in the image, etc. These multiple weight matrices have the same size (rows × columns), and the size of the convolutional feature maps extracted by these multiple weight matrices of the same size is also the same. Then, the multiple convolutional feature maps of the same size that are extracted are combined to form the output of the convolution operation.
[0092] The weight values in these weight matrices need to be obtained through a large amount of training in practical applications. Each weight matrix formed by the weight values obtained through training can be used to extract information from the input image, enabling the convolutional neural network 200 to make correct predictions.
[0093] When the convolutional neural network 200 has multiple convolutional layers, the convolutional layer (such as 221) often extracts more general features, which can also be referred to as low-level features. As the depth of the convolutional neural network 200 increases, the features extracted by the subsequent convolutional layers (such as 226) become increasingly complex, such as high-level semantic features. The higher the semantic features, the more suitable they are for the problem to be solved.
[0094] Pooling layer:
[0095] Since it is often necessary to reduce the number of training parameters, a pooling layer is often introduced periodically after the convolutional layer. Figure 2Each of the layers 221 - 226 shown in 220 can be a convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. During image processing, the sole purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer can include an average pooling operator and / or a max pooling operator for sampling the input image to obtain a smaller-sized image. The average pooling operator can calculate the average value of the pixel values in the image within a specific range as the result of average pooling. The max pooling operator can select the pixel with the maximum value within the specific range as the result of max pooling. Additionally, just as the size of the weight matrix in the convolutional layer should be related to the image size, the operators in the pooling layer should also be related to the image size. The size of the image output after processing by the pooling layer can be smaller than the size of the image input to the pooling layer. Each pixel point in the image output by the pooling layer represents the average or maximum value of the corresponding sub-region of the image input to the pooling layer.
[0096] Fully connected layer 230:
[0097] After being processed by the convolutional / pooling layer 220, the convolutional neural network 200 is still not sufficient to output the required output information. Because as mentioned before, the convolutional / pooling layer 220 only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other relevant information), the convolutional neural network 200 needs to use the fully connected layer 230 to generate one or a set of outputs with the number of classes required. Therefore, the fully connected layer 230 can include multiple hidden layers (such as Figure 3 231, 232 to 23n shown) and the output layer 240. The parameters contained in the multiple hidden layers can be pre-trained according to the relevant training data of the specific task type. For example, the task type can include image recognition, image classification, image super-resolution reconstruction, etc.
[0098] After the multiple hidden layers in the fully connected layer 230, that is, the last layer of the entire convolutional neural network 200 is the output layer 240. The output layer 240 has a loss function similar to categorical cross-entropy, specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network 200 (such as Figure 3 the propagation from 210 to 240 is the forward propagation) is completed, the backpropagation (such as Figure 3 the propagation from 240 to 210 is the backpropagation) will start to update the weight values and biases of the previously mentioned layers to reduce the loss of the convolutional neural network 200, that is, the error between the result output by the convolutional neural network 200 through the output layer and the ideal result.
[0099] It should be noted that Figure 2The convolutional neural network shown is only an example of a convolutional neural network. In specific applications, the convolutional neural network can also exist in the form of other network models.
[0100] The following introduces a product implementation form provided by the embodiments of the present application.
[0101] Figure 3 It is a schematic diagram of a product implementation form provided by the embodiments of the present application.
[0102] A product implementation form of the embodiments of the present application may include program code that is included in a dataset system and deployed on server hardware. The program code of the present application may exist in the data generation module of the dataset system, such as the data generation device 113 in the system architecture 100. This part of the code runs on the host storage (memory or disk) and is used to execute the innovative data automatic generation method.
[0103] Figure 3 A product implementation form of the present application is shown. The product includes a server 310, hardware 320, and software 330. The hardware 330 includes a host memory or disk and is used to run the program code of the present application. The software 320 includes an image recognition device 321 and an image generation device 322. The image recognition device 321 is used to identify the authenticity of the input image. The image generation device 322 is used to generate diverse images of the target semantic information required by the user and store the generated pictures in the host memory or disk for the input image recognition device 321 to perform picture recognition.
[0104] The image generation method and device provided by the embodiments of the present application can be applied to other similar tasks of picture editing and replacement. For example, replacement of vehicles or other target objects in different road conditions in autonomous driving, replacement of hairstyles, hair colors, and headgear of people in video surveillance scenarios, replacement of different product shapes, colors, and layouts in industrial automation quality monitoring scenarios, etc. Specifically, the image generation method of the embodiments of the present application can be applied to the scenario of picture editing and replacement. The embodiments of the present application will introduce the image generation method of the embodiments of the present application in detail taking the autonomous driving scenario as an example.
[0105] To facilitate the understanding of this embodiment, first, a method for generating a sample image disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the sample image generation method provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities. Such a computer device includes, for example: a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the sample image generation method may be implemented by a processor invoking computer-readable instructions stored in a memory.
[0106] Next, Figure 4 the image generation method of the embodiments of the present application will be described. Figure 4 FIG. shows a schematic flowchart of an image generation method 400 provided in the embodiments of the present application. The method may be applied to a data set system, and the data set system includes a detection module and a generation module. Among them, the generation module includes a discriminator and a generator. The generator includes an encoding structure and a decoding structure. The detection model and the generation model may be trained based on a generative adversarial network. The method may be executed by the above-mentioned execution device. Optionally, the method 400 may be processed by a CPU, or jointly processed by a CPU and a GPU. It is also possible not to use a GPU, but to use other processors suitable for neural network computing. The present application does not make any restrictions.
[0107] During training, the detection model first uses a source image with annotations to train a picture recognizer. The source image with annotations refers to pre-annotating the semantic targets included in the pictures in the source image. It can be manually annotated or automatically annotated, without limitation. The purpose of training the picture recognizer is to enable the picture recognizer to recognize the corresponding annotated semantic targets through training. This training method may adopt the training methods in the prior art, or any method that can achieve the purpose of training the picture recognizer. The embodiments of the present application do not make any restrictions on this.
[0108] In the embodiments of the present application, the trained detection model performs semantic target recognition on the source image, and then the generation model edits the recognized semantic targets according to the pre-given conditions to achieve end-to-end automatic image generation and automatic annotation.
[0109] In an embodiment of the present application, the source image is an example of the original picture as the input. The source image includes multiple input pictures. The first semantic target image is an example of the semantic target. The first background image is an example of the image obtained by processing the source image according to the semantic target information to be generated by the user. The target image is an example of the picture generated by the generator.
[0110] It should be understood that the source image can be referred to as the original image, input image, original picture and other similar terms. The target image can also be referred to as the generated image, generated picture and other similar terms. In the embodiments of the present application, the source image and the target image are taken as examples for illustration, and the embodiments of the present application do not limit this.
[0111] Method 400 includes steps S410 to S440. The following is a detailed description of steps S410 to S440.
[0112] S410, the detection model determines the first semantic target image of the source image.
[0113] S420, the generation model determines the first background image according to the first semantic target image.
[0114] S430, the generation model generates the first prior distribution according to the first background image and the semantic target information to be generated.
[0115] S440, the generation model generates the target image according to the first prior distribution and the first background image.
[0116] The following is a detailed description of steps S410 to S440.
[0117] S410, the detection model determines the first semantic target image of the source image.
[0118] The source image is the input image of the dataset system. The detection model of the dataset system includes a picture recognizer, and the picture recognizer can recognize the first semantic target image of the source image.
[0119] In an embodiment of the present application, the image generation method can be used in the autonomous driving scenario. Therefore, the source image can be the image collected by the autonomous driving vehicle, and the first semantic target image can be the vehicle image in the source image, such as a bus, a car, a bicycle, etc.
[0120] Specifically, in a possible implementation manner, the above recognizer is trained, so it can recognize the first semantic target image according to the annotation. Among them, the source image is the original input image, including one or more images, each image includes one or more semantic targets, and the first semantic target image is at least one of them.
[0121] In a possible implementation, the detection model annotates the recognized semantic target image, specifically including annotating the bounding box of the semantic target and the category of the semantic target. For the detected bounding box, its position is recorded and appropriately enlarged to include some of the boundary parts, so as to avoid the generated semantic target from blending with the surrounding environment smoothly.
[0122] For example, the source image is an image of an autonomous driving scenario, and the source image contains multiple semantic targets, such as buses, roads, sky, plants, pedestrians, etc. Among them, the first semantic target image can be a vehicle, such as a bus. The detection model records the position of the first semantic target image and enlarges the bounding box, and at the same time identifies the category of its semantic target as a vehicle.
[0123] It should be understood that the bounding box of the first semantic target image can be of any shape. This arbitrary shape can be understood as appropriately enlarging the detected area containing the first semantic target image to obtain a partial environmental image around the first semantic target image. Therefore, the arbitrary shape includes the complete image of the first semantic target and a partial environmental image around it. The arbitrary shape is, for example: rectangle, circle, trapezoid, etc. The embodiments of the present application do not make any limitations in this regard.
[0124] It should be understood that the above examples are only for illustrative purposes and should not constitute any limitation to the embodiments of the present application. The present application takes image generation in an autonomous driving scenario as an example. In addition, the image generation method provided by the embodiments of the present application can also be applied to any image generation scenario that requires replacing semantic targets.
[0125] S420. The generation model determines a first background image according to the first semantic target image.
[0126] The generation model determines the first semantic target image detected by the detection model, edits the source image according to the first semantic target image to determine the first background image, and performs feature extraction to determine the features of the first background image.
[0127] Specifically, the generation model determines the first semantic target image detected by the detection model, determines the position and annotation information of the first semantic target image, processes the source image according to the annotation information of the first semantic target, thereby determining the first background image. Further, feature extraction is performed on the first background image to obtain the features of the first background image.
[0128] In a possible implementation, smoothing processing is performed on the first semantic target image for enlarging the area of the first semantic target image. It can be understood that after the enlarging processing, the semantic target to be generated can completely cover the area of the first semantic target image, so as to achieve the consistency of the texture of the whole image.
[0129] It should be understood that in this application, after the first semantic target image is smoothed, the area of the first semantic target image can be directly covered by the target semantic image to be generated.
[0130] In a possible implementation, the smoothed first semantic target image is removed from the source image to obtain a first background image.
[0131] In a possible implementation, the semantic target information to be generated is given semantic target information. For example, the semantic target information can be to replace a bus with a truck, or to remove the bus. Similarly, requirements defined according to user needs can all be used as the semantic target information.
[0132] In a possible implementation, the semantic target information to be generated can also be defined by the system, and the embodiments of this application do not limit this.
[0133] Preferably, the semantic target information to be generated may include the label of the semantic target image and the indication information of the semantic target image to be generated. The semantic target can be understood as the semantic target to be generated, and the semantic target to be generated is similar to the semantic target volume and semantic category in the source image.
[0134] For example, when the first semantic target image detected in the source image is a large vehicle, the labels of the semantic targets to be generated can be truck, bus, fire truck, etc. The semantic target label included in the semantic target information is a vehicle, and the semantic target volume description is large. Another example is that when the first semantic target image is a bicycle, the semantic target to be generated can be a motorcycle, a tricycle, etc. The semantic target category included in the semantic target information is a vehicle, and the semantic target volume description is small.
[0135] It should be understood that processing the source image according to the semantic target information to be generated and the annotation information of the first semantic target image includes: smoothing the first semantic target image. For example, if the indication information in the semantic target information to be generated is to replace a bus with a truck, the generation model determines the position and bounding box of the first semantic target in the source image according to the annotation information of the first semantic target, and performs deletion processing or covering processing on the image within the bounding box, so as to obtain a first background image; Another example is that if the indication information in the semantic target information to be generated is to remove the bus, the generation model determines the position and bounding box of the first semantic target, performs deletion processing on the image within the bounding box, and fills the surrounding environmental color blocks at the same time. As an understanding, for example, the road surface color blocks are used to fill the bounding box to obtain a first background image.
[0136] It should be understood that the above examples are only for illustrative purposes and should not constitute any limitation to the embodiments of this application.
[0137] S430, the generation model generates a first prior distribution according to the first background image and the semantic target information to be generated.
[0138] It should be understood that in the embodiments of the present application, the generator includes two parts: an encoding structure and a decoding structure, and the encoder includes a prior condition encoding module.
[0139] Specifically, taking the first background image as the input, the encoder downsamples the image through convolutional layers and pooling layers and extracts the features of the first background image. Subsequently, the features of the first background image are encoded and merged with the semantic target information to be generated, and then input into the prior condition encoding module. Through this prior condition encoding module, a prior condition for the current source image is obtained, and a prior distribution of the noise of the semantic target to be generated is generated according to the prior condition.
[0140] It should be understood that the first prior distribution is only for illustrative purposes. The embodiments of the present application do not limit this.
[0141] It should be understood that the semantic target information to be generated is determined by means such as user definition or system definition. As an understanding, the semantic target to be generated can be, for example, object B, object C, object D, object E, etc. If the first semantic target image is object A, then the semantic target information to be generated can be understood as replacing object A in the source image with object B, or replacing it with object C or object D, etc. Taking the semantic target information as replacing object A in the source image with object B as an example, in this case, the discriminator first discriminates object A in the source image. After inputting it into the generator, the generator first smooths the area of object A to obtain the first background image, and then the encoder extracts features from the first background image through sampling. The features of the first background image are encoded and merged with the information of object B, and then input into the prior condition encoding module to obtain the prior distribution of object B. It can be understood that if it is to be replaced with object C, the prior distribution of object C needs to be generated.
[0142] In the embodiments of the present application, it should be understood that the first background image is used as the input of the encoder. The encoder obtains a length vector and a matrix through convolutional layers and pooling layers. The vector and the matrix are used as the prior conditions of the Gaussian distribution, and sampling is performed from this Gaussian distribution. When a single object A is to be generated, different images of object A such as A1, A2, A3, etc. will be generated according to different random samplings from the Gaussian distribution and different distribution parameters of the Gaussian distribution, so that diverse images of object A can be generated. Among them, the size of the distribution parameter depends on the size of the area of the image to be replaced, that is, the area size of the first semantic target image.
[0143] Further, the noise obtained by Gaussian distribution sampling is combined and decoded with the features of the first background image. For example, when the first background image represents an environment with sufficient light, the target images obtained by combined decoding can be images of object A such as A1, A2, A3, etc. under sufficient light. For another example, when the first background image is an environment with rainy or cloudy weather, the target images obtained by combined decoding can be images of object A such as A1', A2', A3', etc. in rainy or cloudy weather.
[0144] It should be understood that taking the first background image as a prior condition can obtain the environmental information of the source image. This environmental information can be understood as the situation of the surrounding environment excluding the semantic target to be replaced, and thus the environment, texture, and integrity of the generated area can be controlled. Figure 1 Consistent.
[0145] It should be understood that taking the semantic target information to be generated as a condition, including the label of the semantic target to be generated, can control the output type of the semantic target of the generated picture to be similar to or consistent with that of the source image.
[0146] S440. The generation model generates the target image according to the first prior distribution and the first background image.
[0147] Specifically, the generation model samples noise from the above-mentioned first prior distribution, combines it with the features of the first background image extracted by the encoder, and inputs the result into the decoder to generate the target image.
[0148] Optionally, it can be understood that the noise sampled by the generation model from the above prior distribution can be understood as the feature of the position distribution change of the semantic target image to be generated in the first background image based on the Gaussian distribution. Encoding based on this change feature and the first background image feature can obtain multiple different features, and decoding these multiple different features by the decoder can obtain different images.
[0149] It should be understood that the above prior distribution based on the Gaussian distribution is only an exemplary illustration, and this prior distribution can also be in other forms, which is not limited in the embodiments of the present application.
[0150] It should be understood that the above different images are generated based on the sampled noise, and the noise obtained each time of sampling is different, thus realizing the generation of different position distribution change images of the semantic target to be generated in the source image.
[0151] It can be understood that the above different position distributions of the semantic target to be generated in the source image can be understood as changes in angle, position, layout, etc., and moreover, the semantic target to be generated can also be changed in color according to user definition.
[0152] Under the guidance of the above prior distribution, the position of the semantic target image to be generated in the above first background image has a more reasonable and diverse distribution, so that the generated sample data is more reasonable and rich.
[0153] In a possible implementation manner, the generation model generates an annotation of the target image according to the semantic target information to be generated and the information of the first semantic target image, realizing automatic annotation of the semantic target.
[0154] It should be understood that the above generator and discriminator are both trained generators and discriminators. The specific training methods will be introduced in Method 500 below.
[0155] For the convenience of those skilled in the art to understand, the following will be described in combination with Figure 5 the examples in.
[0156] Figure 5 A schematic block diagram of the modules included according to the embodiments of the present application is provided. As Figure 5 shown, it includes: a detection module and a generation module.
[0157] Among them, the detection module includes an identifier; the generation module includes a generator and a discriminator. The generator includes an encoding structure and a decoding structure. The detection model and the generation model can be trained based on a generative adversarial network.
[0158] For the detection module, specifically, the source image is used as the input. The source image is pre-annotated, and the identifier is trained using the annotated source image, so that the identifier can detect the semantic target. This training method can adopt the training methods in the prior art, or any method that can achieve the purpose of training the image identifier. The embodiments of the present application do not limit this.
[0159] It should be understood that the trained identifier can determine the first semantic target image of the source image, as described in step S410 of Method 400, which will not be elaborated here.
[0160] It should be understood that when the data samples are unbalanced, the model will enhance the recognition ability for the categories with rich data volume and weaken the recognition ability for the categories with less data volume. Therefore, the selected bounding boxes have a higher probability of selecting the categories with a larger sample volume. In the semantic target editing module, the selected original semantic target will be replaced by other semantic targets, while the semantic targets and their labels that could not be recognized originally are retained in the original image.
[0161] For the generation module, specifically, in the semantic target editing module, the source image can be processed according to the to-be-generated semantic target information to obtain a first background image. The first background image is input into an editor to obtain first background image features. Then, the source image and the target semantic category are input into the editor to obtain prior conditions, further generate a prior distribution of noise, and then extract noise features from the prior distribution. The noise features and the first background image features are input into the editor for editing and merging, and the obtained features are input into a decoder for decoding to obtain a synthesized image.
[0162] It should be understood that the image synthesized by the generator needs to be input into the discriminator for further judgment of whether it is real.
[0163] It can be understood that the generator and the discriminator are connected in series and the parameters of the discriminator are fixed, and with the goal of "this picture is real", the parameters of the generator are optimized to achieve the purpose that the pictures generated by the generator can "deceive" the discriminator, that is, the pictures generated by the generator are as close as possible to real pictures.
[0164] Specifically, the image synthesized by the generator is input into the discriminator. The discriminator needs to obtain a real picture, that is, the source image. The discriminator judges whether the input image is real based on the source image. According to the discrimination result of the discriminator, the network parameter values of the generator are adjusted. Further, the picture generated by the adjusted generator is used as the input to the discriminator to continue the above steps until it is judged that the output result is real, and the training end conditions of the to-be-trained image generator and the training end conditions of the to-be-trained image discriminator reach a balance.
[0165] In a possible implementation manner, the discriminator performs smoothing processing on the target semantic information included in the target image, and then performs authenticity discrimination on the target image including the smoothed target semantic information. Optionally, one understanding is that the discriminator discriminates the image generated by the recognizer according to the real image. The generated image only has differences from the real image within the first semantic target bounding box, and there are no changes in other semantic targets. Therefore, the main discrimination area of the discriminator is the area within the bounding box of the first semantic target.
[0166] In a possible implementation manner, the discriminator divides the target image into regions and performs authenticity discrimination on different regions.
[0167] Specifically, the to-be-trained image discriminator weights different regions when calculating the loss function in combination with the to-be-generated semantic target information.
[0168] It should be understood that when the discriminator recognizes a region with a relatively high coincidence rate with the semantic target image to be generated, the weight of the loss function will be increased during the calculation of the loss function, so that the discriminative model pays more attention to the generated semantic target image, thereby improving the effect of the model. That is, the authenticity of discrimination within this region can be improved.
[0169] In a possible implementation, the discriminator to be trained magnifies the bounding box including the target semantic information in the image, and discriminates the authenticity of the image within the region of the bounding box.
[0170] It should be understood that in the embodiments of the present application, the discrimination region of the discriminator is expanded, not only within the contour of the semantic target. Therefore, the discriminator also has a discrimination process for the environmental region of the semantic target within the bounding box, and then controls the output of the generation model, so that the local fusion of the target image generated by the trained recognizer is smooth and natural, and the texture of the whole image is highly consistent.
[0171] In a possible implementation, after the backbone network of the discriminator extracts the image features, the features are embedded with the category information map annotated based on the original image, and then a two-dimensional region discrimination result is obtained using a convolutional network to represent the probability that the image in each region is real. In particular, the present application weights the results of the region discrimination to obtain a loss function, so that the model pays more attention to the results of the generated region.
[0172] Preferably, the discriminator of the present application can adopt a PatchGAN structure.
[0173] It can be understood that after the discriminator is trained, its ability to recognize real pictures and forged pictures increases. When optimizing the parameters of the generator, the ability of the generator to generate realistic pictures will be correspondingly improved. Through continuous iterative optimization, the discriminator and the generator gradually reach a balance, thus completing the training of the model.
[0174] It should be understood that the trained recognizer can directly generate the target image.
[0175] Generate the annotation of the target image according to the target semantic information and the information of the first semantic target image, realize the automatic annotation of the semantic target, avoid the workload of manual annotation, and effectively improve the preparation efficiency of the training data set.
[0176] In a possible implementation, Figure 5 the detection module in can be coupled with the generation module, that is, the recognizer and the discriminator are coupled. In this case, it is possible to realize the sharing of feature extraction between the discriminator and the detector, thereby greatly reducing the computational amount of the model and effectively improving the image generation efficiency.
[0177] The image generation method according to an embodiment of the present application, based on the preset semantic target information to be generated, introduces an identifier to automatically identify the semantic target to be replaced in the original image, preprocesses the semantic target, then obtains the processed image features, combines the semantic target information to be generated, generates the conditional distribution of noise and merges it with the image features, and then automatically replaces it with the expected generated semantic target, and finally synthesizes an image that conforms to the image texture information and meets the diversity and automatically annotates it to ensure the diversity of the generated data.
[0178] Figure 6 It is a structural block diagram of an image generation device provided by an embodiment of the present application. The image generation device 600 includes: a detection unit 610, a generation unit 620, and a discrimination unit 630.
[0179] Among them, the detection unit 610 is used to determine the first semantic target image of the source image, and the first semantic target image is at least one of at least one semantic target image included in the source image.
[0180] The generation unit 620 is used to determine the first background image according to the first semantic target image; generate the prior distribution of noise according to the first background image and the semantic target information to be generated; generate the target image according to the prior distribution and the first background image.
[0181] Specifically, the generation unit 620 is used to extract noise from the prior distribution and generate the target image according to the noise and the first background image.
[0182] Specifically, the generation unit 620 is used to perform smoothing processing on the first semantic target image and determine the first background image according to the smoothed first semantic target image.
[0183] In a possible implementation manner, the generation unit 620 is used to remove the smoothed first semantic target image from the source image to obtain the first background image.
[0184] Specifically, the generation unit 620 is further used to complete the annotation of the semantic target of the target image according to the semantic target information.
[0185] The training unit 630 is used to train the image generated by the generation unit to tend to be real.
[0186] Specifically, the target image is used as the input image, and the authenticity of the target image is identified by an image discriminator to be trained; according to the output result of the image discriminator and the input image, the network parameter values of an image generator are adjusted, and the image generator is used to generate the target image; the target image generated by the image generator with adjusted network parameter values is used as the input image, and the identification operation of the image discriminator to be trained is repeated until the training process converges.
[0187] In a possible implementation manner, the image discriminator to be trained discriminates images in different regions included in the target image.
[0188] Optionally, the image discriminator to be trained divides the target image into regions.
[0189] Specifically, the image discriminator to be trained is further configured to smooth the target semantic information included in the target image, and discriminate the authenticity of the target image including the smoothed target semantic information.
[0190] It should be understood that, for example, the smoothing process enlarges the bounding box of the target semantic image, and the image discriminator to be trained discriminates the authenticity of the image within the bounding box region.
[0191] It should be understood that the discriminators in the detection unit 610 and the generation unit 620 can share an image detector, and the image detector is used to extract features from the first semantic target image included in the source image. Thereby, the image generation efficiency can be improved.
[0192] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0193] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0194] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0195] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0196] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0197] In the present application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following items (pieces)" or similar expressions refer to any combination of these items, including any combination of single items (pieces) or plural items (pieces). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or plural.
[0198] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0199] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0200] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0201] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0202] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0203] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0204] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0205] As described above, the above are only specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. An image generation method, characterized in that, it includes: Determine a first semantic target image of the source image, where the first semantic target image is at least one of at least one semantic target image included in the source image; Determine a first background image according to the first semantic target image, where the first background image is the background image of the source image; Generate a first prior distribution according to the first background image and the semantic target information to be generated, where the semantic target information to be generated includes the label of the semantic target image to be generated and the indication information of the semantic target image to be generated, and the first prior distribution is the prior distribution of the noise obtained according to the first background image and the semantic target information to be generated; Generate noise of the semantic target image to be generated according to the first prior distribution; Generate a target image according to the noise of the semantic target image to be generated and the first background image, where the target image includes multiple images containing the semantic target information to be generated.
2. The method according to claim 1, characterized in that, Determining a first background image according to the first semantic target image includes: Performing smoothing processing on the first semantic target image; Determining the first background image according to the first semantic target image after the smoothing processing.
3. The method according to claim 2, characterized in that, The method further includes: Removing the first semantic target image after the smoothing processing from the source image.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: Using the target image as an input image and using a to-be-trained image discriminator to identify the authenticity of the target image; Adjusting the network parameter values of an image generator according to the output result of the to-be-trained image discriminator and the input image, where the image generator is used to generate the target image; Using the target image generated by the image generator with adjusted network parameter values as an input image, and repeating the identification action of the to-be-trained image discriminator until the training process converges.
5. The method according to claim 4, characterized in that, The using a to-be-trained image discriminator to identify the authenticity of the target image includes: The to-be-trained image discriminator discriminates the images in different regions included in the target image, including: The to-be-trained image discriminator weights the different regions when calculating the loss function in combination with the semantic target information to be generated.
6. The method according to claim 5, characterized in that, The method further includes: The to-be-trained image discriminator performs region division on the target image.
7. The method according to claim 4, characterized in that, The method further includes: The to-be-trained image discriminator performs smoothing processing on the semantic target image to be generated included in the target image; The to-be-trained image discriminator discriminates the authenticity of the target image including the smoothed semantic target image to be generated.
8. The method according to any one of claims 1-3, characterized in that, The method further includes: Annotate the semantic target of the target image according to the information of the semantic target information to be generated and the first semantic target image.
9. The method according to claim 4, wherein, the method further includes: The image discriminator to be trained includes an image detector, and the image detector is used to extract features from the first semantic target image included in the source image.
10. An image generation device, wherein, comprising: A detection unit for determining a first semantic target image of a source image, where the first semantic target image is at least one of at least one semantic target image included in the source image; A generation unit for: Determine a first background image according to the first semantic target image, where the first background image is the background image of the source image; Generate a first prior distribution according to the first background image and the semantic target information to be generated, where the semantic target information to be generated includes the label of the semantic target image to be generated and the indication information of the semantic target image to be generated, and the first prior distribution is the prior distribution of the noise obtained according to the first background image and the semantic target information to be generated; Generate noise of the semantic target image to be generated according to the first prior distribution; Generate a target image according to the noise of the semantic target image to be generated and the first background image, where the target image includes multiple images containing the semantic target information to be generated.
11. The image generation device according to claim 10, wherein, The generation unit is further configured to: Perform smoothing processing on the first semantic target image; Determine the first background image according to the first semantic target image after the smoothing processing.
12. The image generation device according to claim 11, wherein, The generation unit is further configured to: Remove the first semantic target image after the smoothing processing from the source image.
13. The image generation device according to any one of claims 10-12, wherein, The image generation device further includes a training unit for: Using the target image as an input image, and using an image discriminator to be trained to identify the authenticity of the target image; Adjust the network parameter value of the image generator according to the output result of the image discriminator to be trained and the input image, where the image generator is used to generate the target image; Using the target image generated by the image generator with the adjusted network parameter value as the input image, and repeating the identification action of the image discriminator to be trained until the training process converges.
14. The image generation device according to claim 13, wherein, The image discriminator to be trained discriminates the images in different regions included in the target image, including: The image discriminator to be trained weights the different regions when calculating the loss function in combination with the semantic target information to be generated.
15. The image generation device according to claim 14, wherein, The image discriminator to be trained divides the target image into regions.
16. The image generation device according to claim 13, wherein, The image discriminator to be trained performs smoothing processing on the to-be-generated semantic target image included in the target image; The image discriminator to be trained performs authenticity discrimination on the target image including the to-be-generated semantic target image that has undergone smoothing processing.
17. The image generation device according to any one of claims 10-12, wherein, The generation unit is further configured to: Complete the annotation of the semantic target of the target image according to the to-be-generated semantic target information and the information of the first semantic target image.
18. The image generation device according to claim 13, wherein, The image discriminator to be trained includes an image detector, and the image detector is configured to perform feature extraction on the first semantic target image included in the source image.
19. An electronic device, wherein, comprising: A processor and a memory, wherein the memory is used to store program instructions, and the processor is used to call the program instructions to execute the method according to any one of claims 1 to 3.
20. A computer-readable storage medium, wherein, The computer-readable medium stores program code for a device to execute, and the program code includes code for executing the method according to any one of claims 1 to 3.
21. A chip, wherein, The chip includes a processor and a data interface, and the processor reads instructions stored on a memory through the data interface to execute the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Local priori distribution-based interactive image segmentation method
CN107610126A
Sample image generation method and device, image processing method and device and intelligent driving control method and device
CN112200889A