Image processing method, device and storage medium

By constructing a generative adversarial network model for text-generated images, using the mutual adjustment mechanism of the generator and discriminator, the problems of slow learning feature speed and poor image accuracy of existing models are solved, and more efficient training and more accurate image generation are achieved.

CN111860555BActive Publication Date: 2025-05-23BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910362077.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-04-30
Publication Date
2025-05-23
Estimated Expiration
2039-04-30

AI Technical Summary

Technical Problem

The existing generative adversarial network model learns features slowly during training, and the generated images have poor accuracy.

Method used

A generative adversarial network model for generating images through text is constructed, including generators and discriminators, and by inputting sample image description text into the generator, generating joint images, and adjusting the generator and discriminators according to the discriminator's discriminatory results to optimize the objective function.

Benefits of technology

It improves the efficiency of network training and the accuracy of generating images, and has good applicability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111860555B_ABST
    Figure CN111860555B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image processing method, device and storage medium, which relates to the field of computer technology, wherein the method includes: constructing a generative adversarial network model, a generator obtains a generated image according to a sample image description text, generates a joint image including a sample sketch and the generated image, generates an objective function according to the discriminator's discrimination result of the joint image and the sample image description text, and generates an image using the generator adjusted based on the objective function. In the image processing method, device and storage medium disclosed in the present disclosure, the sample image description text and the sample sketch can complement each other's information, the model can be adjusted based on the sample image description text and the sample sketch, and a more accurate image is generated by the adjusted model; a new adversarial network model is provided to improve the efficiency of network training and the accuracy of the generated image, and the applicability and robustness are good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to an image processing method, device and storage medium. Background Art

[0002] The Generative Adversarial Network model is a generative model that includes a generator network and a discriminator network. The generator network and the discriminator network compete with each other until a balance is reached. The current generative adversarial network model can only generate images with the same image features as the sketch based on the sketch. During the training process, the existing generative adversarial network model learns features at a relatively slow speed, and the accuracy of the images generated by the generative adversarial network model is relatively poor. Summary of the invention

[0003] In view of this, a technical problem to be solved by the present invention is to provide an image processing method, device and storage medium.

[0004] According to one aspect of the present disclosure, there is provided an image processing method, comprising: constructing a generative adversarial network model for generating images through text; wherein the generative adversarial network model comprises: a generator and a discriminator; inputting a sample image description text into the generator to obtain a generated image output by the generator; generating a joint image containing a sample sketch and the generated image; wherein the image features corresponding to the sample sketch are the same as the image features corresponding to the image sample description text; inputting the joint image and the sample image description text into the discriminator to obtain a discrimination result; generating an objective function of the generative adversarial network model according to the discrimination result, and adjusting the generator and the discriminator based on the objective function; and using the adjusted generator to generate an image based on the image description text and output it.

[0005] Optionally, a first discrimination result is obtained for characterizing the similarity between the generated image and the sample sketch; a second discrimination result is obtained for characterizing that the generated image is false; and a third discrimination result is obtained for characterizing that the generated image is true.

[0006] Optionally, a sketch loss function of the generator is constructed according to the first discrimination result; a semantic loss function of the generator is constructed according to the second discrimination result; a generator objective function is constructed based on the sketch loss function and the semantic loss function; a discriminator loss function is constructed according to the third discrimination result; and the objective function is generated based on the generator loss function and the discriminator loss function.

[0007] Optionally, constructing the sketch loss function of the generator according to the first discrimination result includes: filtering the joint image using an image mask to obtain a generated image sketch corresponding to the generated image and a sample sketch sketch corresponding to the sample sketch; calculating the distance between the generated image sketch and the sample sketch sketch, wherein the distance is the first discrimination result, and the distance includes: KL distance and Euclidean distance; constructing the sketch loss function for determining the similarity between the generated image sketch and the sample sketch sketch based on the distance;

[0008] Optionally, constructing the semantic loss function of the generator based on the second discrimination result includes: obtaining first probability information for discriminating the generated image as false; wherein the first probability information is the second discrimination result; and constructing the semantic loss function based on the first probability information.

[0009] Optionally, constructing the generator objective function based on the sketch loss function and the semantic loss function includes: determining a weighted value corresponding to the sketch loss function or the semantic loss function; and generating the generator objective function based on the sketch loss function, the semantic loss function and the weighted value.

[0010] Optionally, inputting the sample image description text into the generator includes: using an encoder to convert the sample image description text into an input vector for representing image features; obtaining a random variable corresponding to the sample image description text; and combining the input vector and the random variable and inputting them into the generator.

[0011] Optionally, the generator objective function is: L1 = D KL (M⊙y,M⊙G(z,φ(t)))+λlog(1-D(G(z,φ(t)))); where M is the image mask, z is the random variable, φ(t) is the input vector, y is the sample sketch, G(z,φ(t) is the generated image, M⊙y is the sample sketch, M⊙G(z,φ(t)) is the generated image, D KL is the sketch loss function, D(G(z,φ(t)) is the probability that the generated image output by the generator is a real image corresponding to the sample image description text; 1-D(G(z,φ(t)) is the first probability information, and λ is the weighting value.

[0012] Optionally, constructing the discriminator loss function according to the third discrimination result includes: obtaining second probability information for discriminating that the generated image is true; wherein the second probability information is the third discrimination result; and constructing the discriminator loss function based on the second probability information.

[0013] Optionally, the discriminator loss function is L2=logD(G(z,φ(t)); wherein D(G(z,φ(t)) is the second probability information, and the second probability information is the probability that the generated image output by the generator is a real image corresponding to the sample image description text.

[0014] Optionally, the objective function is determined to be min G max D V(D,G)=L2+L1;

[0015] Among them, V(D,G) is the mathematical expectation of the generator and the discriminator after adjustment, D is the discriminator, and G is the generator.

[0016] Optionally, the generator and the discriminator are constructed using a deep convolutional neural network; wherein the generator and the discriminator include: an input layer, a fully connected layer and an output layer.

[0017] According to another aspect of the present disclosure, an image processing device is provided, comprising: a model building module, used to build a generative adversarial network model for generating images through text; wherein the generative adversarial network model comprises: a generator and a discriminator; a sample generation module, used to input a sample image description text into the generator, and obtain a generated image output by the generator; an image union module, used to generate a joint image containing a sample sketch and the generated image; wherein the image features corresponding to the sample sketch are the same as the image features corresponding to the image sample description text; an image discrimination module, used to input the joint image and the sample image description text into the discriminator, and obtain a discrimination result; a model adjustment module, used to generate an objective function of the generative adversarial network model according to the discrimination result, and adjust the generator and the discriminator based on the objective function; and an image generation module, used to use the adjusted generator to generate and output an image based on the image description text.

[0018] Optionally, the image discrimination module is used to obtain a first discrimination result for characterizing the similarity between the generated image and the sample sketch; obtain a second discrimination result for characterizing that the generated image is false; and obtain a third discrimination result for characterizing that the generated image is true.

[0019] Optionally, the model adjustment module includes: a first loss determination unit, used to construct a sketch loss function of the generator according to the first discrimination result; construct a semantic loss function of the generator according to the second discrimination result; construct a generator objective function based on the sketch loss function and the semantic loss function; a second loss determination unit, used to construct a discriminator loss function according to the third discrimination result; and generate the objective function based on the generator loss function and the discriminator loss function.

[0020] Optionally, the first loss determination unit is used to filter the joint image using an image mask to obtain a generated image sketch corresponding to the generated image and a sample sketch sketch corresponding to the sample sketch; calculate the distance between the generated image sketch and the sample sketch sketch, wherein the distance is the first discrimination result, and the distance includes: KL distance and Euclidean distance; construct the sketch loss function for determining the similarity between the generated image sketch and the sample sketch sketch based on the distance;

[0021] Optionally, the first loss determination unit is further used to obtain first probability information for determining whether the generated image is false; wherein the first probability information is the second determination result; and the semantic loss function is constructed based on the first probability information.

[0022] Optionally, the first loss determination unit is further used to determine a weighted value corresponding to the sketch loss function or the semantic loss function; and generate the generator objective function based on the sketch loss function, the semantic loss function and the weighted value.

[0023] Optionally, the sample generation module is further used to use an encoder to convert the sample image description text into an input vector for representing image features; obtain a random variable corresponding to the sample image description text; and combine the input vector and the random variable and input them into the generator.

[0024] Optionally, the generator objective function is: L1 = D KL (M⊙y,M⊙G(z,φ(t)))+λlog(1-D(G(z,φ(t)))); where M is the image mask, z is the random variable, φ(t) is the input vector, y is the sample sketch, G(z,φ(t) is the generated image, M⊙y is the sample sketch sketch, N⊙G(z,φ(t)) is the generated image sketch, D KLis the sketch loss function, D(G(z,φ(t)) is the probability that the generated image output by the generator is a real image corresponding to the sample image description text; 1-D(G(z,φ(t)) is the first probability information, and λ is the weighting value.

[0025] Optionally, the second loss determination unit is used to obtain second probability information for discriminating whether the generated image is true; wherein the second probability information is the third discrimination result; and the discriminator loss function is constructed based on the second probability information.

[0026] Optionally, the discriminator loss function is L2=logD(G(z,φ(t)); wherein D(G(z,φ(t)) is the second probability information, and the second probability information is the probability that the generated image output by the generator is a real image corresponding to the sample image description text.

[0027] Optionally, the model adjustment module includes: an objective function determination unit, configured to determine that the objective function is min G max D V(D,G)=L2+L1; wherein V(D,G) is the mathematical expectation of the generator and the discriminator after adjustment, D is the discriminator, and G is the generator.

[0028] Optionally, the model building module is used to build the generator and the discriminator using a deep convolutional neural network; wherein the generator and the discriminator include: an input layer, a fully connected layer and an output layer.

[0029] According to another aspect of the present disclosure, an image processing device is provided, including: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the method as described above based on instructions stored in the memory.

[0030] According to another aspect of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the instructions are executed by a processor to perform the above method.

[0031] The image processing method, device and storage medium disclosed in the present invention are as follows: a generator of an adversarial network model obtains a generated image according to a sample image description text, generates a joint image including a sample sketch and the generated image, generates an objective function according to a discriminator's discrimination result of the joint image and the sample image description text, and generates an image using a generator adjusted based on the objective function; the sample image description text can describe the category and color of the image, and the sample sketch can describe the specific details, posture, position and size of the image; the model can be adjusted based on the sample image description text and the sample sketch, and a more accurate image is generated by the adjusted model; a new adversarial network model is provided to improve the efficiency of network training and the accuracy of the generated image, and has good applicability and robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0033] Figure 1 is a flowchart of an embodiment of an image processing method according to the present disclosure;

[0034] Figure 2 A schematic diagram of a flow chart of generating an objective function in one embodiment of an image processing method according to the present disclosure;

[0035] Figure 3 is a schematic diagram of a generative adversarial network model in one embodiment of the image processing method according to the present disclosure;

[0036] Figure 4 A schematic diagram of a process for generating a sketch loss function in one embodiment of an image processing method according to the present disclosure;

[0037] Figure 5 is a module schematic diagram of an embodiment of an image processing device according to the present disclosure;

[0038] Figure 6 is a module schematic diagram of a model adjustment module in one embodiment of an image processing device according to the present disclosure;

[0039] Figure 7 FIG. 4 is a module diagram of another embodiment of an image processing device according to the present disclosure. DETAILED DESCRIPTION

[0040] The present disclosure is described more fully below with reference to the accompanying drawings, in which exemplary embodiments of the present disclosure are described. The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure. The technical solutions of the present disclosure are described in many aspects below in conjunction with various figures and embodiments.

[0041] The terms "first", "second", etc. in the following text are only used to distinguish between the two and have no other special meanings.

[0042] Figure 1 FIG. 1 is a flow chart of an embodiment of an image processing method according to the present disclosure. Figure 1 As shown:

[0043] Step 101, constructing a generative adversarial network model for generating images through text, wherein the generative adversarial network model includes: a generator and a discriminator.

[0044] Generative Adversarial Network (GAN) is a deep learning model that produces fairly good outputs through the mutual game learning between the generator model and the discriminator model. Generative adversarial networks usually use deep neural networks as generators G and discriminators D. Based on the idea of ​​"game theory", the generator generates images, and the discriminator judges the input images to determine whether they are images from the dataset or images generated by the generator.

[0045] Step 102: Input the sample image description text into the generator to obtain the generated image output by the generator.

[0046] The sample image description text is a text used to describe the sample image. For example, if the sample image (real image) is an image of a bird, the sample image description text is a text describing the image of the bird, which may be "the bird has black eyes, white feathers, a sharp beak, and red claws".

[0047] Step 103, generating a joint image including the sample sketch and the generated image, wherein the image features corresponding to the sample sketch are the same as the image features corresponding to the image sample description text.

[0048] The sample sketch can be drawn by the user and input by the user. The sample sketch is a sketch corresponding to the image sample description text and the sample image. For example, the user draws a sample sketch by hand, and the sample sketch is a sample sketch corresponding to the image sample description text "the bird has black eyes, white feathers, a sharp beak, and red claws". After the generator outputs the generated image, it obtains the sample sketch and generates a joint image containing the sample sketch and the generated image.

[0049] Step 104: input the joint image and the sample image description text into the discriminator to obtain a discrimination result.

[0050] Step 105: Generate an objective function of the generative adversarial network model according to the discrimination result, and adjust the generator and the discriminator based on the objective function.

[0051] The discriminator is equivalent to a binary classifier, which can distinguish whether the input image comes from a real sample image or an image generated by the generator, and can determine the probability of whether the generated image is a real sample image. The objective function can be determined based on the loss function of the generator and the discriminator. Sample images, image sample description texts, and sample sketches can be used as training sets, and the generative adversarial network model can be trained based on the objective function to adjust the generator and the discriminator. The generator and the discriminator can be adjusted through the existing iterative training to improve the accuracy of the network model.

[0052] For example, the generator G generates images, and the discriminator D determines whether an image is "real". The input of the discriminator is image x, and the output D(x) represents the probability that image x is a real image. If it is 1, it means that image x is 100% a real image, and the output is 0, which means that image x is not a real image.

[0053] The goal of the generator G is to generate realistic images as much as possible to deceive the discriminator D. The goal of the discriminator D is to distinguish the images generated by the generator G from the real images as much as possible. The generator G and the discriminator D constitute a dynamic "game process".

[0054] The result of the game between the generator G and the discriminator D is that in an ideal state, the generator G can generate a picture G(z) that is "indistinguishable from the real thing", while the discriminator D has difficulty determining whether the picture generated by the generator G is real. When D(G(z)) = 0.5, a generator G is obtained that can be used to generate pictures.

[0055] Step 106: Generate an image based on the image description text using the adjusted generator and output it. The trained generator can be used to generate an image based on the image description text.

[0056] In the image processing method in the above embodiment, the sample image description text can describe the category and color of the specific sample image, the sample sketch can describe the specific details, posture, position and size of the sample image, the sample image description text and the sample sketch can complement each other's information, and the model is trained based on the sample image description text and the sample image description text, and the trained model can generate more accurate images.

[0057] There are many ways to input the sample image description text into the generator. For example, use an encoder to convert the sample image description text into an input vector for representing image features, obtain a random variable corresponding to the sample image description text, combine the input vector and the random variable, and input it into the generator. The random variable can be random noise Z~N(0,1) and the like.

[0058] Convolutional neural network is a common deep learning network. The generator and discriminator are constructed using deep convolutional neural network. The generator and discriminator may include input layer, fully connected layer and output layer. For example, the generator network may include an input layer, multiple convolution layers, multiple maximum pooling layers, multiple deconvolution layers and output layer connected in sequence. The input of the input layer is random noise and input vector, and the output layer outputs the generated image.

[0059] In one embodiment, the discriminator may obtain multiple types of discriminant result information, for example, obtaining a first discriminant result for characterizing the similarity between the generated image and the sample sketch, obtaining a second discriminant result for characterizing that the generated image is false, and obtaining a third discriminant result for characterizing that the generated image is true.

[0060] Figure 2 FIG. 1 is a schematic diagram of generating an objective function in an embodiment of an image processing method according to the present disclosure, as shown in FIG. Figure 2 As shown:

[0061] Step 201, constructing a sketch loss function of the generator according to the first discrimination result.

[0062] Step 202, constructing a semantic loss function of the generator according to the second discrimination result.

[0063] Step 203, constructing a generator objective function based on the sketch loss function and the semantic loss function.

[0064] Step 204: construct a discriminator loss function according to the third discrimination result.

[0065] Step 205: Generate an objective function based on the generator loss function and the discriminator loss function.

[0066] like Figure 3As shown in the figure, the generative adversarial network model includes a generator network and a discriminator network. The generator network receives the random variables z and φ combined, and outputs a generated image. A sample sketch is obtained, and a joint image containing the sample sketch and the generated image is generated. The left side of the joint image is the sample sketch, and the right side is the generated image.

[0067] The discriminator network determines whether the generated joint image is real. φ is the input vector representing the image features obtained by processing the sample image description text through the text encoder. The sample image description text can be "a bird with a black head, red and yellow body, gray tail and blue back".

[0068] z is a noise prior that satisfies a normal distribution. The file encoder can be an encoder based on a neural network model, and the generated input vector can be a visual vector. There can be many types of text encoders. For example, the text encoder can be a pre-trained hybrid character-level convolutional-recurrent network model. The text encoder generates a visual vector that corresponds to the sample image description text and can represent the image features.

[0069] Combine the random variable z and the vector generated by encoding the description text t through the encoder φ to generate a new variable, and input the new variable into the generator network. The generator network uses a fully connected convolutional layer and a Leaky ReLUs layer to output an embedding output layer of a specific length.

[0070] The joint image can be a joint image pair of a sample sketch and a generated image, and the joint image pair contains contextual content information. The discriminator in the generative adversarial network model needs to distinguish between generated images and real images, and the generator needs to generate pictures that can deceive the discriminator.

[0071] Figure 4 FIG. 1 is a flow chart of generating a sketch loss function in one embodiment of the image processing method according to the present disclosure, as shown in FIG. Figure 4 As shown:

[0072] Step 401 , use the image mask to filter the joint image to obtain a generated image sketch corresponding to the generated image and a sample sketch sketch corresponding to the sample sketch.

[0073] The image mask may be a binary mask for masking a specified portion of the image, and the joint image may be filtered to generate a stick figure, which is a figure using some simple lines. The obtained stick figures are a generated image stick figure corresponding to the generated image and a sample sketch stick figure corresponding to the sample sketch.

[0074] Step 402, calculate the distance between the generated image sketch and the sample sketch, wherein the distance is the first discrimination result, and the distance includes: KL distance, Euclidean distance, etc. KL distance is Kullback-Leibler difference, and Euclidean distance is Euclidean metric.

[0075] Step 403: construct a sketch loss function based on the distance to determine the similarity between the generated image sketch and the sample sketch.

[0076] There are many ways to construct the semantic loss function of the generator based on the second discrimination result to obtain the first probability information for judging whether the generated image is false. The first probability information is the second discrimination result; the semantic loss function is constructed based on the first probability information. The generated image output by the generator is input to the discriminator. If the discriminator determines that the input image is a sample image corresponding to the sample image description text, the generated image is judged to be true. If the discriminator determines that the input image is the generated image output by the generator, the generated image is judged to be false. Determine the weighted value corresponding to the sketch loss function or the semantic loss function. Generate the generator objective function based on the sketch loss function, the semantic loss function and the weighted value.

[0077] For example, the generator objective function is: L1 = D KL (M⊙y,M⊙G(z,φ(t)))+λlog(1-D(G(z,φ(t)))); where M is the image mask, z is a random variable (random noise), φ(t) is the input vector, y is the sample sketch, G(z,φ(t) is the generated image, M⊙y is the sample sketch, M⊙G(z,φ(t)) is the generated image, and D KL is the sketch loss function, D(G(z,φ(t)) is the probability that the generated image output by the generator is the real image (sample image) corresponding to the sample image description text; 1-D(G(z,φ(t)) is the first probability information, and λ is the weighting value.

[0078] Obtain the second probability information for discriminating the generated image as real, and the second probability information is the third discrimination result. Construct the discriminator loss function based on the second probability information. For example, the discriminator loss function is L2=logD(G(z,φ(t)); where D(G(z,φ(t)) is the second probability information, and the second probability information is the probability that the generated image output by the generator is a real image corresponding to the sample image description text.

[0079] Determine the objective function as min G max D V(D,G)=L2+L1; wherein V(D,G) can be the mathematical expectation of the generator and the discriminator after adjustment, D is the discriminator, and G is the generator.

[0080] In one embodiment, for a variable combining a random variable z and a text encoding φ(t), t is a sample image description text, determining the loss function of the generator includes two loss functions. The first loss function is to minimize the similarity between the sketch part in the generated image G(z, φ(t)) and the sample sketch B in the joint image AB, and the KL distance is used to measure the similarity between the sketch part in the image G(z, φ(t)) and the input sample sketch B.

[0081] Sketch loss function

[0082] L contextual (z,φ(t))=D KL (M⊙y,M⊙G(z,φ(t))); where M is a binary mask and ⊙ is the Hadamard product. M⊙y represents the sample sketch of the joint image AB after being filtered by M. If the sketch part in the image G(z,φ(t)) is the same as the input sample sketch B, then L contextual (z,φ(t)) is 0.

[0083] The second loss function contains the semantic level content. The adversarial loss function of the generator network is: semantic loss function L p-eceptual (z,φ(t))=log(1-D(G(z,φ(t)))). For the input (z,φ(t)), the objective function of the generator becomes the sum of two loss functions:

[0084] L contextual (z,φ(t))+λL p-eceptual (z,φ(t));

[0085] The objective function of the generative adversarial network model is:

[0086] min G max D V(D,G)=logD(()+L conte(tual (z,φ(t))+

[0087] λL preceptual (z,φ(t)).

[0088] The generator and the discriminator can be adjusted based on the objective function using a variety of existing methods. After the adjustment, a generator that meets the expected value is obtained, and an image is generated and output based on the image description text.

[0089] In one embodiment, Figure 5As shown, the present disclosure provides an image processing device 50, including: a model construction module 51, a sample generation module 52, an image combination module 53, an image discrimination module 54, a model adjustment module 55 and an image generation module 56. The model construction module 51 constructs a generative adversarial network model for generating images through text; wherein, the generative adversarial network model includes: a generator and a discriminator. The model construction module 51 uses a deep convolutional neural network to construct a generator and a discriminator, and the generator and the discriminator include: an input layer, a fully connected layer and an output layer.

[0090] The sample generation module 52 inputs the sample image description text into the generator to obtain the generated image output by the generator. The image combination module 53 generates a combined image including the sample sketch and the generated image, and the image features corresponding to the sample sketch are the same as the image features corresponding to the image sample description text.

[0091] The image discrimination module 54 inputs the joint image and the sample image description text into the discriminator to obtain the discrimination result. The model adjustment module 55 generates the objective function of the generative adversarial network model based on the discrimination result, and adjusts the generator and the discriminator based on the objective function. The image generation module 56 uses the adjusted generator to generate an image based on the image description text and outputs it.

[0092] The image discrimination module 54 obtains a first discrimination result for characterizing the similarity between the generated image and the sample sketch, obtains a second discrimination result for characterizing that the generated image is false, and obtains a third discrimination result for characterizing that the generated image is true.

[0093] like Figure 6 As shown, the model adjustment module 55 includes: a first loss determination unit 551, a second loss determination unit 552 and an objective function determination unit 553. The first loss determination unit 551 constructs a sketch loss function of the generator according to the first discrimination result, constructs a semantic loss function of the generator according to the second discrimination result, and constructs a generator objective function based on the sketch loss function and the semantic loss function. The second loss determination unit 552 constructs a discriminator loss function according to the third discrimination result, and generates an objective function based on the generator loss function and the discriminator loss function.

[0094] The first loss determination unit 551 uses the image mask to filter the joint image to obtain a generated image sketch corresponding to the generated image and a sample sketch sketch corresponding to the sample sketch. The first loss determination unit 551 calculates the distance between the generated image sketch and the sample sketch sketch, where the distance is the first discrimination result, and the distance includes: KL distance, Euclidean distance, etc. The first loss determination unit 551 constructs a sketch loss function for determining the similarity between the generated image sketch and the sample sketch sketch based on the distance.

[0095] The first loss determination unit 551 obtains first probability information that the generated image is false, the first probability information is the second discrimination result, and a semantic loss function is constructed based on the first probability information. The first loss determination unit 551 determines a weighted value corresponding to the sketch loss function or the semantic loss function, and generates a generator objective function based on the sketch loss function, the semantic loss function, and the weighted value.

[0096] The sample generation module 52 uses an encoder to convert the sample image description text into an input vector for representing image features. The sample generation module 52 obtains a random variable corresponding to the sample image description text, combines the input vector and the random variable, and inputs them into the generator. The generator objective function is: L1 = D KL (M⊙y,M⊙G(z,φ(t)))+λlog(1-D(G(z,φ(t)))).

[0097] The second loss determination unit 552 obtains the second probability information of determining whether the generated image is true, and the second probability information is the third discrimination result. The discriminator loss function is constructed based on the second probability information. The discriminator loss function is L2=logD(G(z,φ(t)). The objective function determination unit 553 determines the objective function as min G max D V(D,G)=L2+L1; where V(D,G) is the mathematical expectation of the generator and discriminator after adjustment, D is the discriminator, and G is the generator.

[0098] According to another aspect of the present disclosure, there is provided an image processing device, including: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the above method based on instructions stored in the memory.

[0099] In one embodiment, Figure 7 FIG. 1 is a schematic diagram of a module of another embodiment of an image processing device according to the present disclosure. Figure 7 As shown, the apparatus may include a memory 71, a processor 72, a communication interface 73 and a bus 74. The memory 71 is used to store instructions, the processor 72 is coupled to the memory 71, and the processor 72 is configured to execute the above-mentioned image processing method based on the instructions stored in the memory 71.

[0100] The memory 71 may be a high-speed RAM memory, a non-volatile memory, etc., or a memory array. The memory 71 may also be divided into blocks, and the blocks may be combined into virtual volumes according to certain rules. The processor 72 may be a central processing unit CPU, or an application specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the image processing method of the present disclosure.

[0101] In one embodiment, the present disclosure provides a logistics system, including: a robot, and an image processing device as in any of the above embodiments, wherein the image processing device sends three-dimensional center position information and posture information of a target object to the robot.

[0102] In one embodiment, the present disclosure provides a computer-readable storage medium storing computer instructions, which implement the image processing method in any of the above embodiments when executed by a processor.

[0103] The image processing method, device and storage medium provided in the above embodiments construct a generative adversarial network model, in which a generator obtains a generated image based on a sample image description text, generates a joint image including a sample sketch and a generated image, generates an objective function based on the discrimination result of the discriminator for the joint image and the sample image description text, and generates an image using a generator adjusted based on the objective function; the sample image description text can describe the category and color of the image, etc., the sample sketch can describe the specific details, posture, position and size of the image, the sample image description text and the sample sketch can complement each other, and the model can be adjusted based on the sample image description text and the sample sketch, and a more accurate image is generated by the adjusted model; a new adversarial network model is provided to improve the efficiency of network training and the accuracy of the generated images, and has good applicability and robustness.

[0104] The method and system of the present disclosure may be implemented in many ways. For example, the method and system of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.

[0105] The description of the present disclosure is given for the purpose of illustration and description, and is not intended to be exhaustive or to limit the present disclosure to the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present disclosure, and to enable those of ordinary skill in the art to understand the present disclosure and thereby design various embodiments with various modifications suitable for specific uses.

Claims

1. An image processing method, include: Constructing a generative adversarial network model for generating images from text; wherein the generative adversarial network model includes: a generator and a discriminator; Inputting a sample image description text into the generator to obtain a generated image output by the generator; Generate a joint image including a sample sketch and the generated image; wherein the image features corresponding to the sample sketch are the same as the image features corresponding to the sample image description text; after outputting the generated image, the generator obtains the sample sketch and generates a joint image including the sample sketch and the generated image; Inputting the joint image and the sample image description text into the discriminator to obtain a discrimination result; Generate an objective function of the generative adversarial network model according to the discrimination result, and adjust the generator and the discriminator based on the objective function; The adjusted generator is used to generate an image based on the image description text and output it.

2. The method according to claim 1, wherein the identification result is obtained include: Obtaining a first discrimination result for characterizing the similarity between the generated image and the sample sketch; Obtaining a second discrimination result for characterizing that the generated image is false; and, A third discrimination result is obtained, which is used to characterize that the generated image is true.

3. The method according to claim 2, further comprising: include: Constructing a sketch loss function of the generator according to the first discrimination result; Constructing a semantic loss function of the generator according to the second discrimination result; Constructing a generator objective function based on the sketch loss function and the semantic loss function; Constructing a discriminator loss function according to the third discrimination result; The objective function is generated based on the generator objective function and the discriminator loss function.

4. The method of claim 3, wherein the sketch loss function of the generator is constructed according to the first discrimination result include: Using an image mask to filter the combined image to obtain a generated image sketch corresponding to the generated image and a sample sketch sketch corresponding to the sample sketch; Calculating the distance between the generated image sketch and the sample sketch, wherein the distance is the first discrimination result, and the distance includes: KL distance and Euclidean distance; The sketch loss function for determining the similarity between the generated image stick figure and the sample sketch stick figure is constructed based on the distance.

5. The method according to claim 4, wherein the semantic loss function of the generator is constructed according to the second discrimination result include: Obtaining first probability information for determining that the generated image is false; wherein the first probability information is the second determination result; The semantic loss function is constructed based on the first probability information.

6. The method of claim 5, wherein the generator objective function is constructed based on the sketch loss function and the semantic loss function include: Determining a weighted value corresponding to the sketch loss function or the semantic loss function; The generator objective function is generated based on the sketch loss function, the semantic loss function and the weighted value.

7. The method of claim 6, wherein the sample image description text is input into the generator include: Using an encoder, convert the sample image description text into an input vector for representing image features; Obtaining a random variable corresponding to the sample image description text; The input vector and the random variable are combined and input into the generator.

8. The method according to claim 7, in, The generator objective function is: L1 = D KL (M⊙y, M⊙G(z, φ(t))) + λ log(1 - D(G(z, φ(t)))); Wherein, M is the image mask, z is the random variable, φ(t) is the input vector, y is the sample sketch, G(z,φ(t)) is the generated image, M⊙y is the sample sketch sketch, M⊙G(z,φ(t)) is the generated image sketch, D KL is the sketch loss function, D(G(z,φ(t)) is the probability that the generated image output by the generator is a real image corresponding to the sample image description text; 1-D(G(z,φ(t)) is the first probability information, and λ is the weighting value.

9. The method of claim 6, wherein the discriminator loss function is constructed according to the third discrimination result. include: Obtaining second probability information for determining whether the generated image is true; wherein the second probability information is the third determination result; The discriminator loss function is constructed based on the second probability information.

10. The method according to claim 9, in, The discriminator loss function is L2=log D(G(z,φ(t)); wherein D(G(z,φ(t)) is the second probability information, and the second probability information is the probability that the generated image output by the generator is a real image corresponding to the sample image description text.

11. The method according to claim 10, in, Determine the objective function as Among them, V(D,G) is the mathematical expectation of the generator and the discriminator after adjustment, D is the discriminator, and G is the generator.

12. The method according to any one of claims 1 to 11, in, Using a deep convolutional neural network to construct the generator and the discriminator; Wherein, the generator and the discriminator include: an input layer, a fully connected layer and an output layer.

13. An image processing device, include: A model building module, used to build a generative adversarial network model for generating images from text; wherein the generative adversarial network model includes: a generator and a discriminator; A sample generation module, used for inputting a sample image description text into the generator to obtain a generated image output by the generator; An image combination module, used to generate a combined image including a sample sketch and the generated image; wherein the image features corresponding to the sample sketch are the same as the image features corresponding to the sample image description text; after the generator outputs the generated image, it obtains the sample sketch and generates a combined image including the sample sketch and the generated image; An image discrimination module, used for inputting the joint image and the sample image description text into the discriminator to obtain a discrimination result; A model adjustment module, which generates an objective function of the generative adversarial network model according to the discrimination result, and adjusts the generator and the discriminator based on the objective function; The image generation module is used to generate an image based on the image description text using the adjusted generator and output it.

14. An image processing device, include: Memory; and a processor coupled to the memory, the processor being configured to execute the method according to any one of claims 1 to 12 based on instructions stored in the memory. 15 . A computer-readable storage medium storing computer instructions, wherein the instructions are executed by a processor to execute the method according to claim 1 .

Citation Information

Patent Citations

  • A method for cross-modal retrieval of three-dimensional model based on sketch retrieval

    CN109213884A

  • Infinite terrain generation method, system, storage medium and terminal based on cGAN

    CN109215123A