Image generation model training method, service execution method, apparatus and medium

By using noise addition and image restoration techniques within the image generation model, the method addresses the challenge of insufficient image data, enhancing the model's ability to generate high-quality images with similar foreground features.

JP2025092369AActive Publication Date: 2025-06-19ZHEJIANG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024088711
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-05-31
Publication Date
2025-06-19
Estimated Expiration
2044-05-31

AI Technical Summary

Technical Problem

Existing model training technologies face challenges in training image generation models that meet service requirements when the amount of image data is insufficient, leading to issues with expressiveness and realism in generated artistic images.

Method used

The method involves acquiring an original image, performing noise addition processing to obtain a post-noise addition image, and inputting this image into a first image generation model to predict the superimposed noise signal and restore the original image. The model is trained to minimize the difference between the original and restored image foreground features.

Benefits of technology

This approach enables the generation of images with foreground features similar to the original image, even with limited data, improving the training efficiency and quality of the image generation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025092369000001_ABST
    Figure 2025092369000001_ABST
Patent Text Reader

Abstract

To provide an image generation model training method configured to satisfy needs for training a model that satisfies service requirements with a limited training set, thereby improving overall training efficiency of the model, a service execution method, an apparatus, and a medium.SOLUTION: An image generation model training method includes: acquiring an image; carrying out noise addition processing on the acquired original image to obtain a noise-added image; inputting the noise-added image to a first image generation model, denoising the noise-added image using the first image generation model to obtain a restored image, and determining image foreground feature extracted from the restored image; and training the first image generation model with an optimization goal so as to minimize a difference between image foreground feature corresponding to the original image and the image foreground feature extracted from the restored image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and more particularly, to a method for training an image generation model, a service execution method, an apparatus, and a medium.

Background Art

[0002] With the rapid development of the field of computer vision, models based on deep learning technology have been increasingly used in various fields, such as the field of image recognition for determining the content of an image according to user needs, the medical field for generating an image of a lesion based on the obtained lesion information of a patient, and the computer game field for generating a game image according to the exploration status of a player.

[0003] However, in existing model training technologies, in order to train a model that meets service requirements, it is necessary to use a sufficient amount of image data as training data in the training process of the above model. If the amount of image data is not sufficient, it is difficult to train a model that meets service requirements using existing technologies. For example, the higher the expressiveness and realism of the artistic images generated by an artistic image generation model, the more image data is required in the training process. If the image data is insufficient, there will be a problem that the expressiveness and realism of the artistic images generated by the trained artistic image generation model cannot meet the requirements.

[0004] Therefore, how to train a model that meets service requirements based on insufficient image data is an urgent issue.

Summary of the Invention

Problems to be Solved by the Invention

[0005] The present invention provides a method for training an image generation model, a service execution method, an apparatus, and a medium to partially solve the above problems existing in the prior art.

Means for Solving the Problems

[0006] The technical solution adopted by the present invention is as follows.

[0007] The method for training an image generation model provided by the present invention is a step of obtaining an original image, a step of performing noise addition processing on the original image to obtain a post-noise addition image, inputting the post-noise addition image and the number of times the post-noise addition image has been noise-added into a first image generation model, and using the first image generation model to predict the superimposed noise signal used when converting the original image into the post-noise addition image by performing noise addition processing of the number of times on the original image until a restored image is obtained, predicting the (k-1)-th transition image before the k-th noise addition processing based on the superimposed noise signal, predicting the (k-2)-th transition image before the (k-1)-th noise addition processing based on the superimposed noise signal and the (k-1)-th transition image, and determining the image foreground features extracted from the restored image, where k is a positive integer not exceeding the number of times, and the image foreground features are for representing the morphological features of the target object in the image and do not include detailed physical features for representing the target object, and the step, training the first image generation model with the optimization goal of minimizing the difference between the image foreground features corresponding to the original image and the image foreground features extracted from the restored image.

[0008] Optionally, the step of performing noise addition processing on the original image to obtain a post-noise addition image is inputting the original image and the number of noise signals into a pre-constructed second image generation model, and causing the second image generation model to output a post-noise addition image of the original image that has undergone noise addition processing the number of times corresponding to the number of noise signals.

[0009] Optionally, the construction of the second image generation model is a step of obtaining a sample image, A step of adding noise to the image after noise addition that has been added with the (N - 1)-th noise signal using the N-th noise signal to obtain an image after noise addition that has been added with the N-th noise signal, where N is a positive integer greater than or equal to 1, and the image after noise addition that has been added with the 0-th noise signal is the sample image. A step of determining a conversion relationship from the image after noise addition that has been added with the (N - m)-th noise signal to the image after noise addition that has been added with the N-th noise signal based on the image after noise addition that has been added with the N-th noise signal, the image after noise addition that has been added with the (N - m)-th noise signal, the N-th noise signal, and the (N - m + 1)-th noise signal, where m is a positive integer smaller than N. Including the step of constructing the second image generation model based on the conversion relationship.

[0010] The service execution method provided by the present invention includes A step of acquiring an initial image. A step of inputting the initial image into a pre-trained image generation model to output a target image, where the image generation model is a model obtained by training using the above training method. Including the step of constructing a training set based on the initial image and the target image, training a predetermined specified model using the training set, and executing a service using the trained specified model.

[0011] The training device for an image generation model provided by the present invention includes An acquisition module for acquiring an original image. A noise addition module for performing a noise addition process on the original image to obtain an image after noise addition. Input the image after noise addition and the number of times the image after noise addition has been noise-added into a first image generation model, and use the first image generation model to predict the superimposed noise signal used when converting the original image into the image after noise addition by performing the noise addition process of the number of times on the original image until a restored image is obtained. Based on the superimposed noise signal, predict the (k - 1)-th transition image before the k-th noise addition process, and based on the superimposed noise signal and the (k - 1)-th transition image, predict the (k - 2)-th transition image before the (k - 1)-th noise addition process, and an input module for determining the image foreground features extracted from the restored image. A training module for training the first image generation model with the optimization goal of minimizing the difference between the image foreground features corresponding to the original image and the image foreground features extracted from the restored image.

[0012] Optionally, specifically, the noise addition module Inputs the original image and the number of noise signals into a pre-constructed second image generation model, and is used to cause the second image generation model to output the image after noise addition of the original image that has undergone the noise addition process the number of times corresponding to the number of noise signals.

[0013] Optionally, specifically, the noise addition module Obtains a sample image. Uses the N-th noise signal to add noise to the image after noise addition that has been noise-added with the (N - 1)-th noise signal to obtain the image after noise addition that has been noise-added with the N-th noise signal, where N is a positive integer greater than or equal to 1, and the image after noise addition that has been noise-added with the 0-th noise signal is the sample image. Based on the image after noise addition that has been noise-added with the N-th noise signal, the image after noise addition that has been noise-added with the (N - m)-th noise signal, the N-th noise signal, and the (N - m + 1)-th noise signal, determine the conversion relationship from the image after noise addition that has been noise-added with the (N - m)-th noise signal to the image after noise addition that has been noise-added with the N-th noise signal, where m is a positive integer smaller than N. It is used to construct the second image generation model based on the conversion relationship.

[0014] The service execution device provided by the present invention includes an acquisition module for acquiring an initial image, an input module for inputting the initial image into a pre-trained image generation model to output a target image, where the image generation model is a model obtained by training using the above training method, and a training module for constructing a training set based on the initial image and the target image, training a predetermined specified model using the training set, and executing a service using the trained specified model.

[0015] The present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the training method of the above image generation model is implemented.

[0016] The present invention provides an electronic device including a processor and a computer program stored in a memory and operable on the processor, and when the processor executes the computer program, the training method of the above image generation model is implemented.

Advantages of the Invention

[0017] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects.

[0018] In the method for training an image generation model provided by the present invention, a dedicated device performs noise addition processing on the acquired original image using a second image generation model, inputs the noise-added image into a first image generation model, performs noise removal on the noise-added image using the first image generation model to obtain a restored image, determines the image foreground features extracted from the restored image, and uses as an optimization target minimizing the difference between the image foreground features corresponding to the original image and the image foreground features extracted from the restored image to train the first image generation model. The trained first image generation model outputs a restored image having image foreground features similar to the original image based on the input image after noise addition.

[0019] In the service execution method provided by the present invention, after acquiring an initial image, the initial image is input into an image generation model pre-trained using the above training method to output a target image, a training set is constructed based on the initial image and the target image, a predetermined specified model is trained using the training set, and the trained specified model is used to execute a service.

[0020] As can be seen from the above method, by adding noise to the image data of the training set, inputting it into a pre-trained image generation model for noise removal, an image similar to the image foreground features before adding noise but with different other details can be generated. By using the generated image as image data for expanding the training set, the need to train a model that meets service requirements with a limited training set can be satisfied, and the overall training efficiency of the model can be improved.

Brief Description of the Drawings

[0021] The drawings described here are used to provide a further understanding of the present invention, constitute a part of the present invention, and the schematic embodiments and their descriptions of the present invention are used to interpret the present invention and do not constitute an undue limitation on the present invention.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiments for Carrying Out the Invention

[0022] In order to make the object, technical solution and advantages of the present invention clearer, hereinafter, in conjunction with the specific embodiments of the present invention and the corresponding drawings, the technical solution of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] Hereinafter, the technical solutions provided by each embodiment of the present invention will be described in detail.

[0024] FIG. 1 is a schematic diagram showing the flow of the training method for the image generation model provided by the present invention, and includes the following steps.

[0025] In step S101, an original image is acquired.

[0026] The execution subject of the training method for the image generation model provided by the present invention may be any of a terminal device such as a notebook computer or a desktop computer, a client installed on the terminal device, a server, or a dedicated device for training the model. Hereinafter, for the sake of convenience of description, only an example in which the execution subject is a dedicated device will be used to describe the training method for the image generation model provided by the present invention.

[0027] In the current field of computer vision, models based on deep learning technology often require a sufficient amount of image data for training. When the amount of image data is insufficient, it is difficult for the trained model to meet service requirements. Currently, for the problem of insufficient image data, the mainstream method is to perform operations such as rotation, flipping, translation, and filling on the initial image to generate additional image data for model training. However, the image data generated by the above processing methods may have problems such as the structure and content of the image being too similar to the initial image, or the foreground features of the image in the image being too different from the original image, that is, the realism is too low. The generalization ability of the model trained with such potentially problematic image data as the training set is affected, and as a result, the performance of the model cannot meet service requirements.

[0028] Based on this, the present invention provides a method for training an image generation model. A dedicated device acquires the original image, and then the dedicated device inputs the original image with added noise into the first image generation model for noise removal to obtain a restored image. By training the model based on the foreground features of the image extracted from the original image and the foreground features of the obtained restored image, an image generation model capable of determining the restored image from the original image is obtained.

[0029] In the training process of the image generation model, the dedicated device first needs to acquire the original image used as a training sample. Here, the original image may be acquired from a predetermined image set or by a predetermined acquisition device.

[0030] In step S102, perform noise addition processing on the original image to obtain an image after noise addition.

[0031] After the dedicated device acquires the original image, it is necessary to perform noise addition processing on the original image to obtain the image after noise addition. When adding noise to the original image using a plurality of noise signals, it is necessary to sequentially add the noise signals to the original image. That is, after adding noise to the original image with the first noise signal, the second noise signal is used to add noise to the original image that has been added noise with the first noise signal. The same applies hereinafter. After all the noise signals are added to the original image, the image after noise addition is obtained.

[0032] Note that the scale of the calculation performed to sequentially add noise signals is large. Especially when the number of image data is large, if each is calculated multiple times, the calculation load becomes large. Therefore, in the present invention, a method for adding noise to the original image based on a pre-constructed second image generation model is provided. The pre-constructed second image generation model can output an image after noise addition of the original image on which noise addition processing has been performed the number of times corresponding to the number of the noise signals, based on the input original image and the number of noise signals.

[0033] Taking as an example that the noise signals added to the original image are a plurality of Gaussian distribution signals, in the process of constructing the second image generation model, first, it is necessary to obtain a sample image. Next, using the Nth Gaussian distribution signal, noise is added to the image after noise addition that has been added noise with the (N - 1)th Gaussian distribution signal to obtain an image after noise addition that has been added noise with the Nth Gaussian distribution signal. N is a positive integer of 1 or more, and the image after noise addition that has been added noise with the 0th noise signal is the sample image.

[0034] Based on the image after noise addition with the Nth noise signal, the image after noise addition with the (N - m)th noise signal, the Nth noise signal, and the (N - m + 1)th noise signal, determine the conversion relationship from the image after noise addition with the (N - m)th noise signal to the image after noise addition with the Nth noise signal, and construct the second image generation model based on the conversion relationship from the image after noise addition with the (N - m)th noise signal to the image after noise addition with the Nth noise signal. m is a positive integer smaller than N.

[0035] Specifically, in the process of constructing the second image generation model, the dedicated device queries and obtains the additive information of the Gaussian distribution signal based on the received fitting instruction, and based on the obtained additive information of the Gaussian distribution signal, the additive formula of the Gaussian distribution signal for noise addition

Number

[0036] Here, ε is a Gaussian distribution signal, x is the pixel value of the sample image, β is the variance of the Gaussian distribution signal, and it is a coefficient with a value between [0.0, 1.0].

[0037] Next, the dedicated device realizes adding the Gaussian distribution signal to the sample image multiple times by repeating the input additive formula of the Gaussian distribution signal.

[0038] Specifically, taking x0 as the initial sample image, the dedicated device, based on the above additive formula of the Gaussian distribution signal, obtains the superposition formula

Number

[0039] In this way, the dedicated device repeatedly executes the addition formula of the Gaussian distribution signal multiple times based on the fitting instruction, and finally obtains the following superposition formulas.

Number

[0040] Here, ε for each repetition is a random number obtained by resampling according to the Standard Normal Distribution, and 0 < β1 < β2 < β3 …… < β N-3 <β N-2 <β N-1 <β N < 1.

[0041] If α = 1 - β, then

Number

[0042] After the dedicated device determines the calculation process from x0 to x N up to, it may construct a conversion relationship to directly convert x0 into the Nth noisy image x N .

[0043] Specifically, according to the relationship of x N-2 , x N-1、 x Nの relationship

Number

Number

[0044] Simplifying the above formula, we can obtain

Number

[0045] When two normal distributions are convolved, the probability density function after convolution remains a normal distribution. Therefore,

Mathematics

[0046] In the formula, ε N-1 and ε N are two independent random numbers, both of which satisfy the normal distribution. Therefore, based on the above formula, two samplings can be combined into one sampling and sampling can be performed using the overlapping probability distribution.

[0047]

Mathematics

Mathematics

Mathematics

Mathematics

[0048] With μ = 0 and σ = 1,

Mathematics

Mathematics

[0049] Similarly for ε N

Mathematics

[0050] When two normal distributions are superimposed, a new normal distribution N´(0, 1 - α N α N-1 ) can be obtained.

[0051] The dedicated device is equivalent to sampling the original two distributions by randomly sampling the new distribution N´, and the conversion from x N-2 to x N is completed in one sampling. That is,[[]] it is.

[0052] By this conversion method, the dedicated device can determine the conversion relationship N-3 from x N to x .

[0053] In this way, the dedicated device can finally determine the conversion relationship N from x0 to x .

[0054] Assuming that it is N the conversion relationship from x0 to x can be obtained by one sampling.

[0055] ​​​​​Based on the determined conversion relationship, the dedicated device can construct a second image generation model, and thereby can add noise to an image using the second image generation model. That is, the second image generation model can output an image with added noise of the original image after performing noise addition processing for the number of times corresponding to the number of Gaussian distribution signals superimposed on the input original image.

[0056] The dedicated device inputs the original image and the number of noise signals into the second image generation model pre-constructed by the above method, and causes the second image generation model to output an image with added noise of the original image after performing noise addition processing for the number of times corresponding to the number of noise signals. That is, the original image and the number N of noise signals are input into the second image generation model, and the image with added noise output by the second image generation model corresponds to the one obtained by performing noise addition on the original image N times (that is, adding noise to the original image with the first noise signal to obtain an original image with added noise by the first noise signal, and then using the second noise signal to add noise to the original image with added noise by the first noise signal, and so on), thereby greatly improving the efficiency of adding noise to the image. Then, a restored image can be obtained by performing noise removal on the image with added noise.

[0057] In step S103, the image with added noise is input into the first image generation model, and using the first image generation model, noise removal is performed on the image with added noise to obtain a restored image, and the image foreground features extracted from the restored image are determined.

[0058] After the dedicated device acquires the image with added noise output by the second image generation model, it inputs the image with added noise output by the second image generation model and the number value of the number of times the image with added noise is added with noise into the first image generation model, and using the first image generation model, noise removal is performed on the image with added noise to obtain a restored image, and then the image foreground features extracted from the restored image are determined.

[0059] In the process of removing noise from the post-noise addition image by the first image generation model, based on the number of times the post-noise addition image has been subjected to noise addition, for the original image corresponding to the post-noise addition image, when performing the noise addition process corresponding to the number of times, predict the superimposed noise signal used, and based on the predicted superimposed noise signal, predict the transition image before the nth noise addition process of the input post-noise addition image. The above step of predicting the transition image can be used for the post-noise addition image and all transition images of this post-noise addition image. That is, the first image generation model, until a restored image corresponding to the input image is obtained, based on the input image and the number of times k of noise addition corresponding to the image, for the original image corresponding to the input image, predict the superimposed noise signal used when performing the kth noise addition process, and then, based on the superimposed noise signal, predict the (k - 1)th transition image before the kth noise addition process, and based on the superimposed noise signal and the (k - 1)th transition image, the (k - 2)th transition image before the (k - 1)th noise addition process can be predicted. k is a positive integer not exceeding the number of times corresponding to the input image.

[0060] Regarding the removal of the Gaussian distribution signal, as follows, a conversion model may be constructed so that the first image generation model realizes the function of predicting the restored image corresponding to the post-noise addition image based on the number of times of noise addition corresponding to the post-noise addition image.

[0061] x t Let the post-noise addition image with the predicted number of times of noise addition being t be used, and in the reverse direction, obtain the restored image from x t The first image generation model constructs the conversion relationship from x at any time point to x t to x t-1 and determines the conversion relationship from x at any time point to x0 based on the conversion relationship from x at any time point to x t to x t-1 to x t to x0.

[0062] Here, the first image generation model constructs the conversion relationship from x at any point in time t to x t-1 . That is, the procedure for determining the conditional probability P(x t-1 │x t ) is as follows.

[0063] The conversion relationship P(x t │x t-1 ) represented by the conversion relationship with the superimposed noise signal represents P(x t-1 │x t ). This conversion relationship P(x t │x t-1 ) can be determined in the process of constructing the second image generation model.

[0064] x t to x t-1 is a random process, so according to Bayes' theorem

Mathematics

Mathematics

[0065] Here, since both P(x t ) and P(x t-1 ) represent the probabilities of obtaining them from x0,

Mathematics

[0066] Adding the condition of x0 to all indicates that x0 can actually be ignored under the same x0.

[0067] The dedicated device needs to determine a conversion model

Mathematics

[0068] P(x t │x t-1 ) represents the probability that x occurs when x t-1 occurs. t

[0069] In the process of constructing the second image generation model,

Number

Number

[0070] Here, ε t satisfies the distribution N(0,1), and a constant

Number

Number

Number

[0071] P(x t │x0) represents the probability that x occurs when x t occurs. Similarly,

Number

[0072] The dedicated device determines that the probability distribution of P(x t │x0) is

Number

[0073] P(x t-1 │x0) represents the probability that x occurs when x t-1 occurs. Similarly,

Number

[0074] The dedicated device determines that the probability distribution of P(x t-1 │x0) is

Number

[0075] Since the right side of the equation of the conversion model for removing the Gaussian distribution signal is all normal distributions, its parameters can be substituted into the form of the probability density function

Number

[0076] Here, x is the value of the random number, μ is the mean, and σ is the standard deviation.

[0077]

Number

[0078] According to the dedicated device, the above three probability density functions and the conversion model for removing the Gaussian distribution signal from the image after adding noise

Number

Number

[0079] Finally, in a form that conforms to the equation of the normal distribution

Number

[0080] Here, the dedicated device is x t when the condition of x t-1 probability density function and its distribution

Number

[0081] Next, it is necessary to remove the last term x0 of the distribution. The relationship between x0 and x t relationship

Number

Number

Number

[0082] The above x t when the condition of x t-1 Based on the probability distribution, the dedicated device uses the first image generation model to determine x t when the condition of adding the Gaussian distribution signal ε once to x t-1 can be determined, realizing the prediction of the transition image by the first image generation model, and thereby obtaining a restored image.

[0083] In step S104, the first image generation model is trained with the optimization goal of minimizing the difference between the image foreground feature corresponding to the original image and the image foreground feature extracted from the restored image.

[0084] After obtaining a restored image corresponding to the image with added noise using the first image generation model, the optimization goal is to minimize the difference between the image foreground features extracted from the original image corresponding to the image with added noise and the image foreground features extracted from the restored image, and the first image generation model is trained. Thereby, using the trained first image generation model, a restored image having image foreground features similar to the image foreground features of the original image is generated, and the restored image can be used to construct a training set for training other models that require image data for training.

[0085] The trained first image generation model can output a restored image corresponding to the input image with added noise. The image foreground features in the restored image have a high degree of similarity with the image foreground features in the corresponding original image, but other parts have certain differences, achieving the effect of being similar but not exactly the same. Therefore, the restored image can be used as different image data to construct a training set for training the model.

[0086] As can be seen from the above method, through the above model training method, an image generation model capable of obtaining a differentiated image can be obtained, and the image generation model is used to generate training samples used to construct a training set for training the model to be trained. The process is as follows.

[0087] FIG. 2 is a schematic diagram showing the flow of the service execution method provided by the present invention and includes the following steps.

[0088] In step S201, an initial image is obtained.

[0089] The execution entity of the service execution method provided by the present invention may be any of a terminal device such as a notebook computer or a desktop computer, a client installed on the terminal device, or a server. For the sake of convenience of explanation, the service execution method provided by the present invention will be described only using the example where the execution entity is a server.

[0090] Regarding the model required for service execution, it is necessary to obtain a certain number of training data for training the model. Here, the larger the number and the higher the quality of the training data, the higher the ability of the trained model. When the number of training data is not sufficient and the effect of the trained model cannot meet the needs of service execution, the service execution method provided by the present invention can additionally obtain a target image that can be used for constructing a training set from an initial image using the above-described training method of the image generation model.

[0091] Taking the artistic image generation service as an example, the required artistic image generation model (i.e., the designated model to be trained) requires a large amount of artistic image data as learning data. When the number of image data is insufficient, for example, when there are few images of a certain artistic genre, a small number of images of this artistic genre can be used as the initial image, and a target image for training the artistic image generation model can be additionally obtained.

[0092] The server obtains image data that can be used for training the model as the initial image. The initial image may be an image after noise addition or an image without noise addition. Here, for the image after noise addition, it is necessary to carry the number of times of noise addition corresponding to the image after noise addition. Otherwise, it will affect the quality of the restored image output by the pre-trained image generation model. For the image without noise addition, noise addition may be performed using the second image generation model trained using the training method of the image generation model provided by the present invention.

[0093] In step S202, the initial image is input into a pre-trained image generation model to output a target image, and the image generation model is a model obtained by training using the above training method.

[0094] The server inputs the initial image into a pre-trained first image generation model to obtain a target image, and the target image can be used to construct a training set for training a predetermined specified model required for service execution. Here, the pre-trained image generation model is the first image generation model trained using the training method of the image generation model provided by the present invention, and a restored image that can be used for model training can be obtained by performing noise removal on the image after noise addition.

[0095] For example, before training an artistic image generation model (i.e., the specified model to be trained), the images in the training set are used as the initial images. After adding noise to the initial images, the initial images together with the number value of the number of times of noise addition are input into a pre-trained first image generation model to perform noise removal, so as to obtain an image whose image foreground features are similar to those of the initial image as additional image data, and the training set of the artistic image generation model can be expanded.

[0096] Note that by not limiting the number of times of noise addition to the initial image, it is also possible to obtain more target images. For example, for an image with noise added once, not only the target image restored once but also the target image restored twice can be obtained, and this is not particularly limited.

[0097] In step S203, a training set is constructed based on the initial image and the target image, a predetermined specified model is trained using the training set, and the trained specified model is used to execute the service.

[0098] After the server obtains a target image whose foreground image features are similar to those of the initial image, it constructs a training set based on the initial image and the target image, trains a predetermined specified model required for service execution, and executes the service using the trained specified model.

[0099] For example, when the server trains an artistic image generation model (i.e., the specified model to be trained), it constructs a training set for the artistic image generation model based on the initial image and the target image, and trains an artistic image generation model that meets the service requirements with a small number of initial image data, thereby improving the training efficiency of the artistic image generation model.

[0100] In the present invention, the initial image input to the first image generation model may be an image with noise added or an image without noise added. In any case, the first image generation model performs noise removal on the input image as an image with noise added, and obtains a target image whose foreground image features are similar to those of the initial image but have significant differences in other parts. For an initial image without noise added, for an initial image without noise added, the first image generation model performs noise removal by regarding a part of the image data included in the initial image as noise according to the noise removal logic learned in the training process.

[0101] As described above, the training method and service execution method of the image generation model of the present invention have been described. Based on the same concept, the present invention also provides corresponding devices, storage media, and electronic devices.

[0102] FIG. 3 is a schematic diagram showing the structure of a training device for an image generation model provided by the present invention. The device includes: an acquisition module 301 for acquiring an original image; a noise addition module 302 for performing a noise addition process on the original image to obtain a post-noise addition image; Input the image after noise addition and the number of times the image after noise addition has been noise-added into a first image generation model, and use the first image generation model to predict the superimposed noise signal used when converting the original image into the image after noise addition by performing the noise addition process of the number of times on the original image until a restored image is obtained. Based on the superimposed noise signal, predict the (k - 1)-th transition image before the k-th noise addition process, and based on the superimposed noise signal and the (k - 1)-th transition image, predict the (k - 2)-th transition image before the (k - 1)-th noise addition process, and an input module 303 for determining the image foreground features extracted from the restored image. A training module 304 for training the first image generation model with the optimization goal of minimizing the difference between the image foreground features corresponding to the original image and the image foreground features extracted from the restored image.

[0103] Optionally, the noise addition module 302 specifically Inputs the original image and the number of noise signals into a pre-constructed second image generation model, and is used to cause the second image generation model to output the image after noise addition of the original image after the noise addition process of the number of times corresponding to the number of noise signals.

[0104] Optionally, the noise addition module 302 specifically Obtains a sample image. Uses the N-th noise signal to add noise to the image after noise addition that has been noise-added with the (N - 1)-th noise signal to obtain the image after noise addition that has been noise-added with the N-th noise signal, where N is a positive integer greater than or equal to 1, and the image after noise addition that has been noise-added with the 0-th noise signal is the sample image. Based on the image after noise addition with the Nth noise signal, the image after noise addition with the (N - m)th noise signal, the Nth noise signal, and the (N - m + 1)th noise signal, determine the conversion relationship from the image after noise addition with the (N - m)th noise signal to the image after noise addition with the Nth noise signal, where m is a positive integer smaller than N. It is used to construct the second image generation model based on the conversion relationship.

[0105] FIG. 4 is a schematic diagram showing the structure of a service execution device provided by the present invention. The device includes an acquisition module 401 for acquiring an initial image, an input module for inputting the initial image into a pre-trained image generation model to output a target image, where the image generation model is a model obtained by training using the above training method, and the input module is 402, a training module 403 for constructing a training set based on the initial image and the target image, training a predetermined specified model using the training set, and executing a service using the trained specified model.

[0106] The present invention further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the training method of the image generation model provided in FIG. 1 above or the service execution method provided in FIG. 2 above is implemented.

[0107] Based on the training method of the image generation model shown in FIG. 1 and the service execution method shown in FIG. 2, embodiments of the present invention further provide a schematic diagram showing the structure of the electronic device shown in FIG. 5. As shown in FIG. 5, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, an internal memory, and a non-volatile memory. Of course, it may also include other hardware required for other operations. The processor reads the corresponding computer program from the non-volatile memory into the internal memory and executes it to implement the training method of the image generation model shown in FIG. 1 or the service execution method shown in FIG. 2.

[0108] Of course, in addition to being implemented by software, the present invention does not exclude other implementation manners such as logical devices and combinations of hardware and software. That is, the execution subject of the following processing process is not limited to each logical unit, and may be hardware or a logical device.

[0109] In the 1990s, improvements in certain technologies could be clearly distinguished between hardware improvements (such as improvements in circuit structures of diodes, transistors, switches, etc.) and software improvements (such as improvements in method flows). However, with the development of technology, many current improvements in method flows can now be regarded as direct improvements to hardware circuit structures. Designers mostly obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be categorically stated that improvements in method flows cannot be realized by hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit, and its logical function is determined by programming by the user of the device. Instead of chip manufacturers designing and manufacturing dedicated integrated circuit chips, designers program to "integrate" a digital system onto a single PLD.And currently, instead of building integrated circuit chips by hand, this programming is mostly realized using software called a "logic compiler", which is similar to the software compilers used when writing programs. To compile the previous original code, it is necessary to write in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL. There are many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. It will be obvious to those skilled in the art that by simply programming the method flow logically in some of the above hardware description languages and programming it into an integrated circuit, a hardware circuit that realizes the logical method flow can be easily obtained.

[0110] The controller may be implemented in any suitable manner. For example, the controller may be a microprocessor or a processor, and a computer-readable storage medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor. It may also adopt the form of logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of the controller include, but are not limited to, microcontrollers such as ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller may further be implemented as part of the control logic of the memory. In addition to implementing the controller with pure computer-readable program code, it will be apparent to those skilled in the art that by logically programming method steps, the same functions can be made to be executed by the controller in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller may be regarded as a hardware component, and the devices included therein for realizing various functions may also be regarded as the structure within the hardware component. Or, further, the devices for realizing various functions may be regarded as software modules for realizing the method, or may be regarded as the structure within the hardware component.

[0111] The system, apparatus, module, or unit described in the above embodiments may specifically be implemented by a computer chip, an entity, or a product having some functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a mobile phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet, a wearable device, or any combination of several of these devices.

[0112] For the sake of convenience in description, when the above apparatus is described, it is divided into various units according to functions and described respectively. Of course, when implementing the present invention, it is also possible to implement the functions of each unit with the same or a plurality of software and / or hardware.

[0113] As those skilled in the art will understand, the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may be in the form of an embodiment consisting only of hardware, an embodiment consisting only of software, or an embodiment combining software and hardware. Furthermore, the present invention may also be in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0114] The present invention will be described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that an apparatus for realizing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram is generated by instructions executed by the computer or other programmable data processing device's processor.

[0115] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that a product including an instruction apparatus for realizing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram is generated by the instructions stored in the computer-readable memory.

[0116] These computer program instructions may also be loaded onto a computer or other programmable data processing device, thereby generating a process implemented by the computer by executing a series of operation steps on the computer or other programmable device, and thereby providing steps for realizing the functions specified in one or more flows of the flowchart and / or one or more blocks within one or more blocks of the block diagram by instructions executed on the computer or other programmable device.

[0117] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0118] Memory may include forms such as volatile memory, random access memory (RAM), and / or non-volatile memory among computer-readable storage media, for example, read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable storage media.

[0119] Computer-readable storage media includes volatile and non-volatile media, removable and non-removable media, and can implement information storage by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (flash Memory), or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tape, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible from a computing device. According to the definitions herein, computer-readable storage media does not include transitory media, such as modulated data signals and carriers.

[0120] Also, the term "comprising", "containing", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that comprises a list of elements does not include only those elements but may also include other elements not expressly listed or inherent to such process, method, article, or device. Where there are no more limitations, the elements defined by the phrase "comprising one..." do not preclude the presence of additional identical elements in the process, method, article, or device that comprises the said element.

[0121] As will be appreciated by those skilled in the art, embodiments of the present invention may be provided as a method, system, or computer program product. Accordingly, the present invention may take the form of an embodiment consisting of only hardware, an embodiment consisting of only software, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product embodied in one or more computer-readable storage media (including but not limited to magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0122] The present invention may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The present invention may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including memory devices.

[0123] Each embodiment in the present invention is described in a progressive manner. For the same or similar parts between each embodiment, reference may be made to each other, and the differences from other embodiments are emphasized in each embodiment. In particular, for the system embodiment, since it is basically similar to the method embodiment, it will be briefly described, and the relevant parts may refer to the description of a part of the method embodiment.

[0124] The above are only embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various modifications and changes can be made to the present invention. Any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and principle of the present invention should be included in the scope of the claims of the present invention.

Claims

1. acquiring an original image; performing a noise addition process on the original image to obtain a noise-added image; a step of inputting the noise-added image and a number of times the noise-added image is noise-added to a first image generation model, predicting a superimposed noise signal used when performing a noise addition process of the number of times on the original image to convert it into the noise-added image using the first image generation model until a restored image is obtained, predicting a k-1th transition image before performing the k-th noise addition process based on the superimposed noise signal, predicting a k-2th transition image before performing the k-1th noise addition process based on the superimposed noise signal and the k-1th transition image, and determining an image foreground feature extracted from the restored image, where k is a positive integer not exceeding the number of times, the image foreground feature is for representing a morphological feature of a target object in an image, and the image foreground feature does not include detailed physical features for representing the target object; training the first image generation model with an optimization goal of minimizing a difference between image foreground features corresponding to the original image and image foreground features extracted from the restored image; A method for training an image generation model, comprising:

2. The step of performing a noise addition process on the original image to obtain a noise-added image includes: inputting the original image and the number of noise signals into a second image generation model constructed in advance, and causing the second image generation model to output a noise-added image of the original image having been subjected to a noise addition process a number of times corresponding to the number of noise signals; 2. The method of claim 1 .

3. The construction of the second image generation model includes: obtaining a sample image; a step of adding noise to a noise-added image to which noise has been added with an (N-1)th noise signal using an Nth noise signal to obtain a noise-added image to which noise has been added with the Nth noise signal, where N is a positive integer of 1 or more, and the noise-added image to which noise has been added with a 0th noise signal is the sample image; a step of determining a conversion relationship from the noise-added image to which noise has been added with the N-mth noise signal to the noise-added image to which noise has been added with the Nth noise signal, based on the noise-added image to which noise has been added with the N-mth noise signal, the Nth noise signal, and the N-m+1th noise signal, where m is a positive integer smaller than N; and constructing the second image generation model based on the transformation relationship.

3. The method of claim 2 .

4. acquiring an initial image; a step of inputting the initial image into a pre-trained image generation model and outputting a target image, the image generation model being a model obtained by training using the training method according to any one of claims 1 to 3; constructing a training set based on the initial image and the target image, training a predetermined designated model using the training set, and executing a service using the trained designated model. A service execution method comprising:

5. an acquisition module for acquiring an original image; a noise addition module for performing a noise addition process on the original image to obtain a noise-added image; an input module for inputting the noise-added image and a number of times the noise-added image is noise-added to a first image generation model, predicting a superimposed noise signal used when performing a noise-adding process of the number of times on the original image to convert it into the noise-added image using the first image generation model until a restored image is obtained, predicting a k-1th transition image before performing a k-th noise-adding process based on the superimposed noise signal, predicting a k-2th transition image before performing a k-1th noise-adding process based on the superimposed noise signal and the k-1th transition image, and determining image foreground features extracted from the restored image, where k is a positive integer not exceeding the number of times, the image foreground features are for representing morphological features of a target object in an image, and the image foreground features do not include detailed physical features for representing the target object; a training module for training the first image generation model with an optimization objective of minimizing a difference between image foreground features corresponding to the original image and image foreground features extracted from the restored image; An apparatus for training an image generation model, comprising:

6. The noise addition module specifically includes: The original image and the number of noise signals are input to a second image generation model constructed in advance, and the second image generation model is used to output a noise-added image of the original image having been subjected to noise addition processing a number of times corresponding to the number of noise signals.

6. The apparatus according to claim 5 .

7. an acquisition module for acquiring an initial image; an input module for inputting the initial image into a pre-trained image generation model and outputting a target image, the image generation model being a model obtained by training using the training method according to any one of claims 1 to 3; a training module for constructing a training set based on the initial image and the target image, training a predetermined designated model using the training set, and executing a service using the trained designated model; A service execution device comprising:

8. A computer-readable storage medium storing a computer program, the computer program being executed by a processor to perform the method according to any one of claims 1 to 3. A computer-readable storage medium comprising:

9. An electronic device comprising a processor and a computer program stored in a memory and operable on the processor, the electronic device performing the method according to any one of claims 1 to 3 when the processor executes the computer program.

1. An electronic device comprising: