Image generation method and device, electronic equipment and storage medium

By using a generator with convolutional kernels of the same size but different convolutional elements and a self-attention module, as well as a discriminator based on cosine distance and contrastive learning, the problem of image generation quality and speed of generative adversarial networks with a small number of samples is solved, and efficient and accurate generation in skin lesion image generation is achieved.

CN116977653BActive Publication Date: 2025-11-04CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210785925.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-04
Publication Date
2025-11-04
Estimated Expiration
2042-07-04

AI Technical Summary

Technical Problem

Existing generative adversarial networks suffer from poor image quality, long training times, and the need for a large number of samples, especially in the field of medical imaging, such as the generation of skin lesion images, where the number of samples is limited and it is difficult to meet medical requirements.

Method used

The randomly generated image is convolved using N convolution kernels of the same size but different convolution elements to generate a weight map and a second feature map. Combined with a self-attention module and a DPN network, the clarity and realism of the generated image are improved. The discriminator adjusts the generator parameters through cosine distance and contrastive learning to improve the accuracy and efficiency of image generation.

Benefits of technology

Generating skin lesion images that meet medical requirements with a small number of samples improves the quality of image generation and training speed. The images generated by the generator are more accurate and realistic, and the discriminator is more efficient in image differentiation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116977653B_ABST
    Figure CN116977653B_ABST
Patent Text Reader

Abstract

An image generation method of a generative adversarial network performed by a generator provided by the embodiments of the present disclosure comprises: performing convolution processing on a randomly generated first image by using N convolution kernels to obtain N first feature maps; any two of the N convolution kernels have the same size and different convolution elements; generating a weight map according to the first feature map 1 to the first feature map N-1; wherein the size of the weight map is equal to the size of the first feature map; obtaining a second feature map based on the N first feature maps and the weight map; and generating a second image based on the second feature map. Here, the generator can more clearly and comprehensively obtain and filter various features of the image by performing convolution processing and generating a weight map by using N convolution kernels with different convolution elements, thereby obtaining a second image with more significant features and higher correlation, so that the generator can better and faster generate a large number of clear, accurate and real images.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of deep learning, and in particular, to an image generation method and device, an electronic device, and a storage medium. BACKGROUND

[0002] With the rapid development of deep learning technology, an image can be generated by a generative adversarial network model and a network model of the same family as the generative adversarial network. A generative adversarial network (GAN) can be composed of a generator and a discriminator. The generator and the discriminator update each other by iterative training until the generative adversarial network meets a convergence condition. An image can be generated according to the trained generative adversarial network.

[0003] Existing generative adversarial networks often have problems such as poor image quality, the need for a large number of samples for training, long training time, and slow training speed of the generative adversarial network model. SUMMARY

[0004] Therefore, the present disclosure provides an image generation method, an electronic device, and a storage medium.

[0005] According to a first aspect of the present disclosure, an image generation method is provided, which is performed by a generator and includes: performing convolutional processing on a randomly generated first image using N convolutional kernels to obtain N first feature maps; any two of the N convolutional kernels have the same size and different convolutional elements; generating a weight map according to the first feature maps from the first to the N-1th; the size of the weight map is equal to that of the first feature map; obtaining a second feature map based on the N first feature maps and the weight map; and generating a second image based on the second feature map.

[0006] In one embodiment, the generating a weight map according to the first feature maps from the first to the N-1th includes: adjusting the positions of elements of at least part of the first feature maps from the first to the N-1th to obtain N-1 third feature maps; and point multiplying the N-1 third feature maps to obtain the weight map.

[0007] In one embodiment, the adjusting the positions of elements of at least part of the first feature maps from the first to the N-1th to obtain N-1 third feature maps includes: performing transpose processing on at least part of the first feature maps from the first to the N-1th to obtain the N-1 third feature maps.

[0008] In one embodiment, the method further comprises: performing convolution processing on the input parameter to obtain a convolution processing result; performing residual connection on the convolution processing result and the input parameter to obtain a residual feature map; and performing dense connection on the residual feature map, the input parameter, and a parameter obtained by the convolution processing to obtain the first image.

[0009] In a second aspect, the embodiments of the present disclosure provide an image generation method, executed by a discriminator, comprising:

[0010] According to the convolution and pooling processing on the second image and the sample image, a fourth feature map is obtained; wherein the second image is generated by a generator; a feature of the fourth feature map is extracted to obtain a first feature vector; a first cosine distance between the second image and the sample image with the same image content is determined based on the first feature vector; a loss value is determined according to the first cosine distance; and the parameters of the generator are adjusted according to the loss value.

[0011] In one embodiment, the method further comprises: performing image transformation processing on the second image to obtain a third image; performing image transformation processing on the sample image to obtain a fourth image; extracting features of the third image and the fourth image to obtain a second feature vector; determining a second cosine distance between the third image and the fourth image with the same image content according to the second feature vector; and determining the loss value according to the first cosine distance and the second cosine distance.

[0012] In one embodiment, the image transformation processing comprises at least one of the following: image offset; image cropping; image rotation; image filtering; and noise increase.

[0013] In one embodiment, the method further comprises: obtaining a random value; and performing image transformation processing on the second image to obtain a third image comprises: when the random value is greater than a probability value of the image transformation processing, performing image transformation processing on the second image to obtain a third image.

[0014] In a third aspect, the embodiments of the present disclosure provide an image generation device, comprising: a processing module configured to perform convolution processing on a first image randomly generated by using N convolution kernels to obtain N first feature maps; any two convolution kernels in the N convolution kernels have the same size and different convolution elements; a generation module configured to generate a weight map according to the first first feature map to the N-1 first feature map; wherein the size of the weight map is equal to the size of the first feature map; an obtaining module configured to obtain a second feature map based on the N first feature maps and the weight map; and the generation module is further configured to generate a second image based on the second feature map.

[0015] In a fourth aspect, an image generation apparatus is provided, and the apparatus comprises:

[0016] The processing module is configured to obtain a fourth feature map by performing convolution and pooling processing on the second image and the sample image; the extraction module is configured to extract features of the fourth feature map to obtain a first feature vector; the determination module is configured to determine a cosine distance between the second image and the sample image with the same image content based on the first feature vector; determine a loss value based on the first cosine distance; and the adjustment module is configured to adjust parameters of the generator based on the loss value.

[0017] In a fifth aspect, an electronic device is provided, and the electronic device comprises a processor and a memory for storing a computer program capable of running on the processor; when the processor runs the computer program, the steps of the method in the foregoing one or more technical solutions are performed.

[0018] In a sixth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores computer executable instructions; after the computer executable instructions are executed by a processor, the method in the foregoing one or more technical solutions can be implemented.

[0019] The image generation method provided by the embodiments of the present disclosure, the generator can utilize N different convolution kernels with the same size to perform convolution processing on the randomly generated first image, obtain a first feature map, and then generate a weight map according to the first feature map, so as to obtain a second feature map according to the first feature map and the weight map and generate a second image. In this way, the generator can obtain various features of the image more clearly and comprehensively through the convolution processing of the N different convolution kernels and the weight map, so as to screen out the most significant feature with the largest correlation in the image by combining various features, and make the obtained second image more accurate and real.

[0020] The discriminator can more accurately distinguish the features of the second image generated by the generator and the sample image according to the loss value determined by the cosine distance of the feature vector obtained in the contrast learning, compared with ordinary classification, and improve the efficiency of judging the image. The generative adversarial network including the discriminator and the generator trained according to the embodiments of the present disclosure can generate a large number of accurate and clear skin lesion images that meet the medical requirements in the case of having a small number of sample images of skin lesions. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 The flowchart of the first image generation method provided by the embodiments of the present disclosure is shown.

[0022] Figure 2A flowchart of a second image generation method provided by an embodiment of the present disclosure.

[0023] Figure 3 A flowchart of a third image generation method provided by an embodiment of the present disclosure.

[0024] Figure 4 A flowchart of a fourth image generation method provided by an embodiment of the present disclosure.

[0025] Figure 5 A schematic diagram of a skin lesion image provided by an embodiment of the present disclosure.

[0026] Figure 6 A schematic diagram of a grouped convolution provided by an embodiment of the present disclosure.

[0027] Figure 7 A flowchart of obtaining a first image provided by an embodiment of the present disclosure.

[0028] Figure 8 A structural diagram of producing a generative adversarial network provided by an embodiment of the present disclosure.

[0029] Figure 9 A flowchart of obtaining a second image provided by an embodiment of the present disclosure.

[0030] Figure 10 A schematic diagram of a first image generation device provided by an embodiment of the present disclosure.

[0031] Figure 11 A schematic diagram of a second image generation device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] In order to make the purposes, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the drawings, and the described embodiments should not be regarded as limitations of the present disclosure. All other embodiments obtained by those of ordinary skill in the art without creative labor under the premise that there is no conflict, belong to the scope of protection of the present disclosure.

[0033] In the following description, “some embodiments” are described, which describe a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0034] In the following description, the terms "first", "second", "third" are merely used to distinguish similar objects, and do not represent a specific order or sequence of the objects. Understandably, the "first", "second", "third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used in this disclosure is for the purpose of describing embodiments of this disclosure only and is not intended to be limiting of this disclosure.

[0036] In order to better understand the embodiments of the disclosure, the following is described through some scene embodiments:

[0037] With the rapid development of medical technology, in the development of medical artificial intelligence and medical research teaching and many other scenes, due to the privacy ethics and many other problems contained in medical images, it is often difficult to obtain medical images that meet medical requirements. In the field of dermatology, lesion images are more likely to involve sensitive medical issues, and it is more difficult to obtain lesion images.

[0038] Regarding psoriasis in the field of dermatology, psoriasis lesion images that meet medical requirements have many characteristics such as clear morphology, obvious texture features, clear boundary range, and can be combined with many lesions. Although the number of psoriasis patients is large, due to many factors such as sensitive medical information, the number of available psoriasis sample images is small.

[0039] As shown in Figure 1 The embodiments of the disclosure provide an image generation method, executed by a generator, the method comprising:

[0040] Step S101: performing convolution processing on a randomly generated first image using N convolution kernels to obtain N first feature maps; any two convolution kernels in the N convolution kernels have the same size and different convolution elements;

[0041] Step S102: generating a weight map according to the first feature map to the N-1th first feature map; wherein the size of the weight map is equal to the size of the first feature map;

[0042] Step S103: obtaining a second feature map based on the N first feature maps and the weight map;

[0043] Step S104: generating a second image based on the second feature map.

[0044] The generator can be a deep learning model, wherein the deep learning model can include one or more convolutional layers, and the second image with a specific effect can be generated according to the first image generated randomly.

[0045] For example, the generator can be a generator included in a generative adversarial network (GAN).

[0046] For example, the generator can be a generator included in a generative adversarial model built by a dual path network (DPN) as a backbone network architecture. The backbone network can extract image features by convolution and the like. In an embodiment, the backbone network of the generative adversarial network can include a residual network ResNet or a dense network DenseNet. The residual network adds input parameters to the output of convolution, and can reuse the features extracted in the previous level through a residual side branch. The dense network splices the output to the input of each subsequent layer, and can explore new features through a dense connection path.

[0047] In an embodiment of the present disclosure, the backbone networks of the generator and the discriminator of the generative adversarial network can be a DPN network. The DPN network combines the connection methods of residual connection and dense connection. Compared with the residual connection which cannot explore new features and the dense connection which can explore new features but has high redundancy, the DPN network combines the advantages of the two structures, can explore new features without high redundancy, and obtains more accurate image features and faster speed.

[0048] In an embodiment, the generative adversarial network can be used to generate a skin lesion image. For example, the skin lesion image can include a psoriasis lesion image, which can be as shown in Figure 5

[0049] In an embodiment, the generator can include a self-attention mechanism module, and the steps S101 to S104 can be performed by the self-attention module in the generator.

[0050] In an embodiment, the convolution kernel can include a convolution kernel with different convolution elements such as a low-pass filter convolution kernel, a high-pass filter convolution kernel, and / or a differential edge detection convolution kernel. The high-pass filter convolution kernel can extract parts of the image with relatively sharp changes, and can be used for sharpening the image and / or enhancing the edges of objects in the image. The low-pass filter convolution kernel can extract parts of the image with relatively gentle changes, and can be used for blurring the image, smoothing the image, and / or eliminating noise points. The differential edge detection can be used to detect and extract contours and edges in the image. ​

[0051] In some embodiments, the self-attention module in the generator tends to use a convolution kernel of the same convolution element multiple times to extract predetermined types of features in the image. In the embodiments of the present disclosure, convolution according to N convolution kernels with different convolution elements can obtain a plurality of first feature maps with different types of features according to a plurality of different types of features in the first image, which can improve the clarity and authenticity of the image generated by the generator. The second feature map obtained according to the weight map and the first feature map generated according to a plurality of different types of features can include the most obvious feature in the global feature obtained by filtering a plurality of different types of features and the feature with the largest correlation with the first image, which can generate a second image that is more clear and accurate, and can also improve the speed and quality of training the generative adversarial network model.

[0052] In one embodiment, the size of the N convolution kernels with the same size in step S101 can be 1x1, 3x1 or 8x1, etc. For example, the size of the input feature map can be c x w x h, the number of convolution kernels can be 3, and the size of the convolution kernel can be 8x1. Three 8x1 convolution kernels are used to perform convolution processing on the randomly generated first image, and three first feature maps with a size of (8x c) x w x h are obtained.

[0053] In one embodiment, the randomly generated first image can include a randomly generated first image obtained from a random vector input generator. The random vector can pass through the full connection layer, the dimension change layer and the convolution layer in the generator to obtain the randomly generated first image.

[0054] In one embodiment, step S102 can include point multiplication calculation processing between N-1 first feature maps; and classification and normalization processing of the feature maps after the calculation processing to obtain a weight map.

[0055] In one embodiment, the method can further include reshape processing of the feature map, wherein the reshape processing can include dimension increasing processing and dimension decreasing processing. The reshape processing can make the operation of the feature map more convenient and faster.

[0056] In one embodiment, step S102 can include dimension decreasing processing of the first feature map to change the three-dimensional first feature map into a two-dimensional first feature map, and generating a weight map according to the first feature map and the N-1 first feature maps.

[0057] In one embodiment, step S103 can include point multiplication processing of the N first feature maps and the weight map to obtain a second feature map.

[0058] In an embodiment, the step S104 can include performing a dimension transformation on the second feature map to obtain a second image. The dimension transformation can include performing an up-sampling on the second feature map to obtain a three-dimensional second feature map, and performing a down-sampling convolution on the three-dimensional second feature map to obtain the second image with the same size as the first image.

[0059] In some embodiments, as shown in FIG. 2, the generating the weight map according to the first feature map to the N-1th feature map can include: Figure 2

[0060] The step S201 includes adjusting positions of elements in at least part of the first feature map to obtain N-1 third feature maps.

[0061] The step S202 includes point-multiplying the N-1 third feature maps to obtain the weight map.

[0062] In an embodiment, the positions of the elements can include positions of pixels in the feature map and / or positions of parameters in the feature matrix.

[0063] The point-multiplying of the generator after adjusting the positions of the elements in the first feature map can calculate the relationship between any two pixel points in the image, and can learn the overall features, texture features, and structure and geometric features of the lesion image. The texture features can include colors and textures of the lesion area, and the structure and geometric features can include sizes and shapes of the lesion area.

[0064] Thus, compared with the generator without adjusting the positions of the elements in the feature map, the generator can generate a lesion image with clearer and more accurate texture according to the learned texture features, and can better learn the structure and geometric shape features of the lesion area and the features of the connection part between the lesion area and the normal area according to the learned structure and geometric features, thereby avoiding the problem that the generated image has a disjointed lesion area and normal skin area, and generating a more accurate and clear lesion image that meets the medical requirements.

[0065] In an embodiment, the step S202 can include point-multiplying the N-1 third feature maps two by two, and then point-multiplying the feature maps obtained by the point-multiplying to obtain a weight map.

[0066] ​In an embodiment, the step S202 can include multiplying N-1 third feature maps, classifying and normalizing the multiplied feature maps to obtain the weight map. The classification processing can be performed by using an activation function such as softmax, and the normalization processing can be performed by using a normalization layer such as Batch Normalization, Weight Normalization, Layer Normalization, and / or Group Normalization.

[0067] In an embodiment, the normalization layer can include Batch Normalization, which can normalize data by calculating the mean and variance of each batch of data, thereby accelerating the convergence speed of the network model, preventing gradient explosion, and preventing overfitting.

[0068] In some embodiments, the step of adjusting the positions of elements in at least part of the first to N-1th first feature maps to obtain N-1 third feature maps can include:

[0069] In some embodiments, the step of adjusting the positions of elements in at least part of the first to N-1th first feature maps to obtain N-1 third feature maps can include:

[0070] In an embodiment, the transposition processing can include matrix transposition processing of a matrix of feature parameters of the first feature map and / or transposed convolution processing of the first feature map.

[0071] In an embodiment, the step of adjusting the positions of elements in at least part of the first to N-1th first feature maps to obtain N-1 third feature maps can include:

[0072] Thus, performing matrix transpose and dot product on the feature image allows for rapid adjustment and calculation of the relationship between any two pixels in the image, enabling quick acquisition of global features. Generators that do not perform matrix transpose and dot product on the feature image have limited convolution kernel size, meaning each convolution operation can only cover a small neighborhood around the pixel, failing to cover distant features. Therefore, compared to generators that do not perform matrix transpose and dot product on the feature image, the generator in this embodiment can utilize information from distant regions within the image, allowing each location to incorporate information from similar or related regions, resulting in a more consistent generated image.

[0073] In some embodiments, the method further includes:

[0074] Step S301: Perform convolution processing on the input parameters to obtain the convolution processing result;

[0075] Step S302: Perform residual concatenation between the convolution processing result and the input parameters to obtain the residual feature map;

[0076] Step S303: Densely connect the residual feature map, the input parameters, and the parameters obtained by the convolution process to obtain the first image.

[0077] In one embodiment, step S301 may include inputting a random vector, passing the random vector through a fully connected layer and a dimension transformation layer to obtain a random feature map, and then using the random feature map as an input parameter for convolution processing.

[0078] In one embodiment, steps S301 to S303 can be processed by a generator based on a DPN network, wherein the DPN network may include grouped convolution processing of the input parameters. The grouped convolution processing may include: dividing the input dimension of the convolution into N groups, convolving each group using 1 / N of the number of input channels using convolution kernels, and then merging and concatenating them along the channel dimension. This reduces the number of parameters in the convolutional layer, improving the speed of convolution processing while maintaining its quality.

[0079] In one embodiment, such as Figure 6 As shown, the input dimension of the convolution is divided into 3 groups. Each group is convolved by a convolutional kernel with 1 / 3 of the number of input channels through the convolutional layer conv, and then the convolutions are merged and concatenated along the channel dimension.

[0080] In one embodiment, such as Figure 7As shown, the input parameter is x, G is a rectified linear unit (ReLU) activation function, F is a convolution layer function, and d is a hyperparameter used to determine the number of selected features. The input parameter is convolved according to a convolution weight layer and a convolution weight layer and a ReLU activation function to obtain a convolution processing result F(x); the first d feature maps F(x)[:d] of the convolution processing result F(x) are connected in residual learning with the input parameter x to obtain residual feature maps F(x)[d:]+x[d:]]; the residual feature maps, the dense connection part x[d:] of the input parameter x, and the dense connection part F(x)[d:] after convolution are densely connected and merged, and the connected image is passed through an activation function to obtain the first image. The formula of the output can be expressed as:

[0081] y=G([x[:d],F(x)[:d],F(x)[d:]+x[d:]])。

[0082] As shown in Figure 4 The image generation method provided by the embodiment of the disclosure is executed by a discriminator and includes the following steps.

[0083] Step S401: Fourth feature maps are obtained by performing convolution and pooling processing on a second image and a sample image; the second image is generated by a generator.

[0084] Step S402: A first feature vector is obtained by extracting features of the fourth feature maps.

[0085] Step S403: A first cosine distance between the second image and the sample image with the same image content is determined based on the first feature vector.

[0086] Step S404: A loss value is determined according to the first cosine distance.

[0087] Step S405: The parameters of the generator are adjusted according to the loss value.

[0088] In one embodiment, the step S401 can include obtaining feature maps of the second image and the sample image according to convolution processing, and performing pooling processing on the feature maps obtained by the convolution processing to obtain the fourth feature maps.

[0089] In one embodiment, the pooling processing can include maximum pooling processing, average pooling processing, and / or random pooling processing, etc.

[0090] Exemplarily, the step S401 can include: performing first convolution processing on the second image by 32 convolution kernels with a size of 3x3 to obtain a first convolution-processed feature map; performing second convolution processing on the first convolution-processed feature map by a DPN network to obtain a second convolution-processed feature map; and performing 5 times of maximum pooling processing on the second convolution-processed feature map to obtain a fourth feature map. Wherein, the size of the second image is 512x512, the size of the first convolution-processed feature map is 512x512x32, and the size of the second convolution-processed feature map is 16x16x32.

[0091] In one embodiment, the step S402 can include extracting features of the fourth feature map by a multi-layer perceptron based on a hidden layer of the DPN network in the discriminator to obtain a first feature vector.

[0092] In one embodiment, the first feature vector can include feature vectors of multiple dimensions.

[0093] Exemplarily, N second images and N sample images are input into the discriminator to obtain 2N fourth feature maps; a 128-dimensional first feature vector is obtained according to the 2N fourth feature maps; and a regularization processing is performed on the feature vector to convert it into a unit feature vector, wherein the unit feature vector can make the feature vector fall on a hypersphere with a radius of 1, and the feature vector can be represented as {z 1, z2,......,z 2N}.

[0094] In one embodiment, in the step S403, the second image with the same image content and the sample image can be distinguished by a category label.

[0095] In one embodiment, the cosine distance can represent the similarity by the size of the angle between the feature vectors, and the smaller the angle between two feature vectors, the more similar they are. Wherein, the Euclidean distance representing the distance between vectors cannot be kept between 0-1, and the cosine distance can calculate the similarity faster and better represent the similarity between feature vectors compared with the Euclidean distance. The discriminator can more accurately determine the similarity between the second image with the same image content and the sample image by determining the first cosine distance between them based on the first feature vector, and the generator can generate more accurate and real images after adjusting the parameters according to the loss value determined by the first cosine distance.

[0096] In an embodiment, the discriminator can be a discriminator trained by contrastive learning. The contrastive learning can pull the cosine distance between the feature vectors of the second image and the sample image with the same image content closer, and pull the cosine distance between the feature vectors of the second images of different classes farther, so that the generated second image is closer to the sample image in the iterative training.

[0097] In this way, compared with the discriminator based on the ordinary classification network, the discriminator based on the contrastive learning can clearly and intuitively distinguish the clustering of feature vectors of different types on the similar hypersphere, and can more clearly determine the distance and difference between the features of the lesion images. Compared with the ordinary classification network which cannot accurately distinguish and judge the lesion image sample image and the second image generated by the generator without explicit distinguishing features, the discriminator based on the contrastive learning can better distinguish the difference between the lesion sample image and the second image generated by the generator, and can more accurately judge the true and false of the second image generated by the generator, so as to make the image generated by the generator more real and clear, and also can improve the robustness of the generative adversarial network model and avoid the collapse of the generative adversarial network model.

[0098] In an embodiment, the step S404 can include calculating a loss function according to the first cosine distance, and determining a loss value according to the loss function. In an embodiment, the discriminator inputs N real images and N generated images per batch, and for 2N-1 input images other than the input image i, there are N-1 input images of the same class. Assuming that any one input image i, the class label of the input image i is Assuming that there is another input image j, if the input image i and the input image j are input images of the same class, the feature vectors of the input image i and the input image j are made closer, and if the input image i and the input image j are input images of different classes, the feature vectors of the input image i and the input image j are made farther. The formula of the loss function can include:

[0099]

[0100]

[0101] The loss function can represent that for any input image i, the sum of the cosine distances between the feature vectors of all other input images of the same class as the input image i and the feature vector of the input image i is made larger, and the sum of the cosine distances between the feature vectors of all other input images of different classes from the input image i and the feature vector of the input image i is made smaller.

[0102] In this way, the loss function can maintain a relatively stable average loss error under different degrees of image damage, and can completely suppress the image damage caused by the enhancement of a small number of sample images.

[0103] In one embodiment, the parameters of the generator can be adjusted by the loss value obtained by the discriminator until the generative adversarial network meets the convergence condition.

[0104] In one embodiment, the fourth feature map can be discriminated by the full connection layer in the discriminator to obtain a discrimination result.

[0105] In one embodiment, the initial discrimination result of the discriminator can include judging the generated image as false 0 and the sample image as true 1, and the parameters of the generator are adjusted by the loss value obtained by the discriminator to perform iterative training until the convergence condition is that the discriminator cannot distinguish between the generated image and the sample image, that is, the discrimination result for the generated image and the sample image can include equal to or close to 0.5.

[0106] In some embodiments, the method further comprises:

[0107] performing image transformation processing on the second image to obtain a third image;

[0108] performing image transformation processing on the sample image to obtain a fourth image;

[0109] extracting features of the third image and the fourth image to obtain a second feature vector;

[0110] determining a second cosine distance between the third image and the fourth image with the same image content according to the second feature vector;

[0111] The determining a loss value according to the first cosine distance comprises:

[0112] determining a loss value according to the first cosine distance and the second cosine distance.

[0113] In one embodiment, it can include adding different categories of labels to the third image and the fourth image respectively, and determining the second cosine distance between the third image and the fourth image with the same image content according to the second feature vector and the label.

[0114] In one embodiment, the loss value can be determined by calculating a loss function through the first cosine distance and the second cosine distance.

[0115] In an embodiment, the third image can include an image processed by image transformation from the second image and the second image; and the fourth image can include an image processed by image transformation from the sample image and the sample image. The image transformation processing can increase the number of third images and fourth images, which can also be referred to as image data enhancement. In this way, the model with a small number of sample images of skin lesions can increase the amount of sample image data, and can prevent overfitting of the model when there are a small number of sample images of skin lesions.

[0116] In an embodiment, the number of sample images in the generative network adversarial model is often tens of thousands. In the embodiment of the present disclosure, the small number of sample images can include one to two thousand.

[0117] In an embodiment, the image data enhancement of the second image and the sample image can prevent the training direction of the generator from deviating due to the fact that the discriminator discriminates the enhanced part of the sample image as real sample information, compared to only enhancing the sample image. In this way, the generated image of the generator can be more realistic.

[0118] In an embodiment, the image transformation processing of the sample image to obtain the fourth image can include highlight removal processing. The highlight removal processing of the sample image can solve the problem that the psoriasis lesion images taken in places with sufficient light are prone to highlight interference, and can make the generative adversarial network model generate more accurate and clear psoriasis lesion images.

[0119] In an embodiment, the step of performing highlight removal processing on the sample image can include: separating a channel image according to different channels of the sample image; performing multiple filtering processes on the channel images respectively to obtain filtered images; dividing different region images by calculating the filtered images; performing highlight removal processing on the different region images to different degrees; and merging the images after highlight removal processing to obtain the sample image after highlight removal processing.

[0120] In an embodiment, the separation of the channel image according to different channels can include separation according to the R channel, the G channel and the B channel of the image. The filtering process can include filtering by using filters with different sizes and / or parameters.

[0121] In an embodiment, the step of dividing different region images by calculating the filtered images can include: calculating the average value of the filtered images; calculating the threshold value of the filtered images according to the maximum inter-class variance method; and confirming different regions and pixel masks corresponding to different regions according to the average value and the threshold value.

[0122] In one embodiment, the different region images can include a lesion region and a normal skin image. The step of performing highlight removal on the lesion region can include removing highlight specular reflection components within the lesion region, which can include calculating a maximum chroma value Amax and a minimum chroma value Amin for each pixel, which can be expressed as:

[0123] Ai(x) = Ci(x) / (R + G + B);

[0124] where Ci(x) is the {R, G, B} value of the pixel at x, Amax(x) = max(Ai(x)), and Amin = min(Ai(x));

[0125] where C(x) is defined as max(R(x), G(x), B(x)), i.e., the maximum pixel value among the three channels of the pixel at x, and G(x) is defined as the highlight component coefficient, Gi(x) is the highlight component coefficient of the pixel at x in each channel, and G(x) = max(Gi(x)), which can be expressed as:

[0126] Gi(x) = (Ai(x) - Amin(x)) / (1 - 3 * Amin(x));

[0127] The formula for the highlight component GV(x) of the pixel at x can be expressed as:

[0128] GV(x) = (G(x) * (R(x) + G(x) + B(x)) - C(x)) / (3 * G(x) - 1);

[0129] For the pixel at x in the lesion region, the pixel values of the three channels after highlight removal can include the value of Ci(x) - GV(x). Combining the three channels can obtain the lesion region image after highlight removal.

[0130] The step of performing highlight removal on the normal skin region can include defining a diffuse reflection coefficient B(x), a diffuse reflection coefficient Bi(x) of the pixel at x in each channel, and the related formula can be:

[0131] B(x) = max(Bi(x));

[0132] Bi(x) = 1 - (Amax(x) - Ai(x)) / (3 * Amax(x) - 1);

[0133] The formula for the diffuse reflection component BV(x) of the pixel at x in the normal skin region can include:

[0134] BV(x) = (max(B(x) * (R(x) + G(x) + B(x)), Ci(x)) / (3 * B(x) - 1).

[0135] The pixel values of the normal skin region corresponding to the three channels can be expressed by the diffuse reflection component values. The three channels can be merged to obtain a normal skin region image after the highlight removal processing.

[0136] In some embodiments, the image transformation processing includes at least one of image shifting, image cropping, image rotation, image filtering, and / or noise increasing.

[0137] In one embodiment, the image shifting can include shifting the positions of pixels in the image in different directions. For example, the directions of the image shifting can include eight directions, wherein the eight directions can include four directions plus four directions with a 45-degree angle between them.

[0138] In one embodiment, the image cropping can include cropping the edges of the image, and the direction and length of the cropping can be set. For example, the direction of the cropping can include cropping the edges of the image in four directions, and the length of the cropping can include a random value within the range of the shortest side of the image.

[0139] In one embodiment, the image rotation can include rotating the image according to the central axis of the image. For example, the direction of the image rotation can include rotating to the left or rotating to the right.

[0140] In one embodiment, the image filtering can include filtering the image.

[0141] In one embodiment, the image filtering can include band-stop filtering. For example, the steps of the band-stop filtering can include:

[0142] performing Fourier transform on the image; setting the size of a predetermined band-stop filtering annular region; performing band-stop filtering on the Fourier-transformed image according to the size of the predetermined band-stop filtering annular region; and performing inverse Fourier transform on the band-stop filtered image to obtain a filtered image.

[0143] For example, the size of the band-stop annular region can include setting the radius to be a predetermined multiple of the side length of the image, such as 0.23 times the side length of the image to 0.25 times the side length of the image.

[0144] For example, the band-stop filtering processing can preserve the texture features and color information of the lesion region of the lesion image when certain state changes are made to the normal skin region of the lesion image. Therefore, the band-stop filtering processing can enhance the image data while preserving the effective features of the image, thereby improving the efficiency of generating the image.

[0145] In one embodiment, the noise increase can include adding a given type of noise to the image, wherein the noise-increased image can better preserve edge profile features.

[0146] In some embodiments, the method further comprises:

[0147] obtaining a random value;

[0148] The image transformation processing on the second image to obtain a third image comprises:

[0149] When the random value is greater than the probability value of the image transformation processing, the image transformation processing is performed on the second image to obtain a third image.

[0150] In one embodiment, the probability value of the image transformation processing can be a predetermined probability value of the image transformation processing artificially set in the discriminator.

[0151] In one embodiment, the method can include setting a predetermined offset occurrence probability value and / or an offset distance, etc. in the image offset processing; when the random value is greater than the offset occurrence probability value, the image offset processing is performed on the second image to obtain a third image.

[0152] For example, the predetermined image offset occurrence probability value can be 0.25, and when the obtained random value is greater than 0.25, the image offset processing is performed on the image, and when the obtained random value is not greater than 0.25, the image offset processing is not performed on the image.

[0153] The offset distance can include a random value within the range of the shortest side of the entire image. For example, the range of the shortest side of the entire image can be (0, 0.1), and a random value within the range of 0-0.1 is obtained as the offset distance.

[0154] In one embodiment, the method can include setting a predetermined probability value of image rotation occurrence and a rotation angle, etc. in the image rotation processing. The size of the rotation angle can include a random angle value within a predetermined rotation angle size range. For example, the probability value of image rotation occurrence can be 0.25, and the predetermined rotation angle size range can be (0, 15 degrees), and a random value within the range of 0-15 degrees is obtained as the size of the rotation angle.

[0155] In one embodiment, the probability value of the image processing can also include a probability value of image cropping occurrence and / or a probability value of image filtering occurrence.

[0156] In an embodiment, the probability value of the image transformation process, i.e., image data enhancement, can be set to 0.25. Exemplarily, the probability value of the image offset process, the probability value of the image rotation process, the probability value of the image cropping process, and / or the probability value of the image filtering process can each be set to 0.25. When the probability value of the image data enhancement is 0.25, the probability is relatively low, and the low-probability image data enhancement can increase the data quantity without changing the key features of the lesions in the lesion area of the lesion image, such as the morphology, texture, and normal skin area interface morphology of the lesions, and can add details of the sample image to solve the problem of insufficient quantity in generating a large number of real and clear images from a small number of sample images.

[0157] In an embodiment, the generative adversarial network model can be trained by a hardware image processor (GPU, graphics processing unit). Exemplarily, the hardware can be a Nvidia 16GB image processor.

[0158] In an embodiment, the generative adversarial network model can include a fixed-point floating-point quantization process, which can include converting parameters to 8 bits, reducing the model volume, and improving the operation efficiency. The fixed-point floating-point quantization process can include processing by the Ristretto tool.

[0159] In an embodiment, as shown in Figure 8 , a random vector input generator obtains a generated image, the original image, i.e., a sample image, is image-enhanced with the generated image, and the enhanced image is input into a discriminator. Through iterative training, the discriminator cannot distinguish between the generated image and the sample image, i.e., the probability of the result of the judgment output is 0.5 for both the generated image and the sample image.

[0160] In an embodiment, the step of obtaining a second image by the generator can be as shown in Figure 9 .

[0161] Step 1: Convolve the first image with three 8x1 convolution kernels to obtain three first feature maps. The three convolution kernels can be defined as q convolution, k convolution, and v convolution, respectively, and the three obtained feature maps can be q1, k1, and v1. The convolution process can be 8 different feature extractions for each channel of the first image. The size of the first image can be c x w x h, and the size of the three first feature maps can be (8 x c) x w x h, respectively.

[0162] Step 2, reshape processing is performed on the three first feature maps respectively, and the three-dimensional first feature map is reduced in dimension to obtain a two-dimensional first feature map, and the size of the two-dimensional first feature map can be (8 x c) x (w x h);

[0163] Step 3, matrix transposition processing is performed on the first first feature map k1;

[0164] Step 4, the first feature map after transposition processing is multiplied with the second first feature map q1 to obtain a multiplied feature map;

[0165] Step 5, classification and normalization processing are performed on the multiplied feature map to obtain a multi-feature attention weight map a1; wherein the classification can be performed by a softmax function, the normalization processing can be performed by a normalization layer, and the size of the multi-feature attention weight map can be (w x h) x (w x h);

[0166] Step 6, the weight map a1 is multiplied with the third feature map v1 to obtain a second feature map v2; wherein the size of the v2 can be (8 x c) x (w x h);

[0167] Step 7, dimension increasing processing is performed on the second feature map to obtain a three-dimensional second feature map v3, wherein the size of the v3 can be (8 x c) x w x h;

[0168] Step 8, dimension reduction convolution processing is performed on the three-dimensional second feature map v3 to obtain a second image; wherein the size of the second image can be c x w x h, and the second image can be a feature map processed by an improved attention mechanism.

[0169] As shown in Figure 10 The embodiment of the present disclosure provides an image generation device, wherein the device comprises:

[0170] The processing module 10 is configured to perform convolution processing on a randomly generated first image by using N convolution kernels to obtain N first feature maps; any two convolution kernels in the N convolution kernels have the same size and different convolution elements;

[0171] The generation module 20 is configured to generate a weight map according to the first first feature map to the N-1th first feature map; wherein the size of the weight map is equal to the size of the first feature map;

[0172] The obtaining module 30 is configured to obtain a second feature map based on the N first feature maps and the weight map;

[0173] The generation module 20 is further configured to generate a second image based on the second feature map.

[0174] In one embodiment, the generation module 20 further includes an adjustment module 21 configured to adjust positions of elements of at least some of the first feature maps to obtain N-1 third feature maps; and the generation module further includes a point multiplication module 23 configured to multiply the N-1 third feature maps to obtain the weight map.

[0175] In one embodiment, the adjustment module 21 further includes a transpose processing module 22 configured to perform transpose processing on at least some of the first feature maps to obtain the N-1 third feature maps.

[0176] In one embodiment, the obtaining module 30 is further configured to perform convolution processing on the input parameters to obtain a convolution processing result, perform residual connection on the convolution processing result and the input parameters to obtain a residual feature map, and perform dense connection on the residual feature map, the input parameters, and parameters obtained by the convolution processing to obtain the first image.

[0177] As shown in Figure 11 The present disclosure provides an image generation device, which includes:

[0178] A processing module 100 is configured to obtain a fourth feature map by performing convolution and pooling processing on a second image and a sample image, wherein the second image is generated by a generator; and an extraction module 200 is configured to extract features of the fourth feature map to obtain a first feature vector.

[0179] A determination module 300 is configured to determine a first cosine distance between the second image and the sample image with the same image content based on the first feature vector.

[0180] A loss value is determined according to the first cosine distance.

[0181] An adjustment module 400 is configured to adjust parameters of the generator according to the loss value.

[0182] In one embodiment, the processing module 100 is further configured to perform image transformation processing on the second image to obtain a third image, and perform image transformation processing on the sample image to obtain a fourth image.

[0183] The extraction module 200 is further configured to extract features of the third image and the fourth image to obtain a second feature vector.

[0184] The determining module 300 is further configured to determine a second cosine distance between the third image and the fourth image having the same image content according to the second feature map; and determine a loss value according to the first cosine distance and the second cosine distance.

[0185] In an embodiment, the image transformation processing in the processing module 100 includes at least one of the following: image offset; image cropping; image rotation; image filtering; and noise increase.

[0186] In an embodiment, the apparatus further includes an obtaining module 500 configured to obtain a random value; and the processing module 100 is further configured to perform image transformation processing on the second image to obtain a third image when the random value is greater than a probability value of the image transformation processing.

[0187] It should be noted that the method provided by the embodiments of the present disclosure can be executed alone or together with some methods in some methods or related technologies in the art.

[0188] The embodiments of the present disclosure further provide an electronic device, which includes a processor and a memory for storing a computer program capable of running on the processor, and the processor runs the computer program to execute the steps of the method in the foregoing one or more technical solutions.

[0189] The embodiments of the present disclosure further provide a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions are executed by a processor to implement the method in the foregoing one or more technical solutions.

[0190] The computer storage medium provided by the embodiments can be a non-transitory storage medium. In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division, and actual implementation can have another division manner, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be through some interface, indirect coupling or communication connection between devices or units, which can be electrical, mechanical or other forms.

[0191] The units described as separate components above can or can not be physically separate, and the components shown as units can or can not be physical units, that is, can be located in one place or distributed on multiple network units; part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0192] In addition, each functional unit in each embodiment of the present disclosure can be integrated into one processing module, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.

[0193] In some cases, any two of the above technical features can be combined into a new method technical solution without conflict.

[0194] In some cases, any two of the above technical features can be combined into a new device technical solution without conflict.

[0195] Those skilled in the art can understand that all or part of the steps of the above method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program executes the steps of the method embodiments when executed; and the foregoing storage medium includes mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic discs or optical discs, and various media that can store program codes.

[0196] The above is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present disclosure, which should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. An image generation method characterized by, The method is executed by a generator, comprising: performing convolution processing on an input parameter to obtain a convolution processing result; performing residual connection on the convolution processing result and the input parameter to obtain a residual feature map; performing dense connection on the residual feature map, the input parameter and a parameter obtained by the convolution processing to obtain a first image; performing convolution processing on the first image generated randomly by using N convolution kernels to obtain N first feature maps; any two convolution kernels in the N convolution kernels have the same size and different convolution elements; generating a weight map according to the first feature maps from the first to the N-1th; obtaining a second feature map based on the N first feature maps and the weight map; generating a second image based on the second feature map; wherein the convolution processing on the input parameter to obtain the convolution processing result comprises: inputting a random vector, and obtaining a random feature map by passing the random vector through a full connection layer and a transformation dimension layer; performing convolution processing on the random feature map as the input parameter.

2. The method of claim 1, wherein, The method of generating the weight map according to the first feature maps from the first to the N-1th comprises: adjusting the positions of elements in at least part of the first feature maps from the first to the N-1th to obtain N-1 third feature maps after adjustment; point multiplying the N-1 third feature maps to obtain the weight map.

3. The method of claim 2, wherein, The method of adjusting the positions of elements in at least part of the first feature maps from the first to the N-1th to obtain N-1 third feature maps after adjustment comprises: performing transpose processing on at least part of the first feature maps from the first to the N-1th to obtain the N-1 third feature maps.

4. An image generation method characterized by, The method is executed by a discriminator, comprising: obtaining a fourth feature map according to convolution and pooling processing on a second image and a sample image; wherein the second image is generated by the generator according to claim 1; extracting features of the fourth feature map to obtain a first feature vector; determining a first cosine distance between the second image and the sample image with the same image content based on the first feature vector; determining a loss value according to the first cosine distance; adjusting parameters of the generator according to the loss value.

5. The method of claim 4, wherein, The method further comprises: performing image transformation processing on the second image to obtain a third image; performing image transformation processing on the sample image to obtain a fourth image; extracting features of the third image and the fourth image to obtain a second feature vector; determining a second cosine distance between the third image and the fourth image with the same image content according to the second feature vector; The method of determining the loss value according to the first cosine distance comprises: determining the loss value according to the first cosine distance and the second cosine distance.

6. The method of claim 5, wherein, The image transformation processing comprises at least one of the following: image offset; image cropping; image rotation; image filtering; noise increase.

7. The method of claim 6, wherein, The method further comprises: obtaining a random value; The image transformation processing on the second image comprises: When the random value is greater than the probability value of the image transformation processing, the second image is subjected to image transformation processing to obtain a third image.

8. An image generation apparatus characterized by comprising: Comprise: The processing module is configured to: perform convolution processing on the input parameter to obtain a convolution processing result; and perform residual connection on the convolution processing result and the input parameter to obtain a residual feature map; The residual feature map, the input parameter, and the parameter obtained by the convolution processing are densely connected to obtain a first image; the first image is randomly generated and subjected to convolution processing by using N convolution kernels to obtain N first feature maps; any two convolution kernels in the N convolution kernels have the same size and different convolution elements; the convolution processing on the input parameter to obtain the convolution processing result comprises: inputting a random vector, and obtaining a random feature map by passing the random vector through a full connection layer and a transformation dimension layer; and performing convolution processing on the random feature map as the input parameter; The generation module is configured to: generate a weight map according to the first first feature map to the N-1 first feature map; wherein the size of the weight map is equal to the size of the first feature map; The obtaining module is configured to: obtain a second feature map based on the N first feature maps and the weight map; The generation module is further configured to generate a second image based on the second feature map.

9. An image generation apparatus characterized by comprising: Comprise: The processing module is configured to: obtain a fourth feature map according to the convolution and pooling processing on the second image and a sample image; wherein the second image is generated by the image generation apparatus according to claim 8; The extraction module is configured to: extract the features of the fourth feature map to obtain a first feature vector; The determination module is configured to: determine a first cosine distance between the second image and the sample image with the same image content based on the first feature vector; Determine a loss value according to the first cosine distance; The adjustment module is configured to: adjust the parameters of the generator according to the loss value.

10. An electronic device, comprising: The electronic device comprises a processor and a memory for storing a computer program capable of running on the processor, wherein the processor runs the computer program to perform the image generation method of any one of claims 1 to 7.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions; the computer executable instructions are executed by the processor to implement the image generation method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-generator convolution composite image adversarial network algorithm

    CN107563493A

  • An image coloring method based on a self-attention generative adversarial network

    CN109712203A