Image generation and shading method and apparatus
By using a Generative Adversarial Network (GAN) model combined with semantic labels and color features, the problem that existing image generation models cannot edit colors and colorization models lack image generation capabilities is solved, enabling local color control and real-time creation during the image generation process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2021-07-29
- Publication Date
- 2026-04-17
AI Technical Summary
Existing image generation models cannot edit colors locally, and existing coloring models lack image generation capabilities.
By employing a Generative Adversarial Network (GAN) model, combined with semantic labels and color features, and obtaining user input through a drawing interface, images are automatically generated, enabling the editing and control of local colors.
It enables the editing and control of local colors during image generation, providing real-time image creation capabilities and improving the efficiency and flexibility of artistic creation.
Smart Images

Figure CN115812221B_ABST
Abstract
Description
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 060,784, filed August 4, 2020, and U.S. Patent Application No. 17 / 122,680, filed December 15, 2020, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of image processing, specifically to an image generation and coloring method and apparatus. Background Technology
[0003] In existing technologies, image generation models can be used to generate unsupervised or user-input-conditioned images (such as text and semantic tags). However, in these image generation models, once the final image is generated, the colors cannot be edited locally.
[0004] Besides image generation models, colorization has become another popular research area in recent years. Examples include graffiti-based colorization methods that apply graffiti to edge inputs, and color transfer methods that transfer global tones from a reference image to the input, typically in grayscale. However, these models aim at pure colorization and lack image generation capabilities.
[0005] The methods and systems disclosed in this application are intended to solve one or more of the problems described above, as well as other problems. Summary of the Invention
[0006] One aspect of this application provides a method for image generation and colorization. The method includes: displaying a canvas interface; obtaining semantic labels for an image to be generated based on user input on the canvas interface, each semantic label representing the content of a region in the image to be generated; obtaining color features of the image to be generated; and automatically generating the image using a generative adversarial network (GAN) model based on the semantic labels and the color features. The color features are latent vectors input to the GAN model.
[0007] Another aspect of this application provides an apparatus for image generation and colorization. The apparatus includes: a memory; and a processor coupled to the memory. The processor is configured to perform: displaying a canvas interface; obtaining semantic labels of an image to be generated based on user input on the canvas interface, each semantic label representing the content of a region in the image to be generated; obtaining color features of the image to be generated; and automatically generating the image using a GAN model based on the semantic labels and the color features. The color features are latent vectors input to the GAN model.
[0008] Another aspect of this application provides a non-transitory computer-readable storage medium storing computer instructions. When a processor executes the computer instructions, it causes the processor to: display a drawing interface; obtain semantic labels for an image to be generated based on user input on the drawing interface, each semantic label representing the content of a region in the image to be generated; obtain color features of the image to be generated; and automatically generate the image using a GAN model based on the semantic labels and the color features. The color features are latent vectors input to the GAN model.
[0009] Those skilled in the art can understand other embodiments of this application by referring to the specification, claims and drawings of this application. Attached Figure Description
[0010] This application contains at least one color drawing. Copies of the disclosure of this application, as well as the color drawings, will be provided by the Patent Office upon request and payment of the necessary fees.
[0011] The following figures are merely examples illustrating the purpose of this application based on various disclosed embodiments and are not intended to limit the scope of this application.
[0012] Figure 1 This is a block diagram of an exemplary computing system disclosed in the embodiments of this application.
[0013] Figure 2 This is an exemplary image generation and coloring process in the embodiments disclosed in this application.
[0014] Figure 3 This is an exemplary drawing board interface diagram in the embodiments disclosed in this application.
[0015] Figure 4 This is a block diagram illustrating exemplary image generation and coloring in the embodiments disclosed in this application.
[0016] Figure 5 This is another exemplary image generation and coloring process disclosed in the embodiments of this application.
[0017] Figure 6AIt is a semantic diagram of an example consistent with the embodiments disclosed in this application.
[0018] Figure 6B This is used in the embodiments disclosed in this application. Figure 6A The example semantic graph is used as input to the image generation model to automatically generate an image.
[0019] Figure 6C and Figure 6E These are two exemplary stroke images consistent with the embodiments disclosed in this application.
[0020] Figure 6D This is used in the embodiments disclosed in this application. Figure 6A Semantic graph of the example and Figure 6C The example stroke image is used as input to the image generation model to automatically generate the image.
[0021] Figure 6F This is used in the embodiments disclosed in this application. Figure 6A Semantic graph of the example and Figure 6E The example stroke image is used as input to the image generation model to automatically generate the image.
[0022] Figure 7A It is a semantic diagram of an example consistent with the embodiments disclosed in this application.
[0023] Figure 7B This is used in the embodiments disclosed in this application. Figure 7A The example semantic graph is used as input to the image generation model to automatically generate an image.
[0024] Figure 7C and Figure 7E These are two exemplary stroke images consistent with the embodiments disclosed in this application.
[0025] Figure 7D This is used in the embodiments disclosed in this application. Figure 7A Semantic graph of the example and Figure 7C The example stroke image is used as input to the image generation model to automatically generate the image.
[0026] Figure 7F This is used in the embodiments disclosed in this application. Figure 7A Semantic graph of the example and Figure 7E The example stroke image is used as input to the image generation model to automatically generate the image.
[0027] Figure 8A It is a semantic diagram of an example consistent with the embodiments disclosed in this application.
[0028] Figure 8B This is used in the embodiments disclosed in this application. Figure 8AThe example semantic graph is used as input to the image generation model to automatically generate an image.
[0029] Figure 8C , 8E 8G, 8I, 8K and 8M are exemplary stroke images consistent with the embodiments disclosed in this application.
[0030] Figure 8D This is used in the embodiments disclosed in this application. Figure 8A Semantic graph of the example and Figure 8C The example stroke image is used as input to the image generation model to automatically generate the image.
[0031] Figure 8F This is used in the embodiments disclosed in this application. Figure 8A Semantic graph of the example and Figure 8E The example stroke image is used as input to the image generation model to automatically generate the image.
[0032] Figure 8H This is used in the embodiments disclosed in this application. Figure 8A Semantic graph of the example and Figure 8G The example stroke image is used as input to the image generation model to automatically generate the image.
[0033] Figure 8J This is used in the embodiments disclosed in this application. Figure 8A Semantic graph of the example and Figure 8I The example stroke image is used as input to the image generation model to automatically generate the image.
[0034] Figure 8L This is used in the embodiments disclosed in this application. Figure 8A Semantic graph of the example and Figure 8K The example stroke image is used as input to the image generation model to automatically generate the image.
[0035] Figure 8N This is used in the embodiments disclosed in this application. Figure 8A Semantic graph of the example and Figure 8M The example stroke image is used as input to the image generation model to automatically generate the image. Detailed Implementation
[0036] Reference will now be made in detail to embodiments of the present application illustrated in the accompanying drawings. Embodiments consistent with the disclosure of this application will be described below with reference to the accompanying drawings. Where possible, all drawings will use the same reference numerals to refer to the same or similar parts. Obviously, the described embodiments are only a portion, not all, of the embodiments of this application. Other embodiments consistent with this application, obtained by those skilled in the art based on the disclosed embodiments, are within the scope of protection of this invention.
[0037] This application discloses a method and apparatus for image generation and colorization. The disclosed method and / or apparatus can be applied to implement an Artificial Intelligence (AI) canvas based on Generative Adversarial Networks (GANs). The GAN-based canvas is configured to acquire semantic cues (e.g., through segmentation) and hues (e.g., through strokes) from user input and automatically generate an image (e.g., a painting) based on the user input. The disclosed method is built upon a novel and lightweight color feature embedding technique that incorporates colorization effects into the image generation process. Unlike existing GAN-based image generation models that only accept semantic input, the canvas disclosed in this application has the ability to edit local colors of the image after image generation. Color information can be sampled as additional input from the strokes input by the user and fed back into the GAN model for conditional generation. The method and apparatus disclosed in this application can create images or paintings with semantic and color control. That is, the method and apparatus disclosed in this application incorporate colorization into the image generation process, thereby allowing simultaneous control of the position, shape, and color of objects. This application can perform image creation in real time.
[0038] Figure 1 This is a block diagram of a computing system / apparatus implementing the image generation and coloring methods in the embodiments disclosed in this application. Figure 1 As shown, the computing system 100 may include a processor 102 and a storage medium 104. According to some embodiments, the computing system 100 may also include a display 106, a communication module 108, additional peripheral devices 112, and one or more buses 114. The buses 114 couple the aforementioned devices together. Some devices may be omitted, while others may be included within the computing system 100.
[0039] Processor 102 includes any suitable processor. In some embodiments, processor 102 may include multiple cores for multithreaded or parallel processing, and / or a Graphics Processing Unit (GPU). Processor 102 can execute sequences of computer program instructions to perform various processes, such as image generation and colorization programs, GAN model training programs, etc. Storage medium 104 may be a non-transitory computer-readable storage medium and may include storage modules such as ROM, RAM, flash memory modules, and erasable and rewritable memories, as well as large-capacity storage devices such as CD-ROMs, USB flash drives, and hard disks. Storage medium 104 may store computer programs for implementing various processes, which are executed by processor 102. Storage medium 104 may also include one or more databases for storing certain data, such as image data, training datasets, test image datasets, data of trained GAN models, and for performing certain operations on the stored data, such as database searches and data retrieval.
[0040] Communication module 108 may include network devices for establishing connections over a network. Display 106 may include any suitable type of computer display device or electronic device display (e.g., CRT or LCD-based devices, touch screens). Peripheral devices 112 may include I / O devices such as keyboards, mice, etc.
[0041] During operation, the processor 102 can be configured to execute instructions stored on the storage medium 104 and perform various operations related to the image generation and coloring methods as described in detail below.
[0042] Figure 2 An exemplary image generation and coloring process 200 in an embodiment of this application is described. Figure 4 This is a block diagram of an exemplary framework 400 for image generation and colorization according to an embodiment of this application. The image generation and colorization process 200 can be performed by any suitable computing device / server having one or more processors and one or more memories, such as computing system 100 (e.g., processor 102). Framework 400 can be implemented by any suitable computing device / server having one or more processors and one or more memories, such as computing system 100 (e.g., processor 102).
[0043] like Figure 2 As shown, a canvas interface is displayed (S202). The canvas interface may include multiple functions related to image generation and colorization. For example, semantic labels of the image to be generated can be obtained based on user input on the canvas interface (S204). Each semantic label represents the content of a region in the image to be generated. Semantic labels can also be represented as... Figure 4The input label 402 is shown. Input label 402 is one of the inputs to generator 406 (e.g., painting generator 4062).
[0044] Figure 3 This is a schematic diagram of an exemplary drawing board interface in an embodiment disclosed in this application. Figure 3 As shown, the canvas interface may include a label input section (a). The label input may be presented in the form of a semantic map. In this application, a semantic map may refer to a semantic segmentation mask of a target image (e.g., the image to be generated). The semantic map is the same size as the target image and is associated with multiple semantic labels; each pixel in the semantic map has a corresponding semantic label. Pixels with the same semantic label describe the same topic / content. In other words, the semantic map includes multiple regions, and pixels in the same region of the semantic map have the same semantic label indicating the content of that region. For example, the semantic map may include regions labeled as sky, mountain, rock, tree, grass, sea, etc. Such regions can also be called label images. That is, a semantic map may be processed into different label images, each label image representing a different semantic label.
[0045] In some embodiments, a semantic graph can be obtained based on the user's drawing operations on the drawing board interface (e.g., in...). Figure 3 After selecting the "Label" button in the color-label conversion button section (d). For example, the canvas interface can provide drawing tools for users to draw semantic graphs (e.g., Figure 3 (The tools shown in the drawing tools and elements section (c)). Drawing tools can include common drawing functions such as pencil, eraser, zoom in / out, color fill, etc. Users can use drawing tools to outline the contours / edges of different areas describing different components of the desired image. Semantic label text can be assigned by the user to each area of the semantic map. For example, semantic label text can be displayed on the canvas interface, with each semantic label text having a unique format serving as a legend for the corresponding area. Each area of the semantic map displayed in the canvas interface is formatted according to the format of the corresponding semantic label text. A unique format usually means that areas with different semantic labels will not have the same format. The format can be color and / or pattern (e.g., dotted areas, striped areas). For example, when using color to represent object type, the area associated with the semantic label "sky" is light blue, and the area associated with the semantic label "sea" is dark blue. Figure 6A This is a semantic diagram of an example consistent with the embodiments disclosed in this application. Similarly, Figure 7A and Figure 8A These are semantic graphs for two other examples.
[0046] In some embodiments, semantic tags can be obtained based on a semantic graph template. In one example, the canvas interface can provide candidate options for the semantic graph (e.g., ...). Figure 3 (See the label template section (b)). Users can select a desired template from the candidate options. In another example, the canvas interface can read a previously created image from local or online storage and obtain the semantic label text assigned by the user to regions in the previously created image. The labeled image can serve as a semantic graph template. In some embodiments, the semantic graph template can be used directly to obtain semantic labels. In other embodiments, the user can modify the semantic graph, for example, by modifying the semantic label text corresponding to regions (one or more pixels) of the semantic graph, or by modifying the position of the regions corresponding to the semantic label text (using the drawing tools provided by the canvas interface).
[0047] refer to Figure 2 and Figure 4 The color features of the image to be generated are obtained (S206). The color feature embedding unit 404 can extract color features. First, the input color 4042 is obtained based on user input or default settings, and the obtained color features 4044 with color values (e.g., RGB values) and positions are one of the inputs to the generator 406 (e.g., painting generator 4062).
[0048] Custom coloring is implemented based on feature embedding. The user-customized input color 4042 can include a stroke image. The stroke image is the same size as the target image and includes one or more strokes input by the user. Each stroke consists of connected pixels and has a corresponding stroke color. Figure 6C , 6E 7C, 7E, 8C, 8E, 8G, 8I, 8K, and 8M are exemplary stroke images consistent with embodiments of this application. For example, Figure 6C The stroke image shown has the same size as the target image, and the target image has the same... Figure 6A The semantic graphs shown are the same size. The background of the stroke image is black, and four colored strokes (two strokes have the same color, and the other two strokes each have their own color) are located in different areas of the stroke image, indicating the color arrangement desired by the user for these areas.
[0049] In some embodiments, stroke images can be obtained based on user actions on the drawing board interface. For example, candidate colors can be displayed on the drawing board interface (e.g., in...). Figure 3In the color-label conversion button section (d), candidate colors are displayed after selecting the "Color" button. The target color can be obtained based on the user's selection from the candidate colors. Drag operations are detected (e.g., drag operations are detected on the stroke image displayed on the canvas interface), and the colored strokes of the stroke image are recorded according to the target color and the position corresponding to the drag operation. In some embodiments, when no user input is obtained for the stroke image, the color features are extracted from a default stroke image with a uniform color (e.g., black).
[0050] Feature embedding is similar to converting input strokes into high-intensity features. For example, a stroke image is a sparse matrix where only small regions are occupied by stroke color. If this sparse matrix is taken as input according to traditional variational encoder practices, color control may be difficult to achieve. In this case, the encoding module applies the same function to each pixel and makes non-stroke regions dominate. As the output of a traditional variational encoder, only a few differences are observed in the encoder's output.
[0051] In some embodiments, the sparse matrix representing the stroke image is transformed into a strong feature representation for effective control over the results. The color feature embedding process is analogous to object detection based on a regression problem. The input image (i.e., the stroke image) is divided into multiple grids, each grid potentially including zero strokes, a portion of a stroke, a complete stroke, or multiple strokes. In one example, the multiple grids could be S×S grids. In another example, the multiple grids could be S1×S2 grids. Unlike object detection, the color values (e.g., RGB values) of the stroke image can be used as object scores and category scores.
[0052] Mathematically, in some embodiments, taking a scenario where a stroke image is divided into S×S grids as an example, the color features extracted from the stroke image are defined as an array of size (S,S,8), where S×S is the number of grids in the image and 8 is the number of feature channels. Furthermore, for each grid / cell, a tuple (mask,x,y,w,h,R,G,B) is defined to represent the features associated with the grid. Specifically, a maximum rectangular region that can cover one or more strokes corresponding to a grid (e.g., a rectangle whose top-left corner is inside the grid covers any stroke whose top-left corner is inside the corresponding grid) is represented as (x,y,w,h), where x and y are described in Equation 1:
[0053] x = x image -offset x ,y=y image -offset y
[0054] offset xoffset y x is the coordinate of the top left corner of the grid. image y image Let w and h be the coordinates of the top-left corner of the rectangular region, and w and h be the size of the rectangular region (i.e., width and height). (R, G, B) are the average values of each hue within the grid. Furthermore, a mask channel is added to indicate whether the grid contains strokes, to avoid ambiguity from black stroke inputs. For example, the mask value is 0 when the grid does not contain strokes, and 1 when the grid contains pixels of colored strokes (e.g., the color could be black, blue, or any other color).
[0055] Color features associated with stroke images are input into an image generation model to automatically predict / generate a desired image based on stroke images from user input. Training the image generation model also involves these color features. Specifically, when preparing the training dataset, stroke images are generated from original images, one stroke image per original image. Each stroke image generated during training has a valid color in each corresponding grid cell. The color of the grid cells in the stroke image is determined based on the color of the corresponding grid cells in the original images. Furthermore, for each stroke image generated during training, to better simulate the input color, a preset percentage (e.g., 75%) of its total grid cells is set to 0 in all 8 channels, meaning that the remaining percentage (e.g., 25%) of the total grid cells in the stroke image has color. In some embodiments, S is set to 9 and the color feature array is flattened into a latent vector input for the image generator 406. For example, an array of size (S, S, 8) is converted into a one-dimensional latent vector having S×S×8 elements. The operation of obtaining the flattened latent vector (i.e., automatic image generation based on user input) can be performed during the training and testing phases.
[0056] It is understood that steps S202 and S204 do not have a specific execution order. That is, step S202 can be performed before, after, or simultaneously with step S204; no limitation is made here. As long as the two inputs (i.e., semantic labels and color features) are ready, generator 406 can automatically generate an image with custom colors.
[0057] Image generation can be achieved using deep learning methods based on Generative Adversarial Networks (GAN) models. The goal of a GAN model is to find a result that satisfies both the layout indicated by the semantic map and the local color indicated by the stroke image for drawing. That is, images can be automatically generated using a GAN model based on semantic labels and color features (S208). Color features can be latent vectors input to the GAN model. The input to the GAN model includes latent vectors representing color features extracted from the stroke image and semantic labels (e.g., semantic maps). For example, a semantic map with N semantic labels is processed into a label image with N channels (each channel / image represents a semantic label). In embodiments of this application, semantic labels can also refer to images with the same size as the original image or to a generated image depicting pixels corresponding to a single object / entity and labeled with that single object / entity.
[0058] In some embodiments, the GAN model of this application includes a generator and a multi-scale discriminator. The generator takes color features and semantic labels as explicit input. The generator includes multiple cascaded upsampling blocks, each corresponding to a different image resolution, and includes a set of convolutional layers and attention layers that accept semantic labels as shape constraints to implement a coarse-to-fine strategy. That is, during operation, the process sequentially passes through multiple upsampling blocks in order of increasing resolution (i.e., first processes the upsampling blocks corresponding to the coarse resolution, and then processes the upsampling blocks corresponding to the higher resolution).
[0059] The input to the first upsampling block is color features and semantic labels, while the input to all upsampling blocks is the output of their preceding upsampling block. The output of each upsampling block is also used as a hidden feature for the next upsampling block. The output of each upsampling block can include a painting / image with the same resolution as the corresponding upsampling block. In some embodiments, the painting / image from the output of an upsampling block can be resized (e.g., doubled in size, such as resizing an 8x8 image to a 16x16 image) before being input to the next upsampling block with the same configuration.
[0060] In each upsampling block, the semantic labels, serving as shape constraints, are resized to have the same resolution as the hidden features. The generator can employ a spatial adaptation method from the GauGAN model and use the resized semantic labels as attention input to the corresponding upsampling block. For example, the input to the current upsampling block corresponding to a 128*128 resolution might be the output of the previous upsampling block corresponding to a 64*64 resolution; and the size of the semantic labels used for the attention layer in the current upsampling block can be adjusted to 128*128. In some embodiments, the resolution corresponding to each upsampling block can be twice the resolution corresponding to the previous upsampling block. For example, the resolution can include 4*4, 8*8, 16*16, ... up to 1024*1024.
[0061] During training, the generator aims to learn how to output the original painting with colors (i.e., from the corresponding stroke images) and semantics (e.g., the configuration of upsampled blocks). On the other hand, the multi-scale discriminator takes the generated painting / image as input, ensuring that the image generated by the generator is similar to the original image. The multi-scale discriminator is only used during the training phase.
[0062] Compared to existing technologies, such as GauGAN which cannot control the local colors of the generated images, the GAN model disclosed in this application can generate images that take into account both semantic labels and local colors customized by the user through stroke images.
[0063] Following the training phase, only the generator and color feature embedding module (i.e., the module that converts the user's stroke image into a color feature vector) based on the trained model operate, with user-defined strokes and semantic labels as input. The trained model, i.e., the generator, can output pixel values (R, G, B) for each pixel location in the output image, given color and shape constraints (i.e., from the stroke image and semantic map). The image generation process in this application can be implemented in real time. For example, in experimental inference (e.g., when generator 406 generates images based on received user input and the trained GAN model), each image generation takes 0.077 seconds on a Tesla V100 GPU.
[0064] exist Figure 3 On the canvas interface shown, Figure 3 After selecting the "draw" button in section (d) of the color-label conversion button, the resulting image is generated using the revealed GAN model based on the label entered in section (a) and the optional stroke image corresponding to the "color" button function. The resulting image can be displayed in the results section (e) of the canvas interface.
[0065] Figure 6B This is used in the embodiments of this application. Figure 6AThe example semantic graph is used as input to an image generation model (i.e., a GAN model) to automatically generate images (another input to the GAN model is a default stroke image, such as a completely black image). Figure 6D Is using Figure 6A Semantic graph of the example and Figure 6C The example stroke image is automatically generated after being used as input to a GAN model. From Figure 6B and Figure 6D As can be seen, color guidance can produce different resulting images.
[0066] Figure 5 This is another exemplary image generation and coloring process disclosed in the embodiments of this application. In S502, the user inputs a label image (i.e., a semantic label existing in the form of a semantic graph) to generate an image using a default all-black color (i.e., an all-black stroke image with a color feature mask value of 0). Based on the input semantic label, a painting is automatically generated using a GAN model (S504). The generated image (i.e., the painting) can be displayed on the canvas interface. The user may want to modify certain local shapes or colors of an object in the image. That is, partial modification of color and / or label is allowed (S506). Shape modification can be achieved by modifying the desired part in the semantic graph to obtain the current semantic label. Color customization can be achieved by providing a stroke image, which is used to extract the current color features. For example, when the "Draw!" button is selected, the modified painting is returned using a GAN model based on the current semantic label and current color features (S508). The modified painting can be displayed. Steps S506 and S508 can be repeated until the user obtains a satisfactory image.
[0067] In some embodiments, the drawing interface can store images generated using the return drawing element 4064 at each revision (e.g., each time the "Draw!" button is selected). The generated images can be Image(1), Image(2), ..., and Image(N), where Image(N) is the currently generated image. In some scenarios, the user may prefer the previously generated image and give a return instruction, for example, in... Figure 3 In the drawing tools and elements section (c), select the return button. Upon receiving a return command, the canvas interface can display the most recent image, such as Image(N-1). If another return command is received, Image(N-2) will be displayed.
[0068] Figures 6A-6F , Figures 7A-7F and Figure 8A-8N These are three sets of examples showing different results based on different semantic label inputs and color inputs. Specifically, Figure 6A This is a semantic graph of the example. Figure 6B Is using Figure 6AThe example semantic graph is used as input to automatically generate an image. Figure 6C , Figure 6E Is with Figure 6A Two identical exemplary stroke images. Figure 6D Is using Figure 6A Semantic graph of the example and Figure 6C The example stroke image is an automatically generated image. Figure 6F Is using Figure 6A Semantic graph of the example and Figure 6E The example stroke image is an automatically generated image. It is understandable that when the semantic graph is the same (e.g., ...), ... Figure 6A As shown), but the stroke image input by the user (e.g.) Figure 6C and 6E As shown, the resulting images generated from stroke images and semantic maps (as shown) contain the same content arranged in similar layouts, but these contents may have different color features. For example, due to the fact that in Figure 6C The cyan strokes drawn at the top of the stroke image, such as Figure 6D The sky in the resulting image shown is blue and clear, but due to... Figure 6E The red strokes drawn at the top of the stroke image make... Figure 6F The resulting image shows a dark and cloudy scene. Different color characteristics can produce drastically different aesthetic interpretations of paintings depicting the same subject.
[0069] Figure 7A This is another example of a semantic graph. Figure 7B Is using Figure 7A The example semantic graph is used as input to the image generation model to automatically generate images. Figure 7C and Figure 7E These are two exemplary stroke images conforming to embodiments of this application. Figure 7D Is using Figure 7A Semantic graph of the example and Figure 7C The example stroke image is used as input to the image generation model to automatically generate the image. Figure 7F This is used in some embodiments of this application. Figure 7A Semantic graph of the example and Figure 7E The example stroke image is used as input to the image generation model to automatically generate the image. Figure 7D and Figure 7F Two paintings with similar content and layout but different color characteristics were presented.
[0070] Figure 8A It is a semantic graph consistent with some embodiments of this application. Figure 8B This is used according to some embodiments of this application. Figure 8A The example semantic graph is used as input to the image generation model to automatically generate images. Figure 8C ,8E 8G, 8I, 8K, and 8M are exemplary stroke images, and the corresponding resulting images are respectively Figure 8D , 8F The images shown are 8H, 8J, 8L, and 8N. For example, Figure 8G and 8I The stroke image shown includes cyan dots, located in positions corresponding to the sky. However, Figure 8G In the stroke image, both the positions corresponding to the higher and lower skies include cyan dots, while... Figure 8I The stroke image only contains cyan dots in the higher sky regions. Therefore, in the lower sky regions, Figure 8G The result image corresponding to the middle stroke image Figure 8H Compared to Figure 8I The result image corresponding to the middle stroke image Figure 8J Including more color variations. Therefore, the device disclosed in this application can create artistic differences in the resulting image / painting based on simple color strokes, thus providing users with a powerful tool to realize their creative ideas.
[0071] One of the pain points in the art field is that attempting a creative process is time-consuming, and it's not easy to revert to previous versions. The disclosed system / device effectively addresses this issue. Through simple strokes with shapes and colors, the disclosed method helps artists quickly realize their ideas. Furthermore, if the artist is not satisfied with the result, they can quickly return to the previous version by clicking the back button. This can provide an efficient teaching tool in the art field, greatly improving painters' creative efficiency and reducing unnecessary labor.
[0072] In summary, this disclosure provides a GAN-based canvas that helps users simultaneously generate and edit paintings. Painters can repeatedly generate and modify their paintings through simple manipulations of labels and colors. The feature embedding module used is lightweight, and its output serves as input to the latent vectors of the GAN model in the generation task. In some embodiments, feature extraction is restricted to each grid cell corresponding to only one set of color features. In some embodiments, allowing multiple colors within a single grid cell can improve the resulting image.
[0073] Those skilled in the art will understand that all or part of the steps in the above method embodiments can be implemented by a related hardware executable program, which can be stored in a computer-readable storage medium. The program is executed to perform the steps of the method embodiments. The storage medium includes media capable of storing program code, such as mobile storage devices, read-only memory (ROM), disks, or optical discs.
[0074] Optionally, when the integrated unit is implemented as a software functional unit and sold or used as an independent product, the integrated unit can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution, can be implemented in the form of a software product. This software product is stored in a storage medium and includes several instructions for instructing a computer device (e.g., a personal computer, server, or network device) to execute all or part of the steps in the above-described method embodiments of this application. The aforementioned storage medium includes any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, optical disk, etc.
[0075] Other embodiments of this application will be apparent to those skilled in the art based on the specification and examples provided. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
Claims
1. An image generation and shading method applied in a computing device, characterized in that, The method includes: Display the canvas interface; Based on the user input on the drawing board interface, semantic tags of the image to be generated are obtained, and each semantic tag represents the content of a region in the image to be generated; Obtaining the color features of the image to be generated includes: displaying candidate colors on the canvas interface; obtaining a target color from the candidate colors based on user selection; detecting drag operations; and recording colored strokes of a stroke image according to the target color and the position corresponding to the drag operation, wherein the background of the stroke image is black, the colored strokes are located in different regions of the stroke image, and indicate the color arrangement desired by the user in the region; extracting the color features based on the stroke image; and Based on the semantic labels and the color features, an image is automatically generated using a generative adversarial network (GAN) model, wherein the color features are latent vectors input to the GAN model.
2. The image generation and shading method of claim 1, wherein, Obtaining the color features includes: Obtain a stroke image of the same size as the image to be generated, the stroke image including one or more colored strokes input by the user.
3. The image generation and shading method of claim 2, wherein, When the stroke image does not receive user input, the color features are extracted from a default stroke image with uniform color.
4. The image generation and shading method of claim 1, wherein, Obtaining the semantic tags includes: Obtain a semantic map of the same size as the image to be generated. The semantic map includes multiple regions, and pixels in the same region have the same semantic label. The semantic label represents the content of the corresponding region.
5. The image generation and shading method of claim 4, wherein, Obtaining the semantic tags also includes: The semantic graph is generated based on the user's drawing operations on the drawing board interface; Receive semantic label text assigned by the user to multiple regions of the semantic graph; and The semantic tags are obtained based on the semantic graph drawn by the user and the assigned semantic tag text.
6. The image generation and coloring method as described in claim 4, characterized in that, Obtaining the semantic tags also includes: Modifying the semantic graph based on user operations on the canvas interface includes at least one of the following: modifying the first semantic label text corresponding to one or more pixels in the semantic graph, or modifying the position of the region corresponding to the second semantic label text.
7. The image generation and shading method of claim 4, wherein, Obtaining the semantic tags also includes: Multiple semantic graph templates are displayed on the canvas interface; and The semantic label is obtained based on the user's selection of one of multiple semantic graph templates.
8. The image generation and shading method of claim 1, wherein, The method further includes: The automatically generated image is displayed on the canvas interface, and the automatically generated image is the first image. Obtain a modification instruction for at least one of the semantic tags or the color features; and The modified image is generated using the GAN model based on the semantic labels and color features updated according to the modification instructions.
9. The image generation and shading method of claim 8, wherein, The method further includes: The modified image is displayed on the canvas interface; Receive return indication; and In response to the return command, the first image is displayed on the canvas interface.
10. An image generation and shading apparatus characterized by comprising: The device includes: Memory; and A processor, coupled to the memory and configured to execute: Display the canvas interface; Based on the user input on the drawing board interface, semantic tags of the image to be generated are obtained, and each semantic tag represents the content of a region in the image to be generated; Obtaining the color features of the image to be generated includes: displaying candidate colors on the canvas interface; obtaining a target color from the candidate colors based on user selection; detecting drag operations; and recording colored strokes of a stroke image according to the target color and the position corresponding to the drag operation, wherein the background of the stroke image is black, the colored strokes are located in different regions of the stroke image, and indicate the color arrangement desired by the user in the region; extracting the color features based on the stroke image; and Based on the semantic labels and the color features, an image is automatically generated using a generative adversarial network (GAN) model, wherein the color features are latent vectors input to the GAN model.
11. The image generation and shading apparatus of claim 10, wherein, Obtaining the color features includes: Obtain a stroke image of the same size as the image to be generated, the stroke image including one or more colored strokes input by the user.
12. The image generation and shading apparatus of claim 11, wherein, When the stroke image does not receive user input, the color features are extracted from a default stroke image with uniform color.
13. The image generation and shading apparatus of claim 10, wherein, Obtaining the semantic tags includes: Obtain a semantic map of the same size as the image to be generated. The semantic map includes multiple regions, and pixels in the same region have the same semantic label. The semantic label represents the content of the corresponding region.
14. The image generation and shading apparatus of claim 13, wherein, Obtaining the semantic tags also includes: The semantic graph is generated based on the user's drawing operations on the drawing board interface; Receive semantic label text assigned by the user to multiple regions of the semantic graph; and The semantic tags are obtained based on the semantic graph drawn by the user and the assigned semantic tag text.
15. The image generation and shading apparatus of claim 13, wherein, Obtaining the semantic tags also includes: Modifying the semantic graph based on user operations on the canvas interface includes at least one of the following: modifying the first semantic label text corresponding to one or more pixels in the semantic graph, or modifying the position of the region corresponding to the second semantic label text.
16. The image generation and shading apparatus of claim 13, wherein, Obtaining the semantic tags also includes: Multiple semantic graph templates are displayed on the canvas interface; and The semantic label is obtained based on the user's selection of one of multiple semantic graph templates.
17. The image generation and shading apparatus of claim 10, wherein, The processor is also configured to execute: The automatically generated image is displayed on the canvas interface, and the automatically generated image is the first image. Obtain a modification instruction for at least one of the semantic tags or the color features; and The modified image is generated using the GAN model based on the semantic labels and color features updated according to the modification instructions.
18. The image generation and shading apparatus of claim 17, wherein, The processor is also configured to execute: The modified image is displayed on the canvas interface; Receive return indication; and In response to the return command, the first image is displayed on the canvas interface.
19. A non-transitory computer-readable storage medium storing computer instructions that, when executed by a processor, cause the processor to perform: Display the canvas interface; Based on the user input on the drawing board interface, semantic tags of the image to be generated are obtained, and each semantic tag represents the content of a region in the image to be generated; Obtaining the color features of the image to be generated includes: displaying candidate colors on the canvas interface; obtaining a target color from the candidate colors based on user selection; detecting drag operations; and recording colored strokes of a stroke image according to the target color and the position corresponding to the drag operation, wherein the background of the stroke image is black, the colored strokes are located in different regions of the stroke image, and indicate the color arrangement desired by the user in the region; extracting the color features based on the stroke image; and Based on the semantic labels and the color features, an image is automatically generated using a generative adversarial network (GAN) model, wherein the color features are latent vectors input to the GAN model.
Citation Information
Patent Citations
Electronic drawing board copying method and related equipment
CN110244870A
Semantic image synthesis for generating substantially photorealistic images using neural networks
US20200242771A1