Image color filling method and device, equipment and storage medium

By performing initial color filling and area adjustment on the line art, a target image that conforms to semantic information is generated, solving the problem of discrepancies between the generated line art and the generated image, thus improving the success rate and speed.

CN122049098APending Publication Date: 2026-05-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

In existing technologies, when generating line art, the target image does not match the staff's expectations, and the generation speed is slow, making it difficult to effectively improve the success rate.

Method used

By acquiring line art and text prompts, initial coloring is performed to generate an initial image and feature map. The initial coloring area is extracted and adjusted to obtain a baseline coloring area. Target coloring is then performed to generate a target image that conforms to semantic information.

Benefits of technology

It improves the success rate and accuracy of line drawing generation, avoids situations where the generated target image does not match the expectation, and increases the generation speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049098A_ABST
    Figure CN122049098A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and provides an image color filling method and device, equipment and a storage medium. The method can solve the problem of low success rate of line draft drawing in related technologies, and comprises the following steps: obtaining a line draft, obtaining a text prompt set corresponding to the line draft, adopting the text prompt, carrying out initial color filling processing on the line draft, and generating an initial image and a plurality of initial feature maps; based on the plurality of initial feature maps, extracting an initial color filling area corresponding to each cue word in the at least one cue word in the initial image, and when any initial color filling area is not matched with semantic information of the corresponding cue word, performing area division adjustment on the extracted at least one initial color filling area, obtaining a reference color filling area corresponding to each cue word in the at least one cue word; and performing target color filling processing on the line draft based on the obtained reference color filling areas, the line draft and the at least one prompt word to generate a target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and provides an image coloring method, apparatus, device, and storage medium. Background Technology

[0002] In recent years, with the development of deep learning and artificial intelligence technologies, line art has been widely used in illustration, design, animation, and other fields. Line art uses simple lines to outline the shape of an object, while line art refers to the process of filling in the line art with colors to obtain a more colorful image.

[0003] Currently, when generating line art, the process typically involves using a generative model (e.g., StableDiffusion) to fill in the various areas of the line art with color based on the color filling requirements described in the prompt, thereby generating a more complete target image.

[0004] However, in practical applications, when coloring each coloring area in the line drawing according to the text prompts, the generated target image often does not match the expectations of the relevant staff because there is a discrepancy between the coloring area judged by the model and the coloring area indicated in the corresponding text prompts.

[0005] In related technologies, the above problems are usually solved in the following two ways:

[0006] Method 1 involves randomly selecting an art style through a card draw and then regenerating the target image.

[0007] However, this method is highly random and it is difficult to guarantee the generation effect.

[0008] Method 2 involves using separate text prompts for each coloring area in the line art, generating a corresponding area image for each coloring area, and then merging the generated area images into the target image.

[0009] However, this method makes it difficult to select the specific areas to control in the line art because it cannot observe the internal division of the fill regions within the generated model. Furthermore, this method requires generating multiple region images and then fusing them into a target image, resulting in a slow target image generation speed. Additionally, if there are semantic conflicts in the text prompts of different fill regions during the fusion process, the target image generation may fail.

[0010] Therefore, there is a lack of effective methods in related technologies to improve the success rate of line drawing generation. Summary of the Invention

[0011] This application provides an image coloring method, apparatus, device, and storage medium to solve the problem of low success rate of line drawing image generation in related technologies.

[0012] In a first aspect, embodiments of this application provide an image coloring method, including:

[0013] Obtain the line art and obtain the text prompts corresponding to the line art settings, the text prompts including: at least one prompt word for describing the coloring rules of the line art;

[0014] Using the text prompts, the line drawing is initially filled with color to generate an initial image and multiple initial feature maps; each initial feature map contains: the region features corresponding to each prompt word in the at least one prompt word, extracted during one step of the initial color filling process;

[0015] Based on the multiple initial feature maps, in the initial image, an initial coloring region corresponding to each of the at least one prompt words is extracted. When any of the initial coloring regions does not match the semantic information of the corresponding prompt word, the extracted at least one initial coloring region is divided and adjusted to obtain the reference coloring region corresponding to each of the at least one prompt words.

[0016] Based on the obtained baseline coloring areas, the line art, and the at least one prompt word, target coloring processing is performed on the line art to generate a target image.

[0017] Secondly, embodiments of this application also provide an image coloring device, comprising:

[0018] A communication unit is used to acquire a line drawing and to acquire text prompts corresponding to the line drawing, wherein the text prompts include at least one prompt word describing the coloring rules of the line drawing;

[0019] An initial coloring unit is used to perform initial coloring processing on the line drawing using the text prompts, generating an initial image and multiple initial feature maps; each initial feature map includes: region features corresponding to each prompt word in the at least one prompt word extracted during a step of performing the initial coloring processing;

[0020] The region adjustment unit is used to extract the initial coloring region corresponding to each prompt word in the initial image based on the plurality of initial feature maps, and to perform region division adjustment on the extracted at least one initial coloring region when any initial coloring region does not match the semantic information of the corresponding prompt word, so as to obtain the reference coloring region corresponding to each prompt word in the at least one prompt word.

[0021] The target coloring unit is used to perform target coloring processing on the line drawing based on the obtained reference coloring areas, the line drawing, and the at least one prompt word, to generate a target image.

[0022] In one possible implementation, after the communication unit obtains the text prompt corresponding to the line art setting, the prompt word segmentation unit is used to: segment the text prompt to obtain at least one word; divide the at least one word into at least one prompt word according to the position of each word in the text prompt; each prompt word is used to describe the coloring rules of a region in the line art.

[0023] In one possible implementation, the initial coloring unit uses the text prompt to perform initial coloring processing on the line art, generating an initial image and multiple initial feature maps. Specifically, it performs the following operations on multiple sub-models with hierarchical relationships included in the generation model, until the initial image is generated: inputting the line art, the at least one prompt word, and the fusion features obtained by the previous level sub-model into the current level sub-model to extract text features and image features at the current level; performing feature fusion on the text features and image features to obtain fusion features at the current level; and performing dimensionality transformation on the fusion features obtained by each of the multiple sub-models according to the image ratio of the initial image to obtain initial feature maps corresponding to each of the multiple fusion features.

[0024] In one possible implementation, the region adjustment unit, based on the plurality of initial feature maps, extracts the initial fill region corresponding to each of the at least one prompt word in the initial image. Specifically, this is done by: determining an index group for each prompt word based on its position in the text prompt; the index group being used to identify the region corresponding to the prompt word in the initial feature map; fusing the plurality of initial feature maps based on a set intermediate feature map size to obtain an intermediate feature map; splitting the intermediate feature map based on the index group of each prompt word in the at least one prompt word to obtain a region feature map corresponding to each prompt word in the at least one prompt word; and performing region extraction on the initial image based on the obtained at least one region feature map to obtain the initial fill region corresponding to each region feature map in the at least one region feature map.

[0025] In one possible implementation, the initial feature map includes: regional features of multiple feature channels; each feature channel corresponds to a prompt word; the region adjustment unit fuses the multiple initial feature maps based on a set intermediate feature map size to obtain an intermediate feature map, specifically for: for the at least one prompt word, respectively performing: extracting the regional features of the feature channel corresponding to a prompt word from the multiple initial feature maps, and fusing the extracted multiple regional features based on the set intermediate feature map size to obtain the fused regional features corresponding to the prompt word; and obtaining an intermediate feature map based on the fused regional features corresponding to each prompt word in the at least one prompt word.

[0026] In one possible implementation, the region adjustment unit splits the intermediate feature map based on the index group of each of the at least one prompt words to obtain a region feature map corresponding to each of the at least one prompt words. Specifically, it is used to perform the following operations for each of the at least one prompt words: based on the index group of a prompt word, determine the fused region feature corresponding to the prompt word in the intermediate feature map; and use the fused region feature corresponding to the prompt word as the region feature map corresponding to the prompt word.

[0027] In one possible implementation, the region adjustment unit performs region extraction on the initial image based on at least one obtained region feature map to obtain an initial fill region corresponding to each region feature map. Specifically, it performs the following operations on the at least one obtained region feature map: binarizes a region feature map and uses pixels whose pixel values ​​meet the filtering criteria as segmentation points; divides at least one connected region in the region feature map based on the positional relationship of the segmentation points; wherein each connected region contains multiple segmentation points, and the distance between each of the multiple segmentation points and at least one other segmentation point belonging to the same connected region is less than a preset threshold; and performs image segmentation on the initial image according to the at least one connected region to obtain the initial fill region corresponding to the region feature map.

[0028] In one possible implementation, the target coloring unit performs target coloring processing on the line art based on the obtained baseline coloring regions, the line art, and the at least one prompt word to generate a target image. Specifically, it performs the following operations for each of the hierarchical sub-models included in the generation model: inputting the line art, the at least one prompt word, and the updated intermediate feature representation from the previous level into the current level sub-model to extract the intermediate feature representation at the current level; obtaining the estimated coloring region corresponding to each prompt word based on the intermediate feature representation at the current level; updating the intermediate feature representation at the current level based on the regional difference between the estimated coloring region corresponding to each prompt word and the corresponding baseline coloring region to obtain the updated intermediate feature representation at the current level; and coloring the line art based on the updated intermediate feature representation output from the last level to generate the target image.

[0029] In one possible implementation, the target coloring unit updates the intermediate feature representation at the current level based on the regional difference between the baseline coloring region and the estimated coloring region corresponding to each of the at least one prompt words, to obtain the updated intermediate feature representation at the current level. Specifically, for each of the at least one prompt word, the following operations are performed: based on the regional difference between the estimated coloring region corresponding to a prompt word and the baseline coloring region corresponding to the prompt word, the gradient of the intermediate feature representation corresponding to the next prompt word at the current level is determined; based on the gradient, the intermediate feature representation at the current level is updated to obtain the updated intermediate feature representation at the current level.

[0030] In one possible implementation, the target coloring unit fills the line art with color based on the updated intermediate feature representation output from the last level, generating the target image. Specifically, it is used to: obtain the estimated coloring region corresponding to each of the at least one prompt word based on the updated intermediate feature representation output from the last level, and use each of the obtained estimated coloring regions as target coloring regions; obtain the color information corresponding to each target coloring region based on the updated intermediate feature representation output from the last level; and fill the corresponding regions in the line art with color according to the color information corresponding to each target coloring region, generating the target image.

[0031] Thirdly, embodiments of this application also provide a computer device, including a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor performs the steps of any of the above-described image coloring methods.

[0032] Fourthly, embodiments of this application also provide a computer-readable storage medium including program code, which, when the program product is run on a computer device, is used to cause the computer device to perform the steps of any of the above-described image coloring methods.

[0033] Fifthly, embodiments of this application also provide a computer program product, including computer instructions, which are executed by a processor using the steps of any of the above-described image coloring methods.

[0034] The beneficial effects of this application are as follows:

[0035] This application provides an image coloring method, apparatus, device, and storage medium. In this method, an initial coloring process is first performed on a line drawing. By extracting multiple initial feature maps at each step of the initial coloring process, the initial coloring regions corresponding to each prompt word in the text prompt within the initial image generated by the initial coloring process can be obtained. This allows relevant personnel to perceive whether there is a deviation between the coloring region determined by the model and the coloring region indicated by the corresponding prompt word. Therefore, based on whether the semantic information of each initial coloring region matches the prompt word, the division method of the initial coloring region can be adjusted to obtain a baseline coloring region that matches the semantic information.

[0036] Furthermore, since the obtained baseline coloring area matches the semantic information of the corresponding text prompts, the coloring area division method can be dynamically adjusted when performing target coloring processing according to the baseline coloring area. This results in a coloring area division method that better matches the semantic information of the prompts. After coloring the line art based on a coloring area division method that better matches the semantic information of the prompts, a target image that better meets the requirements can be obtained, avoiding situations where the generated target image does not match the expectations of the relevant personnel, thereby improving the success rate and accuracy of the generated line art.

[0037] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0038] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0039] Figure 1 This is a schematic diagram of a line drawing in the related technology provided in the embodiments of this application;

[0040] Figure 2 This is a schematic diagram of an application scenario provided by an embodiment of this application;

[0041] Figure 3 An exemplary flowchart of an image color filling method is provided for embodiments of this application;

[0042] Figure 4 A schematic diagram illustrating the process of determining the index group provided in this application embodiment;

[0043] Figure 5 This is a schematic diagram of the initial coloring process provided in the embodiments of this application;

[0044] Figure 6 A flowchart illustrating the initial fill area extraction method provided in this application embodiment;

[0045] Figure 7 This is a schematic diagram of the intermediate feature map determination process provided in the embodiments of this application;

[0046] Figure 8 This is one of the schematic diagrams of the initial color-filled area provided in the embodiments of this application;

[0047] Figure 9 This is one of the schematic diagrams of the initial color-filled area provided in the embodiments of this application;

[0048] Figure 10 This is one of the schematic diagrams of the initial color-filled area provided in the embodiments of this application;

[0049] Figure 11 This is one of the schematic diagrams of the initial color-filled area provided in the embodiments of this application;

[0050] Figure 12 This is a schematic diagram illustrating the region division adjustment provided in an embodiment of this application;

[0051] Figure 13 This is a schematic diagram illustrating the region division adjustment provided in an embodiment of this application;

[0052] Figure 14A An exemplary flowchart of the initial color filling region extraction process in an image color filling method provided in this application embodiment;

[0053] Figure 14B An exemplary flowchart of the target image guidance generation process in an image coloring method provided in this application embodiment;

[0054] Figure 15 This is a schematic diagram of the target image provided in the embodiments of this application;

[0055] Figure 16This is a schematic diagram of the structure of an image coloring device provided in an embodiment of this application;

[0056] Figure 17 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of this application;

[0057] Figure 18 This is a schematic diagram of the hardware structure of another computer device that applies an embodiment of this application. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0059] The design concept of the embodiments of this application is briefly introduced below:

[0060] Currently, when generating line art, the process typically involves using a generative model (e.g., StableDiffusion) to fill in the various areas of the line art with color based on the color filling requirements described in the prompt, thereby generating a more complete target image.

[0061] However, in practical applications, when coloring each coloring area in the line drawing according to the text prompts, the generated target image often does not match the expectations of the relevant staff because there is a discrepancy between the coloring area judged by the model and the coloring area indicated in the corresponding text prompts.

[0062] For example, see Figure 1 This is a schematic diagram of a line drawing in the related technology provided in this application embodiment. Assuming the text prompt is: "brown hair, brown eyes, red dress," that is, the hair area in the line drawing needs to be filled with brown, the eye area with brown, and the dress area with red. Figure 1 After inputting the line art and text prompts shown in (a) into the model, the following can be obtained: Figure 1The target image is shown in (b). In this image, the diagonal lines represent brown, and the mesh pattern represents red. The color filling results show that regions 1, 2, and 3 have color filling deviations. Region 1 misidentifies the girl's arm area as hair and therefore fills it with brown; regions 2 and 3 misidentify parts of the skirt area as the girl's skin and therefore do not fill that area with red. Therefore, the generated target image does not match the semantic information conveyed by the text prompt.

[0063] In related technologies, the above problems are usually solved in the following two ways:

[0064] Method 1 involves randomly selecting an art style through a card draw and then regenerating the target image.

[0065] However, this method is highly random and it is difficult to guarantee the generation effect.

[0066] Method 2 involves using separate text prompts for each coloring area in the line art, generating a corresponding area image for each coloring area, and then merging the generated area images into the target image.

[0067] However, this method makes it difficult to select the specific areas to control in the line art because it cannot observe the internal division of the fill regions within the generated model. Furthermore, this method requires generating multiple region images and then fusing them into a target image, resulting in a slow target image generation speed. Additionally, if there are semantic conflicts in the text prompts of different fill regions during the fusion process, the target image generation may fail.

[0068] In view of this, embodiments of this application provide an image coloring method, apparatus, device, and storage medium. The method includes: acquiring a line drawing and acquiring text prompts set corresponding to the line drawing; using at least one prompt word included in the text prompts to perform initial coloring processing on the line drawing, generating an initial image and multiple initial feature maps; then, based on the multiple initial feature maps, extracting an initial coloring region corresponding to each prompt word in the initial image; and when any initial coloring region does not match the semantic information of the corresponding prompt word, performing region division adjustment on the extracted at least one initial coloring region to obtain a reference coloring region corresponding to each prompt word; and based on the obtained reference coloring regions, the line drawing, and at least one prompt word, performing target coloring processing on the line drawing to generate a target image.

[0069] This method first performs initial coloring on the line art. By extracting multiple initial feature maps from each step of the initial coloring process, the initial coloring regions corresponding to each prompt word in the text prompts within the initial image generated through the initial coloring process can be obtained. This allows relevant personnel to perceive whether there is a discrepancy between the coloring regions determined by the model and the coloring regions indicated by the corresponding prompt words. Therefore, based on whether the semantic information of each initial coloring region matches the prompt words, the division method of the initial coloring regions can be adjusted to obtain a baseline coloring region that matches the semantic information.

[0070] Furthermore, since the obtained baseline coloring area matches the semantic information of the corresponding text prompts, the coloring area division method can be dynamically adjusted when performing target coloring processing according to the baseline coloring area. This results in a coloring area division method that better matches the semantic information of the prompts. After coloring the line art based on a coloring area division method that better matches the semantic information of the prompts, a target image that better meets the requirements can be obtained, avoiding situations where the generated target image does not match the expectations of the relevant personnel, thereby improving the success rate and accuracy of the generated line art.

[0071] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0072] See Figure 2 This is a schematic diagram of an application scenario provided by an embodiment of this application. In this scenario, there are terminal devices 210 and servers 220. The generated model can be deployed in the server 220. The terminal device 210 establishes a communication connection with the server 220 through a wired network or a wireless network.

[0073] Relevant staff can input line art and corresponding text prompts through terminal device 210.

[0074] Server 220 can acquire the line art and corresponding text prompts set on the line art through a communication connection. Then, using at least one prompt word contained in the text prompt, it performs initial coloring processing on the line art, generating an initial image and multiple initial feature maps. Based on the generated multiple initial feature maps, it extracts the initial coloring area corresponding to each prompt word in the initial image. Then, server 220 can send the extracted initial coloring area corresponding to each prompt word to terminal device 210 and display it on the display screen of terminal device 210, so that relevant personnel can determine whether the initial coloring area matches the semantic information of the corresponding prompt word.

[0075] When any initial fill area does not match the semantic information of the corresponding prompt word, server 220 can adjust the region division of at least one extracted initial fill area to obtain a reference fill area corresponding to each prompt word. Finally, based on the obtained reference fill areas, line art, and at least one prompt word, target fill processing is performed on the line art to generate a target image, which is then sent to terminal device 210 and displayed on the display screen of terminal device 210.

[0076] The terminal device 210 in this application embodiment may be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited to these.

[0077] The server 220 in this application embodiment can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. This application does not impose any restrictions on these services.

[0078] It should be noted that, Figure 2 The application scenarios shown are merely illustrative. In another scenario, the generative model can also be deployed on a terminal device. In this case, the terminal device does not need to interact with the server and can directly execute the image coloring method provided in this application embodiment on the line drawing to generate the target image. This application does not limit the device on which the generative model is deployed.

[0079] This application provides an image coloring method that can be applied to electronic devices that deploy generative models, for example... Figure 2 In server 220 shown. See also Figure 3 An exemplary flowchart of an image color filling method is provided for embodiments of this application. The method may include the following steps 301-303:

[0080] Step 301: Obtain the line art and the corresponding text prompts set for the line art.

[0081] The text prompts include at least one prompt word that describes the coloring rules for the line art.

[0082] In one possible implementation, the text prompt can be as follows: Figure 1As shown, the text prompts can be multiple phrases or sentences separated by separators (such as commas, pauses, etc.) to describe the coloring rules of the line art. In addition to describing the coloring rules, the text prompts can also include multiple phrases or sentences separated by separators that describe the artistic style, structural proportions, content, lighting effects, etc. of the target image to be generated. For example, the text prompts could also be: "1girl, solo, long hair, brown hair, brown eyes, smile, red dress", where "1girl" means a girl, "solo" means the structure of the image contains a subject, "long hair" means the girl has long hair, "brown hair" means the girl's hair is brown, "brown eyes" means the girl's eyes are brown, "smile" means the girl is smiling, and "red dress" means the girl is wearing a red dress.

[0083] In another possible implementation, the text prompt can be a long sentence, such as, "A young girl, wearing a white short-sleeved shirt and blue jeans, with light brown hair." When a richer image is needed, the text prompt can also be a paragraph, such as, "The protagonist is a young girl, wearing a casual white short-sleeved shirt and blue jeans. Her hair is light brown with a soft sheen. Her cheeks are slightly flushed, showing a lively air. The background uses warm sunlight tones, with beige and light gray buildings and dark gray pavement. Some green plants on the street are a vibrant green." In this text prompt, in addition to describing the coloring rules of various areas in the line drawing, the text also describes the lighting effects of the hair area with "a soft sheen" and the artistic style of the target image with "showing a lively air." Therefore, richer text prompts can add more detail to the generated target image, making it more visually appealing.

[0084] It should be noted that the above-mentioned text prompts can be described in other languages, such as English, Chinese, French, etc. This application does not limit the language of the text prompts.

[0085] In one possible implementation, after obtaining the text hints set for the corresponding line art, the text hints can be segmented to obtain at least one word. Then, based on the arrangement position of each word in the text hint, the at least one word is divided into at least one hint word. The hint word can be used to describe the coloring rules of a region in the line art. When the text hint also includes descriptions of other aspects of the target image, the hint word can also be used to describe other types of image generation rules for a certain region of the line art besides the coloring rules.

[0086] Specifically, the text prompt is segmented to obtain at least one segment. Based on the position of each segment in the text prompt, the at least one segment is divided into at least one group, where each group can include at least one segment. Then, the at least one segment included in each group can be used as a prompt word, thus obtaining at least one prompt word contained in the text prompt.

[0087] Furthermore, due to the different presentation formats of text prompts, the way words are grouped according to their position in the text prompt can also be different.

[0088] In one example, when the text prompt consists of multiple phrases or words separated by delimiters, and at least one word is divided into at least one prompt word based on the position of each word in the text prompt, at least one word can be divided into at least one group based on the positional relationship between each word and the delimiter, and the phrases or words composed of at least one word included in each group can be used as a prompt word.

[0089] For example, suppose the text prompt is "brown hair, brown eyes, red dress". We can first segment the text prompt to obtain: "brown", "hair", ",", "brown", "eyes", ",", "red", "dress". The "," is the separator. One separator can be used to separate two groups, that is, to divide "brown hair" into one group, "brown eyes" into another, and "red dress" into a third. Using each group as a prompt word, we can obtain three prompt words: "brown hair", "brown eyes", and "red dress".

[0090] In another example, if the text prompt is a long sentence or a paragraph, at least one segment describing the same feature of the line drawing can be determined based on the context of each segment in the text prompt, and the at least one segment describing the same feature of the line drawing can be divided into a group, and the phrases or words included in the group can be used as a prompt word.

[0091] For example, suppose the text prompt is "A young girl, wearing a white short-sleeved shirt and blue jeans, her hair is light brown." Word segmentation of the text prompt yields: "a," "young," "girl," ",," "wearing," "white," "short-sleeved," "shirt," "and," "blue," "jeans," ",," "her," "hair," "is," "light brown," "of," ".". Then, based on the context of each word in the text prompt, we can determine that "a," "young," and "girl" are used to describe the main features of the line drawing; "wearing," "white," "short-sleeved," "shirt," "and," "blue," and "jeans" are used to describe the clothing features of the main line drawing; and "her," "hair," "is," "light brown," and "of" are used to describe the hair features of the main line drawing. Therefore, the text prompt can be divided into three prompt words: "a young girl," "wearing a white short-sleeved shirt and blue jeans," and "her hair is light brown."

[0092] Based on the above scheme, by dividing the text prompts into at least one prompt word, the generation model can better understand the structure and content of the text prompts during the coloring process, thereby making the generated target image more semantically consistent with the text prompts and improving the user experience for relevant staff.

[0093] In some embodiments, in order to make the segmented prompt words more in line with the needs of relevant staff for generating images, after the server divides at least one word into at least one group, it can also display the grouping results in the corresponding display interface of the terminal device. Relevant staff can adjust one or more words between different groups by dragging, or they can add or delete groups by editing.

[0094] For example, taking the above prompt word segmentation results as an example, suppose the grouping results are: "a young girl", "wearing a white short-sleeved shirt and blue jeans", "her hair is light brown". If the relevant staff wants to further refine the clothing features and divide them into upper garment features and lower garment features, they can add a group through the editing operation to get four groups: "a young girl", "wearing a white short-sleeved shirt", "blue jeans", "her hair is light brown".

[0095] In one possible implementation, in order to determine the region features corresponding to each of the at least one prompt words within the generative model, after dividing the text prompt into at least one prompt word, for each of the at least one prompt word, an index group for that prompt word can be determined based on the arrangement order of the various segments included in the prompt word in the text prompt. The index group can include the indexes of each segment.

[0096] Specifically, when determining the index group corresponding to one of the at least one prompt words, an index query can be performed based on the obtained text prompt to determine the index of each word included in the prompt word in the text prompt. By treating the indices of each word included in the prompt word as an index group, the index group corresponding to the prompt word is obtained. Here, the index refers to the order number of the word in the text prompt. In addition, for ease of description, the index group identifier corresponding to the prompt word can also be determined based on the order of the prompt word in the at least one prompt word.

[0097] For example, see Figure 4 This is a schematic diagram illustrating the process of determining the index group provided in an embodiment of this application. Figure 4 As shown, assuming the text prompt is "brown hair, brown eyes, red dress", the grouped prompt words are "brownhair", "brown eyes", and "red dress". After obtaining each prompt word, an index query can be performed on the obtained text prompt to obtain the index corresponding to each word segment: the index for "brown" is 1, the index for "hair" is 2, the index for "," is 3, the index for "brown" is 4, the index for "eyes" is 5, the index for "," is 6, the index for "red" is 7, and the index for "dress" is 8. Then, based on the indices corresponding to the word segments included in each prompt word, we can obtain the index groups corresponding to each prompt word: the index group corresponding to "brown hair" includes indices 1 and 2, and since "brown hair" is the first of the three prompt words, this index group is identified as 1; the index group corresponding to "brown eyes" includes indices 4 and 5, and since "brown eyes" is the second of the three prompt words, this index group is identified as 2; the index group corresponding to "red dress" includes indices 7 and 8, and since "red dress" is the second of the three prompt words, this index group is identified as 3.

[0098] Optionally, to facilitate the determination of the regional features corresponding to each prompt word, a mapping relationship between the index group identifier and the index can be constructed after determining the index group identifier of each prompt word. For example, in the example above, the indices corresponding to index group identifier 1 are 1 and 2, the indices corresponding to index group identifier 2 are 4 and 5, and the indices corresponding to index group identifier 3 are 7 and 8.

[0099] Step 302: Using text prompts, perform initial coloring on the line art to generate an initial image and multiple initial feature maps.

[0100] Each initial feature map contains: region features corresponding to each prompt word in at least one prompt word, extracted during one step of the initial coloring process. Where the initial coloring process is performed by a generative model that contains multiple sub-models with a hierarchical relationship, one step in the initial coloring process that generates the initial feature map can correspond to one sub-model.

[0101] In some embodiments, the region features corresponding to each prompt word can be composed of the region features corresponding to each word segment included in each prompt word. Therefore, when extracting the region features corresponding to each prompt word, the index group identifier corresponding to each prompt word can be determined first according to the arrangement position of each prompt word in the text prompt. Then, the index corresponding to the prompt word can be determined according to the mapping relationship between the pre-built index group identifier and the index. Finally, the region features corresponding to each index can be combined into the region features corresponding to the prompt word.

[0102] For example, since the index group for the prompt word "brown hair" includes indices 1 and 2, the extracted region features corresponding to "brown hair" can be composed of the region features corresponding to index 1 and the region features corresponding to index 2.

[0103] In one possible implementation, step 302 can be performed as follows: For multiple sub-models with hierarchical relationships contained in the generative model, perform the following operations respectively until the initial image is generated: Input the line art, at least one prompt word, and the fused features obtained from the previous level sub-model into the current level sub-model to extract the text features and image features of the current level. Perform feature fusion on the text features and image features to obtain the fused features of the current level.

[0104] Then, after generating the initial image, the fusion features obtained by each of the multiple sub-models can be transformed into two-dimensional feature maps according to the image ratio of the initial image, thereby obtaining the initial feature maps corresponding to each of the multiple fusion features.

[0105] For details, see Figure 5This is a schematic diagram illustrating the initial coloring process provided in an embodiment of this application. For example... Figure 5 As shown, assuming the generative model comprises k sub-models, where k is a positive integer, after inputting at least one prompt word and line art into the generative model, each sub-model performs feature extraction and feature fusion, resulting in k fused features: C1 S1 C2 S2 C3 S3 ... C k Sk In this diagram, S1, S2, S3, ..., Sk, in superscript, represent the latent attention lengths of the k sub-models. Since the latent attention length represents the sequence length or feature dimension processed by the latent attention mechanism in the corresponding sub-model, S1, S2, S3, ..., Sk also represent the lengths of the corresponding fused features. After k model iterations, the generator model can obtain the initial image.

[0106] Assuming the initial image has a height of H and a width of W, then performing dimensionality transformation on each fused feature can ensure that the image ratio of each initial feature map is H*W. Using C1... S1 For example, during dimension transformation, if S1 = H1 * W1, then C1 can be determined. S1 The corresponding initial feature map is C1 H1*W1 Similarly, if S2 = H2 * W2, then C2 can be determined. S2 The corresponding initial feature map is C2 H2*W2 By analogy, we can obtain k initial feature maps: C1 H1*W1 C2 H2*W2 C3 H3*W3 ... C k Hk*Wk .in, Figure 5 The diagonal padding in each initial feature map is used to indicate the region of interest of the corresponding sub-model, and is only an example.

[0107] It should be noted that the potential attention length corresponding to each sub-model can be set when configuring the model. In addition to the k sub-models, the above-mentioned generated model may also include at least one sub-model for performing other processing steps, which is not limited in this application.

[0108] In one example, the generative model described above can be a stable diffusion model. In this case, the initial coloring process can be the process of inputting the line art and text prompts into the stable diffusion model to generate the initial image. Each sub-model can correspond to a layer in the stable diffusion model that can extract fusion features, such as a cross-attention layer, a convolutional layer, or a decoder layer, etc. By caching the fusion features obtained by each sub-model during the initial image generation process and performing dimensionality transformation on each fusion feature, the corresponding initial feature map can be obtained. The obtained initial feature map can be an attention heatmap. It should be noted that the number of sub-models included in the generative model can be set according to actual conditions or experience. For example, this application does not limit the type or structure of the model used.

[0109] Based on the above scheme, by obtaining the fusion features from each sub-model, the region of interest of each sub-model in the generative model regarding the input data can be obtained. Converting the fusion features into an initial feature map allows for visualization of the region of interest of each step in the generative model regarding the input data, thus presenting the model's decision-making process more intuitively.

[0110] Step 303: Based on multiple initial feature maps, extract the initial fill region corresponding to each prompt word in the initial image. When any initial fill region does not match the semantic information of the corresponding prompt word, perform region division adjustment on the extracted initial fill region to obtain the reference fill region corresponding to each prompt word in the at least one prompt word.

[0111] In one possible implementation, see [link to relevant documentation]. Figure 6 This is a flowchart illustrating the initial fill region extraction method provided in this application embodiment. Based on multiple initial feature maps, when extracting the initial fill region corresponding to each prompt word in at least one prompt word in the initial image, it can be done according to... Figure 6 The process shown is executed, and the process includes:

[0112] Step 601: Determine the index group of the corresponding prompt words based on the position of each prompt word in the text prompt, based on the position of each prompt word in at least one prompt word.

[0113] Different index groups can be used to identify the regions corresponding to different prompt words in the initial feature map. The index group corresponding to each prompt word can be determined after determining the various prompt words included in the text prompt; the determination method can be found in [reference needed]. Figure 4 The relevant descriptions in the method embodiments shown will not be repeated here.

[0114] Step 602: Based on the set intermediate feature map size, fuse multiple initial feature maps to obtain an intermediate feature map.

[0115] In some embodiments, each initial feature map may include: region features of multiple feature channels, each feature channel corresponding to a prompt word. Then step 602 can be performed as follows: for at least one prompt word, perform the following operations respectively to obtain the fused region features corresponding to each prompt word in at least one prompt word: extract the region features of the feature channel corresponding to one prompt word from multiple initial feature maps, and fuse the extracted multiple region features based on the set intermediate feature map size to obtain the fused region features corresponding to one prompt word.

[0116] Then, after obtaining the fusion region features corresponding to each prompt word in at least one prompt word, an intermediate feature map can be obtained based on the obtained fusion region features.

[0117] In one possible implementation, before fusing multiple initial feature maps based on a set intermediate feature map size, the sizes of the multiple initial feature maps can be unified according to the set intermediate feature map size. For example, the size of each initial feature map can be converted into the intermediate feature map size through methods such as bilinear interpolation, nearest neighbor interpolation, or hybrid interpolation. The method used can be selected according to the actual situation or experience, and this application does not limit it.

[0118] In one example, to simplify processing and capture global features in the feature map as much as possible, the size of the intermediate feature map can be set to a height of H / 32 and a width of W / 32. This application does not limit the value of the intermediate feature map size.

[0119] Specifically, when obtaining the fused region features corresponding to each prompt word in at least one prompt word, for each prompt word, the region features of the feature channel corresponding to that prompt word can be extracted separately from each initial feature map after unifying the size. Then, for each initial feature map, the feature weights of the region features extracted from that initial feature map are determined according to the feature intensity of that initial feature map. The extracted multiple region features are then weighted and fused to obtain the fused region features corresponding to that prompt word. Finally, the obtained fused region features are fused to obtain the intermediate feature map.

[0120] For example, see Figure 7 This is a schematic diagram illustrating the intermediate feature map determination process provided in an embodiment of this application. Figure 7 As shown, assuming the prompt words include "brown hair," "brown eyes," and "red dress," the initial feature map includes C1. H1 *W1and C2 H2*W2 The size of the intermediate feature map is H / 32*W / 32. First, the size of the initial feature map is unified to H / 32*W / 32, resulting in an initial feature map C1 with unified dimensions. H / 32*W / 32 and C2 H / 32*W / 32 For "brown hair", then in C1 H / 32*W / 32 Extract the region feature D11 corresponding to the feature channel of "brown hair" in C2. H / 32*W / 32 Extract the region feature D12 of the feature channel corresponding to "brownhair". For "brown eyes", then in C1 H / 32*W / 32 Extract the region feature D21 corresponding to the feature channel of "brown eyes" in C2. H / 32*W / 32 Extract the region feature D22 corresponding to the feature channel of "brown eyes". For "red dress", the feature can be extracted in C1. H / 32*W / 32 Extract the region feature D31 corresponding to the feature channel of "red dress" in C2. H / 32*W / 32 Extract the region feature D32 of the feature channel corresponding to "red dress".

[0121] Then according to C1 H1*W1 The feature strength can be used to determine the feature weights of D11, D21, and D31 as max(C1 H1*W1 According to C2 H2*W2 The feature strengths are used to determine the feature weights of D12, D22, and D32 as: max(C2 H2*W2 ). Where max(C1 H1*W1 ) represents the initial feature map C1 H1*W1 Similarly, the maximum element value in C2 is max(C2). H2*W2 ) represents the initial feature map C2 H2*W2 The maximum element value in the array.

[0122] After determining the feature weights of each region, D11 and D12 can be weighted and fused into the fused region feature D1 corresponding to "brown hair"; D21 and D22 can be weighted and fused into the fused region feature D2 corresponding to "brown eyes"; and D31 and D32 can be weighted and fused into the fused region feature D3 corresponding to "red dress".

[0123] Finally, by fusing the obtained D1, D2, and D3, the intermediate feature map C can be obtained. mid H / 32*W / 32 .

[0124] It should be noted that, Figure 7The feature maps in this example are merely illustrative. When the initial feature map is an attention heatmap, the resulting intermediate feature map can also be an attention heatmap.

[0125] Based on the above scheme, since the intermediate feature map is obtained based on the fusion region features corresponding to each prompt word in at least one prompt word, and the fusion region features are obtained by extracting and fusing the region features of the feature channels corresponding to each prompt word from each initial feature map, it is possible to ensure that when obtaining the region feature map corresponding to each prompt word in at least one prompt word in the future, each region feature map can contain the corresponding region features in each initial feature map.

[0126] Optionally, when fusing multiple initial feature maps based on the set intermediate feature map size, the size weight of each initial feature map can be determined according to the original size of each initial feature map and the size of the intermediate feature map. Then, the intermediate feature map is obtained by weighted fusion based on the respective size weights of the multiple initial feature maps.

[0127] For example, for C1 H1*W1 and C2 H2*W2 When performing fusion, it can be based on C1 H1*W1 The dimensions of C1 and the dimensions of the intermediate feature maps can be used to determine C1. H1*W1 The size weights are: min(H1*32 / H, H*32 / H1), according to C2 H2*W2 The dimensions of C2 and the dimensions of the intermediate feature maps can determine the size of C2. H2*W2 The size weights are: min(H2*32 / H, H*32 / H2), and then the intermediate feature map can be obtained after weighted fusion.

[0128] In some embodiments, when extracting the regional features of the feature channel corresponding to each prompt word in each initial feature map after unifying the size, for each prompt word, at least one index included in the index group corresponding to the prompt word can be determined first, then the regional features of the feature channel corresponding to each of the at least one index can be determined, and the regional features of the feature channel corresponding to each of the at least one index can be fused to obtain the regional features of the feature channel corresponding to the prompt word.

[0129] Step 603: Based on the index group of each prompt word in at least one prompt word, split the intermediate feature map to obtain the region feature map corresponding to each prompt word in at least one prompt word.

[0130] In one possible implementation, step 603 can be performed as follows: For at least one prompt word, perform the following operations respectively: Based on the index group of a prompt word, determine the fusion region features corresponding to that prompt word in the intermediate feature map. Then, use the determined fusion region features as the region feature map corresponding to that prompt word.

[0131] In one example, when determining the fusion region features corresponding to a prompt word in an intermediate feature map based on a prompt word's index group, one can first determine at least one index included in the index group corresponding to the prompt word. Then, one can determine the fusion region features of the feature channels corresponding to each of the at least one index in the intermediate feature map, and merge the fusion region features of the feature channels corresponding to each of the at least one index to obtain the fusion region features of the feature channel corresponding to the prompt word. Finally, the determined fusion region features are extracted from the intermediate feature map as the region feature map corresponding to that prompt word.

[0132] For example, if the index group corresponding to "brown hair" includes indices 1 and 2, then in the intermediate feature map, the fusion region features of the feature channel corresponding to index 1 and the fusion region features of the feature channel corresponding to index 2 can be merged to obtain the fusion region features corresponding to "brown hair". Then, the fusion region features corresponding to "brown hair" can be separated from the intermediate feature map to serve as the region feature map corresponding to "brown hair".

[0133] Based on the above scheme, the region feature maps corresponding to each prompt word are separated from the intermediate feature map, which can more intuitively display the region corresponding to each prompt word from the model's perspective. This makes it easier for relevant staff to confirm whether the region matches the semantics of the prompt word, thereby improving the accuracy of the generated target image.

[0134] Furthermore, to facilitate the differentiation of different region feature maps split from the same intermediate feature map, the corresponding index group identifier for each prompt word can be added to the subscript of the intermediate feature map to distinguish the region feature maps corresponding to different prompt words. For example, the index group identifier for "brown hair" is 1, so the region feature map corresponding to "brown hair" can be identified through C. mid1 H / 32*W / 32 This is used to represent it. The index group identifier for "brown eyes" is 2, so the region feature map corresponding to "browneyes" can be represented by C. mid2 H / 32*W / 32 The index group for "red dress" is identified as 3, therefore the feature map of the region corresponding to "red dress" can be represented by C. mid3 H / 32*W / 32 To express.

[0135] It should be noted that when the intermediate feature map is an attention heatmap, the feature maps of each region obtained can also be attention heatmaps.

[0136] Step 604: Based on the obtained at least one region feature map, perform region extraction on the initial image to obtain the initial color-filled region corresponding to each region feature map in the at least one region feature map.

[0137] Based on the above scheme, since the region feature map is obtained by fusing and splitting the initial feature map, and the initial feature map is generated during the initial coloring process of the line drawing using at least one prompt word, the region feature map can reflect the region corresponding to each prompt word from the model's perspective. This allows relevant staff to perceive the generation model's understanding of the line drawing, and then improve the accuracy of the generated target image by adjusting the initial coloring region to match the semantics of the prompt word.

[0138] In one possible implementation, step 604 can be performed as follows: For the obtained at least one region feature map, perform the following operations respectively: binarize the region feature map and use the pixels in the region feature map whose pixel values ​​meet the screening criteria as segmentation points. Then, based on the positional relationship of each segmentation point in the region feature map, divide the region feature map into at least one connected region. Finally, based on the at least one connected region, perform image segmentation on the initial image to obtain the initial fill region corresponding to the region feature map. Each connected region contains multiple segmentation points, and the distance between each segmentation point and at least one other segmentation point belonging to the same connected region is less than a preset threshold.

[0139] In some embodiments, before binarizing a region feature map, it can be normalized so that the pixel value of each pixel in the region feature map is between 0 and 1. Then, when binarizing a region feature map, pixels with pixel values ​​greater than a filtering threshold can be used as segmentation points. For example, the filtering threshold can be 0.5, then pixels with pixel values ​​greater than 0.5 in a region feature map are segmentation points.

[0140] After determining the segmentation points, based on the positional relationship of each segmentation point in a region feature map, at least one connected region is divided in the region feature map. Based on the at least one connected region, the initial image is segmented to obtain the initial fill region corresponding to a region feature map. When obtaining the initial fill region corresponding to a region feature map, a region feature map can be used as the mask input of the Segment Anything Model (SAM), and the determined segmentation points can be used as the points input of the SAM. Then, the SAM is used to extract regions from the initial image to obtain the initial fill region corresponding to the region feature map, which is also the initial fill region corresponding to the corresponding prompt word.

[0141] When determining whether they belong to the same connected region, the preset threshold used can be configured in the object segmentation model based on experience or actual situation. For example, the preset threshold can be adjusted by setting the adjacency rule to 4 adjacencies (top, bottom, left, and right) or 8 adjacencies (including diagonals). This application does not limit this.

[0142] See Figure 8-10 This is one of the schematic diagrams of the initial color-filled area provided in the embodiments of this application. For example... Figure 8 As shown, based on the region feature map corresponding to "brown hair", after extracting the region from the initial image using the above method, the initial fill region corresponding to "brown hair" is obtained. Figure 8 The gray area in the middle; such as Figure 9 As shown, based on the region feature map corresponding to "brown eyes", after extracting the region from the initial image using the above method, the initial fill region corresponding to "brown eyes" is obtained. Figure 9 The gray area in the middle; such as Figure 10 As shown, based on the region feature map corresponding to "red dress", after extracting the region from the initial image using the above method, the initial fill region corresponding to "red dress" is obtained. Figure 10 The gray area in the text.

[0143] In one example, the prompt word could also include "smile," see [link to example]. Figure 11 This is one of the schematic diagrams of the initial fill area provided in the embodiments of this application. After extracting the region of the initial image using the above method, the initial fill area corresponding to "smile" is obtained. Figure 11 The gray area in the text.

[0144] Based on the above scheme, since the resolution of the regional feature maps split in the intermediate feature map is low, the initial image can be extracted using each regional feature map to obtain a more refined initial coloring region. With the higher resolution initial coloring region, the region division when generating the target image in the later stage can be more accurate, thereby improving the success rate of line drawing.

[0145] In one possible implementation, after extracting the initial fill area corresponding to each prompt word in at least one prompt word, the extracted initial fill areas can be sent to the terminal device and displayed on the terminal device's display interface. This allows relevant personnel to determine whether each initial fill area matches the semantic information of the corresponding prompt word, and to adjust the region division of the initial fill areas that do not match the semantic information of the corresponding prompt word, thereby obtaining the baseline fill area corresponding to each prompt word in at least one prompt word.

[0146] For example, see Figures 12-13 This is a schematic diagram of area division adjustment provided in an embodiment of this application. The terminal device can display, as shown below. Figure 8-10 The initial fill area shown is due to Figure 8 The area around the girl's arm was incorrectly designated as the initial fill area for "brown hair," therefore it could be... Figure 12 As shown, the initial fill area corresponding to "brown hair" is divided and adjusted to obtain the base fill area corresponding to "brown hair". The adjustment details can be found in [reference needed]. Figure 12 The area indicated by the dashed arrow.

[0147] because Figure 9 The initial fill area corresponding to "brown eyes" matches the semantic information of "brown eyes," so it can be directly filled in. Figure 9 The initial fill area corresponding to "brown eyes" shown in the figure is used as the base fill area for "brown eyes".

[0148] because Figure 10 In the example, there was an issue where the clothing area was not fully covered. Specifically, there were two initial fill areas that should have been designated for "reddress," but were not. Therefore, it can be corrected as follows: Figure 13 As shown, the initial fill area corresponding to "red dress" is divided and adjusted to obtain the base fill area corresponding to "red dress". The adjustment details can be found in [reference needed]. Figure 13 The area indicated by the dashed arrow.

[0149] When the prompt word includes "smile", because Figure 11 In the text, the initial fill area corresponding to "smile" matches the semantic information of "smile," so it can be directly filled in. Figure 11 The initial fill area corresponding to "smile" shown in the figure is used as the base fill area for "smile".

[0150] For example, each reference fill region can be represented as a mask, and the size of each reference fill region is the same as the initial image. Therefore, the reference fill region of the prompt word identified as i in the index group is M. ci H*W For example, the base fill area for "brown hair" is M. c1 H*W The base fill area corresponding to "brown eyes" is M. c2 H *W The base fill area corresponding to "red dress" is M. c3 H*W .

[0151] Step 304: Based on the obtained baseline coloring areas, line art, and at least one prompt word, perform target coloring on the line art to generate the target image.

[0152] In one possible implementation, for multiple hierarchical sub-models within the generative model, the following operations are performed to obtain the updated intermediate feature representation of the last level's output: The line art, at least one prompt word, and the updated intermediate feature representation from the previous level are input into the current level's sub-model to extract the intermediate feature representation at the current level. Based on the intermediate feature representation at the current level, the estimated fill area corresponding to each prompt word is obtained. Based on the regional difference between the estimated fill area corresponding to each prompt word and the corresponding baseline fill area, the intermediate feature representation at the current level is updated to obtain the updated intermediate feature representation at the current level.

[0153] Then, after obtaining the updated intermediate feature representation of the last level output, the line art can be filled with color based on this intermediate feature representation to generate the target image.

[0154] Based on the above scheme, by dynamically adjusting the intermediate feature representation, the effect of the corresponding prompt words in the corresponding baseline coloring area can be improved during the image generation process, thereby generating a target image that is more consistent with the semantics of the prompt words.

[0155] The intermediate feature representation obtained at each level can be a latent space vector. After extracting the intermediate feature representation at the current level, a dimensionality transformation can be performed to obtain the attention heatmap at the current level. Based on the mask size of the baseline fill region, a dimensionality transformation is also performed to obtain the attention heatmap at the current level. Then, the attention heatmap is normalized in the cue word dimension using a softmax operation to obtain the estimated fill region corresponding to each cue word in at least one cue word at the current level. For example, assuming the current level is level t and the text cue includes Np cue words, the estimated fill region corresponding to the cue word with index group identifier i can be represented as C. si Ht*Wt*Np .

[0156] In some embodiments, based on the regional difference between the estimated fill region corresponding to each of the at least one prompt words and the corresponding baseline fill region, the intermediate feature representation at the current level is updated to obtain the updated intermediate feature representation at the current level. For at least one prompt word, the following operations are performed respectively: Based on the regional difference between the estimated fill region corresponding to a prompt word and the baseline fill region corresponding to a prompt word, the gradient of the intermediate feature representation corresponding to a prompt word at the current level is determined. Based on the gradient, the intermediate feature representation at the current level is updated to obtain the updated intermediate feature representation at the current level.

[0157] Specifically, the regional difference between the estimated fill area corresponding to a prompt word identified as i in the index group and the baseline fill area corresponding to that prompt word can satisfy formula (1).

[0158]

[0159] In the formula, Loss represents the regional difference, t represents the current level, Ht*Wt represents the size of the current estimated fill area, the size ratio is equal to H*W, and i represents the identifier of an index group of a prompt word. This indicates the estimated fill area corresponding to a prompt word. This represents the base fill area corresponding to a prompt word, and max() indicates taking the maximum value.

[0160] Based on the regional differences obtained by formula (1), the gradient G of the intermediate feature representation corresponding to a prompt word can be obtained. Then, as shown in formula (2), the intermediate feature representation at the current level is updated:

[0161] latents′ t =latents t -λ*G Formula (2)

[0162] In the formula, latents′ tLatents represents the updated intermediate feature representation at the current level. t This represents the intermediate feature representation extracted at the current level, λ represents the learning rate, and G represents the gradient of the intermediate feature representation. The learning rate can be set according to actual conditions or experience, for example, it can be 0.001, etc., and this application does not limit it in this way.

[0163] Based on the above scheme, by updating the intermediate feature representation, we can avoid the situation in related technologies where attention heatmaps are directly controlled without taking into account the complex relationships in the latent space. This can avoid the problem of failed image generation caused by this situation and thus improve the success rate of line drawing image generation.

[0164] In some embodiments, while updating the intermediate feature representation at the current level, the estimated fill region at the current level can also be updated by performing image enhancement on the estimated fill region corresponding to each of the at least one prompt words, and the updated estimated fill region at the current level can be input into the sub-model at the next level.

[0165] Specifically, image enhancement of the estimated fill area corresponding to a prompt word identified as i in the index group can be performed as shown in formula (3):

[0166]

[0167] In the formula, This indicates the updated estimated fill area at the current level corresponding to a prompt word. This represents the estimated fill area at the current level corresponding to a prompt word; β is the enhancement coefficient, satisfying...

[0168] In some embodiments, when filling the line art with color based on the updated intermediate feature representation output from the last level to generate the target image, the estimated filling region corresponding to each of the at least one prompt word can be obtained based on the updated intermediate feature representation output from the last level, and each of the obtained estimated filling regions can be used as the target filling region. Color information corresponding to each target filling region is obtained based on the updated intermediate feature representation output from the last level. The corresponding regions in the line art are filled with color according to the color information corresponding to each target filling region to generate the target image.

[0169] Based on the above scheme, as the regional differences between the estimated fill area and the baseline fill area are gradually reduced during the iterative inference process of multiple sub-models, the estimated fill area obtained through the updated intermediate feature representation output by the last level can better match the semantic information of the prompt words. This results in a better match between the obtained target fill area and the semantic information of the prompt words. With a better match between the target fill area and the semantic information of the prompt words, the target image generated after filling is more in line with the needs of relevant personnel, thereby improving the overall success rate of line drawing images.

[0170] It should be noted that the hierarchical sub-models included in the generative model used to perform the target coloring process may be the same as or different from the hierarchical sub-models included in the generative model used to perform the initial coloring process. This application does not impose any restrictions on this.

[0171] Below, in order to more clearly understand the solution proposed in the embodiments of this application, an image coloring method provided by this application will be introduced in conjunction with specific embodiments.

[0172] The image coloring method provided in this application embodiment can be divided into an initial coloring area extraction process and a target image guided generation process.

[0173] See Figure 14A This is an exemplary flowchart of the initial color filling region extraction process in an image color filling method provided in this application embodiment, specifically including:

[0174] Step 1401a: Obtain the line art.

[0175] Step 1401b: Obtain text hints.

[0176] Among them, 1401a and 1401b can be executed simultaneously, or 1401a can be executed first and then 1401b can be executed, or 1401b can be executed first and then 1401a can be executed.

[0177] Step 1402: Divide the text prompt into at least one prompt word.

[0178] Step 1403: Determine the index group corresponding to each prompt word.

[0179] For each of the at least one suggestion words, perform an index query to determine the index group corresponding to each suggestion word.

[0180] Step 1404: Initial color filling process.

[0181] Step 1405: Extract the region feature map corresponding to each prompt word.

[0182] During the initial coloring process, the region feature map corresponding to each prompt word is extracted based on the index group corresponding to each prompt word.

[0183] Step 1406: Generate the initial image.

[0184] Step 1407: Obtain the initial fill area corresponding to each prompt word.

[0185] The feature map of the region corresponding to each prompt word is refined based on the initial image to obtain the initial coloring region corresponding to each prompt word.

[0186] See Figure 14B This is an exemplary flowchart of the target image guidance generation process in an image coloring method provided in this application embodiment. The process includes:

[0187] Step 1408: Obtain the base color area corresponding to each prompt word based on the initial color area corresponding to each prompt word.

[0188] After obtaining the initial fill area corresponding to each prompt word, if the initial fill area corresponding to any prompt word does not match the semantic information of that prompt word, then proceed to step 1408. If the initial fill area corresponding to each prompt word matches the semantic information of that prompt word, then no further action is required. Figure 14B The process shown allows you to use the initial image as the target image.

[0189] Step 1409: Perform target coloring on the line art based on the index group corresponding to each prompt word.

[0190] Step 1410: Extract intermediate feature representations at each level and obtain the estimated fill area corresponding to each prompt word.

[0191] During the target coloring process, intermediate feature representations at each level can be extracted, and the estimated coloring area corresponding to each prompt word can be obtained based on the intermediate feature identifiers.

[0192] Step 1411: Dynamically adjust the intermediate feature representation and estimate the fill area based on the baseline fill area.

[0193] Step 1412: Generate the target image.

[0194] The above Figure 14A and Figure 14B For the specific implementation methods of each step included, please refer to [link / reference]. Figure 3 The relevant descriptions in the method embodiments shown will not be repeated here.

[0195] See Figure 15This is a schematic diagram of the target image provided for an embodiment of this application. For example... Figure 15 As shown, using the above image coloring method, an initial image can be obtained by first performing initial coloring processing on the line drawing using text descriptions. Since the clothing coverage in the initial image is incomplete, and some arm areas are mistakenly identified as hair areas, the resulting initial image does not match the expectations of the relevant personnel. Therefore, further target coloring processing can be used to dynamically adjust the areas that do not meet expectations, thereby obtaining a target image with accurately defined coloring areas. The adjustment process can be found in [reference needed]. Figure 15 The part indicated by the dashed arrow.

[0196] Based on the above scheme, since the reference color area is used, the regional distribution in the generated model is guided. According to statistics, after implementing the line drawing generation method provided in this application embodiment, the overall output success rate of the line drawing generation can be increased from 70% to 90% without destroying the overall composition of the line drawing.

[0197] Based on the same inventive concept as the above-described method embodiments, this application also provides an image coloring device. For example... Figure 16 The image coloring device 1600 shown may include:

[0198] The communication unit 1601 is used to acquire line art and to acquire text prompts corresponding to the line art settings. The text prompts include at least one prompt word for describing the coloring rules of the line art.

[0199] The initial coloring unit 1602 is used to perform initial coloring processing on the line drawing using the text prompts, generating an initial image and multiple initial feature maps; each initial feature map includes: region features corresponding to each prompt word in the at least one prompt word extracted during a step of performing the initial coloring processing;

[0200] The region adjustment unit 1603 is used to extract the initial coloring region corresponding to each prompt word in the initial image based on the plurality of initial feature maps, and to perform region division adjustment on the extracted at least one initial coloring region when any initial coloring region does not match the semantic information of the corresponding prompt word, so as to obtain the reference coloring region corresponding to each prompt word in the at least one prompt word.

[0201] The target coloring unit 1604 is used to perform target coloring processing on the line drawing based on the obtained reference coloring areas, the line drawing, and the at least one prompt word, to generate a target image.

[0202] In one possible implementation, after the communication unit 1601 obtains the text prompt corresponding to the line drawing settings, the prompt word segmentation unit 1605 is used to: segment the text prompt to obtain at least one word; divide the at least one word into at least one prompt word according to the position of each word in the text prompt; each prompt word is used to describe the coloring rules of a region in the line drawing.

[0203] In one possible implementation, the initial coloring unit 1602 uses the text prompt to perform initial coloring processing on the line art, generating an initial image and multiple initial feature maps. Specifically, it is used to perform the following operations on multiple sub-models with hierarchical relationships included in the generation model, until the initial image is generated: inputting the line art, the at least one prompt word, and the fusion features obtained by the sub-model at the previous level into the current level sub-model, extracting text features and image features at the current level; performing feature fusion on the text features and the image features to obtain the fusion features at the current level; and performing dimensionality transformation on the fusion features obtained by each of the multiple sub-models according to the image ratio of the initial image to obtain the initial feature maps corresponding to each of the multiple fusion features.

[0204] In one possible implementation, the region adjustment unit 1603, based on the plurality of initial feature maps, extracts the initial fill region corresponding to each of the at least one prompt word in the initial image. Specifically, it is used to: determine the index group of the corresponding prompt word based on the position of each of the at least one prompt word in the text prompt; the index group is used to identify the region corresponding to the corresponding prompt word in the initial feature map; fuse the plurality of initial feature maps based on a set intermediate feature map size to obtain an intermediate feature map; split the intermediate feature map based on the index group of each of the at least one prompt word to obtain the region feature map corresponding to each of the at least one prompt word; and extract regions from the initial image based on the obtained at least one region feature map to obtain the initial fill region corresponding to each region feature map in the at least one region feature map.

[0205] In one possible implementation, the initial feature map includes: regional features of multiple feature channels; each feature channel corresponds to a prompt word; the region adjustment unit 1603 fuses the multiple initial feature maps based on a set intermediate feature map size to obtain an intermediate feature map, specifically for: for the at least one prompt word, respectively performing: extracting the regional features of the feature channel corresponding to a prompt word from the multiple initial feature maps, and fusing the extracted multiple regional features based on the set intermediate feature map size to obtain the fused regional features corresponding to the prompt word; and obtaining an intermediate feature map based on the fused regional features corresponding to each prompt word in the at least one prompt word.

[0206] In one possible implementation, the region adjustment unit 1603 splits the intermediate feature map based on the index group of each of the at least one prompt words to obtain a region feature map corresponding to each of the at least one prompt words. Specifically, it is used to perform the following operations for each of the at least one prompt words: based on the index group of a prompt word, determine the fused region feature corresponding to the prompt word in the intermediate feature map; and use the fused region feature corresponding to the prompt word as the region feature map corresponding to the prompt word.

[0207] In one possible implementation, the region adjustment unit 1603 performs region extraction on the initial image based on at least one obtained region feature map to obtain an initial fill region corresponding to each region feature map in the at least one region feature map. Specifically, it performs the following operations on the at least one obtained region feature map: binarizes a region feature map and uses pixels whose pixel values ​​meet the filtering conditions as segmentation points; divides at least one connected region in the region feature map based on the positional relationship of each segmentation point in the region feature map; wherein each connected region contains multiple segmentation points, and the distance between each of the multiple segmentation points and at least one other segmentation point belonging to the same connected region is less than a preset threshold; and performs image segmentation on the initial image according to the at least one connected region to obtain the initial fill region corresponding to the region feature map.

[0208] In one possible implementation, the target coloring unit 1604 performs target coloring processing on the line art based on the obtained baseline coloring regions, the line art, and the at least one prompt word to generate a target image. Specifically, it performs the following operations for each of the multiple sub-models with hierarchical relationships included in the generation model: inputting the line art, the at least one prompt word, and the updated intermediate feature representation from the previous level into the current level sub-model to extract the intermediate feature representation at the current level; obtaining the estimated coloring region corresponding to each prompt word based on the intermediate feature representation at the current level; updating the intermediate feature representation at the current level based on the regional difference between the estimated coloring region corresponding to each prompt word and the corresponding baseline coloring region to obtain the updated intermediate feature representation at the current level; and coloring the line art based on the updated intermediate feature representation output from the last level to generate the target image.

[0209] In one possible implementation, the target coloring unit 1604 updates the intermediate feature representation at the current level based on the regional difference between the baseline coloring region and the estimated coloring region corresponding to each of the at least one prompt words, to obtain the updated intermediate feature representation at the current level. Specifically, for each of the at least one prompt words, the following operations are performed: based on the regional difference between the estimated coloring region corresponding to a prompt word and the baseline coloring region corresponding to the prompt word, the gradient of the intermediate feature representation corresponding to the next prompt word at the current level is determined; based on the gradient, the intermediate feature representation at the current level is updated to obtain the updated intermediate feature representation at the current level.

[0210] In one possible implementation, the target coloring unit 1604 fills the line drawing with color based on the updated intermediate feature representation output from the last level to generate the target image. Specifically, it is used to: obtain the estimated coloring region corresponding to each of the at least one prompt words based on the updated intermediate feature representation output from the last level, and use each of the obtained estimated coloring regions as target coloring regions; obtain the color information corresponding to each target coloring region based on the updated intermediate feature representation output from the last level; and fill the corresponding regions in the line drawing with color according to the color information corresponding to each target coloring region to generate the target image.

[0211] For ease of description, the above sections are divided into modules (or units) according to their functions and described separately. In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program with a predetermined function, which works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functions of that module or unit.

[0212] Having described the image coloring method and apparatus according to exemplary embodiments of this application, we will now describe a computer device according to another exemplary embodiment of this application.

[0213] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."

[0214] Based on the same inventive concept as the above-described method embodiments, this application also provides a computer device. In one embodiment, the computer device may be a server, such as... Figure 2 The server 220 is shown. In this embodiment, the computer device 1700 has the following structure: Figure 17 As shown, it may include at least a memory 1701, a communication module 1703, and at least one processor 1702.

[0215] The memory 1701 is used to store computer programs executed by the processor 1702. The memory 1701 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0216] Memory 1701 may be volatile memory, such as random-access memory (RAM); memory 1701 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 1701 may be any other medium capable of carrying or storing a desired computer program having the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1701 may be a combination of the above-described memories.

[0217] Processor 1702 may include one or more central processing units (CPUs) or digital processing units, etc. Processor 1702 is used to implement the above-described image coloring method when it calls the computer program stored in memory 1701.

[0218] The communication module 1703 is used to communicate with terminal devices and other servers.

[0219] This application embodiment does not limit the specific connection medium between the memory 1701, communication module 1703, and processor 1702. This application embodiment... Figure 17 The memory 1701 and the processor 1702 are connected via a bus 1704, and the bus 1704 is in Figure 17 The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. The 1704 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 17 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.

[0220] The memory 1701 stores a computer storage medium, which stores computer-executable instructions for implementing the image coloring method of this application embodiment. The processor 1702 is used to execute the above-described image coloring method, such as... Figure 3 As shown.

[0221] In another embodiment, the computer device can also be other computer devices, such as... Figure 2 The terminal device 210 is shown. In this embodiment, the structure of the computer device can be as follows: Figure 18As shown, it includes components such as: communication component 1810, memory 1820, display unit 1830, camera 1840, sensor 1850, audio circuit 1860, Bluetooth module 1870, processor 1880, etc.

[0222] The communication component 1810 is used to communicate with the server. In some embodiments, it may include a Wireless Fidelity (WiFi) module, which is a short-range wireless transmission technology, and the electronic device can send and receive information through the WiFi module.

[0223] The memory 1820 can be used to store software programs and data. The processor 1880 executes various functions of the terminal device 210 and performs data processing by running the software programs or data stored in the memory 1820. The memory 1820 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The memory 1820 stores an operating system that enables the terminal device 210 to run. In this application, the memory 1820 may store the operating system and various application programs, and may also store a computer program that executes the image coloring method of the embodiments of this application.

[0224] The display unit 1830 can also be used to display information input by an object or information provided to an object, as well as a graphical user interface (GUI) for various menus of the terminal device 210. Specifically, the display unit 1830 may include a display screen 1832 disposed on the front of the terminal device 210. The display screen 1832 may be configured as a liquid crystal display, a light-emitting diode, or the like. The display unit 1830 can be used to display the initial coloring area, etc., in the embodiments of this application.

[0225] The display unit 1830 can also be used to receive input digital or character information and generate signal inputs related to object settings and function control of the terminal device 210. Specifically, the display unit 1830 may include a touch screen 1831 disposed on the front of the terminal device 210, which can collect touch operations on or near the object, such as clicking a button, dragging a scroll bar, etc.

[0226] The touchscreen 1831 can be placed on top of the display screen 1832, or the touchscreen 1831 and the display screen 1832 can be integrated to realize the input and output functions of the terminal device 210. After integration, it can be referred to as a touch display screen. In this application, the display unit 1830 can display the application and the corresponding operation steps.

[0227] Camera 1840 can be used to capture still images, and objects can publish images captured by camera 1840 through an application. There can be one or multiple cameras 1840. An optical image of an object is generated through a lens and projected onto a photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to processor 1880 to be converted into a digital image signal.

[0228] The terminal device may also include at least one sensor 1850, such as an accelerometer 1851, a proximity sensor 1852, a fingerprint sensor 1853, and a temperature sensor 1854. The terminal device may also be equipped with other sensors such as a gyroscope, barometer, hygrometer, thermometer, infrared sensor, light sensor, and motion sensor.

[0229] Audio circuitry 1860, speaker 1861, and microphone 1862 provide an audio interface between the device and terminal device 210. Audio circuitry 1860 converts received audio data into electrical signals, which are then transmitted to speaker 1861, where they are converted into sound signals for output. Terminal device 210 may also be equipped with volume buttons for adjusting the volume of the sound signal. On the other hand, microphone 1862 converts collected sound signals into electrical signals, which are then received by audio circuitry 1860, converted back into audio data, and output to communication component 1810 for transmission to, for example, another terminal device 210, or to memory 1820 for further processing.

[0230] The Bluetooth module 1870 is used to interact with other Bluetooth devices that also have a Bluetooth module via the Bluetooth protocol. For example, a terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smartwatch) that also has a Bluetooth module through the Bluetooth module 1870, thereby exchanging data.

[0231] The processor 1880 is the control center of the terminal device, connecting various parts of the terminal through various interfaces and lines. It executes various functions and processes data by running or executing software programs stored in the memory 1820 and calling data stored in the memory 1820. In some embodiments, the processor 1880 may include one or more processing units; the processor 1880 may also integrate an application processor and a baseband processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the baseband processor mainly handles wireless communication. It is understood that the baseband processor may not be integrated into the processor 1880. In this application, the processor 1880 can run the operating system, applications, user interface display and touch response, as well as the image coloring method of the embodiments of this application. Furthermore, the processor 1880 is coupled to the display unit 1830.

[0232] Furthermore, it should be noted that in the specific embodiments of this application, object data related to image coloring is involved. When the above embodiments of this application are applied to specific products or technologies, permission or consent from the object is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0233] In some possible implementations, various aspects of the image coloring method provided in this application can also be implemented as a program product, which includes a computer program. When the program product is run on a computer device, the computer program causes the computer device to perform the steps of the image coloring method according to the various exemplary embodiments of this application described above. For example, the computer device can perform actions such as... Figure 3 The steps are shown in the figure.

[0234] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0235] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on an electronic device. However, the program product of this application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with a command execution system, apparatus, or device.

[0236] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.

[0237] Computer programs contained on readable media may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0238] Computer programs for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The computer program can execute entirely on the user's computer device, partially on the user's computer device, as a standalone software package, partially on the user's computer device and partially on a remote computer device, or entirely on a remote computer device. In cases involving remote computer devices, the remote computer device can be connected to the user's computer device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer device (e.g., via the Internet using an Internet service provider).

[0239] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0240] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0241] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing a computer-usable computer program.

[0242] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0243] These computer program commands may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the commands stored in the computer-readable storage medium produce an article of manufacture including command means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0244] These computer program commands can also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing the commands executed on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0245] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0246] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. An image coloring method, characterized in that, include: Obtain the line art and obtain the text prompts corresponding to the line art settings, the text prompts including: at least one prompt word for describing the coloring rules of the line art; Using the text prompts, the line drawing is initially filled with color to generate an initial image and multiple initial feature maps; each initial feature map contains: the region features corresponding to each prompt word in the at least one prompt word, extracted during one step of the initial color filling process; Based on the multiple initial feature maps, in the initial image, an initial coloring region corresponding to each of the at least one prompt words is extracted. When any of the initial coloring regions does not match the semantic information of the corresponding prompt word, the extracted at least one initial coloring region is divided and adjusted to obtain the reference coloring region corresponding to each of the at least one prompt words. Based on the obtained baseline coloring areas, the line art, and the at least one prompt word, target coloring processing is performed on the line art to generate a target image.

2. The method according to claim 1, characterized in that, After obtaining the text prompt corresponding to the line art settings, the method further includes: The text prompt is segmented to obtain at least one word; Based on the arrangement position of each segment in the text prompt, the at least one segment is divided into at least one prompt word; each prompt word is used to describe the coloring rules of a region in the line drawing; For each of the at least one prompt word, perform the following operations: determine the index group of the prompt word based on the order of the various segments included in the prompt word in the text prompt.

3. The method according to claim 1 or 2, characterized in that, The process of using the text prompts to perform initial coloring on the line art, generating an initial image and multiple initial feature maps, includes: For each of the hierarchical sub-models contained in the generative model, perform the following operations until the initial image is generated: The line drawing, the at least one prompt word, and the fusion features obtained from the previous level sub-model are input into the current level sub-model to extract the text features and image features of the current level. The text features and the image features are fused to obtain the fused features at the current level; According to the image ratio of the initial image, the fusion features obtained by each of the multiple sub-models are dimensionally transformed to obtain the initial feature maps corresponding to each of the multiple fusion features.

4. The method according to claim 1 or 2, characterized in that, Based on the multiple initial feature maps, extracting the initial fill region corresponding to each of the at least one prompt word in the initial image includes: Based on the arrangement position of each of the at least one prompting word in the text prompt, an index group for the corresponding prompting word is determined; the index group is used to identify the region corresponding to the corresponding prompting word in the initial feature map; Based on the set intermediate feature map size, the multiple initial feature maps are fused to obtain an intermediate feature map; Based on the index group of each of the at least one prompt words, the intermediate feature map is split to obtain the region feature map corresponding to each of the at least one prompt words; Based on the obtained at least one region feature map, region extraction is performed on the initial image to obtain the initial color-filled region corresponding to each region feature map in the at least one region feature map.

5. The method according to claim 4, characterized in that, The initial feature map includes: regional features of multiple feature channels; each feature channel corresponds to a prompt word; The process of fusing the multiple initial feature maps based on a set intermediate feature map size to obtain an intermediate feature map includes: For each of the at least one prompt word, the following steps are performed: extracting the region features of the feature channel corresponding to the prompt word from the plurality of initial feature maps, and fusing the extracted region features based on the set intermediate feature map size to obtain the fused region features corresponding to the prompt word. An intermediate feature map is obtained based on the fusion region features corresponding to each of the at least one prompt words.

6. The method according to claim 5, characterized in that, The step of splitting the intermediate feature map based on the index group of each of the at least one prompt words to obtain the region feature map corresponding to each of the at least one prompt words includes: For each of the at least one prompt word, perform the following operations: Based on an index group of a prompt word, determine the fusion region features corresponding to the prompt word in the intermediate feature map; The fusion region features corresponding to the prompt word are used as the region feature map corresponding to the prompt word.

7. The method according to claim 4, characterized in that, Based on the obtained at least one region feature map, region extraction is performed on the initial image to obtain the initial fill region corresponding to each region feature map in the at least one region feature map, including: For at least one obtained region feature map, perform the following operations respectively: A region feature map is binarized, and pixels whose pixel values ​​in the region feature map meet the screening criteria are used as segmentation points. Based on the positional relationship of each segmentation point in the feature map of a region, at least one connected region is divided in the feature map of a region; wherein, each connected region contains multiple segmentation points, and the distance between each of the multiple segmentation points and at least one other segmentation point belonging to the same connected region is less than a preset threshold. Based on the at least one connected region, the initial image is segmented to obtain the initial fill region corresponding to the feature map of the region.

8. The method according to claim 1 or 2, characterized in that, The process of performing target coloring on the line art based on the obtained baseline coloring regions, the line art, and the at least one prompt word to generate a target image includes: For each of the hierarchical sub-models contained in the generative model, perform the following operations to obtain the updated intermediate feature representation of the last level output: Input the line drawing, the at least one prompt word, and the updated intermediate feature representation from the previous level into the sub-model of the current level to extract the intermediate feature representation from the current level. Based on the intermediate feature representation at the current level, the estimated fill area corresponding to each of the at least one prompt word is obtained; Based on the regional difference between the estimated fill area corresponding to each of the at least one prompt words and the corresponding baseline fill area, the intermediate feature representation at the current level is updated to obtain the updated intermediate feature representation at the current level. Based on the updated intermediate feature representation output from the last level, the line drawing is filled with color to generate the target image.

9. The method according to claim 8, characterized in that, The step of updating the intermediate feature representation at the current level based on the regional difference between the estimated fill area corresponding to each of the at least one prompt words and the corresponding baseline fill area, to obtain the updated intermediate feature representation at the current level, includes: For each of the at least one prompt word, perform the following operations: Based on the regional difference between the estimated fill area corresponding to a prompt word and the baseline fill area corresponding to the prompt word, the gradient of the intermediate feature representation corresponding to the next prompt word at the current level is determined; Based on the gradient, the intermediate feature representation at the current level is updated to obtain the updated intermediate feature representation at the current level.

10. The method according to claim 8, characterized in that, The process of filling the line art with color based on the updated intermediate feature representation output from the last level to generate the target image includes: Based on the updated intermediate feature representation of the last level output, the estimated coloring region corresponding to each of the at least one prompt words is obtained, and each of the obtained estimated coloring regions is used as the target coloring region. Based on the updated intermediate feature representation output from the last level, the color information corresponding to each target color-filled area is obtained; According to the color information corresponding to each target coloring area, the corresponding areas in the line drawing are filled with color to generate the target image.

11. An image processing apparatus, characterized in that, include: A communication unit is used to acquire a line drawing and to acquire text prompts corresponding to the line drawing, wherein the text prompts include at least one prompt word describing the coloring rules of the line drawing; An initial coloring unit is used to perform initial coloring processing on the line drawing using the text prompts, generating an initial image and multiple initial feature maps; each initial feature map includes: region features corresponding to each prompt word in the at least one prompt word extracted during a step of performing the initial coloring processing; The region adjustment unit is used to extract the initial coloring region corresponding to each prompt word in the initial image based on the plurality of initial feature maps, and to perform region division adjustment on the extracted at least one initial coloring region when any initial coloring region does not match the semantic information of the corresponding prompt word, so as to obtain the reference coloring region corresponding to each prompt word in the at least one prompt word. The target coloring unit is used to perform target coloring processing on the line drawing based on the obtained reference coloring areas, the line drawing, and the at least one prompt word, to generate a target image.

12. A computer device, characterized in that, It includes a processor and a memory, wherein the memory stores program code that, when executed by the processor, causes the processor to perform the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, It includes program code that, when run on a computer device, causes the computer device to perform the steps of the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.