Image processing method, system and apparatus, and electronic device
By acquiring image style information and reference images, and combining the line style control module and structure preservation module, diverse line art images are generated, solving the problem of monotonous line art styles in existing technologies and improving the fun and interactivity of the user experience.
Patent Information
- Application Number
- PCT/CN2025/087508
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-04-07
- Publication Date
- 2026-02-05
AI Technical Summary
Existing technologies cannot incorporate style when generating line art, resulting in line art with a monotonous style that lacks interest and diversity.
By acquiring image style information and combining it with the first image and the reference image, a line drawing image with a specific style is generated. The line style control module and the structure preservation module are used to ensure the aesthetics and structural consistency of the line drawing.
It achieves a diversity of line art styles, increases the fun and interactivity of users generating line art using electronic devices, and is suitable for a wide range of painting enthusiasts to get started quickly.
Smart Images

Figure CN2025087508_05022026_PF_FP_ABST
Abstract
Description
Image processing methods, systems, devices and electronic equipment Technical Field
[0001] This application relates to the field of data processing, and more particularly to an image processing method, system, apparatus, and electronic device. Background Technology
[0002] Line art primarily uses the arrangement and combination of lines to depict the outline, structure, and details of objects. Line art does not contain color; it relies solely on variations in line thickness, curvature, density, and intersections to express the form and spatial relationships of objects. Line art has wide applications in various fields, such as pre-painting artwork, animation and comic production, image restoration, and character design in games.
[0003] Currently, existing technologies generate line art by processing images through editing software, such as desaturating, inverting, and merging; or by inputting images into a generator, which then outputs line art. However, these existing technologies have a significant limitation in the process of generating line art: they cannot incorporate style; this results in line art with a monotonous style, lacking sufficient interest and diversity. Summary of the Invention
[0004] In view of this, this application provides an image processing method, system, apparatus, and electronic device. This image processing method can generate line art images in various styles, achieving diversification of line art styles and increasing interest and interactivity.
[0005] In a first aspect, this application provides an image processing method, which includes: firstly, acquiring a first image; then, acquiring image style information; and subsequently, acquiring a line drawing image of a target object in the first image based on the first image and the image style information.
[0006] The image style information can be the image style information corresponding to the user-set image style, the image style information corresponding to the image style set by the electronic device's system, or the image style information obtained from other electronic devices. Thus, during the line art generation process, the user-set (or system-set, or image style information obtained from other electronic devices) image style can be injected to generate a line art image for the target object in the first image that corresponds to the user-set (or system-set, or image style information obtained from other electronic devices) image style. For example, when the user-set image style or the electronic device's system-set image style is image style 1, this application can generate a line art image with image style 1; when the user-set image style or the electronic device's system-set image style is image style 2, this application can generate a line art image with image style 2. Therefore, this application can generate line art images of various styles, achieving diversification of line art styles and increasing the fun and interactivity of users generating line art using electronic devices.
[0007] For example, the first image can refer to an image of the line drawing of the target object to be generated. For instance, the first image can be an image obtained by taking a photo, an image received from a remote device, an image obtained by taking a screenshot using a screenshot tool, and so on.
[0008] It should be noted that the first image can also be a single image from a video sequence. In other words, this application can also generate line art images of the target object from images within a video sequence.
[0009] For example, the target objects include, but are not limited to, people, animals, objects, etc.
[0010] For example, image style information can be used to indicate the image style of a line drawing image; this application does not limit the specific form of image style information. For example, image style information can be an image style identifier, which is used to indicate an image style.
[0011] For example, when the image processing method is executed by a terminal device, if the user does not select an image style in the interactive interface this time, the image style information obtained by the terminal device may be the image style information set by the system default of the terminal device; or, the image style information corresponding to the image style selected by the user in the last interactive interface; or, an image style information randomly selected from multiple image style information; or, image style information obtained from the server; and so on. When the user selects an image style in the interactive interface this time, the terminal device can obtain the image style information corresponding to the image style selected by the user.
[0012] For example, when the image processing method is executed by a server, the server can obtain image style information from the terminal device. When the server does not obtain image style information from the terminal device, the image style information obtained by the server may be the image style information set by the server's system default settings; or, the image style information obtained from the terminal device last time; or, an image style information randomly selected from multiple image style information; or, image style information obtained from other servers; and so on.
[0013] In one possible approach, “obtaining the line drawing image of the target object in the first image based on the first image and image style information” can be understood as the terminal device generating the line drawing image of the target object in the first image based on the first image and image style information.
[0014] In one possible approach, "obtaining the line drawing image of the target object in the first image based on the first image and image style information" can be understood as the terminal device obtaining the line drawing image of the target object in the first image from the server.
[0015] In one possible approach, “obtaining the line drawing image of the target object in the first image based on the first image and image style information” can be understood as the server generating the line drawing image of the target object in the first image based on the first image and image style information.
[0016] According to the first aspect, obtaining a line drawing image of a target object in the first image based on the first image and the image style information includes: obtaining a reference image based on the first image and the image style information, the reference image being used as a coloring reference for the line drawing image; and obtaining the line drawing image based on the reference image.
[0017] Specifically, "obtaining a reference image based on the first image and the image style information" can generate a reference image whose image style corresponds to the image style information. That is, the image style set by the user (or the system setting or the image style information obtained from other electronic devices) is injected into the reference image. Then, a line drawing image is generated based on the reference image. In this way, the style of the line drawing image can be guaranteed to correspond to the image style information.
[0018] According to the first aspect, or any implementation of the first aspect above, based on the reference image, the line drawing image is obtained, including: obtaining line style information and structural feature information of the reference image; and generating the line drawing image based on the line style information and structural feature information of the reference image.
[0019] The line style information includes information describing the style of lines in a high-quality, aesthetically pleasing line art image (e.g., aesthetically pleasing, pressure-sensitive, highly closed, moderately thick, simple, etc.). Therefore, the line style information controls the clean and simple lines, the high degree of closure, and the pressure sensitivity and aesthetic appeal of the lines in the generated line art image.
[0020] The structural features of the reference image can be used to describe the main structure within the reference image. Therefore, by utilizing the structural features of the reference image, a high degree of alignment between the line drawing and the target object in the reference image can be ensured.
[0021] In this way, high-quality, aesthetically pleasing line drawings can be produced, solving the problem that painting creation must involve professional artists or experts in the field, and requires expensive human resources and time.
[0022] For example, a second network can be used to process the reference image to obtain a line drawing image. This second network may include a second generative model, a line style control module, and a second structure preservation module. It should be noted that the line style control module can be embedded in the second generative model; the output of the second structure preservation module is connected to the output of one or more network layers in the second generative model. The line style control module controls the type, thickness variation, pressure sensitivity, and style and aesthetics of the line drawings; therefore, line style information can be injected into the line drawing image by the line style control module. For example, the second structure preservation module can output structural characteristic information (such as a structural feature map) of the reference image; thus, the second structure preservation module maintains the consistency between the main structure of the line drawing image and the main structure of the reference image.
[0023] According to the first aspect, or any implementation of the first aspect above, a reference image is obtained based on the first image and the image style information, including: generating the reference image based on at least one of the structural feature information of the first image, the detailed feature information of the target object, or the identity feature information of the target object, and the image style information.
[0024] The structural features of the first image can be used to describe its main structure, such as the posture and body shape of a person. Therefore, by using the structural features of the first image, it can be ensured that the main structure of the reference image remains consistent with that of the first image.
[0025] The detailed feature information of the target object can be used to describe details such as clothing and accessories of the target object in the first image. Therefore, by using the detailed feature information of the target object, it can be ensured that the details of clothing, accessories, etc. of the target object in the reference image are consistent with those in the first image.
[0026] The target object's identity features can be used to describe the target object's identity characteristics in the first image, such as appearance, hair color, hairstyle, and whether it is wearing glasses, headphones, or jewelry. Therefore, by using the target object's identity features, consistency between the identity of the person in the reference image and the first image can be maintained.
[0027] For example, a first network can be used to process the second image and output a reference image. The first network may include a first generative model, a style injection module, an identification (ID) preservation module, a first structure preservation module, and a detail preservation module. It should be noted that the ID preservation module, the first structure preservation module, and the detail preservation module can be embedded in the first generative model; the output of the first structure preservation module is connected to the output of one or more network layers in the first generative model, and the output of the detail preservation module is also connected to the output of one or more network layers in the first generative model.
[0028] The style injection module can include multiple sets of network parameters, each set corresponding to a specific image style. When an image style needs to be expanded, the network parameters of the style injection module can be extended, ensuring the image style is scalable. For example, the network parameters of the style injection module can be switched to correspond to an image style set by the system, the user, or obtained from another server. This allows the style injection module to inject (or integrate) the image style corresponding to the system-set, user-set, or server-obtained image style into the reference image, thus stylizing the reference image. Since the line art image is generated based on the reference image, stylization of the line art image can also be achieved.
[0029] The detail preservation module can retain or preserve the detailed characteristics of the target object in the intermediate feature information it processes; thus, the detail preservation module can maintain the consistency of the clothing, accessories and other details of the target object in the reference image and the first image.
[0030] The first structure preservation module can obtain the structural feature information of the first image; and through the first structure preservation module, the consistency of the main structure of the reference image and the first image can be maintained, for example, the consistency of the posture and body shape of the human figure in the reference image and the first image.
[0031] The ID preservation module can retain or preserve the ID feature information of the target object in the intermediate feature information it processes; thus, the ID preservation module can maintain the consistency of the identity of the person in the reference image and the first image. For example, it can maintain the similarity of the target object's face, appearance, hair color, hairstyle, wearing glasses, wearing headphones / jewelry, etc. between the reference image and the first image.
[0032] According to the first aspect, or any implementation of the first aspect above, obtaining a line drawing image of the target object in the first image based on the first image and the image style information, further comprising: preprocessing the first image to obtain a second image; obtaining a reference image based on the first image and the image style information, comprising: obtaining the reference image based on the second image and the image style information.
[0033] For example, the purpose of preprocessing the first image includes, but is not limited to: making the first image conform to specifications, processing the first image into an image suitable (or beneficial) for algorithm processing, so as to make the line drawing effect of the target object in the subsequently generated line drawing image better.
[0034] For example, preprocessing includes, but is not limited to: background removal, brightening, centering the target object or a portion of the target object, magnifying the target object or a portion of the target object, adding filters, etc. For instance, the target object in the first image could be the object with the largest area proportion in the first image; in this case, the area proportion of each object in the first image can be determined, and then the object with the largest area proportion can be used as the target object. Another example is that the target object in the first image could be an object of a preset type in the first image, where the preset type could be such as a person, an animal, etc.; in this case, the type of each object in the first image can be identified, and the object of the preset type can be used as the target object. For example, only people in the first image can be used as target objects; or only cats in the first image can be used as target objects; and so on.
[0035] For example, if the target is a person, preprocessing can also include rejection of small faces or rejection of multiple faces. The preprocessing of the first image can be as follows: Determine if the first image contains multiple faces; if it does, return a rejection message; the terminal device can then display this message to prompt the user to change the first image. If the first image contains only one face, determine if the area of that face is less than an area threshold. If the area of the face in the first image is less than the area threshold, return a rejection message; the terminal device can then display this message to prompt the user to change the first image. If the area of the face in the first image is greater than or equal to the area threshold, the upper body image of the person in the first image can be cropped. Then, the cropped image undergoes a series of operations, including background removal, brightening, local centering of the face, enlarging the face, and adding filters (this application does not limit the execution order of these operations); finally, a clear, distinct face image with a single upper body and no background (i.e., the second image) can be obtained.
[0036] According to the first aspect, or any implementation of the first aspect above, the method further includes: obtaining the color information of the line drawing image. This allows users to refer to the color information when coloring the line drawing image, which is very user-friendly and easy to learn for a wide range of painting enthusiasts and beginners, enabling them to quickly participate and engage in the process.
[0037] It should be noted that the first aspect and any implementation thereof can be executed by the terminal device or by the server.
[0038] Secondly, this application provides an image processing method, which includes: first, in response to a first user operation, acquiring a first image; then, in response to a second user operation, acquiring image style information; next, sending the first image and the image style information to a server; then, receiving a line drawing image of a target object in the first image sent by the server, wherein the style of the line drawing image corresponds to the image style information; and displaying the line drawing image. In this way, a line drawing image corresponding to the desired image style can be generated for the user, which can improve interactivity and engagement.
[0039] In addition, the image styles provided by the interactive interface can be dynamically added, thus allowing the style of the line art images to be dynamically expanded.
[0040] According to the second aspect, the method further includes: receiving a reference image sent by the server, the reference image being used as a coloring reference for the line drawing image; and displaying the reference image. In this way, users can subsequently select colors from the reference image to color the line drawing image; or refer to the color scheme of the reference image to color the line drawing image; this is very user-friendly and easy to learn for a wide range of painting enthusiasts and beginners, allowing them to quickly participate and engage in the process.
[0041] According to the second aspect, or any implementation of the second aspect above, the method further includes: receiving a second image sent by the server, the second image being obtained by preprocessing the first image; and displaying the second image. This allows the user to easily identify the image upon which the line drawing image is generated.
[0042] According to the second aspect, or any implementation thereof, the method further includes: receiving color information of the line drawing image sent by the server; and displaying one or more color options corresponding to the color information. In this way, users can subsequently select one or more colors from the one or more color options corresponding to the color information to color the line drawing image. This is very user-friendly and easy to learn for a wide range of painting enthusiasts and beginners, allowing them to quickly participate and engage in the process.
[0043] According to the second aspect, or any implementation of the second aspect above, the method further includes: placing the line art image onto the canvas in response to a third user operation; and coloring the line art image in response to a fourth user operation. This allows users to color the line art image after it has been generated, increasing interactivity and engagement.
[0044] The second aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the second aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.
[0045] Thirdly, this application provides a terminal device for:
[0046] In response to the first user action, acquire the first image;
[0047] In response to a second user's action, obtain image style information;
[0048] Send the first image and its style information to the server;
[0049] Receive the line drawing image of the target object in the first image sent by the server, the style of the line drawing image corresponding to the image style information;
[0050] This line art image is displayed.
[0051] For example, the terminal device may be a tablet computer, mobile phone, laptop computer, personal computer, in-vehicle computer, wearable device, set-top box, game console, etc.
[0052] It should be understood that the terminal device can be used to execute the methods in the first aspect and any implementation thereof, which will not be elaborated here.
[0053] Fourthly, this application provides an image processing system, which includes: a terminal device and a server, wherein:
[0054] The terminal device is configured to, in response to a first user operation, acquire a first image; in response to a second user operation, acquire image style information; and send the first image and the image style information to the server.
[0055] The server is used to generate a line drawing image of the target object in the first image based on the first image and the image style information; and to send the line drawing image to the terminal device.
[0056] The terminal device is also used to display the line drawing image.
[0057] Fifthly, this application provides an image processing apparatus, the apparatus comprising:
[0058] The image acquisition module is used to acquire the first image;
[0059] The style acquisition module is used to acquire image style information;
[0060] The line art acquisition module is used to acquire the line art image of the target object in the first image based on the first image and the image style information.
[0061] It should be understood that the image processing apparatus of the fifth aspect can execute the methods of the first aspect or any possible implementation thereof, which will not be elaborated here.
[0062] In a sixth aspect, this application provides an electronic device, including: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, which, when executed by the processor, cause the electronic device to perform the method of the first aspect or any possible implementation thereof.
[0063] In a seventh aspect, this application provides a chip including one or more interface circuits and one or more processors; the one or more processors receive or transmit data through the one or more interface circuits, and when the one or more processors execute computer instructions, cause the electronic device to perform the method in the first aspect or any possible implementation of the first aspect.
[0064] Eighthly, this application provides a computer-readable storage medium storing a computer program that, when run on a computer or processor, causes the computer or processor to perform the method of the first aspect or any possible implementation thereof.
[0065] Ninthly, this application provides a computer program product including computer instructions that, when executed by a computer or processor, cause the computer or processor to perform the method in the first aspect or any possible implementation thereof.
[0066] In this embodiment, the electronic device, computer-readable storage medium, computer program product, chip, system, etc. are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding methods provided above.
[0067] The electronic device in this application may refer to a terminal device or a server. Attached Figure Description
[0068] Figure 1A is a schematic diagram of an application scenario provided by an embodiment of this application;
[0069] Figure 1B is a schematic diagram of another application scenario provided by the embodiments of this application;
[0070] Figure 1C is a schematic diagram of another application scenario of the present application embodiment;
[0071] Figure 1D is a schematic diagram of another application scenario of the present application embodiment;
[0072] Figure 2A is a schematic diagram of a tablet computer interface according to an embodiment of this application;
[0073] Figure 2B is a schematic diagram of another tablet computer interface according to an embodiment of this application;
[0074] Figure 2C is a schematic diagram of another tablet computer interface according to an embodiment of this application;
[0075] Figure 2D is a schematic diagram of another tablet computer interface according to an embodiment of this application;
[0076] Figure 2E is a schematic diagram of another tablet computer interface according to an embodiment of this application;
[0077] Figure 2F is a schematic diagram of another tablet computer interface according to an embodiment of this application;
[0078] Figure 2G is a schematic diagram of another tablet computer interface according to an embodiment of this application;
[0079] Figure 2H is a schematic diagram of another tablet computer interface according to an embodiment of this application;
[0080] Figure 3 is a schematic diagram of an image processing procedure 300 according to an embodiment of this application;
[0081] Figure 4A is a schematic diagram of another image processing process 400 according to an embodiment of this application;
[0082] Figure 4B is a schematic diagram of a first image and a second image according to an embodiment of this application;
[0083] Figure 4C is a schematic diagram of the structure of a first network according to an embodiment of this application;
[0084] Figure 4D is a schematic diagram of the structure of a second network according to an embodiment of this application;
[0085] Figure 5 is a schematic diagram of another image processing procedure 500 according to an embodiment of this application;
[0086] Figure 6 is a schematic diagram of the image processing process 600 according to an embodiment of this application;
[0087] Figure 7 is a schematic diagram of a reference image and a line drawing image with different image styles according to an embodiment of this application;
[0088] Figure 8 is a schematic diagram of an image processing apparatus 800 according to an embodiment of this application;
[0089] Figure 9 is a schematic diagram of the structure of a device provided in an embodiment of this application. Detailed Implementation
[0090] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0091] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0092] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.
[0093] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0094] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.
[0095] In the embodiments of this application, the modules / components shown in the framework diagram (or structural diagram or system diagram) are merely examples of this application. The actual framework (or structure or system) may include more or fewer modules / components than those shown in the diagram, or may have different component configurations. Furthermore, the various components / modules shown in the diagrams may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0096] Figure 1A is a schematic diagram of an application scenario according to an embodiment of this application.
[0097] In Figure 1A, when a user wants to generate a line drawing of the person in Image 1, they can perform a line drawing generation operation on the tablet computer. After receiving the user's line drawing generation operation, the tablet computer can respond by sending a line drawing generation request to the cloud server, which includes Image 1. Upon receiving this request, the cloud server can generate a line drawing with the same style as Image 1, i.e., Image 2, and return Image 2 to the tablet computer. The tablet computer can then display Image 2.
[0098] Figure 1B is a schematic diagram of another application scenario of the present application embodiment.
[0099] Figure 1B shows that after the tablet receives the user's line art generation operation for image 3, it can respond to the user's operation by sending a line art generation request to the cloud server. This request includes image 3. Upon receiving the request, the cloud server can generate a line art image (image 4) with the same image style as the default (or preset) image style, and a reference image (image 5) with the same image style as the default (or preset) image style (used as a coloring reference for the line art image). It then returns images 4 and 5 to the tablet. The tablet, upon receiving images 4 and 5, can display them.
[0100] Figure 1C is a schematic diagram of another application scenario of the present application.
[0101] As shown in Figure 1C, after the tablet receives the user's line art generation operation for image 6 (this line art generation operation may include setting the image style, with the image style set to style 1), it can respond to the user's operation by sending a line art generation request to the cloud server. This request includes image 6 and style 1 (which can be understood as the identification information of style 1). After receiving the line art generation request, the cloud server can generate a line art image with style 1, namely image 8, and a reference image with style 1 (used as a coloring reference for the line art image), namely image 7; and return images 7 and 8 to the tablet. After receiving images 7 and 8, the tablet can display images 7 and 8.
[0102] Figure 1D is a schematic diagram of another application scenario of the present application.
[0103] Figure 1D shows that after the tablet receives the user's line art generation operation for image 9 (this line art generation operation may include setting the image style, with the image style set to style 2), it can respond to the user's operation by sending a line art generation request to the cloud server. This request includes image 9 and style 2 (which can be understood as the identification information of style 2). After receiving the line art generation request, the cloud server can generate a line art image with style 2, namely image 11, and a reference image with style 2 (used as a coloring reference for the line art image), namely image 10; and return images 10 and 11 to the tablet. After receiving images 10 and 11, the tablet can display images 10 and 11.
[0104] It should be understood that the tablet computer in Figures 1A to 1D can be replaced by other terminal devices such as mobile phones, laptops, personal computers, in-vehicle computers, wearable devices, set-top boxes, and game consoles. This application uses a tablet computer as an example for illustration.
[0105] It should be understood that the cloud servers in Figures 1A to 1D can be replaced by other types of servers such as physical (dedicated) servers and server clusters.
[0106] Figures 1A to 1D show line art images and reference images generated by a cloud server; it should be understood that line art images and reference images can also be generated locally by a terminal device, and this application does not limit this.
[0107] The following uses the application scenario shown in Figure 1C or Figure 1D as an example to illustrate the process of a user interacting with a tablet computer to obtain line drawing images and reference images.
[0108] Figures 2A to 2H are schematic diagrams of the tablet computer interface according to embodiments of this application.
[0109] Referring to Figure 2A, 201 is the main interface of the tablet computer. The main interface of the tablet computer may include one or more controls, including but not limited to: application icons (e.g., the application icon of Huawei Video application, the application icon of browser application, and the application icon of Born to Draw application 202), network icon, battery icon, etc.
[0110] Referring again to Figure 2A, when a user needs to generate a line drawing of a target object in an image, the user can click the application icon 202 of the "Born to Draw" app. The tablet computer responds to the user's operation and displays the application interface 203 of the "Born to Draw" app, as shown in Figure 2B(1). For example, the application interface 203 of the "Born to Draw" app may include one or more controls, including but not limited to: beginner tutorial options, drawing example options (e.g., Cosmic Symphony option, Spring Scenery option), line drawing practice options (also known as line drawing generation options) 204, etc.
[0111] Referring again to Figure 2B(1), the user can click the line art practice option 204. The tablet computer can respond to the user's operation and display the line art practice interface 205, as shown in Figure 2B(2). For example, the line art practice interface 205 may include one or more controls, including but not limited to: line art example options (e.g., puppy example option, kitten example option), historical line art options (e.g., unnamed 1 option, unnamed 2 option, unnamed 3 option), and the add portrait practice option 206, etc.
[0112] Referring to Figure 2B(2), the user can click the "Add Portrait Practice" option 206. The tablet computer can respond to the user's operation and display the portrait practice interface 207, as shown in Figure 2C. The portrait practice interface 207 may include one or more controls, including but not limited to: image addition option 208, multiple image style options 209 (e.g., cute image style option, minimalist image style option, fashionable image style option, lively image style option, no style option (not shown in the figure, this option is optional) etc.), and get line art option 210.
[0113] Referring again to Figure 2C, in one possible scenario, the user can click the image addition option 208. The tablet computer, in response to the user's action, can enter the main album interface 211, as shown in Figure 2D(1). The user can select an image 212 from the main album interface 211 from which a line drawing is desired to be generated. The tablet computer, in response to the user's action, can display image 212 in the image addition option 208 of the portrait practice interface 207, as shown in Figure 2E. It should be understood that the user can also select a video from the main album interface 211 from which a line drawing is desired to be generated; this application does not limit this selection.
[0114] Referring again to Figure 2C, in one possible scenario, the user can click the image addition option 208, and the tablet computer can respond to the user's action by entering the camera interface 213, as shown in Figure 2D(2). The user can click the camera button in the camera interface 213, and the tablet computer can respond to the user's action by taking a photo and displaying the resulting image in the image addition option 208 of the portrait practice interface 207, as shown in Figure 2E. It should be understood that the user can also shoot a video in the camera interface 213 to generate a line drawing; this application does not limit this.
[0115] Referring to Figure 2E, when a user wants to change the image / video for which the line drawing is to be generated, they can click the image addition option 208 again; the tablet can respond to the user's operation and enter the album main interface 211 or the camera interface 213.
[0116] Referring again to Figure 2E, the user can select the desired image style option for the line art from multiple image style options 209 (e.g., clicking the cute style option). At this time, the tablet can obtain the image style information corresponding to the user's operation (e.g., the image style information is the identifier information corresponding to the cute style option, such as cute). Afterwards, the user can click the "Get Line Art" option 210. The tablet can respond to the user's operation by sending a line art generation request to the cloud server and displaying a waiting interface 214, as shown in Figure 2F.
[0117] It should be noted that, in one possible approach, the above-mentioned line art generation operation can be the operation of clicking the "Get Line Art" option 210. In another possible approach, the above-mentioned line art generation operation can include: clicking the "Add Image" option 208, selecting the image 212 from which the line art is to be generated (or clicking the "Shoot" button), selecting the "Image Style" option 209, and then clicking the "Get Line Art" option 210.
[0118] After the tablet computer receives the line drawing image and reference image sent by the cloud server, in one possible scenario, the tablet computer can display the line drawing display interface 215, as shown in Figure 2G(1) and Figure 2G(2).
[0119] Referring to Figure 2G(1), the line art display interface 215 may include one or more controls, including but not limited to: line art image 216, thumbnail 217 of the original image (i.e., the image selected by the user from the album or the image taken by the user, which is the subsequent first image), thumbnail 218 of the line art image, thumbnail 219 of the reference image, re-acquire option, place canvas option 220, etc.
[0120] Referring to Figure 2G(2), the line art display interface 215 may include one or more controls, including but not limited to: line art image 216, thumbnail 221 of the preprocessed original image (i.e., the subsequent second image), thumbnail 218 of the line art image, thumbnail 219 of the reference image, re-acquire option, and place canvas option 220.
[0121] Referring again to Figure 2G(1), when the user clicks the Canvas Placement option 220, the tablet computer responds to the user's action and displays the canvas interface 218, as shown in Figure 2H(1). For example, the canvas interface 218 may include one or more controls, including but not limited to: line art images, thumbnails of reference images, thumbnails of preprocessed original images, and various painting tool options (such as color options 219, brushes, color fill tools, color pickers, etc.).
[0122] In one possible approach, the user clicks color option 219, and the tablet computer responds to the user's action by displaying multiple color options (not shown in the figure); the user can then select any color option. Afterwards, the user clicks the brush and moves the brush across the desired area of the line art image. The tablet computer responds to the user's action by displaying the color corresponding to the selected color option along the brush's path (this is the smearing / coloring method). Alternatively, after selecting any color option, the user clicks the color fill tool and then clicks on the desired area. The tablet computer responds to the user's action by filling that area of the line art image with the color corresponding to the selected color option.
[0123] In one possible approach, the user can click the color picker, then click on the location of the desired color in the reference image. The tablet can then respond to the user's action and retrieve the color at that location. The user can then choose a brush or fill tool to apply the color, which will not be elaborated further here.
[0124] For example, after receiving the line drawing generation request information, the cloud server can also generate the color scheme information of the line drawing image; then, the cloud server can also send the color scheme information of the line drawing image to the tablet computer. In this way, the tablet computer can also display multiple color options 222 corresponding to the color scheme information in the canvas interface 218, as shown in Figure 2H(2). In one possible way, the user can click on any color option 222, and then the user can choose a brush or fill tool to color, which will not be described in detail here.
[0125] It should be understood that users can also perform other editing operations on the line art image, such as adjusting the position or angle of any line in the line art image, or adding or deleting lines in the line art image, etc., and this application does not impose any restrictions on this.
[0126] It should be understood that this application can also generate line art images of the target object in each video frame, and color the line art images of the target object in each video frame, or adjust, add, or delete lines, etc.
[0127] The process of generating line art images and reference images is explained below.
[0128] Figure 3 is a schematic diagram of an image processing procedure 300 according to an embodiment of this application. The image processing procedure 300 can be executed by a terminal device or by a cloud server.
[0129] S301, acquire the first image.
[0130] For example, the first image can refer to an image of the line drawing of the target object to be generated. For instance, the first image can be an image obtained by taking a photo, an image received from a remote device, an image obtained by taking a screenshot using a screenshot tool, and so on.
[0131] For example, the target objects include, but are not limited to, people, animals, objects, etc.
[0132] For example, when S301 is executed by the cloud server, the cloud server can obtain the first image from the terminal device. When S301 is executed by the terminal device, the terminal device can obtain an image selected by the user from the album or an image obtained by taking a photo, as the first image.
[0133] S302, Obtain image style information.
[0134] For example, when S302 is executed by the terminal device, if the user does not select image style option 209 in Figure 2E this time, the image style information obtained by the terminal device may be the image style information of the terminal device's system default setting; or, the image style information corresponding to an image style option 209 that the user previously selected in Figure 2E; or, an image style information randomly selected from multiple image style information; or, image style information obtained from the server. When the user selects image style option 209 in Figure 2E this time, the terminal device can obtain the image style information corresponding to the image style option 209 selected by the user.
[0135] For example, when S302 is executed by the cloud server, the cloud server can obtain image style information from the terminal device. When the cloud server does not obtain image style information from the terminal device, the image style information obtained by the cloud server may be the image style information set by the cloud server's system default; or, the image style information obtained from the terminal device last time; or, an image style information randomly selected from multiple image style information; or, image style information obtained from other servers.
[0136] For example, image style information can be used to indicate the image style of a line drawing image; this application does not limit the specific form of image style information. For example, image style information can be an image style identifier, which is used to indicate an image style.
[0137] S303, Based on the first image and image style information, obtain the line drawing image of the target object in the first image.
[0138] When S303 is executed by the terminal device, in one possible approach, "obtaining the line drawing image of the target object in the first image based on the first image and image style information" can be understood as the terminal device generating the line drawing image of the target object in the first image based on the first image and image style information. In another possible approach, "obtaining the line drawing image of the target object in the first image based on the first image and image style information" can be understood as the terminal device obtaining the line drawing image of the target object in the first image from the cloud server; that is, in this case, the server generates the line drawing image of the target object in the first image based on the first image and image style information.
[0139] When S303 is executed by the cloud server, the cloud server can generate a line drawing image of the target object in the first image based on the first image and image style information.
[0140] In summary, during the line art generation process, an image style set by the user (or the system setting or image style information obtained from other electronic devices) can be injected to generate a line art image for the target object in the first image that corresponds to the image style set by the user (or the system setting or image style information obtained from other electronic devices). For example, when the image style set by the user or the system setting of the electronic device is image style 1, this application can generate a line art image with image style 1; when the image style set by the user or the system setting of the electronic device is image style 2, this application can generate a line art image with image style 2. Therefore, this application can generate line art images of various styles, achieving diversification of line art styles and increasing the fun and interactivity of users generating line art using terminal devices.
[0141] Figure 4A is a schematic diagram of another image processing procedure 400 according to an embodiment of this application. The image processing procedure 400, based on the image processing procedure 300, describes the process of generating a line drawing image of the target object in the first image. The image processing procedure 400 can be executed by a terminal device or by a cloud server.
[0142] S401, Obtain the first image.
[0143] S402, Obtain image style information.
[0144] For example, S401 and S402 can be referred to the description of S301 to S302 above, and will not be repeated here.
[0145] For example, the process of generating a line drawing image of the target object in the first image based on the first image and image style information may include the following steps S403 to S405:
[0146] S403, preprocess the first image to obtain the second image.
[0147] For example, the purpose of preprocessing the first image includes, but is not limited to: making the first image conform to specifications, processing the first image into an image suitable (or beneficial) for algorithm processing, so as to improve the line drawing effect of the target object in the subsequently generated line drawing image. It should be understood that S403 is an optional step.
[0148] For example, preprocessing includes, but is not limited to: background removal, brightening, centering the target object or a portion of the target object, magnifying the target object or a portion of the target object, adding filters, etc. For instance, the target object in the first image could be the object with the largest area proportion in the first image; in this case, the area proportion of each object in the first image can be determined, and then the object with the largest area proportion can be used as the target object. Another example is that the target object in the first image could be an object of a preset type in the first image, where the preset type could be such as a person, an animal, etc.; in this case, the type of each object in the first image can be identified, and the object of the preset type can be used as the target object. For example, only people in the first image can be used as target objects; or only cats in the first image can be used as target objects; and so on.
[0149] For example, if the target is a person, preprocessing can also include rejection of small faces or rejection of multiple faces. The preprocessing of the first image can be as follows: Determine if the first image contains multiple faces; if it does, return a rejection message; the terminal device can then display this message to prompt the user to change the first image. If the first image contains only one face, determine if the area of that face is less than an area threshold. If the area of the face in the first image is less than the area threshold, return a rejection message; the terminal device can then display this message to prompt the user to change the first image. If the area of the face in the first image is greater than or equal to the area threshold, the upper body image of the person in the first image can be cropped. Then, the cropped image undergoes a series of operations, including background removal, brightening, local centering of the face, enlarging the face, and adding filters (this application does not limit the execution order of these operations); finally, a clear, distinct face image with a single upper body and no background (i.e., the second image) can be obtained.
[0150] Figure 4B is a schematic diagram of a first image and a second image according to an embodiment of this application. Figure 4B(1) is the first image. The second image obtained by preprocessing the first image in Figure 4B(1) is shown in Figure 4B(2). It should be understood that the fact that the first image and the second image shown in Figure 4B have the same size is only an example of this application. The size of the second image in this application may be different from the size of the first image.
[0151] S404, Based on the second image and image style information, obtain the reference image.
[0152] For example, at least one of the structural feature information of the second image, the detailed feature information of the target object, or the identity feature information of the target object can be obtained; then, a reference image is generated based on at least one of the structural feature information of the second image, the detailed feature information of the target object, or the identity feature information of the target object, and image style information.
[0153] For example, a first network can be used to process the second image and output a reference image. The first network may include a first encoder, a second encoder, a third encoder, a first decoder, a first generative model, a style injection module, an identity (ID) preservation module, a first structure preservation module, and a detail preservation module. It should be noted that at least one of the ID preservation module, the first structure preservation module, and the detail preservation module can be embedded in the first generative model; the following explanation uses an example where all three modules are embedded in the first generative model.
[0154] For example, the ID preservation module is used to maintain the consistency of the identity of the person in the reference image and the second image (or the first image). For example, it maintains the consistency of the target object's face in the reference image with the face in the second image, such as similar appearance, consistent hair color, consistent hairstyle, wearing glasses, wearing headphones or jewelry, etc.
[0155] For example, the first structure preservation module is used to maintain the consistency of the main structure of the reference image and the second image (or the first image), such as maintaining the consistency of the pose, body shape, etc. of the human figure in the reference image and the second image.
[0156] For example, the style injection module is used to inject (or integrate) an image style corresponding to system-set, user-set, or image style information obtained from other electronic devices into a reference image.
[0157] For example, the detail preservation module is used to maintain the consistency of details such as clothing and accessories of the target object in the reference image and the first image.
[0158] Figure 4C is a schematic diagram of the structure of a first network according to an embodiment of this application.
[0159] Referring to Figure 4C(1), the first generation model may include multiple network layers (e.g., white rectangles). The style injection module (e.g., gray rectangle) is embedded between each of the N1 (N1 is a positive integer, N1 = 2 in Figure 4C) network layers of the first generation model (each network layer includes two network layers); or, the style injection module is embedded within each of the N2 (N2 is a positive integer) network layers of the first generation model (not shown in the figure). The ID preservation module (e.g., black rectangle) is embedded between each of the N3 (N3 is a positive integer, N3 = 2 in Figure 4C) network layers of the first generation model (each network layer includes two network layers); or, the ID preservation module is embedded within each of the N4 (N4 is a positive integer) network layers of the first generation model (not shown in the figure); or, the ID preservation module is embedded between the network layers of the first generation model and the style injection module.
[0160] The output of the first structure preservation module is connected to the output of one or more network layers in the first generative model. The output of the first structure preservation module and the output of network layer 1 connected to the first structure preservation module in the first generative model serve as the input of the next network layer (i.e., network layer 2) in the first generative model, as shown in Figure 4C(2). The output of the detail preservation module is connected to the output of one or more network layers in the first generative model. The output of the detail preservation module and the output of network layer 2 connected to the detail preservation module in the first generative model serve as the input of the next network layer (i.e., network layer 3) in the first generative model, as shown in Figure 4C(2).
[0161] Optionally, the first generative model is a diffusion network. The chained generation process of the diffusion network can correspond to T (T is a positive integer) + 1 states, and T + 1 states correspond to T + 1 time nodes: ZT (the Tth state) corresponds to time node t = T, ZT-1 (the (T-1)th state) corresponds to time node t = T-1, ..., ZT-H+1 (the (T-H+1)th state) corresponds to time node t = T-H+1, ..., Z1 (the 1st state) corresponds to time node t = 1, and Z0 (the 0th state) corresponds to time node t = 0. The chained generation process of the diffusion network can include T generation processes, where the current time node and its corresponding input information are input into the diffusion network, and the diffusion network performs one generation process to obtain the state corresponding to the next time node; that is, each generation process is located between two adjacent time nodes. For example, the style injection module, ID preservation module, first structure preservation module, and detail preservation module can be activated during M (M is a positive integer, M is less than or equal to T) generation processes in the T generation processes of the chained generation process of the diffusion network. As shown in Figure 4C, the style injection module, ID preservation module, first structure preservation module, and detail preservation module are activated during the H-th generation process of the diffusion network (or at time t = T - H + 1), while they are not activated during other generation processes (or at other times).
[0162] For example, the style injection module can be a neural network such as LoRa. For example, the ID preservation module can be a neural network such as LoRa.
[0163] For example, the first structure preservation module may include one or more structure feature extraction modules, each of which may be the left half of an encoder such as a U-NET network.
[0164] For example, the detail preservation module can be an encoder.
[0165] Referring again to Figure 4C(1), the process of obtaining the reference image based on the second image and image style information can be as follows: The second image can be input into the first encoder to obtain ZT, and ZT and time T can be input into the diffusion network. The diffusion network performs the first generation process to obtain ZT-1. Then, ZT-1 and time T-1 are input into the diffusion network, and the diffusion network performs the second generation process to obtain ZT-2, and so on, until ZT-H+1 is obtained. In each generation process from the first generation process to the H-1th generation process, the style injection module, ID preservation module, first structure preservation module, and detail preservation module are not activated.
[0166] At time t = T - H + 1 or before time t = T - H + 1, the first text prompt can be input to the second encoder to obtain the first feature information; and the second image (or an image of the facial region of the second image) can be input to the third encoder to obtain the second feature information; the first feature information and the second feature information are concatenated to obtain the third feature information. The first text prompt may include text information obtained by the preprocessing module analyzing the content of the second image, and / or image style information.
[0167] Referring to Figure 4C(2), at time t = T - H + 1 or before time t = T - H + 1, the preprocessing module can further preprocess the second image to obtain an image for extracting structural feature information from the second image (hereinafter referred to as the third image, such as the depth map, edge map, key point map, etc. of the second image). At time t = T - H + 1, the first structure preservation module is activated; subsequently, a third image can be input to a structural feature extraction module to obtain structural feature information (which may be a structural feature map).
[0168] Referring to Figure 4C(2), the detail preservation module can be activated at time t = T - H + 1; subsequently, the second image can be input into the detail preservation module to obtain detail feature information (such as a detail feature map).
[0169] Referring to Figure 4C(2), at time t = T-H+1, ZT-H+1 and time T-H+1 are input into the diffusion network, and the style injection module and ID preservation module are activated; then, the diffusion network performs the Hth generation process. The style injection module can include multiple sets of network parameters, each set corresponding to an image style (or an image style information set). After the style injection module is activated, its network parameters can be switched to a set of network parameters corresponding to the image style information. This ensures that the user-set (or system-set or image style information obtained from other electronic devices) image style is injected into the line drawing image.
[0170] Referring to Figure 4C(2), during the H-th generation process of the diffusion network, network layer b1 processes ZT-H+1, the third intermediate feature information, and time H, and outputs intermediate feature information b1 to network layer b2. Each network layer (which can be called network layer bi (i is a positive integer greater than 1 and less than or equal to n)) in network layers b2 to bn (n is a positive integer greater than 1) processes the intermediate feature information b(i-1) and the third feature information output by the previous network layer, and outputs intermediate feature information bi. Among them, network layer bn outputs intermediate feature information bn to the style injection module. The style injection module processes the intermediate feature information bn and the third feature information to obtain intermediate feature information 1 (in this way, the image style information can be integrated into the intermediate feature information 1) and outputs it to network layer 4. Network layer 4 processes intermediate feature information 1 and the third feature information, and outputs intermediate feature information 2 to the ID preservation module. The ID preservation module processes intermediate feature information 2 and the third feature information, outputting intermediate feature information 3 (thus, the ID feature information of the target object can be preserved or retained in intermediate feature information 2) to network layer 1. Network layer 1 processes intermediate feature information 3 and the third feature information, outputting intermediate feature information 4 to network layer 2; and the first structure preservation module outputs structural feature information to network layer 2. Network layer 2 processes intermediate feature information 4, structural feature information, and the third feature information, outputting intermediate feature information 5 to network layer 3; and the detail preservation module outputs detail feature information to network layer 3. Network layer 3 processes intermediate feature information 5, detail feature information, and the third feature information, outputting intermediate feature information 6 to the next network layer; and so on, until the diffusion network completes the Hth generation process to output ZT-H.
[0171] Subsequently, the diffusion network completes the next TH generation processes sequentially (in which the style injection module, ID preservation module, first structure preservation module, and detail preservation module are not activated), and outputs Z0 to the first decoder. The first decoder processes Z0 and outputs the reference image.
[0172] It should be understood that when S403 to S404 are executed by the cloud server, after obtaining the first image and image style information, the cloud server can start the reference image generation thread, which controls the first generation model to perform T generation processes, controls whether the style injection module, ID preservation module, first structure preservation module and detail preservation module are activated, and controls the style injection module to switch network parameters.
[0173] It should be understood that Figure 4C is only an example of this application. The positions in which the style injection module and the ID preservation module are embedded in the diffusion network are not limited, nor are the positions of the network layers to which the first structure preservation module and the detail preservation module are connected.
[0174] In this way, the ID preservation module ensures that the reference image has the characteristic of "likeness" (i.e., the reference image is "like" the first image or the second image); the detail preservation module and the first structure preservation module make the reference image have the characteristic of "consistency" (i.e., the reference image has the same structure and detail as the first image or the second image); the style injection module makes the style of the reference image consistent with the image style information corresponding to the user setting, system setting, or obtained from other electronic devices; and the style injection module, ID preservation module, detail preservation module, and first structure preservation module make the reference image have the characteristic of "aesthetics".
[0175] S405: Obtain the line drawing image based on the reference image.
[0176] For example, line style information and structural feature information of a reference image can be obtained; then, a line drawing image is generated based on the line style information and the structural feature information of the reference image. The line style information includes information describing the line style in a high-quality, aesthetically pleasing line drawing image (e.g., aesthetically pleasing, pressure-sensitive, highly closed, moderately thick, simple, etc.).
[0177] For example, a second network can be used to process the reference image to obtain a line drawing image. The second network may include a fourth encoder, a fifth encoder, a second decoder, a second generative model, a line style control module, and a second structure preservation module. It should be noted that at least one of the line style control module and the second structure preservation module can be embedded in the second generative model; the following explanation uses an example where both the line style control module and the second structure preservation module are embedded in the second generative model.
[0178] For example, the line style control module is used to control the type, thickness variation, pressure sensitivity, style and aesthetics of the line art.
[0179] For example, the second structure preservation module can be used to maintain the consistency between the main structure of the line drawing image and the main structure of the reference image.
[0180] Figure 4D is a schematic diagram of the structure of a second network according to an embodiment of this application.
[0181] Referring to Figure 4D(1), the second generation model may include multiple network layers (such as white rectangles), and the line style control module (such as gray rectangles) is embedded between each group of network layers (a group of network layers includes two network layers) in the N5 (N5 is a positive integer, N5 = 2 in Figure 4D) group of network layers in the second generation model; or, the line style control module is embedded in each of the N6 (N6 is a positive integer) network layers in the diffusion network (not shown in the figure).
[0182] The output of the second structure preservation module is connected to the output of one or more network layers in the second generative model. The output of the second structure preservation module and the output of network layer 1 connected to the structure preservation module in the second generative model are used as the input of the next network layer (i.e. network layer 2) in the second generative model, as shown in Figure 4D(2).
[0183] Optionally, the second generative model is a diffusion network. As shown in Figure 4D, the line style control module and the second structure preservation module are activated during the G-th generation process (G is a positive integer) of the diffusion network (or at time t = T - G + 1), while the line style control module and the second structure preservation module are not activated during other generation processes (or at other times).
[0184] For example, the line style control module can be a neural network such as LoRa.
[0185] For example, the second structure preservation module may include one or more structure feature extraction modules, each of which may be the left half of an encoder such as a U-NET network.
[0186] Referring again to Figure 4D(1), the process of obtaining the line drawing image based on the reference image can be as follows: The reference image can be input into the fourth encoder to obtain ZT, and ZT and time T are input into the diffusion network. The diffusion network performs the first generation process to obtain ZT-1. Then, ZT-1 and time T-1 are input into the diffusion network, and the diffusion network performs the second generation process to obtain ZT-2, and so on, until ZT-G+1 is obtained. In each generation process from the first generation process to the G-1 generation process, the line style control module and the second structure preservation module are not activated.
[0187] Referring to Figure 4D(2), at time t = T - G + 1 or before time t = T - G + 1, the second text prompt can be input to the fifth encoder to obtain the fourth feature information. The second text prompt may include line requirement information for the line drawing image, such as high closure, aesthetic appeal, simplicity, etc. The second text prompt can be system-set or user-inputted (the line requirement option can be displayed on the portrait practice interface 207 in Figure 2C, where the user can input line requirement information), and this application does not impose any restrictions on this.
[0188] Referring to Figure 4D(2), at time t = T - G + 1 or before time t = T - G + 1, the preprocessing module can further preprocess the reference image to obtain an image for extracting structural feature information from the reference image (hereinafter referred to as the fourth image, such as the edge map of the reference image, the inverse color map of the edge map of the reference image, etc.). At time t = T - G + 1, the second structure preservation module is activated; subsequently, a fourth image can be input to a structural feature extraction module to obtain structural feature information (which can be a structural feature map).
[0189] Referring to Figure 4D(2), at time t = T-G+1, ZT-G+1 and time T-G+1 are input into the diffusion network, and the line style control module is activated; then, the diffusion network performs the Gth generation process. During the Gth generation process of the diffusion network, network layer c1 processes ZT-G+1, time T-G+1, and the fourth feature information, and outputs intermediate feature information c1 to network layer c2. Each network layer (called network layer ci (i is a positive integer greater than 1 and less than or equal to n)) from network layer c2 to network layer cn (n is a positive integer greater than 1) processes the intermediate feature information c(i-1) and the fourth feature information output by the previous network layer, and outputs intermediate feature information ci. Among them, network layer cn outputs intermediate feature information cn to the line style control module, and the line style control module processes the intermediate feature information cn and the fourth feature information to obtain intermediate feature information 2 (in this way, the line style information can be integrated into intermediate feature information 1) and outputs it to network layer d1. Each network layer (called network layer di (i is a positive integer greater than 1 and less than or equal to m)) in network layers d1 to dm (m is a positive integer) processes the previous output intermediate feature information d(i-1) and the fourth feature information, and outputs intermediate feature information di. Network layer dm outputs intermediate feature information dm to the line style control module. The line style control module processes the intermediate feature information dm and the fourth feature information, and outputs intermediate feature information 1 to network layer 1. Network layer 1 processes intermediate feature information 1 and the fourth feature information, and outputs intermediate feature information 2 to network layer 2; and the second structure preservation module outputs structural feature information to network layer 2. Network layer 2 processes intermediate feature information 2, structural feature information, and the fourth feature information, and outputs intermediate feature information 3 to network layer 3. Each network layer from network layer 3 to the last network layer of the diffusion network can process the intermediate feature information and the fourth feature information output by the previous network layer, until the diffusion network completes the G-th generation process to output ZT-G.
[0190] Subsequently, the diffusion network sequentially completes the TG generation process (during which the line style control module and the second structure preservation module are not activated), and outputs Z0 to the second decoder. The second decoder processes Z0 and outputs the line art image.
[0191] It should be understood that when S405 is executed by the cloud server, after obtaining the reference image, the cloud server can start the line art image generation thread, which controls the second generation model to perform T generation processes, and controls whether the line style control module and the second structure preservation module are activated.
[0192] It should be understood that Figure 4D is only an example of this application. This application does not limit the position of the line style control module embedded in the diffusion network, nor does it limit which network layer of the diffusion network the second structure holding module is connected to.
[0193] It should be understood that the number of diffusion processes T performed by the first generative model and the second generative model can be the same or different, and this application does not impose any restrictions on this.
[0194] In this way, the line style control module makes the lines of the line art image clean and simple, with high closure, and with pressure sensitivity and aesthetics; the second structure retention module makes the line art image fit the target object in the reference image very well.
[0195] For example, this application also provides an image processing system, which may include a terminal device and a server (cloud server). The terminal device can send a first image and image style information to the server; the server, based on the first image and image style information, generates a line drawing image of the target object in the first image and sends the line drawing image to the terminal device, which then displays the line drawing image. The processing procedure between the terminal device and the server is described below.
[0196] Figure 5 is a schematic diagram of another image processing process 500 according to an embodiment of this application. The image processing process 500 is based on the image processing process 300, and describes the process in which the terminal device obtains a first image and image style information according to the user operation and sends them to the cloud server, the cloud server generates a line drawing image, and the terminal device obtains the line drawing image by interacting with the cloud server.
[0197] S501, the terminal device responds to the first user's operation and acquires the first image.
[0198] Referring again to Figure 2D, in one possible scenario, the first user operation could be the user selecting an image on the main album interface 211, or it could be the user clicking the shutter button on the camera interface 213.
[0199] Referring again to Figures 2C and 2D, in one possible approach, the first user operation may include the user clicking the image addition option 208 in the portrait practice interface 207 and selecting an image in the album main interface 211; or, it may include the user clicking the image addition option 208 in the portrait practice interface 207 and clicking the photo button in the photo capture interface 213.
[0200] S502, the terminal device responds to the second user's operation and obtains image style information.
[0201] Referring again to Figure 2C, for example, the second user operation can be the user selecting any image style option 209 on the portrait practice interface 207.
[0202] S503, the terminal device sends the first image and image style information to the server.
[0203] Referring again to Figure 2C, when the user clicks the "Get Line Art" option 210 (which can be called the fifth user operation) on the portrait practice interface 207, the terminal device can respond to the fifth user operation and generate line art generation request information; the line art generation request information may include the first image and image style information; then, the line art generation request information can be sent to the cloud server.
[0204] S504, the terminal device receives a line drawing image of the target object in the first image sent by the server, and the style of the line drawing image corresponds to the image style information.
[0205] For example, after receiving the line drawing generation request information, the cloud server can extract the first image and image style information from the line drawing generation request information; then it executes the above-described steps S403 to S405 to generate a line drawing image of the target object in the first image and sends the line drawing image to the terminal device. In this way, the terminal device can receive the line drawing image.
[0206] S505, the terminal device displays the line drawing image.
[0207] For example, after receiving the line drawing image, the terminal device can display the line drawing image as shown in Figure 2G.
[0208] Figure 6 is a schematic diagram of the image processing process 600 according to an embodiment of this application. The image processing process 600 describes the interaction process between the terminal device and the server during the image processing, based on the image processing processes 400 and 500.
[0209] S601, the terminal device receives the first user's operation.
[0210] S602, the terminal device responds to the first user operation and acquires the first image.
[0211] S603, the terminal device receives the operation from the second user.
[0212] S604, the terminal device responds to the second user's operation and obtains image style information.
[0213] S605, the terminal device sends the first image and image style information to the server.
[0214] For example, S601 to S605 can be described with reference to the above description of S501 to S503, and will not be repeated here.
[0215] S606, the server preprocesses the first image to obtain the second image.
[0216] S607, the server generates a reference image based on the second image and image style information.
[0217] S608, the server generates a line drawing image based on a reference image.
[0218] For example, S606 to S608 can be described with reference to the above description of S403 to S405, and will not be repeated here.
[0219] S609, The server generates color scheme information for the line drawing image.
[0220] For example, the server can use a third network (neural network or neural network model) to analyze the line drawing image to output its color scheme information. This color scheme information can include multiple color values, such as the RGB values of various colors. For instance, the color scheme information could include: [R: 0, G: 0, B: 0], [R: 255, G: 128, B: 0], [R: 255, G: 255, B: 255], [R: 125, G: 222, B: 130].
[0221] S610: The server sends line drawing images, reference images, second images, and color matching information to the terminal device.
[0222] S611, the terminal device displays one or more color options corresponding to the color matching information of the line drawing image, reference image, second image and line drawing image.
[0223] For example, the display interface of the terminal device can be as shown in Figure 2H(2).
[0224] S612, the terminal device receives operations from a third user.
[0225] Referring again to Figure 2G, the third user operation can be the user clicking the Canvas Placement option 220.
[0226] S613, the terminal device responds to a third user's operation by placing the line drawing image onto the canvas.
[0227] For example, after the terminal device places the line drawing image into the canvas, it displays the canvas interface 218, as shown in Figure 2H.
[0228] S614, the terminal device receives a fourth user operation.
[0229] Referring again to Figure 2H, the fourth user operation can refer to the user moving the brush, or the user clicking the color fill tool on the canvas.
[0230] S615, the terminal device responds to the fourth user's operation and colors the line drawing image.
[0231] Since the line art image generated by this application has good line closure, in addition to smearing and coloring the line art image, it can also be filled with color to fill one or more areas in the line art image.
[0232] Figure 7 is a schematic diagram of a reference image and a line drawing image of a different image style according to an embodiment of this application.
[0233] In Figure 7, the left side shows either the first or second image; each of the four columns within the dashed box on the right represents a reference image and a line drawing image corresponding to an image style generated using the image processing method of this application. In Figure 7, from left to right, the first dashed box includes the reference image and line drawing image corresponding to image style AA, the second dashed box includes the reference image and line drawing image corresponding to image style BB, the third dashed box includes the reference image and line drawing image corresponding to image style CC, and the fourth dashed box includes the reference image and line drawing image corresponding to image style DD.
[0234] It should be understood that this application can also generate reference images and line drawings corresponding to more image styles.
[0235] Figure 8 is a schematic diagram of an image processing apparatus 800 according to an embodiment of this application. This image processing apparatus 800 can be used to execute the methods of the foregoing embodiments; therefore, the beneficial effects it can achieve can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0236] The image processing device 800 may include:
[0237] Image acquisition module 801 is used to acquire the first image;
[0238] Style acquisition module 802 is used to acquire image style information;
[0239] The line art acquisition module 803 is used to acquire the line art image of the target object in the first image based on the first image and image style information.
[0240] For example, the line art acquisition module 803 is used to acquire a reference image based on the first image and image style information, and the reference image is used as a coloring reference for the line art image.
[0241] For example, the line art acquisition module 803 is used to acquire a reference image based on the first image and image style information, the reference image being used as a coloring reference for the line art image; and to acquire the line art image based on the reference image.
[0242] For example, the line art acquisition module 803 is used to acquire line style information and structural feature information of the reference image; and to generate a line art image based on the line style information and structural feature information of the reference image.
[0243] For example, the line drawing acquisition module 803 is used to generate a reference image based on at least one of the structural feature information of the first image, the detailed feature information of the target object, or the identity feature information of the target object, and image style information.
[0244] For example, the line drawing acquisition module 803 preprocesses the first image to obtain the second image; and acquires a reference image based on the second image and image style information.
[0245] For example, the image processing device 800 further includes a color matching acquisition module, which is used to acquire color matching information of the line drawing image.
[0246] In one example, FIG9 shows a schematic block diagram of an apparatus 900 according to an embodiment of the present application. The apparatus 900 may include a processor 901 and a transceiver 902, and optionally, a memory 903.
[0247] The various components of device 900 are coupled together via bus 904, which includes a data bus, a power bus, a control bus, and a status signal bus. However, for clarity, all buses are referred to as bus 904 in the figure.
[0248] Optionally, the memory 903 can be used to store instructions from the foregoing method embodiments. The processor 901 can be used to execute the instructions in the memory 903, and to control the transceiver 902 to receive signals and transmit signals.
[0249] The device 900 may be an electronic device or a chip of an electronic device in the above method embodiments. For example, the electronic device may refer to the terminal device or server described above.
[0250] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0251] This application also provides a chip, including one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the above-described related method are executed to achieve the steps of the method in the above embodiments. The interface circuit is a transceiver 902.
[0252] This embodiment also provides a computer-readable storage medium storing computer instructions. When these computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the methods described in the above embodiments. Exemplarily, the computer-readable storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0253] This embodiment also provides a computer program product containing computer instructions that, when executed by a computer or processor, cause the computer to perform the aforementioned steps to implement the methods described in the above embodiments. Exemplarily, the computer program product can be stored in random access memory (RAM), flash memory, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, read-only optical discs (CD-ROMs), or any other form of storage medium well known in the art.
[0254] In this embodiment, the electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0255] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0256] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0257] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0258] Any content in the various embodiments of this application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.
[0259] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An image processing method, characterized in that, The method includes: Get the first image; Obtain image style information; Based on the first image and the image style information, obtain the line drawing image of the target object in the first image.
2. The method according to claim 1, characterized in that, The step of obtaining the line drawing image of the target object in the first image based on the first image and the image style information includes: Based on the first image and the image style information, a reference image is obtained, which is used as a coloring reference for the line drawing image; The line drawing image is obtained based on the reference image.
3. The method according to claim 2, characterized in that, The step of obtaining the line drawing image based on the reference image includes: Obtain line style information and structural feature information of the reference image; The line drawing image is generated based on the line style information and the structural feature information of the reference image.
4. The method according to claim 2 or 3, characterized in that, The step of obtaining a reference image based on the first image and the image style information includes: The reference image is generated based on at least one of the structural feature information of the first image, the detailed feature information of the target object, or the identity feature information of the target object, and the image style information.
5. The method according to claim 2, characterized in that, The step of obtaining the line drawing image of the target object in the first image based on the first image and the image style information further includes: The first image is preprocessed to obtain the second image; The step of obtaining a reference image based on the first image and the image style information includes: The reference image is obtained based on the second image and the image style information.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Obtain the color scheme information of the line drawing image.
7. An image processing method, characterized in that, The method includes: In response to the first user action, acquire the first image; In response to a second user's action, obtain image style information; Send the first image and the image style information to the server; Receive a line drawing image of the target object in the first image sent by the server, wherein the style of the line drawing image corresponds to the image style information; The line drawing image is displayed.
8. The method according to claim 7, characterized in that, The method further includes: Receive a reference image sent by the server, the reference image being used as a coloring reference for the line drawing image; The reference image is displayed.
9. The method according to claim 7 or 8, characterized in that, The method further includes: Receive a second image sent by the server, the second image being obtained by preprocessing the first image; The second image is displayed.
10. The method according to any one of claims 7 to 9, characterized in that, The method further includes: Receive the color scheme information of the line drawing image sent by the server; Display one or more color options corresponding to the color scheme information.
11. The method according to any one of claims 7 to 10, characterized in that, The method further includes: In response to a third user's action, the line drawing image is placed into the canvas; In response to a fourth user action, the line drawing image is colored.
12. A terminal device, characterized in that, The terminal device is used for: In response to the first user action, acquire the first image; In response to a second user's action, obtain image style information; Send the first image and the image style information to the server; Receive a line drawing image of the target object in the first image sent by the server, wherein the style of the line drawing image corresponds to the image style information; The line drawing image is displayed.
13. An image processing system, characterized in that, The image processing system includes a terminal device and a server, wherein: The terminal device is configured to, in response to a first user operation, acquire a first image; in response to a second user operation, acquire image style information; and send the first image and the image style information to the server. The server is configured to generate a line drawing image of the target object in the first image based on the first image and the image style information; and send the line drawing image to the terminal device. The terminal device is also used to display the line drawing image.
14. An image processing apparatus, characterized in that, The image processing device includes: The image acquisition module is used to acquire the first image; The style acquisition module is used to acquire image style information; The line art acquisition module is used to acquire the line art image of the target object in the first image based on the first image and the image style information.
15. An electronic device, characterized in that, include: A memory and a processor, wherein the memory is coupled to the processor; The memory stores program instructions that, when executed by the processor, cause the electronic device to perform the method as described in any one of claims 1 to 6.
16. A chip, characterized in that, It includes one or more interface circuits and one or more processors; the one or more processors receive or send data through the one or more interface circuits, and when the one or more processors execute computer instructions, the steps of the method as described in any one of claims 1 to 6 are performed.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed on a computer or processor, causes the computer or processor to perform the method as described in any one of claims 1 to 6.
18. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a computer or processor, cause the steps of the method as described in any one of claims 1 to 6 to be performed.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN114926326A
Image processing method and device, equipment and medium
CN115937338A
Image stylization processing method and device, equipment, storage medium and program product
CN116596748A
Image generation method and related device
CN118365752A
Apparatus and methods for processing images
US20110148897A1