Animation image line draft extraction method, system and device and storage medium
The improved U2-Net network extracts multi-scale edge response features of anime images. Combined with deconvolution and SE attention modules, it solves the contour fracture and structural disorder problems of line drawing extraction in existing technologies, and generates high-quality editable line drawings suitable for professional art creation.
Patent Information
- Application Number
- CN202511171344.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-21
AI Technical Summary
When dealing with complex backgrounds and high-density edge scenes, existing technologies for extracting line drafts from animation images are prone to broken contours, blurred edges, or structural disorder. They also lack the ability to process complex textures, making it difficult to generate high-quality structured line drafts.
An improved U2-Net network is used to extract multi-scale edge response features through parallel convolutional layers and skip connections. Deconvolution and SE attention modules are combined for feature fusion. Low-order features are used to correct high-order feature edge distortion, generating editable multi-scale and multi-semantic level feature responses.
The accuracy and precision of line drawing extraction are improved. The generated line drawings can be directly used for professional creation, have structured and semantic coherence, reduce background artifacts, and meet the needs of professional artistic creation.
Smart Images

Figure CN120747540A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and in particular relates to a method, system, device and storage medium for extracting line drafts from animation images. Background Art
[0002] Line drawing extraction, a field at the intersection of computer vision and digital art creation, aims to automatically generate structured and semantically coherent line contours from raw images. This process aims to preserve edge information critical to the visual structure of the image while removing redundant texture and complex background. In the field of artistic creation, line drawings are widely used in animation, illustration design, digital coloring, style transfer, and composition assistance, serving as a highly expressive and instructive intermediate form of expression in visual art. High-quality line drawings not only help clarify the structure and hierarchy of an image but also provide a clear and controllable reference framework for subsequent creative processing. For a long time, line drawings have often relied on manual drawing, a time-consuming and labor-intensive process that requires extremely high painting skills and a strong grasp of structure. This not only raises the bar for artistic creation but also significantly limits creative efficiency and the speed of content production. To this end, in recent years, extensive research has focused on achieving high-quality line drawing extraction through automated algorithms. Deep learning-based methods, in particular, have made progress in image edge detection and semantic modeling. However, existing technologies still have many shortcomings. For example, when processing images with complex backgrounds, strong light and shadow interference, or rich material textures, the extraction results often have problems such as broken contours, blurred edges, or disordered structures. These problems have, to a certain extent, limited the actual application effect of line drawing extraction technology in professional art creation. There is an urgent need to further optimize the model's semantic understanding and structural modeling capabilities to better serve creative generation and artistic expression.
[0003] For example, Chinese invention patent publication number CN119479044A discloses an intelligent comic line draft generation system that preserves facial features. The system includes a backbone network U-NET module and a control network U-NET module to generate a preliminary comic line draft, and a post-line draft adjustment module to fine-tune the preliminary draft to produce the final line draft. Another example is Chinese invention patent publication number CN118968241A, which discloses a mural line draft extraction method based on multi-scale extraction and intra-layer learnable fusion. The system establishes a line draft extraction model, including a multi-scale feature extraction module and an intra-layer learnable fusion module, and uses this line draft extraction model to generate image line drafts.
[0004] However, current line drawing extraction methods have significant limitations in specific areas such as anime-style images, and their processing capabilities for complex textures or high-density edge scenes are insufficient, making them prone to artifacts such as edge adhesion and background artifacts. Summary of the Invention
[0005] In order to solve the problem that shallow features are easily over-abstracted or information is attenuated in deep networks, the present invention provides a method for extracting line drafts from cartoon images, aiming to generate structured line drafts that can be directly used by artists for secondary creation.
[0006] In order to achieve the above object, the present invention provides the following technical solutions: A method for extracting line drawings from an animation image, comprising: Get the anime-style image to be processed; Deconvolution is performed on the high-order features in the hierarchical feature map, and the reconstructed high-order features are channel-joined with the adjacent low-order features. The joined features are fused at the pixel level across hierarchical responses through point convolution. At the same time, the edge distortion of the high-order features is corrected through the geometric constraints of the low-order features to obtain a complementary enhanced feature expression. Dynamically fuse the complementary enhanced feature expressions to obtain multi-scale and multi-semantic level feature responses that integrate pixel-level information; By integrating pixel-level information into multi-scale and multi-semantic feature responses, editable line drawings are output.
[0007] Preferably, the anime style image to be processed is input into the improved U 2 -Net network, outputs multi-scale and multi-semantic level feature responses that complete pixel-level information integration; the improved U 2 -Net network with U 2 -Net nested U-shaped structure is used as the basis, deconvolution is added to the input of each parallel convolutional layer of the symmetric encoder-decoder after the input convolutional layer, and the SE attention module is embedded at the output of each parallel convolutional layer.
[0008] Preferably, the improved U 2 The parallel convolutional layers of the -Net network set up convolution paths with different receptive fields. By adopting convolution kernels of different sizes or different step sizes, convolution operations are performed in parallel on the anime-style images to be processed, features are extracted from images at different scales, feature responses containing multi-scale information are generated, and a hierarchical feature map with spatial perception capabilities is constructed.
[0009] Preferably, the improved U 2 -Net network's skip connection layer passes the low-order features extracted by the shallow network directly to the deep network, bypassing the middle multi-layer network. The shallow network is used to capture the low-order visual features of the image, and the deep network is used to extract the high-order semantic features of the image.
[0010] Preferably, the improved U 2-Net network application SE attention module performs global average pooling operation on the hierarchical feature map in each parallel convolution layer of the symmetric encoder-decoder to compress the global spatial information to generate a channel description vector; the channel description vector is passed through two fully connected layers and the Sigmoid function to generate an attention vector containing the channel importance weight; a 3×3 deep convolution layer is used to independently process the spatial information of each channel of the attention vector to capture the interaction of local neighborhood features; and then a 1×1 convolution layer is used to fuse the cross-channel information point by point on the local neighborhood feature interaction results to integrate multi-scale and multi-semantic level feature responses to obtain multi-scale and multi-semantic level feature responses that complete pixel-level information integration.
[0011] Preferably, the cartoon style image to be processed is input into the improved U 2 -Net network, it also includes training the network through an inhibitory loss function based on response distribution analysis and an improved cross entropy loss function.
[0012] Preferably, the cartoon style image to be processed is input into the improved U 2 Before the -Net network, it also includes preprocessing of the anime-style image to be processed for lighting balance and feature enhancement, unifying the size of the preprocessed image to obtain standardized input data, and inputting the standardized input data into the improved U 2 -Net network.
[0013] The present invention also proposes a line drawing extraction system for anime-style images, comprising: An image acquisition module, used to obtain the anime-style image to be processed; A feature extraction module is used to extract multi-scale edge response features of anime-style images through parallel convolution paths to generate a hierarchical feature map with spatial perception capabilities. The multi-scale edge response features are used to represent low-order and high-order features of the image from different receptive fields. The low-order features include the location, direction, and connectivity of contour lines, and the high-order features include semantic information about the background and characters. The feature enhancement module is used to perform deconvolution reconstruction on high-order features in the hierarchical feature map and channel-join the reconstructed high-order features with adjacent low-order features. The joined features are fused at the pixel level across hierarchical responses through point convolution. At the same time, the edge distortion of the high-order features is corrected through the geometric constraints of the low-order features to obtain a complementary enhanced feature expression. The feature fusion module is used to dynamically fuse the complementary enhanced feature expressions to obtain multi-scale and multi-semantic level feature responses that integrate pixel-level information; The line drawing extraction module is used to output editable line drawings by completing multi-scale and multi-semantic level feature responses that integrate pixel-level information.
[0014] The present invention also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement any one of the steps in the method for extracting line drawings from animation images.
[0015] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is loaded by a processor, it can execute any one of the steps in the method for extracting line drawings from animation images.
[0016] The method for extracting line drawings from cartoon images provided by the present invention has the following beneficial effects: The present invention extracts multi-scale edge response features through parallel convolution paths, combines the construction of hierarchical feature maps, directly captures low-order features and high-order features, and avoids the excessive abstraction of shallow features by deep networks; at the same time, the geometric constraints of low-order features are used to correct the edge distortion of high-order features, alleviating the attenuation of feature information during transmission. After the high-order features are deconvoluted and reconstructed, they are spliced with adjacent low-order feature channels and pixel-level fusion is achieved through point convolution. The detail information of low-order features is used to supplement the detail loss of high-order features, balancing the difference between the high noise of low-order features and the lack of details of high-order features, and achieving complementary enhancement. Through dynamic fusion of complementary enhanced feature expressions, pixel-level information integration is completed, and the different receptive field advantages of multi-scale edge response features are combined to improve the fusion effect of local details and global structure, avoiding feature blur or structural disorder after fusion. Through precise cross-level feature fusion and geometric constraint correction, the interference of background semantic information on the edge in high-order features is reduced, the influence of background artifacts on line drawing extraction is suppressed, and the accuracy and precision of line drawing extraction are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the embodiments of the present invention and its design, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.
[0018] Figure 1 This is a flowchart of the method for extracting line drawings from cartoon images according to Example 1 of the present invention; Figure 2 The improved U 2 -Net model structure diagram; Figure 3 Extract detailed comparison results for model line draft; Figure 4 This is a complete flow chart of the method for extracting line drawings from animation images of the present invention. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the technical solution of the present invention and to be able to implement it, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and are not intended to limit the scope of protection of the present invention.
[0020] Example 1 The present invention provides a method for extracting line drawings from cartoon images, specifically Figure 1 and Figure 4 As shown, the following steps are included: Step 1: Collaborate with a professional animation team to build the first high-precision hand-drawn line drawing dataset ArtLine-2K. This dataset contains 2,000 sets of rendering-line drawing pairs, which were expanded to 10,000 pairs through data augmentation. It covers a variety of animation styles to ensure the diversity and generalization ability of model training.
[0021] Step 2: Perform preprocessing on the artistic image (including lighting balancing and feature enhancement) to improve image quality and enhance feature identifiability, unify the data size, and provide standardized input for subsequent feature extraction and model training.
[0022] Step 3: Construct a multi-branch feature extraction network based on U 2 -Net nested U-shaped structure improvement: parallel convolution layers and skip connections are used to extract multi-scale edge response features, generate hierarchical feature maps, and introduce deconvolution to improve the accuracy of hierarchical feature fusion. 2-Net is a multi-branch feature extraction network with an improved nested U-shaped structure. It extracts multi-scale edge response features from images through parallel convolutional layers and skip connections, generating hierarchical feature maps with spatial awareness. Deconvolution is also introduced to improve the accuracy of hierarchical feature fusion. Specifically, multi-scale edge response features of anime-style images are extracted through parallel convolutional paths to generate hierarchical feature maps with spatial awareness. Multi-scale edge response features are used to represent low-level and high-level features of the image from different receptive fields. Low-level features include the position, direction, and connectivity of contour lines, while high-level features include semantic information about the background and characters. Capturing semantic information improves the accuracy and structure of line drawing extraction, specifically in the following ways: Optimizing edge discrimination: By capturing semantic information from high-level features (such as abstract semantics such as the body structure and clothing outlines of anime characters) and combining it with geometric details from low-level features (such as line direction and local edges), edge distortion in high-level features can be corrected, avoiding spurious edge responses caused by relying solely on low-level features (such as false lines caused by complex texture interference), and enhancing the robustness of valid line discrimination. Ensure the semantic continuity of lines: Line drawings must be structured and semantically coherent (such as the integrity of character outlines and the logic of object boundaries). Capturing semantic information allows the model to understand the semantic associations between different elements in the image (such as the "connection between hair and face" and "matching of clothing wrinkles with body movements"), thereby maintaining the overall structural continuity of the lines when extracting line drawings, avoiding problems such as broken outlines and structural disorder. Improve the ability to handle complex scenes: In the complex backgrounds of anime-style images (such as environmental shadows and rich textures), semantic information can help the model distinguish between "target lines" (such as character edges) and "background interference" (such as irrelevant textures and light and shadow noise), reduce the generation of background artifacts, and ensure that the extracted line drawings focus on edge information that is critical to the visual structure, meeting the demand for high-fidelity line drawings in professional creation.
[0023] To improve U 2 -Net's multi-branch feature extraction network is the backbone, and its feature extraction module can generate each side response, and each side response is expressed as follows . Given that line drawings belong to low-level visual features (such as edge contours, line directions and local textures), it is necessary to avoid excessively deep levels that may lead to over-abstraction of shallow features or information attenuation. Experiments show that when the network reaches the fourth level, the response of the deep convolution kernel is significantly attenuated, but after removing the fourth layer, the feature extraction capability is weakened due to insufficient depth. For this reason, a dynamic fusion strategy for adjacent-level features is proposed: first, a deconvolution operation is performed on the high-order features (learnable upsampling kernel parameters are introduced to enhance reconstruction controllability and reduce edge errors by increasing nonlinearity), and the reconstructed high-order features are channel-spliced with the adjacent low-order features transmitted by jump connections. Finally, pixel-level fusion of cross-level responses is achieved through point convolution. While retaining the multi-scale advantages, low-order geometric constraints are used to correct high-order edge distortions, forming a complementary and enhanced feature expression.
[0024] In the improved model, deconvolution and skip connections are added to U 2 -Net's multi-branch feature extraction network part is specifically distributed among the various levels of its nested U-shaped structure.
[0025] U 2- Net's parallel convolutional layers are distributed across different levels. These layers perform convolution operations on the input image in parallel by setting convolution paths with different receptive fields, such as using convolution kernels of different sizes or different strides. For example, in lower-level parallel convolutional layers, convolution paths with small receptive fields are responsible for capturing local image details, such as subtle details like line turns and endpoints. Meanwhile, at higher levels, convolution paths with large receptive fields focus on global image structural features, such as the overall outline of objects and the general direction of lines. These parallel convolution paths operate simultaneously, extracting features from images at different scales, generating feature responses rich in multi-scale information, and thus constructing hierarchical feature maps with multi-scale perception capabilities.
[0026] Skip connections also exist in U 2 -Net nested between the layers of the U-shaped structure. Its main function is to directly pass the low-level features extracted by the shallow network (containing rich original spatial detail information, such as basic edge contours, the preliminary direction of lines, etc.) to the deep network. For example, in the process of transferring features from the shallow layer to the deep layer of the network, the skip connection will pass the information about the basic structure of the image extracted by the shallow network directly to the deep network, bypassing the intermediate multi-layer network. In this way, while the deep network extracts and abstracts complex features, it can also combine the basic features of the shallow layer, effectively avoiding the problem of information loss or attenuation caused by the continuous abstraction of shallow features in the deep network. The final feature map contains both deep semantic information and shallow spatial details, providing a more comprehensive and accurate information foundation for subsequent feature fusion and processing.
[0027] In the improved model structure, parallel convolutional layers and skip connections work together to acquire and integrate image features from different scales to generate high-quality hierarchical feature maps, providing rich information for subsequent line drawing extraction tasks. Their specific collaborative mechanism is as follows: The parallel convolution layer is responsible for multi-scale feature extraction: the parallel convolution layer sets up convolution paths with different receptive fields, and performs convolution operations on the input image in parallel by adopting convolution kernels of different sizes or different step sizes. At the lower level, the convolution path with a small receptive field focuses on capturing the local details of the image, such as the subtle turns and endpoints of lines; at the higher level, the convolution path with a large receptive field focuses on the global structural features of the image, such as the overall outline of the object and the general direction of the lines. In this way, the parallel convolution layer can extract features from images at different scales, generate feature responses containing rich multi-scale information, and construct a hierarchical feature map with multi-scale perception capabilities.
[0028] Skip connections transfer shallow-layer features: Skip connections transfer low-level features extracted by shallow networks (containing rich raw spatial details, such as basic edge contours and the initial direction of lines) directly to deep networks. As features are transferred from shallow to deep layers, skip connections bypass the intermediate layers and pass shallow-layer basic feature information to the deep network. This effectively avoids the loss or attenuation of shallow-layer features caused by their continuous abstraction in deep networks, allowing deep networks to extract and abstract complex features while still integrating shallow-layer basic features.
[0029] Collaborative optimization of cross-level features: The multi-scale features extracted by the parallel convolutional layers generate corresponding hierarchical feature maps between network layers, while the shallow low-level features transmitted by the jump connection interact with the high-level features (containing more abstract semantic information) extracted by the parallel convolutional layers in the deep network. High-level features can correct their own edge distortion problems with the help of the geometric constraints of low-level features. For example, when generating line drawings, the edges of the lines can be made more accurate. Low-level features can also enhance their focus on key edges under the semantic guidance of high-level features, such as more accurately identifying the edges of object contours. Ultimately, through this collaboration, complementary enhancement of cross-level features is achieved, providing more comprehensive and accurate multi-scale feature inputs for the subsequent dynamic feature fusion module, thereby improving the accuracy and quality of line drawing extraction.
[0030] Multi-scale is achieved with the help of parallel convolution paths. Different convolution paths have different receptive fields (such as through different sizes of convolution kernels or step size settings), and can focus on areas of different sizes in the image. For example, the convolution path with a small receptive field can capture local detail edge features such as subtle turns and endpoints of lines, while the path with a large receptive field can focus on global structural edge features such as the overall outline of the object and the overall direction of the lines. Edge response features refer to the characteristic expression of edge information in the image (such as object contours, line boundaries, etc.). These features can reflect key information such as the existence, position, direction, and connectivity of the edge. Hierarchical feature maps with spatial perception capabilities lay the foundation for subsequent cross-level feature optimization and fusion, ultimately helping to generate high-quality line drafts.
[0031] In this embodiment, deconvolution is introduced to improve the traditional upsampling operation. The traditional upsampling operation is a linear operation. Deconvolution achieves a better mapping from low-resolution feature maps to high-resolution feature maps by introducing nonlinear operations. On the basis of extracting multi-scale edge response features with the help of parallel convolution layers and jump connections to generate hierarchical feature maps, deconvolution upsamples these feature maps, restores the spatial size of the feature maps, and enables the feature maps to contain more spatial detail information. For example, if a feature map of a certain level is reduced in size due to a convolution operation, deconvolution can expand its size, allowing the originally compressed spatial information to be expanded, thereby providing support for the subsequent accurate perception of the spatial position, edge direction, etc. of objects in the image.
[0032] The feature fusion optimization operation is a fusion of the upsampled feature maps generated by the deconvolution operation with feature maps from other levels (such as feature maps transmitted via skip connections). Feature maps at different levels contain information at different levels of abstraction and scale. During the fusion process, the feature map generated by deconvolution can complement the spatial detail features it carries, restored through upsampling, with the information from other feature maps. For example, high-level feature maps contain more semantic information, while low-level feature maps retain more of the original spatial details. The deconvolution-generated feature map participates in the fusion, allowing high-level features to better integrate low-level spatial details while maintaining semantic understanding, thereby optimizing the accuracy of hierarchical feature fusion. Through this fusion optimization, the resulting hierarchical feature map can more comprehensively reflect information such as the spatial layout of objects in the image and the spatial continuity of edges, thereby possessing stronger spatial perception capabilities.
[0033] Step 4: Use a multi-branch feature extraction network to generate a hierarchical feature map with spatial perception capabilities. Optimize cross-level features through scale-adaptive response, local receptive field feature aggregation, and pixel-level semantic alignment. Combine the geometric constraints of low-order features to correct the edge distortion of high-order features and achieve complementary and enhanced feature expression. Through scale-adaptive response, local receptive field feature aggregation, and pixel-level semantic alignment, perform deconvolution on high-order features, channel-join the reconstructed high-order features with the adjacent transmitted low-order features, and then achieve pixel-level fusion of cross-level responses through point convolution. Combine the geometric constraints of low-order features to correct the edge distortion of high-order features and form a complementary and enhanced feature expression. Use the geometric constraints of low-order features to correct the edge distortion of high-order features and achieve complementary and enhanced feature expression to address the differences in responses at different levels.
[0034] Step 5: Apply the dynamic feature fusion module to optimize the integration of cross-level responses through learnable attention weights: embed the SE attention module in each side fusion path, and use parallel two-layer depthwise separable convolution - the first layer 3×3 convolution realizes local feature interaction, and the second layer 1×1 point-by-point convolution establishes cross-channel association to complete the pixel-level information integration of multi-scale features.
[0035] Specifically, to solve the problem of representation differences between low-level features and high-level semantic features, such as Figure 2 As shown, the improved U 2 -Net introduces a dynamic fusion mechanism to achieve more effective feature fusion through adaptive feature guidance and multi-scale interaction. Specifically, the network first introduces an attention mechanism based on the SE module in the side path. This module compresses the input features through global average pooling and applies attention weights on the compressed channel dimension to enhance the importance distinction between channels. After generating the fusion guidance feature, the system further cascades the feature maps of each branch on the channel dimension to form the initial dynamic fusion feature. The constructed fusion features are then adjusted through a dual-branch convolution path for feature mapping. The two branches use different depth-wise separable convolution kernels, 3×3 and 1×1, respectively, to achieve local perception and cross-channel interaction. The calculation method is as follows: , Among them, σ represents the activation function, DWConv n×nrepresents a depthwise separable convolution operation. This structural design enables the network to simultaneously focus on local details and global contextual structure during the fusion process, achieving multi-scale, cross-dimensional semantic compensation and edge fidelity. On this basis, the fused features are input into the subsequent network structure for end-to-end training, effectively improving the model's ability to construct discriminative features across different layers. The proposed dynamic fusion strategy not only enhances the discriminability of feature representation, but also improves the network's understanding of multi-dimensional content in the image and its ability to model the structure.
[0036] Step 6: Design a joint supervision mechanism: To address the problem of background artifacts, an inhibitory loss function based on response distribution analysis is proposed, which imposes a penalty on a pixel-by-pixel basis according to the degree of deviation of the false response; the improved cross-entropy loss function is combined to optimize line quality and improve the model convergence speed, accuracy, and robustness.
[0037] Specifically, when using cross entropy as the basic supervisory signal training model, two key problems are exposed: one is that pseudo-edge fragments unrelated to the target structure are easily generated in the background area; the other is that the edge lines are affected by factors such as color unevenness, and the prediction quality is unstable. In response to the former, the present invention proposes a background suppression loss function based on the distribution of predicted responses. By constructing a pixel-level misjudgment detection mechanism, the network is guided to suppress invalid responses. This mechanism performs statistical analysis on the predicted output of each pixel, identifies its false response degree, and assigns different weights according to the magnitude of its deviation from the true label, thereby realizing a differentiated penalty strategy. The loss term is defined as follows: , where the pixel loss weight function is defined as: , It represents the global expectation of the background pseudo-label, which is used to balance the contribution of each pixel loss under different background complexities. To address the blurring and unevenness problems in line prediction, a collaborative optimization strategy combining temperature scaling and label smoothing is proposed. The specific form is: , , The method first converts the original label y∈{0,1} into a soft label ,in is a small constant used to control the smoothness of the label distribution, which can effectively improve the model's support for extreme value predictions and improve the calibration of probability outputs. At the same time, the model output is scaled by introducing a temperature factor τ at the output end to widen the difference in the probabilities of each category in the predicted distribution, thereby enhancing the model's ability to distinguish in high confidence areas. In order to balance the stability of the model in the early stages of training and the refined convergence in the later stages, this method designs a gradual cooling strategy: the high temperature is set in the initial training stage. The space is explored with expanded parameters, and then the temperature is linearly decayed to every 10 epochs , thereby enhancing the model's ability to learn extreme prediction areas. Finally, the overall training process is optimized by the following joint loss function: ,in, , is an adjustable weight parameter and , which is used to dynamically control the optimal balance between background suppression and edge accuracy.
[0038] Step 7: Train the model using the ArtLine-2K dataset to obtain a high-fidelity line art extraction model. Apply the trained model to actual line art extraction tasks to achieve automated high-fidelity line art output. During training, the model automatically adjusts network parameters by learning from rendering-line art pairs to minimize the loss function.
[0039] The present invention avoids excessively deep hierarchical structures by controlling the network depth and adopting U 2 -Net's encoder-decoder architecture. This architecture uses its nested U-shaped layers and skip connection mechanism to capture shallow features (such as edge contours and line directions) directly from each encoder stage, bypassing the traditional backbone network, reducing computational overhead and alleviating feature loss.
[0040] In view of the differences in responses at different levels (such as high noise in low-order responses and loss of details in high-order responses), a dynamic fusion strategy for inter-level features is proposed: first, deconvolution operation is performed on the high-order features (learnable upsampling kernel parameters are introduced to enhance controllability and nonlinearity, and reduce edge errors), and then the reconstructed high-order features are channel-wise spliced with the low-order features transmitted by adjacent jump connections. Finally, pixel-level fusion of cross-level responses is achieved through point convolution. The edge distortion of high-order features is corrected with the help of the geometric constraints of low-order features, achieving complementary and enhanced feature expression.
[0041] To address the problem of low-level features being affected by similar textures and causing spurious edge responses, a channel attention mechanism (SE module) is embedded in the side aggregation module. This mechanism performs the following steps: a global average pooling (GAP) operation is performed on the input feature map to compress global spatial information and generate a compact channel description vector. This vector is then fed into a bottleneck structure consisting of two fully connected (FC) layers (in between which a nonlinear activation function such as ReLU is used to transform information). Finally, a sigmoid function is used to generate an attention vector containing channel importance weights. This vector is applied to the original input feature map to perform adaptive feature recalibration based on the overall semantic content of the image, significantly suppressing invalid texture responses and enhancing the focus on discriminative edge cues, thereby improving the robustness of edge discrimination of low-level features against complex texture backgrounds.
[0042] To address the issue of insufficient multi-scale feature fusion quality, a two-stream depthwise separable convolution is applied to the spliced composite feature map. The first layer uses a 3×3 depthwise convolution layer to independently process spatial information in each channel, focusing on capturing local neighborhood feature interactions (such as edge orientation and connectivity), maximizing the preservation of original spatial detail while ensuring computational efficiency. The second layer uses a pointwise convolution layer to fuse cross-channel information output by the depthwise convolution, integrating multi-scale and multi-semantic feature responses to enhance the modeling of the overall edge contour and global structure. This two-stream design synergistically optimizes the key requirements of local detail clarity and global structural coherence in edge detection tasks.
[0043] To address the problem that the model generates pseudo-edge fragments (background artifacts) such as discrete point noise and short line segment aggregation in complex background areas (such as clothing wrinkles or environmental shadows), an inhibitory loss function based on response distribution analysis is proposed: first, background pixels are marked by the indicator function I_bg; then the cross-entropy loss L_ce of the background pixels is calculated; finally, a pixel-level dynamic penalty factor α is introduced (calculated based on the deviation |P_i-Y_i| between the predicted value P_i and the true value Y_i), and the loss contribution is weighted with an adjustable penalty coefficient β; this function imposes severe penalties on highly deviated background pixels, significantly suppressing background artifacts and reducing the model's sensitivity to interference such as water stains and shadows.
[0044] To address the problem of low tolerance for sub-pixel position deviations and uneven line color caused by vanishing gradients due to the traditional cross-entropy loss's label binarization (0 / 1), a collaborative solution of temperature scaling and label smoothing is adopted. First, hard labels are converted into soft labels Y_smooth = (1-ε) × Y + ε / K (ε is the smoothing strength parameter), which improves probability calibration and enhances support for correct extreme predictions. Second, a temperature coefficient T is introduced on the prediction side to scale the logical output P_temp = softmax(logits / T), amplifying the difference in probability distribution to improve the model's confidence in clear lines. To address the conflict between training stability and convergence speed in the collaborative optimization of temperature scaling and label smoothing, a phased temperature ramp-up strategy was designed: Initially, a high temperature T is set to broaden the parameter search space, and then T is linearly reduced to the target value every 10 epochs. This strategy allows the model to freely explore the parameter space initially, while later focusing on enhancing the confidence of sharp lines, ultimately achieving efficient convergence in approximately 30 minutes.
[0045] Through the systematic improvements to the aforementioned feature extraction, fusion mechanism, and loss design, the model ultimately achieves high-fidelity line drawing extraction for complex anime-style artistic images. The generated line drawings, after evaluation by professional artists, can be directly used for secondary creation, providing high-quality structured input for subsequent artistic creation tasks such as anime coloring, image style transfer, and digital illustration generation.
[0046] Figure 3 The final rendering is shown in Figure 1. Image is the original anime-style image, label is the line drawing drawn manually by the artist based on the original anime-style image. A 2048×2048 line drawing requires a full day to draw manually. Ours is the line drawing generated by the model of our invention based on the original anime-style image. The red-boxed portion is a comparison of the line drawing drawn by the artist and the line drawing generated by our invention after magnification 10x. The comparison shows only minor differences and can be directly given to the artist for secondary processing.
[0047] The present invention adopts U 2 -Net architecture and jump connection, directly capture shallow features, reduce computational costs and feature loss, and provide more complete basic features for subsequent processing. A dynamic fusion strategy for adjacent features is proposed, and cross-level pixel-level fusion is achieved through deconvolution, channel splicing and point convolution. The edge distortion of high-order features is corrected with the help of low-order features to enhance feature complementarity. The SE attention module is embedded to suppress invalid textures through channel weight adjustment, enhance the ability to focus on edge clues, and improve the robustness of edge discrimination under complex backgrounds. A two-stream deep separable convolution is used to capture local feature interactions and cross-channel information respectively, optimize the balance between local details and global structures, and enhance edge modeling capabilities. Ultimately, end-to-end conversion of complex anime images to structured line drafts is achieved, generating high-fidelity line drafts with continuous lines and no artifacts to meet the needs of professional secondary creation.
[0048] Based on the same inventive concept, the present invention also proposes a line drawing extraction system for anime-style images, comprising: The image acquisition module is used to obtain the anime-style images to be processed.
[0049] The feature extraction module is used to extract multi-scale edge response features of anime-style images through parallel convolution paths to generate hierarchical feature maps with spatial perception capabilities. Multi-scale edge response features are used to represent low-order and high-order features of images from different receptive fields. Low-order features include the position, direction, and connectivity information of contour lines, while high-order features include semantic information of background and characters.
[0050] The feature enhancement module is used to perform deconvolution reconstruction operations on high-order features in the hierarchical feature map, and channel-splicing the reconstructed high-order features with adjacent low-order features; the spliced features are subjected to pixel-level fusion of cross-level responses through point convolution, and the edge distortion of high-order features is corrected through the geometric constraints of low-order features to obtain complementary enhanced feature expression.
[0051] The feature fusion module is used to dynamically fuse complementary enhanced feature expressions to obtain multi-scale and multi-semantic level feature responses that integrate pixel-level information.
[0052] The line drawing extraction module is used to output editable line drawings by completing multi-scale and multi-semantic level feature responses that integrate pixel-level information.
[0053] Each module in the aforementioned line drawing extraction system for anime-style images can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0054] The present invention also provides a computer device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the embodiment of the method for extracting line drawings from an animation image. The specific implementation method can be found in the method embodiment and will not be repeated here.
[0055] Furthermore, the present invention provides a non-transitory computer-readable storage medium containing instructions, wherein the storage medium stores a computer program. For example, this may be a memory device containing instructions, wherein the instructions are executable by a processor of a computer device to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device. When executed by the processor, this computer program is capable of implementing the steps in the embodiments of the method for extracting line art from an animation image. The specific implementation method can be found in the method embodiments and will not be further described here.
[0056] Those skilled in the art will appreciate that embodiments of the present invention may provide methods, systems, or computer program products. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0057] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0058] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0060] It should be pointed out that the specific implementation methods described above can enable those skilled in the art to understand the invention more comprehensively, but do not limit the invention in any way. Therefore, although the present specification and examples have described the invention in detail, those skilled in the art should understand that the invention can still be modified or replaced by equivalents; and all technical solutions and improvements that do not deviate from the spirit and scope of the invention are included in the scope of protection of the patent for the invention. Any figure mark in the claims should not be regarded as limiting the claims involved. Any simple change or equivalent replacement of the technical solution that can be obviously obtained by any person familiar with the art within the technical scope disclosed in the present invention falls within the scope of protection of the present invention.
Claims
1. A method for extracting line drawings from an animation image, characterized in that: include: Get the anime-style image to be processed; Multi-scale edge response features of anime-style images are extracted through parallel convolution paths to generate hierarchical feature maps with spatial perception capabilities. The multi-scale edge response features are used to represent low-order and high-order features of the image from different receptive fields. The low-order features include the location, direction, and connectivity of contour lines, and the high-order features include semantic information about the background and characters. Perform deconvolution reconstruction on the high-order features in the hierarchical feature map, and perform channel splicing on the reconstructed high-order features and adjacent low-order features; The spliced features are fused at the pixel level through point convolution to achieve cross-level responses. At the same time, the edge distortion of high-order features is corrected through the geometric constraints of low-order features to obtain complementary enhanced feature expression. Dynamically fuse the complementary enhanced feature expressions to obtain multi-scale and multi-semantic level feature responses that integrate pixel-level information; By integrating pixel-level information into multi-scale and multi-semantic feature responses, editable line drawings are output.
2. The method for extracting line drawings from an animation image according to claim 1, wherein: Input the anime-style image to be processed into the improved U 2 -Net network, outputs multi-scale and multi-semantic level feature responses that complete pixel-level information integration; the improved U 2 -Net network with U 2 -Net nested U-shaped structure is used as the basis, deconvolution is added to the input of each parallel convolutional layer of the symmetric encoder-decoder after the input convolutional layer, and the SE attention module is embedded at the output of each parallel convolutional layer.
3. The method for extracting line drawings from an animation image according to claim 2, wherein: The improved U 2 The parallel convolutional layers of the -Net network set up convolution paths with different receptive fields. By adopting convolution kernels of different sizes or different step sizes, convolution operations are performed in parallel on the anime-style images to be processed, features are extracted from images at different scales, feature responses containing multi-scale information are generated, and a hierarchical feature map with spatial perception capabilities is constructed.
4. The method for extracting line drawings from an animation image according to claim 2, wherein: The improved U 2 -Net network's skip connection layer passes the low-order features extracted by the shallow network directly to the deep network, bypassing the middle multi-layer network. The shallow network is used to capture the low-order visual features of the image, and the deep network is used to extract the high-order semantic features of the image.
5. The method for extracting line drawings from an animation image according to claim 3, wherein: The improved U 2 -Net network applies SE attention module to perform global average pooling operation on the hierarchical feature map in each parallel convolution layer of symmetric encoder-decoder, compressing global spatial information to generate channel description vector; the channel description vector is passed through two fully connected layers and sigmoid function to generate attention vector containing channel importance weight; 3×3 deep convolution layer is used to independently process the spatial information of each channel of attention vector to capture the interaction of local neighborhood features; Then, the cross-channel information is fused through a 1×1 point-by-point convolutional layer based on the interaction results of local neighborhood features, and the multi-scale and multi-semantic level feature responses are integrated to obtain the multi-scale and multi-semantic level feature responses that complete the pixel-level information integration.
6. The method for extracting line drawings from an animation image according to claim 5, wherein: Input the anime-style image to be processed into the improved U 2 -Net network, it also includes training the network through an inhibitory loss function based on response distribution analysis and an improved cross entropy loss function.
7. The method for extracting line drawings from an animation image according to claim 2, wherein: Input the anime-style image to be processed into the improved U 2 Before the -Net network, it also includes preprocessing of the anime-style image to be processed for lighting balance and feature enhancement, unifying the size of the preprocessed image to obtain standardized input data, and inputting the standardized input data into the improved U 2 -Net network.
8. A line drawing extraction system for anime-style images, characterized by: include: An image acquisition module, used to obtain the anime-style image to be processed; A feature extraction module is used to extract multi-scale edge response features of anime-style images through parallel convolution paths to generate a hierarchical feature map with spatial perception capabilities. The multi-scale edge response features are used to represent low-order and high-order features of the image from different receptive fields. The low-order features include the location, direction, and connectivity of contour lines, and the high-order features include semantic information about the background and characters. The feature enhancement module is used to perform deconvolution reconstruction on the high-order features in the hierarchical feature map and perform channel splicing on the reconstructed high-order features and adjacent low-order features; The spliced features are fused at the pixel level through point convolution to achieve cross-level responses. At the same time, the edge distortion of high-order features is corrected through the geometric constraints of low-order features to obtain complementary enhanced feature expression. The feature fusion module is used to dynamically fuse the complementary enhanced feature expressions to obtain multi-scale and multi-semantic level feature responses that integrate pixel-level information; The line drawing extraction module is used to output editable line drawings by completing multi-scale and multi-semantic level feature responses that integrate pixel-level information.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is loaded into a processor, it is capable of executing the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Mural line draft extraction method based on multi-scale extraction and in-layer learnable fusion
CN118968241A
Intelligent cartoon line manuscript generation system capable of keeping human face features
CN119479044A
Street view image semantic segmentation method based on improved U-Net network
CN117392676A
Multi-scale multi-perception real-time image segmentation method and system, terminal and medium
CN117576118A
Generating digital paintings utilizing an intelligent painting pipeline for improved brushstroke sequences
US20230316590A1