Cartoon character analysis method based on space learning and structure modeling

By adopting spatial learning and structural modeling methods in cartoon animal analysis, combining deformable convolution, multi-task prediction strategies and graph neural networks, efficient capture and analysis of complex features of cartoon animal is achieved, and the problem of low resolution accuracy in the existing technology is solved.

CN120088610APending Publication Date: 2025-06-03SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411274318.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art has low resolution accuracy when processing cartoon animal images with irregular body structures, complex visual appearances and abstract painting styles.

Method used

Using a method based on spatial learning and structural modeling, the spatial characteristics of cartoon animals are obtained through deformable convolution and multi-task center point and edge prediction strategies, and structural information is modeled through graph neural networks and shape-aware relationship networks. At the same time, a cross-attention mechanism is used to achieve consistent learning between spatial and structural features.

Benefits of technology

It improves the accuracy and robustness of cartoon animal analysis, and can better capture and process complex and diverse cartoon animal characteristics, achieving efficient analysis of diverse and complex cartoon animals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088610A_ABST
    Figure CN120088610A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a cartoon role analysis method based on space learning and structure modeling, which is used for cartoon role analysis based on space learning and structure modeling. According to the method, the problems of low analysis accuracy of different cartoon roles and the like in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and specifically, to a cartoon character parsing method based on spatial learning and structure modeling. Background Art

[0002] In recent years, significant progress has been made in human parsing methods, and many methods for parsing human body parts have been proposed, including methods based on spatial learning and methods based on structure modeling. Spatial information plays a crucial role in the parsing task, as it provides important information about the positions of body parts. By considering the spatial distribution of the human body, it helps to identify complex regions, thereby achieving more accurate parsing results. In addition to spatial information, structural information is also important for human parsing. To overcome the limitations of traditional convolutional neural network methods, previous work has proposed structure modeling methods. These methods combine the advantages of convolutional neural networks and graph models to extract visual features and model structural information. However, these techniques customized for human parsing have limited performance when dealing with cartoon animal images with irregular body structures, complex visual appearances, and abstract painting styles.

[0003] In recent years, the field of cartoon parsing has attracted increasing attention. Traditional cartoon parsing methods segment cartoon images into different regions through techniques such as adaptive region propagation merging, but each segmented region lacks semantic distinction. Recently, some pioneering work has applied deep learning-based human parsing techniques to anthropomorphic cartoon characters and cartoon dogs, achieving good results. However, these methods have limitations because they mainly focus on cartoon characters in human form or single-category cartoon characters (such as dogs), and have low parsing accuracy for the appearance changes and structural complexities of different cartoon animals (such as crocodiles and butterflies). Summary of the Invention

[0004] To overcome the deficiencies of the prior art, the present invention provides a cartoon character parsing method based on spatial learning and structure modeling, which solves the problems in the prior art such as low parsing accuracy for different cartoon characters.

[0005] The technical solution adopted by the present invention to solve the above problems is:

[0006] A cartoon character parsing method based on spatial learning and structure modeling, which performs cartoon character parsing based on spatial learning and structure modeling.

[0007] As a preferred technical solution, the features X of the cartoon character are respectively processed by a spatial learning branch and a structure modeling branch, and then spatial structure consistency learning is performed.

[0008] As a preferred technical solution, during the processing of the spatial learning branch, deformable convolution is used to obtain the spatial features of the cartoon character, multi-task edge is used to obtain edge-aware spatial features, and the center point prediction strategy is used to obtain center-aware spatial features.

[0009] As a preferred technical solution, during the processing of the spatial learning branch, the feature X of the cartoon character is obtained through deformable convolution to obtain irregular spatial features, and the center point prediction map is obtained through the center point predictor, and the contour prediction map is obtained through the contour predictor; the features processed by the deformable convolution are input into the channel attention for feature refinement processing to obtain the feature X 1 ; the features from the center point predictor and the contour predictor are refined using channel attention combined with spatial attention to obtain the feature X 2 ; finally, X 1 、X 2 are concatenated and fused to obtain the spatial feature X that can perceive the center point, edge, and irregular shape of the body part sp .

[0010] As a preferred technical solution, during the processing of the structure modeling branch, a graph construction module is used to construct the graph structure representation of the cartoon character, and the shape-aware graph neural network is used to model the structure representation of the cartoon character and obtain the relationship between body parts; among them, the shape-aware graph neural network includes shape convolution and self-attention mechanism, and the shape convolution is used to obtain the shape information of the graph nodes, and the self-attention mechanism is used to associate the graph nodes with the shape information.

[0011] As a preferred technical solution, during the processing of the structure modeling branch, first, the convolutional neural network features are processed, and the graph construction module composed of convolutional layer convolution and feature map division operations is used to construct the graph structure representation of the cartoon character; then, the shape-aware graph neural network is used to model the graph node relationship in the graph structure representation to obtain the shape-aware graph convolution features; then, by concatenating and fusing the paired graph nodes, and then through shape convolution processing to obtain shape features, and then through the self-attention mechanism for graph node information association learning guided by shape information; finally, the graph convolution features are mapped to convolutional features to obtain the structure feature X st .

[0012] As a preferred technical solution, when performing spatial-structure consistency learning, the cross-attention mechanism is used to associate the spatial and structure features, and the consistent feature representation is learned through the fusion of spatial and structure information.

[0013] As a preferred technical solution, when performing spatial-structure consistency learning, the query-key matching mechanism is used to learn the correspondence between the spatial and structure features, and the cross-attention mechanism is used to adjust the fusion weights between the spatial and structure features.

[0014] As a preferred technical solution, when performing spatial structure consistency learning, the spatial information X sp is used as a query, the structural feature X st is used as a key, and the concatenated and fused feature X cat = cat(X sp , X st ) is used as a value. The formula for performing spatial structure consistency learning is:

[0015] Q = X sp W q , K = X st W k , V = X cat W v

[0016] X fused = CrossAtten(Q, K, V)

[0017] where Q represents the query after learnable weight mapping, K represents the key after learnable weight mapping, V represents the value after learnable weight mapping, W q represents the learnable weight corresponding to Q, W k represents the learnable weight corresponding to K, W v represents the learnable weight corresponding to V, CrossAtten represents the cross-attention function, cat represents the concatenation and fusion operation, and X fused represents the spatially and structurally consistent feature representation.

[0018] As a preferred technical solution, the cartoon image is generated successively through a convolutional network and a pyramid pooling module.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] (1) The present invention proposes a novel spatial learning and structure modeling network for cartoon animal parsing, which is dedicated to solving the challenges in spatial learning, structure modeling, and spatial-structure consistency learning;

[0021] (2) The present invention proposes a spatial learning branch with deformable convolution and multi-task center point and edge prediction strategies to capture and learn the complex and inconsistent spatial attributes in various cartoon animals, enhancing the spatial perception ability of the network;

[0022] (3) The present invention proposes a structure modeling branch that combines a graph neural network and a shape-aware relationship network to model structural information and capture the complex structural relationships inside cartoon animals;

[0023] (4) The present invention proposes a spatial-structure consistency learning strategy, which uses a cross-attention mechanism to achieve consistency learning between spatial features and structural features, alleviating the problems brought by the complexity and diversity of cartoon animals;

[0024] (5) The present invention achieves state-of-the-art performance on the cartoon parsing dataset, which proves the effectiveness of the proposed spatial learning and structure modeling network. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a schematic framework diagram of a cartoon character parsing method based on spatial learning and structure modeling according to the present invention;

[0026] Figure 2 is Figure 1 one of the partial enlarged views of

[0027] Figure 3 is Figure 1 the second of the partial enlarged views of DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.

[0029] Embodiment 1

[0030] As Figures 1 to 3 shown, in order to capture and learn the complex and inconsistent spatial attributes in various cartoon animals, a spatial learning branch with deformable convolution and a multi-task center point and edge prediction strategy is designed, which enhances the spatial perception ability of the network.

[0031] In order to model the structural information and capture the complex structural relationships inside the cartoon animals, a structure modeling branch is proposed, which combines a graph neural network and a shape-aware relationship network.

[0032] The present invention proposes a spatial-structure consistency learning strategy, which uses a cross-attention mechanism, aiming to achieve consistency learning between spatial features and structural features, and alleviating the problems brought by the complexity and diversity of cartoon animals.

[0033] To address the challenges in cartoon animal parsing, the present invention proposes a novel spatial learning and structure modeling network for cartoon animal parsing. The network aims to solve the key problems of spatial perception, structure modeling, and spatial-structure consistency learning in cartoon animal parsing. The spatial perception learning module integrates deformable convolutions to learn the spatial features of the irregular body parts of cartoon animals. Meanwhile, the module combines multi-task edge and center point prediction to capture complex spatial information. In addition, a structure modeling method is proposed to model the complex structure information of cartoon animals, which combines graph neural networks with a shape-aware relationship learning module. To alleviate the significant differences between different animals, the present invention proposes a spatial-structure consistency learning mechanism to capture and learn the feature correlations between different cartoon animals.

[0034] The purpose of this solution is to solve the problems of complex and diverse cartoon animal parsing and segmentation. This solution mainly consists of three modules, namely the spatial learning branch, the structure modeling branch, and the spatial-structure consistency learning. The overall architecture is as Figure 1 shown.

[0035] I. Spatial Learning Branch

[0036] Spatial information is very important for cartoon animal parsing. The distribution of the center points and edges of body parts is crucial for depicting the spatial information of cartoon animal images. To learn complex spatial information, the proposed method integrates deformable convolutions to go beyond the rigid receptive field of traditional convolutions and capture the spatial features of abstract cartoon images. In addition, it adopts a multi-task edge and center point prediction strategy to capture edge-aware and center-aware spatial features, enabling the network to adapt to distortions and abstractions in the spatial dimension.

[0037] Given the feature X generated by the convolutional network and the pyramid pooling module (PPM), the irregular spatial feature is obtained through deformable convolution (DC). In addition, the center point prediction map is obtained through the center point predictor, and the contour prediction map is obtained through the contour predictor. The feature processed by the deformable convolution is input into the channel attention (CA) for feature refinement. In addition, the extracted channel attention is combined with the spatial attention (SA) to refine the features from the center point predictor and the contour predictor. Finally, the features processed by the channel attention, the features from the center point predictor and the contour predictor, and the features processed by the channel and spatial attention are concatenated and fused to obtain the spatial feature X that can perceive the center points, edges, and irregular shapes of body parts. sp .

[0038] II. Structure Modeling Branch

[0039] Structure information, including the shapes and relationships of body parts, is also very important for cartoon animal parsing. To model the complex structural features of cartoon animals, a structure modeling module is designed. The proposed method designs a shape-aware graph neural network to model the graph structure representation of cartoon animals. The graph neural network is used to construct the structure of cartoon animals and capture the relationships between body parts. Shape convolution is used to capture the shape information of graph nodes. In addition, a self-attention mechanism is adopted to associate graph nodes with shape information.

[0040] In structure modeling learning, first, the convolutional neural network features are processed, and the graph structure representation is constructed through a graph construction module composed of a convolutional layer and a feature map partitioning operation. Then, the shape-aware graph neural network is used to process the graph node relationships of the graph structure representation, thereby obtaining shape-aware graph convolution features. The shape-aware graph neural network is composed of a combination of shape convolution and a self-attention mechanism. Pairs of graph nodes are first concatenated and fused, and then processed through shape convolution to obtain shape features. After that, the self-attention mechanism is used for graph node information association learning guided by shape information. Finally, the graph convolution features are mapped to convolutional features to obtain structure-aware features X st 。

[0041] III. Spatial-Structural Consistency Learning Module

[0042] The body parts of cartoon animals exhibit significant variations among different animal categories, which are usually caused by complex and inconsistent spatial and structural features. Therefore, a novel Spatial-Structural Consistency Learning (SSCL) mechanism is proposed, aiming to seamlessly integrate spatial and structural features and achieve consistency in spatial and structural dimensions. SSCL uses a cross-attention mechanism to associate spatial and structural features. It learns a consistent feature representation by fully utilizing complementary spatial and structural information. Specifically, spatial information is used as the query, while structural information is used as the key. The principle of this design is to use the query-key matching mechanism, aiming to learn the correspondence between spatial features and structural features. For diverse cartoon animals, their spatial and structural features may vary significantly. The cross-attention mechanism can adaptively adjust the fusion weights between these two features, effectively capturing the association between spatial and structural information. This produces a more consistent feature representation, better handling the overall features of cartoon animals.

[0043] In consistency learning, the spatial feature X sp is regarded as the query, the structural feature X st is used as the key, and the concatenated feature X cat = cat(X sp , X st) is used as the value. The formula for consistency learning is as follows:

[0044] Q = X sp W q , K = X st W k , V = X cat W v

[0045] X fused = CrossAtten(Q, K, V)

[0046] Among them, Q represents the query after being mapped by learnable weights, K represents the key after being mapped by learnable weights, V represents the value after being mapped by learnable weights, and W q represents the learnable weight corresponding to Q, and W k represents the learnable weight corresponding to K, and W v represents the learnable weight corresponding to V. CrossAtten represents the cross-attention function, and cat represents the concatenation and fusion operation. Through this aggregation mechanism, the network fuses spatial and structural information and obtains a spatially and structurally consistent feature representation X fused .

[0047] IV. Experimental Verification and Analysis

[0048] To verify the effectiveness of the proposed method, we conducted verification on a self-built cartoon animal parsing dataset. The experiment used the Mean Intersection over Union (Mean IoU) as the evaluation metric, which calculates the average intersection over union between the predicted body parts and the ground truth labels. In addition, the Mean Accuracy (Mean Acc.) is used to calculate the average accuracy, and the Pixel Accuracy (Pixel Acc.) measures the accuracy of correctly predicted pixels.

[0049] To evaluate the performance of the proposed method, we compared it with the state-of-the-art cartoon animal parsing methods and human parsing methods. The comparison results are listed in Table 1. Existing methods ignore the inherent differences between cartoon animals and real-world humans and also do not notice the differences between cartoon animals and common cartoon characters. Therefore, they have difficulties in cartoon animal parsing. To address the challenges in cartoon animal parsing, the proposed method captures important spatial information and complex structural features by introducing a spatial learning method and a structural modeling method. It deeply explores the basic features of cartoon characters by correlating spatial and structural features and realizes the parsing of diverse and complex cartoon animals. The proposed method achieved the highest results on the cartoon animal parsing dataset, outperforming the comparison methods.

[0050] Table 1 Performance Comparison Table on Cartoon Animal Parsing Dataset

[0051]

[0052] The present invention proposes a parsing solution for cartoon animals based on spatial learning and structural modeling, which improves the robustness of cartoon animal parsing.

[0053] The present invention proposes a spatial learning branch for cartoon animal parsing. This structure is used to capture variable and irregular shape information, center point information, and edge information, and can effectively capture spatial information.

[0054] The present invention proposes a structural modeling branch for cartoon animal parsing. This structure can learn shape information and associated body structures, and effectively model complex cartoon images.

[0055] The present invention proposes a spatial-structure consistency learning strategy. By correlating spatial information and structural information, this strategy can achieve spatial-structure consistency and improve the network's generalization and recognition ability.

[0056] As described above, the present invention can be preferably implemented.

[0057] All features disclosed in all embodiments in this specification, or all steps in any method or process implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or extended and replaced in any manner.

[0058] As mentioned above, it is only a preferred embodiment of the present invention, and there is no any formal limitation to the present invention. According to the technical essence of the present invention, any simple modification, equivalent replacement, and improvement made to the above embodiments within the spirit and principle of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A cartoon character parsing method based on spatial learning and structural modeling, characterized in that: Cartoon character analysis based on spatial learning and structural modeling.

2. A cartoon character parsing method based on spatial learning and structural modeling according to claim 1, characterized in that: The feature X of the cartoon character is processed by the spatial learning branch and the structural modeling branch respectively, and then the spatial structure consistency learning is performed.

3. A cartoon character parsing method based on spatial learning and structural modeling according to claim 2, characterized in that: In the spatial learning branch processing, deformable convolution is used to obtain the spatial features of cartoon characters, multi-task edge is used to obtain edge-aware spatial features, and the center point prediction strategy is used to obtain center-aware spatial features.

4. A cartoon character parsing method based on spatial learning and structural modeling according to claim 3, characterized in that: In the spatial learning branch processing process, the feature X of the cartoon character is used to obtain irregular spatial features through deformable convolution, and the center point prediction map is obtained through the center point predictor, and the contour prediction map is obtained through the contour predictor; the features processed by the deformable convolution are input into the channel attention, and the feature is refined to obtain feature X1; the features from the center point predictor and the contour predictor are refined using channel attention combined with spatial attention to obtain feature X2; finally, X1 and X2 are spliced ​​and fused to obtain the spatial feature X that can perceive the center point, edge and irregular shape of body parts. sp .

5. A cartoon character parsing method based on space learning and structural modeling according to claim 4, characterized in that: During the structural modeling branch processing, a graph construction module is used to construct a graph structure representation of the cartoon character, and a shape-aware graph neural network is used to model the structural representation of the cartoon character and obtain the relationship between body parts. Among them, the shape-aware graph neural network includes shape convolution and self-attention mechanism. The shape convolution is used to obtain the shape information of the graph nodes, and the self-attention mechanism is used to associate the graph nodes with the shape information.

6. A cartoon character parsing method based on spatial learning and structural modeling according to claim 5, characterized in that: In the process of structural modeling branch processing, the convolutional neural network features are first processed, and the graph structure representation of the cartoon character is constructed through the graph construction module composed of convolution layer convolution and feature graph partitioning operations; then the shape-aware graph neural network is used to model the graph node relationship in the graph structure representation to obtain shape-aware graph convolution features; Then, the paired graph nodes are spliced ​​and fused, and then the shape features are obtained through shape convolution. Then, the shape information-guided graph node information association learning is performed through the self-attention mechanism. Finally, the graph convolution features are mapped to convolution features to obtain the structural features X st .

7. A cartoon character parsing method based on spatial learning and structural modeling according to claim 6, characterized in that: When learning spatial-structural consistency, a cross-attention mechanism is used to associate spatial and structural features, and a consistent feature representation is learned by fusing spatial and structural information.

8. A cartoon character parsing method based on spatial learning and structural modeling according to claim 7, characterized in that: When learning spatial structure consistency, the query-key matching mechanism is used to learn the correspondence between spatial features and structural features, and the cross-attention mechanism is used to adjust the fusion weights between spatial features and structural features.

9. A cartoon character parsing method based on spatial learning and structural modeling according to claim 8, characterized in that: When learning spatial structure consistency, the spatial information X sp As a query, the structural feature X st As a key, the concatenated feature X cat =cat(X sp ,X st ) as the value, the formula for learning the consistency of row space structure is: Q=X sp W q ,K=X st W k ,V=X cat W v X fused =CrossAtten(Q,K,V) Among them, Q represents the query after the learnable weight mapping, K represents the key after the learnable weight mapping, V represents the value after the learnable weight mapping, and W q represents the learnable weight corresponding to Q, W k represents the learnable weight corresponding to K, W v represents the learnable weight corresponding to V, CrossAtten represents the cross attention function, cat represents the concatenation and fusion operation, and X fused Represents spatially and structurally consistent feature representations.

10. A cartoon character parsing method based on space learning and structure modeling according to any one of claims 2 to 9, characterized in that: Cartoon images are generated through convolutional networks and pyramid pooling modules in sequence.