Animation image extraction method and device, storage medium and electronic equipment
By using a bidirectional jump-connected encoder and decoder in the animation image extraction model, combining the multiple rounds of feature sharing and fusion encoding and decoding processing of key feature aggregation and semantic feature extraction modules, the problem of incomplete animation image extraction is solved, and the effect and user experience of animation image extraction are improved.
Patent Information
- Application Number
- CN202510492802.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, when extracting animation images from images, there are problems with incomplete or redundant parts, resulting in poor user experience.
The encoder and decoder that establish bidirectional jump connections in the animation image extraction model are adopted. The fusion codec processing of bidirectional feature sharing is carried out through the key feature aggregation module and the semantic feature extraction module, including the bidirectional jump connection between the multi-level encoding module and the decoding module, and multiple rounds of feature sharing and fusion codec are carried out to improve the animation image extraction effect.
It effectively improves the accuracy and completeness of animation image extraction and improves user experience.
Smart Images

Figure CN120472056A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and specifically to a method, device, storage medium and electronic device for extracting an animated image. Background Art
[0002] Animation image extraction is the work of extracting the images of objects such as characters or animals in animation from images. For example, after extracting animation images from cartoons, the background image of the cartoon can be generated based on the extracted animation images.
[0003] The images containing animated characters (such as video frames in cartoons) often have complex colors, and it is often difficult to extract animated characters from these images. In related technologies, animated characters extracted through simple neural network models usually have various defects such as incompleteness or redundant parts.
[0004] Therefore, there is currently a problem of poor animation image extraction effect and poor user experience. Summary of the Invention
[0005] The embodiment of the present application provides a solution that can effectively improve the animation image extraction effect and enhance the user experience.
[0006] The embodiments of the present application provide the following technical solutions:
[0007] According to one embodiment of the present application, a method for extracting an animated image includes: obtaining an image to be processed; using an encoder and a decoder with bidirectional jump connections established in an animated image extraction model to perform bidirectional feature sharing fusion encoding and decoding processing on the image to be processed to obtain an animated image in the image to be processed; wherein a key feature aggregation module and a semantic feature extraction module are set between the encoder and the decoder, the semantic feature extraction module is used to extract a semantic feature map from the encoding feature map output by the encoder, and the key feature aggregation module is used to perform key feature aggregation on the encoding feature map and the semantic feature map to obtain an aggregated feature map and input the aggregated feature map into the decoder.
[0008] In some embodiments of the present application, the key feature aggregation of the encoding feature map and the semantic feature map to obtain an aggregated feature map includes: cascading the upsampled feature map of the encoding feature map and the semantic feature map to obtain a pre-fused feature map; performing weight calculation processing on the pre-fused feature map to obtain a weight value; convolution processing on the upsampled feature map to obtain a convolution feature map; multiplying the convolution feature map by the weight value to obtain a first weighted feature map; multiplying the semantic feature map by the complement of the weight value to obtain a second weighted feature map; and fusing the first weighted feature map with the second weighted feature map to obtain the aggregated feature map.
[0009] In some embodiments of the present application, the weight calculation processing is performed on the pre-fused feature map to obtain a weight value, including: performing convolution processing on the pre-fused feature map to obtain a convolution feature map; performing global average pooling processing, convolution processing, batch normalization processing and activation processing on the convolution feature map in sequence to obtain the weight value.
[0010] In some embodiments of the present application, the extracting of a semantic feature map from the encoding feature map output by the encoder includes: performing convolution processing on the encoding feature map output by the encoder to obtain a first convolution feature; performing atrous spatial pyramid pooling processing on the first convolution feature to obtain a pyramid pooling feature; performing channel weight adjustment processing on the pyramid pooling feature through an effective channel attention mechanism to obtain a channel weight adjustment feature; performing convolution processing on the channel weight adjustment feature to obtain a second convolution feature; and performing upsampling processing on the second convolution feature to obtain the semantic feature map.
[0011] In some embodiments of the present application, the encoder includes multiple levels of encoding modules, and the decoder includes multiple levels of decoding modules, wherein a bidirectional jump connection is established between the encoding modules and decoding modules of the same level; the bidirectional feature sharing fusion encoding and decoding processing of the image to be processed to obtain the animated image in the image to be processed includes: performing a first round of bidirectional feature sharing fusion encoding and decoding processing on the image to be processed through the multiple levels of encoding modules and the multiple levels of decoding modules to obtain a first decoding output feature map output by the last decoding module; performing a second round of bidirectional feature sharing fusion encoding and decoding processing on the first decoding output feature map through the multiple levels of encoding modules and the multiple levels of decoding modules to obtain the animated image.
[0012] In some embodiments of the present application, the first round of bidirectional feature sharing fusion encoding and decoding processing is performed on the image to be processed by the encoding modules of the multiple levels and the decoding modules of the multiple levels to obtain the first decoding output feature map output by the last decoding module, including: encoding the image to be processed in turn by the encoding modules of the multiple levels to obtain the first shared encoding feature map output by the encoding modules of each level and the first output encoding feature map output by the last encoding module, and sharing the first shared encoding feature map of each level to the decoding modules of the same level respectively; processing the first output encoding feature map by the key feature aggregation module and the semantic feature extraction module to obtain the first aggregated feature map; performing fusion decoding processing on the first aggregated feature map combined with the shared first shared encoding feature map in turn by the decoding modules of the multiple levels to obtain the first shared decoding feature map of each level and the first decoding output feature map output by the last decoding module, and sharing the first shared decoding feature map of each level to the encoding modules of the same level respectively.
[0013] In some embodiments of the present application, the first decoding output feature map is subjected to a second round of bidirectional feature sharing fusion encoding and decoding processing through the encoding modules of the multiple levels and the decoding modules of the multiple levels to obtain an animated image, including: through the encoding modules of the multiple levels, the first decoding output feature map is sequentially subjected to fusion encoding processing in combination with the shared first shared decoding feature map to obtain the second shared encoding feature map output by the encoding modules of each level and the second output encoding feature map output by the last encoding module, and the second shared encoding feature map of each level is respectively shared with the decoding modules of the same level; the second output encoding feature map is processed through the key feature aggregation module and the semantic feature extraction module to obtain a second aggregated feature map; through the decoding modules of the multiple levels, the second aggregated feature map is sequentially subjected to fusion decoding processing in combination with the shared second shared encoding feature map to obtain an animated image in the image to be processed.
[0014] According to one embodiment of the present application, an animation image extraction device includes: an image acquisition unit, used to: acquire an image to be processed; a model calling unit, used to: use an encoder and a decoder with bidirectional jump connections established in an animation image extraction model to perform bidirectional feature sharing fusion encoding and decoding processing on the image to be processed to obtain an animation image in the image to be processed; wherein a key feature aggregation module and a semantic feature extraction module are set between the encoder and the decoder, the semantic feature extraction module is used to extract a semantic feature map from the encoding feature map output by the encoder, and the key feature aggregation module is used to perform key feature aggregation on the encoding feature map and the semantic feature map to obtain an aggregated feature map and input the aggregated feature map into the decoder.
[0015] According to another embodiment of the present application, a storage medium stores a computer program thereon. When the computer program is executed by a processor of a device, the device executes the method described in the embodiment of the present application.
[0016] According to another embodiment of the present application, an electronic device may include: a memory storing a computer program; and a processor reading the computer program stored in the memory to execute the method described in the embodiment of the present application.
[0017] According to another embodiment of the present application, a computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the device to perform the methods provided in various optional implementations described in the embodiments of the present application.
[0018] In an embodiment of the present application, an image to be processed is obtained; an encoder and a decoder with bidirectional jump connections established in an animation image extraction model are used to perform bidirectional feature sharing fusion encoding and decoding processing on the image to be processed to obtain an animation image in the image to be processed; wherein a key feature aggregation module and a semantic feature extraction module are set between the encoder and the decoder, the semantic feature extraction module is used to extract a semantic feature map from the encoding feature map output by the encoder, and the key feature aggregation module is used to perform key feature aggregation on the encoding feature map and the semantic feature map to obtain an aggregated feature map and input the aggregated feature map into the decoder.
[0019] In this manner of the embodiment of the present application, during the process of fusion encoding and decoding processing of the image to be processed by the encoder and decoder through bidirectional feature sharing, the semantic feature extraction module extracts the semantic feature map from the encoding feature map output by the encoder, and the key feature aggregation module performs key feature aggregation on the encoding feature map and the semantic feature map to obtain the aggregated feature map and inputs it into the decoder. Accordingly, the animation image extraction model extracts the animation image from the image to be processed through deep feature fusion analysis, effectively improving the animation image extraction effect and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 A flowchart of an animation image extraction method according to an embodiment of the present application is shown.
[0022] Figure 2 A framework diagram of an animation image extraction module according to an embodiment of the present application is shown.
[0023] Figure 3 A framework diagram of an animation image extraction module according to another embodiment of the present application is shown.
[0024] Figure 4 A framework diagram of an encoding module and a decoding module according to an embodiment of the present application is shown.
[0025] Figure 5 A framework diagram of a semantic feature extraction module according to an embodiment of the present application is shown.
[0026] Figure 6 A framework diagram of a key feature aggregation module according to an embodiment of the present application is shown.
[0027] Figure 7 A block diagram of an animation image extraction device according to an embodiment of the present application is shown.
[0028] Figure 8 A block diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0029] The present disclosure will be further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the examples provided herein are merely for explaining the present disclosure and are not intended to limit the present disclosure. In addition, the examples provided below are partial examples for implementing the present disclosure, rather than providing all examples for implementing the present disclosure. In the absence of conflict, the technical solutions described in the examples of the present disclosure may be implemented in any combination.
[0030] It should be noted that, in the embodiments of the present disclosure, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a method or apparatus comprising a series of elements includes not only the elements explicitly stated, but also other elements not explicitly listed, or also includes elements inherent to the implementation of the method or apparatus. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other related elements (such as steps in the method or units in the apparatus, for example, a unit may be part of a circuit, part of a processor, part of a program or software, etc.) in the method or apparatus comprising the element.
[0031] For example, the animation image extraction method provided by the embodiment of the present disclosure includes a series of steps, but the animation image extraction method provided by the embodiment of the present disclosure is not limited to the recorded steps. Similarly, the animation image extraction device provided by the embodiment of the present disclosure includes a series of units, but the device provided by the embodiment of the present disclosure is not limited to including the units explicitly recorded, and can also include units that need to be set up to obtain relevant information or perform processing based on information.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure pertains. The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure.
[0033] It is understandable that in the specific implementation of this application, relevant data is involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0034] Figure 1 The flowchart of the method for extracting an animated image according to one embodiment of the present application is schematically shown. The execution subject of the method for extracting an animated image can be a device with processing capabilities, such as a TV, computer, mobile phone, tablet, smart watch, and home appliance.
[0035] like Figure 1 As shown, the animation image extraction method may include steps S110 to S120.
[0036] Step S110, obtaining an image to be processed;
[0037] Step S120, using the encoder and decoder with bidirectional skip connection established in the animation image extraction model to perform bidirectional feature sharing fusion encoding and decoding processing on the image to be processed, to obtain the animation image in the image to be processed;
[0038] Among them, a key feature aggregation module and a semantic feature extraction module are set between the encoder and the decoder, the semantic feature extraction module is used to extract a semantic feature map from the encoding feature map output by the encoder, and the key feature aggregation module is used to perform key feature aggregation on the encoding feature map and the semantic feature map to obtain an aggregated feature map and input the aggregated feature map into the decoder.
[0039] See Figure 2 The animation image extraction model may include an encoder 210, a decoder 220, a key feature aggregation module 230 and a semantic feature extraction module 240, wherein a bidirectional jump connection is established between the encoder 210 and the decoder 220 (e.g. Figure 2 The output of the encoder 210 is connected to the input of the key feature aggregation module 230 and the semantic feature extraction module 240 respectively, the output of the semantic feature extraction module 240 is connected to the input of the key feature aggregation module 230, and the input and output of the key feature aggregation module 230 are connected to the input of the decoder 220.
[0040] The image to be processed 200 is input into the animation image extraction model. The animation image extraction model can perform bidirectional feature sharing encoding and decoding processing on the image to be processed based on the encoder 210 and the decoder 220 that have established bidirectional jump connections. During the encoding and decoding process, the encoder 210 encodes the image to be processed to obtain a coding feature map, which will be transmitted to the key feature aggregation module 230 and the semantic feature extraction module 240. The semantic feature extraction module 240 extracts the semantic feature map from the coding feature map output by the encoder and transmits it to the key feature aggregation module 230. The key feature aggregation module 230 performs key feature aggregation on the coding feature map and the semantic feature map to obtain an aggregated feature map, and then inputs the aggregated feature map into the decoder 220. Then, the decoder 220 decodes the aggregated feature map.
[0041] The encoder 210 performs encoding processing by fusing the encoded feature maps shared backward by the decoder 220 via a bidirectional skip connection, and the decoder 220 performs decoding processing by fusing the decoded feature maps shared forward by the encoder 210 via a bidirectional skip connection. Ultimately, the decoder 220 outputs the animated image 260 in the image to be processed 200. The encoding process performed by the encoder is a downsampling process, and the decoding process performed by the decoder is an upsampling process.
[0042] In this manner of the embodiment of the present application, during the process of fusion encoding and decoding processing of the image to be processed by the encoder and decoder through bidirectional feature sharing, the semantic feature extraction module extracts the semantic feature map from the encoding feature map output by the encoder, and the key feature aggregation module performs key feature aggregation on the encoding feature map and the semantic feature map to obtain the aggregated feature map and inputs it into the decoder. Accordingly, the animation image extraction model extracts the animation image from the image to be processed through deep feature fusion analysis, effectively improving the animation image extraction effect and enhancing the user experience.
[0043] Described below Figure 1 When performing animation image extraction under the embodiment, further optional specific embodiments are provided for each step performed.
[0044] In one embodiment, the encoder includes multiple levels of encoding modules, and the decoder includes multiple levels of decoding modules, wherein bidirectional jump connections are established between encoding modules and decoding modules at the same level; the bidirectional feature sharing fusion encoding and decoding processing of the image to be processed to obtain an animated image in the image to be processed may specifically include: performing a first round of bidirectional feature sharing fusion encoding and decoding processing on the image to be processed through the multiple levels of encoding modules and the multiple levels of decoding modules to obtain a first decoding output feature map output by the last decoding module; performing a second round of bidirectional feature sharing fusion encoding and decoding processing on the first decoding output feature map through the multiple levels of encoding modules and the multiple levels of decoding modules to obtain an animated image.
[0045] In this embodiment, in the animation image extraction model, the encoder includes multiple levels of encoding modules, the decoder includes multiple levels of decoding modules, and bidirectional jump connections are established between the encoding modules and decoding modules at the same level. Figure 3The encoder 210 includes a first-level encoding module 211, a second-level encoding module 212, a third-level encoding module 213, and a fourth-level encoding module 214. The decoder 220 includes a first-level decoding module 221, a second-level decoding module 222, a third-level decoding module 223, and a fourth-level decoding module 224. Bidirectional skip connections 251, 252, 253, and 254 are established between the encoding modules and decoding modules at the same level, respectively. Thus, the animated image extraction model is a U-Net neural network model that introduces a key feature aggregation module 230 and establishes bidirectional skip connections.
[0046] In this manner, a first round of bidirectional feature sharing fusion encoding and decoding processing is performed on the image to be processed through the encoding modules and decoding modules of the multiple levels, thereby obtaining a first decoded output feature map output by the last decoding module. A second round of bidirectional feature sharing fusion encoding and decoding processing is performed on the first decoded output feature map through the encoding modules and decoding modules of the multiple levels, and an animated image can be predicted based on the feature map output by the last decoding module. Through two rounds of fusion encoding and decoding processing, an animated image with good effects can be reliably extracted.
[0047] Furthermore, in one embodiment, performing a first round of bidirectional feature sharing fusion encoding and decoding processing on the image to be processed through the encoding modules of the multiple levels and the decoding modules of the multiple levels to obtain a first decoding output feature map output by the last decoding module may specifically include:
[0048] The image to be processed is encoded in sequence by the encoding modules of the multiple levels, and a first shared encoding feature map output by the encoding modules of each level and a first output encoding feature map output by the last encoding module are obtained, and the first shared encoding feature map of each level is shared with the decoding module of the same level respectively;
[0049] Processing the first output encoding feature map by the key feature aggregation module and the semantic feature extraction module to obtain a first aggregated feature map;
[0050] Through the decoding modules of the multiple levels, the first aggregated feature map is combined with the shared first shared coding feature map for fusion decoding processing in turn to obtain the first shared decoding feature map of each level and the first decoding output feature map output by the last decoding module, and the first shared decoding feature map of each level is shared with the coding module of the same level respectively.
[0051] For example, see Figure 3, through the encoding modules of the 1st to 4th levels, the image to be processed is encoded in sequence, and the first shared encoding feature maps output by the encoding modules of the 1st to 4th levels can be obtained. Further, the first shared encoding feature map output by the encoding module 211 of the 1st level is shared with the decoding module 221 of the 1st level through the bidirectional jump connection 251. Similarly, the first shared encoding feature maps output by the encoding modules of other levels are shared with the decoding modules of the same level. Among them, the encoding modules of each level further downsample the first shared encoding feature map output by the encoding modules to obtain the first output encoding feature map. The first output encoding feature map output by the encoding modules from the 1st to the 4th level continuously becomes smaller, and the first output encoding feature map output by the last encoding module (i.e., the encoding module of the 4th level) is the smallest.
[0052] Then, the first output encoding feature map output by the last encoding module 214 is processed by the key feature aggregation module 230 and the semantic feature extraction module 240 to obtain the first aggregated feature map of the first round. Specifically, the semantic feature extraction module 240 extracts the first semantic feature map from the first output encoding feature map output by the last encoding module 214, and the key feature aggregation module performs key feature aggregation on the first output encoding feature map and the first semantic feature map to obtain the first aggregated feature map of the first round.
[0053] Then, through the decoding modules from the 4th level to the 1st level, the first aggregated feature map of the first round is combined with the first shared coding feature map shared by each decoding module to perform fusion decoding processing in turn, and the first shared decoding feature map output by the decoding modules of each level is obtained. The decoding modules of each level share the first shared decoding feature map with the coding modules of the same level. For example, the first shared decoding feature map output by the decoding module 224 of the 4th level is reversely shared with the decoding module 224 of the 4th level through the bidirectional jump connection 254. Similarly, the first shared decoding feature maps output by the decoding modules of other levels are shared with the coding modules of the same level. Among them, the decoding modules of each level further upsample the first shared decoding feature map output by the upsampling layer to obtain the first output decoding feature map. The first output decoding feature map output by the decoding modules from the 4th level to the 1st level continues to increase, and the first output decoding feature map output by the last decoding module (i.e., the decoding module of the 1st level) is the largest.
[0054] Specifically, in one embodiment, see Figure 4Each encoding module may include a fusion layer (concat) 310, a convolution layer (conv) 320, a convolution layer (conv) 330, and a downsampling layer (down) 340. Each decoding module may include a fusion layer (concat) 410, a convolution layer (conv) 420, a convolution layer (conv) 430, and an upsampling layer (up) 440. The fusion layer (concat) 310 may specifically fuse two feature maps in a cascade manner.
[0055] Each encoding module can first combine the input feature data (Ain) with the shared first shared decoding feature map (f dec ) are fused into fused coding features, and then two layers of convolution processing are performed through the convolution layer (conv) 320 and the convolution layer (conv) 330 to obtain the first shared coding feature map (f enc The first shared encoding feature map (f enc ) is also forward-shared to the decoding module at the same level through a bidirectional jump connection. Each encoding module further downsamples its output first shared encoding feature map through a downsampling layer (down) to obtain the first output encoding feature map (Aout) finally output by each encoding module.
[0056] The input feature data (Ain) in the fusion layer of each coding module is the first output coding feature map (Aout) finally output by the previous coding module, wherein, in the first coding processing step (encoding the image to be processed in sequence through the coding modules of the 1st to 4th levels): the input feature data (Ain) in the fusion layer of the coding module 211 of the first level is the image to be processed, the input feature data (Ain) in the fusion layer of the coding module 212 of the second level is the first output coding feature map (Aout) finally output by the coding module 211, and so on for the input feature data (Ain) of other coding modules; and in this process, since the decoding process has not yet been performed, each coding module has not yet received the first shared decoding feature map (f dec ), that is, in this process, the fusion layer (concat) in the encoding module fuses the input feature data (Ain) with the empty first output encoding feature map.
[0057] Furthermore, the semantic feature extraction module extracts a first semantic feature map from the first output encoding feature map output by the last encoding module, and the key feature aggregation module performs key feature aggregation on the first output encoding feature map and the first semantic feature map to obtain a first aggregated feature map of the first round.
[0058] Each decoding module can first combine the input feature data (Bin) with the shared first shared encoding feature map (f enc ) is fused into a fused decoding feature, and then two layers of convolution processing are performed through the convolution layer (conv) 420 and the convolution layer (conv) 430 to obtain the first shared decoding feature map (f dec The first shared decoding feature map (f dec ) is also forward-shared to the encoding module at the same level through a bidirectional skip connection. Each decoding module further upsamples its output first shared decoding feature map through the upsampling layer (up) to obtain the first output decoding feature map (Bout) ultimately output by each decoding module.
[0059] The input feature data (Bin) in the fusion layer of each decoding module is the first output decoding feature map (Bout) finally output by the previous decoding module, wherein, in the first fusion decoding processing step (through the decoding modules from the 4th level to the 1st level, the first aggregated feature map of the first round is combined with the first shared coding feature map shared by each decoding module in turn for fusion decoding processing): the input feature data (Bin) in the fusion layer of the 4th level decoding module 224 is the first aggregated feature map, the input feature data (Bin) in the fusion layer of the 3rd level decoding module 223 is the first output decoding feature map (Bout) finally output by the decoding module 224, and so on for the input feature data (Bin) of other decoding modules.
[0060] In one embodiment, performing a second round of bidirectional feature sharing fusion encoding and decoding processing on the first decoded output feature map through the multiple levels of encoding modules and the multiple levels of decoding modules to obtain an animated image may specifically include:
[0061] Through the encoding modules of the multiple levels, the first decoding output feature map is combined with the shared first shared decoding feature map in sequence to perform fusion encoding processing, thereby obtaining the second shared encoding feature map output by the encoding modules of each level and the second output encoding feature map output by the last encoding module, and sharing the second shared encoding feature map of each level with the decoding modules of the same level respectively;
[0062] Processing the second output encoding feature map by the key feature aggregation module and the semantic feature extraction module to obtain a second aggregated feature map;
[0063] Through the decoding modules of the multiple levels, the second aggregated feature map is sequentially combined with the shared second shared coding feature map to perform fusion decoding processing to obtain the animated image in the image to be processed.
[0064] For example, see Figure 3 , through the encoding modules of the 1st to 4th levels, the first decoding output feature map output by the last decoding module is combined with the shared first shared decoding feature map for fusion coding processing in turn to obtain the second shared coding feature map output by the encoding modules of each level. Furthermore, the second shared coding feature map output by the encoding module 211 of the 1st level is shared with the decoding module 221 of the 1st level through the bidirectional jump connection 251. Similarly, the second shared coding feature maps output by the encoding modules of other levels are shared with the decoding modules of the same level. Among them, the encoding modules of each level further downsample the second shared coding feature map output by the encoding modules to obtain the second output coding feature map. The second output coding feature map output by the encoding modules from the 1st to the 4th level continuously becomes smaller, and the second output coding feature map output by the last encoding module (i.e., the encoding module of the 4th level) is the smallest.
[0065] Then, the second output encoding feature map output by the last encoding module 214 is processed by the key feature aggregation module 230 and the semantic feature extraction module 240 to obtain a second aggregated feature map of the second round. Specifically, the semantic feature extraction module 240 extracts a second semantic feature map from the second output encoding feature map output by the last encoding module 214, and the key feature aggregation module performs key feature aggregation on the second output encoding feature map output by the last encoding module 214 and the second semantic feature map to obtain a second aggregated feature map of the second round.
[0066] Then, through the decoding modules from the 4th level to the 1st level, the second aggregated feature map is combined with the second shared coding feature map shared by each decoding module in turn for fusion decoding processing, and the second shared decoding feature map output by the decoding modules of each level is obtained. The decoding modules of each level further upsample the second shared decoding feature map output by the upsampling layer to obtain the second output decoding feature map. The second output decoding feature map output by the decoding modules from the 4th level to the 1st level continuously increases. According to the second output decoding feature map output by the last decoding module (i.e., the decoding module of the 1st level), the animated image in the image to be processed can be predicted.
[0067] Specifically, in one embodiment, see Figure 4 Each encoding module may include a fusion layer (concat) 310, a convolution layer (conv) 320, a convolution layer (conv) 330, and a downsampling layer (down) 340. Each decoding module may include a fusion layer (concat) 410, a convolution layer (conv) 420, a convolution layer (conv) 430, and an upsampling layer (up) 440. The fusion layer (concat) 310 may specifically fuse two feature maps in a cascade manner.
[0068] Each encoding module can first combine the input feature data (Ain) with the shared first shared decoding feature map (f dec ) are fused into fused coding features, and then two layers of convolution processing are performed through the convolution layer (conv) 320 and the convolution layer (conv) 330 to obtain the second shared coding feature map (f enc The second shared encoding feature map (f enc ) is also forward-shared to the decoding module at the same level through a bidirectional jump connection. Each encoding module further downsamples its output second shared encoding feature map through the downsampling layer (down) to obtain the second output encoding feature map (Aout) finally output by each encoding module.
[0069] The input feature data (Ain) in the fusion layer of each coding module is the second output coding feature map (Aout) finally output by the previous coding module, wherein, in the first fusion coding processing step (through the coding modules of the 1st to 4th levels, the first decoding output feature map output by the last decoding module is combined with the shared first shared decoding feature map for fusion coding processing): the input feature data (Ain) in the fusion layer of the coding module 211 of the 1st level is the "first decoding output feature map output by the last decoding module", the input feature data (Ain) in the fusion layer of the coding module 212 of the second level is the second output coding feature map (Aout) finally output by the coding module 211, and so on for the input feature data (Ain) of other coding modules.
[0070] Furthermore, the semantic feature extraction module extracts a second semantic feature map from the second output encoding feature map output by the last encoding module, and the key feature aggregation module performs key feature aggregation on the second output encoding feature map output by the last encoding module and the second semantic feature map to obtain a second aggregated feature map of the second round.
[0071] Each decoding module can first combine the input feature data (Bin) with the shared second shared encoding feature map (f enc ) is fused into a fused decoding feature, and then two layers of convolution processing are performed through the convolution layer (conv) 420 and the convolution layer (conv) 430 to obtain the second shared decoding feature map (f dec ). Each decoding module further performs upsampling processing on the second shared decoding feature map output by it through the upsampling layer (up), thereby obtaining the second output decoding feature map (Bout) finally output by each decoding module.
[0072] The input feature data (Bin) in the fusion layer of each decoding module is the second output decoding feature map (Bout) finally output by the previous decoding module, wherein, in the second fusion decoding processing step (through the decoding modules from the 4th level to the 1st level, the second aggregated feature map is sequentially combined with the second shared coding feature map shared by each decoding module for fusion decoding processing): the input feature data (Bin) in the fusion layer of the 4th level decoding module 224 is the second aggregated feature map, the input feature data (Bin) in the fusion layer of the 3rd level decoding module 223 is the second output decoding feature map (Bout) finally output by the decoding module 224, and so on for the input feature data (Bin) of other decoding modules.
[0073] In one embodiment, extracting a semantic feature map from the encoding feature map output by the encoder may include: performing convolution processing on the encoding feature map output by the encoder to obtain a first convolution feature; performing atrous spatial pyramid pooling processing on the first convolution feature to obtain a pyramid pooling feature; performing channel weight adjustment processing on the pyramid pooling feature through an effective channel attention mechanism to obtain a channel weight adjustment feature; performing convolution processing on the channel weight adjustment feature to obtain a second convolution feature; and performing upsampling processing on the second convolution feature to obtain the semantic feature map.
[0074] The semantic feature extraction module is used to extract the semantic feature map from the encoding feature map output by the encoder, see Figure 5 In this embodiment, the semantic feature extraction module may include a convolution layer (conv) 510, an atrous spatial pyramid pooling layer (ASPP) 520, an effective channel attention layer (ECA) 530, a convolution layer (conv) 540 and an upsampling layer (up) 550. Thus, the semantic feature extraction module in this embodiment is a middle block.
[0075] The encoding feature map output by the encoder is convolved by a convolution layer (conv) 510 to obtain a first convolution feature; the first convolution feature is subjected to atrous spatial pyramid pooling by an atrous spatial pyramid pooling layer (ASPP) 520 to obtain a pyramid pooling feature; the pyramid pooling feature is subjected to channel weight adjustment by an effective channel attention mechanism by an effective channel attention layer (ECA) 530 to obtain a channel weight adjustment feature; the channel weight adjustment feature is convolved by a convolution layer (conv) 540 to obtain a second convolution feature; the second convolution feature is upsampled by an upsampling layer (up) 550 to obtain a semantic feature map. The semantic feature extraction module adopts this structure to achieve deep semantic feature extraction processing. The extracted semantic feature map can be used in the embodiments of the present application to further ensure the animation extraction effect.
[0076] In one embodiment, performing key feature aggregation on the encoding feature map and the semantic feature map to obtain an aggregated feature map may include:
[0077] The upsampled feature map of the encoding feature map and the semantic feature map are fused to obtain a pre-fused feature map; the pre-fused feature map is weighted and processed to obtain a weight value; the upsampled feature map is convolved to obtain a convolution feature map; the convolution feature map is multiplied by the weight value to obtain a first weighted feature map; the semantic feature map is multiplied by the complement of the weight value to obtain a second weighted feature map; the first weighted feature map is fused with the second weighted feature map to obtain the aggregated feature map.
[0078] See Figure 6 The key feature aggregation module may include a pre-fusion layer 610, a weight calculation layer 620, a first multiplication layer (Multiply) 630, a second multiplication layer (Multiply) 640, a final fusion layer 650, an upsampling layer 660, and a third convolution layer 670.
[0079] The upsampled feature map of the coding feature map F1 (i.e., the upsampled feature map obtained by upsampling the coding feature map through the upsampling layer 660) and the semantic feature map F2 can be fused through the pre-fusion layer 610 (the fusion method can be a cascade of two feature maps or other optional methods) to obtain a pre-fused feature map; then, the pre-fused feature map can be weighted by the weight calculation layer 620 to obtain a weight value; the upsampled feature map can be convolved by the third convolution layer 670 to obtain a convolution feature map; the convolution feature map is multiplied by the weight value (α) through the first multiplication layer (Multiply) 630 to obtain a first weighted feature map; the semantic feature map is multiplied by the complement of the weight value (1-α) through the second multiplication layer (Multiply) 640 to obtain a second weighted feature map; finally, the final fusion layer 650 obtains an aggregated feature map by fusing the first weighted feature map with the second weighted feature map (the fusion method can be a cascade of two feature maps or other optional methods).
[0080] In the manner of this embodiment, the key feature aggregation module adopts this structure to perform key feature aggregation on feature maps of two different dimensions, the encoding feature map and the semantic feature map, to obtain an aggregated feature map with better content representation effect. The aggregated feature map used in the embodiment of this application can further improve the animation image extraction effect.
[0081] Furthermore, in one embodiment, the weight calculation processing of the pre-fused feature map to obtain the weight value may include: performing convolution processing on the pre-fused feature map to obtain a convolution feature map; and performing global average pooling processing, convolution processing, batch normalization processing and activation processing on the convolution feature map in sequence to obtain the weight value.
[0082] See Figure 6 The weight calculation layer 620 may include a fourth convolution layer 621, a global average pooling (Avgpool) layer 622, a fifth convolution layer 623, a batch normalization (BatchNorm) layer 624, and an activation (sigmoid) layer 625. Through the fourth convolution layer 621, the global average pooling (Avgpool) layer 622, the fifth convolution layer 623, the batch normalization (BatchNorm) layer 624, and the activation (sigmoid) layer 625, the pre-fused feature map can be sequentially subjected to convolution processing, global average pooling processing, convolution processing, batch normalization processing, and activation processing to obtain a weight value. The weight value obtained in this way is used in the aforementioned embodiment to obtain an aggregated feature map, which can further improve the animation image extraction effect.
[0083] To facilitate better implementation of the animated image extraction method provided in the embodiments of this application, the embodiments of this application also provide an animated image extraction device based on the aforementioned animated image extraction method. The meanings of the terms herein are the same as those in the aforementioned animated image extraction method, and specific implementation details can be found in the description of the method embodiment. Figure 7 A block diagram of an animation image extraction device according to an embodiment of the present application is shown.
[0084] like Figure 7 As shown, the animation image extraction device 700 may include: an image acquisition unit 710 can be used to: acquire the image to be processed; a model calling unit 720 can be used to: use an encoder and a decoder with bidirectional jump connections established in the animation image extraction model to perform bidirectional feature sharing fusion encoding and decoding processing on the image to be processed to obtain the animation image in the image to be processed; wherein a key feature aggregation module and a semantic feature extraction module are set between the encoder and the decoder, the semantic feature extraction module is used to extract a semantic feature map from the encoding feature map output by the encoder, and the key feature aggregation module is used to perform key feature aggregation on the encoding feature map and the semantic feature map to obtain an aggregated feature map and input the aggregated feature map into the decoder.
[0085] In some embodiments of the present application, the key feature aggregation module can be used to: cascade fuse the upsampled feature map of the encoded feature map and the semantic feature map to obtain a pre-fused feature map; perform weight calculation processing on the pre-fused feature map to obtain a weight value; perform convolution processing on the upsampled feature map to obtain a convolution feature map; multiply the convolution feature map by the weight value to obtain a first weighted feature map; multiply the semantic feature map by the complement of the weight value to obtain a second weighted feature map; and fuse the first weighted feature map with the second weighted feature map to obtain the aggregated feature map.
[0086] In some embodiments of the present application, the key feature aggregation module can be used to: perform convolution processing on the pre-fused feature map to obtain a convolution feature map; perform global average pooling processing, convolution processing, batch normalization processing and activation processing on the convolution feature map in sequence to obtain the weight value.
[0087] In some embodiments of the present application, the semantic feature extraction module can be used to: perform convolution processing on the encoded feature map output by the encoder to obtain a first convolution feature; perform atrous spatial pyramid pooling processing on the first convolution feature to obtain a pyramid pooling feature; perform channel weight adjustment processing on the pyramid pooling feature through an effective channel attention mechanism to obtain a channel weight adjustment feature; perform convolution processing on the channel weight adjustment feature to obtain a second convolution feature; and perform upsampling processing on the second convolution feature to obtain the semantic feature map.
[0088] In some embodiments of the present application, the encoder includes multiple levels of encoding modules, and the decoder includes multiple levels of decoding modules, wherein bidirectional jump connections are established between encoding modules and decoding modules at the same level; the multiple levels of encoding modules and the multiple levels of decoding modules can be used to: perform a first round of bidirectional feature sharing fusion encoding and decoding processing on the image to be processed to obtain a first decoding output feature map output by the last decoding module; and perform a second round of bidirectional feature sharing fusion encoding and decoding processing on the first decoding output feature map to obtain an animated image.
[0089] In some embodiments of the present application, when the first round of bidirectional feature sharing fusion encoding and decoding processing is performed on the image to be processed by the encoding modules of the multiple levels and the decoding modules of the multiple levels to obtain the first decoding output feature map output by the last decoding module, the encoding modules of the multiple levels can be used to: encode the image to be processed in sequence to obtain the first shared encoding feature map output by the encoding modules of each level and the first output encoding feature map output by the last encoding module, and share the first shared encoding feature map of each level to the decoding modules of the same level respectively; the key feature aggregation module and the semantic feature extraction module can be used to: process the first output encoding feature map to obtain the first aggregated feature map; the decoding modules of the multiple levels can be used to: perform fusion decoding processing on the first aggregated feature map combined with the shared first shared encoding feature map in sequence to obtain the first shared decoding feature map of each level and the first decoding output feature map output by the last decoding module, and share the first shared decoding feature map of each level to the encoding modules of the same level respectively.
[0090] In some embodiments of the present application, when the first decoding output feature map is subjected to a second round of bidirectional feature sharing fusion encoding and decoding processing through the encoding modules of the multiple levels and the decoding modules of the multiple levels to obtain an animated image, the encoding modules of the multiple levels can be used to: sequentially perform fusion encoding processing on the first decoding output feature map in combination with the shared first shared decoding feature map to obtain the second shared encoding feature map output by the encoding modules of each level and the second output encoding feature map output by the last encoding module, and respectively share the second shared encoding feature map of each level to the decoding modules of the same level; the key feature aggregation module and the semantic feature extraction module can be used to: process the second output encoding feature map to obtain a second aggregated feature map; the decoding modules of the multiple levels can be used to: sequentially perform fusion decoding processing on the second aggregated feature map in combination with the shared second shared encoding feature map to obtain an animated image in the image to be processed.
[0091] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0092] In addition, the present invention also provides an electronic device, such as Figure 8 As shown, Figure 8A block diagram of an electronic device according to an embodiment of the present application is shown, specifically:
[0093] The electronic device may include one or more processors 801, one or more computer-readable storage media memories 802, a power supply 803, an input unit 804, and other components. It will be understood by those skilled in the art that Figure 8 The electronic device structure shown in the figure does not constitute a limitation to the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0094] in:
[0095] Processor 801 is the control center of the electronic device, connecting the various components of the entire computer device using various interfaces and lines. By running or executing software programs and / or modules stored in memory 802 and accessing data stored in memory 802, it performs various computer device functions and processes data, thereby monitoring the electronic device as a whole. Optionally, processor 801 may include one or more processing cores; preferably, processor 801 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interfaces, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 801.
[0096] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.
[0097] The electronic device also includes a power supply 803 for supplying power to various components. Preferably, the power supply 803 can be logically connected to the processor 801 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 803 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0098] The electronic device may further include an input unit 804, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0099] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically, in this embodiment, the processor 801 in the electronic device loads the executable files corresponding to one or more computer program processes into the memory 802 according to the instructions, and the processor 801 runs the computer program stored in the memory 802, thereby realizing the various functions in the aforementioned embodiments of the present application.
[0100] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by a computer program, or by controlling related hardware through a computer program. The computer program may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0101] To this end, an embodiment of the present application further provides a storage medium storing a computer program, which can be loaded by a processor to execute the steps of any method provided in the embodiment of the present application.
[0102] The storage medium may be a computer-readable storage medium, and the storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0103] Since the computer program stored in the storage medium can execute the steps of any method provided in the embodiments of the present application, the beneficial effects that can be achieved by the method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0104] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.
[0105] It should be understood that the present application is not limited to the embodiments that have been described above and shown in the accompanying drawings, but various modifications and changes may be made without departing from the scope thereof.
Claims
1. A method for extracting an animated image, characterized in that: include: Get the image to be processed; Using an encoder and a decoder with bidirectional skip connections established in an animation image extraction model, the image to be processed is subjected to a fusion encoding and decoding process with bidirectional feature sharing to obtain an animation image in the image to be processed; Among them, a key feature aggregation module and a semantic feature extraction module are set between the encoder and the decoder, the semantic feature extraction module is used to extract a semantic feature map from the encoding feature map output by the encoder, and the key feature aggregation module is used to perform key feature aggregation on the encoding feature map and the semantic feature map to obtain an aggregated feature map and input the aggregated feature map into the decoder.
2. The method according to claim 1, characterized in that The step of performing key feature aggregation on the encoding feature map and the semantic feature map to obtain an aggregated feature map includes: Cascadingly fusing the upsampled feature map of the encoding feature map and the semantic feature map to obtain a pre-fused feature map; Performing weight calculation processing on the pre-fused feature map to obtain a weight value; Performing convolution processing on the upsampled feature map to obtain a convolution feature map; Multiplying the convolution feature map by the weight value to obtain a first weighted feature map; Multiplying the semantic feature map by the complement of the weight value to obtain a second weighted feature map; The first weighted feature map is fused with the second weighted feature map to obtain the aggregated feature map.
3. The method according to claim 2, characterized in that The performing weight calculation processing on the pre-fusion feature map to obtain a weight value includes: Performing convolution processing on the pre-fused feature map to obtain a convolution feature map; The convolution feature map is sequentially subjected to global average pooling processing, convolution processing, batch normalization processing and activation processing to obtain the weight value.
4. The method according to claim 1, wherein The step of extracting a semantic feature map from the encoding feature map output by the encoder comprises: Performing convolution processing on the encoding feature map output by the encoder to obtain a first convolution feature; Performing a dilated spatial pyramid pooling process on the first convolutional features to obtain a pyramid pooling feature; Performing channel weight adjustment processing on the pyramid pooling feature through an effective channel attention mechanism to obtain a channel weight adjustment feature; Convolution is performed on the channel weight adjustment feature to obtain a second convolution feature; The second convolutional feature is upsampled to obtain the semantic feature map.
5. The method according to claim 1, wherein The encoder includes multiple levels of encoding modules, and the decoder includes multiple levels of decoding modules, wherein bidirectional jump connections are established between encoding modules and decoding modules at the same level; The step of performing bidirectional feature sharing fusion encoding and decoding processing on the image to be processed to obtain an animated image in the image to be processed includes: Performing a first round of bidirectional feature sharing fusion encoding and decoding processing on the image to be processed through the encoding modules of the multiple levels and the decoding modules of the multiple levels, to obtain a first decoding output feature map output by the last decoding module; Through the encoding modules of the multiple levels and the decoding modules of the multiple levels, the first decoding output feature map is subjected to a second round of bidirectional feature sharing fusion encoding and decoding processing to obtain an animated image.
6. The method according to claim 5, characterized in that The step of performing a first round of bidirectional feature sharing fusion encoding and decoding processing on the image to be processed by the encoding modules of the multiple levels and the decoding modules of the multiple levels to obtain a first decoding output feature map output by the last decoding module includes: The image to be processed is encoded in sequence by the encoding modules of the multiple levels, and a first shared encoding feature map output by the encoding modules of each level and a first output encoding feature map output by the last encoding module are obtained, and the first shared encoding feature map of each level is shared with the decoding module of the same level respectively; Processing the first output encoding feature map by the key feature aggregation module and the semantic feature extraction module to obtain a first aggregated feature map; Through the decoding modules of the multiple levels, the first aggregated feature map is combined with the shared first shared coding feature map for fusion decoding processing in turn to obtain the first shared decoding feature map of each level and the first decoding output feature map output by the last decoding module, and the first shared decoding feature map of each level is shared with the coding module of the same level respectively.
7. The method according to claim 6, characterized in that The step of performing a second round of bidirectional feature sharing fusion encoding and decoding processing on the first decoded output feature map through the encoding modules of the multiple levels and the decoding modules of the multiple levels to obtain an animated image includes: Through the encoding modules of the multiple levels, the first decoding output feature map is combined with the shared first shared decoding feature map in sequence to perform fusion encoding processing, thereby obtaining the second shared encoding feature map output by the encoding modules of each level and the second output encoding feature map output by the last encoding module, and sharing the second shared encoding feature map of each level with the decoding modules of the same level respectively; Processing the second output encoding feature map by the key feature aggregation module and the semantic feature extraction module to obtain a second aggregated feature map; Through the decoding modules of the multiple levels, the second aggregated feature map is sequentially combined with the shared second shared coding feature map to perform fusion decoding processing to obtain the animated image in the image to be processed.
8. An animation image extraction device, characterized in that: include: An image acquisition unit, configured to: acquire an image to be processed; The model calling unit is used to: use the encoder and decoder with bidirectional jump connection established in the animation image extraction model to perform bidirectional feature sharing fusion encoding and decoding processing on the image to be processed to obtain the animation image in the image to be processed; wherein, a key feature aggregation module and a semantic feature extraction module are set between the encoder and the decoder, the semantic feature extraction module is used to extract a semantic feature map from the encoding feature map output by the encoder, and the key feature aggregation module is used to perform key feature aggregation on the encoding feature map and the semantic feature map to obtain an aggregated feature map and input the aggregated feature map into the decoder.
9. A storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of a device, the device is caused to perform the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: a memory storing a computer program; A processor reads a computer program stored in a memory to execute the method according to any one of claims 1 to 7.