An automatic image cutting method based on adaptive feature extraction and semantic guidance

By building an automatic picture cutting method of multi-branch collaborative architecture, the problem of poor results in the existing technology in complex backgrounds and details is solved, and adaptive training and fine picture cutting for multiple types of prospects are achieved, which significantly improves the picture cutting effect and computing efficiency.

CN119832253BActive Publication Date: 2025-05-13MO NI XUEDI (JIANGXI) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510309879.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-05-13
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

Existing automatic image cutting methods have poor results when dealing with complex backgrounds and details, especially under conditions without auxiliary input, making it difficult to achieve effective work on various types of images.

Method used

An automatic cutout method for adaptive feature extraction and semantic guidance is proposed. By building a multi-branch collaborative architecture, including adaptive multi-source feature capture branches, local enhanced self-attention branches, local understanding branches, cross-scale hollow convolution pooling branches and global semantic guidance branches, it realizes adaptive training and fine cutouts for multi-type prospects.

Benefits of technology

Without auxiliary input, the capture capability of irregular edges is significantly improved, computing redundancy in complex scenarios is reduced, pixel-level transparency prediction is achieved, and industrial-grade real-time needs are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832253B_ABST
    Figure CN119832253B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic cutout method for adaptive feature extraction and semantic guidance, comprising the following steps: constructing an automatic cutout model, the automatic cutout model consisting of an adaptive multi-source feature capture branch, a local enhanced self-attention branch, a local understanding branch, a cross-scale hole convolution pooling branch, a global semantic guidance branch and an aggregation module; inputting an object image into each branch in the automatic cutout model for processing, and outputting a transparency mask; the present invention constructs a new multi-branch collaborative architecture, proposes an adaptive multi-source feature capture branch consisting of a deformable convolution, a residual structure and a multi-level pooling, and significantly improves the ability to capture irregular edges by dynamically adjusting the convolution sampling area and the feature capture method, while reducing the computational redundancy in complex scenes, and combines the detail enhancement module of the local understanding branch with the recursive cross attention of the global semantic guidance branch to achieve pixel-level transparency prediction without the need for auxiliary input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of image processing, and in particular to an automatic image cutting method based on adaptive feature extraction and semantic guidance. Background Art

[0002] Image matting aims to extract the foreground from the image and separate it from the background. This task is usually regarded as an under-constrained problem. In color images, matting involves 3 known numbers and 7 unknown numbers, so most traditional methods require additional auxiliary inputs (such as trimap, graffiti or background images) to guide the matting process. Although this auxiliary information can simplify the solution of the problem, it also increases the complexity of user operations and limits the wide applicability of the method in non-interactive applications. With the development of deep learning technology, especially convolutional neural networks (CNNs) and Transformers, image matting methods based on these technologies have made significant progress in image segmentation tasks. Although most existing methods rely on additional prior information to improve the matting performance to a certain extent, fully automatic image matting (AIM) still faces challenges, especially when dealing with complex backgrounds and details, and the problem of automation without auxiliary input has not been fully solved.

[0003] Automatic image matting focuses on automatically extracting soft foregrounds from arbitrary natural images without any auxiliary input. However, most existing automatic image matting methods focus on scenes with salient opaque foregrounds, such as the matting of objects such as humans or animals. These methods perform well when dealing with simple scenes, but perform poorly when faced with complex or transparent and detailed foregrounds (such as glass, plastic bags, etc.) or non-salient foregrounds (such as smoke, grids, raindrops, etc.). The fundamental reason is that existing methods usually rely on explicit semantic features or implicit feature representations learned for specific types of images, which makes them lack versatility and robustness when dealing with different types of images. Therefore, how to extend existing automatic image matting methods so that they can work effectively in a variety of different types of images, especially in complex backgrounds, remains an important challenge in current research. Summary of the invention

[0004] In view of the deficiencies in the prior art, the present invention provides an automatic image cutting method based on adaptive feature extraction and semantic guidance, which aims to solve the problems in the background technology.

[0005] To achieve the above object, the present invention provides the following technical solution: an automatic image cutting method based on adaptive feature extraction and semantic guidance, comprising the following steps:

[0006] Step S1: construct a cutout dataset, which includes a number of object images;

[0007] Step S2: construct an automatic cutout model, which consists of an adaptive multi-source feature capture branch, a local enhanced self-attention branch, a local understanding branch, a cross-scale dilated convolution pooling branch, a global semantic guidance branch and an aggregation module;

[0008] The local understanding branch consists of a detail enhancement module and multiple local understanding decoding blocks in series; the global semantic guidance branch consists of a recursive cross attention and multiple semantic guidance modules in series, and a global guidance decoding block is set before each semantic guidance module;

[0009] Step S3: input the object image into the adaptive multi-source feature capture branch to extract a multi-level feature map;

[0010] Step S4: input the extracted multi-level feature map into the local enhanced self-attention branch to obtain a local enhanced feature map;

[0011] Step S5: the detail enhancement module DEM receives the output of the last layer of the adaptive multi-source feature capture branch and processes it together with the local enhancement feature map, and then inputs it into the subsequent local understanding decoding block for layer-by-layer processing to obtain the local understanding decoding feature map;

[0012] Step S6: input the extracted multi-level feature map into the cross-scale dilated convolution pooling branch to obtain a cross-scale dilated convolution pooling branch feature map;

[0013] Step S7: Recursive cross attention receives the output of the last layer of the adaptive multi-source feature capture branch and processes it with the cross-scale dilated convolution pooling branch features. Figure 1 The subsequent global guided decoding block and semantic guided module process the same input layer by layer to obtain the global semantic guided decoding feature map; wherein the semantic guided module receives the output of the global guided decoding block and recursive cross attention of the previous layer;

[0014] Step S8: Aggregate the local understanding decoding feature map and the global semantic guided decoding feature map through the aggregation module to obtain a transparency mask, which is the final output.

[0015] Furthermore, the adaptive multi-source feature capture branch consists of five encoding blocks, including the initial encoding block , the first coding block , the second coding block , the third coding block and the fourth coding block ; First, the image That is, the object image input initial encoding block , get the initial feature map , the initial feature map Enter the first encoding block , get the first feature map , the first feature map Enter the second code block , and get the second feature map , the second feature map Enter the third encoding block , and get the third feature map , the third feature map Enter the fourth code block , and get the fourth feature map That is, the final output of the adaptive multi-source feature capture branch.

[0016] Furthermore, the local enhanced self-attention branch is composed of the parallel first local enhanced self-attention module , the second local enhanced self-attention module And the third local enhanced self-attention module Composition: The first local enhanced self-attention module Receive the first feature map And process it to get the first local enhanced feature map ; Second local enhanced self-attention module Receive the second feature map And process it to get the second local enhanced feature map ; The third local enhanced self-attention module Receive the third feature map And process it to get the third local enhanced feature map .

[0017] Furthermore, the local understanding branch includes a detail enhancement module DEM, which is connected to the initial local understanding decoding block in sequence. , the first local understanding decoding block , the second local understanding decoding block , the third local understanding decoding block and the fourth local understanding decoding block ; First, the detail enhancement module receives the fourth feature map Process and output detail enhancement feature map , the detail enhancement feature map and the fourth feature map Input initial local understanding decoding block Get the initial local understanding decoding feature map , the initial local understanding decoding feature map And the third local enhanced feature map Input the first local understanding decoding block The first local understanding decoding feature map is obtained , decode the first local feature map and the second local enhanced feature map Input the second local understanding decoding block The second local understanding decoding feature map is obtained , decode the second local feature map And the first local enhanced feature map Input the third local understanding decoding block The third local understanding decoding feature map is obtained , decode the third local feature map and the initial feature map Enter the fourth local understanding decoding block The fourth local understanding decoding feature map is obtained That is, the final output of the local understanding branch;

[0018] The detail enhancement module DEM consists of a parallel dilated convolution sub-channel and a pooled convolution sub-channel; the dilated convolution sub-channel consists of three layers of dilated convolution layers with a dilation rate of 2 connected in series, and each dilated convolution layer is followed by a batch normalization layer BN and a ReLU activation function; the pooled convolution sub-channel consists of a maximum pooling layer, a 1×3 convolution layer, and a 3×1 convolution layer connected in series;

[0019] The detail enhancement module receives the fourth feature map Process and output detail enhancement feature map The specific process is as follows: In the dilated convolution subchannel, the fourth feature map After three layers of 3×3 dilated convolution layers with a dilation rate of 2, the dilated feature map is obtained. ; In the pooled convolution subchannel, the fourth feature map The pooled feature map is obtained by sequentially passing through the maximum pooling layer, 1×3 convolution layer and 3×1 convolution layer. , the dilated feature map And pooling feature map After splicing in the channel dimension, the detail enhancement feature map is obtained .

[0020] Furthermore, the cross-scale atrous convolution pooling branch is composed of the parallel first cross-scale atrous convolution pooling module , the second cross-scale dilated convolution pooling module And the third cross-scale dilated convolution pooling module Composition: The first cross-scale dilated convolution pooling module Receive the first feature map And process to get the first cross-scale feature map ; The second cross-scale dilated convolution pooling module Receive the second feature map And process to get the second cross-scale feature map ; The third cross-scale dilated convolution pooling module Receive the third feature map And process to get the third cross-scale feature map ; The first cross-scale feature map , the second cross-scale feature map And the third cross-scale feature map After splicing in the channel dimension, we get the cross-scale dilated convolutional pooling branch feature map. That is, the output of the cross-scale dilated convolution pooling branch;

[0021] The first cross-scale dilated convolutional pooling module , the second cross-scale dilated convolution pooling module , the third cross-scale hole convolution pooling module The structures are the same, consisting of a 1×3 convolution layer, a 3×1 convolution layer, a 1×3 dilated convolution layer, a 3×1 dilated convolution layer, a 7×7 average pooling layer, a 11×11 average pooling layer, a 17×17 average pooling layer, a 23×23 average pooling layer, two 1×1 convolution layers and a 3×3 convolution layer;

[0022] The first cross-scale dilated convolutional pooling module Receive the first feature map And process to get the first cross-scale feature map The specific process is as follows: First, the first feature map After a 1×3 convolutional layer, the first convolution feature map is obtained ; The first feature map After a 1×3 hole convolution layer, the first hole convolution feature map is obtained. ; The first feature map After a 3×1 hole convolution layer, the second hole convolution feature map is obtained ; The first feature map After a 3×1 convolution layer, the second convolution feature map is obtained ; The first convolution feature map And the first feature map After splicing in the channel dimension, the first convolution splicing feature map is obtained ; Splice the first convolution feature map Input to the 7×7 average pooling layer for processing to obtain the first pooling feature map ; The first hole convolution feature map And the first feature map After splicing in the channel dimension, the second convolution splicing feature map is obtained , concatenate the second convolution feature map Input to the 11×11 average pooling layer for processing to obtain the second pooling feature map ; The second hole convolution feature map And the first feature map After splicing in the channel dimension, the third convolution splicing feature map is obtained , concatenate the third convolution feature map Input to the 17×17 average pooling layer for processing to obtain the third pooling feature map ; The second convolution feature map And the first feature map After splicing in the channel dimension, the fourth convolution splicing feature map is obtained , concatenate the fourth convolution feature map Input to the 23×23 average pooling layer for processing to obtain the fourth pooling feature map ; The first pooling feature map And the second pooling feature map After splicing in the channel dimension, the first pooled splicing feature map is obtained , the first pooling splicing feature map After a 1×1 convolutional layer, the first fused feature map is obtained ; The third pooling feature map And the fourth pooling feature map After splicing by channel, we get the second pooled splicing feature map , the second pooling splicing feature map After a 1×1 convolution layer, the second fusion feature map is obtained ; The first feature map , the first fusion feature map and the second fusion feature map After splicing in the channel dimension, the pooled fusion feature map is obtained , the pooled fusion feature map obtained by splicing After a 3×3 convolutional layer, the first cross-scale feature map is obtained .

[0023] Furthermore, the global semantic guidance branch includes a recursive criss-cross attention module, which is sequentially connected to the initial global guidance decoding block. , initial semantic guidance module, first global guidance decoding block , first semantic guidance module, second global guidance decoding block , the second semantic guidance module, the third global guidance decoding block and the third semantic guidance module; first, the fourth feature map is processed by a recursive cross-attention module , get the attention enhancement feature map , the attention-enhanced feature map And cross-scale dilated convolution pooling branch feature map Input initial global boot decoding block The initial global guided decoding module feature map is obtained from , the initial global guided decoding module feature map and attention-enhanced feature maps Input into the initial semantic guidance module to obtain the initial semantic guidance feature map, and input the initial semantic guidance feature map into the first global guidance decoding block The first global guided decoding feature map is obtained from , the first global guided decoding feature map and attention-enhanced feature maps Input into the first semantic guidance module to obtain the first semantic guidance feature map, and input the first semantic guidance feature map into the second global guidance decoding block , get the second global guided decoding feature map , the second global guided decoding feature map and attention-enhanced feature maps Input into the second semantic guidance module to obtain the second semantic guidance feature map, and input the second semantic guidance feature map into the third global guidance decoding block The third global guided decoding feature map is obtained , the third global guided decoding feature map and attention-enhanced feature maps The third semantic guidance feature map is input into the third semantic guidance module, which is the final output of the global semantic guidance branch.

[0024] The initial semantic guidance module, the first semantic guidance module, the second semantic guidance module and the third semantic guidance module have the same structure, which consists of two 1×1 convolutional layers and one gated convolutional layer;

[0025] The initial global guided decoding module feature map and attention-enhanced feature maps Input into the initial semantic guidance module to obtain the initial semantic guidance feature map The specific process is as follows: First, the initial global guided decoding module feature map After a 1×1 convolution layer and interpolation upsampling operation, the upsampled feature map is obtained. , adjust the upsampled feature map The number of channels and size of the attention-enhanced feature map Consistent and attention-enhanced feature maps Splice by channel to get the spliced ​​feature map , the concatenated feature map Attention-enhanced feature map after difference upsampling The fused output is then fed into a gated convolutional layer for fusion. The fused output is then passed through a 1×1 convolutional layer to further adjust the number of channels and obtain the initial semantically guided feature map. , expressed as:

[0026] ;

[0027] ;

[0028] ;

[0029] In the formula, Indicates difference upsampling; represents a 1×1 convolutional layer; represents a gated convolutional layer.

[0030] Furthermore, the initial coding block It includes a 3×3 convolution layer, a batch normalization layer, a ReLU activation function and a maximum pooling operation layer; first, the object image is processed by the 3×3 convolution layer, the batch normalization layer, and the ReLU activation function in turn to obtain the initial intermediate feature map , then, the initial intermediate feature map The initial feature map is obtained by spatial downsampling through the maximum pooling operation layer ;

[0031] First coded block , the second coding block , the third coding block and the fourth coding block The structures of are the same, consisting of a maximum pooling layer, a deformable convolutional layer, and a ResNet-34 residual block;

[0032] The initial feature map Enter the first encoding block , get the first feature map The specific process is: the initial feature map Enter the first encoding block The first feature map is obtained by sequentially processing the maximum pooling layer, deformable convolution layer and ResNet-34 residual block. .

[0033] Furthermore, the first local enhanced self-attention module , the second local enhanced self-attention module And the third local enhanced self-attention module The structure is the same as that of the first local enhanced self-attention module Receive the first feature map And process it to get the first local enhanced feature map The specific process is as follows: First, the first feature map The key matrix is ​​obtained by processing the convolution layer with a convolution kernel size of 3×3. ; The first feature map After processing by a convolution layer with a convolution kernel size of 5×5, the value matrix is ​​obtained ; Perform square sum and mutual dot product processing inside the bond matrix to obtain the affinity matrix ; The affinity matrix Perform normalization operation to obtain the normalized matrix ; Normalize the matrix With value matrix Perform matrix multiplication to obtain weighted feature map ; The weighted feature map With value matrix Concatenate in the channel dimension to get the first local enhanced feature map .

[0034] Furthermore, the initial global guided decoding block , the first global guide decoding block , Second global guide decoding block and the third global boot decoding block The structures are the same, consisting of three 3×3 convolutional layers, three batch normalization layers, three ReLU activation functions, and an interpolation upsampling operation stacked together;

[0035] Attention-enhanced feature map And cross-scale dilated convolution pooling branch feature map Input initial global boot decoding block The initial global guided decoding module feature map is obtained from The specific process is as follows: First, the attention enhancement feature map And cross-scale dilated convolution pooling branch feature map After splicing in the channel dimension, we get the attention splicing feature map ; Attention splicing feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the first global feature map is obtained. ; The first global feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the second global feature map is obtained. ; The second global feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the third global feature map is obtained. ; The third global feature map After an interpolation upsampling process, the initial global guided decoding module feature map is obtained .

[0036] Furthermore, the initial local understanding decoding block , the first local understanding decoding block , the second local understanding decoding block , the third local understanding decoding block and the fourth local understanding decoding block The structures are the same, consisting of three 3×3 convolutional layers, three batch normalization layers BN, three ReLU activation functions and a maximum pooling layer stacked;

[0037] The detail enhancement feature map and the fourth feature map Input initial local understanding decoding block Get the initial local understanding decoding feature map The specific process is as follows: First, the detail enhancement feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the first local feature map is obtained. ; The first local feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the second local feature map is obtained. ; The second local feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the third local feature map is obtained. ; The third local feature map After a maximum pooling layer, the initial local understanding decoding feature map is obtained .

[0038] Compared with the existing technology, the present invention has the following beneficial effects:

[0039] (1) This paper constructs a new multi-branch collaborative architecture. From the perspective of multi-branch collaboration of adaptive multi-scale feature extraction and global-local feature enhancement, it proposes an adaptive multi-source feature capture branch composed of deformable convolution, residual structure and multi-level pooling. By dynamically adjusting the convolution sampling area and feature capture method, it significantly improves the ability to capture irregular edges and reduces computational redundancy in complex scenes. Combining the detail enhancement module DEM of the local understanding branch with the recursive cross attention RCCA of the global semantic guidance branch, pixel-level transparency prediction is achieved without the need for auxiliary input.

[0040] (2) This invention adopts the innovative design of the local enhanced self-attention branch module, uses convolutions of different sizes to generate key matrices and value matrices of different dimensions, and performs square sum and mutual dot product processing inside the key matrix to obtain the affinity matrix, which greatly reduces the computational complexity of the standard self-attention while retaining the local context relevance. Combined with the cross-level feature interaction mechanism, the accuracy of high-frequency detail recovery is improved with the same number of parameters.

[0041] (3) The present invention constructs a multi-layer cross-scale processing operation through a multi-level feature fusion design of a cross-scale dilated convolution pooling branch, and processes the encoding features at different levels respectively. Through multiple layers of convolution, dilated convolution and pooling layers, the receptive field is expanded and contextual information at different levels is captured, achieving the complementarity of micro-texture and macro-structure.

[0042] (4) The multi-type foreground adaptive training strategy proposed in the present invention divides the synthetic image into three categories: significantly opaque, significantly transparent, and non-significant according to the foreground saliency and transparency characteristics, and designs a differentiated auxiliary image generation strategy to solve the problem that traditional technologies are difficult to finely cut out the foreground of multiple types of objects. The present invention does not require complex and tedious processes, meets industrial-level real-time requirements, and can achieve automatic image cutout. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic diagram of the structure of the automatic cutout model of the present invention.

[0044] Figure 2 It is a schematic diagram of the structure of the detail enhancement module of the present invention.

[0045] Figure 3 This is a schematic diagram of the structure of the cross-scale atrous convolution pooling module of the present invention.

[0046] Figure 4 It is a schematic diagram of the structure of the semantic guidance module of the present invention. DETAILED DESCRIPTION

[0047] The present invention provides a technical solution: an automatic image cutting method based on adaptive feature extraction and semantic guidance, comprising the following steps:

[0048] Step S1: construct a cutout dataset, which includes several object images.

[0049] The cutout dataset includes the DUTS dataset and the splicing dataset. The DUTS dataset contains 10,553 high-quality object images, covering a variety of backgrounds and foregrounds, and each image is finely annotated for the pre-training stage; the splicing dataset includes the foreground images in the Composition-1k dataset, the HAtt dataset, the AM-2k dataset, the tree image cutout dataset, and the background images in the BG-20K dataset. Then, each foreground and background image are synthesized to obtain a synthetic image; an auxiliary image is generated for each synthetic image that is only used for training. The value of the auxiliary image is only The auxiliary map has three values: 0, 0.5, and 1; according to the feature differences of each foreground, the synthetic images are divided into three categories, including foreground significant opaque synthetic images, foreground significant transparent synthetic images, and foreground non-significant synthetic images; for foreground significant opaque synthetic images, the auxiliary map has three values ​​0, 0.5, and 1; for foreground significant transparent synthetic images, the auxiliary map has two values ​​0 and 0.5; for foreground non-significant synthetic images, the value of the auxiliary map is all 0; finally, 11297 images are generated for fine-tuning; the AIM-500 and tree image test sets are used in the test phase, and the total number of test image data sets is 792 images.

[0050] Step S2: Construct an automatic cutout model, such as Figure 1 As shown in the figure, the automatic cutout model consists of an adaptive multi-source feature capture branch, a local enhanced self-attention branch, a local understanding branch, a cross-scale atrous convolution pooling branch, a global semantic guidance branch and an aggregation module; the local understanding branch consists of a detail enhancement module and multiple local understanding decoding blocks in series; the global semantic guidance branch consists of a recursive cross-attention and multiple semantic guidance modules in series, and a global guidance decoding block is set in front of each semantic guidance module.

[0051] Step S3: Input the object image into the adaptive multi-source feature capture branch to extract multi-level feature maps.

[0052] Among them, the adaptive multi-source feature capture branch consists of five encoding blocks, including the initial encoding block , the first coding block , the second coding block , the third coding block and the fourth coding block ; First, the image (Object image) Input initial encoding block , get the initial feature map , the initial feature map Enter the first encoding block , get the first feature map , the first feature map Enter the second code block , and get the second feature map , the second feature map Enter the third encoding block , and get the third feature map , the third feature map Enter the fourth code block , and get the fourth feature map That is, the final output of the adaptive multi-source feature capture branch.

[0053] Among them, the initial coding block It includes a 3×3 convolution layer, a batch normalization layer (BatchNormalization), a ReLU activation function and a maximum pooling operation layer; first, the image (Object image) is processed by 3×3 convolution layer, batch normalization layer, and ReLU activation function in sequence to obtain the initial intermediate feature map , then, the initial intermediate feature map The initial feature map is obtained by spatial downsampling through the maximum pooling operation layer , to reduce the spatial size of the feature map. Among them, the 3×3 convolution layer uses the convolution kernel to Slide up to extract features, and its operation can be expressed by the following formula:

[0054] (1);

[0055] In the formula, Represents the output of the 3×3 convolutional layer Row and The value of the column; is an image Middle position The value at Represents the convolution kernel The row index of Represents the convolution kernel Column index of Represents the convolution kernel Line The value of the column; and are the height and width of the convolution kernel respectively.

[0056] Among them, the first coding block , the second coding block , the third coding block and the fourth coding block The structures of are the same, which are composed of a maximum pooling layer, a deformable convolutional layer and a ResNet-34 residual block.

[0057] Among them, the initial feature map Enter the first encoding block , get the first feature map The specific process is: the initial feature map Enter the first encoding block The first feature map is obtained by sequentially processing the maximum pooling layer, deformable convolution layer and ResNet-34 residual block. ; Second coding block , the third coding block , the fourth coding block and the first coded block The structure is the same as that of , and the processing flow will not be described in detail here.

[0058] Step S4: Input the multi-level feature map extracted by the adaptive multi-source feature capture branch into the local enhanced self-attention branch to obtain a local enhanced feature map.

[0059] Among them, the local enhanced self-attention branch consists of the first parallel local enhanced self-attention module , the second local enhanced self-attention module And the third local enhanced self-attention module Composition: The first local enhanced self-attention module Receive the first feature map And process it to get the first local enhanced feature map ; Second local enhanced self-attention module Receive the second feature map And process it to get the second local enhanced feature map ; The third local enhanced self-attention module Receive the third feature map And process it to get the third local enhanced feature map .

[0060] Among them, the first local enhanced self-attention module , the second local enhanced self-attention module And the third local enhanced self-attention module The structure is the same.

[0061] The local enhanced self-attention module is different from the standard self-attention mechanism. In the standard self-attention mechanism, the input feature map is projected into the query matrix, key matrix and value matrix through linear transformation. The query matrix, key matrix and value matrix are calculated through dot product to calculate the attention matrix, which makes the dimensions of each query matrix, key matrix and value matrix the same and fixed. The local enhanced self-attention module customizes a new self-attention weight calculation method based on the core idea of ​​the self-attention mechanism, and performs feature projection of the key matrix and value matrix through convolution operation; it reduces the computational complexity while ensuring the local context understanding ability of the model.

[0062] Among them, the first local enhanced self-attention module Receive the first feature map And process it to get the first local enhanced feature map The specific process is as follows: First, the first feature map The key matrix is ​​obtained by processing the convolution layer with a convolution kernel size of 3×3. ; The first feature map After processing by a convolution layer with a convolution kernel size of 5×5, the value matrix is ​​obtained ; Perform square sum and mutual dot product processing inside the bond matrix to obtain the affinity matrix ; The affinity matrix Perform normalization operation to obtain the normalized matrix ; Normalize the matrix With value matrix Perform matrix multiplication to obtain weighted feature map ; The weighted feature map With value matrix Concatenate in the channel dimension to get the first local enhanced feature map , which can be expressed as:

[0063] (2);

[0064] (3);

[0065] (4);

[0066] (5);

[0067] (6);

[0068] (7);

[0069] In the formula, represents the bond matrix; represents a 3×3 convolutional layer; represents a 5×5 convolutional layer; Represents the key matrix Middle The sum of the squares of the component vectors; Indicates the key matrix The vector and The dot product of the component vectors; Represents the key matrix Middle The sum of the squares of the component vectors; Indicates the key matrix The vector and The affinity of the component vectors; Represents a normalization operation; Represents a concatenation operation in the channel dimension; Represents the dimension of each component vector in the key matrix. The dimension of each component vector in the key matrix is ​​the same.

[0070] Step S5: The detail enhancement module DEM receives the output of the last layer of the adaptive multi-source feature capture branch and processes it together with the local enhancement feature map, and then inputs it into the subsequent local understanding decoding block for layer-by-layer processing to obtain the local understanding decoding feature map.

[0071] The local understanding branch includes a detail enhancement module DEM, which is followed by the initial local understanding decoding block. , the first local understanding decoding block , the second local understanding decoding block , the third local understanding decoding block and the fourth local understanding decoding block ; First, the detail enhancement module receives the fourth feature map Process and output detail enhancement feature map , the detail enhancement feature map and the fourth feature map Input initial local understanding decoding block Get the initial local understanding decoding feature map , the initial local understanding decoding feature map And the third local enhanced feature map Input the first local understanding decoding block The first local understanding decoding feature map is obtained , decode the first local feature map and the second local enhanced feature map Input the second local understanding decoding block The second local understanding decoding feature map is obtained , decode the second local feature map And the first local enhanced feature map Input the third local understanding decoding block The third local understanding decoding feature map is obtained , decode the third local feature map and the initial feature map Enter the fourth local understanding decoding block The fourth local understanding decoding feature map is obtained That is the final output of the local understanding branch.

[0072] like Figure 2As shown in Figure 2, the detail enhancement module DEM extracts local detail information through parallel expansion convolution sub-channels and pooling convolution sub-channels. The outputs of the expansion convolution sub-channel and the pooling convolution sub-channel are spliced ​​through channels to obtain a detail enhancement feature map. This is the final output of the detail enhancement module.

[0073] Among them, the dilated convolution sub-channel is composed of three layers of dilated convolution layers with a dilation rate of 2 connected in series, and a batch normalization layer BN and ReLU activation function are added after each dilated convolution layer to ensure the stability of the training process and the nonlinear expression ability of the network; the pooled convolution sub-channel is composed of a maximum pooling layer, a 1×3 convolution layer and a 3×1 convolution layer connected in sequence. The maximum pooling layer reduces the spatial complexity and increases the receptive field, and the ordinary convolution layer is used to further extract detail features; the detail enhancement module receives the fourth feature map Process and output detail enhancement feature map The specific process is as follows: In the dilated convolution subchannel, the fourth feature map After three layers of 3×3 dilated convolution layers with a dilation rate of 2, the dilated feature map is obtained. ; In the pooled convolution subchannel, the fourth feature map The pooled feature map is obtained by sequentially passing through the maximum pooling layer, 1×3 convolution layer and 3×1 convolution layer. , the dilated feature map And pooling feature map After splicing in the channel dimension, the detail enhancement feature map is obtained , expressed as:

[0074] (8);

[0075] (9);

[0076] (10);

[0077] (11);

[0078] (12);

[0079] (13);

[0080] In the formula, is a 3×3 dilated convolutional layer with a dilation rate of 2; is the output of the first dilated convolution layer in the dilated convolution subchannel; is the output of the second dilated convolution layer in the dilated convolution subchannel; is the output of the third dilated convolution layer in the dilated convolution subchannel; is the maximum pooling layer; It is a 3×1 convolution layer; It is a 1×3 convolutional layer; It represents a splicing operation in the channel dimension; by fusing information of two different scales, the ability to restore image details can be improved.

[0081] Among them, the initial local understanding decoding block , the first local understanding decoding block , the second local understanding decoding block , the third local understanding decoding block and the fourth local understanding decoding block The structures are the same, consisting of three 3×3 convolutional layers, three batch normalization layers (BN), three ReLU activation functions and a maximum pooling layer stacked together.

[0082] Among them, the detail enhancement feature map and the fourth feature map Input initial local understanding decoding block Get the initial local understanding decoding feature map The specific process is as follows: First, the detail enhancement feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the first local feature map is obtained. ; The first local feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the second local feature map is obtained. ; The second local feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the third local feature map is obtained. ; The third local feature map After a maximum pooling layer, the initial local understanding decoding feature map is obtained .

[0083] Step S6: input the multi-level feature map extracted by the adaptive multi-source feature capture branch into the cross-scale dilated convolution pooling branch to obtain a cross-scale dilated convolution pooling branch feature map.

[0084] Among them, the cross-scale dilated convolution pooling branch consists of the parallel first cross-scale dilated convolution pooling module , the second cross-scale dilated convolution pooling module And the third cross-scale dilated convolution pooling module Composition: The first cross-scale dilated convolution pooling module Receive the first feature map And process to get the first cross-scale feature map ; The second cross-scale dilated convolution pooling module Receive the second feature map And process to get the second cross-scale feature map ; The third cross-scale dilated convolution pooling module Receive the third feature map And process to get the third cross-scale feature map ; The first cross-scale feature map , the second cross-scale feature map And the third cross-scale feature map After splicing in the channel dimension, we get the cross-scale dilated convolutional pooling branch feature map. That is, the output of the cross-scale atrous convolution pooling branch.

[0085] like Figure 3 As shown, the first cross-scale hole convolution pooling module , the second cross-scale dilated convolution pooling module , the third cross-scale hole convolution pooling module The structures are the same, consisting of a 1×3 convolutional layer, a 3×1 convolutional layer, a 1×3 dilated convolutional layer, a 3×1 dilated convolutional layer, a 7×7 average pooling layer, a 11×11 average pooling layer, a 17×17 average pooling layer, a 23×23 average pooling layer, two 1×1 convolutional layers and one 3×3 convolutional layer.

[0086] Among them, the first cross-scale hole convolution pooling module Receive the first feature map And process to get the first cross-scale feature map The specific process is as follows: First, the first feature map After a 1×3 convolutional layer, the first convolution feature map is obtained ; The first feature map After a 1×3 hole convolution layer, the first hole convolution feature map is obtained. ; The first feature map After a 3×1 hole convolution layer, the second hole convolution feature map is obtained ; The first feature map After a 3×1 convolution layer, the second convolution feature map is obtained ; The first convolution feature map And the first feature map After splicing in the channel dimension, the first convolution splicing feature map is obtained ; Splice the first convolution feature map Input to the 7×7 average pooling layer for processing to obtain the first pooling feature map ; The first hole convolution feature map And the first feature map After splicing in the channel dimension, the second convolution splicing feature map is obtained , concatenate the second convolution feature map Input to the 11×11 average pooling layer for processing to obtain the second pooling feature map ; The second hole convolution feature map And the first feature map After splicing in the channel dimension, the third convolution splicing feature map is obtained , concatenate the third convolution feature map Input to the 17×17 average pooling layer for processing to obtain the third pooling feature map ; The second convolution feature map And the first feature map After splicing in the channel dimension, the fourth convolution splicing feature map is obtained , concatenate the fourth convolution feature map Input to the 23×23 average pooling layer for processing to obtain the fourth pooling feature map ; The first pooling feature map And the second pooling feature map After splicing in the channel dimension, the first pooled splicing feature map is obtained , the first pooling splicing feature map After a 1×1 convolutional layer, the first fused feature map is obtained ; The third pooling feature map And the fourth pooling feature map After splicing by channel, we get the second pooled splicing feature map , the second pooling splicing feature map After a 1×1 convolution layer, the second fusion feature map is obtained ; The first feature map , the first fusion feature map and the second fusion feature map After splicing in the channel dimension, the pooled fusion feature map is obtained , the pooled fusion feature map obtained by splicing After a 3×3 convolutional layer, the first cross-scale feature map is obtained , which can be expressed as:

[0087] (14);

[0088] (15);

[0089] (16);

[0090] (17);

[0091] (18);

[0092] (19);

[0093] (20);

[0094] (twenty one);

[0095] (twenty two);

[0096] (twenty three);

[0097] (twenty four);

[0098] (25);

[0099] (26);

[0100] (27);

[0101] In the formula, It is a 3×1 convolution layer; It is a 1×3 convolutional layer; It is a 1×3 hole convolution layer; It is a 3×1 hole convolution layer; It is a 7×7 average pooling layer; It is an 11×11 average pooling layer; It is a 17×17 average pooling layer; It is a 23×23 average pooling layer.

[0102] Step S7: Recursive cross attention receives the output of the last layer of the adaptive multi-source feature capture branch and processes it with the cross-scale dilated convolution pooling branch features. Figure 1 The subsequent global guided decoding block and semantic guided module process the same input layer by layer to obtain the global semantic guided decoding feature map; among them, the semantic guided module receives the output of the global guided decoding block and recursive cross attention of the previous layer.

[0103] The global semantic guidance branch includes a recurrent criss-cross attention module (RCCA), which is connected to the initial global guidance decoding block in sequence. , initial semantic guidance module, first global guidance decoding block , first semantic guidance module, second global guidance decoding block , the second semantic guidance module, the third global guidance decoding block and the third semantic guidance module; first, the fourth feature map is processed by a recursive cross-attention module The irregular structure in the image is used to obtain the attention-enhanced feature map , the attention-enhanced feature map And cross-scale dilated convolution pooling branch feature map Input initial global boot decoding block The initial global guided decoding module feature map is obtained from , the initial global guided decoding module feature map and attention-enhanced feature maps Input into the initial semantic guidance module to obtain the initial semantic guidance feature map, and input the initial semantic guidance feature map into the first global guidance decoding block The first global guided decoding feature map is obtained from , the first global guided decoding feature map and attention-enhanced feature maps Input into the first semantic guidance module to obtain the first semantic guidance feature map, and input the first semantic guidance feature map into the second global guidance decoding block , get the second global guided decoding feature map , the second global guided decoding feature map and attention-enhanced feature maps Input into the second semantic guidance module to obtain the second semantic guidance feature map, and input the second semantic guidance feature map into the third global guidance decoding block The third global guided decoding feature map is obtained , the third global guided decoding feature map and attention-enhanced feature maps The third semantic guidance feature map obtained by inputting it into the third semantic guidance module is the final output of the global semantic guidance branch.

[0104] Among them, the initial global guided decoding block , the first global guide decoding block , Second global guide decoding block and the third global boot decoding block The structures are the same, consisting of three 3×3 convolutional layers, three batch normalization layers, three ReLU activation functions, and an interpolation upsampling operation stacked together.

[0105] Among them, the attention enhancement feature map And cross-scale dilated convolution pooling branch feature map Input initial global boot decoding block The initial global guided decoding module feature map is obtained from The specific process is as follows: First, the attention enhancement feature map And cross-scale dilated convolution pooling branch feature map After splicing in the channel dimension, we get the attention splicing feature map ; Attention splicing feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the first global feature map is obtained. ; The first global feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the second global feature map is obtained. ; The second global feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the third global feature map is obtained. ; The third global feature map After an interpolation upsampling process, the initial global guided decoding module feature map is obtained .

[0106] like Figure 4 As shown, the initial semantic guidance module, the first semantic guidance module, the second semantic guidance module and the third semantic guidance module have the same structure, which are composed of two 1×1 convolutional layers and one gated convolutional layer.

[0107] The initial global guided decoding module feature map and attention-enhanced feature maps Input into the initial semantic guidance module to obtain the initial semantic guidance feature map The specific process is as follows: First, the initial global guided decoding module feature map After a 1×1 convolution layer and interpolation upsampling operation, the upsampled feature map is obtained. , adjust the upsampled feature map The number of channels and size of the attention-enhanced feature map Consistent and attention-enhanced feature maps Splice by channel to get the spliced ​​feature map , the concatenated feature map Attention-enhanced feature map after difference upsampling The fused output is then fed into a gated convolutional layer for fusion. The fused output is then passed through a 1×1 convolutional layer to further adjust the number of channels and obtain the initial semantically guided feature map. , expressed as:

[0108] (28);

[0109] (29);

[0110] (30);

[0111] In the formula, Indicates difference upsampling; represents a 1×1 convolutional layer; represents a gated convolutional layer.

[0112] Step S8: Aggregate the local understanding decoding feature map and the global semantic guided decoding feature map through the aggregation module to obtain a transparency mask, which is the final output.

[0113] The aggregation module is used to fuse the outputs of the local understanding branch and the global semantic guidance branch to generate the final transparency mask; the aggregation module receives the fourth local understanding decoding feature map Combined with the third semantically guided feature map and weighted operation to obtain the transparency mask , expressed as:

[0114] (31);

[0115] In the formula, Represents the third semantic guided feature map The foreground mask of , that is, the predicted foreground area.

[0116] The automatic cutout model is optimized in various ways during training, and the overall loss of the automatic cutout model It consists of multiple losses, including the loss of the local understanding branch , the loss of the global semantics-guided branch and the loss of the aggregation module , the overall loss of the automatic cutout model It is expressed as:

[0117] (32).

[0118] The local understanding branch focuses on estimating the transparency of the details in the image. The loss of this branch is composed of Alpha loss and Laplace loss It consists of two parts, and the formula is as follows:

[0119] (33);

[0120] (34);

[0121] (35);

[0122] In the formula, is the number of samples, that is, the total number of pixels of the input object image; Represents the predicted transparency value of the i-th pixel in the input object image; Represents the true transparency value of the i-th pixel in the input object image; Represents the weight value obtained according to the auxiliary map. For the area with a value of 0 in the auxiliary map, ; For the area with value 1 in the auxiliary graph, , the rest of the area ; Represents the scale number of the Laplace pyramid; Indicated in scale Laplace operation on ; Indicated in scale Forecast transparency chart Gradient information of the Laplace operation performed; Indicated in scale The real transparency map above Gradient information of the Laplacian operation performed.

[0123] The global semantic guidance branch is used to optimize the classification of the foreground, background, and uncertain areas of the image; the loss of the global semantic guidance branch It is expressed as:

[0124] (36);

[0125] In the formula, is the number of categories, including background, foreground, and uncertain area; It is The true label of each pixel; Indicates that for pixels belong to the class The predicted probability value.

[0126] The aggregation module is used to supervise the final transparency map so that it can maintain high quality at both global and local levels. It is expressed as:

[0127] (37);

[0128] In the formula, Represents the blended transparency mask The predicted transparency value at the i-th pixel position.

[0129] The automatic cutout model is trained in a supervised learning manner using an annotated automatic cutout dataset. In order to improve training efficiency and avoid training the model from scratch, a transfer learning strategy is adopted. During the training process, the Adam optimizer (Adaptive Moment Estimation) is used to update the weights of the entire network. In the transfer learning process, the pre-training phase is first carried out, and the learning rate of the pre-training is set to 0.0001, the training batch is 16, and the number of training rounds is 100. In this phase, the weights of the pre-trained model are used for initialization, which can accelerate the convergence of the model and improve performance. After that, the fine-tuning training phase is entered, and the learning rate is further reduced to 0.000001, the fine-tuning training batch is 16, and the number of training rounds is set to 50. The main purpose of the fine-tuning phase is to adjust the network parameters to make it better adapt to the data of the target task. After all training is completed, the optimal parameters of the model will be saved, predicted and evaluated on the test set.

[0130] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An automatic image cutting method based on adaptive feature extraction and semantic guidance, characterized in that: The steps include: Step S1: construct a cutout dataset, which includes a number of object images; Step S2: construct an automatic cutout model, which consists of an adaptive multi-source feature capture branch, a local enhanced self-attention branch, a local understanding branch, a cross-scale dilated convolution pooling branch, a global semantic guidance branch and an aggregation module; The local understanding branch consists of a detail enhancement module and multiple local understanding decoding blocks in series; the global semantic guidance branch consists of a recursive cross attention and multiple semantic guidance modules in series, and a global guidance decoding block is set before each semantic guidance module; Step S3: input the object image into the adaptive multi-source feature capture branch to extract a multi-level feature map; Step S4: input the extracted multi-level feature map into the local enhanced self-attention branch to obtain a local enhanced feature map; Step S5: the detail enhancement module DEM receives the output of the last layer of the adaptive multi-source feature capture branch and processes it together with the local enhancement feature map, and then inputs it into the subsequent local understanding decoding block for layer-by-layer processing to obtain the local understanding decoding feature map; Step S6: input the extracted multi-level feature map into the cross-scale dilated convolution pooling branch to obtain a cross-scale dilated convolution pooling branch feature map; Step S7: After the recursive cross attention receives the output of the last layer of the adaptive multi-source feature capture branch, it is input together with the cross-scale dilated convolution pooling branch feature map into the subsequent global guide decoding block and the semantic guide module for layer-by-layer processing to obtain a global semantic guide decoding feature map; wherein the semantic guide module receives the output of the global guide decoding block and the recursive cross attention of the previous layer; Step S8: Aggregate the local understanding decoding feature map and the global semantic guided decoding feature map through the aggregation module to obtain a transparency mask, which is the final output.

2. The method for automatic image cutting based on adaptive feature extraction and semantic guidance according to claim 1, characterized in that: The adaptive multi-source feature capture branch consists of five encoding blocks, including the initial encoding block , the first coding block , the second coding block , the third coding block and the fourth coding block ; First, the image That is, the object image input initial encoding block , get the initial feature map , the initial feature map Enter the first encoding block , get the first feature map , the first feature map Enter the second code block , and get the second feature map , the second feature map Enter the third encoding block , and get the third feature map , the third feature map Enter the fourth code block , and get the fourth feature map That is, the final output of the adaptive multi-source feature capture branch.

3. The method for automatic image cutting based on adaptive feature extraction and semantic guidance according to claim 2, characterized in that: The local enhanced self-attention branch consists of the first local enhanced self-attention module in parallel , the second local enhanced self-attention module And the third local enhanced self-attention module Composition: The first local enhanced self-attention module Receive the first feature map And process it to get the first local enhanced feature map ; Second local enhanced self-attention module Receive the second feature map And process it to get the second local enhanced feature map ; The third local enhancement self-attention module Receive the third feature map And process it to get the third local enhanced feature map .

4. The method for automatic image cutting based on adaptive feature extraction and semantic guidance according to claim 3, characterized in that: The local understanding branch includes a detail enhancement module DEM, which is followed by the initial local understanding decoding block , the first local understanding decoding block , the second local understanding decoding block , the third local understanding decoding block and the fourth local understanding decoding block ; First, the detail enhancement module receives the fourth feature map Process and output detail enhancement feature map , the detail enhancement feature map and the fourth feature map Input initial local understanding decoding block Get the initial local understanding decoding feature map , the initial local understanding decoding feature map And the third local enhanced feature map Input the first local understanding decoding block The first local understanding decoding feature map is obtained , decode the first local feature map and the second local enhanced feature map Input the second local understanding decoding block The second local understanding decoding feature map is obtained , decode the second local feature map And the first local enhanced feature map Input the third local understanding decoding block The third local understanding decoding feature map is obtained , decode the third local feature map and the initial feature map Enter the fourth local understanding decoding block The fourth local understanding decoding feature map is obtained That is, the final output of the local understanding branch; The detail enhancement module DEM consists of a parallel dilated convolution sub-channel and a pooled convolution sub-channel; the dilated convolution sub-channel consists of three layers of dilated convolution layers with a dilation rate of 2 connected in series, and each dilated convolution layer is followed by a batch normalization layer BN and a ReLU activation function; the pooled convolution sub-channel consists of a maximum pooling layer, a 1×3 convolution layer, and a 3×1 convolution layer connected in series; The detail enhancement module receives the fourth feature map Process and output detail enhancement feature map The specific process is as follows: In the dilated convolution subchannel, the fourth feature map After three layers of 3×3 dilated convolution layers with a dilation rate of 2, the dilated feature map is obtained. ; In the pooled convolution subchannel, the fourth feature map The pooled feature map is obtained by sequentially passing through the maximum pooling layer, 1×3 convolution layer and 3×1 convolution layer. , the dilated feature map And pooling feature map After splicing in the channel dimension, the detail enhancement feature map is obtained .

5. The method for automatic image cutting based on adaptive feature extraction and semantic guidance according to claim 4, characterized in that: The cross-scale dilated convolutional pooling branch consists of the parallel first cross-scale dilated convolutional pooling module , the second cross-scale dilated convolution pooling module And the third cross-scale dilated convolution pooling module Composition: The first cross-scale dilated convolution pooling module Receive the first feature map And process to get the first cross-scale feature map ; The second cross-scale dilated convolutional pooling module Receive the second feature map And process to get the second cross-scale feature map ; The third cross-scale hole convolution pooling module Receive the third feature map And process to get the third cross-scale feature map ; The first cross-scale feature map , the second cross-scale feature map And the third cross-scale feature map After splicing in the channel dimension, we get the cross-scale dilated convolutional pooling branch feature map. That is, the output of the cross-scale dilated convolution pooling branch; The first cross-scale dilated convolutional pooling module , the second cross-scale dilated convolution pooling module , the third cross-scale hole convolution pooling module The structures are the same, consisting of a 1×3 convolution layer, a 3×1 convolution layer, a 1×3 dilated convolution layer, a 3×1 dilated convolution layer, a 7×7 average pooling layer, a 11×11 average pooling layer, a 17×17 average pooling layer, a 23×23 average pooling layer, two 1×1 convolution layers and a 3×3 convolution layer; The first cross-scale dilated convolutional pooling module Receive the first feature map And process to get the first cross-scale feature map The specific process is as follows: First, the first feature map After a 1×3 convolutional layer, the first convolution feature map is obtained ; The first feature map After a 1×3 hole convolution layer, the first hole convolution feature map is obtained ; The first feature map After a 3×1 hole convolution layer, the second hole convolution feature map is obtained ; The first feature map After a 3×1 convolution layer, the second convolution feature map is obtained ; The first convolution feature map And the first feature map After splicing in the channel dimension, the first convolutional splicing feature map is obtained ; Splice the first convolution feature map Input to the 7×7 average pooling layer for processing to obtain the first pooling feature map ; The first hole convolution feature map And the first feature map After splicing in the channel dimension, the second convolution splicing feature map is obtained , concatenate the second convolution feature map Input to the 11×11 average pooling layer for processing to obtain the second pooling feature map ; The second hole convolution feature map And the first feature map After splicing in the channel dimension, the third convolution splicing feature map is obtained , concatenate the third convolution feature map Input to the 17×17 average pooling layer for processing to obtain the third pooling feature map ; The second convolution feature map And the first feature map After splicing in the channel dimension, the fourth convolution splicing feature map is obtained , concatenate the fourth convolution feature map Input to the 23×23 average pooling layer for processing to obtain the fourth pooling feature map ; The first pooling feature map And the second pooling feature map After splicing in the channel dimension, the first pooled splicing feature map is obtained , the first pooling splicing feature map After a 1×1 convolutional layer, the first fused feature map is obtained ; The third pooling feature map And the fourth pooling feature map After splicing by channel, we get the second pooled splicing feature map , the second pooling splicing feature map After a 1×1 convolution layer, the second fusion feature map is obtained ; The first feature map , the first fusion feature map and the second fusion feature map After splicing in the channel dimension, the pooled fusion feature map is obtained , the pooled fusion feature map obtained by splicing After a 3×3 convolutional layer, the first cross-scale feature map is obtained .

6. The method for automatic image cutting based on adaptive feature extraction and semantic guidance according to claim 5, characterized in that: The global semantic guidance branch consists of a recursive criss-cross attention module, which is followed by the initial global guidance decoding block. , initial semantic guidance module, first global guidance decoding block , first semantic guidance module, second global guidance decoding block , the second semantic guidance module, the third global guidance decoding block and the third semantic guidance module; first, the fourth feature map is processed by a recursive cross-attention module , get the attention enhancement feature map , the attention-enhanced feature map And cross-scale dilated convolution pooling branch feature map Input initial global boot decoding block The initial global guided decoding module feature map is obtained from , the initial global guided decoding module feature map and attention-enhanced feature maps Input into the initial semantic guidance module to obtain the initial semantic guidance feature map, and input the initial semantic guidance feature map into the first global guidance decoding block The first global guided decoding feature map is obtained from , the first global guided decoding feature map and attention-enhanced feature maps Input into the first semantic guidance module to obtain the first semantic guidance feature map, and input the first semantic guidance feature map into the second global guidance decoding block , get the second global guided decoding feature map , the second global guided decoding feature map and attention-enhanced feature maps Input into the second semantic guidance module to obtain the second semantic guidance feature map, and input the second semantic guidance feature map into the third global guidance decoding block The third global guided decoding feature map is obtained , the third global guided decoding feature map and attention-enhanced feature maps The third semantic guidance feature map is input into the third semantic guidance module, which is the final output of the global semantic guidance branch. The initial semantic guidance module, the first semantic guidance module, the second semantic guidance module and the third semantic guidance module have the same structure, which consists of two 1×1 convolutional layers and one gated convolutional layer; The initial global guided decoding module feature map and attention-enhanced feature maps Input into the initial semantic guidance module to obtain the initial semantic guidance feature map The specific process is as follows: First, the initial global guided decoding module feature map After a 1×1 convolution layer and interpolation upsampling operation, the upsampled feature map is obtained. , adjust the upsampled feature map The number of channels and size of the attention-enhanced feature map Consistent and attention-enhanced feature maps Splice by channel to get the spliced ​​feature map , concatenate the feature maps Attention-enhanced feature map after difference upsampling The fused output is then fed into a gated convolutional layer for fusion. The fused output is then passed through a 1×1 convolutional layer to further adjust the number of channels and obtain the initial semantically guided feature map. , expressed as: ; ; ; In the formula, Indicates difference upsampling; represents a 1×1 convolutional layer; represents a gated convolutional layer.

7. The method for automatic image cutting based on adaptive feature extraction and semantic guidance according to claim 6, characterized in that: Initial Encoding Block It includes a 3×3 convolution layer, a batch normalization layer, a ReLU activation function and a maximum pooling operation layer; first, the object image is processed by the 3×3 convolution layer, the batch normalization layer, and the ReLU activation function in turn to obtain the initial intermediate feature map , then, the initial intermediate feature map The initial feature map is obtained by spatial downsampling through the maximum pooling operation layer ; First coded block , the second coding block , the third coding block and the fourth coding block The structures of are the same, consisting of a maximum pooling layer, a deformable convolutional layer, and a ResNet-34 residual block; The initial feature map Enter the first encoding block , get the first feature map The specific process is: the initial feature map Enter the first encoding block The first feature map is obtained by sequentially processing the maximum pooling layer, deformable convolution layer and ResNet-34 residual block. .

8. The method for automatic image cutting based on adaptive feature extraction and semantic guidance according to claim 7, characterized in that: The first local enhanced self-attention module , the second local enhanced self-attention module And the third local enhanced self-attention module The structure is the same as that of the first local enhanced self-attention module Receive the first feature map And process it to get the first local enhanced feature map The specific process is as follows: First, the first feature map The key matrix is ​​obtained by processing the convolution layer with a convolution kernel size of 3×3. ; The first feature map After processing by a convolution layer with a convolution kernel size of 5×5, the value matrix is ​​obtained ; Perform square sum and mutual dot product processing inside the bond matrix to obtain the affinity matrix ; The affinity matrix Perform normalization operation to obtain the normalized matrix ; Normalize the matrix With value matrix Perform matrix multiplication to obtain weighted feature map ; The weighted feature map With value matrix Concatenate in the channel dimension to get the first local enhanced feature map .

9. The method for automatic image cutting based on adaptive feature extraction and semantic guidance according to claim 8, characterized in that: Initial global boot decoding block , the first global guide decoding block , Second global boot decoding block and the third global boot decoding block The structures are the same, consisting of three 3×3 convolutional layers, three batch normalization layers, three ReLU activation functions, and an interpolation upsampling operation stacked together; Attention-enhanced feature map And cross-scale dilated convolution pooling branch feature map Input initial global boot decoding block The initial global guided decoding module feature map is obtained from The specific process is as follows: First, the attention enhancement feature map And cross-scale dilated convolution pooling branch feature map After splicing in the channel dimension, we get the attention splicing feature map ; Attention splicing feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the first global feature map is obtained. ; The first global feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the second global feature map is obtained. ; The second global feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the third global feature map is obtained. ; The third global feature map After an interpolation upsampling process, the initial global guided decoding module feature map is obtained .

10. The method for automatic image cutting based on adaptive feature extraction and semantic guidance according to claim 9, characterized in that: Initial local understanding decoding block , the first local understanding decoding block , the second local understanding decoding block , the third local understanding decoding block and the fourth local understanding decoding block The structures are the same, consisting of three 3×3 convolutional layers, three batch normalization layers BN, three ReLU activation functions and a maximum pooling layer stacked; The detail enhancement feature map and the fourth feature map Input initial local understanding decoding block Get the initial local understanding decoding feature map The specific process is as follows: First, the detail enhancement feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the first local feature map is obtained. ; The first local feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the second local feature map is obtained. ; The second local feature map After a 3×3 convolution layer, a batch normalization layer, and a ReLU activation function, the third local feature map is obtained. ; The third local feature map After a maximum pooling layer, the initial local understanding decoding feature map is obtained .

Citation Information

Patent Citations

  • Double-branch portrait automatic matting model based on lightweight visual self-attention network

    CN117252892A

  • Automatic animal image matting method

    CN119540252A