Image Processing Method, Apparatus, Device, Storage Medium and Product
By adopting multi-scale feature extraction and classification technology in image processing, the problem of poor image cutting effect in the prior art is solved, and higher-precision image segmentation and cutting effect are achieved.
Patent Information
- Application Number
- CN202411259358.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-09-09
AI Technical Summary
The existing semantic segmentation method for image cutout is not effective in practical applications, especially when dealing with objects of different scales, the segmentation of edges and details is not accurate enough, and it is difficult to accurately distinguish between the cutout target and the background under complex backgrounds.
An image processing method is proposed, which extracts images multi-scale features through an encoder to generate shallow, middle-level and advanced feature maps, and classifies feature points in the advanced feature map through the decoder to construct a target feature map, and finally generates a cutout image through multi-scale feature fusion.
This method can maintain high-precision segmentation effect when processing objects of different scales, retain more details, and more accurately distinguish between the cut-out target and the background in complex backgrounds.
Smart Images

Figure CN119206219B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular, to an image processing method, apparatus, device, storage medium, and product. Background Art
[0002] With the rapid development of computer vision and deep learning technologies, semantic segmentation has become an important research direction in the field of image processing. Semantic segmentation aims to classify each pixel in an image into a specific category, thereby enabling detailed analysis of the image content. Especially in the task of image matting (such as portrait matting), separating a specific object from the background is a key step, which is of great significance in applications such as promotional information production, film and television special effects, virtual reality, and augmented reality.
[0003] However, existing methods for semantic segmentation used in image matting have certain defects in practical applications and the effects are not good:
[0004] (1) Most are based on single-scale feature extraction and cannot fully fuse multi-scale features when processing objects of different scales, resulting in inaccurate segmentation of edges and details when processing matting targets with different scales.
[0005] (2) During the process of processing multi-scale feature extraction, the attention mechanism cannot be effectively used to enhance the attention to important features while suppressing the interference of unimportant features, resulting in deficiencies in the segmentation of detailed information and edge sharpness.
[0006] (3) When dealing with complex backgrounds, it is often difficult to accurately distinguish the matting target from the background. Especially when the background and the matting target have similar colors or textures, this background interference will lead to inaccurate segmentation results and affect the final image matting effect. Summary of the Invention
[0007] The main purpose of this application is to provide an image processing method, apparatus, device, storage medium, and product, aiming to solve the technical problem of poor actual use effect of image matting in practical applications.
[0008] To achieve the above object, this application proposes an image processing method applied to an image matting model, and the image matting model includes an encoder and a decoder;
[0009] The image processing method includes:
[0010] Performing multi-scale feature extraction on the to-be-processed image including the to-be-matted target through the encoder to generate a shallow feature map, a middle feature map, and a high-level feature map;
[0011] Classify the feature points in the high-level feature map through the decoder to obtain a feature classification result;
[0012] Construct a target feature map through the decoder according to the feature classification result and the high-level feature map;
[0013] Perform multi-scale feature fusion on the target feature map, the shallow feature map, and the middle-level feature map through the decoder to generate a cropped image of the target to be cropped.
[0014] Optionally, the encoder includes a first convolutional group, a pooling layer, a second convolutional group, and a third convolutional group;
[0015] The step of performing multi-scale feature extraction on the image to be processed containing the target to be cropped through the encoder to generate a shallow feature map, a middle-level feature map, and a high-level feature map includes:
[0016] Perform feature extraction on the image to be processed containing the target to be cropped through the first convolutional group to generate a shallow feature map;
[0017] Perform secondary feature extraction on the shallow feature map through the pooling layer and the second convolutional group to generate a middle-level feature map;
[0018] Perform depth feature extraction on the middle-level feature map through the third convolutional group to generate a high-level feature map;
[0019] Wherein, the size of the middle-level feature map is smaller than the size of the shallow feature map, and the size of the high-level feature map is smaller than the size of the middle-level feature map.
[0020] Optionally, the encoder further includes a parallel attention module and a multi-scale extraction module;
[0021] The step of performing depth feature extraction on the middle-level feature map through the third convolutional group to generate a high-level feature map includes:
[0022] Perform depth feature extraction on the middle-level feature map through the third convolutional group to generate a depth feature map;
[0023] Set weights for the depth feature map through the parallel attention module to generate a weighted feature map, and perform parallel sampling extraction on the depth feature map through the multi-scale extraction module to generate a parallel sampling feature map;
[0024] Fuse the weighted feature map and the parallel sampling feature map to generate a high-level feature map.
[0025] Optionally, the parallel attention module includes a position attention module and a channel attention module;
[0026] The step of setting weights on the depth feature map by a parallel attention module to generate a weight feature map includes:
[0027] The depth feature map is weighted by the position attention module to generate a first weighted feature map, and the depth feature map is weighted by the channel attention module to generate a second weighted feature map;
[0028] The first weight setting feature map and the second weight setting feature map are fused to generate a weight feature map.
[0029] Optionally, the multi-scale extraction module includes a point convolution kernel, a plurality of dilation degrees of hole convolutions and a global pooling layer;
[0030] The performing parallel sampling extraction on the depth feature map by the multi-scale extraction module to generate a parallel sampling feature map includes:
[0031] Sampling and extracting the depth feature map through the point convolution kernel to obtain a first sampling feature map;
[0032] Sampling and extracting the depth feature map by using the plurality of dilated convolutions with different dilation degrees respectively to obtain a plurality of second sampling feature maps;
[0033] Performing global pooling processing on the depth feature map through the global pooling layer to generate a third sampling feature map;
[0034] The first sampling feature map, the multiple second sampling feature maps and the third sampling feature map are fused to generate a parallel sampling feature map.
[0035] Optionally, performing multi-scale feature fusion on the target feature map, the shallow feature map and the middle feature map through the decoder to generate a cutout image of the target to be cutout includes:
[0036] Performing upsampling processing on the target feature map based on a first upsampling ratio by the decoder to generate an upsampled target feature map;
[0037] The decoder fuses the expanded target feature map with the middle-level feature map to generate a first spliced feature map;
[0038] Performing upsampling processing on the first splicing feature map based on a second upsampling ratio by the decoder to generate an upsampled splicing feature map;
[0039] The decoder fuses the expanded sample splicing feature map with the shallow feature map to generate a second splicing feature map;
[0040] The decoder performs upsampling processing on the second stitched feature map based on the third upsampling ratio to generate a matte image of the target to be extracted.
[0041] Optionally, before the encoder performs multi-scale feature extraction on the image to be processed containing the target to be extracted to generate a shallow feature map, a middle-level feature map, and a high-level feature map, it further includes:
[0042] Training the initial matte model with a model training set;
[0043] When the initial matte model meets the preset convergence condition, verifying the initial matte model with a model validation set to generate a model evaluation index value;
[0044] If the model evaluation index value meets the preset evaluation condition, using the initial matte model as the image matte model.
[0045] Optionally, the initial matte model includes an encoder and a decoder;
[0046] After obtaining the model evaluation index value corresponding to the initial matte model when the initial matte model meets the preset convergence condition, it further includes:
[0047] When the initial matte model does not meet the preset convergence condition, obtaining a first loss value of the encoder in the initial matte model and a second loss value of the decoder;
[0048] Constructing an overall model loss according to a preset loss weight, the first loss value, and the second loss value;
[0049] Adjusting the parameters of the initial matte model according to the overall model loss and returning to the step of training the initial matte model with the model training set.
[0050] Optionally, before training the initial matte model with the model training set, it further includes:
[0051] Obtaining a sample image set, where the sample image set contains sample images of multiple different scenes, and the sample images include scene images and target images;
[0052] Performing pixel-level annotation on each sample image in the sample image set to obtain an annotated image set;
[0053] Dividing the annotated image set to generate a model training set and a model validation set.
[0054] Optionally, before dividing the annotated image set to generate a model training set and a model validation set, it further includes:
[0055] Perform size unification processing on each image in the labeled image set to generate a processed labeled image set;
[0056] Partition the processed labeled image set to generate a model training set and a model validation set.
[0057] Optionally, before partitioning the labeled image set to generate a model training set and a model validation set, it further includes:
[0058] Perform data augmentation processing on each image in the labeled image set to obtain an augmented labeled image set, where the data augmentation includes at least one of rotation, flipping, cropping, and brightness adjustment;
[0059] Partition the augmented labeled image set to generate a model training set and a model validation set.
[0060] In addition, to achieve the above object, the present application also proposes an image processing device applied to an image matting model, and the image matting model includes an encoder and a decoder;
[0061] The image processing device includes:
[0062] An extraction module for performing multi-scale feature extraction on a to-be-processed image containing a to-be-matted target through the encoder to generate a shallow feature map, a middle feature map, and a high-level feature map;
[0063] A classification module for classifying feature points in the high-level feature map through the decoder to obtain a feature classification result;
[0064] A construction module for constructing a target feature map through the decoder according to the feature classification result and the high-level feature map;
[0065] A generation module for performing multi-scale feature fusion on the target feature map, the shallow feature map, and the middle feature map through the decoder to generate a matting image of the to-be-matted target.
[0066] Optionally, the encoder includes a first convolutional group, a pooling layer, a second convolutional group, and a third convolutional group;
[0067] The extraction module is further configured to perform feature extraction on a to-be-processed image containing a to-be-matted target through the first convolutional group to generate a shallow feature map; perform secondary feature extraction on the shallow feature map through the pooling layer and the second convolutional group to generate a middle feature map; perform depth feature extraction on the middle feature map through the third convolutional group to generate a high-level feature map; where the size of the middle feature map is smaller than the size of the shallow feature map, and the size of the high-level feature map is smaller than the size of the middle feature map.
[0068] Optionally, the encoder further includes a parallel attention module and a multi-scale extraction module;
[0069] The extraction module is further configured to perform deep feature extraction on the middle-layer feature map through a third convolutional group to generate a deep feature map; perform weight setting on the deep feature map through the parallel attention module to generate a weighted feature map, perform parallel sampling extraction on the deep feature map through the multi-scale extraction module to generate a parallel sampling feature map; fuse the weighted feature map and the parallel sampling feature map to generate a high-level feature map.
[0070] Optionally, the parallel attention module includes a position attention module and a channel attention module;
[0071] The extraction module is further configured to perform weight setting on the deep feature map through the position attention module to generate a first weight-setting feature map, and perform weight setting on the deep feature map through the channel attention module to generate a second weight-setting feature map; fuse the first weight-setting feature map and the second weight-setting feature map to generate a weighted feature map.
[0072] Optionally, the multi-scale extraction module includes a point convolution kernel, multiple dilated convolutions with different dilation rates, and a global pooling layer;
[0073] The extraction module is further configured to perform sampling extraction on the deep feature map through the point convolution kernel to obtain a first sampling feature map; perform sampling extraction on the deep feature map through the multiple dilated convolutions with different dilation rates respectively to obtain multiple second sampling feature maps; perform global pooling processing on the deep feature map through the global pooling layer to generate a third sampling feature map; fuse the first sampling feature map, the multiple second sampling feature maps, and the third sampling feature map to generate a parallel sampling feature map.
[0074] Optionally, the generation module is further configured to perform upsampling on the target feature map by the decoder based on a first upsampling ratio to generate an upsampled target feature map; fuse the upsampled target feature map and the middle-layer feature map by the decoder to generate a first concatenated feature map; perform upsampling on the first concatenated feature map by the decoder based on a second upsampling ratio to generate an upsampled concatenated feature map; fuse the upsampled concatenated feature map and the shallow-layer feature map by the decoder to generate a second concatenated feature map; perform upsampling on the second concatenated feature map by the decoder based on a third upsampling ratio to generate a cropped image of the target to be cropped.
[0075] In addition, to achieve the above object, the present application further provides an image processing device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the image processing method as described above.
[0076] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the image processing method as described above.
[0077] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the image processing method as described above.
[0078] One or more technical solutions provided by the present application have at least the following technical effects:
[0079] Due to the adoption of multi-scale feature extraction, feature maps of different scales are generated, and the target feature maps obtained after classification based on the high-level feature maps are fused with the multi-scale feature maps to generate a matte image, which ensures that more details can be retained, so that high-precision segmentation effects can be maintained when dealing with large ranges and small areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0081] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0082] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the image processing method of the present application;
[0083] Figure 2 It is a schematic flowchart provided for Embodiment 2 of the image processing method of the present application;
[0084] Figure 3 It is a schematic diagram of the overall architecture of the image matte model according to an embodiment of the present application;
[0085] Figure 4 It is a schematic flowchart provided for Embodiment 3 of the image processing method of the present application;
[0086] Figure 5 Schematic diagram of the module structure of the image processing device according to an embodiment of the present application;
[0087] Figure 6 Schematic diagram of the device structure of the hardware operating environment involved in the image processing method according to an embodiment of the present application.
[0088] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0089] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0090] In order to better understand the technical solutions of the present application, the following will be described in detail with reference to the accompanying drawings of the specification and specific embodiments.
[0091] Based on this, an embodiment of the present application provides an image processing method, which is applied to an image matting model. The image matting model includes an encoder and a decoder. Refer to Figure 1 , Figure 1 Schematic flowchart of the first embodiment of the image processing method of the present application.
[0092] In this embodiment, the image processing method includes steps S10 to S40:
[0093] Step S10: Perform multi-scale feature extraction on the to-be-processed image containing the to-be-matted target through the encoder to generate a shallow feature map, a middle feature map, and a high-level feature map.
[0094] It should be noted that the execution subject of this embodiment can be the image matting model or an image processing device deployed with the image matting model. The image processing device can be an electronic device such as a personal computer or a server, or other devices that can achieve the same or similar functions. This embodiment does not limit this. In this embodiment and the following embodiments, the image processing method of the present invention will be described by taking the image processing device as an example.
[0095] It should be noted that the to-be-matted target can be the target of the corresponding image that needs to be extracted from the to-be-processed image through image matting processing. The shallow feature map can contain features such as relatively simple and basic visual patterns, such as edges, colors, textures, etc.; the middle feature map can contain information that is more abstract than the shallow feature map, such as features of simple shapes, parts of objects, etc.; the high-level feature map can contain highly abstract information, such as features of complete objects, scenes, etc.
[0096] In actual use, a feature extraction network with a relatively deep depth can be set in the encoder. Through this feature extraction network, multi-scale feature extraction is performed on the image to be processed, generating shallow feature maps, middle-level feature maps, and high-level feature maps. For example, a relatively deep ResNet-101 network is adopted, with a residual structure. Through the convolutional layer and pooling layer in this network, deep convolutional processing is performed to achieve feature extraction.
[0097] Step S20: Classify the feature points in the high-level feature map through the decoder to obtain a feature classification result.
[0098] It should be noted that the high-level feature map can contain multiple feature points, and each feature point corresponds to one or more pixels in the image to be processed. When the encoder generates the high-level feature map, a corresponding classification weight value will be set for each feature point in the high-level feature map. Among them, the classification weight value is used to characterize the relevance between the feature point and the target to be extracted. The higher the relevance between the feature point and the target to be extracted, the higher the corresponding classification weight value.
[0099] In actual use, the decoder can read the classification weight values of each feature point in the high-level feature map, divide the feature points in the high-level feature map whose corresponding classification weight values are greater than the preset weight threshold into partial feature points related to the target to be extracted, and divide the other feature points into partial feature points unrelated to the target to be extracted, and generate a corresponding feature classification result.
[0100] Step S30: Construct a target feature map through the decoder according to the feature classification result and the high-level feature map.
[0101] In actual use, the decoder can extract the partial feature points related to the target to be extracted in the high-level feature map according to the classification result of the previous classification to construct a target feature map, or add corresponding classification labels to each feature point in the high-level feature map, and use the feature map with labels added as the target feature map.
[0102] Step S40: Perform multi-scale feature fusion on the target feature map, the shallow feature map, and the middle-level feature map through the decoder to generate a cropped image of the target to be extracted.
[0103] It should be noted that when performing multi-scale feature fusion on the target feature map, the shallow feature map, and the middle-level feature map, the feature information of the feature points corresponding to the same pixel in each feature map can be fused by means such as summation, weighted summation, or weighted averaging, so as to generate an image with the same size specification as the image to be processed. According to the classification labels of the feature points corresponding to each pixel in this image, the pixel points unrelated to the target to be extracted in the generated image are determined. Then, the color values of the pixel points unrelated to the target to be extracted in the generated image are set to the preset color value, so as to generate a cropped image of the target to be extracted.
[0104] Among them, the preset color value can be set in advance by the administrator of the image processing device. For example, when performing image matting, if the color of the background part is set to black, the color value corresponding to black can be set as the preset color value.
[0105] In actual use, due to the different processing depths of the target feature map, the shallow feature map, and the middle feature map, their size specifications may vary. When performing multi-scale fusion, upsampling processing can also be combined to adjust the size. After adjusting the sizes to be the same, fusion is then carried out.
[0106] This embodiment provides an image processing method. Due to the use of multi-scale feature extraction, feature maps of different scales are generated, and based on the multi-scale feature map fusion of the target feature map obtained after classification based on the high-level feature map, a matting image is generated, ensuring that more details can be retained, so that high-precision segmentation effects can be maintained when processing large ranges and small areas.
[0107] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be elaborated hereinafter. On this basis, please refer to Figure 2 , the encoder may include a first convolutional group, a pooling layer, a second convolutional group, and a third convolutional group;
[0108] Step S10 includes steps S101 to S103:
[0109] Step S101: Perform feature extraction on the image to be processed containing the target to be extracted through the first convolutional group to generate a shallow feature map.
[0110] It should be noted that the first convolutional group can be a convolutional group for preliminary feature extraction of the image to be processed, which is set in advance by the administrator of the image processing device. The first convolutional group may include at least one convolutional network. The size of the shallow feature map can be smaller than the size of the image to be processed. For example, the size of the shallow feature map can be 1 / 2 of the size of the image to be processed.
[0111] Step S102: Perform secondary feature extraction on the shallow feature map through the pooling layer and the second convolutional group to generate a middle feature map.
[0112] It should be noted that the pooling layer can perform any one of pooling operations such as max pooling, average pooling, global pooling, and adaptive pooling. This embodiment does not limit this. The second convolutional group may include at least one convolutional network.
[0113] In actual use, the shallow feature map can be first pooled through the pooling layer, and then the shallow feature map after pooling processing can be used by the second convolutional group for feature extraction to achieve secondary feature extraction and generate the middle-level feature map.
[0114] Of course, it is also possible to first perform feature extraction through the second convolutional group, and then pool the feature map generated by the feature extraction through the pooling layer, and use the feature map after pooling processing as the middle-level feature map. This embodiment does not limit this.
[0115] Among them, the size of the middle-level feature map can be smaller than the size of the shallow feature map. For example, the size of the shallow feature map can be 1 / 2 of the size of the image to be processed, while the size of the middle-level feature map can be 1 / 4 of the size of the image to be processed.
[0116] Step S103: Perform deep feature extraction on the middle-level feature map through the third convolutional group to generate a high-level feature map.
[0117] In actual use, the second convolutional group can include at least one convolutional network. The size of the high-level feature map can be smaller than the size of the middle-level feature map. For example, if the size of the shallow feature map is 1 / 2 of the size of the image to be processed and the size of the middle-level feature map is 1 / 4 of the size of the image to be processed, then the size of the high-level feature map can be 1 / 16 of the size of the image to be processed at this time.
[0118] In specific implementation, in order to ensure the actual use effect of the high-level feature map, an attention mechanism and multi-scale extraction can also be additionally added. At this time, the encoder in this embodiment can also include a parallel attention module and a multi-scale extraction module. Step S103 in this embodiment can include:
[0119] Perform deep feature extraction on the middle-level feature map through the third convolutional group to generate a deep feature map;
[0120] Set the weights of the deep feature map through the parallel attention module to generate a weighted feature map, and perform parallel sampling extraction on the deep feature map through the multi-scale extraction module to generate a parallel sampling feature map;
[0121] Fuse the weighted feature map and the parallel sampling feature map to generate a high-level feature map.
[0122] It should be noted that additional processing can be introduced. At this time, after the middle-level feature map is subjected to depth feature extraction through the third convolutional group, the obtained feature map can be used as the depth feature map. Then, the parallel attention module is used to set weights for the depth feature map to generate a weighted feature map, and at the same time, the multi-scale extraction module is used to perform parallel sampling extraction on the depth feature map to generate a parallel sampling feature map. Finally, the weighted feature map and the parallel sampling feature map are fused to generate a high-level feature map.
[0123] Among them, when fusing the weighted feature map and the parallel sampling feature map, it can be to sum the features of the corresponding feature points in the weighted feature map and the parallel sampling feature map. This summation can be direct summation or weighted summation. The corresponding feature points can be the feature points with the same corresponding pixels.
[0124] It should be noted that the parallel attention module can include various attention mechanisms for parallel processing. Among them, the attention mechanism can include a position attention mechanism (or spatial attention mechanism), a self-attention mechanism, a multi-head attention mechanism, a channel attention mechanism, etc. Of course, there can be more, and this embodiment does not limit this.
[0125] In a specific implementation, the parallel attention module can include a position attention module (Position Attention Module, PAM) and a channel attention module (Channel Attention Module, CAM). Then, the step of setting weights for the depth feature map through the parallel attention module to generate a weighted feature map in this embodiment can include:
[0126] Setting weights for the depth feature map through the position attention module to generate a first weighted setting feature map, and setting weights for the depth feature map through the channel attention module to generate a second weighted setting feature map;
[0127] Fusing the first weighted setting feature map and the second weighted setting feature map to generate a weighted feature map.
[0128] It should be noted that the position attention module can set corresponding classification weights for each feature point in the depth feature map based on the position attention mechanism. At this time, the depth feature map with classification weights can be used as the first feature map; the channel attention module can set corresponding classification weights for each feature point in the depth feature map based on the channel attention mechanism. At this time, the depth feature map with classification weights can be used as the second feature map.
[0129] In actual use, the classification weights of the matching feature points in the first weight setting feature map and the second weight setting feature map can be summed to achieve feature map fusion, thereby generating a weight feature map. Among them, the summation can be a direct summation or a weighted summation, and the matching feature points can be feature points with the same corresponding pixels.
[0130] It should be noted that by combining the position attention mechanism and the channel attention mechanism to form a dual attention mechanism, problems such as slow fitting speed, inaccurate edge target segmentation, inconsistent large-scale target segmentation, and omissions during image matting can be solved, helping the entire network obtain better robustness, better retaining spatial details, and making feature extraction richer.
[0131] In a specific implementation, the multi-scale extraction module may include a point convolution kernel, multiple dilated convolutions with different dilation rates, and a global pooling layer. The step of generating a parallel sampling feature map by performing parallel sampling extraction on the depth feature map through the multi-scale extraction module in this embodiment may include:
[0132] Sampling and extracting the depth feature map through the point convolution kernel to obtain a first sampling feature map;
[0133] Sampling and extracting the depth feature map through the multiple dilated convolutions with different dilation rates respectively to obtain multiple second sampling feature maps;
[0134] Performing global pooling processing on the depth feature map through the global pooling layer to generate a third sampling feature map;
[0135] Fusing the first sampling feature map, the multiple second sampling feature maps, and the third sampling feature map to generate a parallel sampling feature map.
[0136] It should be noted that the point convolution kernel can be a 1×1 specification convolution kernel; the dilated convolution can be a 3×3 depthwise separable dilated convolution for expanding the receptive field; the global pooling layer can be a network layer for performing global pooling processing.
[0137] Among them, the dilation rates of the respective dilated convolutions can be different. For example, assuming there are 3 dilated convolutions, different dilation rates (such as 6, 12, 18) can be set for the three dilated convolutions at this time.
[0138] In actual use, the generation of the first sampled feature map, the second sampled feature map, and the third sampled feature map can be processed in parallel, that is, the first sampled feature map, the second sampled feature map, and the third sampled feature map are generated in parallel by means of multi-threading or similar methods. The implementation of fusing the first sampled feature map, multiple second sampled feature maps, and the third sampled feature map can also be to sum the feature information of the matching feature points in each feature map, so as to achieve feature map fusion. Among them, the summation can be direct summation or weighted summation, and the matching feature points can be the feature points with the same corresponding pixels.
[0139] It should be noted that by using a point convolution kernel, multiple dilated convolutions with different dilation rates, and a global pooling layer to perform parallel sampling on the depth feature map, various context information at different scales can be captured, with less information loss, thereby ensuring the improvement of the image quality of the generated matte image.
[0140] In this embodiment, step S40 may include:
[0141] The decoder performs upsampling on the target feature map based on the first upsampling ratio to generate an upsampled target feature map;
[0142] The decoder fuses the upsampled target feature map with the middle layer feature map to generate a first spliced feature map;
[0143] The decoder performs upsampling on the first spliced feature map based on the second upsampling ratio to generate an upsampled spliced feature map;
[0144] The decoder fuses the upsampled spliced feature map with the shallow layer feature map to generate a second spliced feature map;
[0145] The decoder performs upsampling on the second spliced feature map based on the third upsampling ratio to generate a matte image of the target to be extracted.
[0146] It should be noted that the size of the target feature map is actually the same as that of the high-level feature Figure 1 To ensure that the target feature map can be fused with the middle layer feature map, the target feature map can be first upsampled based on the first upsampling ratio to generate an upsampled target feature map with the same size as the middle layer feature map, and then fused. Among them, the fusion of the upsampled target feature map and the middle layer feature map can be to use a splicing function (Add()) to superimpose the same pixels in the upsampled target feature map and the middle layer feature map, so as to increase the amount of information. The upsampling algorithm used for upsampling the target feature map can be bilinear interpolation upsampling.
[0147] In actual use, the upsampling algorithms used for upsampling the first spliced feature map and the second spliced feature map can be the same as the algorithm used for upsampling the target feature map, which will not be elaborated here. Similarly, when fusing the upsampled spliced feature map with the shallow feature map, the method used is the same as that used when fusing the upsampled target feature map with the middle feature map, which will not be elaborated here.
[0148] In specific applications, the first upsampling ratio, the second upsampling ratio, and the third upsampling ratio can be set according to the size differences of the image to be processed, the shallow feature map, the middle feature map, and the high-level feature map. For example: the size of the shallow feature map can be 1 / 2 of the size of the image to be processed, and the size of the middle feature map can be 1 / 4 of the size of the image to be processed. Then, at this time, the size of the high-level feature map can be 1 / 16 of the size of the image to be processed. Then, at this time, the first upsampling ratio can be 4 times, the second upsampling ratio can be 2 times, and the third upsampling ratio can be 2 times.
[0149] For the sake of easy understanding, it is now combined with Figure 3 for illustration, but it does not limit this solution. Figure 3 This is the overall architecture schematic diagram of the image matting model in this embodiment.
[0150] As Figure 3 shown, the image matting model is divided into two major modules: an encoder and a decoder. The encoder includes ResNet-101, a dual attention mechanism module (i.e., a parallel attention module), and DS-ASPP (i.e., a multi-scale extraction module). ResNet-101 includes multiple convolutional groups layer0 - layer4 and a pooling layer (Maxpool). Among them, layer0 is the first convolutional group described above, layer1 is the second convolutional group described above, and the combination of layer2 - layer4 constitutes the third convolutional group described above;
[0151] When the decoder obtains the high-level feature map processed by the encoder, it will perform 4-fold upsampling on it, and then splice it with the middle feature map output by layer1 to generate the first spliced feature map. After that, it will perform 2-fold upsampling on the first spliced feature map and splice the upsampled result with the shallow feature map to generate the second spliced feature map. Finally, it will perform 2-fold upsampling on the second spliced feature map to generate the matting image of the target to be extracted.
[0152] This embodiment provides an image processing method. Since a network with a relatively deep depth is used for feature extraction during feature extraction, it is ensured that feature extraction at different scales can be achieved through a network with a relatively deep depth, which helps to significantly improve the segmentation accuracy and the retention effect of edge details.
[0153] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 4 , before the step S10 described in this embodiment, the following steps may further be included:
[0154] Step S01: Train the initial extraction model with a model training set.
[0155] It should be noted that the model training set may include multiple model training samples. The model training samples may be sample images with classification annotations. The classification annotations may be pixel-level annotations, that is, corresponding classification annotations are set for each pixel in the sample image. The classification annotations are used to characterize the relevance between the pixel and the target to be extracted in the sample image. For example, if the classification annotation is 1, it means it is relevant to the target to be extracted, and if the classification annotation is 0, it means it is not relevant to the target to be extracted.
[0156] In actual use, the initial extraction model can be trained with a model training set. Among them, the number of iteration rounds and the learning rate during the training process can be set in advance by the management personnel of the image processing device.
[0157] Step S02: When the initial extraction model meets the preset convergence condition, verify the initial extraction model with a model validation set to generate a model evaluation index value.
[0158] It should be noted that the preset convergence condition can be set in advance by the management personnel of the image processing device. For example, the preset convergence condition can be set as the training iteration rounds reaching the preset iteration rounds, or the preset convergence condition can be set as that the model loss value does not decrease for multiple consecutive rounds.
[0159] In actual use, if the initial extraction model meets the preset convergence condition, it is necessary to verify whether the trained model meets the actual use standard. Therefore, the initial extraction model can be verified with a model validation set, and a model evaluation index value is generated during the verification process.
[0160] Among them, the model evaluation index can be an index value calculated according to a preset evaluation index set in advance. The preset evaluation index may include at least one of accuracy, precision, recall, intersection over union (IoU), F1 score (F1), and Bayesian error rate (BER).
[0161] In actual use, after validating the initial extraction model through the model validation set, according to the difference between the model prediction result and the true result, true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN) can be obtained. The calculation formulas for the index values corresponding to various preset evaluation metrics are as follows:
[0162]
[0163] Step S03: If the model evaluation metric value meets the preset evaluation conditions, then use the initial extraction model as the image extraction model.
[0164] In actual use, if the model evaluation metric value meets the preset evaluation conditions, it can be determined that the initial extraction model trained to convergence can be put into actual use. Therefore, the initial extraction model can be used as the image extraction model.
[0165] In actual application, the preset evaluation conditions can be set in advance by the management personnel of the image processing device. For example, the preset evaluation conditions can be set as the model evaluation metric value being greater than the corresponding index passing threshold. Among them, if multiple preset evaluation metrics are adopted, the corresponding index passing thresholds can be set for each preset evaluation metric. Only when the model evaluation metric values corresponding to each preset evaluation metric are all greater than the corresponding index passing thresholds, it is determined that the preset evaluation conditions are met.
[0166] In specific implementation, to ensure that the parameters in the initial extraction model can be reasonably adjusted, after step S02 of this embodiment, the following is further included:
[0167] When the initial extraction model does not meet the preset convergence conditions, obtain the first loss value of the encoder in the initial extraction model and the second loss value of the decoder;
[0168] Construct the overall model loss according to the preset loss weight, the first loss value, and the second loss value;
[0169] Adjust the parameters of the initial extraction model according to the overall model loss, and return to the step of training the initial extraction model through the model training set.
[0170] It should be noted that the loss values corresponding to the encoder and the decoder can be calculated separately. Constructing the overall model loss according to the preset loss weight, the first loss value, and the second loss value can be to perform weighted summation on the first loss value and the second loss value according to the preset loss weight, and use the weighted sum value as the overall model loss.
[0171] For example: Suppose the first loss value of the encoder is L encoder , and the second loss value of the decoder is L decoder, where α1 is the preset loss weight corresponding to the decoder and α2 is the preset loss weight corresponding to the encoder. Then the overall loss Loss of the model at this time MSFF-AENet = α1 * L decoder + α2 * L encoder .
[0172] Among them, the first loss value of the encoder can be calculated based on the cross-entropy loss function, calculated based on the mean-square loss function, or calculated by combining the cross-entropy loss function and the mean-square loss function together (such as the sum of the loss value calculated by the cross-entropy loss function and the loss value calculated by the mean-square loss function as the first loss value of the encoder); the calculation method of the second loss value of the decoder can be similar to that of the first loss value of the encoder, which will not be elaborated here.
[0173] In actual use, since image matting is essentially a binary classification task, when using the cross-entropy loss function, it can be the binary cross-entropy loss function. Then the formula of the cross-entropy loss function can be:
[0174]
[0175] σ(x) = 1 / (1 + exp(-x))
[0176] In the formula, Loss BCE can be the cross-entropy loss value, y i represents the label of sample i, where the positive example is 1 and the negative example is 0; the number of pixels is N, and p i represents the probability that sample i is predicted as the positive class. σ(x) is the Sigmoid activation function, and its role is to make the predicted value between 0 and 1, and x is the predicted classification probability.
[0177] In specific implementation, in order to construct the model training set and the model validation set, before step S01 described in this embodiment, it may further include:
[0178] Obtain a sample image set;
[0179] Perform pixel-level annotation on each sample image in the sample image set to obtain an annotated image set;
[0180] Divide the annotated image set to generate a model training set and a model validation set.
[0181] It should be noted that the sample image set contains sample images of multiple different scenes. The sample images include scene images and target images. Among them, the scene image is an image related to the scene that has nothing to do with the target to be extracted, and the target image is an image related to the target to be extracted.
[0182] In actual use, multiple images containing the target to be extracted can be retrieved from an image library as sample images, and the retrieved sample images can be aggregated into a sample image set. Here, the image library can be a database for storing images.
[0183] In actual use, pixel-level annotation is performed on each sample image in the sample image set to obtain an annotated image set. This can be achieved by manually annotating each sample image to obtain the annotated image set, or by using a pre-trained large model or annotation algorithm to perform pixel-level annotation on the sample images to obtain the annotated image set. Here, pixel-level annotation can involve classifying each pixel in the sample image separately.
[0184] In actual use, the annotated image set is divided to generate a model training set and a model validation set. This can be done by dividing the annotated image set proportionally to generate the model training set and the model validation set. For example, the annotated image set can be divided into a model training set and a model validation set at a ratio of 7:3.
[0185] In a specific implementation, a model test set can also be set up to further test the model after it has been verified by the model validation set to determine its specific model performance. In this case, the annotated image set can be divided proportionally to generate a model training set, a model validation set, and a model test set.
[0186] For example: The annotated image set is divided into a training set, a validation set, and a test set at a ratio of 8:1:1.
[0187] In actual use, to ensure the model training effect, it is necessary to ensure that the sizes of the images used during training are consistent. Before the step of dividing the annotated image set to generate a model training set and a model validation set described in this embodiment, the following steps can also be included:
[0188] Perform size unification processing on each image in the annotated image set to generate a processed annotated image set;
[0189] Divide the processed annotated image set to generate a model training set and a model validation set.
[0190] It should be noted that size unification processing can stretch the sizes of the images in the annotated image set to be consistent through methods such as image stretching. Existing software tools can be used to achieve image stretching. For example: Image stretching can be performed using OpenCV or Pillow to achieve image scaling.
[0191] In actual use, to ensure the model generalization ability, before the step of dividing the annotated image set to generate a model training set and a model validation set described in this embodiment, the following steps can also be included
[0192] Perform data augmentation on each image in the labeled image set to obtain an augmented labeled image set;
[0193] Partition the augmented labeled image set to generate a model training set and a model validation set.
[0194] It should be noted that data augmentation may include at least one of rotation, flipping, cropping, and brightness adjustment. When performing data augmentation on an image, data augmentation operations can be synchronously performed on the image and the mask corresponding to the image.
[0195] It can be understood that through data augmentation, the scenes or environments involved in the images can be made more diverse, so that the trained model can handle more scenarios.
[0196] It should be noted that size unification processing and data augmentation processing are combined. For example: first perform size unification processing on each image in the labeled image set, then further perform data augmentation processing on each image in the processed labeled image set, and then partition the augmented labeled image set.
[0197] This embodiment provides an image processing method. Since the model has been pre-trained, verified, and tested, it ensures that the subsequent image matting model put into use has a good matting effect.
[0198] This application also provides an image processing device. Please refer to Figure 5 , the image processing device includes:
[0199] An extraction module 10, configured to perform multi-scale feature extraction on a to-be-processed image including a to-be-extracted target through the encoder to generate a shallow feature map, a middle feature map, and a high-level feature map;
[0200] A classification module 20, configured to classify feature points in the high-level feature map through the decoder to obtain a feature classification result;
[0201] A construction module 30, configured to construct a target feature map through the decoder according to the feature classification result and the high-level feature map;
[0202] A generation module 40, configured to perform multi-scale feature fusion on the target feature map, the shallow feature map, and the middle feature map through the decoder to generate a matting image of the to-be-extracted target.
[0203] The image processing device provided by this application adopts the image processing method in the above embodiment, and can solve the technical problem that the actual use effect of image matting in practical applications is not good. Compared with the prior art, the beneficial effects of the image processing device provided by this application are the same as those of the image processing method provided by the above embodiment, and other technical features in the image processing device are the same as the features disclosed in the method of the above embodiment, which will not be elaborated here.
[0204] This application provides an image processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image processing method in Embodiment 1 above.
[0205] Refer to the following Figure 6 , which shows a schematic structural diagram of an image processing device suitable for implementing the embodiments of this application. The image processing device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The image processing device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0206] As shown in Figure 6As shown, the image processing device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the image processing device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the image processing device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an image processing device having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or had alternatively.
[0207] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0208] The image processing device provided by the present application adopts the image processing method in the above-mentioned embodiment, and can solve the technical problem that the actual use effect of image matting in actual applications is not good. Compared with the prior art, the beneficial effects of the image processing device provided by the present application are the same as those of the image processing method provided by the above-mentioned embodiment, and other technical features in the image processing device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0209] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0210] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0211] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the image processing method in the above embodiments.
[0212] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0213] The above computer-readable storage medium can be included in the image processing device; or it can exist separately and not be assembled into the image processing device.
[0214] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the image processing device, the image processing device is caused to:
[0215] The encoder performs multi-scale feature extraction on the to-be-processed image containing the target to be cropped, generating a shallow feature map, a middle-level feature map, and a high-level feature map; the decoder classifies the feature points in the high-level feature map to obtain a feature classification result; the decoder constructs a target feature map according to the feature classification result and the high-level feature map; the decoder performs multi-scale feature fusion on the target feature map, the shallow feature map, and the middle-level feature map to generate a cropped image of the target to be cropped.
[0216] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0217] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0218] The modules described in the embodiments of the present application may be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0219] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above image processing method, which can solve the technical problem that the actual use effect of image matting in practical applications is not good. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the image processing method provided by the above embodiments, and will not be elaborated here.
[0220] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the image processing method as described above.
[0221] The computer program product provided by this application can solve the technical problem that the actual use effect of image matting in practical applications is not good. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as those of the image processing method provided by the above embodiments, and will not be elaborated here.
[0222] The above are only partial embodiments of this application, and do not limit the patent scope of this application. Any equivalent structural transformation made under the technical concept of this application by using the content of the specification and drawings of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.
[0223] This application discloses A1. An image processing method is applied to an image matting model, and the image matting model includes an encoder and a decoder;
[0224] The image processing method includes:
[0225] Performing multi-scale feature extraction on the to-be-processed image including the to-be-matted target through the encoder to generate a shallow feature map, a middle feature map, and a high-level feature map;
[0226] Classifying the feature points in the high-level feature map through the decoder to obtain a feature classification result;
[0227] Constructing a target feature map through the decoder according to the feature classification result and the high-level feature map;
[0228] Performing multi-scale feature fusion on the target feature map, the shallow feature map, and the middle feature map through the decoder to generate a matting image of the to-be-matted target.
[0229] A2. The image processing method as described in A1, where the encoder includes a first convolutional group, a pooling layer, a second convolutional group, and a third convolutional group;
[0230] The step of performing multi-scale feature extraction on the image to be processed containing the target to be cropped through the encoder to generate a shallow feature map, a middle-level feature map, and a high-level feature map includes:
[0231] Performing feature extraction on the image to be processed containing the target to be cropped through the first convolutional group to generate a shallow feature map;
[0232] Performing secondary feature extraction on the shallow feature map through a pooling layer and a second convolutional group to generate a middle-level feature map;
[0233] Performing depth feature extraction on the middle-level feature map through a third convolutional group to generate a high-level feature map;
[0234] Among them, the size of the middle-level feature map is smaller than the size of the shallow feature map, and the size of the high-level feature map is smaller than the size of the middle-level feature map.
[0235] A3. The image processing method as described in A2, wherein the encoder further includes a parallel attention module and a multi-scale extraction module;
[0236] The step of performing depth feature extraction on the middle-level feature map through a third convolutional group to generate a high-level feature map includes:
[0237] Performing depth feature extraction on the middle-level feature map through a third convolutional group to generate a depth feature map;
[0238] Performing weight setting on the depth feature map through a parallel attention module to generate a weight feature map, and performing parallel sampling extraction on the depth feature map through a multi-scale extraction module to generate a parallel sampling feature map;
[0239] Fusing the weight feature map and the parallel sampling feature map to generate a high-level feature map.
[0240] A4. The image processing method as described in A3, wherein the parallel attention module includes a position attention module and a channel attention module;
[0241] The step of performing weight setting on the depth feature map through a parallel attention module to generate a weight feature map includes:
[0242] Performing weight setting on the depth feature map through a position attention module to generate a first weight-setting feature map, and performing weight setting on the depth feature map through the channel attention module to generate a second weight-setting feature map;
[0243] Fusing the first weight-setting feature map and the second weight-setting feature map to generate a weight feature map.
[0244] A5. The image processing method as described in A3, wherein the multi-scale extraction module includes a point convolution kernel, a plurality of dilation degrees of the dilated convolutions, and a global pooling layer;
[0245] The performing parallel sampling extraction on the depth feature map by the multi-scale extraction module to generate a parallel sampling feature map includes:
[0246] Sampling and extracting the depth feature map through the point convolution kernel to obtain a first sampling feature map;
[0247] Sampling and extracting the depth feature map by using the plurality of dilated convolutions with different dilation degrees respectively to obtain a plurality of second sampling feature maps;
[0248] Performing global pooling processing on the depth feature map through the global pooling layer to generate a third sampling feature map;
[0249] The first sampling feature map, the multiple second sampling feature maps and the third sampling feature map are fused to generate a parallel sampling feature map.
[0250] A6. The image processing method as described in A1, wherein the target feature map, the shallow feature map and the middle feature map are subjected to multi-scale feature fusion by the decoder to generate a cutout image of the target to be cutout, comprising:
[0251] Performing upsampling processing on the target feature map based on a first upsampling ratio by the decoder to generate an upsampled target feature map;
[0252] The decoder fuses the expanded target feature map with the middle-level feature map to generate a first spliced feature map;
[0253] Performing upsampling processing on the first splicing feature map based on a second upsampling ratio by the decoder to generate an upsampled splicing feature map;
[0254] The decoder fuses the expanded sample splicing feature map with the shallow feature map to generate a second splicing feature map;
[0255] The decoder performs upsampling processing on the second spliced feature map based on a third upsampling ratio to generate a cutout image of the target to be cutout.
[0256] A7. The image processing method as described in A1, wherein before the encoder performs multi-scale feature extraction on the image to be processed containing the target to be extracted and generates a shallow feature map, a middle feature map and a high-level feature map, the method further comprises:
[0257] Train the initial extraction model through the model training set;
[0258] When the initial extraction model meets the preset convergence condition, the initial extraction model is verified by a model validation set to generate a model evaluation index value;
[0259] If the model evaluation index value meets the preset evaluation condition, the initial extraction model is used as an image extraction model.
[0260] A8. The image processing method as described in A7, wherein the initial extraction model includes an encoder and a decoder;
[0261] After obtaining the model evaluation index value corresponding to the initial extraction model when the initial extraction model meets the preset convergence condition, the method further includes:
[0262] When the initial extraction model does not meet the preset convergence condition, obtain a first loss value of the encoder in the initial extraction model and a second loss value of the decoder;
[0263] Construct an overall model loss according to a preset loss weight, the first loss value, and the second loss value;
[0264] Adjust the parameters of the initial extraction model according to the overall model loss, and return to the step of training the initial extraction model by a model training set.
[0265] A9. The image processing method as described in A7, before training the initial extraction model by a model training set, further includes:
[0266] Obtain a sample image set, where the sample image set contains sample images of multiple different scenes, and the sample images include scene images and target images;
[0267] Perform pixel-level annotation on each sample image in the sample image set to obtain an annotated image set;
[0268] Divide the annotated image set to generate a model training set and a model validation set.
[0269] A10. The image processing method as described in A9, before dividing the annotated image set to generate a model training set and a model validation set, further includes:
[0270] Perform size unification processing on each image in the annotated image set to generate a processed annotated image set;
[0271] Divide the processed annotated image set to generate a model training set and a model validation set.
[0272] A11. The image processing method as described in A9, before dividing the annotated image set to generate a model training set and a model validation set, further includes:
[0273] Perform data augmentation on each image in the labeled image set to obtain an augmented labeled image set, where the data augmentation includes at least one of rotation, flipping, cropping, and brightness adjustment;
[0274] Partition the augmented labeled image set to generate a model training set and a model validation set.
[0275] This application also discloses B12, an image processing device applied to an image matting model, where the image matting model includes an encoder and a decoder;
[0276] The image processing device includes:
[0277] An extraction module for performing multi-scale feature extraction on a to-be-processed image containing a to-be-matted target through the encoder to generate a shallow feature map, a middle feature map, and a high-level feature map;
[0278] A classification module for classifying feature points in the high-level feature map through the decoder to obtain a feature classification result;
[0279] A construction module for constructing a target feature map through the decoder according to the feature classification result and the high-level feature map;
[0280] A generation module for performing multi-scale feature fusion on the target feature map, the shallow feature map, and the middle feature map through the decoder to generate a matting image of the to-be-matted target.
[0281] B13. The image processing device according to B12, where the encoder includes a first convolutional group, a pooling layer, a second convolutional group, and a third convolutional group;
[0282] The extraction module is further configured to perform feature extraction on a to-be-processed image containing a to-be-matted target through the first convolutional group to generate a shallow feature map; perform secondary feature extraction on the shallow feature map through the pooling layer and the second convolutional group to generate a middle feature map; perform depth feature extraction on the middle feature map through the third convolutional group to generate a high-level feature map; where the size of the middle feature map is smaller than the size of the shallow feature map, and the size of the high-level feature map is smaller than the size of the middle feature map.
[0283] B14. The image processing device according to B13, where the encoder further includes a parallel attention module and a multi-scale extraction module;
[0284] The extraction module is also used to perform deep feature extraction on the middle-layer feature map through the third convolution group to generate a deep feature map; set weights on the deep feature map through the parallel attention module to generate a weighted feature map; perform parallel sampling extraction on the deep feature map through the multi-scale extraction module to generate a parallel sampling feature map; and fuse the weighted feature map with the parallel sampling feature map to generate a high-level feature map.
[0285] B15. The image processing device as described in B14, wherein the parallel attention module includes a position attention module and a channel attention module;
[0286] The extraction module is also used to set the weight of the depth feature map through the position attention module to generate a first weight setting feature map, and to set the weight of the depth feature map through the channel attention module to generate a second weight setting feature map; the first weight setting feature map and the second weight setting feature map are merged to generate a weight feature map.
[0287] B16. The image processing device as described in B14, wherein the multi-scale extraction module includes a point convolution kernel, a plurality of dilation degrees of the dilated convolutions, and a global pooling layer;
[0288] The extraction module is further used to sample and extract the depth feature map through the point convolution kernel to obtain a first sampling feature map; sample and extract the depth feature map through the multiple hole convolutions with different expansion degrees respectively to obtain multiple second sampling feature maps; perform global pooling processing on the depth feature map through the global pooling layer to generate a third sampling feature map; and fuse the first sampling feature map, the multiple second sampling feature maps and the third sampling feature map to generate a parallel sampling feature map.
[0289] B17. In the image processing device as described in B12, the generating module B is also used to upsample the target feature map based on the first upsampling ratio through the decoder to generate an enlarged target feature map; fuse the enlarged target feature map with the middle-layer feature map through the decoder to generate a first splicing feature map; upsample the first splicing feature map based on the second upsampling ratio through the decoder to generate an enlarged splicing feature map; fuse the enlarged splicing feature map with the shallow-layer feature map through the decoder to generate a second splicing feature map; upsample the second splicing feature map based on the third upsampling ratio through the decoder to generate a cutout image of the target to be cutout.
[0290] The present application also discloses C18, an image processing device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the image processing method as described above.
[0291] The present application also discloses D19, a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the image processing method as described above.
[0292] The present application also discloses E20, a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the image processing method as described above.
Claims
1. An image processing method, characterized in that: Applied to an image extraction model, the image extraction model includes an encoder and a decoder, the encoder includes a first convolution group, a pooling layer, a second convolution group and a third convolution group, the encoder also includes a parallel attention module and a multi-scale extraction module, the multi-scale extraction module includes a point convolution kernel, a plurality of hole convolutions with different dilation degrees and a global pooling layer; The image processing method comprises: Performing feature extraction on the image to be processed containing the target to be extracted through the first convolution group to generate a shallow feature map; Performing secondary feature extraction on the shallow feature map through a pooling layer and a second convolution group to generate a middle feature map; Performing deep feature extraction on the middle-layer feature map through a third convolution group to generate a deep feature map; The depth feature map is weighted by a parallel attention module to generate a weighted feature map, and the depth feature map is sampled and extracted in parallel by a multi-scale extraction module to generate a parallel sampling feature map; The weight feature map is fused with the parallel sampling feature map to generate a high-level feature map, wherein the size of the middle-level feature map is smaller than the size of the shallow-level feature map, and the size of the high-level feature map is smaller than the size of the middle-level feature map; Classifying the feature points in the high-level feature graph by the decoder to obtain a feature classification result; Constructing a target feature map according to the feature classification result and the high-level feature map by the decoder; The decoder performs multi-scale feature fusion on the target feature map, the shallow feature map and the middle feature map to generate a cutout image of the target to be cutout; The parallel sampling extraction is to process the depth feature map respectively through a point convolution kernel, a plurality of hole convolutions with different dilation degrees and a global pooling layer, and fuse the feature maps obtained by the respective processing.
2. The image processing method according to claim 1, characterized in that: The parallel attention module includes a position attention module and a channel attention module; The step of setting weights on the depth feature map by a parallel attention module to generate a weight feature map includes: The depth feature map is weighted by the position attention module to generate a first weighted feature map, and the depth feature map is weighted by the channel attention module to generate a second weighted feature map; The first weight setting feature map and the second weight setting feature map are fused to generate a weight feature map.
3. The image processing method according to claim 1, characterized in that: The performing parallel sampling extraction on the depth feature map by the multi-scale extraction module to generate a parallel sampling feature map includes: Sampling and extracting the depth feature map through the point convolution kernel to obtain a first sampling feature map; Sampling and extracting the depth feature map by using the plurality of dilated convolutions with different dilation degrees respectively to obtain a plurality of second sampling feature maps; Performing global pooling processing on the depth feature map through the global pooling layer to generate a third sampling feature map; The first sampling feature map, the multiple second sampling feature maps and the third sampling feature map are fused to generate a parallel sampling feature map.
4. The image processing method according to claim 1, wherein: The step of performing multi-scale feature fusion on the target feature map, the shallow feature map and the middle feature map through the decoder to generate a cutout image of the target to be cutout includes: Performing upsampling processing on the target feature map based on a first upsampling ratio by the decoder to generate an upsampled target feature map; The decoder fuses the expanded target feature map with the middle-level feature map to generate a first spliced feature map; Performing upsampling processing on the first splicing feature map based on a second upsampling ratio by the decoder to generate an upsampled splicing feature map; The decoder fuses the expanded sample splicing feature map with the shallow feature map to generate a second splicing feature map; The decoder performs upsampling processing on the second spliced feature map based on a third upsampling ratio to generate a cutout image of the target to be cutout.
5. An image processing device, characterized in that: Applied to an image extraction model, the image extraction model includes an encoder and a decoder, the encoder includes a first convolution group, a pooling layer, a second convolution group and a third convolution group, the encoder also includes a parallel attention module and a multi-scale extraction module, the multi-scale extraction module includes a point convolution kernel, a plurality of hole convolutions with different dilation degrees and a global pooling layer; The image processing device comprises: An extraction module, used for performing multi-scale feature extraction on the image to be processed containing the target to be extracted through the encoder, and generating a shallow feature map, a middle feature map and a high-level feature map; A classification module, used for classifying the feature points in the high-level feature map through the decoder to obtain a feature classification result; A construction module, configured to construct a target feature map according to the feature classification result and the high-level feature map through the decoder; A generation module, used for performing multi-scale feature fusion on the target feature map, the shallow feature map and the middle feature map through the decoder to generate a cutout image of the target to be cutout; Among them, the extraction module is also used to perform feature extraction on the image to be processed containing the target to be extracted through the first convolution group to generate a shallow feature map; perform secondary feature extraction on the shallow feature map through the pooling layer and the second convolution group to generate a middle feature map; perform deep feature extraction on the middle feature map through the third convolution group to generate a deep feature map; perform weight setting on the deep feature map through the parallel attention module to generate a weighted feature map, perform parallel sampling extraction on the deep feature map through the multi-scale extraction module to generate a parallel sampling feature map; fuse the weighted feature map with the parallel sampling feature map to generate a high-level feature map, the size of the middle feature map is smaller than the size of the shallow feature map, and the size of the high-level feature map is smaller than the size of the middle feature map; The parallel sampling extraction is to process the depth feature map respectively through a point convolution kernel, a plurality of hole convolutions with different dilation degrees and a global pooling layer, and fuse the feature maps obtained by the respective processing.
6. An image processing device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the image processing method according to any one of claims 1 to 4.
7. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 4 are implemented.
8. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Medical image segmentation method fusing multi-scale features and attention mechanism
CN114119638A