Method, system and related device for separating foreground and background in image
By converting images to RGB and HSV color spaces, extracting and weighting feature maps, and performing multiple spatial transformations and classifications, the problem of inaccurate separation of foreground and background in green screen matting is solved, achieving higher separation accuracy and resistance to external interference.
Patent Information
- Application Number
- CN202511275332.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-25
AI Technical Summary
The accuracy of separating the foreground and background in green screen keying is greatly affected by external conditions such as lighting and occlusion, resulting in poor separation results.
The image is converted to RGB and HSV color spaces, feature maps are extracted and stitched together, channel weights are calculated, and the feature maps are weighted and fused through an attention mechanism. Multiple spatial dimension transformations are performed and the images are classified to generate a mask that separates the foreground and background.
It improves the accuracy of separating the foreground and background in the image, reduces the influence of external conditions on the separation effect, and achieves more accurate image segmentation.
Smart Images

Figure CN121010764A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method, system and related apparatus for separating the foreground and background in an image. Background Technology
[0002] Green screen keying, also known as chroma keying, is a technique that separates the foreground from a single-color background (such as a green or blue background) by defining a range of colors. Specifically, when the color value of a region in an image or video falls within a defined color range, that region is identified as the background area of the image or video.
[0003] When shooting images or videos, external conditions (such as lighting, occlusion, etc.) can easily affect the color values of some areas in each frame of the image or video, causing the color values of some areas that should fall within the set color range to fall outside the set color range, thus affecting the accuracy of separating the foreground and background in green screen keying. Summary of the Invention
[0004] In view of the above problems, this application provides a method, system, and related apparatus for separating the foreground and background in an image, so as to improve the accuracy of foreground and background separation in green screen keying. The specific solution is as follows:
[0005] The first aspect of this application provides a method for separating the foreground and background in an image, comprising:
[0006] The current color space of the image to be processed is converted to the RGB color space to obtain a first image represented by the RGB color space, and the current color space of the image to be processed is converted to the HSV color space to obtain a second image represented by the HSV color space. The image to be processed includes a foreground and a background of a single color.
[0007] Image feature extraction is performed on the first image to obtain a first feature map of the first image. Image feature extraction is performed on the second image to obtain a second feature map of the second image. Based on the first feature map and the second feature map, a stitched feature map of the image to be processed is obtained. The image channels of the stitched feature map include: the RGB channel of the first feature map and the HSV channel of the second feature map.
[0008] Calculate the RGB channel weight and HSV channel weight of each pixel based on the channel features of each pixel in the stitched feature map, and obtain the RGB channel weight of each pixel in the first feature map and the HSV channel weight of each pixel in the second feature map.
[0009] The pixel value of each pixel in the first feature map is weighted according to the RGB channel weight of each pixel in the first feature map, and the pixel value of each pixel in the second feature map is weighted according to the HSV channel weight of each pixel in the second feature map. The weighted feature maps are then fused to obtain the weighted feature map of the image to be processed.
[0010] The weighted feature map of the image to be processed is subjected to multiple spatial dimension transformations to obtain the channel feature vector of each pixel in the weighted feature map, and the pixels are divided into foreground category and background category according to the channel feature vector;
[0011] The foreground and background in the image to be processed are separated based on the classification results of the pixels in the weighted feature map of the image to be processed.
[0012] In one possible implementation, for each pixel in the stitched feature map...
[0013] The step of calculating the RGB channel weight and HSV channel weight of each pixel based on the channel features of each pixel in the stitched feature map includes:
[0014] Perform a 1×1 convolution on the spliced feature map to obtain the intermediate feature map of the spliced feature map;
[0015] The intermediate feature map is then subjected to nonlinear activation;
[0016] The intermediate feature map after the nonlinear activation is convolved again with 1×1 to obtain the attention map of the spliced feature map. In the attention map, each pixel includes a first channel, an attention score corresponding to the first channel, a second channel, and an attention score corresponding to the second channel. The first channel corresponds to the RGB channel, and the second channel corresponds to the HSV channel.
[0017] The two attention scores of each pixel in the attention map are normalized to obtain the RGB channel weights and the HSV channel weights. The sum of the weights of the RGB channel weights and the HSV channel weights is 1.
[0018] In one possible implementation, fusing the weighted feature maps to obtain a weighted feature map of the image to be processed includes:
[0019] The pixel values of pixels at the same position in the weighted first feature map and the weighted second feature map are added together to obtain the weighted feature map of the image to be processed.
[0020] In one possible implementation, the channel feature vector includes: feature values of the RGB channel and feature values of the HSV channel.
[0021] The step of classifying pixels into foreground and background categories based on the channel feature vector includes:
[0022] The channel feature vector of each pixel is converted into a category score vector by a linear transformation, the category score vector including the initial score of the foreground category and the initial score of the background category;
[0023] The category score vector of each pixel is converted into a category probability distribution for each pixel, wherein the category probability distribution includes the probability value of the foreground category and the probability value of the background category;
[0024] From the probability values of the foreground category and the probability values of the background category for each pixel, select the category corresponding to the highest probability value as the category of the pixel.
[0025] One possible implementation also includes:
[0026] In the process of performing multiple spatial dimension transformations on the weighted feature map of the image to be processed to obtain the feature vector of each pixel in the weighted feature map, the channel weight and spatial weight of each pixel in the spatial feature map are generated, and the pixel value of each pixel in the spatial feature map is weighted according to the channel weight and spatial weight of each pixel. The spatial feature map is the feature map generated after each spatial dimension transformation in the process of multiple spatial dimension transformations.
[0027] In one possible implementation, separating the foreground and background in the image to be processed based on the classification results of pixels in the weighted feature map of the image to be processed includes:
[0028] Traverse each pixel in the image to be processed, and set the mask value of each pixel according to the classification result of each pixel in the weighted feature map to generate a mask for the image to be processed. In the mask, the mask value of the pixel whose classification result is foreground is different from the mask value of the pixel whose classification result is background.
[0029] For each pixel at the same position in the image to be processed and the mask, the pixel value is multiplied with the corresponding mask value, and the background and the foreground in the image to be processed are separated based on the pixel value of the image to be processed after the multiplication process.
[0030] A second aspect of this application provides a system for separating the foreground and background in an image, comprising:
[0031] An image processing unit is configured to convert the current color space of the image to be processed into the RGB color space to obtain a first image represented by the RGB color space, and to convert the current color space of the image to be processed into the HSV color space to obtain a second image represented by the HSV color space. The image to be processed includes a foreground and a background of a single color.
[0032] The feature extraction unit is used to extract image features from the first image to obtain a first feature map of the first image, extract image features from the second image to obtain a second feature map of the second image, and obtain a stitched feature map of the image to be processed based on the first feature map and the second feature map. The image channels of the stitched feature map include: the RGB channel of the first feature map and the HSV channel of the second feature map.
[0033] The weight calculation unit is used to calculate the RGB channel weight and HSV channel weight of each pixel based on the channel features of each pixel in the stitched feature map, so as to obtain the RGB channel weight of each pixel in the first feature map and the HSV channel weight of each pixel in the second feature map.
[0034] The feature update unit is used to weight the pixel value of each pixel in the first feature map according to the RGB channel weight of each pixel in the first feature map, and to weight the pixel value of each pixel in the second feature map according to the HSV channel weight of each pixel in the second feature map, and to fuse the weighted feature maps to obtain the weighted feature map of the image to be processed.
[0035] The feature classification unit is used to perform multiple spatial dimension transformations on the weighted feature map of the image to be processed, obtain the channel feature vector of each pixel in the weighted feature map, and classify the pixels into foreground category and background category according to the channel feature vector;
[0036] The foreground separation unit is used to separate the foreground and background in the image to be processed based on the classification results of pixels in the weighted feature map of the image to be processed.
[0037] In one possible implementation, for each pixel in the stitched feature map...
[0038] The weight calculation unit is specifically configured as follows:
[0039] A 1×1 convolution is performed on the stitched feature map to obtain an intermediate feature map; non-linear activation is applied to the intermediate feature map; a 1×1 convolution is then performed on the intermediate feature map after non-linear activation to obtain an attention map of the stitched feature map. In the attention map, each pixel includes a first channel, an attention score corresponding to the first channel, a second channel, and an attention score corresponding to the second channel. The first channel corresponds to the RGB channel, and the second channel corresponds to the HSV channel. The two attention scores of each pixel in the attention map are normalized to obtain the RGB channel weight and the HSV channel weight. The sum of the weights of the RGB channel weight and the HSV channel weight is 1.
[0040] A third aspect of this application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0041] The memory is used to store computer programs;
[0042] The processor is configured to execute the computer program to enable the electronic device to implement the method for separating the foreground and background in an image as described in the first aspect or any implementation thereof.
[0043] A fourth aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the method for separating the foreground and background in an image as described in the first aspect or any implementation thereof.
[0044] Based on the above technical solution, this application provides a method, system, and related apparatus for separating foreground and background in an image. The method converts the acquired image into a first image represented in RGB color space and a second image represented in HSV color space. A first feature map of the first image and a second feature map of the second image are extracted respectively. A stitched feature map of the image to be processed is obtained based on the first and second feature maps. The stitched feature map can include RGB and HSV channels. An attention mechanism is used to assign RGB and HSV channel weights to each pixel in the stitched feature map. The pixel values of each pixel in the first and second feature maps are weighted using the RGB and HSV channel weights respectively. The weighted feature maps are then fused to obtain a weighted feature map of the image to be processed. Finally, each pixel in the weighted feature map of the image to be processed is classified into foreground and background categories, facilitating the separation of foreground and single-color background in the image to be processed based on the pixel classification results, thus achieving image segmentation. This method converts images to different color spaces for representation, obtaining more detailed color information such as color, saturation, and brightness. It no longer relies solely on the current color value of the image for background and foreground separation, thus reducing the influence of external conditions. Furthermore, it uses an attention mechanism to dynamically determine the color information that different regions of the image rely on most, effectively addressing the impact of external conditions such as lighting on the color of different areas of the image. Therefore, this method can effectively improve the accuracy of foreground and background separation in images. Attached Figure Description
[0045] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0046] Figure 1 A flowchart illustrating a method for separating the foreground and background in an image, provided in an embodiment of this application;
[0047] Figure 2 This is a schematic diagram of the structure of a U-Net provided in an embodiment of this application;
[0048] Figure 3 This is a schematic diagram of the structure of a first downsampling block provided in an embodiment of this application;
[0049] Figure 4 This is a schematic diagram of the structure of a first upsampling block provided in an embodiment of this application;
[0050] Figure 5This application provides a schematic diagram of the structure of a system for separating the foreground and background in an image, as illustrated in an embodiment of the present application.
[0051] Figure 6 This application provides a hardware structure block diagram of an electronic device. Detailed Implementation
[0052] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0053] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0054] The terms "first," "second," etc., used in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0055] To address the issue of low accuracy in separating foreground and background during green screen keying, this application provides a method for separating foreground and background in an image. The method for separating foreground and background in an image according to this application will be described in detail below with reference to the accompanying drawings.
[0056] Reference Figure 1 , Figure 1 A flowchart illustrating a method for separating the foreground and background in an image, as provided in an embodiment of this application, is shown below. Figure 1 As shown in the embodiment of this application, a method for separating the foreground and background in an image may include steps S10 to S15, which are described in detail below.
[0057] S10. Convert the current color space of the image to be processed to the RGB color space to obtain a first image represented by the RGB color space, and convert the current color space of the image to be processed to the HSV color space to obtain a second image represented by the HSV color space. The image to be processed includes a foreground and a background of a single color.
[0058] The image to be processed can refer to an object that requires background and foreground separation. Optionally, the image to be processed can be a single image or a video frame from a video. The image to be processed can include a single-color background and foreground. In this embodiment, the foreground can refer to an area in the image that is of interest, highlighted, or needs to be analyzed, such as a person or object in the image. The background can refer to all unimportant areas other than the foreground, areas that are usually ignored, or areas that want to be removed. In the application scenario of this embodiment, the background can be a specific single color, such as a green screen background or a blue screen background. The current color space of the image to be processed can refer to the color representation method currently used by the image. In this embodiment, the current color space of the image to be processed can be the RGB color space, the HSV color space, or a color space other than RGB and HSV.
[0059] The RGB color space refers to a color representation method that uses red, green, and blue to represent all colors. The RGB color space can provide accurate color and texture information in an image. The HSV color space refers to a color representation method that uses hue, saturation, and lightness to represent all colors. The HSV color space can provide more stable color characteristics and illumination invariance in an image (the color representation remains constant under arbitrary changes in light). Hue refers to the basic identity of a color, such as the collective term for pure colors like red, orange, and yellow, without any grayscale or brightness. Saturation refers to the purity or vividness of a color; higher saturation results in a purer color, while lower saturation results in a grayer color. Lightness refers to the brightness of a color; higher lightness makes it closer to white, while lower lightness makes it closer to black. Therefore, in the first image represented using the RGB color space, the color of each pixel is represented by different combinations of red, green, and blue values. In the second image represented using the HSV color space, the color of each pixel is represented by different combinations of hue, saturation, and lightness values.
[0060] The conversion between color spaces is essentially a calculation process that maps one set of numerical representations of the current colors in an image to another set of numerical representations. For example, when converting from RGB color space to HSV color space, trigonometric functions or polar coordinate transformations can be used to convert the numerical values of the three RGB colors to the corresponding numerical values of the three HSV colors.
[0061] S11. Perform image feature extraction on the first image to obtain a first feature map of the first image. Perform image feature extraction on the second image to obtain a second feature map of the second image. Based on the first feature map and the second feature map, obtain a stitched feature map of the image to be processed. The image channels of the stitched feature map include: the RGB channel of the first feature map and the HSV channel of the second feature map.
[0062] The first feature map can be the feature map obtained after extracting image features from the first image, and the second feature map can be the feature map obtained after extracting image features from the second image. The image feature extraction method used in this embodiment is convolution, and the size of the convolution kernel used in the convolution can be 1×1.
[0063] This embodiment uses a feature stitching method to superimpose and stitch together a first feature map and a second feature map to obtain a stitched feature map. Superimposing and stitching feature maps can refer to stitching them together along the channel dimension, so that the stitched feature map has the same dimensions as the original feature map, but with more channels. Specifically, in this embodiment, since the first image and the second image have the same dimensions as the image to be processed, and since the same size convolution kernel is used to convolve the first image and the second image respectively to obtain the first feature map and the second feature map, the dimensions of the first feature map and the second feature map are the same. Since the first feature map is obtained from the first image represented in RGB color space, it has three channels: R, G, and B channels. Since the second feature map is obtained from the second image represented in HSV color space, it has three channels: H, S, and V channels. Therefore, after superimposing and stitching the first feature map and the second feature map, the dimensions of the stitched feature map are the same as those of the first feature map and the second feature map. Figure 1 However, the stitched feature map can have six channels: R, G, B, H, S, and V channels. Alternatively, in another optional embodiment, the RGB channels of the first feature map and the HSV channels of the second feature map can be combined using methods such as linear combination or dot product.
[0064] In another optional embodiment, during the process of obtaining the stitched feature map, in addition to the feature maps corresponding to the RGB color space and the HSV color space, feature maps of the image to be processed in other color spaces can be added according to different actual needs to obtain more information about the image to be processed. For example, in actual needs, it is also necessary to pay attention to the color difference of the image. Since the color representation of the LAB color space mainly includes complementary colors, when the first feature map (corresponding to RGB) and the second feature map (corresponding to HSV) are stitched together, the feature map obtained after converting the image to be processed to the LAB color space can also be added. This not only obtains the RGB and HSV information of the image to be processed, but also the color difference information of the image to be processed. The channels of the finally obtained feature stitched map can include: R channel, G channel, B channel, H channel, S channel, V channel, L channel, A channel, and B channel.
[0065] S12. Calculate the RGB channel weight and HSV channel weight of each pixel based on the channel features of each pixel in the stitched feature map, and obtain the RGB channel weight of each pixel in the first feature map and the HSV channel weight of each pixel in the second feature map.
[0066] After obtaining the stitched feature map in this embodiment, since the R, G, and B channels are three parallel channels, and the H, S, and V channels are also parallel channels, this embodiment uses an attention mechanism to calculate the attention weights between the RGB and HSV channels of each pixel in the stitched feature map based on the channel features of each pixel, thus obtaining the RGB and HSV channel weights of each pixel in the stitched feature map. Here, the channel features of each pixel can refer to the feature values of each pixel in different channels in the stitched feature map; in this embodiment, these are the RGB and HSV channels.
[0067] Specifically, the calculation process for the attention weights between the RGB and HSV channels of each pixel can be as follows:
[0068] A 1×1 convolution is performed on the stitched feature map to obtain an intermediate feature map. Non-linear activation is then applied to the intermediate feature map. A 1×1 convolution is performed again on the non-linearly activated intermediate feature map to obtain an attention map of the stitched feature map. In the attention map, each pixel includes a first channel, an attention score corresponding to the first channel, and attention scores corresponding to the second and third channels. The first channel corresponds to the RGB channel, and the second channel corresponds to the HSV channel. The two attention scores of each pixel in the attention map are normalized to obtain the RGB channel weight and the HSV channel weight. The sum of the weights of the RGB channel weight and the HSV channel weight is 1.
[0069] Among them, performing a 1×1 convolution on the stitched feature map can obtain an intermediate feature map with the same length and width as the stitched feature map. The 1×1 convolution can reduce the number of channels in the stitched feature map, while learning the relationship between the channels in each pixel and extracting more meaningful intermediate features.
[0070] The intermediate feature map contains higher-level feature representations extracted from the RGB and HSV channel features in the stitched feature map.
[0071] To extract more complex and richer feature relationships from the intermediate feature map, this embodiment performs nonlinear activation on the intermediate feature map. Nonlinear activation refers to introducing nonlinear characteristics to learn and represent complex nonlinear relationships. This embodiment mainly uses ReLU (Rectified Linear Unit) as the activation function to perform nonlinear activation processing on the intermediate feature map. The specific process can be as follows: traverse each pixel of the intermediate feature map. For each pixel value in the intermediate feature map, if the pixel value is negative, set the pixel value to 0; if the pixel value is positive, leave the pixel value unchanged. Of course, in another optional embodiment, nonlinear activation can be achieved using other activation functions, such as the Sigmoid function, Tanh function, etc.
[0072] After nonlinear activation, a 1×1 convolution is used again to convert the intermediate feature map after nonlinear activation into an attention map with two channels. In the attention map, each pixel includes two channels, and the feature value of each channel can be used as an attention score in a color space. In this embodiment, the feature value of the first channel can be used as the attention score of the RGB channel, and the feature value of the second channel can be used as the attention score of the HSV channel. To ensure that the attention scores are within a reasonable range and to effectively utilize them, this embodiment needs to normalize the attention scores of the RGB channel and the HSV channel respectively. The normalized attention scores can be used as weights, thereby obtaining the RGB channel weights and HSV channel weights. In this embodiment, the sum of the weights of the RGB color space and the weights of the HSV color space is 1, thereby obtaining the RGB channel weights and HSV channel weights of each pixel in the stitched feature map. Here, normalization processing can refer to a data preprocessing technique that can transform the data into a uniform range. This embodiment mainly uses the softmax function for normalization processing. The softmax function converts the attention score into a probability value, and this probability value is used as the weight.
[0073] Furthermore, in addition to the method described above for calculating the RGB channel weights and HSV channel weights of each pixel in the stitched feature map, this embodiment can also provide another method for calculating the RGB channel weights and HSV channel weights, the specific process of which is as follows:
[0074] From the perspective of RGB channels, calculate the self-attention score of RGB channels and the first attention score of RGB channels on HSV channels. Then, use softmax to normalize the attention weights of the self-attention scores and first attention scores of RGB channels to obtain the RGB channel weights and HSV channel weights.
[0075] Alternatively, one can choose to start from the perspective of the HSV channel, calculate the self-attention score of the HSV channel and the second attention score of the HSV channel on the RGB channel, and normalize the attention weights of the self-attention score and the second attention score of the HSV channel using softmax, thereby obtaining the RGB channel weights and HSV channel weights.
[0076] Taking the RGB channel as an example, the specific process of calculating the self-attention score of the RGB channel and the first attention score of the RGB channel on the HSV channel can be as follows: convert the RGB channel and the HSV channel into corresponding vectors respectively, calculate the dot product of the corresponding vector of the RGB channel with itself, and use it as the self-attention score of the RGB channel; calculate the dot product of the corresponding vector of the RGB channel with the corresponding vector of the HSV channel, and use it as the first attention score of the RGB channel on the HSV channel.
[0077] Taking the RGB channel as an example, the specific process of softmax normalizing attention weights can be as follows: the self-attention score and the first attention score of the RGB channel are exponentialized respectively. The self-attention score and the first attention score of the exponentialized RGB channel are added together, and the result is the normalization constant. The ratio of the self-attention score (numerator) of the exponentialized RGB channel to the normalization constant (denominator) is used as the RGB channel weight, and the ratio of the first attention score (numerator) of the exponentialized RGB channel to the normalization constant (denominator) is used as the HSV channel weight.
[0078] S13. The pixel value of each pixel in the first feature map is weighted according to the RGB channel weight of each pixel in the first feature map, and the pixel value of each pixel in the second feature map is weighted according to the HSV channel weight of each pixel in the second feature map. The weighted feature maps are then fused to obtain the weighted feature map of the image to be processed.
[0079] Since the dimensions of the stitched feature map are identical to those of the first and second feature maps, after obtaining the RGB and HSV channel weights of each pixel in the stitched feature map, the pixel values in the first and second feature maps can be weighted. Specifically, for each pixel in the first feature map, the pixel value can be weighted according to the RGB channel weights of pixels at the same position in the stitched feature map to update the first feature map. For each pixel in the second feature map, the pixel value can be weighted according to the HSV channel weights of pixels at the same position in the stitched feature map to update the second feature map, thereby obtaining the updated first and second feature maps. In this embodiment, the weighting of pixel values can refer to multiplying the channel weights corresponding to the color space with the pixel values of the feature map. Specifically, the RGB channel weights corresponding to each pixel in the first feature map can be multiplied by the pixel values of each pixel in the first feature map; the HSV channel weights corresponding to each pixel in the second feature map can be multiplied by the pixel values of each pixel in the second feature map to update the pixel values of the second feature map. Of course, in another alternative embodiment, the pixel values of each pixel in the first feature map or the pixel values of each pixel in the second feature map can be processed (e.g., nonlinear transformation) before weighted calculation.
[0080] After weighted calculations are performed on the pixel values in both the first and second feature maps, the weighted feature maps can be fused to obtain a weighted feature map of the image to be processed. Feature map fusion refers to the process of combining feature maps from different sources that represent different information to generate a richer and more expressive feature representation. In this embodiment, a weighted summation method is used to achieve the fusion of the weighted feature maps. Specifically, since the first and second weighted feature maps have the same width and height, this embodiment can directly add the pixel values of pixels at the same position in the first and second weighted feature maps to obtain the weighted feature map of the image to be processed. Alternatively, in another optional embodiment, this embodiment can first superimpose and concatenate the updated first and second feature maps, and then process them using convolution to achieve the fusion of the weighted first and second feature maps.
[0081] In another optional embodiment, during the fusion process of the weighted feature map corresponding to the RGB color space and the weighted feature map corresponding to the HSV color space, this embodiment can add additional weighted feature maps of the image to be processed in other color spaces according to different actual needs, so that the weighted feature map can contain weights of more color spaces.
[0082] Since each pixel in the weighted first feature map has an RGB channel weight, and each pixel in the weighted second feature map has an HSV channel weight, each pixel in the weighted feature map obtained by fusing the weighted first and second feature maps can have both RGB and HSV channel weights. This allows the setting of RGB and HSV channel weights in the subsequent image classification process to dynamically determine whether to rely more on RGB or HSV features at each pixel location. This enables better analysis of color information in different regions and reduces the influence of external conditions such as lighting. For example, in areas with clear edges, more RGB features can be relied upon, while in areas with changing lighting, more HSV features can be relied upon.
[0083] S14. Perform multiple spatial dimension transformations on the weighted feature map of the image to be processed, obtain the channel feature vector of each pixel in the weighted feature map, and classify the pixels into foreground and background categories based on the channel feature vectors.
[0084] S15. Separate the foreground and background in the image to be processed based on the classification results of the pixels in the weighted feature map of the image to be processed.
[0085] Spatial dimension transformation refers to adjusting the height and width of the feature map to change its spatial resolution. By performing multiple spatial dimension transformations, feature information at different scales can be combined to obtain richer feature representations. Therefore, in this embodiment, the channel feature vector of each pixel obtained after multiple spatial dimension transformations can possess multi-scale and rich feature information, facilitating pixel classification. Specifically, this embodiment obtains the category score for each pixel by performing a 1×1 convolution on the channel feature vector of each pixel. Since the channel feature vector can include feature values from both the RGB and HSV channels, the specific process of the 1×1 convolutional layer classifying pixels based on the channel feature vector can be as follows:
[0086] A linear transformation converts the channel feature vector of each pixel into a category score vector. This category score vector can include initial scores for the foreground and background categories. Specifically, a 1×1 convolutional layer includes a weight matrix and a bias vector, forming a linear relationship between the channel feature vector and the category score vector. The 1×1 convolutional layer uses this linear relationship to convert the channel feature vector into a category score vector. The weight matrix maps the feature vector to the category space by multiplying it with the feature vector. The bias vector provides a base score for each category. Adjusting the bias vector adjusts the category distribution; for example, a higher bias vector value for a certain category will result in a higher final category score, making that category more likely to be selected as the predicted category in the final probability distribution, all other things being equal.
[0087] To ensure that the category scores are within a reasonable range, this embodiment can convert the category score vector of each pixel into a category probability distribution for each pixel, which serves as the final category score for each pixel. Specifically, this embodiment can use the softmax function to implement the category probability conversion of the category score vector. The category probability distribution can include the probability values of the foreground category and the background category. After obtaining the category probability distribution for each pixel, the category corresponding to the largest probability value (the larger of the foreground and background probability values) can be selected as the pixel's category. For example, if a pixel has a foreground category probability value of 0.6 and a background category probability value of 0.4, and 0.6 is the maximum of the two, then the pixel can be identified as belonging to the foreground category.
[0088] After obtaining the category result of each pixel in the image to be processed, the foreground and background in the image to be processed can be separated according to the category result of the pixel. Specifically, in this embodiment, each pixel in the image to be processed can be traversed, and a mask value of each pixel can be set according to the classification result of each pixel in the weighted feature map to generate a mask for the image to be processed. For each pixel in the image to be processed and at the same position in the mask, the pixel value is multiplied with the corresponding mask value. The background and foreground in the image to be processed are separated according to the pixel value of the image to be processed after multiplication. Among them, the mask value of the pixel whose classification result is foreground category can be different from the mask value of the pixel whose classification result is background category. Then the mask can be a binary mask. Since the mask value can be the pixel value in the mask, in this binary mask, for each mask pixel (pixel in the mask) that has the same pixel position as the pixel in the weighted feature map or has a corresponding relationship, the mask value of the pixel in the weighted feature map can be the pixel value of the mask pixel in the mask. For example, pixel A is a pixel in the weighted feature map, and pixel B is a pixel in the mask. Pixel A's position in the weighted feature map is the same as pixel B's position in the mask. When pixel A's category is foreground and the mask value is set to 1, then in the mask, pixel B's pixel value is the same as pixel A's mask value, both being 1. Furthermore, in this binary mask, since the pixel values of the mask pixels corresponding to foreground category pixels are different from those corresponding to background category pixels, when the mask pixel value is multiplied by the pixel values of the image to be processed, the pixel values of pixels belonging to the foreground category are different from those belonging to the background category, facilitating the separation of foreground and background. For example, when traversing the weighted feature map of the image to be processed, if the pixel category is foreground, the mask value can be set to 1; if the pixel category is background, the mask value can be set to 0. This generates a binary mask with a mask value of 0-1. In this mask, the pixel value corresponding to the foreground category is 1, and the pixel value corresponding to the background category is 0. Multiplying the pixel value of this binary mask with the pixel value of the image to be processed, since the pixel value corresponding to the foreground category in the mask is 1 and the pixel value corresponding to the background category is 0, in the image to be processed, the pixel value multiplied by 0 becomes 0, and the pixel value multiplied by 1 remains unchanged. This ensures that the pixel value of the background category pixels in the image to be processed is 0, while the pixel value of the foreground category pixels remains unchanged. Therefore, the separation of background and foreground in the image to be processed can be achieved directly based on the pixel value. Of course, in another optional embodiment, an alpha mask can also be generated based on the pixel classification result to make the background transparent, thereby separating the background and foreground. Alpha masking is a special type of mask that uses the alpha channel (grayscale channel, which stores the transparency of each pixel) of a pixel to control the transparency of the pixel, and is used to precisely separate the foreground and background, or to achieve a local transparency effect.
[0089] Furthermore, this embodiment can employ U-Net to implement the aforementioned multiple spatial dimension transformations, pixel classification, and mask generation. U-Net (U-shaped convolutional neural network) can refer to a deep learning network architecture that may include an encoder part (multiple downsampling modules), a decoder part (multiple upsampling modules), and skip connections. Downsampling refers to the process of reducing the spatial dimension (spatial size) of a feature map while preserving the main features; upsampling refers to the process of increasing the spatial dimension (spatial size) of a feature map while preserving the main features. Skip connections are the core of U-Net, directly connecting the feature maps of the encoder part (downsampling) to the corresponding positions in the decoder part (upsampling), allowing low-level detail information to directly participate in the decoding process of high-level features.
[0090] To improve the accuracy of background and foreground separation in detailed regions (such as human hair and object outlines) of the image being processed, this embodiment incorporates CBAM (Convolutional Block Attention Module) into each upsampling and downsampling module of U-Net. CBAM can be defined as an attention mechanism module that enhances the feature representation capabilities of convolutional neural networks. By introducing channel attention and spatial attention, CBAM can better focus on important feature regions. Specifically, the structure of CBAM can include a channel attention module and a spatial attention module. The channel attention module can dynamically adjust the importance weight of each channel, allowing the convolutional neural network to focus on more important channel features, while the spatial attention module can dynamically adjust the importance weight of each spatial location, allowing the convolutional neural network to focus more on important spatial regions. In this embodiment, CBAM assigns weights to the feature map during U-Net processing, enabling U-Net to determine whether to focus more on channels or spatial locations in different regions of the feature map. Therefore, during the process of U-Net performing multiple spatial dimension transformations on the weighted feature map of the image to be processed, CBAM can generate the channel weight and spatial weight of each pixel in the spatial feature map in each spatial dimension transformation. It then weights the pixel value of each pixel in the spatial feature map based on these channel and spatial weights, allowing U-Net to determine the features of interest. Here, the spatial feature map refers to the feature map generated after each spatial dimension transformation during the multiple transformations. In U-Net, the spatial feature map can be the feature map output after each downsampling or each upsampling.
[0091] The structural diagram of the U-Net in this embodiment can be shown as follows: Figure 2As shown, the system includes three downsampling modules, an intermediate layer, three upsampling modules, skip connections, a convolutional layer, and a classification layer. Each upsampling module and each downsampling module may contain a CBAM (Convolutional Beam Array). In this embodiment, after obtaining the weighted feature map of the image to be processed, the feature map is subjected to multiple convolutions and pooling processes through a downsampling path (including a first downsampling block, a second downsampling block, and a third downsampling block). Then, the downsampling feature map is deconvolved through an upsampling path (including a first upsampling block, a second upsampling block, and a third upsampling block) to finally obtain a mask for the image to be processed.
[0092] The first, second, and third downsampling blocks have the same structure and processing procedure. Therefore, the structure and processing procedure of the downsampling block will be explained using the first downsampling block as an example. Figure 3 The diagram shows the structure of the first downsampling block, which may include a double convolutional layer, a CBMA (Convolutional Block Attention Module), and a max pooling layer. The double convolutional layer extracts features through two consecutive operations, the CBMA improves feature representation through channel attention and spatial attention modules, and the max pooling layer reduces computational complexity and preserves important features through downsampling.
[0093] The first, second, and third upsampling blocks have the same structure and the same processing procedure. Therefore, taking the first upsampling block as an example, the structure and processing procedure of the upsampling block will be explained as follows: Figure 4 The diagram shows the structure of the first upsampling block, which may include a deconvolution layer, a feature concatenation layer, and a double convolutional layer with CBMA (Convolutional Block Attention Module). The deconvolution layer increases the resolution of the feature map through deconvolution, the feature concatenation layer fuses features from different paths (features in upsampling (indicated by solid arrows) and features in downsampling (indicated by dashed arrows)) through feature concatenation, and the double convolutional layer with CBMA further enhances the feature representation through double convolution and the CBMA module.
[0094] Specifically, such as Figure 2As shown, the first, second, and third downsampling blocks are sequentially convolved and pooled. The output of the third downsampling block is input to the intermediate layer for convolution calculation. The output of the intermediate layer and the output of the third downsampling block are simultaneously input to the third upsampling block. The output of the second downsampling block is skipped, and the output of the third upsampling block and the output of the second downsampling block are used together as the input of the second upsampling block. The output of the first downsampling block is skipped, and the output of the second upsampling block and the output of the first downsampling block are used together as the input of the first upsampling block. The output of the first upsampling block is a feature map. This feature map is convolved with a 1×1 convolution, mapping the number of channels of each pixel in the feature map to two channels. Each channel corresponds to a category, and the feature value of each channel can correspond to an initial category score. Finally, the initial category score is converted into a probability value through an activation function. Based on the probability value, the pixel is classified into background and foreground, and a mask is output to separate the foreground and background in the image to be processed.
[0095] Thus, this embodiment provides a model for separating the foreground and background in an image. The architecture of this model may include a data processing module, a color fusion module, and a U-Net main module.
[0096] During training, the training data for this model can include multiple green screen images to be processed and their corresponding masks (as labeled data for the green screen images). The training data can be processed with color perturbations, blur perturbations, and random flipping to expand the training dataset. Color perturbations can enhance the model's adaptability to tonal variations on specific single-color backgrounds (such as green or blue backgrounds), while blur perturbations can enhance the model's robustness to motion blur or lens defocus, thereby improving the model's generalization ability after training.
[0097] In the trained model, the data processing module can perform data augmentation on the image to be processed to improve the image quality, convert the current color space of the image to be processed to the RGB color space to obtain the first image represented by the RGB color space, and convert the current color space of the image to the HSV color space to obtain the second image represented by the HSV color space.
[0098] The color fusion module can extract image features from the first image and the second image respectively, obtain the first feature map and the second feature map, and then overlay and stitch them together to obtain the stitched feature map. The attention mechanism is used to calculate the RGB channel weight and HSV channel weight of each pixel in the stitched feature map. The first feature map and the second feature map are weighted and then fused to obtain the weighted feature map of the image to be processed.
[0099] The main U-Net module can perform multiple downsampling and upsampling on the weighted feature map of the image to be processed. The CBAM module within it improves the segmentation processing of U-Net in the detailed region. The final output of the upsampled data is subjected to 1×1 convolution and binary classification. The output is used as a mask to separate the foreground and background in the image to be processed.
[0100] This application provides a method for separating foreground and background in an image. The method converts the acquired image into a first image represented in RGB color space and a second image represented in HSV color space. A first feature map of the first image and a second feature map of the second image are extracted. A stitched feature map of the image to be processed is obtained based on the first and second feature maps. The stitched feature map can include RGB and HSV channels. An attention mechanism is used to assign RGB and HSV channel weights to each pixel in the stitched feature map. The RGB and HSV channel weights of each pixel are then used to weight the pixel values of each pixel in the first and second feature maps, respectively. The weighted feature maps are then fused to obtain a weighted feature map of the image to be processed. Finally, each pixel in the weighted feature map of the image to be processed is classified into foreground and background categories, facilitating the separation of the foreground and a single-color background in the image to be processed based on the pixel classification results, thus achieving image segmentation. This method converts images to different color spaces for representation, obtaining more detailed color information such as color, saturation, and brightness. It no longer relies solely on the current color value of the image for background and foreground separation, thus reducing the influence of external conditions. Furthermore, it uses an attention mechanism to dynamically determine the color information that different regions of the image rely on most, effectively addressing the impact of external conditions such as lighting on the color of different areas of the image. Therefore, this method can effectively improve the accuracy of foreground and background separation in images.
[0101] The above describes a method for separating the foreground and background in an image according to embodiments of this application. The following will describe a system that applies the above method for separating the foreground and background in an image.
[0102] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a system for separating the foreground and background in an image, provided as an embodiment of this application. Figure 5 As shown, the system for separating the foreground and background in an image includes:
[0103] The image processing unit 100 is used to convert the current color space of the image to be processed into the RGB color space to obtain a first image represented by the RGB color space, and to convert the current color space of the image to be processed into the HSV color space to obtain a second image represented by the HSV color space. The image to be processed includes a foreground and a background of a single color.
[0104] The feature extraction unit 110 is used to extract image features from the first image to obtain a first feature map of the first image, extract image features from the second image to obtain a second feature map of the second image, and obtain a stitched feature map of the image to be processed based on the first feature map and the second feature map. The image channels of the stitched feature map include: the RGB channel of the first feature map and the HSV channel of the second feature map.
[0105] The weight calculation unit 120 is used to calculate the RGB channel weight and HSV channel weight of each pixel based on the channel features of each pixel in the stitched feature map, so as to obtain the RGB channel weight of each pixel in the first feature map and the HSV channel weight of each pixel in the second feature map.
[0106] The feature update unit 130 is used to weight the pixel value of each pixel in the first feature map according to the RGB channel weight of each pixel in the first feature map, and to weight the pixel value of each pixel in the second feature map according to the HSV channel weight of each pixel in the second feature map, and to fuse the weighted feature maps to obtain a weighted feature map of the image to be processed.
[0107] The feature classification unit 140 is used to perform multiple spatial dimension transformations on the weighted feature map of the image to be processed, obtain the channel feature vector of each pixel in the weighted feature map, and classify the pixels into foreground category and background category according to the channel feature vector;
[0108] The foreground separation unit 150 is used to separate the foreground and background in the image to be processed based on the classification results of pixels in the weighted feature map of the image to be processed.
[0109] In one possible implementation, for each pixel in the stitched feature map...
[0110] The weight calculation unit 120 can be specifically configured as follows:
[0111] A 1×1 convolution is performed on the stitched feature map to obtain an intermediate feature map. Non-linear activation is then applied to the intermediate feature map. A 1×1 convolution is performed again on the non-linearly activated intermediate feature map to obtain an attention map of the stitched feature map. In the attention map, each pixel includes a first channel, an attention score corresponding to the first channel, and attention scores corresponding to the second and third channels. The first channel corresponds to the RGB channel, and the second channel corresponds to the HSV channel. The two attention scores of each pixel in the attention map are normalized to obtain the RGB channel weight and the HSV channel weight. The sum of the weights of the RGB channel weight and the HSV channel weight is 1.
[0112] In one possible implementation, the feature update unit 130 fuses the weighted feature maps to obtain a weighted feature map of the image to be processed. Specifically, this can be configured as follows:
[0113] The pixel values of pixels at the same position in the weighted first feature map and the weighted second feature map are added together to obtain the weighted feature map of the image to be processed.
[0114] In one possible implementation, the channel feature vector includes: feature values of the RGB channels and feature values of the HSV channels.
[0115] In feature classification unit 140, pixels are divided into foreground and background categories based on channel feature vectors. Specifically, this can be configured as follows:
[0116] The channel feature vector of each pixel is transformed into a category score vector through linear transformation. The category score vector includes the initial score of the foreground category and the initial score of the background category. The category score vector of each pixel is then transformed into a category probability distribution for each pixel, which includes the probability value of the foreground category and the probability value of the background category. From the probability values of the foreground category and the probability values of the background category of each pixel, the category corresponding to the highest probability value is selected as the category of the pixel.
[0117] In one possible implementation, the feature classification unit 140 can also be configured as follows:
[0118] In the process of performing multiple spatial dimension transformations on the weighted feature map of the image to be processed to obtain the feature vector of each pixel in the weighted feature map, the channel weight and spatial weight of each pixel in the spatial feature map are generated, and the pixel value of each pixel in the spatial feature map is weighted according to the channel weight and spatial weight of each pixel. The spatial feature map is: the feature map generated after each spatial dimension transformation in the process of multiple spatial dimension transformations.
[0119] In one possible implementation, the foreground separation unit 150 can be specifically configured as follows:
[0120] Iterate through each pixel in the image to be processed, and set a mask value for each pixel based on the classification result of each pixel in the weighted feature map to generate a mask for the image to be processed. In the mask, the mask value of the pixel classified as foreground is different from the mask value of the pixel classified as background. For each pixel at the same position in the image to be processed and in the mask, multiply the pixel value with the corresponding mask value. Separate the background and foreground in the image to be processed based on the pixel value of the image to be processed after the multiplication process.
[0121] This application also provides an electronic device in its embodiments. (See reference...) Figure 6 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0122] like Figure 6 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0123] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, memory cards, hard drives, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0124] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the methods for separating the foreground and background in an image provided in this application.
[0125] This application also provides a computer-readable storage medium carrying one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device is able to implement any of the methods for separating the foreground and background in an image provided in this application.
[0126] It should also be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the system embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0127] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0128] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0129] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0130] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0131] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0132] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for separating the foreground and background in an image, characterized in that, include: The current color space of the image to be processed is converted to the RGB color space to obtain a first image represented by the RGB color space, and the current color space of the image to be processed is converted to the HSV color space to obtain a second image represented by the HSV color space. The image to be processed includes a foreground and a background of a single color. Image feature extraction is performed on the first image to obtain a first feature map of the first image. Image feature extraction is performed on the second image to obtain a second feature map of the second image. Based on the first feature map and the second feature map, a stitched feature map of the image to be processed is obtained. The image channels of the stitched feature map include: the RGB channel of the first feature map and the HSV channel of the second feature map. Calculate the RGB channel weight and HSV channel weight of each pixel based on the channel features of each pixel in the stitched feature map, and obtain the RGB channel weight of each pixel in the first feature map and the HSV channel weight of each pixel in the second feature map. The pixel value of each pixel in the first feature map is weighted according to the RGB channel weight of each pixel in the first feature map, and the pixel value of each pixel in the second feature map is weighted according to the HSV channel weight of each pixel in the second feature map. The weighted feature maps are then fused to obtain the weighted feature map of the image to be processed. The weighted feature map of the image to be processed is subjected to multiple spatial dimension transformations to obtain the channel feature vector of each pixel in the weighted feature map, and the pixels are divided into foreground category and background category according to the channel feature vector; The foreground and background in the image to be processed are separated based on the classification results of the pixels in the weighted feature map of the image to be processed.
2. The method for separating the foreground and background in an image according to claim 1, characterized in that, For each pixel in the stitched feature map, The step of calculating the RGB channel weight and HSV channel weight of each pixel based on the channel features of each pixel in the stitched feature map includes: Perform a 1×1 convolution on the spliced feature map to obtain the intermediate feature map of the spliced feature map; The intermediate feature map is then subjected to nonlinear activation; The intermediate feature map after the nonlinear activation is convolved again with 1×1 to obtain the attention map of the spliced feature map. In the attention map, each pixel includes a first channel, an attention score corresponding to the first channel, a second channel, and an attention score corresponding to the second channel. The first channel corresponds to the RGB channel, and the second channel corresponds to the HSV channel. The two attention scores of each pixel in the attention map are normalized to obtain the RGB channel weights and the HSV channel weights. The sum of the weights of the RGB channel weights and the HSV channel weights is 1.
3. The method for separating the foreground and background in an image according to claim 1, characterized in that, The step of fusing the weighted feature maps to obtain a weighted feature map of the image to be processed includes: The pixel values of pixels at the same position in the weighted first feature map and the weighted second feature map are added together to obtain the weighted feature map of the image to be processed.
4. The method for separating the foreground and background in an image according to claim 1, characterized in that, The channel feature vector includes: feature values of the RGB channel and feature values of the HSV channel. The step of classifying pixels into foreground and background categories based on the channel feature vector includes: The channel feature vector of each pixel is converted into a category score vector by a linear transformation, the category score vector including the initial score of the foreground category and the initial score of the background category; The category score vector of each pixel is converted into a category probability distribution for each pixel, wherein the category probability distribution includes the probability value of the foreground category and the probability value of the background category; From the probability values of the foreground category and the probability values of the background category for each pixel, select the category corresponding to the highest probability value as the category of the pixel.
5. The method for separating the foreground and background in an image according to claim 1, characterized in that, Also includes: In the process of performing multiple spatial dimension transformations on the weighted feature map of the image to be processed to obtain the feature vector of each pixel in the weighted feature map, the channel weight and spatial weight of each pixel in the spatial feature map are generated, and the pixel value of each pixel in the spatial feature map is weighted according to the channel weight and spatial weight of each pixel. The spatial feature map is the feature map generated after each spatial dimension transformation in the process of multiple spatial dimension transformations.
6. The method for separating the foreground and background in an image according to claim 1, characterized in that, The step of separating the foreground and background in the image to be processed based on the classification results of pixels in the weighted feature map of the image to be processed includes: Traverse each pixel in the image to be processed, and set the mask value of each pixel according to the classification result of each pixel in the weighted feature map to generate a mask for the image to be processed. In the mask, the mask value of the pixel whose classification result is foreground is different from the mask value of the pixel whose classification result is background. For each pixel at the same position in the image to be processed and the mask, the pixel value is multiplied with the corresponding mask value, and the background and the foreground in the image to be processed are separated based on the pixel value of the image to be processed after the multiplication process.
7. A system for separating the foreground and background in an image, characterized in that, include: An image processing unit is configured to convert the current color space of the image to be processed into the RGB color space to obtain a first image represented by the RGB color space, and to convert the current color space of the image to be processed into the HSV color space to obtain a second image represented by the HSV color space. The image to be processed includes a foreground and a background of a single color. The feature extraction unit is used to extract image features from the first image to obtain a first feature map of the first image, extract image features from the second image to obtain a second feature map of the second image, and obtain a stitched feature map of the image to be processed based on the first feature map and the second feature map. The image channels of the stitched feature map include: the RGB channel of the first feature map and the HSV channel of the second feature map. The weight calculation unit is used to calculate the RGB channel weight and HSV channel weight of each pixel based on the channel features of each pixel in the stitched feature map, so as to obtain the RGB channel weight of each pixel in the first feature map and the HSV channel weight of each pixel in the second feature map. The feature update unit is used to weight the pixel value of each pixel in the first feature map according to the RGB channel weight of each pixel in the first feature map, and to weight the pixel value of each pixel in the second feature map according to the HSV channel weight of each pixel in the second feature map, and to fuse the weighted feature maps to obtain the weighted feature map of the image to be processed. The feature classification unit is used to perform multiple spatial dimension transformations on the weighted feature map of the image to be processed, obtain the channel feature vector of each pixel in the weighted feature map, and classify the pixels into foreground category and background category according to the channel feature vector; The foreground separation unit is used to separate the foreground and background in the image to be processed based on the classification results of pixels in the weighted feature map of the image to be processed.
8. The system for separating the foreground and background in an image according to claim 7, characterized in that, For each pixel in the stitched feature map, The weight calculation unit is specifically configured as follows: The stitched feature map is convolved with a 1×1 convolution to obtain an intermediate feature map; the intermediate feature map is then non-linearly activated; the intermediate feature map after non-linear activation is convolved with a 1×1 convolution again to obtain an attention map of the stitched feature map. In the attention map, each pixel includes a first channel, an attention score corresponding to the first channel, a second channel, and an attention score corresponding to the second channel. The first channel corresponds to the RGB channel, and the second channel corresponds to the HSV channel. The two attention scores of each pixel in the attention map are normalized to obtain the RGB channel weights and the HSV channel weights. The sum of the weights of the RGB channel weights and the HSV channel weights is 1.
9. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is configured to execute the computer program to enable the electronic device to implement the method for separating the foreground and background in an image as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the method for separating the foreground and background in an image as described in any one of claims 1 to 6.