Image background feature extraction method, device, equipment and medium
By segmenting and filling the image, generating a non-background mask and labeling the attention tag, the problem of inaccurate background feature extraction in the prior art is solved, and a higher accuracy and processing efficiency of background feature extraction are achieved.
Patent Information
- Application Number
- CN202210869075.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-22
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-07-22
AI Technical Summary
In the prior art, the accuracy of image background feature extraction is low, and is affected and interfered by the large area of the portrait part in the image and the location information, resulting in the CNN model being unable to accurately extract background features.
The image is segmented by a preset segmentation model, a portrait mask is generated and filled, and expanded into a non-background mask. The attention tag is marked with a self-attention matrix, and the position information of the portrait part is blocked, and only the background area features are extracted.
The accuracy of the preset self-attention model on the background area features of the image is improved, the interference of portrait parts on background feature extraction is reduced, and the processing efficiency is improved.
Smart Images

Figure CN115376186B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to an image background feature extraction method, device, equipment and medium. Background Art
[0002] In the business process, it is necessary to analyze the self-taken images submitted by customers. The self-taken images mainly include a portrait part and a background part, and background information is extracted from the images to analyze customer information.
[0003] Generally, in the prior art, portrait segmentation is performed on an image, and the numerical values of the portrait part of the image are uniformly set to 0 to remove the pixel information of the portrait part. However, the position information of the portrait part still remains in the image. When a CNN model extracts features from the segmented image, affected and interfered by the large area ratio of the portrait part in the image and the position information of the portrait part, the network convolution kernel of the CNN model cannot accurately extract the background features of the image. Summary of the Invention
[0004] In view of the above, the present invention provides an image background feature extraction method, device, equipment and medium, aiming to solve the technical problem of low accuracy in extracting image background features in the prior art.
[0005] To achieve the above object, the present invention provides an image background feature extraction method, which includes:
[0006] Using a preset segmentation model to segment a to-be-processed image to obtain a portrait mask of the to-be-processed image, and combining the portrait mask with the to-be-processed image to obtain a portrait segmentation map;
[0007] Filling the portrait segmentation map to obtain a filled image, segmenting the filled image to obtain a portrait filled segmentation map of n x n small blocks, and expanding the n x n small blocks into a 1 x n 2 non-background mask, where the portrait filled segmentation map includes a background area and a non-background area;
[0008] Performing matrix calculation on the non-background mask to obtain a non-background self-attention mask, generating a self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filled segmentation map, and marking attention labels for each pair of small blocks in the background area according to the non-background self-attention mask;
[0009] Inputting the self-attention matrix and the segmentation map matrix into a preset self-attention model for matrix multiplication, and extracting the features of the small blocks corresponding to all elements with a value of 1 in the portrait filled segmentation map from the result of the matrix multiplication according to the attention labels as the background features of the to-be-processed image.
[0010] Preferably, the step of using a preset segmentation model to segment the image to be processed to obtain a portrait mask of the image to be processed includes:
[0011] Performing feature extraction on the image to be processed to obtain a feature map of the image to be processed;
[0012] Segmenting the portrait part of the feature map according to a preset portrait area selection frame to obtain the portrait mask.
[0013] Preferably, the step of segmenting the portrait part of the feature map according to a preset portrait area selection frame to obtain the portrait mask includes:
[0014] Performing upsampling on the feature map to expand the feature map to a preset resolution to obtain an expanded feature map;
[0015] According to the portrait area selection frame, distinguishing each pixel of the expanded feature map to obtain the portrait part of the feature map and performing segmentation, and generating a portrait mask of the image to be processed according to the segmented portrait part.
[0016] Preferably, the step of filling the portrait segmentation map to obtain a filled image includes:
[0017] Reading the numerical values of the side lengths of each side of the portrait segmentation map, and selecting the side with the largest numerical value;
[0018] Scaling the portrait segmentation map proportionally with the selected side with reference to a preset square frame so that the side length of the selected side is equal to the side length of the preset square frame;
[0019] Performing image filling on the area of the portrait segmentation map that does not exceed the preset square frame to obtain the filled image.
[0020] Preferably, the step of performing matrix calculation on the non-background mask to obtain a non-background self-attention mask and generating a self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filled segmentation map includes:
[0021] Extracting the features of each small block of the non-background mask to perform matrix calculation to obtain a non-background attention mask;
[0022] And extracting the features of each small block of the portrait filled segmentation map to generate the segmentation map matrix;
[0023] Reading each element of the mask matrix and each element of the segmentation map matrix to perform matrix multiplication to generate the self-attention matrix.
[0024] Preferably, marking attention labels for each pair of small blocks in the background area according to the non-background self-attention mask includes:
[0025] Judging the relevance between each small block of the portrait filling segmentation map according to the non-background self-attention mask;
[0026] If it is judged that the values of the small blocks in pairs are both greater than or equal to the preset value, then the small blocks in pairs are used as the background area and marked as attention labels.
[0027] Preferably, after judging that the values of the small blocks in pairs are both greater than or equal to the preset value, then using the small blocks in pairs as the background area and marking them as attention labels, the method includes:
[0028] If it is judged that the values of the small blocks in pairs are both less than the preset value, then the small blocks in pairs are used as the non-background area and marked as non-attention labels.
[0029] To achieve the above object, the present invention also provides an image background feature extraction device, and the device includes:
[0030] A segmentation module: used to segment the image to be processed by using a preset segmentation model to obtain a portrait mask of the image to be processed, and merge the portrait mask and the image to be processed to obtain a portrait segmentation map;
[0031] A sorting module: used to fill the portrait segmentation map to obtain a filled image, segment the filled image to obtain a portrait filling segmentation map of n x n small blocks, and expand the n x n small blocks into a 1 x n 2 non-background mask, and the portrait filling segmentation map includes a background area and a non-background area;
[0032] A marking module: used to perform matrix calculation on the non-background mask to obtain a non-background self-attention mask, generate a self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filling segmentation map, and according to the non-background self-attention mask, mark attention labels for each pair of small blocks in the background area;
[0033] An output module: used to input the self-attention matrix and the segmentation map matrix into a preset self-attention model for matrix multiplication, and extract the features of the small blocks corresponding to all elements with a value of 1 in the portrait filling segmentation map from the result of the matrix multiplication according to the attention labels as the background features of the image to be processed.
[0034] To achieve the above object, the present invention also provides an electronic device, and the electronic device includes:
[0035] at least one processor; and,
[0036] a memory communicatively connected to the at least one processor; wherein,
[0037] the memory stores a program executable by the at least one processor, and when the program is executed by the at least one processor, the at least one processor is enabled to execute the image background feature extraction method according to any one of claims 1 to 7.
[0038] To achieve the above object, the present invention further provides a computer-readable medium storing image background features, and when the image background features are executed by a processor, the steps of the image background feature extraction method according to any one of claims 1 to 7 are implemented.
[0039] The present invention segments the image to be processed to obtain a portrait mask and a portrait segmentation map. After filling and segmentation processing on the portrait segmentation map, a filled and segmented portrait map is obtained, and each small block of the filled and segmented portrait map is expanded into a one-dimensional shape to generate a non-background mask. By using the multiple small blocks obtained by segmenting the filled and segmented portrait map, the problem that the area of the portrait part in the image is relatively large in the prior art is solved, enabling the preset self-attention model to effectively distinguish and analyze each pixel point area.
[0040] A segmentation map matrix is generated for the filled and segmented portrait map, matrix calculation is performed on the non-background mask to obtain a non-background self-attention mask, a self-attention matrix is generated according to the non-background self-attention mask and the segmentation map matrix, and according to the non-background self-attention mask, the small blocks in the background area are marked with attention labels in pairs. According to the attention labels, the position information of the non-background area of the filled and segmented portrait map is masked, so that the position information of the portrait part in the image to be processed cannot interfere with the feature extraction of the background area, improving the accuracy of the features of the background area of the image to be processed by the preset self-attention model. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flowchart schematic diagram of a preferred embodiment of the image background feature extraction method of the present invention;
[0042] Figure 2 It is a module schematic diagram of a preferred embodiment of the image background feature extraction device of the present invention;
[0043] Figure 3 It is a schematic diagram of a preferred embodiment of an electronic device of the present invention;
[0044] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0046] The embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0047] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0048] The present invention provides a method for extracting image background features. Refer to Figure 1 As shown, it is a schematic flowchart of the method of the embodiment of the method for extracting image background features of the present invention. This method can be executed by an electronic device, and the electronic device can be implemented by software and / or hardware. The method for extracting image background features includes the following steps S10 - S40:
[0049] Step S10: Use a preset segmentation model to segment the image to be processed, obtain the portrait mask of the image to be processed, and merge the portrait mask with the image to be processed to obtain a portrait segmentation map.
[0050] In this embodiment, in order to meet the requirement of accurately extracting the background information of the image to be processed, the image to be processed is input into a preset segmentation model for segmentation. The preset segmentation model refers to a pre-trained portrait segmentation model (maskRegion-CNN), which is an algorithm model that applies deep learning to object detection. The image to be processed refers to an image containing an object and a background. The object includes, but is not limited to, a portrait, an animal, etc., and the background refers to the elements or scenes that set off the object in the image to be processed.
[0051] After segmenting the image to be processed using a preset segmentation model, the background part and the portrait part of the image to be processed are obtained. Although the background part and the portrait part have been segmented, there is still a correlation in position information between the background part and the portrait part. This correlation will affect the feature extraction of the portrait part or the background part, and it is easy to result in a low accuracy of the analyzed data.
[0052] Extract the features of the portrait part to generate a portrait mask for the image to be processed, and align the portrait mask with the position of the portrait part of the image to be processed for merging to obtain a portrait segmentation map. That is to say, in order for the portrait part of the image to be processed not to participate in the feature extraction process, the portrait part of the image to be processed is filled with an image using the shielding property of the portrait mask (for example, the portrait part is uniformly filled with black in the image filling). In fact, a mask is a two-dimensional matrix array. A portrait mask refers to selecting and shielding the portrait part so that the portrait part does not participate in the calculation of processing parameters, reducing the procedures in the image processing process and facilitating the improvement of the processing efficiency of image processing.
[0053] In one embodiment, the using the preset segmentation model to segment the image to be processed to obtain the portrait mask of the image to be processed includes:
[0054] Extract the features of the image to be processed to obtain the feature map of the image to be processed;
[0055] Segment the portrait part of the feature map according to a preset portrait area selection box to obtain the portrait mask.
[0056] Read the feature parameters (N×H×W) of the image to be processed and input them into the preset segmentation model. Process the feature parameters according to the feature segmentation module of the preset segmentation model, and output the feature map of the image to be processed. The feature map of the image to be processed includes N'×H'×W'. Among them, N represents the number of image channels of the image to be processed, H represents the height of the image to be processed, W represents the width of the image to be processed, N' represents the number of channels of the feature map, H' represents the height of the feature map, and W' represents the width of the feature map. Usually, H' is less than H, and W' is less than W. The specific structure of this feature segmentation module can select network extraction structures such as the VGG network and the Deep Residual Network (ResNet).
[0057] The preset portrait region selection box is an anchor box label trained by slicing positive and negative samples of a preset number of feature maps using a region network (e.g., RegionProposalNet - RPN network). This anchor box label can automatically identify the true border of the portrait part in the image and segment out the portrait part. Other networks based on region candidate extraction can also be used. The embodiments of this application do not limit the extraction network of the portrait region selection box, as long as it can extract the portrait region selection box from the feature map. The portrait part of the image to be processed is segmented using the portrait region selection box to obtain the portrait mask of the image to be processed.
[0058] In one embodiment, the segmenting the portrait part of the feature map according to the preset portrait region selection box to obtain the portrait mask includes:
[0059] Performing upsampling processing on the feature map to expand the feature map to a preset resolution to obtain an expanded feature map;
[0060] According to the portrait region selection box, each pixel of the expanded feature map is distinguished to obtain the portrait part of the feature map and segmented, and the portrait mask of the image to be processed is generated according to the segmented portrait part.
[0061] Performing upsampling processing on the feature map includes: According to the transposed convolution of the region network, expanding the resolution of the feature map to a preset resolution (e.g., the preset resolution is 1920 or 2K).
[0062] According to the portrait region selection box, the difference in the gray - scale value of each pixel of the expanded feature map and the HARR feature (region pixel sum feature) are used to distinguish and segment the portrait part. The tensor input of the segmented portrait part is multiplied element - by - element with each element of the kernel tensor of the preset segmentation model and placed in the corresponding matrix position for feature mapping to generate the portrait mask of the image to be processed. To better detect and distinguish the portrait part and the background part, the portrait mask is extracted down to each pixel, which is beneficial to the accuracy of background feature extraction.
[0063] Before using the above - mentioned preset segmentation model to segment the image to be processed, a feature segmentation module needs to be established in advance, then it is trained using sample images, and finally the trained feature segmentation module is used to extract features from the image to be processed to obtain the corresponding feature map.
[0064] Step S20: Fill the portrait segmentation map to obtain a filled image, segment the filled image to obtain a portrait fill - segmentation map of n x n small blocks, and expand the n x n small blocks into a 1 x n 2 non - background mask. The portrait fill - segmentation map includes a background region and a non - background region.
[0065] In this embodiment, before filling the portrait segmentation map to obtain a filled image, a regular square box needs to be preset as a reference object in advance, and the portrait segmentation map is scaled proportionally to reach the size of the preset regular square box. The size of the preset regular square box is such that the side length is divisible by 16. For the convenience of illustration in the present invention, the size of the preset regular square box is set to 224x224 pixels, and the related embodiments are also based on the size of the preset regular square box being set to 224x224 pixels, but this does not limit the size of the preset regular square box. If the portrait segmentation map is a square, it can be directly scaled proportionally to reach the size of the preset regular square box. However, in actual scenarios, most selfie images are in the 16:9 format. If it is determined that the portrait segmentation map is a rectangle, then the area that does not meet the preset regular square box is filled with an image (for example, the area that does not exceed the preset regular square box is uniformly filled with black), and the filled area is merged with the portrait segmentation map to obtain a filled image. That is, the size of the obtained filled image is also 224x224 pixels, and the filled image contains two parts: the filled area and the portrait segmentation map.
[0066] The filled image is divided n x n times, where n is a positive integer. The n x n times of division is more conducive to the analysis and processing of the encoder and decoder of the model compared to asymmetric division. For example, the size of the filled image is divided by 16 respectively, which is equivalent to 14x14 times of division. The filled image is non-overlappingly divided into 196 small blocks, each small block is 16x16 pixels, and each small block also contains 3 color channels of RGB. That is to say, the total number of pixels in each small block is (16x16x3 pixels). According to these 196 small blocks of images, a portrait filling segmentation map is obtained. According to the self-attention network, the 196 small blocks of the portrait filling segmentation map are sorted in the order from left to right and top to bottom in the portrait segmentation map and unfolded into a one-dimensional shape of 1xN 2 to generate a non-background mask (1, 196). If the 16x16 block is a filled area or a portrait area, the element with a value of 0 corresponds to the position in the non-background mask, otherwise it is 1. The non-background mask refers to masking the portrait part and the uniformly formatted filled part behind, so that the portrait part and the filled part do not participate in the calculation of processing parameters, reducing the procedures in the image processing process and facilitating the improvement of the processing efficiency of image processing. That is, in the portrait filling segmentation map, the area other than the non-background mask is the background area. The portrait filling segmentation map contains a background area and a non-background area, and the non-background area contains a portrait area and a filled area.
[0067] In order to solve the problem in the prior art that the proportion of the non-background area in the image to be processed is relatively large, and the network convolution kernel of the CNN model cannot accurately extract the background features of the image, the present invention divides the filled image into 196 non-overlapping small blocks, each small block being 16x16 pixels, and each small block also including 3 color channels of RGB, which can enable the network to distinguish the area represented by each pixel point and has a higher accuracy in extracting the features of the non-background area.
[0068] In one embodiment, the filling the portrait segmentation map to obtain a filled image includes:
[0069] Reading the numerical values of the side lengths of the portrait segmentation map and selecting the side with the largest numerical value;
[0070] Scaling the portrait segmentation map proportionally with the selected side with reference to a preset square box so that the side length of the selected side is equal to the side length of the preset square box;
[0071] Performing image filling on the area of the portrait segmentation map that does not exceed the preset square box to obtain the filled image.
[0072] Read the height and width of the portrait segmentation map and compare their sizes, select the side with the largest numerical value (for example, the unit of the numerical value is pixels). If the height of the rectangle is greater than the width, scale the portrait segmentation map proportionally so that the height of the rectangle is equal to the side length of the preset square box, and perform image filling on the area of the portrait segmentation map that does not exceed the preset square box (for example, the image filling uniformly fills the area that does not exceed the preset square box with black) to obtain the filled image, so that the height and width of the image to be processed reach 224x224 pixels respectively. In order to unify the format of each image to be processed when inputting the model to extract features, it can improve the model recognition efficiency and reduce the time for extracting features, and also improve the processing time for a large number of images to be processed.
[0073] Step S30: Performing matrix calculation on the non-background mask to obtain a non-background self-attention mask, generating a self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filling segmentation map, and marking attention labels for each pair of small blocks in the background area according to the non-background self-attention mask.
[0074] In this embodiment, after obtaining the non-background self-attention mask (1, 196, 196), the non-background self-attention mask judges the correlation between 196 small blocks of the portrait filling segmentation map, so as to determine which two small blocks need to be marked with attention labels and which two small blocks need to be marked with non-attention labels. For the two small blocks corresponding to the row and column where the element with a value of 1 in the non-background self-attention mask is located, the correlation needs to be concerned, and for the two small blocks corresponding to the row and column where the element with a value of 0 is located, the correlation does not need to be concerned.
[0075] In one embodiment, the matrix calculation of the non-background mask to obtain the non-background self-attention mask, and generating the self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filling segmentation map includes:
[0076] Extracting the features of each small block of the non-background mask for matrix calculation to obtain the non-background attention mask;
[0077] And extracting the features of each small block of the portrait filling segmentation map to generate the segmentation map matrix;
[0078] Reading each element of the mask matrix and each element of the segmentation map matrix for matrix multiplication to generate the self-attention matrix.
[0079] Extracting the features of each small block of the non-background mask for matrix calculation to obtain the non-background attention mask. The features of the non-background mask refer to which blocks are background regions and which blocks are non-background regions among 196 small blocks. Extracting the features (1, 196, 16x16x3 pixels) of the portrait filling segmentation map to generate the segmentation map matrix. The feature values of the portrait filling segmentation map refer to 196 small blocks in 1 dimension, and the pixels of each small block are 16x16x3. 3 represents that each small block also contains 3 color channels of RGB.
[0080] Obtaining the mask matrix according to the two-dimensional matrix array of the non-background self-attention mask, and reading each element of the mask matrix and each element of the segmentation map matrix for matrix multiplication to generate the self-attention matrix. Specifically, it includes: according to the calculation unit of the matrix, performing matrix calculation on the segmentation map matrix to establish the relationship between features and obtain the attention matrix. According to the dot product unit of the attention matrix, multiplying the values at the corresponding positions of the self-attention matrix and the mask matrix obtained from the segmentation map matrix, and inputting the product of the multiplied values at each corresponding position into the attention matrix to obtain the self-attention matrix (1, 196, 196 pixels). The role of this self-attention matrix is to use the inner product of the Query (giver) and Key (receiver) of the matrix divided by the square root of its dimension. Each small block uses Query to match the Key of the target small block as the attention, so as to generate attention to all small blocks.
[0081] In one embodiment, marking attention labels for each pair of small blocks in the background area according to the non-background self-attention mask includes:
[0082] Judging the relevance between each small block of the portrait filling segmentation map according to the non-background self-attention mask;
[0083] If it is judged that the values of the small blocks in pairs are both greater than or equal to a preset value, then the small blocks in pairs are used as the background area and marked as attention labels.
[0084] In one embodiment, after the step of if it is judged that the values of the small blocks in pairs are both greater than or equal to a preset value, then the small blocks in pairs are used as the background area and marked as attention labels, the method includes:
[0085] If it is judged that the values of the small blocks in pairs are both less than the preset value, then the small blocks in pairs are used as the non-background area and marked as non-attention labels.
[0086] Judging the relevance between all (196) small blocks according to the non-background self-attention mask. The rule for judging relevance is: the elements of the values of two small blocks will have the following four situations: (0, 0) or (0, 1) or (1, 1) or (1, 0). Among them, (0, 0) represents the portrait area and the portrait area (for example, an image to be processed contains multiple portraits), the portrait area and the filling area, and (0, 1) represents the portrait area and the background area, the filling area and the background area. The two areas refer to not distinguishing the adjacent relationship of the areas, which means that any two areas, whether far apart or close, will establish an association. The associations established between the non-background and the background area (0, 1) and the associations established by the background area (0, 0) are both marked as non-attention. The areas marked as non-attention labels do not need to be concerned, and the network does not need to learn them.
[0087] Step S40: Input the self-attention matrix and the segmentation map matrix into a preset self-attention model for matrix multiplication, and extract the features of the small blocks corresponding to all elements with a value of 1 in the portrait filling segmentation map from the result of the matrix multiplication according to the attention labels, as the background features of the image to be processed.
[0088] In this embodiment, the preset self-attention model includes multiple self-attention modules and an output fully connected layer. Each self-attention module includes a first feature network composed of three fully connected linear units, a first matrix calculation unit, a dot product unit, and a Softmax unit, and a second feature network composed of a second matrix calculation unit, a first fully connected unit, an activation unit, a second fully connected unit, and a normalization unit. After the self-attention matrix and the segmentation map matrix are respectively subjected to feature calculations through the first feature network and the second feature network, they are then input into the output fully connected layer for processing.
[0089] Taking the first self-attention module as an example, after the self-attention matrix and the portrait filling segmentation map are input into the first feature network of the first self-attention module, the features of the portrait filling segmentation map (1, 196, 16x16x3 pixels) are copied into three copies, and the features of each copy are (1, 196, 16x16x3 pixels). The first copy of the features is input into the first fully connected linear unit to change the feature dimension (such as compression and expansion), and the output features are (1, 196, 512). The second copy of the features is input into the second fully connected linear unit to change the feature dimension, and the output features are (1, 512, 196). The features (1, 196, 512) and the features (1, 512, 196) are input into the first matrix calculation unit for matrix operation to establish the relationship between the features and obtain the attention matrix (1, 196, 196). According to the dot product unit, the corresponding position elements of the attention matrix (1, 196, 196) and the self-attention matrix (1, 196, 196) are multiplied, and a matrix (1, 196, 196) with the same size as the input matrix is output.
[0090] In this embodiment, according to the attention label, the function of the Softmax unit is used to calculate the relative attention values at each position of the output matrix (1, 196, 196). Only the features of the small blocks corresponding to all elements with a value of 1 in the result of the matrix multiplication in the portrait filling segmentation map are calculated to obtain the first output result (1, 196, 512).
[0091] In other embodiments, according to the non-attention label, the function of the Softmax unit can also be used to calculate the relative attention values at each position of the output matrix (1, 196, 196). The features of the small blocks corresponding to all elements with a value of 0 in the result of the matrix multiplication in the portrait filling segmentation map are ignored, so as to ignore the attention values of the small blocks, realize shielding the position information of the non-background areas of the portrait filling segmentation map, and only calculate the features of the small blocks corresponding to the values of 1 to obtain the first output result (1, 196, 512). That is to say, the attention label or the non-attention label can be selected to extract or shield the corresponding areas.
[0092] In the second feature network of the first self-attention module, the result of inputting the first output result and the replicated third copy of the feature into the third fully-connected linear unit is input into the second matrix calculation unit pair to obtain the second output result (1, 196, 512). The second output result is then input into the first fully-connected unit, activation unit, second fully-connected unit, and normalization unit for processing to obtain the processed matrix (1, 196, 786) of the portrait filling segmentation map. Among them, the activation unit performs non-linear processing on the input second output result, and the normalization unit performs normalization operation on the input second output result to change the input distribution, making it easier for the network to learn, without changing the matrix size of the input second output result.
[0093] Finally, the processed matrix (1, 196, 786) of the portrait filling segmentation map and the attention matrix (1, 196, 196) output by the first self-attention module are input into the second self-attention module to the last self-attention module, and all processing steps of the first feature network and the second feature network of the first self-attention module are executed. Finally, the feature (1, 512) of the background region of the image to be processed is output through the output fully-connected layer.
[0094] Before using the above preset self-attention model to extract the features of the background region, the deep network structure of the self-attention model is built through a deep learning framework, and the deep learning network model is trained through the collected selfie image dataset until the self-attention model converges. The input of the self-attention model is the portrait filling segmentation map (1, 196, 786) after the preprocessing step, where 786 = 16x16x3 (pixel points), and the output is the feature of (1, 512) dimensions.
[0095] The technical concept of the present invention for extracting the background features of the image to be processed is as follows: The image to be processed is uniformly sized and formatted, and the image to be processed is filled, segmented, etc. to obtain a portrait filling segmentation map. Each small block of the portrait filling segmentation map is expanded into a one-dimensional shape to generate a non-background mask. By using the multiple small blocks obtained by segmenting the portrait filling segmentation map, the problem that the area of the portrait part in the image accounts for a relatively large proportion in the prior art is solved, enabling the preset self-attention model to effectively distinguish and analyze each pixel point area.
[0096] A segmentation map matrix is generated for the portrait filling segmentation map, a non-background self-attention mask is obtained by performing matrix calculation on the non-background mask, a self-attention matrix is generated based on the non-background self-attention mask and the segmentation map matrix, and according to the non-background self-attention mask, the small blocks of the background region are marked with attention labels in pairs. According to the attention labels, the position information of the non-background region of the portrait filling segmentation map is masked, so that the position information of the portrait part in the image to be processed cannot interfere with the feature extraction of the background region, improving the accuracy of the features of the background region of the image to be processed by the preset self-attention model.
[0097] Instead of using the CNN model in the prior art to extract features, the present invention avoids the problem that the CNN model is not accurate enough in the attention judgment of each small block. By mainly marking the background area and non-background area with labels, the preset self-attention model only extracts features from the areas with attention labels and shields the position information of the areas without attention labels.
[0098] Refer to Figure 2 As shown, it is a schematic diagram of the functional modules of the image background feature extraction device 100 of the present invention.
[0099] The image background feature extraction device 100 of the present invention can be installed in an electronic device. According to the functions achieved, the image background feature extraction device 100 may include a segmentation module 110, a segmentation module 20, a marking module 130, and an output module 140. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by the processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0100] In this embodiment, the functions of each module / unit are as follows:
[0101] The segmentation module 110: is used to segment the image to be processed by using a preset segmentation model to obtain the portrait mask of the image to be processed, and merge the portrait mask and the image to be processed to obtain a portrait segmentation map;
[0102] The sorting module 120: is used to fill the portrait segmentation map to obtain a filled image, segment the filled image to obtain a portrait filled segmentation map of n x n small blocks, and expand the n x n small blocks into a 1 x n 2 non-background mask, and the portrait filled segmentation map includes a background area and a non-background area;
[0103] The marking module 130: is used to perform matrix calculation on the non-background mask to obtain a non-background self-attention mask, generate a self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filled segmentation map, and mark attention labels for each pair of small blocks in the background area according to the non-background self-attention mask;
[0104] The output module 140: is used to input the self-attention matrix and the segmentation map matrix into a preset self-attention model for matrix multiplication, and extract the features of the small blocks corresponding to all elements with a value of 1 in the portrait filled segmentation map from the result of the matrix multiplication according to the attention labels as the background features of the image to be processed.
[0105] In one embodiment, the step of segmenting the image to be processed by using a preset segmentation model to obtain a portrait mask of the image to be processed includes:
[0106] Performing feature extraction on the image to be processed to obtain a feature map of the image to be processed;
[0107] Segmenting the portrait part of the feature map according to a preset portrait area selection box to obtain the portrait mask.
[0108] In one embodiment, the step of segmenting the portrait part of the feature map according to a preset portrait area selection box to obtain the portrait mask includes:
[0109] Performing upsampling on the feature map to expand the feature map to a preset resolution to obtain an expanded feature map;
[0110] According to the portrait area selection box, distinguishing each pixel of the expanded feature map to obtain the portrait part of the feature map and performing segmentation, and generating a portrait mask of the image to be processed according to the segmented portrait part.
[0111] In one embodiment, the step of filling the portrait segmentation map to obtain a filled image includes:
[0112] Reading the numerical values of the side lengths of each side of the portrait segmentation map, and selecting the side with the largest numerical value;
[0113] Scaling the portrait segmentation map proportionally with reference to a preset square box by using the selected side, so that the side length of the selected side is equal to the side length of the preset square box;
[0114] Performing image filling on the area of the portrait segmentation map that does not exceed the preset square box to obtain the filled image.
[0115] In one embodiment, the step of performing matrix calculation on the non-background mask to obtain a non-background self-attention mask, and generating a self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filled segmentation map includes:
[0116] Extracting the features of each small block of the non-background mask to perform matrix calculation to obtain a non-background attention mask;
[0117] And extracting the features of each small block of the portrait filled segmentation map to generate the segmentation map matrix;
[0118] Reading each element of the mask matrix and each element of the segmentation map matrix to perform matrix multiplication to generate the self-attention matrix.
[0119] In one embodiment, marking attention tags for each pair of small blocks in the background area according to the non-background self-attention mask includes:
[0120] Judging the relevance between each small block of the portrait filling segmentation map according to the non-background self-attention mask;
[0121] If it is judged that the values of the small blocks in each pair are both greater than or equal to a preset value, then the small blocks in each pair are used as the background area and marked with attention tags.
[0122] In one embodiment, after the step of if it is judged that the values of the small blocks in each pair are both greater than or equal to a preset value, then the small blocks in each pair are used as the background area and marked with attention tags, the method includes:
[0123] If it is judged that the values of the small blocks in each pair are both less than the preset value, then the small blocks in each pair are used as the non-background area and marked with non-attention tags.
[0124] Refer to Figure 3 As shown, it is a schematic diagram of a preferred embodiment of the electronic device 1 of the present invention.
[0125] The electronic device 1 includes but is not limited to: a memory 11, a processor 12, a display 13, and a network interface 14. The electronic device 1 is connected to the network through the network interface 14 to obtain raw data. Among them, the network can be an enterprise internal network (Intranet), the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, a call network, or other wireless or wired networks.
[0126] Among them, the memory 11 includes at least one type of readable medium, and the readable medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the hard disk or memory of the electronic device 1. In other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped with the electronic device 1. Of course, the memory 11 may also include both the internal storage unit of the electronic device 1 and its external storage device. In this embodiment, the memory 11 is generally used to store the operating system installed in the electronic device 1 and various application software, such as the program code of the image background feature extraction 10. In addition, the memory 11 can also be used to temporarily store various data that have been output or will be output.
[0127] In some embodiments, the processor 12 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 12 is generally used to control the overall operation of the electronic device 1, such as performing control and processing related to data interaction or communication. In this embodiment, the processor 12 is used to run the program code stored in the memory 11 or process data, such as running the program code of the image background feature extraction 10.
[0128] The display 13 may be referred to as a display screen or a display unit. In some embodiments, the display 13 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an organic light-emitting diode (OLED) toucher, etc. The display 13 is used to display the information processed in the electronic device 1 and to display a visual working interface, such as displaying the result of data statistics.
[0129] The network interface 14 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), and the network interface 14 is generally used to establish a communication connection between the electronic device 1 and other electronic devices.
[0130] Figure 3Only the electronic device 1 with components 11 - 14 and image background feature extraction 10 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0131] Optionally, the electronic device 1 may further include a user interface, and the user interface may include a display, an input unit such as a keyboard. Optionally, the user interface may further include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an organic light - emitting diode (OLED) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.
[0132] The electronic device 1 may further include a radio frequency (RF) circuit, a sensor, an audio circuit, etc., which will not be elaborated here.
[0133] In the above - mentioned embodiment, when the processor 12 executes the image background feature extraction 10 stored in the memory 11, the following steps may be implemented:
[0134] Segment the image to be processed by using a preset segmentation model to obtain a portrait mask of the image to be processed, and merge the portrait mask and the image to be processed to obtain a portrait segmentation map;
[0135] Fill the portrait segmentation map to obtain a filled image, segment the filled image to obtain a portrait filled segmentation map of n x n small blocks, and expand the n x n small blocks into a 1 x n 2 non - background mask. The portrait filled segmentation map includes a background area and a non - background area;
[0136] Perform matrix calculation on the non - background mask to obtain a non - background self - attention mask, generate a self - attention matrix according to the mask matrix corresponding to the non - background self - attention mask and the segmentation map matrix corresponding to the portrait filled segmentation map, and mark attention labels for each pair of small blocks in the background area according to the non - background self - attention mask;
[0137] Input the self - attention matrix and the segmentation map matrix into a preset self - attention model for matrix multiplication, and extract the features of the small blocks corresponding to all elements with a value of 1 in the portrait filled segmentation map from the result of the matrix multiplication according to the attention labels, as the background features of the image to be processed.
[0138] The storage device may be the memory 11 of the electronic device 1, or may be other storage devices communicatively connected to the electronic device 1.
[0139] For a detailed introduction to the above steps, please refer to the above Figure 2 Regarding the functional module diagram of the embodiment of the image background feature extraction device 100 and Figure 1 Regarding the flowchart description of the embodiment of the image background feature extraction method.
[0140] In addition, an embodiment of the present invention also proposes a computer-readable medium, which may be non-volatile or volatile. The computer-readable medium may be any one or any combination of a hard disk, a multimedia card, an SD card, a flash card, an SMC, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, and the like. The computer-readable medium includes a storage data area and a storage program area. The storage data area stores data created according to the use of the blockchain node, and the storage program area stores the image background feature 10. When the image background feature extraction 10 is executed by a processor, the following operations are implemented:
[0141] Use a preset segmentation model to segment the image to be processed to obtain a portrait mask of the image to be processed, and merge the portrait mask with the image to be processed to obtain a portrait segmentation map;
[0142] Fill the portrait segmentation map to obtain a filled image, segment the filled image to obtain a portrait filled segmentation map of n x n small blocks, and expand the n x n small blocks into a 1 x n 2 non-background mask. The portrait filled segmentation map includes a background area and a non-background area;
[0143] Perform matrix calculation on the non-background mask to obtain a non-background self-attention mask, generate a self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filled segmentation map, and mark attention labels for each pair of small blocks in the background area according to the non-background self-attention mask;
[0144] Input the self-attention matrix and the segmentation map matrix into a preset self-attention model for matrix multiplication, and extract the features of the small blocks corresponding to all elements with a value of 1 in the portrait filled segmentation map from the result of the matrix multiplication according to the attention label as the background feature of the image to be processed.
[0145] The specific implementation manner of the computer-readable medium of the present invention is substantially the same as the specific implementation manner of the above image background feature extraction method, and will not be elaborated here.
[0146] In another embodiment, for the image background feature extraction method provided by the present invention, to further ensure the privacy and security of all the data appearing above, all the above data can also be stored in a node of a blockchain. For example, attention tags and background features can all be stored in the blockchain node.
[0147] It should be noted that the blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.
[0148] It should be noted that the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments. And the term "including" or "comprising" or any other variant thereof in this article is intended to cover a non-exclusive inclusion, so that a process, device, article, or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or the method further includes elements inherent to such process, device, article, or method. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, device, article, or method including that element.
[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a medium as described above (such as ROM / RAM, magnetic disk, optical disc), including several instructions to enable a terminal device (which can be a mobile phone, computer, electronic device, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0150] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. An image background feature extraction method, characterized in that The method includes: Segmenting the image to be processed using a preset segmentation model to obtain a portrait mask of the image to be processed, and merging the portrait mask and the image to be processed to obtain a portrait segmentation map; Fill the portrait segmentation map to obtain a filled image, segment the filled image to obtain a portrait fill segmentation map of n x n small blocks, and expand the n x n small blocks into a 1 x n 2 non-background mask. The portrait fill segmentation map includes a background area and a non-background area; Performing matrix calculation on the non-background mask to obtain a non-background self-attention mask, generating a self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filling segmentation map, and marking attention labels for each pair of small blocks in the background area according to the non-background self-attention mask; Inputting the self-attention matrix and the segmentation map matrix into a preset self-attention model for matrix multiplication, and extracting the features of the small blocks corresponding to all elements with a value of 1 in the portrait filling segmentation map from the result of the matrix multiplication according to the attention labels as the background features of the image to be processed; Among them, performing matrix calculation on the non-background mask to obtain a non-background self-attention mask, and generating a self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filling segmentation map includes: extracting the features of each small block of the non-background mask for matrix calculation to obtain a non-background attention mask; and extracting the features of each small block of the portrait filling segmentation map to generate the segmentation map matrix; reading each element of the mask matrix and each element of the segmentation map matrix for matrix multiplication to generate the self-attention matrix; The marking attention labels for each pair of small blocks in the background area according to the non-background self-attention mask includes: judging the relevance between each small block of the portrait filling segmentation map according to the non-background self-attention mask; if it is judged that the values of the two small blocks in a pair are both greater than or equal to a preset value, then taking the two small blocks in a pair as the background area and marking them as attention labels.
2. The image background feature extraction method according to claim 1, wherein The segmenting the image to be processed using a preset segmentation model to obtain a portrait mask of the image to be processed includes: Performing feature extraction on the image to be processed to obtain a feature map of the image to be processed; Segmenting the portrait part of the feature map according to a preset portrait area selection box to obtain the portrait mask.
3. The image background feature extraction method according to claim 2, wherein The segmenting the portrait part of the feature map according to a preset portrait area selection box to obtain the portrait mask includes: Performing upsampling processing on the feature map to expand the feature map to a preset resolution to obtain an expanded feature map; According to the portrait area selection box, distinguishing each pixel of the expanded feature map to obtain the portrait part of the feature map and performing segmentation, and generating a portrait mask of the image to be processed according to the segmented portrait part.
4. The image background feature extraction method according to claim 1, wherein, The filling the portrait segmentation map to obtain a filled image includes: Reading the values of the side lengths of each side of the portrait segmentation map and selecting the side with the largest value; Scaling the portrait segmentation map proportionally with the selected side as a reference according to a preset square box so that the side length of the selected side is equal to the side length of the preset square box; Performing image filling on the area of the portrait segmentation map that does not exceed the preset square box to obtain the filled image.
5. The image background feature extraction method according to claim 1, wherein After determining that the values of the small blocks in pairs are both greater than or equal to the preset value, and taking the small blocks in pairs as the background area and marking them with the attention label, the method includes: If it is determined that the values of the small blocks in pairs are both less than the preset value, the small blocks in pairs are taken as the non-background area and marked with the non-attention label.
6. An image background feature extraction device for implementing the image background feature extraction method according to any one of claims 1 to 5, characterized in that, The device includes: A segmentation module: configured to segment the image to be processed by using a preset segmentation model to obtain a portrait mask of the image to be processed, and merge the portrait mask and the image to be processed to obtain a portrait segmentation map; Sorting module: used to fill the portrait segmentation map to obtain a filled image, segment the filled image to obtain a portrait fill segmentation map of n x n small blocks, and expand the n x n small blocks into 1 x n 2 non-background mask, and the portrait fill segmentation map includes a background area and a non-background area; A marking module: configured to perform matrix calculation on the non-background mask to obtain a non-background self-attention mask, generate a self-attention matrix according to the mask matrix corresponding to the non-background self-attention mask and the segmentation map matrix corresponding to the portrait filling segmentation map, and mark the attention label for the small blocks in pairs in the background area according to the non-background self-attention mask; An output module: configured to input the self-attention matrix and the segmentation map matrix into a preset self-attention model for matrix multiplication, and extract the features of the small blocks corresponding to all elements with a value of 1 in the portrait filling segmentation map from the result of the matrix multiplication according to the attention label as the background features of the image to be processed.
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a program executable by the at least one processor, and the program is executed by the at least one processor so that the at least one processor can execute the image background feature extraction method according to any one of claims 1 to 5.
8. A computer-readable medium, characterized in that, The computer-readable medium stores image background features, and when the image background features are executed by a processor, the image background feature extraction method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Method for improving pedestrian attribute recognition accuracy, terminal and medium
CN113221757A
Generating and ordering tags for an image using subgraph of concepts
US20200050895A1