A training data augmentation method based on global matting
Through the training data enhancement method based on global cutout, infrared images are cut out and guide filtered, which solves the problem that infrared images do not stand out in target objects under complex backgrounds, and achieves the effect of reducing the difficulty of feature extraction of neural networks and reducing the complexity of model.
Patent Information
- Application Number
- CN202210899071.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-28
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-07-28
AI Technical Summary
Existing infrared images do not stand out in complex backgrounds, and neural networks are difficult to learn effective features.
The training data enhancement method based on global cutout is adopted, and the original image is cut out by sampling cutout method, and the transparency mask obtained from cutout is filtered to generate the enhanced training data.
It reduces the difficulty of neural networks to extract effective features, highlights the details of the target object and filters irrelevant backgrounds, thereby reducing model complexity and calculation time.
Smart Images

Figure CN115359242B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image training data enhancement, and specifically to a training data enhancement method based on global matting. Background Art
[0002] In the era of big data, every piece of image data is hard-won, and rich and diverse data has increasingly become an intangible wealth. Compared with visible light, infrared rays have stronger adaptability and anti-interference ability, so infrared imaging technology has a wide range of applications in many fields. However, the infrared radiation transmission characteristics in complex environments cause problems such as low signal-to-noise ratio, weak detail blurring, and low contrast in infrared images. Image matting is a digital image processing technology that separates the part of interest (i.e., the foreground part) from other parts of a digital image. The synthesis formula of a digital image is described as: C = αF + (1 - α)B, where C, F, and B represent the synthesized image, the foreground image, and the background image respectively. Currently, matting algorithms are mainly divided into three categories: sampling-based methods, propagation-based methods, and deep learning-based methods.
[0003] In the past, when solving the problem that the target object in an infrared image is not prominent in a complex background, a deeper neural network was usually used for processing, which brought higher computational complexity at the same time. Suppressing the cluttered background can improve the network performance, but this aspect of work has rarely been emphasized before. Summary of the Invention
[0004] The purpose of the present invention is to overcome the problems that the target object in the existing infrared image is not prominent and it is difficult for the neural network to learn, and to provide a training data enhancement method based on global matting. The present invention enhances the part of interest in the original image by matting, thereby reducing the difficulty for the neural network to extract effective features and realizing the enhancement of training data. Two parts are required to implement this solution. One is to perform matting on the original image according to the specified trimap by using the sampling matting method, and the other is to perform guided filtering on the transparency mask obtained by matting.
[0005] The present invention is realized by the following technical solutions:
[0006] A training data enhancement method based on global matting includes the following steps:
[0007] (a) The algorithm takes a picture as input, scales the input picture proportionally, and scales the picture so that the length of the long side is 64 pixels.
[0008] (b) According to the input trimap, perform edge sampling on the pixels in the foreground and background regions of the image;
[0009] (c) Traverse each unknown pixel in the image according to the input tripartite graph. For each unknown pixel, calculate the evaluation function value with all sampled pixel pairs, and select the pixel pair with the optimal evaluation function value;
[0010] (d) Each unknown pixel calculates the transparency mask value according to the optimal pixel pair;
[0011] (e) Use the original image to perform guided filtering on the obtained transparency mask image.
[0012] (f) Output the enhanced image.
[0013] In step (b) of the above training data enhancement method based on global matting, it is mainly divided into two steps:
[0014] (b-1) Traverse the foreground pixels one by one. If there are unknown pixels in the four-neighborhood of the pixel, add it to the foreground sampling subset;
[0015] (b-2) Traverse the background pixels one by one. If there are unknown pixels in the four-neighborhood of the pixel, add it to the background sampling subset.
[0016] In step (c) of the above training data enhancement method based on global matting, it is mainly divided into two steps:
[0017] (c-1) If the original image is a color image, the calculated evaluation function includes a color criterion and a spatial proximity criterion. The specific calculation formula is: Where:
[0018]
[0019]
[0020]
[0021] k is the kth unknown pixel, and i and j are the corresponding ith foreground pixel and jth background pixel, C j (B) are the colors of the kth unknown pixel, its corresponding ith foreground pixel, and jth background pixel respectively, S j (B) are the coordinates of the kth unknown pixel, the ith foreground pixel, and its corresponding jth background pixel respectively, and σ c 、σ s are the penalty factors of the color criterion and the spatial proximity criterion respectively;
[0022] (c-2) If the original image is a grayscale image, the calculated evaluation function only includes the spatial proximity criterion. The specific calculation formula is: The definition is shown in Formulas (2) and (3).
[0023] In step (c-1) of the above training data augmentation method based on global matting, the penalty factor σ c = 0.5, σ s = 0.5. That is, the proportion of the color criterion and the spatial criterion in the evaluation function is 1:1.
[0024] In step (d) of the above training data augmentation method based on global matting, the formula for calculating the transparency mask value of each unknown pixel according to the optimal pixel pair is where α z is the transparency value of this unknown pixel, I z is the color of this unknown pixel, F z is the color of the foreground pixel in the optimal pixel pair corresponding to this unknown pixel, and B z is the color of the background pixel in the optimal pixel pair corresponding to this unknown pixel.
[0025] In step (e) of the above training data augmentation method based on global matting, the original image is used as the guidance image, the transparency mask value is used as the filtering input, and the guided filter is used to output the filtered image. Guided filtering means that for an input image p, through the guidance image I, the output image q is obtained after filtering, where both p and I are inputs of the algorithm. Guided filtering defines a linear filtering process. For the pixel point at position i, the obtained filtering output is a weighted average: q i = ∑ j W ij (I)p j , where i and j respectively represent the pixel subscripts, and W ij is the filter kernel related only to the guidance image I.
[0026] The training data augmentation method based on global matting provided by the present invention first uses the sampling matting algorithm to perform matting augmentation operations on the input image according to the given tripartite graph to obtain the transparency mask map of the original image, and then uses guided filtering to perform filtering operations on the transparency mask map with the original image as the guidance image to obtain the final augmented training data. The augmented training pictures have more prominent features than the original pictures, can highlight the details of the target object while filtering out the irrelevant background, so that deep learning methods such as neural networks can learn the features of the object without a complex network structure, which is beneficial to reducing the complexity and calculation time of the model.
[0027] Compared with the neural network directly trained using the original data, the present invention has the following advantages and technical effects:
[0028] The present invention uses a sampling-based matting method, which to a certain extent avoids the problem of inability to perform matting due to poor quality of the trimap, and has a certain robustness for the matting problem of infrared images in complex environments; by using different evaluation functions for color images and grayscale images, it not only ensures the accuracy of matting under color images, but also ensures that the algorithm speed can be improved when the image is a grayscale (single-channel) image; by using a guided filter to filter the matting result after matting, the matting result can be made smoother while ensuring the edges. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a flowchart of the training data enhancement method based on global matting of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0030] The present invention will be further described in detail below with reference to specific embodiments.
[0031] As Figure 1 shown, the present invention discloses a training data enhancement method based on global matting; the main process includes the following steps:
[0032] In the first step, read in the training image and preprocess the image. Specifically, scale the image proportionally to a long side of 64, determine whether the image type is a color image or a grayscale image, and convert the image type to ubyte8. Prepare for subsequent operations.
[0033] The algorithm reads in the training image and performs a scaling operation on the image. The purpose is to scale the image to a unified specification, which is beneficial for subsequent algorithms to process and can ensure the stability of the algorithm speed.
[0034] In the second step, sample the foreground pixels and background pixels. The purpose of this step is to reduce the search space.
[0035] Perform edge sampling on the foreground pixels and background pixels respectively: for the foreground pixels, traverse each foreground pixel one by one. If there are unknown pixels within the four-neighborhood range of the pixel, add it to the foreground sample set, and use the same method to sample the background pixels. Finally, obtain the sampled subsets of the foreground pixels and background pixels.
[0036] If the foreground sampled subset is empty, then use all foreground pixels as the sampled subset. If it is still empty, then randomly select one-tenth of the total number of pixels as the foreground sampled subset. The background pixels are processed in the same way.
[0037] In the third step, perform matting enhancement on the image. The purpose of this step is to obtain preliminary enhanced training data.
[0038] First, determine whether the image type is a color image or a grayscale image. If it is a color image, the evaluation function is If it is a grayscale image, the evaluation function is
[0039] Subsequently, process each unknown pixel one by one. For each unknown pixel, calculate the evaluation function values of all sampled foreground and background pixel pairs. According to the pixel pair with the optimal evaluation function value, calculate the transparency mask value at this unknown pixel according to where α z is the transparency value of this unknown pixel, I z is the color of this unknown pixel, F z is the color of the foreground pixel in the optimal pixel pair corresponding to this unknown pixel, and B z is the color of the background pixel in the optimal pixel pair corresponding to this unknown pixel.
[0040] In the fourth step, perform a guided filtering operation on the transparency mask image obtained by matting using the original image as the guidance image. The purpose of this step is to smooth the transparency mask image while maintaining the image edges.
[0041] Obtain the sum of the transparency mask images of matting. Take the transparency mask image p as the input and the original image as the guidance image I. After guided filtering, obtain the output image q. Here, both p and I are inputs of the algorithm. Guided filtering defines a linear filtering process. For the pixel at position i, the obtained filtering output is a weighted average: q i = ∑ j W ij (I)p j , where i and j respectively represent the pixel subscripts, and W ij is a filter kernel that only depends on the guidance image I.
[0042] As described above, the present invention enhances the training data by matting and efficiently solves the matting problem through sampling. The present invention is simple to use, has a small amount of calculation, has no special requirements for the input training pictures, and can quickly and effectively perform data enhancement on the training pictures.
[0043] The implementation manners of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention.
Claims
1. A training data augmentation method based on global matting, characterized in that, it includes the following steps: (a) The algorithm reads in the data to be trained. First, the long side of the image is scaled to a ratio of 64, and the image is scaled proportionally. (b) According to the input trimap, edge sampling is performed on the pixels in the foreground and background regions of the image. (c) According to the input trimap, each unknown pixel in the image is traversed. For each unknown pixel, the evaluation function values of all sampled pixel pairs are calculated, and the pixel pair with the optimal evaluation function value is selected. (d) Each unknown pixel calculates the transparency mask value according to the optimal pixel pair. (e) The original image is used to perform guided filtering on the obtained transparency mask image. (f) The enhanced image is output. Step (c) includes the following sub-steps: (c-1) If the original image is a color image, the calculated evaluation function includes a color criterion and a spatial proximity criterion. The specific calculation formula is: k is the k-th unknown pixel, and i, j are its corresponding i-th and j-th foreground and background pixels, The colors of the k-th unknown pixel, its corresponding i-th foreground pixel, and j-th background pixel respectively, are the coordinates of the k-th unknown pixel, the i-th foreground pixel, and its corresponding j-th background pixel respectively, σ c and σ s are the penalty factors of the color criterion and the spatial proximity criterion respectively; (c-2) If the original image is a grayscale image, the calculated evaluation function only includes a spatial proximity criterion. The specific calculation formula is: In step (d), the formula for calculating the transparency mask value of each unknown pixel based on the optimal pixel pair is as follows: where α z is the transparency value of this unknown pixel, I z is the color of this unknown pixel, F z is the color of the foreground pixel in the optimal pixel pair corresponding to this unknown pixel, and B z is the color of the background pixel in the optimal pixel pair corresponding to this unknown pixel.
2. The training data augmentation method based on global matting according to claim 1, characterized in that, step (b) includes the following sub-steps: (b-1) Each foreground pixel is traversed one by one. If there is an unknown pixel in the four-neighborhood of the pixel, it is added to the foreground sampling subset. (b-2) Each background pixel is traversed one by one. If there is an unknown pixel in the four-neighborhood of the pixel, it is added to the background sampling subset.
3. The training data augmentation method based on global matting according to claim 1, characterized in that, in step (e), the original image is used as the guidance image, the transparency mask value is used as the filtering input, and the filtered image is output using a guided filter.
4. The training data augmentation method based on global matting according to claim 3, characterized in that, the guided filtering means that for an input image p, through a guidance image I, an output image q is obtained after filtering, where both p and I are inputs of the algorithm.
Citation Information
Patent Citations
Method for extracting a clean foreground from a sectional drawing task and a model training method
CN109829925A
Full-automatic matting method based on non-local attention mechanism
CN113012169A