An Otsu thresholding method with a mask
By introducing mask design into the Otsu algorithm, only the region of interest is processed for maximum inter-class variance calculation, the problem of poor thresholding effect in complex image processing is solved, and more efficient and accurate image segmentation is achieved.
Patent Information
- Application Number
- CN202110991862.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-08-27
AI Technical Summary
When the existing Otsu algorithm processes complex targets and images with more interference, the thresholding operation effect is not ideal and is sensitive to background interference, resulting in inaccurate segmentation results.
The masked Otsu thresholding method is used to mask and penetrate the image through personalized mask design, and the maximum inter-class variance is calculated based on the mask area as the object to determine the threshold.
It improves the accuracy and efficiency of image segmentation, reduces the impact of background interference on segmentation results, simplifies the thresholding algorithm for complex grayscale distribution characteristics, and improves the shortcomings of the existing Otsu algorithm.
Smart Images

Figure CN113870182B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital image processing in machine vision, and particularly relates to an Otsu thresholding method with a masking function. Background Art
[0002] With the rapid development of technologies such as machine vision, computer vision, and artificial intelligence, digital image processing technology has been applied more and more widely, such as in the fields of industrial non-contact detection, medical image processing, remote sensing image processing, military security, intelligent transportation, and smart home. In the process of digital image processing, since the accuracy of the segmentation result plays a key role in the processing result, it is of utmost importance to segment real and effective target information from the image. Common traditional image segmentation methods include region segmentation, threshold segmentation, and edge detection, etc. Among them, threshold segmentation has the widest application range because of its simple operation and fast calculation speed.
[0003] In threshold segmentation, the determination of the threshold is particularly important. An inappropriate threshold selection will lead to inaccurate target segmentation, resulting in over-segmentation, that is, the same object is thresholded and segmented into two or more parts, and under-segmentation, that is, two or more objects are glued together to form a whole after thresholding. Currently, the methods for determining the threshold mainly include fixed threshold method, local threshold method, and global threshold method.
[0004] The fixed threshold method performs thresholding operations on the image with a specific threshold. Its advantages are convenience and speed, while its disadvantages are that the threshold selection is affected by human subjective consciousness, it is difficult to evaluate the rationality, and the algorithm robustness is poor.
[0005] The local threshold method divides an image into multiple sub-images, and correspondingly determines the threshold for each sub-image, or divides the image into regular equal-sized regions for adaptive threshold calculation. The advantage of the local threshold is that it can reduce the influence of uneven illumination and background interference, etc. The disadvantages are that the two complex operations of image splitting and merging after thresholding increase the computational complexity of the algorithm, and it has poor adaptability to targets with a relatively large area ratio and irregular shapes in the image.
[0006] The global threshold method first calculates the threshold based on the overall gray information of the image using different mathematical theories, and then performs thresholding operations based on this threshold. The calculation method of the threshold is the key to the algorithm. Common threshold calculation theories include minimum error, maximum between-class variance, and maximum entropy, etc. The Otsu global threshold algorithm adopts the maximum between-class variance theory and is widely used in digital image processing because of its simple principle and convenient operation, etc.
[0007] The Otsu algorithm is applied to digital images. Limited by the image acquisition method and the two-dimensional matrix representation form, it only processes the gray values in the image or within a rectangular area in the image. At the same time, since the algorithm uses the maximum between-class variance value as the measure of the threshold calculation result, it usually produces a good thresholding effect on the gray set with the "hump" distribution characteristic. However, in practical applications, the geometric shape and gray distribution of the target object are often relatively complex, and it is difficult to obtain the "hump" distribution characteristic in the whole image or rectangular area. Moreover, there are often many interferences in the image background. Therefore, the thresholding operation effect of the algorithm is often not ideal.
[0008] Chinese Patent Application CN201811294422.9 discloses a method for using a mask to extract the image of the region of interest and perform Otsu thresholding, including making an annular mask using the inner and outer diameter parameters of the steel pipe obtained by least square fitting; performing an intersection operation between the annular mask and the original image to extract the image of the region of interest that only contains the end face and does not contain the chamfer; performing Otsu thresholding operation on the extracted image of the region of interest to obtain a binary image. However, the purpose of setting the mask is to extract the annular region of interest and set the pixel values of the image background outside the annulus to 0. After extraction, the Otsu calculation is still performed on all the pixels in the whole image. It only uses the Otsu algorithm without any improvement to the algorithm. The distribution of 0 pixel values in the background is still objectively included in the calculation process of the Otsu algorithm. Although good results are obtained in the implementation cases, after analysis, it is because the gray values of the defects in the implementation cases are quite different from the 0 values, and the 0 value distribution does not have a great impact on the segmentation result. If the gray values of the defect part are close to the 0 value, this method will cause serious missegmentation phenomena. To sum up, it is necessary to improve and innovate the existing Otsu algorithm. Summary of the Invention
[0009] The purpose of the present invention is to provide an Otsu thresholding method with a mask. First, the image is masked and penetrated through personalized mask design, then the target area is extracted using the Hadamard product, and finally, the threshold is determined only for the mask area using the maximum between-class variance calculation theory. The principle is simple and the practicability is good.
[0010] To achieve the above purpose, the solution of the present invention is:
[0011] An Otsu thresholding method with a mask includes the following steps:
[0012] Step 1, determine the number of masks f m and the number of penetrated areas r within the mask m , and make an equal-sized template M according to the original image S m×n m×n, and in the shielding area M in the template 0 and the penetration area M 1 are respectively filled with the grayscale shielding factor w 0 and the penetration factor w 1 , where w 0 ∈[0,1], w 1 ∈[0,1], and w 0 and w 1 are not both 0 or 1 at the same time;
[0013] Step 2, take the Hadamard product of the image S m×n and the mask M m×n as the extracted target area;
[0014] Step 3, calculate the maximum between-class variance in the target area to obtain the threshold for the thresholding operation of the target area;
[0015] Step 4, perform threshold binary processing with the threshold obtained in Step 3.
[0016] In the above Step 1, if the original image is a RectN image, according to the grayscale of the image including black-gray, dark-gray, and off-white, and the target object is R off-white graphics located in the dark-gray image block, set f m =1, r m =R; use the Otsu algorithm to obtain the binary result of the black-gray and dark-gray parts, and this result is used as the template M of the mask after one image closing operation. Fill w 1 =1 in its white area and fill w 0 =0 in the black area.
[0017] In the above Step 1, if the original image is a Branch image, according to the image including a dark-gray background and an off-white foreground, and the grayscale distribution shows a "hump" feature; set f m =1, r m =1; use the edge detection result as the template M, fill w 1 =1 inside the edge and fill w 0 =0 in the rest, and perform dilation on the filling result as the mask design result.
[0018] In the above Step 1, if the original image is a Circle image, first determine the number of masks f m according to the nature of the target object and the image processing task, and the number of penetration areas r m in each mask takes 1; use the Sobel operator to solve the gradient, then binaryize the gradient value, and finally use the Hough circle detection to extract each circle respectively. Fill w 1 =1 in the intersecting part of the circles and fill w 0 =0 in the rest.
[0019] In the above step 2, the Hadamard product is expressed as follows:
[0020]
[0021]
[0022]
[0023] where p (i,j) represents the gray value of the i-th row and j-th column of the image.
[0024] The specific process of the above step 3 is as follows:
[0025] Step 31, calculate the number and probability distribution of each gray-level pixel;
[0026] Step 32, divide the gray-scale set of the target area into two categories C 0 and C 1 with the parameter t as the threshold, where t ∈ [0, 255];
[0027] Step 33, calculate the within-class probabilities of the two categories C 0 and C 1 respectively, and then calculate the gray-scale means of C 0 and C 1 ;
[0028] Step 34, calculate the within-class variances of the two categories C 0 and C 1 respectively, as well as the between-class variance between the two;
[0029] Step 35, take t' as the threshold, where t' ∈ [0, t) ∪ (t, 255], and repeat steps 32 - 34 until the maximum between-class variance appears, and its corresponding value t is the required threshold T; if f m = 1, the threshold solving ends; if f m > 1, the thresholds need to be obtained separately for the remaining masks.
[0030] In the above step 4, if f m = 1, the image threshold segmentation is completed; if f m > 1, the remaining target areas are binarized separately, and then the intersection and union operations are performed on each binary image, that is, the image threshold segmentation is completed.
[0031] At the theoretical level, the maximum inter-class variance theory can be applied to the binary classification task of any numerical set. Therefore, the present invention proposes a maximum inter-class variance thresholding algorithm with a mask, which is called the MaskOtsu algorithm. During the calculation of the maximum inter-class variance of the image grayscale, the MaskOtsu algorithm masks and penetrates the processed image by using masks (Mask) of any shape and quantity, so that it can only process the area that the user is most interested in, can achieve the thresholding operation of any shaped area, reduce the influence of the background grayscale distribution on the peak value of the foreground grayscale distribution, simplify the complexity of the thresholding algorithm whose grayscale distribution presents a multi-peak distribution characteristic, thereby improving the deficiencies of the existing Otsu algorithm; and by using the mask, the influence of the background pixels on the calculation of the maximum inter-class variance can also be removed, greatly reducing the computational amount of the algorithm for threshold search in the grayscale exhaustive space, and improving the efficiency and accuracy of the algorithm.
[0032] After adopting the above scheme, the principle of the present invention is simple. Since the "attention" of solving the maximum inter-class variance is concentrated on the area where the target object is located, the performance of the algorithm such as anti-interference ability, adaptability, accuracy, and running efficiency can be improved. Brief Description of the Drawings
[0033] Figure 1 is the flowchart of the present invention;
[0034] Figure 2 is the factor filling schematic diagram in the mask design;
[0035] Among them, (a) is the template M m×n schematic diagram, (b) is w 0 and w 1 factor filling schematic diagram;
[0036] Figure 3 is the mask design schematic diagram of Experiment 1;
[0037] Among them, (a) is the experimental material RectN image, (b) is the mask design result;
[0038] Figure 4 is the mask design schematic diagram of Experiment 2;
[0039] Among them, (a) is the experimental material Branch image, (b) is the mask design result;
[0040] Figure 5 is the mask design schematic diagram of Experiment 3;
[0041] Among them, (a) is the experimental material Circle image, (b) is the grayscale distribution histogram;
[0042] Figure 6 is Figure 5Schematic diagram of the design results of three masks;
[0043] Figure 7 It is a schematic diagram of the result of applying the present invention;
[0044] Among them, (a) is the result of Experiment 1, (b) is the result of Experiment 2, and (c) is the result of Experiment 3. Detailed implementation manners
[0045] Hereinafter, the technical solutions and beneficial effects of the present invention will be described in detail with reference to the accompanying drawings.
[0046] As Figure 1 shown, the present invention provides a Mask Otsu method, which combines a mask in the application of the Otsu method. We call it the MaskOtsu algorithm. The core idea is to selectively shield and penetrate the image by using the mask, so that the algorithm only counts the pixel values inside the mask when calculating the between-class variance, and the pixel values outside the mask are no longer input into the algorithm, so as to achieve the purpose of divide and conquer in the thresholding process. The present invention specifically includes the following steps:
[0047] Step 1: Design the mask.
[0048] First, determine the number of masks according to the properties of the target object and the image processing task, denoted as f m ; then determine the number of penetration regions inside the mask according to the number of target objects, denoted as r m ; then make an equal-sized template M according to the size (m×n) of the original image m×n ; finally, fill the shielding region M 0 and the penetration region M 1 with the gray-scale shielding factor w 0 (w 0 ∈[0,1]) and the penetration factor w 1 (w 1 ∈[0,1]). When w 0 is 0, it means that the gray scale is completely shielded, and when w 0 is 1, it means no shielding; when w 1 is 0, it means that the transmittance of the gray scale is 0, and when w 1 is 1, it means the transmittance is 100%. Figure 2 It is a schematic diagram for filling the shielding factor and the penetration factor.
[0049] To illustrate the specific design method of the mask, the mask design methods in the cases of regular regions, irregular regions, and gray-scale multimodal distribution characteristics are respectively introduced in the following three experiments.
[0050] Experiment 1: Image segmentation and positioning
[0051] Experimental purpose: Use the MaskOtsu algorithm to segment 5 small off-white squares that are on a black-gray background and embedded in a dark-gray image block; and use the connected component analysis method to locate their positions to verify the accuracy of the algorithm segmentation. The experimental material image is as Figure 3 shown in (a), and the mask design result is as Figure 3 shown in (b).
[0052] The mask design steps are as follows:
[0053] (1). Image feature analysis. From the RectN image, it can be seen that its gray levels can be divided into three categories: black-gray, dark-gray, and off-white. The target objects are mainly located in the dark-gray image block. Therefore, the black-gray part is the interference part in the thresholding operation, and a mask needs to be designed to shield it.
[0054] (2). Target object attribute analysis. The gray level distributions of the target objects are the same, and the local backgrounds they are in are the same. Therefore, a single mask is made, that is, f m = 1.
[0055] (3). Target object position distribution. The target objects are distributed at 5 different positions in the image. Therefore, the number of penetration regions r m = 5.
[0056] (4). Mask production. The target objects are small and have a low probability density in the image. Therefore, the existing Otsu algorithm is used to obtain the binary result of the black-gray and dark-gray parts of the image. This result is used as the template M of the mask after one image closing operation. Fill w 1 = 1 in its white area and fill w 0 = 0 in its black area. The design result is as Figure 3 shown in (b).
[0057] Experiment 2: Image segmentation and target object extraction
[0058] Experimental purpose: Use the MaskOtsu algorithm to segment and extract irregular curves from an image with strong interference background. The experimental material image is as Figure 4 shown in (a), and the mask design result is as follows Figure 4 shown in (b).
[0059] The analysis idea of the mask design is the same as that of Experiment 1. From the analysis of the Branch image characteristics, it can be seen that the image is mainly composed of a dark-gray background and an off-white foreground, and the gray level distribution shows a "hump" feature. However, due to the presence of noise interference that is difficult to suppress in the dark-gray background, this part needs to be shielded during thresholding; the number of masks designed f m = 1, the number of penetration regions r m = 1; since the target edges are relatively obvious, the edge detection result is used as the template M, and w is filled inside the edge1 = 1, and the rest is filled with w 0 = 0, and the dilated result is used as the mask design result. Experiments show that a 1.5-fold dilation can completely cover the target area with the penetration factor and introduce less background interference.
[0060] Experiment 3: Image Segmentation and Geometric Feature Calculation
[0061] Experiment purpose: Use the MaskOtsu algorithm to segment the intersecting parts of the three circles pairwise and calculate their areas with the help of Blob analysis to verify the applicability of the algorithm to the feature of multi-modal gray distribution of the image. The experimental material image is as Figure 5 (a) shown, and its gray distribution histogram is as Figure 5 (b) shown.
[0062] The mask design analysis idea is the same as above. It can be seen from the gray distribution histogram of Circle that the image gray presents a multi-modal distribution feature. If better thresholding effect is to be obtained, multiple masks need to be added. The gray distribution of the target object is adjusted to a bimodal distribution through the action of different masks. Since the target object presents three kinds of gray levels, the number of masks f m is taken as 3, and the number of penetration areas r m in each mask is taken as 1. Since there are large gray level mutations at the edges of each circle, the Sobel operator is first used to solve the gradient, then the gradient value is binarized, and finally the Hough circle detection is used to extract the three circles respectively. In the intersecting parts of the circles, fill w 1 = 1, and the rest is filled with w 0 = 0, and the design results of the three masks are as Figure 6 shown.
[0063] Step 2, extract the target area from the original image.
[0064] The image and the mask are respectively represented as two-dimensional matrices S m×n and M m×n , and p (i,j) represents the gray value of the i-th row and j-th column of the image. The two matrices are equal in the number of rows and columns respectively, and the extraction operation can be represented by the multiplication of the corresponding elements of the matrices. Therefore, the Hadamard product is used to represent as follows:
[0065]
[0066]
[0067]
[0068] When w 0 = 0, w 1When \(w = 0\), the Hadamard product is an \(m\times n\) zero matrix, that is, the extraction result is an image with a gray value of 0.
[0069] When \(w\) 0 = 1, \(w\) 1 = 1, the Hadamard product is still the original image, that is, an invalid mask.
[0070] When \(w\) 0 = 0, \(w\) 1 ≠0, the Hadamard product is the product of the original gray value and \(w\) 1 in the region where the penetration factor acts, and 0 in other regions. The extraction result is the region where the penetration factor acts.
[0071] When \(w\) 0 ≠0, \(w\) 1 = 0, the result at this time is the reverse extraction when \(w\) 0 = 0, \(w\) 1 ≠0.
[0072] When \(w\) 0 ≠0, \(w\) 1 ≠0, the entire image is extracted at this time, but the gray value has changed according to the mask design result. This situation belongs to a valid mask.
[0073] Through the above analysis, to make the mask design result have practical significance, it is required that \(w\) 0 and \(w\) 1 are not both 0 or 1 at the same time.
[0074] Step 3, calculation of the maximum between-class variance under the mask. The calculation process is divided into the following steps:
[0075] (1). Calculate the number of pixels and probability distribution of each gray level. Let the number of gray levels after the Hadamard product of the original image and the mask be \(L\) m , the number of pixels of the \(i\)-th gray level be \(n\) mi , and the total number of pixels The shielding factor and the number of penetration factors can be calculated from the mask design result, denoted as and and The three satisfy the relationship:
[0076] When \(w\) 0 = 0, \(w\) 1 ≠0, the probability distribution of the \(i\)-th gray level can be expressed as:
[0077]
[0078] When \(w\) 0 ≠0, \(w\) 1When w = 0, the i-th order gray-level probability distribution can be expressed as:
[0079]
[0080] When w 0 ≠ 0 and w 1 ≠ 0, the i-th order gray-level probability distribution can be expressed as:
[0081]
[0082] (2). Calculate the within-class probability distribution. Let the parameter t (t ∈ [0, 255]). Divide the gray-level set of the extracted target region into C 0 and C 1 into two classes. Among them, C 0 represents the gray-level set with gray-level values [0,..., t], and C 1 represents the gray-level set with gray-level values [t + 1,..., L m - 1]. Then the within-class probability distributions are respectively:
[0083]
[0084] (3). Calculate the within-class gray-level mean. The gray-level means of classes C 0 and C 1 are calculated according to the following formula:
[0085]
[0086] (4). Calculate the within-class variance. The calculation formula is as follows:
[0087]
[0088] (5). Calculate the between-class variance. The calculation formula is as follows:
[0089]
[0090] Step 4, determine the threshold. Take another t' as the threshold, where t' ∈ [0, t) ∪ (t, 255]. Repeat steps (2) - (5) in step 3 until the maximum between-class variance appears, and its corresponding value t is the required threshold T. If f m = 1, the threshold solving ends; if f m > 1, the thresholds need to be obtained for the remaining masks respectively.
[0091] Step 5, threshold the target region image. Use the threshold T obtained in step 4 as the threshold for the thresholding operation of the target region for threshold binary processing. Specifically, the thresholding function cv::thershold() or other thresholding methods can be selected. If f m = 1, the algorithm ends here; if fm If it is >1, then perform binaryzation processing on the remaining target areas respectively, and then perform simple intersection and union operations on each binary image to complete image threshold segmentation.
[0092] The experimental results can be referred to as Figure 7 shown. By using the method provided by the present invention, the off-white square target was accurately segmented in Experiment 1; in Experiment 2, the irregular curve was successfully extracted. Although there are still sporadic interferences in the results, they can be removed through further simple analysis and processing; in Experiment 3, multiple masks were used, and the divide-and-conquer idea was adopted to simplify complex objects into simple objects, achieving remarkable results.
[0093] Parameter w 0 and w 1 can be selected according to requirements. According to existing experience, under normal circumstances, w 0 =0 and w 1 =1 are preferably selected.
[0094] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention fall within the protection scope of the present invention.
Claims
1. A Otsu thresholding method with a mask, characterized in that it includes the following steps: Step 1: Determine the number of masks f m and the number of penetration areas in the mask r m , according to the original image S m×n Make the same size template M m×n , and in the masked area M of the template 0 and the penetration area M 1 Fill in the grayscale masking factor w respectively 0 and permeability factor w 1 , where w 0 ∈[0,1],w 1 ∈[0,1], and w 0 and w 1 Not 0 or 1 at the same time; In the first step, if the original image is a RectN image, since the gray scale of this image includes black - gray, dark - gray, and off - white, and the target object is R off - white graphics located in the dark - gray image blocks, set f m = 1, r m = R; Use the Otsu algorithm to obtain the binarization result of the black - gray and dark - gray parts. After one image closing operation on this result, it is used as the template M of the mask. Fill the white area with w 1 = 1, and fill the black area with w 0 = 0; In the first step, if the original image is a Branch image, according to the image containing a dark gray background and an off-white foreground, the gray-scale distribution has a "hump" feature; set f m = 1, r m = 1; use the edge detection result as the template M, fill w 1 = 1 inside the edge, and fill w 0 = 0 for the rest, and use the filling result for dilation as the mask design result; In the first step, if the original image is a Circle image, first determine the number of masks \(f\) according to the properties of the target object and the image processing task m , the number of penetration regions \(r\) in each mask m Take 1; use the Sobel operator to solve the gradient, then binarize the gradient values, and finally use the Hough circle detection to extract each circle respectively. Fill \(w\) in the intersecting part of the circles 1 = 1, and fill \(w\) in the remaining parts 0 = 0; Step 2: Use the Hadamard product of image S m×n and mask M m×n as the target region to be extracted; Step 3, calculate the maximum between-class variance in the target area to obtain the threshold for the thresholding operation in the target area; Step 4: Perform threshold binaryzation using the threshold obtained in Step 3; use the threshold T obtained in Step 3 as the threshold for the thresholding operation of the target region to perform threshold binaryzation, specifically select the thresholding function cv::thershold(); if f m = 1, then the image threshold segmentation is completed; if f m > 1, then perform binaryzation processing on the remaining target regions respectively, and then perform intersection and union operations on each binary image, that is, complete the image threshold segmentation.
2. The Otsu thresholding method with a mask according to claim 1, characterized in that: In step 2, the Hadamard product is expressed as follows: where p (i,j) represents the gray value of the i-th row and j-th column of the image.
3. The Otsu thresholding method with a mask according to claim 1, characterized in that: The specific process of step 3 is: Step 31, calculate the number and probability distribution of each grayscale pixel; Step 32: Divide the grayscale set of the target region into two categories, C 0 and C 1 , with parameter t as the threshold, where t ∈ [0, 255]; 0 and 1 Step 33, calculate C respectively 0 , C 1 intra-class probabilities of the two categories, and then calculate the grayscale means of C 0 and C 1 ; Step 34, calculate C 0 and C 1 respectively for the within-class variances of the two categories and the between-class variance between the two; Step 35, take t ′ as the threshold, where t ′ ∈[0, t) ∪ (t, 255], repeat steps 32 - 34 until the maximum between-class variance appears, and the corresponding value t is the desired threshold T; if f m = 1, the threshold solution ends; if f m > 1, the thresholds need to be obtained separately for the remaining masks.
Citation Information
Patent Citations
A method for extracting and identifying the end surface defects of steel pipes
CN109544513A
Image semantic segmentation-based sea-sky-line detection method
CN107808386A
Target detection tracking method
CN111383244A