Image segmentation algorithm based on multiple clustering

The image is denoised by the DBSCAN algorithm, and the clustering centers and number generated are used to provide initial conditions for the K-means algorithm, solving the local optimal problem and noise sensitivity caused by the randomness of the initial condition in image segmentation, achieving a better image segmentation effect.

CN120182294APending Publication Date: 2025-06-20XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510271730.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-09
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing K-means clustering method is prone to local optimization due to random selection of the initial clustering center during image segmentation, and is sensitive to noise, making it difficult to obtain ideal segmentation results.

Method used

The image is denoised by the DBSCAN algorithm, and the central data points and number K of each generated cluster are recorded, and image segmentation is performed by the K-means clustering algorithm based on this information.

Benefits of technology

Through the preprocessing of the DBSCAN algorithm, the error of the K-means algorithm when selecting the initial clustering center and the number of clusters is reduced, and the accuracy and robustness of image segmentation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182294A_ABST
    Figure CN120182294A_ABST
Patent Text Reader

Abstract

The invention relates to an image segmentation algorithm based on multiple clustering, and the method comprises the steps: carrying out the denoising processing of an image through a DBSCAN algorithm, recording the central data points of each generated cluster and the number K of clusters, and carrying out the segmentation processing of the denoised image through a K-means clustering algorithm based on the central data points and the number K of clusters. The image segmentation algorithm not only can be used for noise image segmentation, but also can realize a good image segmentation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to an image segmentation algorithm based on multiple clustering. Background Art

[0002] Image segmentation refers to dividing an image into several non-overlapping regions with distinct features, and extracting the region of interest as the target. Image segmentation is a key step from image processing to image analysis, and the quality of the image segmentation result directly affects the subsequent image analysis.

[0003] In the post-print inspection stage of the printing process, it is necessary to analyze the printed products to confirm the printing quality, especially to sample and confirm whether there are defects during the printing process. Therefore, it is necessary to perform image analysis on the printed products, and in this process, image segmentation technology is the most important basic link, which is a key step from image processing to image analysis.

[0004] Currently, the K-means clustering method is a widely used clustering image segmentation method, but it has the following obvious disadvantages: (1) The selection of the initial clustering center has a great influence on the image segmentation result. If the initial clustering center is not selected well, the method will fall into a local optimum and cannot obtain an ideal segmentation result; (2) This clustering method is sensitive to noise. When processing noisy images, it cannot obtain satisfactory segmentation results. Summary of the Invention

[0005] In view of the above technical problems, the present invention proposes an image segmentation algorithm based on multiple clustering. This method performs denoising processing on the image through the DBSCAN algorithm, and records the central data points of each generated cluster and the number of clusters K. Based on the central data points and the number of clusters K, the denoised image is segmented through the K-means clustering algorithm. This image segmentation algorithm reduces the errors generated by the K-means clustering algorithm when randomly selecting the initial clustering center and the number of clusters, and can better remove noise points, thereby enabling better image segmentation processing.

[0006] To achieve the above technical objectives, the present invention adopts the following technical solutions:

[0007] The present invention specifically relates to an image segmentation algorithm based on multiple clustering, which specifically includes: Step 1, obtaining the image to be processed and converting it to the Lab color space; Step 2, performing denoising processing on the converted image through the DBSCAN clustering algorithm, and recording the central data points T k (k = 1, 2,..., K) and the number of clusters K; Step 3, based on the central data points T generated in Step 2 k(k = 1, 2, ……, K) and the number of clusters K, and the denoised image is segmented by the K - means clustering algorithm.

[0008] Further, in step 2, denoising is performed by the DBSCAN clustering algorithm, and the central data points T of each generated cluster are recorded k (k = 1, 2, ……, K) and the number of clusters K, specifically including:

[0009] Step 21, define the feature vector of each pixel point as (L i , a i , b i ), and the feature vectors of each pixel point form the data points f i = (L i , a i , b i ), where i is from 1 to M * N, and M and N are the width and height of the image respectively;

[0010] Step 22, perform normalization processing on the data points f i constructed in step 21 to obtain the first data set F1;

[0011] Step 23, determine the neighborhood radius ∈ and the minimum number of points M in , specifically including:

[0012] Step 231, select the minimum number of points M in ;

[0013] Step 232, for each data point f i calculate the average weighted Euclidean distance AD i from this point f in to the M i nearest data points, where f j represents the 1 - M i nearest data points around f in , and d(f i , f j ) represents the weighted Euclidean distance between the data points f i and f j , where ɑ, β, γ are preset coefficients;

[0014] Step 233, select the neighborhood radius where δ is a preset coefficient, and δ ∈ [1.2, 1.5];

[0015] Step 24, perform clustering processing on the normalized first data set F1 by the DBSCAN clustering algorithm, specifically including:

[0016] Step 241: For each data point f in the standardized first data set F1 i , calculate the number of data points within the neighborhood of the neighborhood radius ∈ determined in Step 233. If the number of data points is greater than or equal to the minimum number of points M in , then mark this point as a core point, create a new cluster, and add this point and all data points within its ∈-neighborhood to the cluster;

[0017] Step 242: If the number of data points within the neighborhood of the neighborhood radius ∈ is less than the minimum number of points M in , and there are core points within the ∈-neighborhood radius of this data point, then mark this point as a boundary point;

[0018] Step 243: If the number of data points within the ∈-neighborhood of a data point is less than the minimum number of points M in , and there are no core points within the ∈-neighborhood radius of this data point, then mark this point as a noise point and record the number of noise points; replace the pixel points marked as noise points with the mean of all non-noise points within the ∈-neighborhood of this noise point;

[0019] Step 244: Repeat Steps 241 - 243 until all data points are marked as core points or boundary points, thus obtaining K clusters each containing core points and boundary points, and record the central data points T of each cluster k (k = 1, 2, ……, K);

[0020] Step 245: Generate a denoised image based on the processed data points.

[0021] Furthermore, based on the central data points T generated in Step 2 k (k = 1, 2, ……, K) and the number of clusters K, perform segmentation processing on the denoised image through the K-means clustering algorithm, specifically including:

[0022] Step 31: Normalize the data points of each pixel of the denoised image to obtain the second data set F2;

[0023] Step 32: Use the central data points T of each cluster determined in Step 246 k as the initial cluster centers of K-means, and use the number of clusters K as the number of initial clusters;

[0024] Step 33: Calculate the Euclidean distance D(F i ' of each data point F' in the second data set F2 to each initial cluster center T k (k = 1, 2, ……, p, …… K) respectively i ′, Tk ) If D(F i ′, T p ) ≤ D(F i ′, T k ), then assign F' i point to the p-th cluster;

[0025] Step 34, recalculate the center positions of each cluster, O k is the number of data points in the current k-th cluster, k = 1, 2,..., K, C l is the data point value in the current k-th cluster;

[0026] Step 35, perform convergence judgment. Through the loop calculation of Step 33 and Step 34, until the center points t k of each cluster basically no longer change, then the clustering division ends;

[0027] Step 36, extract the pixel brightness values of the center points t k of each cluster, and sort them according to the brightness value to obtain t 1L <t 2L <... < t KL , calculate the segmentation threshold S K = ((t 1L + t 2L ) / 2, (t 2L + t 3L ) / 2,..., (t (K - 1 )L + t KL ) / 2):

[0028] Step 37, divide the denoised image into K regions according to the threshold S K calculated in Step 36.

[0029] Furthermore, after the clustering process of the standardized first data set F1 by the DBSCAN algorithm in Step 24, it further includes: when the number of clusters K generated in Step 246 is greater than 2 times the minimum number of points Min, merge similar clusters to simplify the number of clusters, specifically including: calculate the Euclidean distance between the center data points of each cluster respectively. When the Euclidean distance between any two center data points of the clusters is less than the preset threshold, then merge the two clusters.

[0030] On the other hand, the present invention also discloses a terminal, including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the above method.

[0031] The present invention performs denoising and image segmentation processing on an image by using the DBSCAN clustering algorithm and the K-means clustering algorithm successively. It can give full play to the advantages of the DBSCAN clustering algorithm, perform noise reduction processing on the image in advance, and pre-determine the initial centers of clustering, thereby reducing the errors generated by the K-means clustering algorithm when randomly selecting the initial clustering centers and the number of clusters, and enabling better image segmentation processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The drawings described herein are used to provide a further understanding of the present disclosure and form a part of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:

[0033] Figure 1 A flowchart showing a method for adjusting the image contrast according to an embodiment of the present invention is shown;

[0034] Figure 2 An original image provided by an embodiment of the present invention is shown;

[0035] Figure 3 A segmented image obtained by applying the K-means clustering algorithm to the original image provided by an embodiment of the present invention is shown;

[0036] Figure 4 A segmented image obtained by applying the method of the present invention to the original image provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention.

[0038] Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0039] As Figure 1 shown, the present invention provides an image segmentation algorithm based on multiple clustering, which specifically includes:

[0040] Step 1, image preprocessing: Obtain the image to be processed, convert it to the Lab color space, and extract its luminance channel value L, color channel values a and b.

[0041] Step 2: Denoise the transformed image using the DBSCAN clustering algorithm and record the central data points T of each generated cluster k (k = 1, 2, ……, K) and the number of clusters K, specifically including:

[0042] Step 21: Define the feature vector of each pixel point as (L i , a i , b i ), and the feature vectors of each pixel point constitute the data points f i = (L i , a i , b i ), where i is from 1 to M*N, and M and N are the width and height of the image respectively.

[0043] Step 22: Standardize each data point f i constructed in Step 21 to obtain the first data set F1. The standardization process here can use methods such as mean normalization, Z-score normalization, and maximum-minimum normalization. The present invention does not make specific limitations.

[0044] Step 23: Determine the neighborhood radius ∈ and the minimum number of points M in , specifically including:

[0045] Step 231: Select the minimum number of points M in . In a preferred embodiment, M in ∈ [5, 15]. In a preferred embodiment, M in is 6.

[0046] Step 232: For each data point f i calculate the average weighted Euclidean distance AD i from this point f in to the M i nearest neighbor data points, where f j represents the 1-M i nearest neighbor data points around f in , and d(f i , f j ) represents the weighted Euclidean distance between the data points f i and f j , where ɑ, β, and γ are preset coefficients. In a preferred embodiment, ɑ = 0.3, β = 0.4, and γ = 0.3.

[0047] Step 233: Initially select the neighborhood radius Where δ is a preset coefficient, δ∈[1.2,1.5]. In a preferred embodiment, the value of δ is 1.25. As the value of δ increases, the radius of the neighborhood increases, and more data points are included in the neighborhood, thereby reducing noise marking. By selecting δ∈[1.2,1.5], the present invention can avoid mistakenly marking ordinary data points as noise points as much as possible.

[0048] Step 24, clustering the standardized first data set F1 using the DBSCAN algorithm, specifically includes:

[0049] Step 241, for each data point f in the standardized first data set F1 i , calculate the number of data points in the neighborhood of the neighborhood radius ∈ determined in step 233, if the number of data points is greater than or equal to the minimum number of points M in , then mark the point as a core point, create a new cluster, and add the point and all the data points in its ∈-neighborhood to the cluster.

[0050] Step 242: If the number of data points in the neighborhood of the neighborhood radius ∈ is less than the minimum number of points M in , and there is a core point within the ∈-neighborhood radius of the data point, then the point is marked as a boundary point;

[0051] Step 243, if the number of data points in the ∈-neighborhood of the data point is less than the minimum number of points M in , and there is no core point within the ∈-neighborhood radius of the data point, then mark the point as a noise point and record the number of noise points; replace the pixel points marked as noise points with the mean of all non-noise points in the ∈-neighborhood of the noise point.

[0052] Step 244, repeat steps 241-243 until all data points are marked as core points or boundary points, thereby obtaining K clusters each containing core points and boundary points, and recording the central data point T of each cluster. k (k=1, 2, ..., K).

[0053] In a preferred embodiment, the central data point T of each cluster k (i=1, 2, ..., p, ..., K) can be the geometric center of the cluster; or the point closest to all other points in the cluster can be calculated as the center of the cluster. The specific formula is not specifically limited here.

[0054] Step 245, generating a denoised image based on the processed data points.

[0055] In another embodiment, when the number of noise points is greater than or equal to 25% of the total number of all data points, the size of the minimum number of points Min is decreased by 1, and step 24 is performed again. In another embodiment, when the number of noise points is less than 3% of the total number of all data points, the size of the minimum number of points Min is increased by 1, and step 24 is performed again. Through the above adaptive adjustment process, the size of the minimum number of points Min can be better optimized, so that noise points and normal data points can be better distinguished.

[0056] It should be noted that the present invention first performs noise reduction processing on the image through the DBSCAN clustering algorithm. The pre-denoising by DBSCAN can reduce the sensitivity to outliers and improve the robustness of the final segmentation. In addition, the present invention also pre-determines the initial center points of clustering and the initial number of clusters through the DBSCAN clustering algorithm, so as to provide the initial center points of clustering and the initial number of clusters for the subsequent K-means clustering algorithm, avoiding the conventional automatic generation steps and greatly improving the image segmentation ability of the algorithm.

[0057] In another preferred embodiment, after step 24 performs clustering processing on the standardized first data set F1 through the DBSCAN algorithm, it further includes: when the number of clusters K generated in step 246 is greater than 2 times the minimum number of points Min, similar clusters are merged to simplify the number of clusters, specifically including: calculating the Euclidean distance between the data points of each cluster center respectively, and when the Euclidean distance between any two cluster center data points is less than the preset threshold, then the two clusters are merged.

[0058] In a preferred embodiment, the preset threshold is 2 times the neighborhood radius ∈.

[0059] Step 3, based on the center data points T k (k = 1, 2, ……, K) generated in step 2 and the number of clusters K, perform segmentation processing on the denoised image through the K-means clustering algorithm, specifically including:

[0060] Step 31, perform normalization processing on the data points of each pixel of the denoised image to obtain a second data set F2.

[0061] Step 32, based on the center data points T k determined in step 246 as the initial cluster centers of K-means, and taking the number of clusters K as the number of initial clusters.

[0062] Step 33, calculate the Euclidean distance D(F i from each data point F' in the second data set F2 to each initial cluster center T k (k = 1, 2, ……, p, …… K) respectively. i′, T k ), if D(F i ′, T p ) ≤ D(F i ′, T k ), then assign F' i point to the p-th cluster, that is, when the data point F' i to one of the cluster centers T p has the minimum Euclidean distance D, then assign the data point F' i to the p-th cluster.

[0063] In another preferred embodiment, the above weighted Euclidean distance formula can also be used for calculation, which will not be elaborated here. In this preferred embodiment, ɑ = 0.7, β = 0.15, γ = 0.15.

[0064] Step 34, recalculate the center positions of each cluster, O k is the number of data points in the current k-th cluster, k = 1, 2,..., K, C l is the data point value in the current k-th cluster;

[0065] Step 35, perform convergence judgment. Through the loop calculation of steps 33 and 34, until each cluster center t k basically no longer changes, then the clustering division ends.

[0066] This convergence condition, that is, basically no longer changes, can be that when the change value of the cluster center t i is less than 0.05%, or when the number of iterations is greater than 100. The present invention does not make specific limitations here.

[0067] Step 36, extract the pixel brightness values of each cluster center t k and sort them according to this brightness value to get t 1L < t 2L <... < t KL , calculate the segmentation threshold S K = ((t 1L + t 2L ) / 2, (t 2L + t 3L ) / 2,..., (t (K -1 )L + t KL ) / 2).

[0068] Step 37, divide the denoised image into K regions according to the threshold S K calculated in step 36. It should be noted that for how to divide according to the threshold S iDividing a grayscale image into K-1 regions can be done using common practices in the art, which is not the inventive point of the present invention and will not be elaborated here.

[0069] As described above, the present invention performs denoising and image segmentation on an image by successively using the DBSCAN clustering algorithm and the K-means clustering algorithm. It can not only utilize the advantages of the DBSCAN clustering algorithm to perform noise reduction on the image in advance and pre-determine the initial centers of clustering, thereby reducing the errors generated by the K-means clustering algorithm when randomly selecting the initial clustering centers and the number of clusters, but also achieve better image segmentation processing.

[0070] Figure 2 This is the original image provided for an embodiment of the present invention. Figure 3 It shows the segmented image obtained by the K-means clustering algorithm for the original image provided by an embodiment of the present invention. Figure 4 It shows the segmented image obtained by the method of the present invention for the original image provided by an embodiment of the present invention.

[0071] As Figure 3 and Figure 4 shown, it can be seen that compared with Figure 3 the segmented image obtained by the K-means clustering algorithm in Figure 4 the final image segmentation effect shown in

[0072] obtained by the image segmentation method provided by the present invention is better, and the noise in the image can be removed better.

[0073] In summary, the present invention provides an image segmentation algorithm based on multiple clustering. This method performs denoising on the image through the DBSCAN algorithm, records the central data points of each generated cluster and the number of clusters K, and based on the central data points and the number of clusters K, performs segmentation processing on the denoised image through the K-means clustering algorithm. This image segmentation algorithm reduces the errors generated by the K-means clustering algorithm when randomly selecting the initial clustering centers and the number of clusters, and can better remove noise points, thereby enabling better image segmentation processing.

[0074] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

Claims

1. An image segmentation algorithm based on multiple clustering, characterized in that: Specifically include: Step 1, obtain the image to be processed and convert it to Lab color space; Step 2: De-noise the converted image using the DBSCAN clustering algorithm and record the central data point T of each generated cluster. k (k=1, 2, ..., K) and the number of clusters K; Step 3: Based on the central data point T generated in step 2 k (k=1, 2, ..., K) and the number of clusters K, and the denoised image is segmented using the K-means clustering algorithm.

2. The image segmentation algorithm according to claim 1, characterized in that: In step 2, the DBSCAN clustering algorithm is used to perform denoising and record the central data point T of each generated cluster. k (k=1, 2, ..., K) and the number of clusters K, specifically including: Step 21, define the feature vector of each pixel as (L i , a i , b i ), the feature vector of each pixel constitutes the feature vector of each data point f i =(L i , a i , b i ), i is 1~M*N, M and N are the width and height of the image respectively; Step 22, for each data point f constructed in step 21 i Perform standardization processing to obtain a first data set F1; Step 23, determine the neighborhood radius ∈ and the minimum number of points M in , specifically including: Step 231, select the minimum number of points M in ; Step 232, for each data point f i Calculate the point f i To M in The average weighted Euclidean distance AD ​​of the nearest data points i , where f j represents f i Around 1-M in The nearest data point, d(f i ,f j ) represents the data point f i and f j The weighted Euclidean distance between Among them, ɑ, β, and γ are preset coefficients; Step 233, select the neighborhood radius Where δ is the preset coefficient, δ∈[1.2,1.5]; Step 24, clustering the standardized first data set F1 using the DBSCAN clustering algorithm, specifically includes: Step 241, for each data point f in the standardized first data set F1 i , calculate the number of data points in the neighborhood of the neighborhood radius ∈ determined in step 233, if the number of data points is greater than or equal to the minimum number of points M in , then mark the point as a core point, create a new cluster, and add the point and all the data points in its ∈-neighborhood to the cluster; Step 242: If the number of data points in the neighborhood of the neighborhood radius ∈ is less than the minimum number of points M in , and there is a core point within the ∈-neighborhood radius of the data point, then the point is marked as a boundary point; Step 243, if the number of data points in the ∈-neighborhood of the data point is less than the minimum number of points M in , and there is no core point within the ∈-neighborhood radius of the data point, then mark the point as a noise point and record the number of noise points; replace the pixel points marked as noise points with the mean of all non-noise points in the ∈-neighborhood of the noise point; Step 244, repeat steps 241-243 until all data points are marked as core points or boundary points, thereby obtaining K clusters each containing core points and boundary points, and recording the central data point T of each cluster. k (k=1,2,……,K); Step 245, generating a denoised image based on the processed data points.

3. The image segmentation algorithm according to claim 2, characterized in that: The step 3 is based on the central data point T generated in step 2. k (k=1, 2, ..., K) and the number of clusters K, the denoised image is segmented by the K-means clustering algorithm, specifically including: Step 31, normalizing the data point of each pixel of the denoised image to obtain a second data set F2; Step 32, based on the central data point T of each cluster determined in step 246 k As the initial cluster center of K-means, and the number of clusters K as the number of initial clusters; Step 33, respectively calculate each data point F' in the second data set F2 i To each initial cluster center T k (k=1,2,……,p,……K) i , T k ), if D(F′ i , T p )≤D(F′ i , T k ), then assign F' i Point to the pth cluster; Step 34, recalculate the center position of each cluster, O k is the number of data points in the current k-th cluster, k = 1, 2, ..., K, C l is the data point value in the current k-th cluster; Step 35, perform convergence judgment, and repeat the calculations of steps 33 and 34 until each cluster center t k If there is basically no change, the clustering is completed; Step 36, extract the center point t of each cluster k The pixel brightness value is sorted by the brightness value to get t 1L <t 2L <…… <t KL , calculate the segmentation threshold S K =((t 1L +t 2L ) / 2,(t 2L +t 3L ) / 2,……,(t (K-1)L +t KL ) / 2): Step 37, based on the threshold value S calculated in step 36 K The denoised image is divided into K regions.

4. The image segmentation algorithm according to claim 3, characterized in that: Also includes: The step 24, after clustering the standardized first data set F1 by the DBSCAN algorithm, further includes: when the number of clusters K generated in step 246 is greater than 2 times the minimum number of points Min, merging similar clusters to simplify the number of clusters, specifically including: calculating the Euclidean distance between each cluster center data point respectively, and when the Euclidean distance between any two cluster center data points is less than a preset threshold, merging the two clusters.

5. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 4.