A robust greedy clustering method
By adopting a clustering method based on greedy ideas in three-dimensional scanning measurement, using center of gravity calculation and noise removal steps, the problems of unstable and noise sensitivity of clustering processing in the prior art are solved, and efficient and automated clustering effect is achieved.
Patent Information
- Application Number
- CN202211566272.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-12-07
AI Technical Summary
In the prior art, in three-dimensional scanning measurement, clustering algorithms require user interaction, unstable processing, and are sensitive to noise, and cannot effectively avoid noise interference.
The clustering method based on greedy ideas is adopted to realize the clustering of points in the two-dimensional space domain through the plane distance criterion, and the center of gravity calculation and noise removal steps are used to improve robustness and automation.
It reduces the number of user interactions, improves the degree of automation of the system, can effectively avoid noise interference, and improves the accuracy and robustness of clustering results.
Smart Images

Figure CN115761292B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of three-dimensional scanning measurement and is widely used in calibration and measurement processes. Specifically, it relates to a point clustering algorithm. The algorithm is based on a greedy idea and takes plane distance as a criterion to realize clustering of points in a two-dimensional space domain. The algorithm has certain robustness and can also complete clustering well in the case of a small amount of noise points. Background Art
[0002] In the field of object geometric dimension measurement, optical 3D scanners have been widely used. To ensure the scanning accuracy of 3D scanners, it is necessary to calibrate the scanner and the scanner tracking and positioning system before scanning. Figure 1 As shown in the figure, reflective marker points are often used as feature points during calibration, and some special coded patterns with coded information are formed to facilitate the recognition and matching of points with the same name; and in order to assist the 3D scanner to complete the scanning and measurement tasks, multiple groups of reflective marker point patterns with coded information are often used in the scanning process to assist the positioning of the 3D scanner and the splicing and fusion of the point cloud.
[0003] like Figure 2 The figure shows an image of a calibration plate taken during the prior art calibration process. There are multiple groups of coded patterns with coded information on the calibration plate. The coding principle shown in the figure is that each coded pattern is composed of eight points, and the relative positions of some points are different. According to the linearity and cross ratio invariance of the mapped projection line, the decoding operation of different patterns is completed. In order to complete the decoding of each pattern, all the reflective mark points in the entire image need to be clustered according to the two-dimensional image plane distance criterion, and the division between different coded patterns is completed, and then the subsequent decoding operation is performed for each coded pattern.
[0004] Common clustering algorithms, such as k-means, often take as input the points to be clustered and the number of categories after clustering. However, the number of categories is often different for different images, so continuous user interaction is required during processing, which affects the fluency and automation of the detection process. In addition, the k-means algorithm needs to be iterated continuously to complete clustering, and it is highly dependent on the selection of random initial values. If the initial value is not selected well, the number of iterations will increase, the operation time will become longer, and in the case of noise points, the noise points will also be clustered, which cannot avoid the noise points well. Finally, when there are a small amount of noise points in the image, such as Figure 3 As shown, the points marked by the red frame are noise points, and the existing technology cannot effectively avoid the influence of noise points on the clustering results. Summary of the invention
[0005] The purpose of the present invention is to provide a clustering method for points in a two-dimensional space domain based on a greedy idea and with plane distance as a criterion, which serves the field of three-dimensional shape measurement, is widely used in calibration and measurement processes, and the clustering method of the present invention facilitates the decoding of each coding pattern and has good robustness. It can also complete clustering well in the case of a small amount of noise points.
[0006] To achieve the purpose of the invention, the technical solution provided by the invention is: a robust clustering method based on greedy thinking, comprising the following steps:
[0007] Step S1, firstly, the input conditions are given, the set of points to be classified C in the image and the number of individuals N in each class;
[0008] Step S2, create a new point set C i , take a point p from the point set C to be classified i Add, point set C i ;
[0009] Step S3, calculate the point set C i The center of gravity p hi ;
[0010] Step S4, extract the distance p from the point set C to be classified hi The nearest point p i Join point set C i ;
[0011] Step S5, recalculate and update the center of gravity p hi ;
[0012] Step S6, determine the point set C at this time i Whether the number of midpoints reaches N, if the point set C i If the number of midpoints is less than N, then according to the greedy idea, return to S4 and continue to take out the distance p from the set of points to be classified C. hi The nearest point is added to the point set C i , and recalculate and update the center of gravity p hi ; If the point set C i The number of midpoints is equal to N, and the process goes to step S7;
[0013] Step S7, now the point set C i Considered as the initial i-th class, calculate the point set C i The center of gravity p hi With point set C i Center of gravity at mid-range p hi The farthest point p fi ;
[0014] Step S8, if the number n of remaining points in the point set C to be classified is greater than N, return to Step S2; continuously take points from the point set C to be classified until C is an empty set or the number of remaining points in C satisfies 0 < n < N, and complete the initialization of each category;
[0015] Among them, the point set C to be classified represents the original point set, N represents the number of individuals in each category after clustering, and p h represents the centroid position of each category, the subscript i represents the i-th category, and n represents a natural number;
[0016] The centroid p h (x h , y h ) coordinate calculation formula:
[0017] Among them, x h is the abscissa of the centroid p h , y h is the ordinate of the centroid p h , x i is the abscissa of the i-th point in this category, and y i is the ordinate of the i-th point in this category.
[0018] The preferred technical solution provided by the present invention is:
[0019] It further includes the following steps:
[0020] Step S9, the remaining points in the point set C to be classified are regarded as pseudo-noise points, save the data in the point set C to be classified as a queue, and take out a pseudo-noise point p e from the head of the queue;
[0021] Step S10, calculate the distance l e from p hi to the centroid point p he of each category respectively, and compare it with the distance l fi from the farthest point p hi to the centroid point p hf (i = 1, 2, 3... n), where n represents a natural number; if the calculated l he for each category is greater than l hf , then the point p e is a true noise point, and p e is removed from C;
[0022] If when traversing to the i-th category, l he is less than l hf , it means that p e is a false noise point and p fi is a pseudo-noise point, then add p e to this category, and store the p fi of this category at the tail of the queue, and recalculate the p of this categoryhi With p fi ; Continuously take points from the head of the queue, determine whether they are real noise points or false noise points, and continuously remove false noise points;
[0023] Among them, p e represents pseudo noise, p h Represents the center of gravity of the class, p f Indicates the point farthest from the centroid in the cluster, l he Indicates the distance from the centroid to the pseudo noise point, l hf It represents the distance from the centroid to the farthest point in the class, and the subscript i represents the i-th class;
[0024] l he The distance calculation formula is:
[0025] l hf The distance calculation formula is:
[0026] Among them, x h The center of gravity p h The horizontal axis, y h The center of gravity p h The vertical coordinate, x e is the horizontal coordinate of the pseudo noise point, y e is the ordinate of the pseudo noise point, x f is the horizontal coordinate of the farthest point in this class, y f is the ordinate of the farthest point of this class;
[0027] Step S11, determine whether the point set C to be classified is an empty set. If so, the entire algorithm is completed, the clustering result is correct, the noise points are not classified into any category, and the interference of the noise points is avoided; if not, return to step S9.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] The input of the algorithm of the present invention is the points to be clustered and the number of points in each cluster after clustering, so that batch processing of images containing coded patterns can be completed in the entire scanning process, the number of user interactions can be reduced, and the degree of automation of the system can be improved. In addition, the k-means algorithm needs to be continuously iterated to complete clustering, and it is highly dependent on the selection of random initial values. If the initial value is not selected well, the number of iterations will increase, the operation time will become longer, and in the case of noise points, the noise points will also be clustered, and the noise points cannot be avoided well. The algorithm proposed in this patent does not require iteration, the calculation process is simple, and it has good robustness. When there are a small amount of noise points in the image, such as Figure 3 As shown in the figure, the points marked by the red box are noise points and will not affect the clustering results. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a schematic diagram of the application scenario of the existing three-dimensional scanning technology;
[0031] Figure 2 Noise-free original images taken for prior art calibration;
[0032] Figure 3 For relative Figure 2 The original noisy image;
[0033] Figure 4 It is a flowchart of the present invention;
[0034] Figure 5 This is a schematic diagram for comparing the initial clustering results without noise;
[0035] Figure 6 This is a schematic diagram for comparing the initial clustering results with noise;
[0036] Figure 7 This is a schematic diagram of the final denoising clustering result with noise points. DETAILED DESCRIPTION
[0037] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings.
[0038] Figure 4 It is a flowchart of the present invention, combined with Figure 4 The present invention is specifically described as follows:
[0039] A robust greedy clustering method includes the following steps:
[0040] Step S1, firstly, the input conditions are given, the set of points to be classified C in the image and the number of individuals N in each class;
[0041] Step S2, create a new point set C i , take a point p from the point set C to be classified i Add, point set C i ;
[0042] Step S3, calculate the point set C i The center of gravity p hi ;
[0043] Step S4, extract the distance p from the point set C to be classified hi The nearest point p i Join point set C i ;
[0044] Step S5, recalculate and update the center of gravity p hi ;
[0045] Step S6, determine the point set C at this time iWhether the number of midpoints reaches N. If the point set C i the number of midpoints is less than N, then according to the greedy idea, return to S4, and continuously take out the point from the point set C to be classified that is hi closest to p i (The narrow sense of distance refers to the physical distance in two-dimensional space. Here, it is the pixel distance in the image coordinate system, and any form of distance expression can be customized according to different problems), and recalculate and update the centroid p hi ; if the point set C i the number of midpoints is equal to N, enter step S7;
[0046] Step S7, at this time, regard the point set C i as the initial i-th class, and calculate the centroid p i of the point set C hi and the point in the point set C i farthest from the centroid p hi p fi (p fi may be a noise point);
[0047] Step S8, if the number n of remaining points in the point set C to be classified is greater than N, return to step S2; continuously take points from the point set C to be classified until C is an empty set or the number of remaining points 0 < n < N in C, and complete the initialization of each category;
[0048] Among them, the point set C to be classified represents the original point set, N represents the number of individuals in each category after clustering, p h represents the centroid position of each category, the subscript i represents the i-th category, and n represents a natural number;
[0049] The centroid p h (x h , y h ) coordinate calculation formula:
[0050] In the ideal case (in the case of no noise points), steps S1 to S8 can complete the clustering, and the clustering result is as Figure 5 shown. The left is the original picture, and the right is the result after clustering (points of different colors represent different categories).
[0051] However, for the case where there are noise points, the clustering may be incorrect, as Figure 6 shown. The noise points occupy the positions of the normal points in the category, and the white points are the remaining pseudo-noise points after clustering, not real noise points. After clustering in steps S1 to S8, there may be noise points mixed in some categories, resulting in the abandonment of valid points and deviation in the clustering result. To enhance the robustness of the algorithm of the present invention, it is necessary to continue clustering to remove noise points.
[0052] Step S9: The remaining points in the point set C to be classified are regarded as pseudo noise points. The data in the point set C to be classified are saved as a queue, and a pseudo noise point p is taken from the head of the queue. e ;
[0053] Step S10, calculate p e To each type of centroid p hi The distance l he , and the farthest point p fi To the center of gravity p hi The distance l hf Compare (i = 1, 2, 3...n), if the calculated l of each type he Both are greater than l hf , then point p e is the real noise point, and p e Remove from C;
[0054] If we traverse to the i-th category, l he Less than l hf , indicating that p e is a false noise point, p fi is a pseudo noise point, then p e Add this class and change the p fi Store it at the end of the queue and recalculate the p of this class hi With p fi ; Continuously take points from the head of the queue, determine whether they are real noise points or false noise points, and continuously remove false noise points.
[0055] Among them, p e represents pseudo noise, p h Represents the center of gravity of the class, p f Indicates the point farthest from the centroid in the cluster, l he Indicates the distance from the centroid to the pseudo noise point, l hf It represents the distance from the centroid to the farthest point in the class, and the subscript i represents the i-th class.
[0056] l he The distance calculation formula is:
[0057] l hf The distance calculation formula is:
[0058] Among them, x h The center of gravity p h The horizontal axis, y h The center of gravity p h The vertical coordinate, x e is the horizontal coordinate of the pseudo noise point, y e is the ordinate of the pseudo noise point, x f is the horizontal coordinate of the farthest point in this class, y fis the ordinate of the farthest point of this class.
[0059] The pseudo noise mentioned above means that it is uncertain whether it is real noise or false noise. Real noise is determined to be real noise, and false noise is determined not to be noise. Therefore, pseudo noise and false noise are two different concepts.
[0060] Step S11, determine whether the point set C to be classified is an empty set. If so, complete the entire algorithm. The clustering result is as follows: Figure 7 As shown, the clustering result is correct, the noise points are not classified into any category, and the interference of the noise points is well avoided; if not, return to step S9.
[0061] The described embodiments are only a part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
Claims
1. A robust greedy clustering method, characterized by: It includes the following steps: Step S1: First, given the input conditions, the point set C to be classified in the image and the number of individuals N in each category; Step S2, create a new point set C i , take a point p from the point set C to be classified i Add, point set C i ; Step S3, calculate the point set C i The center of gravity p hi ; Step S4, extract the distance p from the point set C to be classified hi The nearest point p i Join point set C i ; Step S5, recalculate and update the center of gravity p hi ; Step S6, determine the point set C at this time i Whether the number of midpoints reaches N, if the point set C i If the number of midpoints is less than N, then according to the greedy idea, return to S4 and continue to take out the distance p from the set of points to be classified C. hi The nearest point is added to the point set C i , and recalculate and update the center of gravity p hi ; If the point set C i The number of midpoints is equal to N, and the process goes to step S7; Step S7, now the point set C i Considered as the initial i-th class, calculate the point set C i The center of gravity p hi With point set C i Center of gravity at mid-range p hi The farthest point p fi ; Step S8: If the number of remaining points n in the point set C to be classified is n > N, then return to Step S2; continuously take points from the point set C to be classified until C is an empty set or 0 < n < N for the number of remaining points in C, and complete the initialization of each category; Among them, the point set to be classified C represents the original point set, N represents the number of individuals in each class after clustering, and p h represents the centroid position of each category, the subscript i represents the i-th category, and n represents a natural number; Center of gravity h (x h ,y h )Coordinate calculation formula: Among them, x h The center of gravity p h The horizontal axis, y h The center of gravity p h The vertical coordinate, x i is the horizontal coordinate of the i-th point in this class, y i is the ordinate of the i-th point in this class.
2. A robust greedy clustering method according to claim 1, characterized in that: It also includes the following steps: Step S9: The remaining points in the point set C to be classified are regarded as pseudo noise points. The data in the point set C to be classified are saved as a queue, and a pseudo noise point p is taken from the head of the queue. e ; Step S10, calculate p e To each type of centroid p hi The distance l he , and the farthest point p fi To the center of gravity p hi The distance l hf Compare (i = 1, 2, 3 ... n), n is a natural number; if the l calculated for each type he Both are greater than l hf , then point p e is the real noise point, and p e Remove from C; If we traverse to the i-th category, l he Less than l hf , indicating that p e is a false noise point, p fi is a pseudo noise point, then p e Add this class and change the p fi Store it at the end of the queue and recalculate the p of this class hi With p fi ; Continuously take points from the head of the queue, determine whether they are real noise points or false noise points, and continuously remove false noise points; Among them, p e represents pseudo noise, p h Represents the center of gravity of the class, p f Indicates the point farthest from the centroid in the cluster, l he Indicates the distance from the centroid to the pseudo noise point, l hf It represents the distance from the centroid to the farthest point in the class, and the subscript i represents the i-th class; l he The distance calculation formula is: l hf The distance calculation formula is: Among them, x h The center of gravity p h The horizontal axis, y h The center of gravity p h The vertical coordinate, x e is the horizontal coordinate of the pseudo noise point, y e is the ordinate of the pseudo noise point, x f is the horizontal coordinate of the farthest point in this class, y f is the ordinate of the farthest point of this class; Step S11: Determine whether the point set C to be classified is an empty set. If so, complete the entire algorithm, the clustering result is correct, and the noise points are not classified into any category, avoiding the interference of noise points; if not, return to Step S9.
Citation Information
Patent Citations
Improved method of DPC clustering algorithm
CN113780437A
Clustering method for rosette scan image
KR1020020013067A