Traffic Sign Detection Method and Device Based on Cluster Analysis
Through the traffic sign detection method based on cluster analysis, the problem of difficulty in detecting anti-attacks in the existing technology is solved, and accurate identification and high-accuracy detection of traffic signs are achieved.
Patent Information
- Application Number
- CN202510483080.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing traffic sign detection technology is difficult to detect unknown confrontational attacks, resulting in insufficient accuracy in sign detection results.
The traffic sign detection method based on cluster analysis is adopted, and the traffic sign detection model is preprocessed by obtaining the multi-perturbation traffic sign data set, data translation and cluster analysis are performed, the unconnected data set and unsupervised contour coefficient are calculated, and the clustering results are iteratively optimized. Finally, the traffic sign detection model is trained based on the annotated data.
It improves the accuracy of traffic sign detection, reduces the difficulty of detection under counterattack, and realizes accurate identification of traffic signs in the presence of interference.
Smart Images

Figure CN120014604B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic sign detection, and particularly to a traffic sign detection method and device based on clustering analysis. Background Art
[0002] Clustering Analysis (CA) is an important research direction in the field of machine learning, aiming to divide a data set into different categories to reflect the attribution of samples. This simple and effective process plays an important role in many fields such as bioinformatics, information theory, and pattern recognition. However, the data overlap problem existing in the data set often has an adverse impact on the performance of clustering analysis technology, which is also one of the most complex and challenging tasks in the current clustering analysis field. The reason for this problem is that most data sets contain a certain number of special samples, which belong to different categories but have very similar data characteristics, so that the clustering analysis model can hardly distinguish the differences between them, and thus cannot cluster the data correctly.
[0003] In recent years, a lot of work has been done in solving the data overlap problem. Most methods can be summarized into three categories: density-based preset clustering algorithms, dimensionality reduction techniques, and feature selection. Density-based preset clustering algorithms (such as DBSCAN and OPTICS) improve data overlap by identifying high-density regions, but these algorithms are sensitive to parameter selection and may not work well when dealing with non-uniform density data. Dimensionality reduction techniques (such as PCA and t-SNE) can effectively reduce data overlap caused by irrelevant features, but may lose key information useful for clustering analysis. Feature selection can select features that are more informative for clustering analysis, but this requires sufficient domain knowledge and computing resources, and sometimes it is difficult to find the exact feature combination to distinguish overlapping data.
[0004] Adversarial attacks in machine learning refer to the behavior of making the model make wrong predictions by tampering with input data. Adversarial attacks usually generate images with only tiny perturbations, resulting in an overlap between clean images and tampered images, which makes unsupervised pattern recognition models may not be able to detect these differences. Especially when the input data is subject to unknown adversarial attacks, the detection difficulty will be greater. Such attacks pose a serious threat to transportation applications. For example, they may lead to incorrect traffic sign recognition, incorrect vehicle route planning, or even the failure of the autonomous driving navigation system. These threats may trigger traffic accidents, endanger lives, and cause significant economic losses and infrastructure damage.
[0005] Therefore, it is necessary to construct and design a traffic sign detection and recognition method that can accurately identify to reduce the detection difficulty and improve the detection accuracy. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a traffic sign detection method and device based on clustering analysis to solve the problem that the existing traffic sign detection technology has great difficulty in detecting unknown adversarial attacks, resulting in inaccurate sign detection results.
[0007] The technical solution adopted by the present invention is as follows:
[0008] In a first aspect, the present invention provides a traffic sign detection method based on clustering analysis, including:
[0009] S1. Obtain a multi-perturbation traffic sign data set, preprocess the multi-perturbation traffic sign data set to obtain a numerical data set;
[0010] S2. Perform data translation on the numerical data set, and perform clustering analysis on the data set after data translation to determine the first centroid of the data set;
[0011] S3. Based on the numerical data set, calculate an unconnected data set through an Euclidean distance calculation function, and sort the first centroid to obtain a second centroid;
[0012] S4. Use the unconnected data set to update the data set after data translation to obtain a first data set, pull the data points of the first data set towards the second centroid, update the data after data translation based on the first data set, and perform aggregation and separation processing on the data points of the first data set to obtain a second data set;
[0013] S5. Use a preset clustering algorithm to cluster the second data set, and calculate the unsupervised silhouette coefficient of the clustering result;
[0014] S6. Iteratively execute steps S3 - S5, determine the optimal unsupervised silhouette coefficient from multiple unsupervised silhouette coefficients, and determine the typical sizes and feature distributions of various traffic signs based on the clustering results of the multi-perturbation traffic sign data set corresponding to the optimal unsupervised silhouette coefficient. At the same time, label the unlabeled original data to obtain a labeled multi-perturbation traffic sign data set;
[0015] S7. Train a preset traffic sign detection model based on the labeled multi-perturbation traffic sign data set, and use the trained traffic sign detection model to identify the traffic sign to be recognized to confirm the category of the traffic sign to be recognized.
[0016] Further, in S1, obtaining a multi-perturbation traffic sign data set and preprocessing the multi-perturbation traffic sign data set to obtain a numerical data set includes:
[0017] S11: Obtain a multi-perturbation traffic sign data set , and define the following image data expansion formula:
[0018]
[0019] Among them, represents the height of the image, represents the width of the image, represents the number of color channels RGB, and flatten represents the flatten function; using the above expansion formula, for each three-dimensional image data in the dataset is successively expanded into one-dimensional numerical data ;
[0020] S12: Combine the expanded one-dimensional numerical data into a new numerical dataset , which contains samples, and each sample represents an image and contains features.
[0021] Furthermore, in S2, the numerical dataset is subjected to data translation, and the dataset after data translation is subjected to clustering analysis to determine the first centroid of the dataset, including:
[0022] S21: According to the formula , the numerical dataset is translated as a whole in the positive direction so that all feature values in the dataset are non-negative, and the dataset after data translation is obtained β ;
[0023] S22: Calculate k centroids from the dataset β through a preset clustering algorithm, and record the index values of the k centroids.
[0024] Furthermore, in S3, based on the numerical dataset, the unconnected dataset is calculated through the Euclidean distance calculation function, and the first centroid is sorted to obtain the second centroid, including:
[0025] S31: Define the following Euclidean distance calculation function:
[0026]
[0027] Among them, is the Euclidean distance between the data point and , and m is the number of features of the data points in the dataset ;
[0028] S32: Calculate the Euclidean distance between each data point in the dataset and the reference point through the Euclidean distance calculation function, and sort the dataset Reorder the data points in it, and an unconnectable data set is obtained after reordering ; The reference point is used to calculate the Euclidean distance to generate an unconnectable data set ψ , and α takes the minimum centroid of the sum of feature values ;
[0029] S33: According to the formula , reorder the k centroids to obtain the second centroid ; Among them, argsort represents the argsort function, represents the centroid with the smallest sum of feature values; distance represents the Euclidean distance calculation function; I represents the index value after centroid sorting, represents the second centroid obtained after reordering.
[0030] Furthermore, in S4, the unconnectable data set is used to update the data set after data translation to obtain the first data set, and the data points of the first data set are pulled towards the second centroid. Based on the first data set, the data after data translation is updated, and the data points of the first data set are aggregated and separated to obtain the second data set, including:
[0031] S41: Store the data set into a temporary data set ;
[0032] S42: Based on the unconnectable data set , according to the formula , update the data set to obtain the first data set, and pull all the data points in the first data set towards the second centroid ; Among them, represents the th data point in the data set i ; represents the th data point in the data set i ; represents the second centroid in the data set corresponding data point; represents the feature mean of all data points in the data set ;
[0033] S43: For each second centroid , execute step S42 once, and repeat k times in total;
[0034] S44: According to the formula , update the data set , separating data points of different categories and clustering data points of the same category to obtain a second data set; representing a data set in the i th
[0035] Further, the using a preset clustering algorithm to cluster the second data set and calculating the unsupervised silhouette coefficient of the clustering result includes:
[0036] Using a preset clustering algorithm to perform clustering processing on the data points of the second data set to obtain a clustering result ;
[0037] According to the data set and the clustering result , calculating the unsupervised silhouette coefficient of this clustering.
[0038] In a second aspect, the present invention provides a traffic sign detection device based on clustering analysis, including:
[0039] A data set preprocessing module, configured to obtain a multi-disturbed traffic sign data set, perform preprocessing on the multi-disturbed traffic sign data set to obtain a numerical data set;
[0040] A data translation analysis module, configured to perform data translation on the numerical data set, and perform clustering analysis on the data set after data translation to determine the first centroid of the data set;
[0041] A centroid sorting module, configured to calculate a non-connectable data set based on the numerical data set through an Euclidean distance calculation function, and sort the first centroid to obtain a second centroid;
[0042] A data set update module, configured to use the non-connectable data set to update the data set after data translation to obtain a first data set, pull the data points of the first data set towards the second centroid, update the data after data translation based on the first data set, perform aggregation and separation processing on the data points of the first data set to obtain a second data set;
[0043] A data clustering processing module, configured to use a preset clustering algorithm to cluster the second data set and calculate the unsupervised silhouette coefficient of the clustering result;
[0044] A sign clustering analysis module, configured to determine an optimal unsupervised silhouette coefficient from multiple unsupervised silhouette coefficients, determine the typical sizes and feature distributions of various traffic signs based on the clustering result of the multi-disturbed traffic sign data set corresponding to the optimal unsupervised silhouette coefficient, and at the same time label the unlabeled original data to obtain a labeled multi-disturbed traffic sign data set;
[0045] A logo detection module is used to train a preset traffic sign detection model based on an annotated multi-disturbance traffic sign data set, and use the trained traffic sign detection model to identify the traffic sign to be recognized and confirm the category of the traffic sign to be recognized.
[0046] In summary, the beneficial effects of the present invention are as follows:
[0047] A traffic sign detection method based on clustering analysis provided by the present invention, in the process of processing a multi-disturbance traffic sign data set, introduces an unconnectable data set based on distance calculation, and changes the data distribution through repeated aggregation processing and separation processing, so as to obtain the optimal unsupervised silhouette coefficient for accurately identifying the sizes of various traffic signs, solves the problem of data overlap in the clustering analysis of traffic sign data, and has a great detection difficulty for unknown adversarial attacks, resulting in inaccurate sign detection results. At the same time, the present invention also trains a preset traffic sign detection model based on the sizes of various traffic signs, uses the trained traffic sign detection model to identify the traffic sign to be recognized, and confirms the category of the traffic sign to be recognized, which can achieve accurate detection and recognition of traffic signs in the case of interference in the original traffic sign data. Description of the Drawings
[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts, and these are all within the protection scope of the present invention.
[0049] Figure 1 It is a flowchart of a traffic sign detection method based on clustering analysis of the present invention;
[0050] Figure 2 It is a schematic diagram of the data iterative processing process in the present invention;
[0051] Figure 3 It is a functional module diagram of a traffic sign detection device based on clustering analysis in the present invention. Detailed Embodiments
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. If there is no conflict, the various features in the present invention and its embodiments can be combined with each other, and all are within the protection scope of the present invention.
[0053] Embodiment 1:
[0054] Please refer to Figure 1 , Figure 1 , which is a schematic flowchart of a traffic sign detection method based on clustering analysis in Embodiment 1 of the present invention. As Figure 1 shown, the method provided by the present invention includes:
[0055] S1. Obtain a multi-disturbed traffic sign data set, preprocess the multi-disturbed traffic sign data set to obtain a numerical data set;
[0056] S2. Perform data translation on the numerical data set, and perform clustering analysis on the data set after data translation to determine the first centroid of the data set;
[0057] S3. Based on the numerical data set, calculate the unconnectable data set through the Euclidean distance calculation function, and sort the first centroid to obtain the second centroid;
[0058] S4. Use the unconnectable data set to update the data set after data translation to obtain the first data set, and pull the data points of the first data set towards the second centroid. Based on the first data set, update the data after data translation, and perform aggregation and separation processing on the data points of the first data set to obtain the second data set;
[0059] S5. Use a preset clustering algorithm to cluster the second data set, and calculate the unsupervised silhouette coefficient of the clustering result;
[0060] S6. Iteratively execute steps S3-S5, determine the optimal unsupervised silhouette coefficient from multiple unsupervised silhouette coefficients, and determine the typical sizes and feature distributions of various traffic signs based on the clustering results of the multi-disturbed traffic sign data set corresponding to the optimal unsupervised silhouette coefficient. At the same time, label the unlabeled original data to obtain the labeled multi-disturbed traffic sign data set;
[0061] S7. Train a preset traffic sign detection model based on the labeled multi-disturbed traffic sign dataset, and use the trained traffic sign detection model to identify the traffic sign to be recognized, and confirm the category of the traffic sign to be recognized.
[0062] Further, in the embodiment of the present invention, in S1, a multi-disturbed traffic sign dataset is obtained, and the multi-disturbed traffic sign dataset is preprocessed to obtain a numerical dataset, which specifically includes the following sub-steps:
[0063] S11: Obtain a multi-disturbed traffic sign dataset , and define the following image data expansion formula:
[0064]
[0065] where represents the height of the image, represents the width of the image, represents the number of color channels RGB, and flatten represents the flatten function; using the above expansion formula, each three-dimensional image data in the dataset is sequentially expanded into one-dimensional numerical data ;
[0066] S12: Combine the expanded one-dimensional numerical data into a new numerical dataset , which contains samples, each sample represents an image and contains features.
[0067] In the embodiment of the present invention, by converting three-dimensional image data into one-dimensional numerical data for processing, it can facilitate the subsequent data processing process to quickly complete data clustering and analysis, and improve the data processing speed.
[0068] Further, in the embodiment of the present invention, in S2, the numerical dataset is translated, and the dataset after data translation is subjected to clustering analysis to determine the first centroid of the dataset, which specifically includes the following sub-steps:
[0069] S21: According to the formula , translate the numerical dataset as a whole in the positive direction so that all feature values in the dataset are non-negative, and obtain the dataset after data translation β。
[0070] S22: Calculate k centroids from the dataset β through a preset clustering algorithm, and record the index values of the k centroids . In addition, k centroids can also be manually selected from the dataset β .
[0071] Among them, the preset clustering algorithm can adopt the k-means clustering algorithm, that is, the K-means clustering algorithm.
[0072] Furthermore, in the embodiment of the present invention, in S3, based on the numerical dataset, the unconnectable dataset is calculated through the Euclidean distance calculation function, and the first centroids are sorted to obtain the second centroids, including but not limited to the following sub-steps:
[0073] S31: Define the following Euclidean distance calculation function:
[0074]
[0075] Among them, is the Euclidean distance between the data point and , m is the number of feature of the data points in the dataset ;
[0076] S32: Calculate the Euclidean distance between each data point in the dataset and the reference point through the Euclidean distance calculation function, and reorder the data points in the dataset according to the calculated Euclidean distance. After sorting, the unconnectable dataset is obtained; the reference point is used to calculate the Euclidean distance to generate the unconnectable dataset ψ , and α takes the centroid with the smallest sum of feature values ;
[0077] S33: According to the formula , reorder the k centroids to obtain the second centroids ; where argsort represents the argsort function, represents the centroid with the smallest sum of feature values; distance represents the Euclidean distance calculation function; I represents the index value after the centroid sorting, represents the second centroids obtained after reordering.
[0078] Furthermore, in the embodiment of the present invention, in S4, the unconnectable dataset is used to update the dataset after data translation to obtain the first dataset, and the data points of the first dataset are pulled towards the second centroids. Based on the first dataset, the data after data translation is updated, and the data points of the first dataset are aggregated and separated to obtain the second dataset, including:
[0079] S41: Construct a temporary dataset, and store the dataset into a temporary dataset .
[0080] S42: Based on the non-connectable data set , according to the formula , update the data set , obtain the first data set, and pull all data points in the first data set towards the second centroid ; where represents the th data point in the data set i ; represents the th data point in the data set i ; represents the second centroid 's corresponding data point in the data set ; represents the feature mean of all data points in the data set .
[0081] S43: For each second centroid , execute step S42 once, and repeat it k times in total. Among them, the iterative processing process of the data is as Figure 2 shown.
[0082] S44: According to the formula , update the data set , separate data points of different classes, aggregate data points of the same class, and obtain the second data set; represents the th data point in the data set i .
[0083] Further, in the embodiment of the present invention, in S5, a preset clustering algorithm is used to cluster the second data set, and the unsupervised silhouette coefficient of the clustering result is calculated, which specifically includes the following sub-steps:
[0084] S51: Use a preset clustering algorithm to perform clustering processing on the data points of the second data set to obtain a clustering result ;
[0085] S52: Based on the data set and the clustering result , calculate the unsupervised silhouette coefficient of this clustering.
[0086] Specifically, in the embodiments of the present invention, through iterative clustering of the dataset, multiple unsupervised silhouette coefficients are obtained. At this time, through coefficient optimization analysis of the unsupervised silhouette coefficients, the optimal unsupervised silhouette coefficient is selected from multiple unsupervised silhouette coefficients, and based on the clustering results of the multi-disturbed traffic sign dataset corresponding to the optimal unsupervised silhouette coefficient, the typical sizes (such as aspect ratio, pixel range) and feature distributions (such as color, clustering center) of various traffic signs are determined. At the same time, the original unlabeled data is labeled (i.e., pseudo-labels are added) to obtain the labeled multi-disturbed traffic sign dataset. If the original data is already labeled, the noise labels can be corrected in combination with the clustering results.
[0087] Finally, the present invention trains a preset traffic sign detection model based on the labeled multi-disturbed traffic sign dataset, and uses the trained traffic sign detection model to identify the traffic sign to be recognized, so as to confirm the category of the traffic sign to be recognized. Among them, the preset traffic sign detection model can be pre-trained using the existing Yolov5 model.
[0088] Embodiment 2: Refer to Figure 3 As shown, the present invention provides a traffic sign detection device based on clustering analysis, including:
[0089] A dataset preprocessing module, configured to obtain a multi-disturbed traffic sign dataset, preprocess the multi-disturbed traffic sign dataset to obtain a numerical dataset;
[0090] A data translation analysis module, configured to perform data translation on the numerical dataset, and perform clustering analysis on the dataset after data translation to determine the first centroid of the dataset;
[0091] A centroid sorting module, configured to calculate an unconnectable dataset based on the numerical dataset through an Euclidean distance calculation function, and sort the first centroid to obtain a second centroid;
[0092] A dataset update module, configured to use the unconnectable dataset to update the dataset after data translation to obtain a first dataset, pull the data points of the first dataset towards the second centroid, update the data after data translation based on the first dataset, and perform aggregation and separation processing on the data points of the first dataset to obtain a second dataset;
[0093] A data clustering processing module, configured to cluster the second dataset using a preset clustering algorithm and calculate the unsupervised silhouette coefficient of the clustering result;
[0094] A flag clustering analysis module, which is used to determine an optimal unsupervised silhouette coefficient from multiple unsupervised silhouette coefficients, and determine the typical sizes and feature distributions of various traffic signs based on the clustering results of the multi-disturbed traffic sign dataset corresponding to the optimal unsupervised silhouette coefficient. At the same time, the original unlabeled data is labeled to obtain a labeled multi-disturbed traffic sign dataset;
[0095] A flag detection module, which is used to train a preset traffic sign detection model based on the labeled multi-disturbed traffic sign dataset, and use the trained traffic sign detection model to identify the traffic sign to be recognized, and confirm the category of the traffic sign to be recognized.
[0096] Specifically, in the embodiment of the present invention, on the one hand, the device neither requires operations of feature selection and data dimensionality reduction, nor solves the problem of data overlap. On the other hand, the device greatly improves the application scope of each clustering analysis process in the device and reduces data interference in the dataset through unlabeled data and an unsupervised verification process.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A traffic sign detection method based on cluster analysis, characterized in that: include: S1. Obtain a multi-disturbance traffic sign dataset, and pre-process the multi-disturbance traffic sign dataset to obtain a numerical dataset; S2, performing data translation on the numerical data set, and performing cluster analysis on the data set after the data translation to determine the first centroid of the data set; S3. Based on the numerical data set, the unconnectable data set is calculated by using the Euclidean distance calculation function, and the first centroids are sorted to obtain the second centroids; S4, using the unconnectable data set to update the data set after the data translation to obtain a first data set, and pulling the data points of the first data set to the second centroid, updating the data after the data translation based on the first data set, and aggregating and separating the data points of the first data set to obtain a second data set; S5. Clustering the second data set using a preset clustering algorithm, and calculating an unsupervised silhouette coefficient of the clustering result; S6, iteratively executing steps S3-S5, determining an optimal unsupervised contour coefficient from multiple unsupervised contour coefficients, and determining typical sizes and characteristic distributions of various types of traffic signs based on clustering results of the multi-disturbance traffic sign data set corresponding to the optimal unsupervised contour coefficient, and annotating the unannotated original data to obtain an annotated multi-disturbance traffic sign data set; S7. Train a preset traffic sign detection model based on the labeled multi-disturbance traffic sign dataset, and use the trained traffic sign detection model to identify the traffic sign to be identified, and confirm the category of the traffic sign to be identified.
2. The traffic sign detection method based on cluster analysis according to claim 1, characterized in that: In S1, a multi-disturbance traffic sign dataset is obtained, and the multi-disturbance traffic sign dataset is preprocessed to obtain a numerical dataset, including: S11: Obtaining a multi-perturbation traffic sign dataset , define the following image data expansion formula: in, Represents the height of the image, Represents the width of the image, represents the number of color channels RGB, flatten represents the flatten function; using the above expansion formula, the data set Each three-dimensional image data in is expanded into one-dimensional numerical data in turn. ; S12: Combine the expanded one-dimensional numerical data into a new numerical data set , which contains samples, each sample represents an image and contains Features.
3. The traffic sign detection method based on cluster analysis according to claim 1, characterized in that: In S2, data translation is performed on the numerical data set, and cluster analysis is performed on the data set after the data translation to determine the first centroid of the data set, including: S21: According to the formula , translate the numerical data set in the positive direction as a whole, so that all the eigenvalues in the data set are non-negative, and obtain the data set after data translation β ; S22: Using a preset clustering algorithm from the data set β Calculate k centroids and record the index values of k centroids .
4. The traffic sign detection method based on cluster analysis according to claim 1, characterized in that: In the S3, based on the numerical data set, the unconnectable data set is calculated by the Euclidean distance calculation function, and the first centroid is sorted to obtain the second centroid, including: S31: Define the Euclidean distance calculation function: in, It is a data point and The Euclidean distance between m It is a dataset The number of features of the data points in ; S32: Calculate the data set using the Euclidean distance calculation function Each data point and reference point The Euclidean distance of Reorder the data points in to obtain a non-connectable data set The reference point Used to calculate Euclidean distance to generate unconnectable datasets ψ , α is the minimum centroid of the sum of the eigenvalues ; S33: According to the formula , reorder the k centroids to get the second centroid ; Among them, argsort represents the argsort function, represents the centroid with the smallest sum of eigenvalues; distance Represents the Euclidean distance calculation function; I Represents the index value after centroid sorting, Represents the second centroid after reordering.
5. The traffic sign detection method based on cluster analysis according to claim 1, characterized in that: In S4, the unconnectable data set is used to update the data set after the data translation to obtain a first data set, and the data points of the first data set are pulled to the second centroid, the data after the data translation is updated based on the first data set, and the data points of the first data set are aggregated and separated to obtain a second data set, including: S41: Dataset Store in a temporary dataset middle; S42: Based on unconnectable datasets , according to the formula , update the dataset , get the first data set, and pull all data points in the first data set to the second centroid ;in, Representation dataset The i data points; Representation dataset The i data points; Represents the second centroid In the dataset The corresponding data points in ; Representation dataset The feature mean of all data points in ; S43: For each second centroid , execute step S42 once, and repeat k times in total; S44: According to the formula , update the dataset , separate data points of different categories, aggregate data points of the same category, and obtain the second data set; Representation dataset The i data points.
6. The traffic sign detection method based on cluster analysis according to claim 1, characterized in that: The method of clustering the second data set using a preset clustering algorithm and calculating an unsupervised silhouette coefficient of the clustering result includes: Use the preset clustering algorithm to cluster the data points of the second data set to obtain the clustering results ; Based on the data set And clustering results , calculate the unsupervised silhouette coefficient of this clustering.
7. A traffic sign detection device based on cluster analysis, characterized in that: include: A data set preprocessing module is used to obtain a multi-disturbance traffic sign data set, preprocess the multi-disturbance traffic sign data set, and obtain a numerical data set; The data translation analysis module is used to translate the numerical data set, and perform cluster analysis on the data set after the data translation to determine the first centroid of the data set; The centroid sorting module is used to calculate the unconnectable data set based on the numerical data set through the Euclidean distance calculation function, and sort the first centroid to obtain the second centroid; A data set updating module, used to update the data set after data translation by using the unconnectable data set to obtain a first data set, and pull the data points of the first data set to the second centroid, update the data after data translation based on the first data set, and aggregate and separate the data points of the first data set to obtain a second data set; A data clustering processing module, used to cluster the second data set using a preset clustering algorithm and calculate an unsupervised silhouette coefficient of the clustering result; A sign clustering analysis module is used to determine the optimal unsupervised silhouette coefficient from multiple unsupervised silhouette coefficients, and determine the typical size and characteristic distribution of various types of traffic signs based on the clustering results of the multi-disturbance traffic sign data set corresponding to the optimal unsupervised silhouette coefficient, and annotate the unannotated original data to obtain the annotated multi-disturbance traffic sign data set; The sign detection module is used to train a preset traffic sign detection model based on the labeled multi-disturbance traffic sign data set, and use the trained traffic sign detection model to identify the traffic sign to be identified, and confirm the category of the traffic sign to be identified.
Citation Information
Patent Citations
Model training method and device
CN115731530A
Accelerating convolutional neural network computation throughput
US20190102640A1