Traffic sign recognition method, electronic device, vehicle and storage medium
Patent Information
- Application Number
- CN202110109832.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-27
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2041-01-27
AI Technical Summary
[0003]在现有识别交通标志物的方法中,需要计算场景中全部物体的情况或使用过程复杂的可见光图像识别算法,计算量较大,不适于需要快速识别出交通标志物的实际场景中
[0053] Understandably, the electronic device provided in the second aspect, the vehicle provided in the third aspect, the traffic sign recognition device provided in the fourth aspect, and the computer storage medium provided in the fifth aspect are all used to execute the method provided in the first aspect and any possible implementation thereof. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
Smart Images

Figure CN114821212B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method for recognizing traffic signs, electronic devices, vehicles, and storage media. Background Technology
[0002] In typical traffic scenarios, traffic signs are placed on roads or lanes to guide vehicles. Real-time detection of traffic signs is one of the fundamental tasks in traffic safety fields such as autonomous driving and driver assistance systems. In recent years, traffic sign detection technology applied to traffic safety and navigation has become a research hotspot in the field of traffic safety. Related technologies can be divided into two main categories: On the one hand, sensors such as millimeter-wave radar, lidar, and binocular stereo cameras are used to acquire information such as the distance, speed, azimuth, and 3D point cloud or depth information of objects in the surrounding environment. Based on physical or geometric principles, the objects in the scene are analyzed and detected to obtain results. On the other hand, visible light images are acquired through cameras, and based on the imaging characteristics of target objects, objects in the visible light images are identified to obtain results.
[0003] Existing methods for identifying traffic signs require calculating the situation of all objects in the scene or using complex visible light image recognition algorithms, which involve a large amount of computation and are not suitable for real-world scenarios where rapid identification of traffic signs is required. Summary of the Invention
[0004] This application discloses a method for recognizing traffic signs, an electronic device, a vehicle, and a storage medium. It can identify traffic signs in an image based on their color, geometric, and texture features. The method has low computational complexity and is suitable for scenarios where vehicles can quickly identify traffic signs while in motion.
[0005] In a first aspect, embodiments of this application provide a method for recognizing traffic signs. This method is applied to an electronic device and includes: identifying pixels of a target color from an image to be recognized, obtaining at least one block, the block consisting of multiple connected pixels of the same color, the target color being one or more colors of a target traffic sign; clustering the at least one block according to its position to obtain at least one candidate instance, the candidate instance being a combination of one or more blocks from the at least one block; extracting features from each candidate instance; selecting a target instance from the at least one candidate instance based on its respective features; and recognizing the traffic sign corresponding to the target instance.
[0006] Implementing the method provided in the first aspect, the electronic device identifies blocks containing pixels of the target color from the image to be recognized based on the color features of traffic signs. These blocks are then clustered to obtain candidate instances. Further, the electronic device filters out target instances from the candidate instances based on the features of the traffic signs. At this point, the target instance is determined, and finally, the traffic sign corresponding to the target instance is identified, and the recognition result of the target instance is output. In this way, since each step only requires calculation based on the results obtained from the previous step, the computational load in each step is relatively small, thus shortening the overall recognition time for traffic signs in the image.
[0007] In conjunction with the first aspect, in some embodiments, the features include geometric features, and the step of filtering target instances from the at least one candidate instances based on their respective features includes: determining a first matching degree between the geometric features of each candidate instance and the geometric template corresponding to each candidate instance, wherein the geometric template is a geometric feature used to describe the target traffic sign; identifying candidate instances among the at least one candidate instances whose first matching degree is greater than a first target threshold as target instances; and identifying the traffic sign corresponding to the target instance includes: determining the identification result of the target instance based on the traffic sign described by the geometric template corresponding to the target instance.
[0008] By matching the geometric features of candidate instances with geometric templates, there is no need to train on the geometric features of the target traffic sign. That is, there is no need to collect data and label ground truth for the target traffic sign and its category. Thus, the geometric template matching method has a low dependence on the amount and distribution of data. Moreover, it can still obtain relatively accurate matching results when the overall shape of the target traffic sign does not change significantly.
[0009] In conjunction with the first aspect, in some embodiments, the features include geometric features, and the method further includes: extracting the geometric features of each candidate instance; the step of selecting a target instance from the at least one candidate instance based on the features of each candidate instance includes: inputting the geometric features of each candidate instance into a first classifier to obtain a first classification result for each candidate instance, the first classification result being used to indicate the type of traffic sign corresponding to the input geometric features, the first classifier being trained with a first sample instance as input, the true result of the first sample instance being obtained by label training, the true result being used to indicate the traffic sign corresponding to the first sample instance; determining the candidate instances among the at least one candidate instances whose first classification result is a traffic sign as the target instance; the step of identifying the traffic sign corresponding to the target instance includes: determining the identification result of the target instance based on the first classification result of the target instance.
[0010] In conjunction with the first aspect, in some embodiments, the features include texture features, and the step of filtering target instances from the at least one candidate instances based on their respective features includes: determining a texture template corresponding to each of the at least one candidate instances; the texture template being a texture feature used to describe the target traffic sign; determining a second matching degree between the texture feature of each candidate instance and the texture template corresponding to each candidate instance, the texture template being the texture feature of the traffic sign; identifying candidate instances among the at least one candidate instances whose second matching degree is greater than a second target threshold as target instances; and identifying the traffic sign corresponding to the target instance includes: determining the identification result of the target instance based on the traffic sign described by the texture template corresponding to the target instance.
[0011] By matching the texture features of candidate instances with texture templates, there is no need to train on the texture features of the target traffic sign. That is, there is no need to collect data and label ground truth for the target traffic sign and its category. Thus, the texture template matching method has a low dependence on the amount and distribution of data. Furthermore, even if there are some changes in the appearance of the target traffic sign, as long as the texture does not change significantly, a relatively accurate matching result can still be obtained.
[0012] In conjunction with the first aspect, in some embodiments, the features include texture features, and the method further includes: extracting texture features of each candidate instance; selecting a target instance from the at least one candidate instance based on the features of each candidate instance includes: inputting the texture features of each candidate instance into a second classifier to obtain a second classification result for each candidate instance, the second classification result indicating the type of traffic sign corresponding to the input texture features, the second classifier being trained using the texture features of a second sample instance as input and the true result of the second sample instance as the label, the true result being used for the traffic sign corresponding to the second sample instance; determining the candidate instances among the at least one candidate instances whose second classification result is a traffic sign as the target instance; and identifying the traffic sign corresponding to the target instance includes: determining the identification result of the target instance based on the second classification result of the target instance.
[0013] In conjunction with the first aspect, in some embodiments, the features include geometric features and texture features, and the method further includes: extracting the geometric features and texture features of each candidate instance; the step of filtering out target instances from the at least one candidate instance based on the features of each candidate instance includes: inputting the geometric features and texture features of each candidate instance into a third classifier to obtain a third classification result for each candidate instance, wherein the third classification result is used to indicate the type of traffic sign corresponding to the input geometric features and texture features, the third classifier is trained with the geometric features and texture features of a third sample instance as input, and the true result of the third sample instance is used to indicate the traffic sign corresponding to the third sample instance; the candidate instances whose third classification result is a traffic sign among the at least one candidate instances are determined as target instances; the step of identifying the traffic sign corresponding to the target instance includes: determining the identification result of the target instance as the traffic sign indicated by the third classification result of the target instance.
[0014] By extracting the geometric and texture features of candidate instances and inputting these features into a classifier, the classification results of the candidate instances can be obtained directly, enabling rapid identification of target traffic signs.
[0015] In conjunction with the first aspect, in some embodiments, determining the identification result of the target instance based on the traffic sign described by the geometric template corresponding to the target instance includes: determining the identification result of the target instance as the traffic sign described by the geometric template corresponding to the target instance.
[0016] In conjunction with the first aspect, in some embodiments, determining the recognition result of the target instance based on the traffic sign described by the geometric template corresponding to the target instance includes: extracting the texture features of the target instance; inputting the texture features of the target instance into a fourth classifier to obtain a fourth classification result of the target instance, wherein the fourth classification result is used to indicate the type of traffic sign corresponding to the input texture features, the fourth classifier is trained with the texture features of a fourth sample instance as input, and the true result of the fourth sample instance is used to indicate the traffic sign corresponding to the fourth sample instance; when the traffic sign indicated by the fourth classification result is the same as the traffic sign described by the geometric template corresponding to the target instance, the recognition result of the target instance is determined to be the traffic sign described by the geometric template corresponding to the target instance.
[0017] In conjunction with the first aspect, in some embodiments, determining the first matching degree between the geometric features of each candidate instance and the geometric template corresponding to each candidate instance includes: determining the aspect ratio of the convex hull of each candidate instance; determining the aspect ratio of the geometric template corresponding to each candidate instance; and determining a first ratio of the aspect ratio of the convex hull of each candidate instance to the aspect ratio of the geometric template corresponding to each candidate instance as the first matching degree.
[0018] In conjunction with the first aspect, in some embodiments, determining the first matching degree between the geometric features of each candidate instance and the geometric template corresponding to each candidate instance includes: determining the area of the convex hull of each instance; determining the area of the geometric template corresponding to each candidate instance; and determining a second ratio of the area of the convex hull of each candidate instance to the area of the geometric template corresponding to each candidate instance as the first matching degree.
[0019] In conjunction with the first aspect, in some embodiments, determining the first matching degree between the geometric features of each candidate instance and the geometric template corresponding to each candidate instance in the at least one candidate instance includes: aligning the convex hull of each instance with the geometric template corresponding to each candidate instance, such that the midpoint of the bottom edge of the convex hull of each instance coincides with the geometric template corresponding to each candidate instance; determining the intersection area of the area of the convex hull of each instance and the area of the geometric template corresponding to each candidate instance; determining the union area of the area of the convex hull of each instance and the area of the geometric template corresponding to each candidate instance; and determining a third ratio of the intersection area to the union area as the first matching degree.
[0020] In conjunction with the first aspect, in some embodiments, determining the first matching degree between the geometric features of each candidate instance and the geometric template corresponding to each candidate instance in the at least one candidate instance includes: extracting the first Huvgeny invariant moments of the geometric features of each candidate instance; extracting the second Huvgeny invariant moments of the geometric template corresponding to each candidate instance; and determining the similarity between the first Huvgeny invariant moments and the second Huvgeny invariant moments as the first matching degree.
[0021] In conjunction with the first aspect, in some embodiments, the method further includes: extracting the convex hull of each candidate instance, wherein the convex hull of each candidate instance is a geometric feature of each candidate instance.
[0022] In conjunction with the first aspect, in some embodiments, the method further includes: using the geometric template corresponding to the position of each candidate instance in the lookup table as the geometric template corresponding to each candidate instance, wherein the lookup table includes multiple positions and geometric templates corresponding to the multiple positions respectively.
[0023] Storing a lookup table can quickly determine the geometric template corresponding to the candidate instance. This process involves relatively little computation and is beneficial for obtaining recognition results quickly.
[0024] In conjunction with the first aspect, in some embodiments, the image to be identified is captured by a target device, and the method further includes: determining a second position of each candidate instance in the world coordinate system based on the position of the target device in the world coordinate system, a first position of each candidate instance in the image to be identified, and depth of field; determining a transformation relationship between the coordinate system of the image to be identified and the world coordinate system based on the first position and the second position; and determining the geometric template corresponding to each candidate instance as the projection of the three-dimensional model of the traffic sign in the image to be identified based on the three-dimensional model of the traffic sign at the second position and the transformation relationship.
[0025] By generating geometric templates corresponding to candidate instances in real time, the geometric templates for candidate instances are determined based on the transformation relationship between the position of the candidate instance in the image to be recognized and its position on the ideal ground plane. This method can quickly obtain the geometric templates corresponding to candidate instances and save memory space.
[0026] In conjunction with the first aspect, in some embodiments, the step of clustering the at least one block according to its location to obtain at least one candidate instance includes: determining the distance between every two blocks in the at least one block according to its location; and grouping two blocks with a distance less than a first threshold into one candidate instance to obtain the at least one candidate instance.
[0027] In conjunction with the first aspect, in some embodiments, determining the recognition result of the target instance based on the traffic sign described by the geometric template corresponding to the target instance includes: extracting the texture features of the target instance; determining the texture template of the target instance; determining a third matching degree between the texture features of the target instance and the texture template corresponding to the target instance, wherein the texture template is the texture features of the traffic sign; determining a first matching result of the target instance where the third matching degree is greater than a third target threshold, wherein the matching result is the traffic sign described by the texture template corresponding to the target instance; and determining the recognition result of the target instance as the traffic sign described by the geometric template corresponding to the target instance when the first matching result is the same as the traffic sign described by the geometric template corresponding to the target instance.
[0028] After determining the matching degree between the target instance and the corresponding geometric template, the classification result of the target instance is confirmed by the classifier. Only when the traffic sign indicated by the classification result matches the traffic sign described by the geometric template corresponding to the target instance is the recognition result of the target instance determined. This improves the accuracy of recognition.
[0029] In conjunction with the first aspect, in some embodiments, determining the recognition result of the target instance based on the traffic sign described by the texture template corresponding to the target instance includes: extracting the geometric features of the target instance; determining a fourth matching degree between the geometric features of the target instance and the geometric template corresponding to the target instance, wherein the geometric template is the geometric features of the traffic sign; determining a second matching result of the target instance where the fourth matching degree is greater than a fourth target threshold, wherein the second matching result is the traffic sign described by the geometric template corresponding to the target instance; and determining the recognition result of the target instance as the traffic sign described by the texture template corresponding to the target instance when the second matching result is the same as the traffic sign described by the texture template corresponding to the target instance.
[0030] In conjunction with the first aspect, in some embodiments, determining the identification result of the target instance based on the second classification result of the target instance includes: extracting the geometric features of the target instance; determining a fourth matching degree between the geometric features of the target instance and the geometric template corresponding to the target instance, wherein the geometric template is the geometric features of the traffic sign; determining a second matching result of the target instance where the fourth matching degree is greater than a fourth target threshold, wherein the second matching result is the traffic sign described by the geometric template corresponding to the target instance; and determining the identification result of the target instance as the traffic sign indicated by the second classification result when the traffic sign indicated by the second matching result and the second classification result are the same.
[0031] In conjunction with the first aspect, in some embodiments, determining the recognition result of the target instance based on the traffic sign described by the texture template corresponding to the target instance includes: extracting the geometric features of the target instance; inputting the geometric features of the target instance into a fifth classifier to obtain a fifth classification result of the target instance, wherein the fifth classification result is used to indicate the type of traffic sign corresponding to the input geometric features, the fifth classifier is obtained by training the true result of the fifth sample instance with labels, and the true result is used to indicate the traffic sign corresponding to the fifth sample instance; when the traffic sign indicated by the fifth classification result is the same as the traffic sign described by the texture template corresponding to the target instance, the recognition result of the target instance is determined to be the traffic sign described by the texture template corresponding to the target instance.
[0032] In conjunction with the first aspect, in some embodiments, determining the identification result of the target instance based on the second classification result of the target instance includes: extracting the geometric features of the target instance; inputting the geometric features of the target instance into a sixth classifier to obtain a sixth classification result of the target instance, wherein the sixth classification result is used to indicate the type of traffic sign corresponding to the input geometric features, the sixth classifier is obtained by training the true result of the sixth sample instance with labels, and the true result is used to indicate the traffic sign corresponding to the sixth sample instance; when the traffic sign indicated by the fifth classification result is the same as the traffic sign indicated by the second classification result of the target instance, the identification result of the target instance is determined to be the traffic sign indicated by the second classification result of the target instance.
[0033] The traffic sign recognition method provided by the first aspect or any possible implementation thereof does not rely on the optical flow of traffic signs in front of or in other directions while the vehicle is in motion. Effective detection of traffic signs can still be achieved when the vehicle is stationary and other stationary objects in the scene cannot generate optical flow, or when the vehicle is moving and the traffic sign is directly in front of the vehicle and its motion is stationary or parallel to the vehicle's direction of motion. Because this method does not rely on the motion state of the traffic sign relative to the vehicle, but rather utilizes the inherent properties and characteristics of the traffic sign, it exhibits good stability.
[0034] In a second aspect, embodiments of this application provide an electronic device, including: one or more processors and one or more memories, the one or more memories being coupled to the one or more processors respectively; the one or more memories being used to store computer program code, the computer program code including computer instructions; the processor being used to invoke the computer instructions to execute: a method for traffic sign recognition as described in the first aspect and any possible implementation thereof.
[0035] Thirdly, embodiments of this application provide a vehicle, including: one or more processors and one or more memories, the one or more memories being coupled to the one or more processors respectively; the one or more memories being used to store computer program code, the computer program code including computer instructions; the processor being used to invoke the computer instructions to execute: a method for traffic sign recognition as described in the first aspect and any possible implementation thereof.
[0036] Fourthly, embodiments of this application provide a traffic sign recognition device, comprising the following units: The first recognition unit 1101 is used to recognize pixels of a target color from an image to be recognized, and obtain at least one block, wherein the block is composed of multiple connected pixels, the multiple pixels have the same color, and the target color is one or more colors of a target traffic sign. Clustering unit 1102 is configured to cluster the at least one block according to the location of the at least one block to obtain at least one candidate instance, wherein the candidate instance is a combination of one or more blocks among the at least one block; Extraction unit 1103 is used to extract features of each candidate instance for each of the at least one candidate instance; The filtering unit 1104 is used to filter out a target instance from the at least one candidate instance based on the characteristics of each of the at least one candidate instance; The second identification unit 1105 is used to identify the traffic signs corresponding to the target instance.
[0037] In one possible implementation, the features include geometric features. The filtering unit 1104 is specifically used to: determine a first matching degree between the geometric features of each candidate instance in the at least one candidate instance and the geometric template corresponding to each candidate instance, wherein the geometric template is a geometric feature used to describe the target traffic sign; and determine the candidate instances in the at least one candidate instance whose first matching degree is greater than a first target threshold as target instances. The filtering unit 1104 is specifically used to: determine the recognition result of the target instance based on the traffic signs described by the geometric template corresponding to the target instance.
[0038] In one possible implementation, the features include geometric features. The extraction unit 1103 is specifically used to: extract the geometric features of each candidate instance; The filtering unit 1104 is specifically used for: inputting the geometric features of each candidate instance into a first classifier to obtain a first classification result for each candidate instance, wherein the first classification result is used to indicate the type of traffic sign corresponding to the input geometric features, the first classifier is trained with a first sample instance as input and the true result of the first sample instance as a label, wherein the true result is used to indicate the traffic sign corresponding to the input geometric features; and determining the candidate instance whose first classification result is a traffic sign among the at least one candidate instance as the target instance; The second identification unit 1105 is specifically used to: determine the identification result of the target instance based on the first classification result of the target instance.
[0039] In one possible implementation, the features include texture features. The filtering unit 1104 is specifically used for: determining a texture template corresponding to each candidate instance in the at least one candidate instance; the texture template is a texture feature used to describe the target traffic sign; determining a second matching degree between the texture feature of each candidate instance and the texture template corresponding to each candidate instance; and determining the candidate instances in the at least one candidate instance whose second matching degree is greater than a second target threshold as target instances; The second identification unit 1105 is specifically used to: determine the identification result of the target instance based on the traffic sign described by the texture template corresponding to the target instance.
[0040] In one possible implementation, the extraction unit 1103 is specifically used to: extract the texture features of each candidate instance; The filtering unit 1104 is specifically used for: inputting the texture features of each candidate instance into a second classifier to obtain a second classification result for each candidate instance, wherein the second classification result is used to indicate the type of traffic sign corresponding to the input texture features, the second classifier is trained using the texture features of a second sample instance as input and the true result of the second sample instance as the label, wherein the true result is used to indicate the traffic sign corresponding to the second sample instance; and determining the candidate instance whose second classification result is a traffic sign among the at least one candidate instance as the target instance; The second identification unit 1105 is specifically used to: determine the identification result of the target instance based on the second classification result of the target instance.
[0041] In one possible implementation, the extraction unit 1103 is specifically used to: extract the geometric features and texture features of each candidate instance; The filtering unit 1104 is specifically used to: input the geometric features and texture features of each candidate instance into a third classifier to obtain a third classification result for each candidate instance, wherein the third classification result is used to indicate the type of traffic sign corresponding to the input geometric features and texture features, wherein the third classifier is trained with the geometric features and / or texture features of a third sample instance as input, and the true result of the third sample instance is used to indicate the traffic sign corresponding to the third sample instance; and to determine the candidate instance whose third classification result is a traffic sign among the at least one candidate instance as the target instance; The second identification unit 1105 is specifically used to: determine the identification result of the target instance as a traffic sign indicated by the third classification result of the target instance.
[0042] In one possible implementation, the second identification unit 1105 is specifically used to: determine the identification result of the target instance as the traffic sign described by the geometric template corresponding to the target instance.
[0043] In one possible implementation, the extraction unit 1103 is specifically used to: extract the texture features of the target instance; The second identification unit 1105 is specifically used for: inputting the texture features of the target instance into a fourth classifier to obtain a fourth classification result of the target instance, wherein the fourth classification result is used to indicate the type of traffic sign corresponding to the input texture features, wherein the fourth classifier is obtained by training the true result of the fourth sample instance with labels, and the true result is used to indicate the traffic sign corresponding to the fourth sample instance; and, when the traffic sign indicated by the fourth classification result is the same as the traffic sign described by the geometric template corresponding to the target instance, determining that the identification result of the target instance is the traffic sign described by the geometric template corresponding to the target instance.
[0044] In one possible implementation, the filtering unit 1104 is specifically configured to: determine the aspect ratio of the convex hull of each candidate instance; determine the aspect ratio of the geometric template corresponding to each candidate instance; and determine a first ratio of the aspect ratio of the convex hull of each candidate instance to the aspect ratio of the geometric template corresponding to each candidate instance as the first matching degree.
[0045] In one possible implementation, the filtering unit 1104 is specifically configured to: determine the area of the convex hull of each instance; determine the area of the geometric template corresponding to each candidate instance; and determine a second ratio of the area of the convex hull of each candidate instance to the area of the geometric template corresponding to each candidate instance as the first matching degree.
[0046] In one possible implementation, the filtering unit 1104 is specifically configured to: align the convex hull of each instance with the geometric template corresponding to each candidate instance, such that the midpoint of the bottom edge of the convex hull of each instance coincides with the geometric template corresponding to each candidate instance; determine the intersection area of the area of the convex hull of each instance and the area of the geometric template corresponding to each candidate instance; determine the union area of the area of the convex hull of each instance and the area of the geometric template corresponding to each candidate instance; and determine a third ratio of the intersection area to the union area as the first matching degree.
[0047] In one possible implementation, the filtering unit 1104 is specifically used to: extract the first Huvren invariant moment of the geometric features of each candidate instance; extract the second Huvren invariant moment of the corresponding geometric template of each candidate instance; and determine the similarity between the first Huvren invariant moment and the second Huvren invariant moment as the first matching degree.
[0048] In one possible implementation, the features include geometric features, and the extraction unit 1103 is specifically used to: extract the convex hull of each candidate instance, wherein the convex hull of each candidate instance is the geometric feature of each candidate instance.
[0049] In one possible implementation, the filtering unit 1104 is further specifically used to: use the geometric template corresponding to the position of each candidate instance in the lookup table as the geometric template corresponding to each candidate instance, wherein the lookup table includes multiple positions and geometric templates corresponding to the multiple positions respectively.
[0050] In one possible implementation, the image to be identified is captured by a target device, and the filtering unit 1104 is further configured to: determine a second position of each candidate instance in the world coordinate system based on the position of the target device in the world coordinate system, the first position of each candidate instance in the image to be identified, and the depth of field; determine the transformation relationship between the coordinate system of the image to be identified and the world coordinate system based on the first position and the second position; and determine the geometric template corresponding to each candidate instance as the projection of the three-dimensional model in the image to be identified based on the three-dimensional model of the traffic sign at the second position and the transformation relationship.
[0051] In one possible implementation, the clustering unit 1102 is specifically used to: determine the distance between every two blocks in the at least one block according to the location of the at least one block; and divide two blocks with a distance less than a first threshold into a candidate instance to obtain the at least one candidate instance.
[0052] Fifthly, embodiments of this application provide a computer storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform a traffic sign recognition method as described in the first aspect and any possible implementation thereof.
[0053] Understandably, the electronic device provided in the second aspect, the vehicle provided in the third aspect, the traffic sign recognition device provided in the fourth aspect, and the computer storage medium provided in the fifth aspect are all used to execute the method provided in the first aspect and any possible implementation thereof. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0054] Figure 1A This is a schematic diagram of the system involved in a traffic sign recognition method provided in an embodiment of this application; Figure 1B This is a functional block diagram of a vehicle provided in an embodiment of this application; Figures 2A-2B This is a schematic flowchart of a traffic sign recognition method provided in an embodiment of this application; Figures 3A-3G This is a schematic diagram illustrating the principle of some traffic cone recognition methods provided in the embodiments of this application; Figures 4A-4H This is a schematic diagram illustrating the principle of some triangular warning sign recognition methods provided in the embodiments of this application; Figure 5 This is a schematic diagram of the HSV color space involved in a traffic sign recognition method provided in an embodiment of this application; Figure 6 These are example diagrams of traffic cones and warning triangles provided in embodiments of this application; Figure 7 These are example diagrams of traffic cones and warning triangles provided in embodiments of this application, and their corresponding geometric templates; Figures 8A-8B This is a schematic diagram of a real-time geometric template generation process provided in an embodiment of this application; Figure 9 This is a schematic diagram illustrating how to determine a target instance and its classification result using the geometric and texture features of candidate instances, as provided in an embodiment of this application. Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 11 This is a schematic diagram of the structure of a traffic sign recognition device provided in an embodiment of this application. Detailed Implementation
[0055] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.
[0056] The concepts and terms used in this application are described below.
[0057] RGB color space: RGB stands for the three primary colors: red, green, and blue. It is a method of representing color. The common RGB color space is a cube of one unit length. The eight common colors, black, blue, green, cyan, red, purple, yellow, and white, are located at the eight vertices of the cube. Black is usually placed at the origin of a three-dimensional rectangular coordinate system, and red, green, and blue are placed on the three coordinate axes respectively.
[0058] HSV color space: HSV stands for Hue, Saturation, and Value, and is a method of representing color. The common HSV color space is an inverted cone, with the central axis of the cone representing black at the bottom and white at the top, and gray in between. The angle around this axis corresponds to "hue", and the distance from this axis corresponds to "saturation".
[0059] Density-based spatial clustering of applications with noise (DBSCAN) is a clustering algorithm. This algorithm divides regions with sufficient density into clusters and discovers clusters of arbitrary shapes in a noisy spatial database. It defines a cluster as the largest set of density-connected points. The algorithm utilizes the concept of density-based clustering, which requires that the number of objects (points or other spatial objects) contained within a certain region of the cluster space is not less than a given threshold. The significant advantages of the DBSCAN algorithm are its fast clustering speed and its ability to effectively handle noisy points and discover spatial clusters of arbitrary shapes.
[0060] Convex hull: Given a set of points, the convex hull of this set is the convex polygon with the smallest area containing all the points in the set. Intuitively, a convex polygon is a polygon without any indentations.
[0061] Haar-like features are digital image features used for object detection and recognition. They are named for their striking similarity to the Haar wavelet transform and were initially applied to real-time face detection. Haar features use neighboring rectangles at a specified location within the detection window, calculating the sum of pixels for each rectangle and taking the difference. These differences are then used to classify sub-regions of the image. The main advantage of Haar features is their extremely fast computation; using a structure called an integral image, Haar features of any size can be computed in constant time.
[0062] Connected component analysis: A connected component in an image is a region consisting of pixels with the same pixel value and adjacent positions. Connected component analysis is the process of finding and marking mutually independent connected components in an image.
[0063] Histogram of Oriented Gradient (HOG) features are feature descriptors used in computer vision and image processing for object detection. They are constructed by calculating and statistically analyzing the histograms of gradient orientations in local image regions. HOG features combined with SVM classifiers have been widely applied in image recognition, achieving great success, particularly in pedestrian detection.
[0064] Support Vector Machine (SVM) classifier: It is a binary classification model. Its basic model is defined as a linear classifier with the largest margin in the feature space. Its learning strategy is to maximize the margin, which can be transformed into solving a convex quadratic programming problem.
[0065] Lab color space: a method of representing color. Lab consists of one luminance channel and two color channels. In the Lab color space, each color is represented by three numbers: L, a, and b. The meanings of each component are: L Represents brightness, a Representing the proportions from green to red, b It represents the range from blue to yellow.
[0066] HDBSCAN Clustering Algorithm: An optimized algorithm based on DBSCAN clustering. It eliminates the need for manual selection of neighborhood radii R and MinPts; most of the time, only the minimum generated cluster size needs to be selected, and the algorithm automatically recommends the optimal cluster results. It also defines a new distance metric that better reflects point density.
[0067] Hu's invariant moments: Moments are a concept in probability and statistics, representing a numerical characteristic of a random variable. For an image, if we consider the coordinates of pixels as a two-dimensional random variable, then a grayscale image can be represented by a two-dimensional grayscale density function. Therefore, moments can be used to describe the characteristics of a grayscale image. Hu constructed seven invariant moments using second- and third-order normalized central moments. Invariant moments are highly condensed image features, exhibiting translation, grayscale, scale, and rotation invariance in continuous images.
[0068] The following is a schematic diagram illustrating the traffic sign recognition system according to an embodiment of this application, such as... Figure 1A As shown, the system includes a mobile device 10 and a server 20. The mobile device can be a vehicle 11, a vehicle 12, a robot 13, a mobile terminal 14, or other mobile devices.
[0069] The following uses a vehicle 11 as an example of a mobile device 10 to illustrate the traffic sign recognition system. The vehicle 11 can acquire images in real time from the direction it is traveling in, its surrounding environment, or its direction it is traveling in. In this recognition system, traffic signs can be recognized from multiple consecutive frames of images, or a single frame can be selected from these multiple frames according to a preset number of frames for recognition. This application embodiment does not limit the scope of the method. This application embodiment uses one frame as an example to illustrate the method for recognizing traffic signs in each frame.
[0070] In some implementations, vehicle 11 can identify traffic signs in the acquired image to be identified. Vehicle 11 uses the color of the traffic sign as the target color in the image to be identified, obtaining candidate instances. Then, vehicle 11 further filters and classifies the candidate instances based on the geometric and / or texture features of the traffic sign, ultimately determining the identification result of the image to be identified. This result can be whether the image to be identified contains traffic signs and the type of traffic sign. Specific identification methods can be found below. Figure 2A and Figure 2B The steps and related descriptions shown will not be repeated here.
[0071] The method for recognizing traffic signs in the image to be recognized can be performed by the vehicle 11 itself or by a device installed on the vehicle 11, such as one or more of an on-board unit, a vehicle navigation device, a vehicle driver assistance device, a vehicle automatic driving device, and a driving recorder.
[0072] In other implementations, vehicle 11 sends the acquired image to be identified to server 20. Server 20 identifies traffic signs in the received image and sends the identification result back to vehicle 11. Specific identification methods are detailed below. Figure 2A and Figure 2B The steps and related descriptions shown will not be repeated here.
[0073] It should be understood that traffic signs may include, for example: Figure 6 The traffic cone 101 and warning triangle 102 shown may also include prohibition signs and instruction signs, etc. This application embodiment uses... Figure 6 The traffic cone 101 and the warning triangle 102 shown are used as examples to illustrate the point. This application does not limit the type of traffic signs.
[0074] Please see Figure 1B , Figure 1B This application provides a functional block diagram of a vehicle 002 according to an embodiment. The vehicle 002 can be as described above. Figure 1A Vehicle 11 in the system shown.
[0075] In one embodiment, vehicle 002 can be configured in a fully or partially autonomous driving mode. For example, vehicle 002 can control itself while in autonomous driving mode, and can determine the current state of the vehicle and its surrounding environment through human intervention, determine the possible behaviors of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the probability of that other vehicle performing the possible behavior, and control vehicle 002 based on the determined information. When vehicle 002 is in autonomous driving mode, vehicle 002 can be configured to operate without human interaction.
[0076] Vehicle 002 may include various subsystems, such as a mobility system 202, a sensor system 204, a control system 206, one or more peripheral devices 208, a power supply 210, a computer system 212, and a user interface 216. Optionally, vehicle 002 may include more or fewer subsystems, and each subsystem may include multiple components. Furthermore, each subsystem and component of vehicle 002 may be interconnected via wired or wireless means.
[0077] The propulsion system 202 may include components that provide powered motion to the vehicle 002. In one embodiment, the propulsion system 202 may include an engine 218, an energy source 219, a transmission 220, and wheels 221. The engine 218 may be an internal combustion engine, an electric motor, an air-compressed engine, or other types of engine combinations, such as a hybrid engine consisting of a gasoline engine and an electric motor, or a hybrid engine consisting of an internal combustion engine and an air-compressed engine. The engine 218 converts the energy source 219 into mechanical energy.
[0078] Examples of energy sources 219 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. Energy source 219 may also provide energy to other systems of vehicle 002.
[0079] The transmission 220 can transmit mechanical power from the engine 218 to the wheels 221. The transmission 220 may include a gearbox, a differential, and a drive shaft. In one embodiment, the transmission 220 may also include other components, such as a clutch. The drive shaft may include one or more axles that can be coupled to one or more wheels 221.
[0080] Sensor system 204 may include several sensors for sensing information about the environment surrounding vehicle 002. For example, sensor system 204 may include a global positioning system 222 (which may be GPS, BeiDou, or other positioning systems), an inertial measurement unit (IMU) 224, radar 226, a laser rangefinder 228, and a camera 230. Sensor system 204 may also include sensors for the internal systems of the monitored vehicle 002 (e.g., in-vehicle air quality monitor, fuel gauge, oil temperature gauge, etc.). Sensor data from one or more of these sensors can be used to detect objects and their corresponding characteristics (position, shape, orientation, speed, etc.). This detection and identification is a key function for the safe operation of the autonomous vehicle 002.
[0081] The Global Positioning System 222 can be used to estimate the geographic location of vehicle 002. The IMU 224 is used to sense changes in the position and orientation of vehicle 002 based on inertial acceleration. In one embodiment, IMU 224 can be a combination of an accelerometer and a gyroscope. For example, IMU 224 can be used to measure the curvature of vehicle 002.
[0082] Radar 226 can use radio signals to sense objects in the surrounding environment of vehicle 002. In some embodiments, in addition to sensing objects, radar 226 can also be used to sense the speed and / or direction of travel of objects.
[0083] The laser rangefinder 228 can use lasers to sense objects in the environment in which the vehicle 002 is located. In some embodiments, the laser rangefinder 228 may include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components.
[0084] Camera 230 can be used to capture multiple images of the surrounding environment of vehicle 002. Camera 230 can be a still camera or a video camera, or a visible light camera or an infrared camera; it can be any camera used to acquire images, and this application embodiment does not limit this.
[0085] In this embodiment, the camera 230 can be mounted on the front, rear, and left and right sides of the vehicle 002. The camera 230 can be a camera whose shooting angle can be adjusted by rotation. Additionally, the camera in this embodiment can also be used to acquire images to be identified.
[0086] The control system 206 controls the operation of the vehicle 002 and its components. The control system 206 may include various elements, including a steering unit 232, a throttle 234, a braking unit 236, a sensor fusion algorithm unit 238, a computer vision system 240, a route control system 242, and an obstacle avoidance system 244.
[0087] The steering unit 232 is operable to adjust the forward direction of the vehicle 002. For example, in one embodiment, it can be a steering wheel system.
[0088] Throttle 234 is used to control the operating speed of engine 218 and thus the speed of vehicle 002.
[0089] Braking unit 236 is used to control the deceleration of vehicle 002. Braking unit 236 can use friction to slow down wheel 221. In other embodiments, braking unit 236 can convert the kinetic energy of wheel 221 into electric current. Braking unit 236 may also take other forms to slow down the rotational speed of wheel 221 to control the speed of vehicle 002.
[0090] The computer vision system 240 is operable to process and analyze images captured by the camera 230 to identify objects and / or features in the environment surrounding the vehicle 002. The objects and / or features may include traffic signals, road boundaries, and obstacles. The computer vision system 240 may use object recognition algorithms, Structure from Motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, the computer vision system 240 may be used to map the environment, track objects, estimate object speeds, etc. In embodiments of this application, the computer vision system 240 can also be used to identify traffic signs; specific methods for identifying traffic signs are described below. Figure 2A The relevant descriptions in the illustrated embodiments.
[0091] The route control system 242 is used to determine the driving route of the vehicle 002. In some embodiments, the route control system 242 may combine data from the sensor fusion algorithm unit 238, the global positioning system 222, and one or more predetermined maps to determine the driving route for the vehicle 002.
[0092] The obstacle avoidance system 244 is used to identify, assess and avoid or otherwise traverse potential obstacles in the environment of the vehicle 002.
[0093] Of course, in one instance, the control system 206 may include additional or alternative components besides those shown and described. Alternatively, some of the components shown above may be reduced.
[0094] Vehicle 002 interacts with external sensors, other vehicles, other computer systems, or users via peripheral device 208. Peripheral device 208 may include wireless communication system 246, on-board computer 248, microphone 250, and / or speaker 252.
[0095] In some embodiments, peripheral device 208 provides a means for a user of vehicle 002 to interact with user interface 216. For example, on-board computer 248 may provide information to the user of vehicle 002. User interface 216 may also operate on-board computer 248 to receive user input. On-board computer 248 may be operated via touchscreen. In other cases, peripheral device 208 may provide a means for vehicle 002 to communicate with other devices located within the vehicle. For example, microphone 250 may receive audio (e.g., voice commands or other audio input) from the user of vehicle 002. Similarly, speaker 252 may output audio to the user of vehicle 002.
[0096] Wireless communication system 246 can communicate wirelessly with one or more devices directly or via a communication network. Examples include code division multiple access (CDMA), enhanced versatile disk (EVD), global system for mobile communications (GSM) / general packet radio service (GPRS), 4G cellular communication such as long term evolution (LTE), 5G cellular communication, new radio (NR) systems, or future communication systems. Wireless communication system 246 can communicate using WiFi and wireless local area network (WLAN). In some embodiments, wireless communication system 246 can communicate directly with devices using an infrared link, Bluetooth, or ZigBee. Other wireless protocols, such as various vehicle communication systems, are also possible. For example, wireless communication system 246 may include one or more dedicated short range communications (DSRC) devices that can enable public and / or private data communication between vehicles and / or roadside stations.
[0097] Power source 210 can provide power to various components of vehicle 002. In one embodiment, power source 210 can be a rechargeable lithium-ion or lead-acid battery. One or more such battery packs can be configured to provide power to various components of vehicle 002. In some embodiments, power source 210 and energy source 219 can be implemented together, as is the case in some fully electric vehicles.
[0098] Some or all of the functions of vehicle 002 are controlled by computer system 212. Computer system 212 may include at least one processor 213, which executes instructions 215 stored in a non-transitory computer-readable medium such as a data storage device. Computer system 212 may also be multiple computing devices that control individual components or subsystems of vehicle 002 in a distributed manner.
[0099] Processor 213 can be any conventional processor, such as a commercially available CPU. Alternatively, the processor can be a special-purpose device such as an ASIC or other hardware-based processor. Although Figure 1B The illustrations functionally depict a processor, memory, and other components of a computer within the same block; however, those skilled in the art will understand that the processor, computer, or memory may actually include multiple processors, computers, or memories that may or may not be stored in the same physical enclosure. For example, memory may be a hard disk drive or other storage media located in an enclosure different from that of the computer. Therefore, references to a processor or computer will be understood to include references to a collection of processors or computers or memories that may or may not operate in parallel. Unlike using a single processor to perform the steps described herein, some components, such as steering and deceleration assemblies, may each have their own processor that performs calculations only relevant to the component's specific function.
[0100] In the various aspects described herein, the processor may be located remotely from the vehicle and communicate wirelessly with the vehicle. In other aspects, some of the processes described herein are executed on a processor located within the vehicle, while others are executed by a remote processor, including taking the necessary steps to perform a single operation.
[0101] In some embodiments, memory 214 may contain instructions 215 (e.g., program logic) that can be executed by processor 213 to perform various functions of vehicle 002, including those described above. The data storage device may also contain additional instructions, including instructions to send data to, receive data from, interact with, and / or control one or more of the mobility system 202, sensor system 204, control system 206, and peripheral devices 208.
[0102] In addition to instruction 215, memory 214 may also store data such as road maps, route information, vehicle position, direction, speed, and other vehicle data, as well as other information. This information can be used by vehicle 002 and computer system 212 during autonomous, semi-autonomous, and / or manual operation. For example, the current speed and curvature of the vehicle can be fine-tuned based on road information of the target road segment and the received vehicle speed and curvature ranges to keep the intelligent vehicle's speed and curvature within the vehicle speed and curvature ranges.
[0103] User interface 216 is used to provide information to or receive information from a user of vehicle 002. Optionally, user interface 216 may include one or more input / output devices within a set of peripheral devices 208, such as wireless communication system 246, on-board computer 248, microphone 250, and speaker 252.
[0104] Computer system 212 can control the functions of vehicle 002 based on input received from various subsystems (e.g., driving system 202, sensor system 204, and control system 206) and from user interface 216. For example, computer system 212 can utilize input from control system 206 to control steering unit 232 to avoid obstacles detected by sensor system 204 and obstacle avoidance system 244. In some embodiments, computer system 212 is operable to provide control over many aspects of vehicle 002 and its subsystems.
[0105] Optionally, one or more of these components may be installed separately from or associated with vehicle 002. For example, memory 214 may exist partially or completely separately from vehicle 002. The components may be communicatively coupled together in a wired and / or wireless manner.
[0106] Optionally, the components described above are merely examples. In actual applications, components in each of the above modules may be added or removed as needed. Figure 1B This should not be construed as a limitation on the embodiments of this application.
[0107] Autonomous vehicles traveling on roads, such as vehicle 002 above, can identify objects in their surrounding environment to determine adjustments to their current speed. These objects can be other vehicles, traffic control equipment, traffic signs, or other types of objects. In some examples, each identified object can be considered independently, and based on the object's individual characteristics, such as its current speed, acceleration, and distance from the vehicle, the speed adjustment to be made by the autonomous vehicle can be determined.
[0108] Optionally, the autonomous vehicle 002 or the computing device associated with the autonomous vehicle 002 (such as...) Figure 1B The computer system 212, computer vision system 240, and memory 214 can predict the behavior of the identified objects based on the characteristics of the identified objects and the state of the surrounding environment (e.g., traffic, rain, ice on the road, etc.). Optionally, each identified object depends on the behavior of each other, so all identified objects can also be considered together to predict the behavior of a single identified object. The vehicle 002 can adjust its speed based on the predicted behavior of the identified objects. In this process, other factors can also be considered to determine the speed of the vehicle 002, such as the lateral position of the vehicle 002 in the road, the curvature of the road, the proximity of static and dynamic objects, etc.
[0109] In addition to providing instructions to adjust the speed of the autonomous vehicle, the computing device can also provide instructions to modify the steering angle of the vehicle 002 so that the autonomous vehicle follows a given trajectory and / or maintains a safe lateral and longitudinal distance from objects near the autonomous vehicle (e.g., cars in adjacent lanes on the road).
[0110] The aforementioned vehicle 002 can be a car, truck, motorcycle, bus, ship, airplane, helicopter, lawnmower, recreational vehicle, amusement park vehicle, construction equipment, tram, golf cart, train, and handcart, etc., and this application embodiment does not impose any special limitations.
[0111] Understandable, Figure 1B The intelligent vehicle functional diagram in this application is only an exemplary implementation of the present application. The intelligent vehicle in the present application includes, but is not limited to, the above structure.
[0112] The following is combined Figures 2A-2B The method for identifying traffic signs provided in the embodiments of this application will be described. Figures 2A-2B This is a flowchart illustrating a method for recognizing traffic signs, which can be derived from the above... Figure 1A The system shown can be implemented, or it can be executed by an electronic device, which can be one of the aforementioned... Figure 1A The mobile device 10 or server 20 shown, or the one described above Figure 1B The vehicle shown is 002. This application embodiment uses... Figures 3A-3G The diagram illustrates the principle of the traffic cone identification method. Figures 4A-4H The schematic diagram illustrating the principle of the triangular warning sign recognition method is used as an example. It should be understood that the method for recognizing traffic signs provided in this application is not limited to traffic cones and triangular warning signs; it can also be used to recognize other traffic signs. For example... Figure 2A As shown, the method may include, but is not limited to, the following steps: S201. The electronic device acquires the image to be recognized.
[0113] In some implementations, the electronic device can be a vehicle, which can acquire images of the direction in front of the vehicle, the surrounding environment, or the rear of the vehicle in real time via a camera at a preset frequency. Traffic signs can be identified from multiple consecutive frames of images. The image to be identified can be any frame in the image set; or it can be a selected frame, for example, selecting one frame every 10 frames as the image to be identified. This application uses one frame (i.e., the image to be identified) as an example for illustration. The method for identifying traffic signs in each frame is the same as the method for identifying the image to be identified, and will not be repeated here.
[0114] In some implementations, the electronic device can be a server, which can receive images acquired by the vehicle through its camera in real time or periodically. For example, the vehicle sends an image frame to the server every 0.5 seconds, and the image to be identified can be any frame of the images received by the server; or, for another example, the vehicle sends each frame of the image to the server in real time, and the server can select from them, for example, selecting one frame every 10 frames as the image to be identified.
[0115] In some embodiments, after the electronic device acquires or receives the image to be identified, the image can be preprocessed to obtain an image with a specified resolution. For example, the original size of the image to be identified is 1920. 1080, downsampled by 2x to a size of 960 Image 540. (e.g.) Figure 3A and Figure 3B The images shown are all downsampled images to be identified. This application's embodiments are used to identify... Figure 3A Traffic cones and identification Figure 3B Let's take the triangular warning sign in the picture as an example.
[0116] S202. The electronic device identifies pixels of a target color from the image to be identified, and obtains at least one block. The block consists of multiple connected pixels with the same color. The target color is one or more colors of the target traffic sign.
[0117] The target traffic sign can be one or more traffic signs, i.e., the traffic signs that need to be identified. "Multiple" means two or more.
[0118] When the target traffic object is a traffic sign, the target color can be one or more colors included in the colors contained in that traffic sign, for example, such as... Figure 6As shown, the target traffic sign is a traffic cone. The national standard specifies that traffic cones should be decorated with alternating orange-red and white stripes. In this case, the target color can be orange-red, or it can be a combination of orange-red and white. Similarly, if the target traffic sign is a warning triangle, the national standard specifies that the warning triangle should be red. In this case, the target color can also be red.
[0119] When the target traffic sign consists of multiple traffic signs, such as traffic cones and warning triangles, the target color can include orange-red and red.
[0120] It should be noted that a color corresponds to a range of colors in a color space, and colors within that range are considered to be the same color. For example, in the RGB color space, the range of red color values can be P ≤ red color value ≤ Q. Colors with color values between P and Q are considered to be the same color and are all called red.
[0121] In one implementation, the electronic device identifies pixels of the target color from the image to be identified and obtains at least one block. One implementation can be: retaining pixels of the target color in the image to be identified, filtering out pixels of a color other than the target color, and forming a block with the retained connected pixels of the same color, thereby obtaining the above-mentioned at least one block.
[0122] It should be understood that in some embodiments, after the electronic device identifies the target color pixels from the image to be identified, the electronic device can also filter the obtained blocks. For example, the electronic device determines the number of pixels included in each block and filters out blocks with a number of pixels less than m, such as m being 5. The obtained blocks are the blocks that need to be clustered.
[0123] For example, converting the color space of the image to be recognized from RGB to HSV. Figure 5 This is a schematic diagram of the HSV color space. Figure 5 In the diagram, 'a' represents the cone model of the HSV color space. The HSV color space is described by three dimensions: hue (H), saturation (S), and lightness (V), as shown in the figure. The central axis of the cone extends from black at the bottom to white at the top, with gray in between. Lightness (V) is measured along this axis, with the angle around this axis corresponding to hue (H) and the distance to this axis corresponding to saturation (S). Hue represents color information, specifically the position of a color within the spectrum, such as... Figure 5 As shown in b, hue is measured in degrees, ranging from 0 to 360°; saturation represents how close a color is to a spectral color, ranging from 0 to 1; and lightness determines the brightness or darkness of a color in the color space, ranging from 0 to 1. For example... Figure 6As shown, the national standard specifies that traffic cones include orange-red in color, and the target color is orange-red when identifying traffic cones. In the HSV color space, the color range of orange-red includes a range across three dimensions: hue, saturation, and lightness. Figure 5 It is known that the hue, saturation, and lightness ranges for orange-red are 0° ≤ H ≤ 10° or 160° ≤ H ≤ 180°; 70 ≤ S ≤ 255; 100 ≤ V ≤ 255. When identifying pixels of the target color from the image to be identified, a pixel is retained only if its hue, saturation, and lightness fall within the above ranges. Figure 3B As shown, Figure 3B In order to address the above Figure 3A The result obtained after performing target color recognition on the image to be recognized in the HSV color space.
[0124] For example, the image to be recognized is an image in the RGB color space. Based on the national standard for traffic sign colors, the electronic device determines the weight range of the target color in the R, G, and B color channels, obtaining the minimum color value T_lower = a'×R + b'×G + c'×B and the maximum color value T_upper = a”×R + b”×G + c”×B. Therefore, the range of the target color value is determined as: T_lower ≤ T ≤ T_upper. Here, (a', b', c') represents the weight of the minimum color value of the target color in the R, G, and B color channels, and (a”, b”, c”) represents the weight of the maximum color value of the target color in the R, G, and B color channels. a', b', c', a”, b”, c” are all rational numbers. Further, the electronic device determines the color value Y = a×R + b×G + ... c×B, where (a, b, c) are the weights of the R, G, and B color channels, respectively, and a, b, and c are rational numbers. A pixel is retained only if its color value Y satisfies T_lower ≤ Y ≤ T_upper. For example... Figure 4B As shown, Figure 4B In order to address the above Figure 4A The result obtained after performing target color recognition on the image to be recognized in the RGB color space.
[0125] It should be noted that in traffic scenarios, the color information of traffic signs is standardized. Therefore, the colors in the image to be identified can be filtered based on the color characteristics of the traffic signs to determine the areas where the traffic signs may exist. In situations where color information is unstable due to changes in lighting conditions, such as severe weather, strong daylight, or insufficient nighttime illumination, the image to be identified can be converted to the HSV or Lab color space. The colors in the HSV and Lab color spaces are independent of light, significantly reducing the influence of lighting. Therefore, color space conversion can be used to obtain more accurate target color recognition results.
[0126] S203. The electronic device clusters the at least one block according to its location to obtain at least one candidate instance, wherein the at least one candidate instance includes M candidate instances, where M is a positive integer.
[0127] One possible clustering method is that the electronic device determines the distance between any two blocks based on the location of each block; then, two blocks with a distance less than a first threshold are grouped into the same candidate instance.
[0128] In specific implementations, algorithms such as density-based spatial clustering of applications with noise (DBSCAN), HDBSCAN, connected component analysis (CCA), and K-means clustering can be used to cluster blocks based on their location.
[0129] The DBSCAN algorithm is used by electronic devices to calculate the density value of blocks and classify density-connected blocks into a single candidate instance. For example, Figure 3C In order to address the above Figure 3B The results of clustering blocks in the data, such as Figure 3C As shown, three candidate instances 801, 802 and 803 are obtained.
[0130] Connected component analysis is a technique used by electronic devices to identify regions in an image that are adjacent to each other and have the same pixel value. These regions are called connected components, and independent connected components are then identified as at least one candidate instance. For example... Figure 4C As shown, Figure 4B The grid is divided into equal-sized cells, and then connected component analysis is used to group adjacent cells into the same category, resulting in the following: Figure 4D Candidate instances 1601 and 1602 are shown.
[0131] The K-means algorithm involves an electronic device first identifying several blocks as a cluster, initializing each cluster with a random center vector value, and then calculating the distance of each block from each center, assigning it to the cluster represented by the nearest cluster center. Next, the electronic device recalculates the center vector of each cluster based on the assignment of all blocks; this center vector is the mean of all blocks in that cluster. This process is repeated until convergence, resulting in at least one cluster, or at least one candidate instance.
[0132] S204. For the above M candidate instances, the electronic device extracts the features of each candidate instance, where candidate instance F i Let i be one of the M candidate instances, where i is a positive integer not greater than M.
[0133] The features can be geometric features, texture features, or a combination of geometric and texture features.
[0134] S205. The electronic device selects the target instance from the M candidate instances based on the characteristics of each candidate instance.
[0135] S206. Electronic devices identify traffic signs corresponding to target instances.
[0136] It should be noted that the candidate instance and target instance described in the embodiments of this application are images or data of the regions corresponding to the candidate instance and target instance in the image to be identified, respectively.
[0137] The specific implementation of steps S204-S206 above is described below through three implementation methods: Implementation method (1): In implementation method (1), the electronic device determines the target instance and the traffic sign corresponding to the target instance by the geometric features of the candidate instances. Specifically, the target instance can be determined from the candidate instances by two methods: geometric template matching (method A1) and a first classifier (method A2). Furthermore, the traffic sign corresponding to the target instance can be identified by two methods: B1 and B2.
[0138] The following describes methods A1 and A2 respectively: Method A1: Determine the target instance through geometric template matching. This process may include the following steps: A1a. The electronic device extracts the convex hull of the above M candidate instances, where candidate instance F i The convex hull is the candidate instance F i Geometric features.
[0139] Here, the convex hull is a convex polygon formed by connecting the outermost points of the candidate instances. For example, Figure 3D In order to address the above Figure 3CThe convex hull extraction of the three candidate instances shown yields the convex hulls 901, 902, and 903 for candidate instances 801, 802, and 803. Figure 4E In order to address the above Figure 4D The convex hulls of the two candidate instances shown are extracted to obtain the convex hulls 1701 and 1702 of candidate instances 1601 and 1602.
[0140] A1b, The electronic device determines the geometric template for each candidate instance.
[0141] The following is an example of candidate instance F. i The following is an example of how to determine the geometric template for each candidate instance F. i Two methods for creating geometric templates: Method 1: Target traffic signs Figure 7 Taking the traffic cone 101 shown as an example, according to the national standard description of the shape of a traffic cone, namely a cone, geometric features are extracted, which serve as the geometric template for the traffic cone and a reference object for matching. Since the image of a cone in an image is roughly triangular, an isosceles acute triangle can be used as the geometric template when identifying traffic cones in the image to be identified, such as... Figure 7 The isosceles acute triangle 103 is shown. Based on the principle of perspective (objects appear larger when closer and smaller when farther away), the size of the geometric template is determined by its position in the image. In this embodiment, a lookup table can be established, which can include multiple geometric templates corresponding to different positions. For example, the image can be divided into n equally spaced rows, with each row corresponding to a geometric template. For instance, the size of the geometric template corresponding to the j-th row is w. j ×h j , where w j h j These are the pixel width and pixel height of the geometric template corresponding to the j-th row, respectively. For example... Figure 7 The triangular warning sign 102 shown is designed with a geometric template based on the national standard description of a triangular warning sign, namely an equilateral triangle. This template serves as the reference object for matching the triangular warning sign. Accordingly, this embodiment uses an equilateral triangle as the geometric template, such as... Figure 7 The equilateral triangle 104 is shown.
[0142] In some embodiments, the electronic device will look up candidate instance F in the table. i The geometric template corresponding to the position is used as a candidate instance F i The lookup table contains multiple locations and their corresponding geometric templates. These multiple locations can be described by the row numbers in the aforementioned image, for example, candidate instance F. i The position can be a candidate instance F iThe number of rows X in the image to be identified i Candidate instance F i The geometric template is the row number X. i The corresponding geometric template.
[0143] For example, Figure 3E The above includes Figure 3C The three candidate instances 801, 802, and 803 correspond to geometric templates 1001, 1002, and 1003, which are based on the following: Figure 6 The traffic cone shown is designed with the geometric features of an isosceles acute triangle. For example, Figure 4F The above includes Figure 4D The geometric templates 1801 and 1802 correspond to 1601 and 1602.
[0144] Method 2: like Figure 8A and Figure 8B As shown, the real-time generation method of the geometric template is illustrated using a traffic cone as an example. First, the description of the geometric shape and size of the actual traffic cone 101 in the national standard is extracted, and a three-dimensional model is established ( Figure 8A The three-dimensional model 201 is described in a coordinate system 301 with the central axis of the three-dimensional model 201 as the z-axis and the base of the traffic cone as the xy-plane. Then, the image coordinates of a candidate instance are back-projected onto an ideal ground plane in the world coordinate system to obtain the world coordinates of the candidate instance, i.e., its position on the ideal ground plane. This back-projection process is a conversion of two-dimensional data to three-dimensional data, also known as a 2D-to-3D back-projection process, where the image coordinates are the coordinates on the plane of the image to be identified. Further, the three-dimensional model 201 is placed on the ideal ground plane at the position of the candidate instance. Finally, the three-dimensional model placed on the ideal ground plane is projected onto the plane of the image to be identified to obtain the geometric template corresponding to the candidate instance. This projection process is a conversion of three-dimensional data to two-dimensional data, also known as a 3D-to-2D projection process. Thus, geometric template matching can be further performed on the candidate instance. For example, as... Figure 8B As shown, Figure 8B In the image ①, the input image to be recognized first undergoes target color recognition and block clustering to obtain... Figure 8B The clustering result shown in ② is a number of candidate instances. Then, by back-projecting the image coordinates of each candidate instance in ②, the positions of the multiple candidate instances on the ideal ground plane are obtained as shown in ③. Further, a three-dimensional model is placed at the position of each candidate instance on the ideal ground plane, as shown in ④. Finally, the three-dimensional model placed on the ideal ground plane is projected onto the plane of the image to be identified, as shown in ⑤.
[0145] In a specific implementation, the image to be identified is captured by the target device. The electronic device can determine the image based on the target device's position in the world coordinate system and candidate instances F. i Based on information such as the first position and depth in the image to be identified, candidate instances F are generated. i The corresponding geometric template, the process may include: (1) the electronic device based on the position of the target device in the world coordinate system, the candidate instance F i Based on the first position and depth of field in the image to be identified, determine the candidate instance F. i In the second position in the world coordinate system, the target device is a camera that captures the image to be identified. The position of the target device can be determined based on the position of the electronic device and the positional relationship between the camera and the electronic device; (2) The electronic device determines the transformation relationship between the coordinate system of the image to be identified and the world coordinate system based on the first position and the second position; (3) The electronic device determines the candidate instance F based on the three-dimensional model of the traffic sign and the transformation relationship at the second position. i The geometric template is the projection of the 3D model onto the image to be recognized. The projection process of the 3D model involves rotation and translation calculations based on the aforementioned transformation relationships. The 3D model is designed according to the description of the geometric features of traffic signs in national standards.
[0146] It should be understood that, in one implementation, candidate instance F i The position can be a candidate instance F i The position of the midpoint of the bottom edge of the convex hull can be the coordinate of the midpoint in the coordinate system of the image to be recognized.
[0147] A1c, Electronic device determines candidate instance F i Geometric features and candidate instances F i The matching degree of the corresponding geometric template.
[0148] Optionally, candidate instance F i The matching degree of the corresponding geometric template is determined by one of the following: the ratio of the aspect ratio, area ratio, intersection-union area ratio, and similarity of Hughes' invariant moments of the two geometric features; or by multiple similarities of the above aspect ratio, area ratio, intersection-union area ratio, and Hughes' invariant moments. The following uses candidate instance F as an example. i Taking candidate instance F as an example, we will introduce it separately. i The specific implementation of the similarity of the corresponding geometric template in terms of aspect ratio, area ratio, intersection-merger area ratio, and Hu's invariant moments.
[0149] (1) Aspect ratio: Electronic devices determine candidate instance F i The aspect ratio of the convex hull; determining candidate instances F i The aspect ratio of the corresponding geometric template; further, determine the candidate instance F. iThe ratio of the aspect ratio of the convex hull to the aspect ratio of its corresponding geometric template is also called the first ratio.
[0150] The aspect ratio of the convex hull or geometric template can be the aspect ratio of the circumscribed rectangle of the convex hull or geometric template.
[0151] (2) Area ratio: Electronic devices determine candidate instance F i The area of the convex hull; determine candidate instances F. i The area of the corresponding geometric template; determine the candidate instance F. i The area of the convex hull and candidate instance F i The ratio of the areas of the corresponding geometric templates is also called the second ratio.
[0152] (3) Intersection and merger area ratio: The electronic device will select candidate instance F i Convex hull and candidate instance F i The corresponding geometric templates are aligned, and the candidate instance F is obtained after alignment. i The midpoint of the bottom edge of the convex hull and the candidate instance F i The corresponding geometric templates overlap. In a specific implementation, the electronic device selects candidate instance F. i The position of the corresponding geometric template in the coordinate system is translated, and the candidate instance F is calculated. i The corresponding geometric template is translated to the candidate instance F. i The position of the candidate instance F in the coordinate system; i The area of the convex hull and candidate instance F i The intersection area of the corresponding geometric templates; determining the candidate instance F. i The area of the convex hull and candidate instance F i The area of the union of the areas of the corresponding geometric templates; determine the ratio of the intersection area to the union area, which is also called the third ratio.
[0153] (4) Similarity of Hu's invariant moments: Extraction of candidate instances F by electronic devices i The first Hugh invariant moment of the geometric features; extract candidate instances F i The second Huv's invariant moment of the corresponding geometric template; determine the similarity between the first and second Huv's invariant moments.
[0154] It should be noted that one or more of the above implementation methods (1) to (4) can be executed. The execution order of the above implementation methods (1) to (4) can be arbitrary, and they can be executed simultaneously or sequentially. This application embodiment does not limit this.
[0155] A1d: Electronic devices identify candidate instances with a matching degree greater than the target threshold as target instances.
[0156] In some embodiments, the matching degree can be one of the following: the ratio of aspect ratio, area ratio, intersection-union area ratio, and similarity of Hu's invariant moments.
[0157] In some embodiments, the matching degree includes multiple similarities such as the aspect ratio, area ratio, intersection-union area ratio, and Hu's invariant moment. These aspect ratios, area ratios, intersection-union area ratios, and Hu's invariant moments can each correspond to a target threshold. A candidate instance is determined to be a target instance only when the aspect ratio, area ratio, intersection-union area ratio, and Hu's invariant moment of the candidate instance and its geometric template are all greater than their corresponding target thresholds.
[0158] For example, electronic devices perform calculations, such as Figure 3E The matching degree between the three candidate instances shown and their corresponding geometric templates is determined. Figure 3E The aspect ratio, area ratio, and intersection / merging area ratio of candidate instance 1004 and its corresponding geometric template 1001 are all greater than their corresponding target thresholds. Therefore, as Figure 3F As shown, candidate instance 1004 is determined as target instance 1101. For example, Figure 4F The aspect ratio, area ratio, and intersection / merging area ratio of candidate instance 1803 and its corresponding geometric template 1801 are all greater than their corresponding target thresholds. Therefore, as Figure 4G As shown, candidate instance 1803 is determined to be target instance 1901.
[0159] In some embodiments, candidate instance F i The matching degree of the corresponding geometric template can be a weighted sum of various factors, including the aspect ratio, area ratio, intersection-union area ratio, and similarity of Hu's invariant moments.
[0160] The method described in A1 above, which uses geometric template matching to determine target instances, requires no training and no data collection or ground truth labeling for the target traffic signs and their categories. Even if the target traffic signs of the same type have certain variations in appearance, as long as their color and overall shape have not changed significantly (and the deviation from the national standard is not significant), they can still be detected.
[0161] Method A2: Determine the target instance using the first classifier.
[0162] A first classifier can be trained. This classifier is trained using the geometric features of sample instances as input and the true results of the sample instances as labels. The true results are used to indicate the traffic signs corresponding to the input sample instances.
[0163] For each candidate instance, its geometric features can be extracted and input into the first classifier to predict the classification result of the candidate instance.
[0164] With candidate instance F i For example: Extracting candidate instances F from electronic devices i Geometric features; candidate instance F i The geometric features are input into the first classifier to obtain candidate instances F. i The classification result indicates the type of traffic sign corresponding to the input geometric features. If candidate instance F i If the classification result indicates that it is the target traffic sign or one of the target traffic signs, then the candidate instance F is determined. i For the target instance, if the candidate instance F i If the classification result is not a traffic sign, then determine whether the next candidate instance is the target instance.
[0165] After identifying the target instance, the traffic signs corresponding to that target instance can be identified using two methods, B1 and B2. Methods B1 and B2 are described below: Method B1: The electronic device directly uses the traffic sign described by the geometric template corresponding to the target instance in A1 or the classification result in A2 as the recognition result of the target instance.
[0166] Method B2: After identifying a target instance, the electronic device can further confirm the traffic sign corresponding to the target instance based on texture features to improve the accuracy of recognition. Only when the traffic sign of the target instance identified based on texture features is the same as the traffic sign of the target instance identified based on geometric features, is the recognition result of the target instance determined to be the traffic sign of the target instance identified based on geometric features, that is, the traffic sign described by the geometric template corresponding to the target instance in A1 or the classification result in A2.
[0167] Methods for identifying traffic signs corresponding to a target instance based on texture features can include two approaches: using a fourth classifier and texture template matching, which are described below: Through the second classifier: First, a second classifier can be trained. This second classifier is trained using the texture features of the sample instance as input and the true result of the sample instance as the label. The true result is used to indicate the traffic sign corresponding to the sample instance.
[0168] Further, the electronic device extracts the texture features of the target instance; inputs the texture features of the target instance into the second classifier to obtain the classification result of the target instance; when the classification result of the target instance is the same as the traffic sign corresponding to the target instance, the electronic device determines that the recognition result of the target instance is the traffic sign corresponding to the target instance. Conversely, when the classification result of the target instance is different from the traffic sign corresponding to the target instance, the recognition result of the target instance cannot be determined. The classification result is used to indicate the type of traffic sign corresponding to the input texture features. It should be noted that the traffic sign corresponding to the target instance is the traffic sign identified based on the geometric features of the target instance. When the electronic device executes A1 to obtain the target instance, the traffic sign corresponding to the target instance is the traffic sign described by the geometric template corresponding to the target instance in A1; when the electronic device executes A2 to obtain the target instance, the traffic sign corresponding to the target instance is the traffic sign indicated by the classification result in A2.
[0169] Optionally, the texture features can be Haar features, and the second classifier can be a Haar feature-based classifier. The electronic device may include a texture feature extraction module, which takes the Haar features of the target instance as input and then obtains the classification result of the target instance through a fourth classifier. For example, the above... Figure 3F The image of target instance 1101 in the image (i.e. Figure 3F The image within the rectangular bounding box is input into a second classifier, resulting in the classification result of target instance 1101 as a traffic cone, as shown below. Figure 3G The identification result of the traffic cone shown is 1201.
[0170] Optionally, the texture features can be histogram of oriented gradient (HOG) features, and the fourth classifier can be a support vector machine (SVM) classifier based on HOG features. The electronic device may include a texture feature extraction module, which takes the HOG features of the target instance as input and then uses the fourth classifier to obtain the classification result of the target instance. For example, the above... Figure 4G The image of target instance 1901 in the image (i.e. Figure 4G The image within the rectangular bounding box is input into an SVM classifier based on HOG features, resulting in the classification result of target instance 1901 as a triangular warning sign. Figure 4H The recognition result of the triangular warning sign shown is 2001.
[0171] It should be understood that texture features can also be other representations, such as local binary pattern (LBP) features, and other classifiers can also be selected. This application does not limit the specific classifiers used.
[0172] Texture template matching: The texture template is the texture feature of the target traffic sign. The texture feature of the target traffic sign can be extracted and used as the texture template. Similar to the geometric template, based on the principle of perspective (objects appear larger when closer and smaller when farther away), the size of the texture template is determined by its position in the image. The texture template corresponding to the target instance can be selected based on the target instance's position in the image.
[0173] Furthermore, the electronic device can calculate the matching degree between the texture features of the target instance and the texture template corresponding to the target instance. When the matching degree of the target instance is greater than a preset threshold, the matching result of the target instance is determined to be the traffic sign described by the texture template corresponding to the target instance. When the matching result is the same as the traffic sign corresponding to the target instance, the recognition result of the target instance is determined to be the traffic sign corresponding to the target instance. Conversely, when the classification result of the target instance and the traffic sign corresponding to the target instance are different, the recognition result of the target instance cannot be determined. It should be noted that the traffic sign corresponding to the target instance is the traffic sign identified based on the geometric features of the target instance. Specifically, when the electronic device executes A1 to obtain the target instance, the traffic sign corresponding to the target instance is the traffic sign described by the geometric template corresponding to the target instance in A1; when the electronic device executes A2 to obtain the target instance, the traffic sign corresponding to the target instance is the traffic sign indicated by the classification result in A2.
[0174] The method described in B2 above, which uses texture template matching to determine target instances, requires no training and no data collection or ground truth labeling for the target traffic signs and their categories, enabling rapid determination of target instances.
[0175] Implementation Method (II): In implementation method (ii), the electronic device determines the target instance and the traffic sign corresponding to the target instance by the texture features of the candidate instances. Specifically, the target instance can be determined from the candidate instances by two methods: texture template matching (method C1) and a second classifier (method C2). Furthermore, the traffic sign corresponding to the target instance can be identified by two methods: D1 and D2.
[0176] The following describes methods C1 and C2 respectively: Method C1 Texture features of target traffic signs can be extracted and used as texture templates. Similar to geometric templates, based on the principle of perspective (objects appear larger when closer and smaller when farther away), the size of the texture template is determined by its position in the image. Texture templates corresponding to candidate instances can be selected from those candidate instances based on their position in the image.
[0177] Furthermore, for each candidate instance, the matching degree between the texture features of the candidate instance and the texture template corresponding to the candidate instance is determined, also known as the second matching degree. Candidate instances with a second matching degree greater than a second target threshold are determined as target instances.
[0178] Method C2 Train a second classifier. This second classifier is trained using the texture features of the sample instances as input and the true results of the sample instances as labels. The true results are used to indicate the traffic signs corresponding to the sample instances.
[0179] Furthermore, for each candidate instance, its texture features can be extracted and input into the second classifier to predict the classification result of the candidate instance.
[0180] With candidate instance F i For example: Extracting candidate instances F from electronic devices i Texture features; candidate instance F i The texture features are input into the second classifier to obtain candidate instances F. i The classification result, whereby the classification result is used to indicate the type of traffic sign corresponding to the input texture feature; if candidate instance F i If the classification result indicates that the candidate instance is a target traffic sign or one of the target traffic signs, then the candidate instance is determined to be a target instance. i If the classification result does not indicate the target traffic sign, then determine whether the next candidate instance is the target instance.
[0181] For information on texture features and methods for extracting texture features, please refer to the relevant description in Method B2 above, which will not be repeated here.
[0182] Similar to implementation method (i) above, after determining the target instance, the traffic signs corresponding to the target instance can be identified using two methods, D1 and D2. Methods D1 and D2 are described below: D1. The electronic device directly uses the traffic signs described by the texture template corresponding to the target instance in C1 or the classification result in C2 as the recognition result of the target instance.
[0183] D2. After identifying a target instance, the electronic device can further confirm the traffic sign corresponding to that target instance based on geometric features to improve the accuracy of recognition. It should be noted that the traffic sign corresponding to the target instance is the sign identified based on the texture features of the target instance. Only when the traffic sign of the target instance identified based on geometric features is the same as the traffic sign of the target instance identified based on texture features in method C1 or method C2, is the recognition result of the target instance determined to be the traffic sign of the target instance identified based on texture features, i.e., the traffic sign described by the texture template corresponding to the target instance in C1 or the classification result in C2.
[0184] The method for identifying the traffic sign corresponding to the target instance based on geometric features can include two methods: a first classifier and geometric template matching. For the first classifier, please refer to the relevant description in A2 of the above-mentioned implementation method (I), and for geometric template matching, please refer to the relevant description in A1 of the above-mentioned implementation method (I), which will not be repeated here.
[0185] Implementation method (3): In implementation method (iii), the electronic device determines the target instance and the traffic sign corresponding to the target instance by using the geometric and texture features of the candidate instance.
[0186] A third classifier is trained, which is trained with the geometric and / or texture features of the sample instance as input and the true result of the sample instance as the label. The true result is used to indicate the traffic sign corresponding to the sample instance.
[0187] Furthermore, for each candidate instance, it is input into a third classifier to obtain the classification result of each candidate instance. This classification result is used to indicate the type of traffic sign corresponding to the input geometric and texture features. Further, the electronic device identifies the candidate instances among the M candidate instances whose classification result is either the target traffic sign or one of the target traffic signs as the target instance. Then, the identification result of the target instance is determined to be the traffic sign indicated by the classification result.
[0188] like Figure 9 As shown, Figure 9 This is a schematic diagram illustrating a method for determining a target instance and its classification result using the geometric and texture features of candidate instances. Candidate instance F... i For example, the process may include: (1) Extracting candidate instances F from electronic devices i The geometric and texture features. For the method of extracting geometric features, please refer to the relevant description in A2 of Implementation (I), and for the implementation of extracting texture features, please refer to the relevant description in B2 of Implementation (I).
[0189] (2) The electronic device will select candidate instance F i The geometric and texture features are input into a third classifier to obtain candidate instances F. i The classification results (3) If candidate instance F i If the classification result is a target traffic sign or one of the target traffic signs, then the candidate instance is determined as the target instance, and (4) is executed; if the candidate instance F i If the classification result is not a traffic sign, then the next candidate instance can be identified.
[0190] (4) Determine the target instance F i The identification result is the traffic sign indicated by the classification result.
[0191] The aforementioned traffic sign recognition method identifies candidate instances by recognizing the colors in the image to be recognized based on the color features of the traffic signs. Then, it filters or classifies the candidate instances based on the geometric and / or texture features of the traffic signs. This method can quickly identify traffic signs in the image to be recognized, thereby providing timely reminders or planning and controlling the vehicle's driving direction, which can improve the safety of the vehicle during driving.
[0192] Furthermore, when traffic signs of the same type exhibit certain variations in appearance, recognition can still be achieved as long as their color, overall shape, and texture remain largely unchanged (with minimal deviation from national standards). Moreover, effective detection of traffic signs can still be achieved even when the vehicle is stationary and other stationary objects in the scene cannot generate optical flow, or when the vehicle is moving and the traffic sign is directly in front of the vehicle and its motion is stationary or parallel to the vehicle's direction of movement. Because this method does not rely on the motion of the traffic sign relative to the vehicle but utilizes the inherent properties and characteristics of the traffic sign, it exhibits good stability.
[0193] The electronic devices involved in the embodiments of this application are described below.
[0194] Please see Figure 10 , Figure 10 This is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of this application. The electronic device 100 can be... Figure 1A The mobile device or server in it can also be Figure 1B Vehicles, or computer identification systems or computing devices within vehicles, etc. For example... Figure 10As shown, the electronic device includes a processor 110, internal memory 120, external memory interface 130, camera 140, display screen 150, communication module 160, audio module 170, etc. The internal memory 120 or the external memory connected via the external memory interface 130 can store computer program code, which includes computer instructions. The processor 110 can call these computer instructions to enable the electronic device 100 to perform the aforementioned functions. Figure 2A The method for identifying any of the traffic signs shown.
[0195] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware. Detailed descriptions of each part are as follows.
[0196] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0197] The controller can serve as the nerve center and command center of an electronic device. Based on the instruction opcode and timing signals, the controller generates operation control signals to control the fetching and execution of instructions.
[0198] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0199] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0200] Internal memory 120 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 120. Internal memory 120 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phone book, etc.). In addition, internal memory 120 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0201] The external storage interface 130 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 130 to perform data storage functions. For example, recorded audio and video files can be saved on the external memory card.
[0202] Camera 140 is used to capture still images or videos of the surrounding environment of a vehicle or other mobile platform along its travel path. Camera 140 may include one or more of a monocular camera 141, a binocular camera 142, a multi-view camera, or a surround-view camera 143. An object is projected onto a photosensitive element in camera 140 through a lens to generate an optical image. The photosensitive element may be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, electronic device 100 may include one or N cameras, where N is a positive integer greater than 1.
[0203] Display screen 150 is used to display images, videos, etc., such as images of the surrounding environment of a vehicle or other mobile platform in its travel path. Display screen 150 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 150, where N is a positive integer greater than 1.
[0204] The communication module 160 may include a wireless communication module and a wired communication module. The wireless communication module can implement 2G / 3G / 4G / 5G mobile communication, wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), and other short-range wireless communication technologies. The wired communication module may include tangible media such as metal wires or optical fibers, enabling the electronic device 100 to transmit information to other devices via this tangible medium.
[0205] A communication bus can include a pathway for transmitting information between the aforementioned components; this bus can be a PCI bus or an EISA bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one thick line is used to represent them in the diagram, but this does not imply that there is only one bus or one type of bus.
[0206] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0207] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can play audio files through the speaker 170A.
[0208] Microphone 170B, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. Electronic device 100 can use microphone 170B to collect sound from the environment in which the vehicle or other mobile platform is located. Electronic device 100 may have at least one microphone 170B. In some embodiments, electronic device 100 may have two microphones 170B, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170B, enabling sound signal collection, noise reduction, sound source identification, and directional recording, among other functions.
[0209] The following describes a traffic sign recognition device according to an embodiment of this application.
[0210] Please see Figure 11 , Figure 11 This is a schematic diagram of the structure of a traffic sign recognition device 1100 provided in an embodiment of this application. The recognition device can be... Figure 1A The mobile device or server in it can also be Figure 1B The vehicle in question, or the computer identification system or computing device within the vehicle, etc. The identification device includes some or all of the following units: The first recognition unit 1101 is used to recognize pixels of a target color from an image to be recognized, and obtain at least one block, wherein the block is composed of multiple connected pixels, the multiple pixels have the same color, and the target color is one or more colors of a target traffic sign. Clustering unit 1102 is configured to cluster the at least one block according to the location of the at least one block to obtain at least one candidate instance, wherein the candidate instance is a combination of one or more blocks among the at least one block; Extraction unit 1103 is used to extract features of each candidate instance for each of the at least one candidate instance; The filtering unit 1104 is used to filter out a target instance from the at least one candidate instance based on the characteristics of each of the at least one candidate instance; The second identification unit 1105 is used to identify the traffic signs corresponding to the target instance.
[0211] In one possible implementation, the features include geometric features. The filtering unit 1104 is specifically used to: determine a first matching degree between the geometric features of each candidate instance in the at least one candidate instance and the geometric template corresponding to each candidate instance, wherein the geometric template is a geometric feature used to describe the target traffic sign; and determine the candidate instances in the at least one candidate instance whose first matching degree is greater than a first target threshold as target instances. The filtering unit 1104 is specifically used to: determine the recognition result of the target instance based on the traffic signs described by the geometric template corresponding to the target instance.
[0212] In one possible implementation, the features include geometric features. The extraction unit 1103 is specifically used to: extract the geometric features of each candidate instance; The filtering unit 1104 is specifically used for: inputting the geometric features of each candidate instance into a first classifier to obtain a first classification result for each candidate instance, wherein the first classification result is used to indicate the type of traffic sign corresponding to the input geometric features, the first classifier is trained with a first sample instance as input and the true result of the first sample instance as a label, wherein the true result is used to indicate the traffic sign corresponding to the input geometric features; and determining the candidate instance whose first classification result is a traffic sign among the at least one candidate instance as the target instance; The second identification unit 1105 is specifically used to: determine the identification result of the target instance based on the first classification result of the target instance.
[0213] In one possible implementation, the features include texture features. The filtering unit 1104 is specifically used for: determining a texture template corresponding to each candidate instance in the at least one candidate instance; the texture template is a texture feature used to describe the target traffic sign; determining a second matching degree between the texture feature of each candidate instance and the texture template corresponding to each candidate instance; and determining the candidate instances in the at least one candidate instance whose second matching degree is greater than a second target threshold as target instances; The second identification unit 1105 is specifically used to: determine the identification result of the target instance based on the traffic sign described by the texture template corresponding to the target instance.
[0214] In one possible implementation, the extraction unit 1103 is specifically used to: extract the texture features of each candidate instance; The filtering unit 1104 is specifically used for: inputting the texture features of each candidate instance into a second classifier to obtain a second classification result for each candidate instance, wherein the second classification result is used to indicate the type of traffic sign corresponding to the input texture features, the second classifier is trained using the texture features of a second sample instance as input and the true result of the second sample instance as the label, wherein the true result is used to indicate the traffic sign corresponding to the second sample instance; and determining the candidate instance whose second classification result is a traffic sign among the at least one candidate instance as the target instance; The second identification unit 1105 is specifically used to: determine the identification result of the target instance based on the second classification result of the target instance.
[0215] In one possible implementation, the extraction unit 1103 is specifically used to: extract the geometric features and texture features of each candidate instance; The filtering unit 1104 is specifically used to: input the geometric features and texture features of each candidate instance into a third classifier to obtain a third classification result for each candidate instance, wherein the third classification result is used to indicate the type of traffic sign corresponding to the input geometric features and texture features, wherein the third classifier is trained with the geometric features and / or texture features of a third sample instance as input, and the true result of the third sample instance is used to indicate the traffic sign corresponding to the third sample instance; and to determine the candidate instance whose third classification result is a traffic sign among the at least one candidate instance as the target instance; The second identification unit 1105 is specifically used to: determine the identification result of the target instance as a traffic sign indicated by the third classification result of the target instance.
[0216] In one possible implementation, the second identification unit 1105 is specifically used to: determine the identification result of the target instance as the traffic sign described by the geometric template corresponding to the target instance.
[0217] In one possible implementation, the extraction unit 1103 is specifically used to: extract the texture features of the target instance; The second identification unit 1105 is specifically used for: inputting the texture features of the target instance into a fourth classifier to obtain a fourth classification result of the target instance, wherein the fourth classification result is used to indicate the type of traffic sign corresponding to the input texture features, wherein the fourth classifier is obtained by training the true result of the fourth sample instance with labels, and the true result is used to indicate the traffic sign corresponding to the fourth sample instance; and, when the traffic sign indicated by the fourth classification result is the same as the traffic sign described by the geometric template corresponding to the target instance, determining that the identification result of the target instance is the traffic sign described by the geometric template corresponding to the target instance.
[0218] In one possible implementation, the filtering unit 1104 is specifically configured to: determine the aspect ratio of the convex hull of each candidate instance; determine the aspect ratio of the geometric template corresponding to each candidate instance; and determine a first ratio of the aspect ratio of the convex hull of each candidate instance to the aspect ratio of the geometric template corresponding to each candidate instance as the first matching degree.
[0219] In one possible implementation, the filtering unit 1104 is specifically configured to: determine the area of the convex hull of each instance; determine the area of the geometric template corresponding to each candidate instance; and determine a second ratio of the area of the convex hull of each candidate instance to the area of the geometric template corresponding to each candidate instance as the first matching degree.
[0220] In one possible implementation, the filtering unit 1104 is specifically configured to: align the convex hull of each instance with the geometric template corresponding to each candidate instance, such that the midpoint of the bottom edge of the convex hull of each instance coincides with the geometric template corresponding to each candidate instance; determine the intersection area of the area of the convex hull of each instance and the area of the geometric template corresponding to each candidate instance; determine the union area of the area of the convex hull of each instance and the area of the geometric template corresponding to each candidate instance; and determine a third ratio of the intersection area to the union area as the first matching degree.
[0221] In one possible implementation, the filtering unit 1104 is specifically used to: extract the first Huvren invariant moment of the geometric features of each candidate instance; extract the second Huvren invariant moment of the corresponding geometric template of each candidate instance; and determine the similarity between the first Huvren invariant moment and the second Huvren invariant moment as the first matching degree.
[0222] In one possible implementation, the features include geometric features, and the extraction unit 1103 is specifically used to: extract the convex hull of each candidate instance, wherein the convex hull of each candidate instance is the geometric feature of each candidate instance.
[0223] In one possible implementation, the filtering unit 1104 is further specifically used to: use the geometric template corresponding to the position of each candidate instance in the lookup table as the geometric template corresponding to each candidate instance, wherein the lookup table includes multiple positions and geometric templates corresponding to the multiple positions respectively.
[0224] In one possible implementation, the image to be identified is captured by a target device, and the filtering unit 1104 is further configured to: determine a second position of each candidate instance in the world coordinate system based on the position of the target device in the world coordinate system, the first position of each candidate instance in the image to be identified, and the depth of field; determine the transformation relationship between the coordinate system of the image to be identified and the world coordinate system based on the first position and the second position; and determine the geometric template corresponding to each candidate instance as the projection of the three-dimensional model in the image to be identified based on the three-dimensional model of the traffic sign at the second position and the transformation relationship.
[0225] In one possible implementation, the clustering unit 1102 is specifically used to: determine the distance between every two blocks in the at least one block according to the location of the at least one block; and divide two blocks with a distance less than a first threshold into a candidate instance to obtain the at least one candidate instance.
[0226] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0227] It is understood that those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in the various embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0228] Those skilled in the art will appreciate that the functionality described by the various illustrative logic blocks, modules, and algorithmic steps disclosed in connection with the various embodiments of this application can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality described by the various illustrative logic blocks, modules, and steps can be stored or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may comprise a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium, or a communication medium that includes any medium that facilitates the transfer of a computer program from one place to another (e.g., according to a communication protocol). In this way, the computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium, such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this application. The computer program product may comprise a computer-readable medium.
[0229] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0230] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0231] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0232] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0233] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0234] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for identifying traffic signs, characterized in that, The method is applied to an electronic device, and the method includes: Identify pixels of a target color from the image to be identified to obtain at least one block, the block being composed of multiple connected pixels of the same color, the target color being one or more colors of a target traffic sign; the target traffic sign being a type of traffic sign. Based on the location of the at least one block, the at least one block is clustered to obtain at least one candidate instance, wherein the candidate instance is a combination of one or more blocks among the at least one block; For each candidate instance among the at least one candidate instance, features of each candidate instance are extracted; the features include geometric features; the geometric features are convex hulls. A first matching degree is determined between the geometric features of each candidate instance in the at least one candidate instance and the geometric template corresponding to each candidate instance, wherein the geometric template is a geometric feature used to describe the target traffic sign; the first matching degree is determined based on at least one of the similarity between the candidate instance and the corresponding geometric template in terms of aspect ratio, area ratio, intersection-union area ratio and Hu's invariant moment. Among the at least one candidate instance, the candidate instance whose first matching degree is greater than the first target threshold is determined as the target instance; the two candidate instances are in different positions and the geometric templates corresponding to the two candidate instances are different. Extract the texture features of the target instance; Traffic signs are identified based on the texture features of the target instance; If the traffic sign identified based on the texture features of the target instance is the same as the traffic sign described by the geometric template corresponding to the target instance, then the identification result of the target instance is determined to be the traffic sign described by the geometric template corresponding to the target instance.
2. The method according to claim 1, characterized in that, The method of identifying traffic signs based on the texture features of the target instance includes: Determine the texture template of the target instance; Determine the third matching degree between the texture features of the target instance and the texture template corresponding to the target instance, wherein the texture template is the texture feature of the traffic sign; A first matching result is determined for the target instance whose third matching degree is greater than the third target threshold. The first matching result is the traffic sign described by the texture template corresponding to the target instance. The first matching result is the traffic sign identified based on the texture features of the target instance.
3. The method according to claim 1, characterized in that, The method of identifying traffic signs based on the texture features of the target instance includes: The texture features of the target instance are input into a fourth classifier to obtain a fourth classification result for the target instance. The fourth classification result is used to indicate the type of traffic sign corresponding to the input texture features. The fourth classifier takes the texture features of a fourth sample instance as input, and the true result of the fourth sample instance is obtained by label training. The true result is used to indicate the traffic sign corresponding to the fourth sample instance. The traffic sign indicated by the fourth classification result is a traffic sign identified based on the texture features of the target instance.
4. The method according to any one of claims 1-3, characterized in that, Determining the first matching degree between the geometric features of each candidate instance and the geometric template corresponding to each candidate instance in the at least one candidate instance includes: Determine the aspect ratio of the convex hull of each candidate instance; Determine the aspect ratio of the geometric template corresponding to each candidate instance; The first matching degree is determined by the ratio of the aspect ratio of the convex hull of each candidate instance to the aspect ratio of the geometric template corresponding to each candidate instance.
5. The method according to any one of claims 1-3, characterized in that, Determining the first matching degree between the geometric features of each candidate instance and the geometric template corresponding to each candidate instance in the at least one candidate instance includes: Determine the area of the convex hull for each instance; Determine the area of the geometric template corresponding to each candidate instance; The first matching degree is determined as a second ratio of the area of the convex hull of each candidate instance to the area of the geometric template corresponding to each candidate instance.
6. The method according to any one of claims 1-3, characterized in that, Determining the first matching degree between the geometric features of each candidate instance and the geometric template corresponding to each candidate instance in the at least one candidate instance includes: Align the convex hull of each instance with the geometric template corresponding to each candidate instance. After alignment, the midpoint of the bottom edge of the convex hull of each instance coincides with the geometric template corresponding to each candidate instance. Determine the intersection area of the area of the convex hull of each instance and the area of the geometric template corresponding to each candidate instance; Determine the area of the union of the area of the convex hull of each instance and the area of the geometric template corresponding to each candidate instance; The third ratio of the intersection area to the union area is determined as the first matching degree.
7. The method according to any one of claims 1-3, characterized in that, Determining the first matching degree between the geometric features of each candidate instance and the geometric template corresponding to each candidate instance in the at least one candidate instance includes: Extract the first Hugh invariant moments of the geometric features of each candidate instance; Extract the second Hugh invariant moment of the corresponding geometric template for each candidate instance; The similarity between the first Huvgeny invariant moment and the second Huvgeny invariant moment is determined as the first matching degree.
8. The method according to any one of claims 1-3, characterized in that, The features include geometric features, and the extraction of features for each candidate instance includes: Extract the convex hull of each candidate instance, where the convex hull of each candidate instance is the geometric feature of each candidate instance.
9. The method according to any one of claims 1-3, characterized in that, The method further includes: The geometric template corresponding to the position of each candidate instance in the lookup table is used as the geometric template corresponding to each candidate instance. The lookup table includes multiple positions and geometric templates corresponding to the multiple positions respectively.
10. The method according to any one of claims 1-3, characterized in that, The image to be identified is captured by the target device, and the method further includes: Based on the position of the target device in the world coordinate system, the first position of each candidate instance in the image to be identified, and the depth of field, determine the second position of each candidate instance in the world coordinate system; Based on the first position and the second position, determine the transformation relationship between the coordinate system of the image to be identified and the world coordinate system; Based on the 3D model of the traffic sign at the second position and the transformation relationship, the geometric template corresponding to each candidate instance is determined as the projection of the 3D model onto the image to be identified.
11. The method according to any one of claims 1-3, characterized in that, The step of clustering the at least one block based on its location to obtain at least one candidate instance includes: Based on the location of the at least one block, determine the distance between every two blocks in the at least one block; Two blocks with a distance less than a first threshold are grouped into one candidate instance to obtain at least one candidate instance.
12. An electronic device, characterized in that, include: One or more processors and one or more memories, wherein the one or more memories are respectively coupled to the one or more processors; the one or more memories are used to store computer program code, the computer program code including computer instructions; The processor is used to invoke the computer instructions to execute: the traffic sign recognition method as described in any one of claims 1-11.
13. A vehicle, characterized in that, include: One or more processors and one or more memories, wherein the one or more memories are respectively coupled to the one or more processors; the one or more memories are used to store computer program code, the computer program code including computer instructions; The processor is used to invoke the computer instructions to execute: the traffic sign recognition method as described in any one of claims 1-11.
14. A computer storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the traffic sign recognition method as described in any one of claims 1-11.
Citation Information
Patent Citations
Traffic sign detection method and equipment
CN103020623A
Method and system for navicular identification
CN103514448A
Temporary license plate detection method and device
CN107194393A
Fingerprint image recognition method based on detail points and texture features
CN107748877A