Traffic sign detection method and training method of traffic sign detection model

By using a traffic sign detection model that combines detection points and detection lines to represent the location of traffic signs, the problem of poor detection performance of distant signs by autonomous vehicles has been solved, achieving efficient and accurate detection of traffic signs and supporting the safe driving of autonomous vehicles.

CN116071284BActive Publication Date: 2026-04-07BEIJING TUSEN ZHITU TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, autonomous vehicles have difficulty effectively detecting traffic signs on the road, especially when dealing with distant traffic signs. The detection performance is limited, and the location of multiple traffic signs cannot be accurately represented.

Method used

A traffic sign detection model is adopted, which represents the position of traffic signs by combining detection points and detection lines. A convolutional neural network is used to generate response maps and train the model to improve detection accuracy.

Benefits of technology

It enables accurate detection of traffic signs, regardless of their distance from the vehicle, ensuring the accuracy and clarity of the detection results and supporting the safe operation of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116071284B_ABST
    Figure CN116071284B_ABST
Patent Text Reader

Abstract

This disclosure relates to a traffic sign detection method and a training method for a traffic sign detection model, relating to the field of intelligent transportation technology, and particularly to autonomous driving technology. The traffic sign detection method includes: acquiring a target image containing traffic signs; and inputting the target image into a traffic sign detection model to obtain detection markers corresponding to the traffic signs; wherein the detection markers include at least one of detection points and detection lines, used to characterize the position of the traffic signs in the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of intelligent transportation, and particularly to autonomous driving technology, specifically to a training method, apparatus, electronic device, computer-readable storage medium, and computer program product for traffic sign detection methods and traffic sign detection models. Background Technology

[0002] Traffic signs, such as traffic cones, are commonly found on roads as road separation and warning facilities. Autonomous vehicles need to have the ability to detect traffic signs in real time.

[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0004] According to one aspect of this disclosure, a traffic sign detection method is provided, comprising: acquiring a target image containing traffic signs; and inputting the target image into a traffic sign detection model to obtain detection markers corresponding to the traffic signs, wherein the detection markers include at least one of detection points and detection lines, used to characterize the position of the traffic signs in the target image.

[0005] According to another aspect of this disclosure, a method for training a traffic sign detection model is provided, comprising: acquiring a training image and a labeled image corresponding to the training image, the training image including traffic signs, the labeled image including labels corresponding to the traffic signs in the training image, the labels being used to mark the positions of the traffic signs in the corresponding training image, and the labels including at least one of labeled points and labeled lines; inputting the training image into a traffic sign detection model to obtain predicted labels of the traffic signs in the training image; and training the traffic sign detection model based on the predicted labels of the training image and the labels of the corresponding labeled image.

[0006] According to another aspect of this disclosure, a traffic sign detection apparatus is provided, comprising: an image acquisition unit configured to acquire a target image containing traffic signs; and a detection unit configured to input the target image into a traffic sign detection model to obtain detection marks corresponding to the traffic signs, wherein the detection marks include detection points for characterizing the position of the traffic signs in the target image, and the detection marks include at least one of detection points and detection lines.

[0007] According to another aspect of this disclosure, an apparatus for training a traffic sign detection model is provided, comprising: a first acquisition unit configured to acquire a training image and a labeled image corresponding to the training image, the training image including traffic signs, the labeled image including labels corresponding to the traffic signs in the training image, the labels being used to mark the positions of the traffic signs in the corresponding training image, and the labels including at least one of labeled points and labeled lines; a model prediction unit configured to input the training image into a traffic sign detection model to obtain predicted labels corresponding to the traffic signs in the training image; and a model training unit configured to train the traffic sign detection model based on the predicted labels of the training image and the labels of the corresponding labeled image.

[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the methods described in this disclosure.

[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods described in this disclosure.

[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described in this disclosure.

[0011] According to another aspect of this disclosure, a vehicle is provided, including the electronic devices as described above.

[0012] According to one or more embodiments of this disclosure, by inputting a target image containing traffic signs into a traffic sign detection model, detection points or detection lines corresponding to the traffic signs are obtained, so that traffic signs on the road can be detected regardless of their form or distance from the camera device on the vehicle used to acquire the target image.

[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0014] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0015] Figure 1 A schematic diagram illustrating an application scenario according to an exemplary embodiment;

[0016] Figure 2 A flowchart of a traffic sign detection method according to an embodiment of the present disclosure is shown;

[0017] Figure 3 A schematic diagram showing an image including traffic signs according to an embodiment of the present disclosure is illustrated;

[0018] Figure 4 A flowchart is shown illustrating the process of inputting a target image into a traffic sign detection model to obtain the detection mark corresponding to the traffic sign in a traffic sign detection method according to an embodiment of the present disclosure;

[0019] Figure 5 A flowchart illustrating a training method for a traffic sign detection model according to an embodiment of the present disclosure is shown;

[0020] Figure 6 A schematic diagram showing labeled images according to embodiments of the present disclosure is provided.

[0021] Figure 7 A flowchart is shown illustrating the process of training a traffic sign detection model based on predicted markers of training images and labeled markers of labeled images in a training method for a traffic sign detection model according to an embodiment of the present disclosure.

[0022] Figure 8 An exemplary block diagram of a traffic sign detection device according to an embodiment of the present disclosure is shown;

[0023] Figure 9 An exemplary block diagram of a training apparatus for a traffic sign detection model according to an embodiment of the present disclosure is shown; and

[0024] Figure 10 This is a structural block diagram illustrating an exemplary computing device that can be applied to exemplary embodiments. Detailed Implementation

[0025] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0026] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0027] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0028] The embodiments of this disclosure provide a novel method for traffic sign detection and a method for training a traffic sign detection model. The principles of this disclosure will be described below with reference to the accompanying drawings.

[0029] Figure 1 A schematic diagram illustrating an application scenario 100 according to an exemplary embodiment is shown. For example... Figure 1 As shown, this application scenario 100 may include vehicle 110 (e.g., Figure 1 The system includes a vehicle (shown in the diagram), a network 120, a server 130, and a database 140. The vehicle 110 may be equipped with an electronic system for autonomous driving.

[0030] Vehicle 110 may be coupled with a camera device to acquire image data of the road, wherein the image data includes target images containing traffic signs. Vehicle 110 may be an autonomous vehicle, such as an autonomous truck.

[0031] The vehicle 110 may also include a communication unit, and may send the target image to the server 130 via the network 120, and the server 130 may execute the traffic sign detection method provided in this application.

[0032] In some implementations, server 130 can execute the traffic sign detection method provided in this application using a built-in application. In other implementations, server 130 can execute the traffic sign detection method provided in this application by calling an application stored externally on the server.

[0033] Network 120 can be a single network or a combination of at least two different networks. For example, network 120 can be one or more of the following: local area network, wide area network, public network, private network, etc.

[0034] Server 130 can be built into vehicle 110 as an on-board server; alternatively, it can be an external device independent of vehicle 110, used to remotely provide services to vehicle 110. Server 110 can be a single server or a server cluster, with servers within the cluster connected via wired or wireless networks. A server cluster can be centralized, such as in a data center, or distributed. Server 130 can be local or remote.

[0035] Database 140 can refer to any device with storage capabilities. Database 140 is primarily used to store various data used, generated, and output during the operation of vehicle 110 and server 130. Database 140 can be local or remote. Database 140 may include various types of memory, such as Random Access Memory (RAM) and Read Only Memory (ROM). The storage devices mentioned above are just a few examples; the storage devices that can be used in this system are not limited to these.

[0036] Database 140 can be interconnected or communicate with server 130 or a part thereof via network 120, or directly interconnected or communicate with server 130, or a combination of the above two methods.

[0037] In some embodiments, database 140 may be a standalone device. In other embodiments, database 140 may be integrated into at least one of vehicle 110 and server 130. For example, database 140 may be located on vehicle 110 or on server 130. Alternatively, database 140 may be distributed, with one portion located on vehicle 110 and another portion on server 130.

[0038] Figure 2 A flowchart of a traffic sign detection method according to an embodiment of the present disclosure is shown. It can be utilized... Figure 1 The server 130 shown in the figure is performing Figure 2 Method 200 is shown in the figure.

[0039] like Figure 2 As shown, the traffic sign detection method 200 according to some embodiments of this disclosure includes:

[0040] Step S210: Acquire a target image containing traffic signs; and

[0041] Step S220: Input the target image into the traffic sign detection model to obtain the detection mark corresponding to the traffic sign, wherein the detection mark includes at least one of detection point and detection line, used to characterize the position of the traffic sign in the target image.

[0042] According to one or more embodiments of this disclosure, by inputting a target image containing traffic signs into a traffic sign detection model, detection points or detection lines corresponding to the traffic signs are obtained, so that traffic signs on the road can be detected regardless of their form or distance from the camera device on the vehicle used to acquire the target image.

[0043] In related technologies, during the detection of traffic signs on roads, a rectangular frame surrounding the traffic sign is used to represent its position in the image. When an image contains multiple traffic signs on the road, it is often impossible to clearly represent each one, thus limiting the detection of traffic signs.

[0044] Figure 3 A schematic diagram showing an image containing traffic signs according to some embodiments is provided. Figure 3 Images of multiple traffic cones, such as traffic cone 301, placed along the road's extension direction and captured by an onboard camera device of a vehicle traveling on the road, often appear blurry or obscured due to the near-object-far-object imaging characteristic of the camera device. This is because multiple traffic cones 301 located further away from the vehicle in the image frequently obscure each other or appear blurry. Figure 3(As shown in box A). In this case, a rectangular frame surrounding the traffic cone 301 is used to represent the position of the traffic cone 301 in the image. This causes multiple detection boxes corresponding to multiple traffic cones 301 that are far from the vehicle to overlap and become indistinguishable or unusable for representation, thus limiting the detection of distant traffic cones 301. In embodiments according to this disclosure, the traffic cone 301 can be represented by detection points and / or detection lines. Specifically, all of them can be represented as detection points, or the detection points can be fitted to detection lines, or a combination of detection points and detection lines can be used. For example, traffic cones 301 that are close to the measurement can be represented by detection points, while multiple traffic cones 301 that are far from the vehicle can be represented by detection lines, thereby enabling the representation of all traffic cones in the image and, consequently, the detection of all traffic cones in the image.

[0045] It should be understood that the embodiment includes multiple traffic cones placed along the road's extension direction. Figure 3 The method disclosed herein is merely an example to illustrate how the representation of multiple traffic cones located at a distance from a camera device on a vehicle can be achieved. Those skilled in the art should understand that the method disclosed herein can represent traffic signs of any placement (e.g., placed across the road, located on both sides of the road, etc.) and thus enable the detection of traffic signs.

[0046] At the same time, it should be understood that the embodiments are for the purpose of... Figure 3 The use of detection points to represent individual traffic cones 301 closer to the vehicle and detection lines to represent multiple traffic cones 301 farther from the vehicle is merely exemplary. Those skilled in the art should understand that the method according to this disclosure can also be used for... Figure 3 The multiple traffic cones 301 in the system are characterized by detection points or by the same detection line.

[0047] For example, in some embodiments, Figure 3 Each traffic cone 301 is represented by a detection point. Visually, the detection points of multiple traffic cones 301 densely packed in the distance can be approximated as a line. In other embodiments, the detection points of multiple traffic cones 301 densely packed in the distance can be connected into a detection line. In other embodiments, a fitted line is generated by clustering and fitting the detection points of these traffic cones 301, thereby representing these traffic cones 301 in the form of a fitted line.

[0048] In some embodiments, traffic signs include various types of traffic sign elements, such as traffic cones, water-filled barriers, utility poles, lane lines, or trees along the roadside, etc., and are not limited thereto.

[0049] In some embodiments, the target image includes multiple traffic sign elements of the same type, such as multiple traffic cones placed along the road extension direction, and multiple water-filled barriers placed across the road.

[0050] In some embodiments, the traffic sign includes a plurality of first traffic sign elements, wherein the pixel distance between any two adjacent first traffic sign elements is greater than or equal to a preset threshold, and the detection marker includes a detection point corresponding to each first traffic sign element. The pixel distance here can be the pixel distance between the center points of the first traffic sign elements, or the pixel distance between vertices at the same location, or the distance between the adjacent left and right boundaries of two adjacent first traffic sign elements.

[0051] For multiple first traffic sign elements, since the pixel distance between adjacent first traffic sign elements is greater than or equal to a preset threshold, each of the multiple first traffic sign elements can be clearly distinguished in the target image. Each first traffic sign element is marked with a detection point, ensuring that the position of the first traffic sign element represented by the detection point in the target image is accurate, clear, and distinguishable, which helps vehicles adjust their driving path based on the first traffic sign elements marked by the detection points.

[0052] In some embodiments, the traffic sign includes a plurality of second traffic sign elements, wherein the pixel distance between any two adjacent second traffic sign elements is less than a preset threshold, and the marking includes detection lines corresponding to the plurality of second traffic sign elements. Similarly, the pixel distance here may be the pixel distance between the center points of the first traffic sign elements, or the pixel distance between vertices at the same location, or the distance between the adjacent left and right boundaries of two adjacent first traffic sign elements.

[0053] For multiple second traffic sign elements, since the pixel distance between adjacent second traffic sign elements is less than a preset threshold, they may overlap in the target image and cannot be clearly distinguished. By using detection line marking, the overlapping second traffic sign elements in the target image can still be represented, and the position of the second traffic sign elements represented by the detection line in the target image is accurate, clear and distinguishable. This is helpful for vehicles to predict the driving path that needs to be adjusted based on the second traffic sign elements marked by the detection points.

[0054] In one example, traffic sign elements are traffic cones, and multiple traffic cones are arranged along the road's extension direction, such that multiple first traffic sign elements and multiple second traffic sign elements in the target image correspond to multiple first traffic cones and multiple second traffic cones, respectively. The multiple first traffic cones are located on the side of the target image closer to the image acquisition device used to acquire the target image, and the multiple second traffic cones are located on the side of the target image farther from the image acquisition device. For the multiple first traffic cones, the pixel distance between any two adjacent traffic cones is greater than or equal to a preset threshold; for the multiple second traffic cones, the pixel distance between any two adjacent traffic cones is less than the preset threshold. In embodiments according to this disclosure, each of the multiple first traffic cones is marked with a detection point, and the multiple second traffic cones are marked with a detection line, thereby obtaining a detection mark that simultaneously includes both detection points and detection lines.

[0055] It should be understood that the above examples illustrating the inclusion of detection points and detection lines in the detection markings are merely exemplary. Those skilled in the art should understand that detection markings that include only detection points, only detection lines, a combination of detection points and detection frames, or a combination of detection lines, detection points, and detection frames, etc., are all applicable to the embodiments of this disclosure. In other embodiments, the detection markings may also include a detection surface, that is, a plane representing the traffic sign.

[0056] It should be understood that the above example, using traffic cones as a traffic sign element to illustrate the detection points and lines obtained when detecting multiple first traffic sign elements and multiple second traffic sign elements, is merely exemplary. Those skilled in the art should understand that the multiple first traffic sign elements and multiple second traffic sign elements can be of the same type or different types; similarly, the individual first traffic sign elements can be of the same type or different types, and the individual second traffic sign elements can be of the same type or different types, all applicable to the embodiments of this disclosure. For example, multiple first traffic signs may include multiple traffic cones and multiple utility poles, and multiple second traffic signs may include multiple traffic cones and multiple water-filled barriers, etc.

[0057] In some embodiments, the target image is acquired by an image acquisition device on the vehicle, and the detection marker is located on the side of the traffic sign closest to the vehicle. Further, the detection marker is located at the vertex corner of the traffic sign closest to the vehicle. For example, assuming the vehicle with the image acquisition device is located to the left of the traffic sign, if the vehicle is moving backward, the detection marker can be located to the lower left of the traffic sign; if the vehicle is moving forward, the detection marker can be located to the upper left of the traffic sign. This facilitates the calculation of the closest distance between the vehicle and the traffic sign based on the detection marker, thus avoiding a collision.

[0058] Based on the target image obtained by the image acquisition device on the vehicle, a detection mark is obtained on the side closest to the vehicle. The vehicle's driving path can be adjusted based on this detection mark, for example, the driving path can be adjusted to avoid the detection mark.

[0059] In some embodiments, such as Figure 4 As shown, step S220, inputting the target image into the traffic sign detection model to obtain the detection markers corresponding to the traffic signs, includes:

[0060] Step S410: Obtain the response map corresponding to the target image, wherein the response value at each location in the response map is used to characterize the confidence level that a traffic sign exists at each location; and

[0061] Step S420: Generate the detection marker based on the response values ​​at each position in the response graph.

[0062] In some embodiments, the response map corresponds to the target image. The response map may correspond one-to-one with the pixel size of the target image, or it may be a scaled image of the target image at a predetermined ratio, as long as the corresponding point in the response map can be found based on the pixels in the target image. This disclosure does not impose any limitations on this.

[0063] In some embodiments, the response map corresponds to the area in the target image where the traffic sign is located. For example, the response map appears in the target image corresponding to the area where the traffic sign is located. Further, the response map corresponds to the area in the target image where the detection marker of the traffic sign is located. For example, if the detection marker of the traffic sign is in the lower left corner of the traffic sign, the response map mainly displays the response values ​​within a predetermined range in that lower left corner.

[0064] In some embodiments, generating the detection marker based on the response values ​​at each position in the response graph includes: performing a sliding window operation on the response graph using a preset sliding window region; if the response value of the target point in the sliding window region meets a predetermined condition, then generating a detection marker at the target point in the sliding window region; wherein, the predetermined condition is that the response value of the target point in the sliding window region is greater than or equal to a preset threshold and is a local maximum value of the sliding window region.

[0065] Since the response value at each location in the response map represents the confidence level of the presence of traffic signs at each location, the local maxima of the sliding window region that are greater than or equal to the threshold are used as the target points for generating detection markers. That is, detection markers are generated at the locations with the highest confidence level of the presence of traffic signs within the sliding window region, so that the generated detection markers are accurate.

[0066] It should be noted that the term "local maximum" in this disclosure refers to the maximum response value among all response values ​​at various locations within the sliding window region of the response map.

[0067] In some embodiments, the traffic sign includes multiple types of traffic sign elements, and the response map includes multiple sub-response maps corresponding to the multiple types of traffic sign elements, and the response value at each location in each of the multiple sub-response maps is used to characterize the confidence level that the corresponding type of traffic sign element exists at each location.

[0068] In embodiments according to this disclosure, for various types of traffic sign elements, corresponding sub-response maps are generated respectively, so that detection marks for various types of traffic sign elements can be generated based on the corresponding sub-response maps, making the generated detection marks accurate.

[0069] For example, traffic signs include two types of traffic sign elements: traffic cones and water-filled barriers. The generated response map includes a first sub-response map corresponding to the traffic cone and a second sub-response map corresponding to the water-filled barrier. The first sub-response map can be used to generate a detection mark corresponding to the traffic cone, and the second sub-response map can be used to generate a detection mark corresponding to the water-filled barrier, making the generated detection mark accurate.

[0070] In some embodiments, the detection marker includes multiple types of detection marker identifiers, and step S320, generating the detection marker based on the response values ​​at each position in the response graph, includes:

[0071] For each of the multiple sub-response maps, a detection mark corresponding to the type of traffic sign element is generated based on the response value at each location in that sub-response map.

[0072] For each type of traffic sign element among various types of traffic sign elements, a corresponding type of detection mark is generated based on the corresponding sub-response map, so that the corresponding type of traffic sign element can be identified based on the detection mark, and then the vehicle driving path can be adjusted according to the corresponding type of traffic sign element.

[0073] In some embodiments, the multiple types of detection markers can be detection markers of different colors, such as red detection points and detection lines, green detection points and detection lines, and blue detection points and detection lines.

[0074] For example, for traffic cones and water-filled barriers in traffic signs, red detection lines and green detection lines are generated as their respective detection markers, so that the traffic sign element in the traffic sign can be identified as a traffic cone based on the red detection line, and the traffic sign element in the traffic sign can be identified as a water-filled barrier based on the green detection line.

[0075] In some embodiments, the various types of detection markers can be detection markers of various shapes, such as circular detection points, diamond detection points, and triangular detection points.

[0076] In some embodiments, the various types of detection markers can be a combination of multiple color types and multiple shapes.

[0077] According to another aspect of this disclosure, a training method for a traffic sign detection model is also provided. It should be understood that those skilled in the art can set the model structure, parameters, and hyperparameters of the traffic sign detection model as needed, as long as it can output a corresponding response map and detection labels based on the input target image. This disclosure does not impose specific limitations on the structure and parameters of the network model. For example, the traffic sign detection model may include a convolutional neural network and a label detection module. The convolutional neural network may be a network structure such as UNet (U-shaped network) / HRNet (high-resolution network) to generate a response map corresponding to the target image, and the label detection module is used to output the final detection labels based on the response map. Those skilled in the art can also set the structure and parameters of these two sub-networks as needed, and this disclosure does not impose any limitations in this regard.

[0078] See Figure 5 The training method 500 for a traffic sign detection model according to some embodiments of the present disclosure includes:

[0079] Step S510: Obtain a training image and a labeled image corresponding to the training image. The training image includes traffic signs, and the labeled image includes labels corresponding to the traffic signs in the training image. The labels are used to mark the position of the traffic signs in the corresponding training image, and the labels include at least one of label points and label lines.

[0080] Step S520: Input the training image into the traffic sign detection model to obtain the predicted labels of the traffic signs in the training image; and

[0081] Step S530: Train the traffic sign detection model based on the predicted labels of the training images and the corresponding labeled labels of the labeled images.

[0082] The location of traffic signs in the training image is marked by at least one of the annotation marks, including annotation points and annotation lines. The detection marks used to characterize the location of traffic signs by the traffic sign detection model trained based on the training image are at least one of the points or lines. Thus, for traffic signs on the road, regardless of their form or distance from the camera device on the vehicle used to acquire the target image, the trained traffic sign detection model can detect traffic signs.

[0083] In some embodiments, traffic signs include various types of traffic sign elements, such as traffic cones, water-filled barriers, utility poles, lane lines, or trees on both sides of the road, etc., which are not limited herein.

[0084] In some embodiments, a traffic sign includes a plurality of first traffic sign elements, wherein the pixel distance between any two adjacent first traffic sign elements is greater than or equal to a preset pixel distance, and the label includes a label point corresponding to each first traffic sign element.

[0085] For multiple first traffic sign elements, since the pixel distance between adjacent first traffic sign elements is greater than or equal to a preset threshold, each of the multiple first traffic sign elements can be clearly distinguished in the target image. Each first traffic sign element is labeled with a marker point, ensuring that the position of the first traffic sign element represented by the marker point in the training image is accurate. This allows the traffic sign detection model trained on the labeled image to accurately obtain the position of the first traffic sign element in the training image.

[0086] In some embodiments, a traffic sign includes a plurality of second traffic sign elements, wherein the pixel distance between any two adjacent second traffic sign elements is less than a preset pixel distance, and the label includes a label line corresponding to the plurality of second traffic sign elements.

[0087] For multiple second traffic sign elements, since the pixel distance between adjacent second traffic sign elements is less than a preset threshold, they may overlap in the training image and cannot be clearly distinguished. By using annotation lines, the overlapping second traffic sign elements in the training image can still be represented, and the second traffic sign elements represented by the annotation lines are accurately positioned in the training image. This enables the traffic sign detection model trained on the annotated image including the annotations to accurately obtain the position of the second traffic sign elements in the training image.

[0088] In one example, such as Figure 6As shown, the traffic signs in the training images obtained by the vehicle-mounted camera device include multiple traffic cones 601, and the multiple traffic cones are arranged along the road extension direction, so that the training traffic signs included in the training images include multiple first traffic cones (such as...). Figure 6 (as shown in box C) and multiple second traffic cones (such as Figure 6 (As shown in box B). Multiple first traffic cones are located on the side of the training image closer to the image acquisition device used to acquire the training image, and multiple second traffic cones are located on the side of the training image farther from the image acquisition device. This is such that for the multiple first traffic cones, the pixel distance between any two adjacent first traffic cones is greater than or equal to a preset threshold; and for the multiple second traffic cones, the pixel distance between any two adjacent second traffic cones is less than a preset threshold. In embodiments according to this disclosure, such as... Figure 6 As shown, each of the multiple first traffic cones is labeled with a label point 620, and each of the multiple second traffic cones is labeled with a label line 610, thus obtaining a label that includes both label point 610 and label line 620.

[0089] It should be noted that the use of marking marks that include both marking points and marking lines in this embodiment is merely an example. Traffic signs can also be marked using only marking points or only marking lines.

[0090] In some embodiments, step S520, inputting the training image into a traffic sign detection model to obtain predicted labels for the traffic signs in the training image, involves the traffic sign detection model obtaining predicted labels corresponding to the traffic signs in the training image based on the training image. This includes the traffic sign detection model acquiring a predicted response map corresponding to the training image, the predicted response map including predicted response values ​​at each location, and wherein, for example... Figure 7 As shown, step S530, training the traffic sign detection model based on the predicted labels of the training images and the labeled labels of the labeled images, includes:

[0091] Step S710: Obtain the corresponding truth response map of the labeled image, wherein the truth response map includes the truth response value at each location; and

[0092] Step S720: Train the traffic sign detection model based on the predicted response value of the predicted response map and the true response value of the true response map.

[0093] The traffic sign detection model is trained using the predicted response map of the training image and the ground truth response map of the labeled image. Since the predicted response value at each location in the predicted response map represents the confidence level of the presence of a traffic sign at that location, and the ground truth response value at each location in the ground truth response map also represents the confidence level of the presence of a traffic sign at that location, the traffic sign detection model is trained based on the predicted response values ​​of the predicted response map and the ground truth response map. This enables the trained traffic sign detection model to output a predicted response map corresponding to the ground truth response map, and then the predicted label is obtained based on the predicted response map, thereby realizing the training of the traffic sign detection model.

[0094] In some embodiments, in the truth response map, the truth response values ​​at each position within a predetermined range (e.g., within a predetermined range of m pixels) of the annotation mark present a predetermined distribution (e.g., a Gaussian distribution, but not limited thereto), so that the truth response value at the first position corresponding to the location of the annotation mark is greater than the truth response value at the second position which is different from the first position.

[0095] For example, in the process of obtaining the truth response map, the truth response value of the first position in the labeled image corresponding to the location of the label mark is set to 1. Based on the distance to the first position, the truth response values ​​of the positions near the first position are set to corresponding values ​​less than 1. Specifically, the truth response value of the second position whose distance to the first position is less than or equal to the first distance S is set to 0.8, the truth response value of the third position whose distance to the first position is less than or equal to 2S (twice the first distance) is set to 0.6, and so on. The response value of the sixth position whose distance to the first position is more than 5S (five times the first distance) is set to 0.

[0096] In some embodiments, in the truth response graph, the truth response value at the marked location is a first value, and the truth response value at the unmarked location is a second value that is different from the first value.

[0097] For example, in the process of obtaining the truth response map, the truth response value of the first position in the labeled image corresponding to the location of the label mark is set to 1, and the truth response values ​​of other positions in the labeled image (positions not where the label mark is located) are set to 0.

[0098] In some embodiments, the traffic signs include multiple types of traffic sign elements, and obtaining the truth response map corresponding to the labeled image includes: obtaining multiple truth sub-response maps corresponding to the multiple types of traffic sign elements, wherein, for each of the multiple truth sub-response maps, the truth response value at each position in the truth sub-response map is used to characterize the confidence that a traffic sign element of the corresponding type exists at each position.

[0099] For multiple types of traffic sign elements, the truth sub-response maps corresponding to each type of traffic sign element are obtained respectively, so that for each type of traffic sign element among the multiple types of traffic sign elements, the corresponding detection mark can be generated based on its corresponding truth sub-response map, so that the corresponding type of traffic sign element can be obtained according to the corresponding detection mark.

[0100] In some embodiments, the training images are obtained by an image acquisition device on a vehicle, and the label is located on the side of the traffic sign closer to the vehicle.

[0101] The training images obtained from the image acquisition device on the vehicle are used to make the output of the traffic sign detection model located on the side of the image closer to the vehicle. This allows the vehicle's driving path to be adjusted based on the detection marks, for example, to avoid the detection marks.

[0102] According to another aspect of this disclosure, a traffic sign detection device is also provided.

[0103] See Figure 8 According to some embodiments of the present disclosure, a traffic sign detection apparatus 800 includes: an image acquisition unit 810 configured to acquire a target image containing traffic signs; and a detection unit 820 configured to input the target image into a traffic sign detection model to obtain a detection mark corresponding to the traffic sign, wherein the detection mark includes a detection point for characterizing the position of the traffic sign in the target image, and the detection mark includes at least one of a detection point and a detection line.

[0104] Here, the operation of each of the above-mentioned units 810 to 820 of the traffic sign detection device 800 is similar to the operation of steps S210 to S220 described above, and will not be repeated here.

[0105] According to another aspect of this disclosure, an apparatus for training a traffic sign detection model is also provided.

[0106] See Figure 9 The apparatus 900 for training a real-time traffic sign detection model according to this disclosure includes:

[0107] A first acquisition unit 910 is configured to acquire a training image and a corresponding labeled image, wherein the training image includes traffic signs, and the labeled image includes labels corresponding to the traffic signs in the training image. The labels are used to mark the positions of the traffic signs in the corresponding training image, and the labels include at least one of labeled points and labeled lines. A model prediction unit 920 is configured to input the training image into a traffic sign detection model to obtain predicted labels corresponding to the traffic signs in the training image. A model training unit 930 is configured to train the traffic sign detection model based on the predicted labels of the training image and the corresponding labeled labels of the labeled image.

[0108] Here, the operation of each of the above units 910 to 930 of the device 900 for training traffic sign detection models is similar to the operation of steps S510 to S530 described above, and will not be repeated here.

[0109] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the method as described above.

[0110] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to cause the computer to perform the method described above.

[0111] According to another aspect of this disclosure, a computer program product is also provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the method described above.

[0112] According to another aspect of this disclosure, a vehicle is also provided, including the electronic equipment described above.

[0113] Reference Figure 10 Electronic device 1000 will now be described, which is an example of a hardware device that can be applied to various aspects of this disclosure. Electronic device 1000 can be any machine configured to perform processing and / or calculations, and can be, but is not limited to, a workstation, server, desktop computer, laptop computer, tablet computer, personal digital assistant, smartphone, in-vehicle computer, or any combination thereof. The aforementioned traffic data visualization device can be implemented wholly or at least partially by electronic device 1000 or similar devices or systems.

[0114] Electronic device 1000 may include elements that are connected to or communicate with bus 1002 (possibly via one or more interfaces). For example, electronic device 1000 may include bus 1002, one or more processors 1004, one or more input devices 1006, and one or more output devices 1008. The one or more processors 1004 may be any type of processor and may include, but are not limited to, one or more general-purpose processors and / or one or more dedicated processors (e.g., special-purpose chips). Input devices 1006 may be any type of device capable of inputting information to electronic device 1000 and may include, but are not limited to, a mouse, keyboard, touchscreen, microphone, and / or remote control. Output devices 1008 may be any type of device capable of presenting information and may include, but are not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Electronic device 1000 may also include or be connected to a non-transitory storage device 1010. The non-transitory storage device may be any storage device that is non-transitory and capable of storing data, and may include, but is not limited to, disk drives, optical storage devices, solid-state storage, floppy disks, flexible disks, hard disks, magnetic tapes or any other magnetic media, optical discs or any other optical media, ROM (read-only memory), RAM (random access memory), cache memory and / or any other memory chip or cartridge, and / or any other medium from which a computer can read data, instructions, and / or code. The non-transitory storage device 1010 may be detachable from an interface. The non-transitory storage device 1010 may have data / programs (including instructions) / code for implementing the methods and steps described above. Electronic device 1000 may also include a communication device 1012. The communication device 1012 may be any type of device or system enabling communication with external devices and / or with a network, and may include, but is not limited to, modems, network interface cards, infrared communication devices, wireless communication devices and / or chipsets, such as Bluetooth. TM Devices, 1302.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.

[0115] Electronic device 1000 may also include working memory 1014, which may be any type of working memory that can store programs (including instructions) and / or data useful to the operation of processor 1004, and may include, but is not limited to, random access memory and / or read-only memory devices.

[0116] The software elements (programs) may reside in the working memory 1014, including but not limited to the operating system 1016, one or more application programs 1018, drivers, and / or other data and code. Instructions for performing the methods and steps described above may be included in one or more application programs 1018, and the aforementioned traffic data visualization device can be implemented by the processor 1004 reading and executing the instructions of one or more application programs 1018. The executable code or source code of the instructions of the software elements (programs) may be stored in a non-transitory computer-readable storage medium (e.g., the aforementioned storage device 1010) and may be stored in the working memory 1014 during execution (possibly compiled and / or installed). The executable code or source code of the instructions of the software elements (programs) may also be downloaded from a remote location.

[0117] It should also be understood that various modifications can be made depending on specific requirements. For example, custom hardware can also be used, and / or specific elements can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. For example, some or all of the disclosed methods and apparatus can be implemented by programming hardware (e.g., programmable logic circuits including field-programmable gate arrays (FPGAs) and / or programmable logic arrays (PLAs)) using logic and algorithms according to this disclosure in assembly language or hardware programming languages ​​(such as Verilog, VHDL, C++).

[0118] It should also be understood that the aforementioned methods can be implemented using a server-client model. For example, the client can receive user input data and send it to the server. Alternatively, the client can receive user input data, perform a portion of the processing described in the aforementioned methods, and send the resulting data to the server. The server can receive data from the client, execute the aforementioned methods or a portion thereof, and return the execution result to the client. The client can receive the execution result from the server and, for example, present it to the user via an output device.

[0119] It should also be understood that the components of electronic device 1000 can be distributed across a network. For example, some processing can be performed by one processor, while other processing can be performed by another processor located far away from that processor. Other components of computing system 1000 can also be distributed similarly. In this way, electronic device 1000 can be interpreted as a distributed computing system that performs processing in multiple locations.

[0120] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A method for detecting traffic signs, comprising: Acquire a target image containing traffic signs; as well as The target image is input into a traffic sign detection model to obtain detection markers corresponding to the traffic signs. The detection markers include at least one of detection points and detection lines, which are used to characterize the position of the traffic signs in the target image. The traffic signs include multiple traffic sign elements. The target image is input into the traffic sign detection model to obtain the detection markers corresponding to the traffic signs, including: Based on whether the pixel distance between any two adjacent traffic sign elements is less than a preset pixel distance, a detection line or detection point is determined to be output for the multiple traffic sign elements.

2. The method as described in claim 1, wherein, The step of inputting the target image into the traffic sign detection model to obtain the detection markers corresponding to the traffic signs includes: Obtain the response map corresponding to the target image, where the response value at each location in the response map is used to characterize the confidence level that a traffic sign exists at each location; and The detection marker is generated based on the response value at each position in the response graph.

3. The method as described in claim 2, wherein, The step of generating the detection marker based on the response values ​​at each position in the response graph includes: A sliding window operation is performed on the response map using a preset sliding window area. If the response value of the target point in the sliding window area meets a predetermined condition, a detection mark is generated at the target point in the sliding window area. The predetermined condition is that the response value of the target point in the sliding window area is greater than or equal to a preset threshold and is a local maximum value of the sliding window area.

4. The method of claim 2, wherein, The traffic signs include various types of traffic sign elements, and the process of obtaining the response map corresponding to the target image includes: Multiple sub-response maps corresponding to the various types of traffic sign elements are obtained, and the response value of each position in each sub-response map is used to characterize the confidence that the corresponding type of traffic sign element exists at each position.

5. The method of claim 4, wherein, The detection markers include multiple types of detection marker identifiers, and generating the detection markers based on the response values ​​at each position in the response graph includes: For each of the multiple sub-response maps, based on the response values ​​at each location in that sub-response map, a corresponding type of detection mark is generated for the traffic sign element of that type.

6. The method according to any one of claims 1-5, wherein, The target image is obtained by an image acquisition device on the vehicle, and the detection mark is located on the side of the traffic sign closest to the vehicle.

7. A training method for a traffic sign detection model, comprising: Acquire a training image and a corresponding labeled image. The training image includes traffic signs, and the labeled image includes labels corresponding to the traffic signs in the training image. The labels are used to mark the positions of the traffic signs in the corresponding training image, and the labels include at least one of labels and labels. The training image is input into the traffic sign detection model to obtain the predicted labels of the traffic signs in the training image; as well as The traffic sign detection model is trained based on the predicted labels of the training images and the corresponding labeled labels of the labeled images; The traffic signs include multiple traffic sign elements. The multiple traffic sign elements in the labeled image are marked as label lines or label points based on whether the pixel distance between any two adjacent traffic sign elements is less than a preset pixel distance.

8. The method of claim 7, wherein, The traffic sign includes multiple first traffic sign elements, wherein the pixel distance between any two adjacent first traffic sign elements is greater than or equal to a preset pixel distance, and the labeling includes a labeling point corresponding to each first traffic sign element.

9. The method of claim 7 or 8, wherein, The traffic sign includes multiple second traffic sign elements, wherein the pixel distance between any two adjacent second traffic sign elements is less than a preset pixel distance, and the label includes a label line corresponding to the multiple second traffic sign elements.

10. The method of claim 7, wherein, The step of inputting the training image into the traffic sign detection model to obtain the predicted labels of the traffic signs in the training image includes: The traffic sign detection model obtains a predicted response map corresponding to the training image based on the training image, the predicted response map including the predicted response value at each location; and wherein, Training the traffic sign detection model based on the predicted labels of the training images and the labeled labels of the labeled images includes: Obtain the corresponding ground truth map of the labeled image, the ground truth map including the ground truth response value at each location; and The traffic sign detection model is trained based on the predicted response values ​​of the predicted response map and the true response values ​​of the true response map.

11. The method of claim 10, wherein, The truth response diagram: The true value response values ​​at each location within the predetermined range of the marked markers exhibit a predetermined distribution; or the true value response value at the marked markers is a first value, and the true value response value at the unmarked markers is a second value that is different from the first value.

12. The method of claim 10, wherein, The traffic signs include various types of traffic sign elements, and obtaining the ground truth response map corresponding to the labeled image includes: Multiple truth sub-response maps corresponding to the various types of traffic sign elements are obtained, wherein, for each of the multiple truth sub-response maps, the truth response value at each position in the truth sub-response map is used to characterize the confidence that the corresponding type of traffic sign element exists at each position.

13. The method of claim 7, wherein, The training images were obtained by an image acquisition device on the vehicle, and the markings were located on the side of the traffic sign closest to the vehicle.

14. A traffic sign detection device, comprising: The image acquisition unit is configured to acquire a target image containing traffic signs; as well as The detection unit is configured to input the target image into the traffic sign detection model to obtain the detection markers corresponding to the traffic signs. The detection marker includes a detection point, which is used to characterize the position of the traffic sign in the target image, and the detection marker includes at least one of a detection point and a detection line; The traffic sign includes multiple traffic sign elements, and the detection unit is further used for: Based on whether the pixel distance between any two adjacent traffic sign elements is less than a preset pixel distance, a detection line or detection point is determined to be output for the multiple traffic sign elements.

15. An apparatus for training a traffic sign detection model, comprising: The first acquisition unit is configured to acquire a training image and a labeled image corresponding to the training image. The training image includes traffic signs, and the labeled image includes labels corresponding to the traffic signs in the training image. The labels are used to mark the position of the traffic signs in the corresponding training image, and the labels include at least one of label points and label lines. The model prediction unit is configured to input the training image into the traffic sign detection model to obtain the predicted label corresponding to the traffic sign in the training image; as well as The model training unit is configured to train the traffic sign detection model based on the predicted labels of the training images and the corresponding labeled labels of the labeled images; The traffic signs include multiple traffic sign elements. The multiple traffic sign elements in the labeled image are marked as label lines or label points based on whether the pixel distance between any two adjacent traffic sign elements is less than a preset pixel distance.

16. An electronic device comprising: At least one processor; as well as At least one memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the method of any one of claims 1-13.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-13.

18. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-13.

19. A vehicle comprising the electronic equipment as claimed in claim 16.

Citation Information

Patent Citations

  • Traffic sign recognition method and related device

    CN112052778A

Cited By

  • Traffic marker detection method and training method for traffic marker detection model

    EP4170601A1