Traffic light detection and recognition method and device, mobile tool, and storage medium
By introducing the concept of regional feature combination in the traffic light detection and recognition algorithm, and combining the lamp position frame and the area frame, the problems of low accuracy and high error detection rate of traffic light detection and recognition models in the prior art are solved, and reliable detection and recognition of traffic light signal status are achieved and the accuracy of traffic lights is improved.
Patent Information
- Application Number
- CN202210530117.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-05-16
AI Technical Summary
In the prior art, the accuracy of the traffic light detection and identification model is not high, and missed detection and missed detection are prone to occur, and it is difficult to avoid the problem of high false detection rates caused by interference factors such as traffic lights, poor resolution of light frames, street lights, headlights, and taillights at night.
By introducing the concept of regional feature combination in the traffic light detection and recognition algorithm, the traffic light detection and recognition method is designed and optimized by combining two regional features of the light disk position frame and the area frame to improve the accuracy and recall of detection and recognition, and reduce false detection and missed detection.
It realizes reliable detection and identification of traffic light signal status in the autonomous driving system, reduces false detection and missed detection rates, improves the accuracy and recall of traffic light detection and identification, and especially effectively avoids mis-detection of vehicle rearview mirrors, rear taillights, and non-traffic lights at night.
Smart Images

Figure CN114972731B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a traffic light detection and recognition method, a traffic light detection and recognition device, a mobile tool and a storage medium. Background Art
[0002] In the field of automatic or semi-automatic driving, the perception ability of the driving vehicle is an important foundation and basis for achieving safe and reliable automatic or semi-automatic driving decisions. At present, vehicles are generally equipped with multiple sensors including lidar, ultrasonic radar, cameras, etc. to provide road perception capabilities for the vehicle. Among them, the camera provides visual perception capabilities for the automatic or semi-automatic driving system. By collecting and analyzing the image data collected by the camera, it can provide a reliable status basis for downstream decision-making control.
[0003] In automatic or semi-automatic driving systems, the timely and accurate detection and recognition of traffic light signals by automatic or semi-automatic driving vehicles is an important basis for downstream decision-making and control, and is a guarantee for ensuring that the vehicle travels safely, smoothly and compliantly on the road. With the gradual maturity of target detection methods based on deep learning, the use of various methods to solve the problem of traffic light detection and recognition has gradually become a mainstream trend. How to ensure that the traffic light detection and recognition model outputs the signal state in real time and accurately, especially how to avoid the adverse effects on the reliability of the signal state caused by image false detection and missed detection problems, and how to avoid the high false detection rate of traffic lights caused by interference factors such as traffic lights at night, poor light frame resolution, street lights, headlights, taillights, etc., has become a new problem that the industry is constantly exploring and urgently needs to solve. Summary of the invention
[0004] The embodiment of the present invention provides a traffic light detection and recognition solution to solve the problem that the traffic light detection and recognition model used in the prior art has low accuracy and is prone to missed detection and false detection.
[0005] In a first aspect, an embodiment of the present invention provides a traffic light detection and recognition method, which comprises:
[0006] Inputting the collected image to be identified into a pre-trained target detection model to obtain prediction information of the collected image, wherein the target detection model is trained based on a combination of at least two image features, and the at least two image features include a light panel position frame for representing the position of the traffic light and a region frame for representing the region of interest;
[0007] The recognition result of the traffic light in the collected image is determined according to the prediction information, wherein the recognition result includes the light panel position frame information corresponding to the traffic light and its color status.
[0008] In a second aspect, an embodiment of the present invention provides a traffic light recognition and detection device, which comprises
[0009] A detection and recognition module, used for inputting the collected image to be recognized into a pre-trained target detection model to obtain prediction information of the collected image, wherein the target detection model is trained based on a combination of at least two image features, and the at least two image features include a light panel position frame for representing the position of the traffic light and a region frame for representing the region of interest;
[0010] The result determination module is used to determine the recognition result of the traffic light in the collected image according to the prediction information, wherein the recognition result includes the light panel position frame information corresponding to the traffic light and its color status.
[0011] In a third aspect, an embodiment of the present invention provides another traffic light recognition and detection device, which includes a memory for storing executable instructions; and
[0012] The processor is used to execute the executable instructions stored in the memory, and the executable instructions, when executed by the processor, implement the steps of the traffic detection and identification method of any embodiment of the present invention.
[0013] In a fourth aspect, an embodiment of the present invention provides a mobile tool, which includes: a traffic light recognition and detection device according to the third aspect of the present invention.
[0014] In a fifth aspect, the present invention provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the traffic detection and identification method of any embodiment of the present invention.
[0015] In a sixth aspect, the present invention provides a computer program product, comprising a computer program stored on a non-volatile computer-readable storage medium, wherein the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer executes the traffic detection and identification method of any embodiment of the present invention.
[0016] The beneficial effects of the embodiments of the present invention are as follows: the method provided by the embodiments of the present invention introduces the concept of regional feature combination into the traffic light detection and recognition algorithm, and designs and optimizes the traffic light detection and recognition method by combining two regional features, namely the light panel position frame and the regional frame. It can take into account high recall and traffic light detection and recognition accuracy on the basis of the algorithm based on image detection and recognition of traffic lights, reduce false detections and missed detections, and especially avoid false detections of vehicle rearview mirrors, rear tail lights, non-traffic lights at night, etc. through regional feature combination, which helps to provide reliable traffic light signal status for the automatic driving system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0018] Figure 1 A flow chart of a traffic light detection and recognition method according to an embodiment of the present invention;
[0019] Figure 2 A flow chart of a method for training a target detection model according to an embodiment of the present invention;
[0020] Figure 3 This is a diagram showing the effect of region marking according to an embodiment of the present invention. Figure 3 A is the traffic light status effect diagram when it is not marked. Figure 3 B is the effect picture after marking the position frame of the lamp panel. Figure 3 C is Figure 3 The effect diagram after continuing to mark the area frame based on B;
[0021] Figure 4 A flow chart of a method for optimizing prediction information according to an embodiment of the present invention;
[0022] Figure 5 A schematic diagram of a network structure of a target detection model according to an embodiment of the present invention;
[0023] Figure 6 The schematic diagram shows a functional block diagram of a traffic light detection and recognition device according to an embodiment of the present invention;
[0024] Figure 7 It is a principle block diagram of a traffic light detection and identification device according to another embodiment of the present invention;
[0025] Figure 8 This is a principle block diagram of a mobile tool according to an embodiment of the present invention;
[0026] Fig. 9 It is a schematic structural diagram of an embodiment of a traffic light detection and identification device of the present invention. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application may be combined with each other.
[0029] The present invention may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0030] In the present invention, "module", "device", "system" and the like refer to related entities applied to computers, such as hardware, a combination of hardware and software, software or software in execution, etc. In detail, for example, an element can be, but is not limited to, a process, a processor, an object, an executable element, an execution thread, a program and / or a computer running on a processor. In addition, an application or script program running on a server, a server can all be an element. One or more elements can be in an execution process and / or thread, and an element can be localized on a computer and / or distributed between two or more computers, and can be operated by various computer-readable media. An element can also communicate through local and / or remote processes according to a signal with one or more data packets, for example, a signal from a data that interacts with another element in a local system, a distributed system, and / or a network on the Internet through a signal to interact with other systems.
[0031] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include" and "comprise" include not only those elements, but also other elements that are not explicitly listed, or also include elements inherent to such processes, methods, articles or equipment. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the existence of other identical elements in the process, method, article or equipment that includes the elements.
[0032] The traffic light detection and recognition method in the embodiment of the present invention can be applied to a traffic light detection and recognition device, so that a user or an autonomous driving device can use the traffic light detection and recognition device to obtain the detection and recognition result of the traffic light, so as to control the autonomous driving device to perform the corresponding driving action according to the recognition result. These traffic light detection and recognition devices include, for example, but are not limited to, detection identifiers on autonomous driving vehicles, perception modules on autonomous driving vehicles, smart tablets, personal PCs, computers, cloud servers, etc. In particular, the traffic light detection and recognition method in the embodiment of the present invention can also be directly applied to autonomous driving devices such as autonomous driving vehicles, and the present invention does not limit this.
[0033] Figure 1 The traffic light detection and recognition method according to an embodiment of the present invention is schematically shown. The execution subject of the method may be an autonomous driving domain controller or computing system mounted on a robot, autonomous driving vehicle or other equipment, or a cloud server that wirelessly communicates with the robot or autonomous driving vehicle, etc., and the embodiment of the present invention is not limited to this. Figure 1 As shown, the method of the embodiment of the present invention includes:
[0034] Step S10: inputting the collected image to be identified into a pre-trained target detection model to obtain prediction information of the collected image, wherein the target detection model is trained based on a combination of at least two image features, and the at least two image features include a light panel position frame for representing the position of the traffic light and a region frame for representing the region of interest;
[0035] Step S11: determining a recognition result of a traffic light in the collected image according to the prediction information, wherein the recognition result includes a light panel position frame information corresponding to the traffic light and its color state;
[0036] In step S10, the collected image to be identified refers to the collected image that needs to be detected and identified by the traffic light, which can be a real-time image obtained from the perception module. For example, it can be a real-time image collected by the RGB camera on the autonomous driving device. The collected image to be identified can be obtained by directly reading the real-time collected image from the perception module, so as to detect and identify the input collected image using the pre-trained target detection model to obtain the prediction information of the traffic light in the collected image. Among them, the detection and identification capabilities that can be provided by the target detection model are generated by pre-training it.
[0037] As a preferred implementation example, the prediction information obtained by the target detection model in step S10 can be implemented as including the area frame information in the captured image for representing the area of interest in the image, the light panel position frame information for representing the position of the traffic light in the captured image, and the color state for representing the color of the traffic light corresponding to the light panel position frame. Figure 2 The method process for training a target detection model capable of providing the recognition capability according to an embodiment of the present invention is schematically shown. Figure 2 As shown, the target detection model used in step S10 can be exemplarily implemented as trained by the following method:
[0038] Step S101: forming training sample data by annotating image features in the captured image, wherein the image features include a light panel position frame formed based on the position annotations of each traffic light in the image and an area frame formed by annotating a group of light panel position frames that meet the conditions according to a predefined area frame division condition, and the training sample data includes area frame information corresponding to each image, light panel position frame information and color status corresponding to the light panel position frame;
[0039] Step S102: training the selected network model using the training sample data to determine the model parameters of the network model;
[0040] Step S103: forming a trained target detection model according to the determined model parameters and the selected network model.
[0041] In step S101, training sample data can be formed by annotating the image features of the collected image. Specifically, it can include two annotation processes, namely, the annotation of the light panel position frame according to the position of the traffic light in the image and the annotation of the regional frame according to the light panel position frame. Through the two annotation processes, the embodiment of the present invention can not only realize the feature extraction of the light panel position of the traffic light, but also divide the light panel position into regions to define the region of interest, thereby realizing the fusion of the light panel position and regional feature information of the traffic light, and the target detection model trained thereby can simultaneously fuse the two image features for detection and recognition. Since the light panel combination feature has better stability than the single light panel feature, the false detection rate is much lower than that of detecting a single light panel. Therefore, the target detection model trained by the embodiment of the present invention can effectively improve the accuracy of the traffic light detection and recognition results by combining the light panel position and regional feature information of the traffic light. Exemplarily, the marking of the position frame in the captured image can be specifically implemented as follows: first, determining whether the number of traffic light panels in the image exceeds two, and when the number exceeds two, marking the panel position frame for each traffic light panel in the image; after marking the panel position frame, then selecting a group of panel position frames that meet the conditions according to the predefined area frame division condition for area frame marking, so as to add the selected group of panel position frames that meet the conditions to the same area frame. As a preferred implementation example, the marked area frame can be defined as the minimum circumscribed rectangular frame of the selected group of panel frames that meet the conditions. Exemplarily, the predefined area frame division condition can be to divide at least two traffic light panels directly in front of the vehicle at the same traffic intersection into a group, that is, after marking the panel position frames of these traffic light panels, the panel position frames corresponding to the traffic light panels divided into a group will be marked in the same area frame, such as selecting the minimum circumscribed rectangular frame of the panel position frame corresponding to the group of traffic light panels as the area frame for marking. Figure 3 The process of region labeling is schematically shown, Figure 3 As shown, in Figure 3 A shows the traffic light status in an image. In this example scenario, the image collected at the traffic light intersection includes three traffic light panels. Each traffic light panel corresponds to a traffic light. Each traffic light panel a is independent. By judging the number of traffic light panels in the image, the panel position frame b will be marked for each traffic light panel a in the image. Figure 3 As shown in FIG. 1 , by marking the light panel position frame b, each traffic light panel a is taken as an interesting feature extracted independently from the image; then, the marked light panel position frame b in the image is classified into regions according to the pre-defined region frame division conditions, and the light panel position frames that meet the preset conditions are divided into the same region frame, such as Figure 3As shown in C, the three light panel position frames b are divided into the same area frame c, and at least two traffic light panels a are regionally combined by marking the area frame c, and the information after the regional combination is also extracted as an independent feature of interest. In this way, the feature extraction of all traffic light panels in the image is completed, so that each traffic light panel not only has the feature label of the light panel position frame, but also is added to different area frames respectively to realize the regional feature combination of the traffic light panel. The training sample data obtained by the regional feature combination method simultaneously includes the regional frame information corresponding to each image, the light panel position frame information and the color state corresponding to the light panel position frame. In steps S102 and S103, the pre-selected network model is trained using the training sample data obtained by the regional combination method, and a target detection model that can simultaneously integrate the light panel position and area category of the traffic light can be obtained. Exemplarily, the target detection model trained in this way is a target detection network model that takes the collected original image as input and outputs the regional frame information, the light panel position frame information of the traffic light and its color state as prediction information.
[0042] In order to ensure a better model training effect, illustratively, the data source of the training sample data in step S101, that is, the image collected to form the training sample data through annotation processing, can be an image with a resolution of not less than 720P collected by using an RGB camera in an open scene, and the image acquisition environment conditions include daytime and nighttime, and there is no clear occlusion data. More preferably, before annotation, the collected images can also be preprocessed to enhance the diversity and richness of the sample data, thereby ensuring the recognition accuracy of the trained target detection model. Exemplarily, the preprocessing can be implemented to include left and right rotation of the image, 90-degree rotation of the image, adaptive image scaling, image stitching, etc. Specifically, the above-mentioned preprocessing method can be used to perform random operations on the selected image data of a preset proportion, such as randomly selecting half of the data in the image data, and performing any one of the above four operations or a combination of two or more of them on the selected image data to enhance the richness of the data.
[0043] As another preferred embodiment, the marked area box can also be defined as the minimum circumscribed circular box of a group of selected lamp panel frames that meet the conditions. As long as the feature combination of the position and area type of the lamp panel can be achieved to further improve the accuracy of the recognition result by combining the features, the embodiment of the present invention does not limit the specific shape of the area box.
[0044] As a preferred embodiment, the region frame labeling process of the embodiment of the present invention can be completed by automated calculation. Exemplarily, the automated calculation process may include first counting the number n of traffic light panels in an image, and then judging the number of traffic light panels in the image. If it is judged that n>=2, then it is calculated whether the current n light panels are approximately on the same straight line. If they are on a straight line, then the current light panels are considered to be a group. The position coordinates of the n light panels and the labeled light panel position frame can be used to calculate the minimum circumscribed rectangle surrounding the n light panels as the region frame of the n light panels, and the n light panels are labeled with region frames. More preferably, after completing the automated calculation process, manual assistance can be used to eliminate erroneous labeling items, thereby completing the entire region frame labeling process. Using an automated calculation process to label the region frames can greatly improve the labeling speed, save labor costs, and improve the labeling efficiency of training sample data.
[0045] Thus, by pre-training the model to form a target detection model that can provide specific functional services, in step S10, the real-time acquired image to be identified can be directly input into the trained target detection model to obtain the desired prediction information. Exemplarily, the prediction information output by the target detection model can be in the following data form: light panel position frame information-red (Red), light panel position frame information-yellow (Yellow), light panel position frame information-green (Green), light panel position frame information-black (Black), region frame information (ROI). Among them, red (Red), yellow (Yellow), green (Green) and black (Black) are the current color states of the traffic lights corresponding to each light panel position frame. In a specific implementation example, the color state of the traffic light can also be other desired colors according to the needs, and the embodiment of the present invention is not regarded as a limitation on the color state type. Exemplarily, the light panel position frame information and the area frame information can be specifically implemented as information that can be used to characterize the position range of the light panel position frame and the area frame in the captured image, for example, it can be information represented by coordinates (x, y, w, h), where x, y are the upper left corner coordinates of the light panel position frame and the area frame, and w, h are the width and height of the light panel position frame and the area frame, thereby defining the position area range of the light panel position frame and the area frame to accurately identify a light panel and an area category. In other embodiments, the light panel position frame information and the area frame information can also include the coordinates and area area of the frame, etc., which is not limited by the embodiment of the present invention. It should be noted that the default order of the light panel position frames in the prediction information output by the embodiment of the present invention is formed by sorting from left to right in sequence based on the vehicle as the coordinate system, and this order can correspond to the coordinate order in the captured image, and thus can accurately correspond to the order and position of the light panels of each traffic light in the captured image.
[0046] Exemplarily, the target detection model in the embodiment of the present invention can be any network model for target detection in the target detection networks such as Yolo series networks, SSD, CenterNet, RCNN, Fast RCNN, Faster RCNN, etc. Among them, the target detection model preferably used in the embodiment of the present invention is the Yolov5 network model.
[0047] Therefore, the area frame features in the captured image provided in the embodiment of the present invention can be integrated with the light panel position frame features, that is, the area frame information can be used to characterize the area category of the light panel position frame. Therefore, in step S11, the embodiment of the present invention can optimize the detection and recognition of traffic lights by utilizing the area combination features of the light group of traffic lights, such as using the area frame information as verification information for regional redundancy cross-check of the light panel position frame information, so as to filter and screen the prediction information output by the target detection model through cross-check of the area combination features, avoid false detection and wrong detection, and improve the accuracy of detection and recognition of traffic lights.
[0048] Specifically, in one of the embodiments of the present invention, taking the case where the prediction information includes the area frame information, the light panel position frame information and the color state corresponding to the light panel position frame in the acquired image as an example, in step S11, after the prediction information is obtained using the target detection model, the area frame information and the light panel position frame information in the prediction information are also used to optimize the prediction results detected and identified using the target detection model, and the valid light panel position frame is screened out, and the light panel position frame information and its color state of the screened out valid light panel position frame are used as the recognition result of the traffic light in the acquired image, so as to improve the reliability of the output traffic light position and color state. Preferably, the embodiment of the present invention performs false detection filtering on the light panel position frame of the output traffic light by means of regional combination, so as to achieve optimization processing of the prediction results by false detection filtering. Exemplarily, the embodiment of the present invention achieves false detection filtering by determining the positional relationship between each light panel position frame and the area frame, Figure 4 The optimization process is schematically shown as follows: Figure 4 As shown, its implementation includes:
[0049] Step S111: determining a first overlap degree corresponding to each lamp panel position frame according to the lamp panel position frame information and the area frame information;
[0050] Step S112: determining a second overlap degree corresponding to the area frame according to the area frame information;
[0051] Step S113: Filter out valid lamp panel position frame information from the lamp panel position frame information according to the first overlap, the preset first overlap threshold, the second overlap and the preset second overlap threshold.
[0052] After the above-mentioned optimization processing and screening out the valid light panel position frame information, the embodiment of the present invention uses the screened out valid light panel position frame information and its corresponding color status as the final recognition result of the traffic light in the captured image to be recognized, thereby obtaining the traffic light position and color status with higher accuracy, and realizing more precise detection and recognition of the traffic light.
[0053] Among them, Figure 4 In the illustrated embodiment, the positional relationship between each light panel position frame and the area frame is characterized by two overlap degrees, namely, a first overlap degree and a second overlap degree, wherein, exemplarily, the first overlap degree is used to characterize the validity of the light panel position frame information predicted by the target detection model, and the second overlap degree is used to characterize the validity of the area frame information predicted by the target detection model.
[0054] As a preferred implementation example, in step S111, an embodiment of the present invention can determine the first degree of overlap corresponding to each light panel position frame by judging whether the light panel position frame falls within the range of the area frame. Taking the first degree of overlap determined based on the positional relationship between the center point of each light panel position frame and the area frame, and the light panel frame position information and the area frame information are information represented by coordinates (x, y, w, h) as an example, an embodiment of the present invention can determine the center point of each light panel position frame by the formula center_x=x+w / 2, center_y=y+h / 2, wherein the center point of each determined light panel position frame is represented by the coordinates (center_x, center_y), center_x is the coordinate of the center point in the x direction, and center_y is the coordinate of the center point in the y direction. After determining the center point of each lamp panel position frame, the positional relationship and the first degree of overlap between each lamp panel position frame and the area frame can be determined based on the center point coordinates and the area frame information. For example, the first degree of overlap can be determined by comparing the center point coordinates with the area frame coordinates to determine whether the center point coordinates fall within the area frame. For example, the positional relationship when the center point coordinates fall within the area frame is determined as a first degree of overlap of 1, and the situation when the center point coordinates do not fall within the area frame, that is, are outside the area frame, is determined as a first degree of overlap of 0.
[0055] As a preferred embodiment, in step S112, the second degree of overlap can be represented by the intersection-over-union (IoU) of the region box. For example, the second degree of overlap can be determined by calculating the intersection-over-union (IoU) of the region box. Wherein, the intersection-over-union (IoU) is a concept used in target detection, which refers to the overlap rate between the candidate box predicted by the target detection model and the original marked box, that is, the ratio of the intersection and union of the candidate box and the original marked box. The most ideal situation is complete overlap, that is, the ratio is 1. Therefore, in multi-target tracking, the intersection and union can be used to determine the similarity between the tracking box and the target detection box. It should be noted that in the embodiment of the present invention, when calculating the intersection and union of the region box, that is, the second degree of overlap, the region box information in the prediction information can be selected as the candidate box, and the projection frame information corresponding to the region box information can be selected as the original marked box, so as to determine the second degree of overlap of the region box by calculating the ratio of the intersection and union of the region box information and the projection frame information. Exemplarily, the acquisition method of the projection frame information corresponding to the area frame information can be determined based on the traffic light position in the high-precision map, such as by querying the traffic light position in the high-precision map that is pre-marked and stored, and generating the area frame corresponding to the traffic light group according to the traffic light position in the high-precision map, and then projecting the generated area frame into the current captured image to be identified, so as to obtain the projection frame information corresponding to the area frame corresponding to the traffic light group, and use the projection frame information as the original marking frame corresponding to the area frame information in the prediction information. In another embodiment, the acquisition of the projection frame information can also be implemented in a target tracking manner, which can specifically include selecting the area frame information in the captured image of the previous frame of the current frame that has been identified as the projection frame information corresponding to the area frame information in the prediction information of the captured image of the current frame, that is, as the original marking frame corresponding to the area frame information in the prediction information of the captured image of the current frame. Thus, the second overlap degree of the area frame can be calculated according to the area frame information in the prediction information and the corresponding projection frame information obtained, and the specific calculation method can be implemented with reference to the prior art, and the embodiment of the present invention will not be repeated here.
[0056] The embodiment of the present invention will pre-set a first overlap threshold used to characterize whether the lamp panel position frame exceeds the range of the regional frame and a second overlap threshold used to characterize the validity of the regional frame. After calculating the first overlap and the second overlap, in step S114, the embodiment of the present invention will implement false detection filtering based on the comparison result of the first overlap and the first overlap threshold and the comparison result of the second overlap and the second overlap threshold. Among them, the embodiment of the present invention preferably determines the lamp panel position frame whose center point exceeds the range of the regional frame as not being within the regional frame, that is, it means that the prediction of the lamp panel position frame is wrong and the lamp panel position frame information is invalid, and determines the regional frame information whose intersection-over-union ratio of the regional frame is less than the preset value as the prediction of the regional frame is wrong, that is, the predicted regional frame information is invalid. Therefore, the first overlap is recorded as D center , the second overlap is denoted as D IoU , the first overlap threshold is denoted as β, and the second overlap threshold is denoted as α. By judging whether D IoU ≥α, and delete those that do not satisfy D IoU ≥α, we can filter out the cases where the predicted region frame is considered invalid because the region frame intersection ratio is small; and by judging whether there is D center ≥β, and delete those that do not satisfy D center ≥β condition, the predicted lamp panel position frame information can filter out the situation where the lamp panel position frame is not within the area frame and thus the predicted lamp panel position frame is regarded as a false detection item, thereby filtering out the valid lamp panel position frame information. Among them, in the preferred embodiment of the present invention, the valid lamp panel position frame information is the one that satisfies D center ≥β condition and the area frame where the lamp panel position frame is located also satisfies D IoU ≥α condition, thereby detecting the valid light panel position frame based on the validity of the two features of light panel position and area, and further improving the accuracy of traffic light detection and recognition.
[0057] Since the lamp panel position frame whose center point exceeds the range of the area frame is not within the area frame, the lamp panel position frame that is not within the area frame can be deleted from the prediction result of the target detection model as a false detection item by judging the first overlap; and since the intersection of the area frames is relatively low, indicating that the correlation between the predicted range and the true range is low, the invalid area frame can be deleted from the prediction output result of the target detection model as a false detection item by judging the second overlap. In order to realize the regional feature combination, the embodiment of the present invention also preferably requires that the area frame that meets the second overlap requirement must also include the lamp panel position frame, that is, the area frame without the lamp panel position frame is also deleted as an invalid situation. Therefore, by filtering the false detection of the lamp panel position frame and the area frame separately and in combination, the two false detection scenarios of invalid area frame and lamp panel position frame not being in the area frame can be filtered out, and the scenarios that meet both D and D requirements can be filtered out.IoU ≥α and D center ≥β condition, that is, the light panel position frame is within the range of the area frame and the area frame meets the intersection and union ratio requirements, and those light panel position frames are output as valid detection and recognition results. Therefore, the light panel position frame information used to identify the location of the traffic light and the color state used to identify the current color of the traffic light will be more accurate.
[0058] Therefore, the embodiment of the present invention performs target detection and recognition on the captured image, and after the image is detected and recognized, the recognition result is filtered and optimized by introducing a regional combination method, so that the false detection items can be filtered out, and thus the detection and recognition results of the traffic light are more accurate. And by introducing the regional combination method, the false detection situation is effectively filtered out, and the false detection rate is effectively reduced, so that the false detection situation caused by the vehicle rearview mirror, rear taillight, night non-traffic light, etc. can be avoided, and a reliable traffic light signal state detection and recognition result is provided for the automatic driving system.
[0059] In other implementation examples, the second overlap in the embodiment of the present invention can be represented by IoU, and can also be replaced by GIoU, DIoU, CIoU, etc., as long as the invalid area frame can be detected and filtered out. The embodiment of the present invention does not limit the specific overlap standard parameters selected. Among them, the specific calculation method of GIoU, DIoU, and CIoU can refer to the relevant existing technology, and the embodiment of the present invention will not be repeated here.
[0060] As another preferred embodiment of the present invention, the target detection model trained in step S10 can be a network model that can provide the above-mentioned functional services, and can also be implemented as a target detection model that takes the collected image to be identified as input and the light panel position frame information and color state of the traffic light as output. In this case, the selected network model can be trained in advance to introduce the regional combination features formed by the light panel position frame and the region frame into the target detection network model, so as to train a target detection model that can achieve high-precision end-to-end detection and recognition of the traffic light signal state. In such an embodiment, in step S10, it is only necessary to input the collected image to be identified into the trained target detection model to obtain the prediction information including the light panel position frame information and its color state corresponding to the traffic light. Therefore, in step S11, the prediction information output by the target detection module can be directly determined as the recognition result of the traffic light in the collected image, so as to improve the accuracy of the end-to-end output of the traffic light panel position frame and its color state, and improve the accuracy of traffic light detection and recognition, and reduce the false detection rate.
[0061] Exemplarily, the embodiment of the present invention can train such a target detection model based on the principle of the attention mechanism. Specifically, the two image features of the light panel position box and the area box can be extracted, and the feature vector weighting of the two image features can be performed to achieve the combination of regional features of the traffic light, thereby improving the interest in the regional combination features in the internal structure of the network model, thereby training a network model that is more sensitive to the area where the light panel is located, so as to improve the accuracy of end-to-end detection. Therefore, the embodiment of the present invention can achieve the desired training of a target detection model that meets the end-to-end detection recognition features by integrating the two regional combination features of the light panel position box and the area box into the network structure of the target detection model. In one of the embodiments of the present invention, in order to train such a target detection model, the selected network model needs to satisfy the following structural requirements in addition to the following requirements: including a backbone network for extracting features from an input image to be identified and generating a first feature map output, a fully connected module for performing a fully connected operation, and a regression module for performing a regression operation; it also needs to satisfy the following requirements: after the backbone network and before the fully connected operation, the feature map is weighted using the regional combination coding information to achieve shared reuse of the image depth feature map, so that the traffic light target detection task and the classification and recognition task can be completed simultaneously through an end-to-end network, thereby improving the accuracy of the end-to-end output of the traffic light panel frame and the color status of the light panel. In a specific implementation, the network model selected to meet this requirement is preferably a Yolov5 network model. Of course, it can also be a target detection network model selected from other Yolo series networks, SSD, CenterNet, RCNN, Fast RCNN or Faster RCNN, etc., and the embodiment of the present invention does not limit this. Taking the network model selected by the present invention as a yolov5 network as an example, Figure 5 The network structure design diagram of the trained target detection model according to one embodiment of the present invention is schematically shown. Figure 5 As shown, the network structure of the target detection model in the embodiment of the present invention includes:
[0062] A backbone network 50 for extracting features from an input image to be identified and generating a first feature map output;
[0063] Used to perform RoI Pooling operation on the first feature map to generate a mapping pool 51 output based on the second feature map of the region of interest, wherein the region of interest is the region frame information in the collected image to be identified;
[0064] A gradient module 52 for performing gradient sorting on the second feature map;
[0065] A first weighting module 53 for weighting the gradient information to the second feature map;
[0066] A second weighting module 54 for weighting the second feature map to the first feature map;
[0067] A fully connected module 55 for performing a fully connected operation;
[0068] and a regression module 56 for performing a regression operation.
[0069] Among them, in the above-mentioned network structure of the embodiment of the present invention, the first feature map is a feature map that uses the light panel position frame as the feature extraction, and the second feature map is a feature map that uses the area frame as the feature extraction. Since the area frame usually has a large gradient, the gradient module of the embodiment of the present invention will sort the second feature map by gradient from high to low according to the gradient size of the feature map, thereby improving the sensitivity to the area frame information. In addition, the target detection model of the embodiment of the present invention is also designed with a two-layer weighting module, which weights the gradient information to the second feature map and the second feature map to the first feature map respectively. Through the two-layer weighting, the regional feature vector of the traffic light group is also weighted to the first feature map, that is, the area frame feature is weighted to the light panel position frame feature, thereby improving the sensitivity of the target detection model to the divided traffic light group area, so that the detection and recognition result output by the model is more accurate and the false detection rate is lower.
[0070] Preferably, in an embodiment of the present invention, the regression operation performed by the regression module of the target detection model is a regression operation of calculating the color state classification and the traffic light panel position frame information, so that the output result of the target detection model is the traffic light panel position frame information and its corresponding color state.
[0071] Therefore, the embodiment of the present invention realizes the embedding of the regional feature combination concept into the yolov5 network, improves the accuracy of the end-to-end output of the traffic light position frame and color state, and reduces the false detection rate. In addition, the regional combination encoding method of the embodiment of the present invention can realize the weighted interest level of the traffic light group area, which is essentially consistent with the attention mechanism (Attention). Therefore, by weighting the regional feature vector of the traffic light group, the sensitivity of the network to the divided traffic light group area is improved, and the false detection rate is reduced.
[0072] In other embodiments, the network design structure of the target detection model that can achieve this goal may not be limited to Figure 5 The network structure shown can be replaced by other network design structures, as long as it can achieve the fusion and attention of the regional features of the divided traffic light groups, so as to improve the sensitivity to the divided areas by weighting the regional feature vectors of the traffic light groups, thereby reducing the false detection rate. Exemplarily, such a target detection model can also be replaced by a transform design network.
[0073] Figure 6 A traffic light recognition and detection device according to an embodiment of the present invention is schematically shown. Figure 6 As shown, the device comprises:
[0074] The detection and recognition module 30 is used to input the collected image to be recognized into a pre-trained target detection model to obtain prediction information of the collected image, wherein the target detection model is trained based on a combination of at least two image features, and the at least two image features include a light panel position frame for representing the position of the traffic light and a region frame for representing the region of interest;
[0075] The result determination module 31 is used to determine the recognition result of the traffic light in the collected image according to the prediction information, wherein the recognition result includes the light panel position frame information corresponding to the traffic light and its color status.
[0076] As a preferred embodiment, the pre-trained target detection model is a target detection network model that takes the original image collected as input and outputs the region frame information, the light panel position frame information of the traffic light and its color status as prediction information. After the detection and recognition module 30 inputs the collected image acquired in real time as the collected image to be recognized into the target detection model, the prediction information obtained includes the region frame information, the light panel position frame information of the traffic light and its color status. In the result determination module 31, a regional feature redundancy cross-check in a regional combination manner will be performed based on the region frame information in the prediction information and the light panel position frame information of the traffic light to screen out the valid light panel position frame information, and the screened out valid light panel position frame information and its corresponding color status will be used as the location information of the traffic light in the collected image and the current color status of the traffic light that are finally detected and recognized.
[0077] As another preferred embodiment, the pre-trained target detection model is an end-to-end target detection network model that uses the collected original image as input and the traffic light panel position frame information and its color state as the prediction information output. After the detection and recognition module 30 inputs the collected image acquired in real time as the collected image to be recognized into the target detection model, the obtained prediction information includes the traffic light panel position frame information and its color state. In the result determination module 31, the prediction information will be directly used as the position information of the traffic light in the collected image and the current color state of the traffic light in the final detection and recognition. In this embodiment, specifically, the regional features of the light panel position frame and the regional frame are weightedly combined within the target detection network model, so that the trained end-to-end target detection model itself has already taken into account the regional combination features of the traffic light panel, so that the output prediction information has a higher detection and recognition accuracy. Exemplarily, the network structure of the target detection model of the embodiment of the present invention may include:
[0078] A backbone network for extracting features from an input image to be identified and generating a first feature map output, wherein the feature extracted from the first feature map is a lamp panel position frame in the image to be identified;
[0079] Used to perform a RoI Pooling operation on the first feature map to generate a mapping pool output based on the second feature map of the region of interest, wherein the region of interest is a region box in the image to be identified;
[0080] a gradient module for performing gradient sorting on the second feature map;
[0081] A first weighting module for weighting gradient information to a second feature map;
[0082] a second weighting module for weighting the second feature map to the first feature map;
[0083] A fully connected module for performing fully connected operations; and
[0084] Regression module for performing regression operations.
[0085] Among them, the specific implementation process of the detection and identification module 30 and the result determination module 31, as well as the training process of the target detection model, can refer to the description of the method part above, and will not be repeated here.
[0086] Figure 7 The traffic light recognition and detection device according to another embodiment of the present invention is schematically shown. Figure 7 As shown, the device includes
[0087] A memory 40 for storing executable instructions; and
[0088] The processor 41 is used to execute the executable instructions stored in the memory, and the executable instructions implement the steps of the method of any of the above-mentioned embodiments of the invention when executed by the processor.
[0089] Figure 8 A real-time mobile tool of the present invention is schematically shown, such as Figure 8 As shown, the mobile tool includes the traffic light recognition and detection device 70 according to any of the above-mentioned embodiments, so that the mobile tool can detect and identify traffic lights using the traffic light recognition and detection device arranged thereon, and then perform subsequent control such as direction, acceleration, throttle, brake, etc. based on the determined traffic light position and color status.
[0090] Optionally, in actual applications, the mobile tool may also include a perception and recognition module and other planning and control modules, such as a path planning controller, a bottom-level controller, etc. The functions of the traffic light recognition and detection device 70 may also be implemented in the perception and recognition module or the planner, etc. The embodiment of the present invention does not limit this.
[0091] The “mobile tool” referred to in the embodiment of the present invention may be a vehicle with an L0-L5 autonomous driving technology level as established by the Society of Automotive Engineers International (SAE International) or the Chinese national standard “Automotive Driving Automation Classification”.
[0092] Exemplarily, the mobile tool may be a vehicle device or a robot device having the following various functions:
[0093] (1) Passenger-carrying function, such as family cars and buses;
[0094] (2) Cargo-carrying functions, such as ordinary trucks, van trucks, trailer trucks, closed trucks, tank trucks, flatbed trucks, container trucks, dump trucks, trucks with special structures, etc.;
[0095] (3) Tool functions, such as logistics delivery vehicles, automated guided vehicles (AGVs), patrol cars, cranes, hoists, excavators, bulldozers, forklifts, road rollers, loaders, off-road engineering vehicles, armored engineering vehicles, sewage treatment vehicles, sanitation vehicles, vacuum trucks, floor scrubbers, water sprinklers, sweeping robots, food delivery robots, shopping guide robots, lawn mowers, golf carts, etc.;
[0096] (4) Entertainment functions, such as entertainment vehicles, amusement park self-driving devices, balance vehicles, etc.;
[0097] (5) Special rescue functions, such as fire trucks, ambulances, power repair vehicles, engineering rescue vehicles, etc.
[0098] In some embodiments, an embodiment of the present invention provides a non-volatile computer-readable storage medium, which stores one or more programs including execution instructions, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute the traffic light detection and identification method of any of the above embodiments of the present invention.
[0099] In some embodiments, an embodiment of the present invention further provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the traffic light detection and identification method of any one of the above embodiments.
[0100] In some embodiments, an embodiment of the present invention also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the traffic light detection and identification method of any of the above embodiments.
[0101] In some embodiments, an embodiment of the present invention further provides a storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the traffic light detection and recognition method of any of the above embodiments.
[0102] Fig. 9 FIG. 1 is a schematic diagram of the hardware structure of an electronic device for executing a traffic light detection and recognition method provided by another embodiment of the present invention. Fig. 9 As shown, the device includes:
[0103] One or more processors 610 and memory 620, Fig. 9 A processor 610 is taken as an example.
[0104] The device for executing the traffic light detection and recognition method may further include: an input device 630 and an output device 640 .
[0105] The processor 610, the memory 620, the input device 630 and the output device 640 may be connected via a bus or other means. Fig. 9 The example of connecting through bus is taken in the following.
[0106] The memory 620 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the traffic light detection and recognition method in the embodiment of the present invention. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 620, that is, the traffic light detection and recognition method in the above method embodiment is implemented.
[0107] The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the traffic light detection and recognition method, etc. In addition, the memory 620 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 620 may optionally include a memory remotely arranged relative to the processor 610, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0108] The input device 630 may receive input digital or character information and generate signals related to user settings and function control of the image processing device. The output device 640 may include a display device such as a display screen.
[0109] The memory 620 may be configured to store instructions executable by the processor 610 to perform various functions, including but not limited to positioning fusion, perception, driving state determination, navigation module, decision making, driving control, task reception, etc.
[0110] The processor 610 may be configured to execute programs (instructions) stored in the memory 620 to perform various functions.
[0111] The one or more modules are stored in the memory 620, and when executed by the one or more processors 610, perform the traffic light detection and recognition in any of the above method embodiments.
[0112] The above product can execute the method provided by the embodiment of the present invention, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not described in detail in this embodiment, please refer to the method provided by the embodiment of the present invention.
[0113] The electronic device of the embodiments of the present invention may exist in various forms, including but not limited to an autonomous driving domain controller or a computing system mounted on a robot, an autonomous driving vehicle and the like, wherein the computing system may include multiple computing devices that control individual components or individual systems of the robot or the autonomous driving vehicle in a distributed manner.
[0114] The electronic device of the embodiment of the present invention may also be a cloud server that wirelessly communicates with a robot or an autonomous driving vehicle.
[0115] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0116] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A traffic light detection and recognition method, characterized in that: The method comprises: Inputting the collected image to be identified into a pre-trained target detection model to obtain prediction information of the collected image, wherein the target detection model is trained based on a combination of at least two image features, the at least two image features include a light panel position frame for representing the position of the traffic light and an area frame for representing the area of interest, and the prediction information includes the area frame information in the collected image, the light panel position frame information and the color state corresponding to the light panel position frame; Determine a first overlap degree corresponding to each lamp panel position frame according to the lamp panel position frame information and the area frame information; Determining a second overlap degree corresponding to the area frame according to the area frame information and an original marking frame corresponding to the area frame information; Filtering out valid lamp panel position frame information from the lamp panel position frame information according to the first overlap, the preset first overlap threshold, the second overlap, and the preset second overlap threshold; The filtered out valid light panel position frame information and its color status are used as the recognition result of the traffic light in the collected image.
2. The method according to claim 1, characterized in that The training process of the target detection model includes: The training sample data is formed by annotating image features in the captured image, wherein the image features include a light panel position frame formed based on the position annotations of each traffic light in the image and an area frame formed by annotating a group of light panel position frames that meet the conditions according to a predefined area frame division condition, and the training sample data includes area frame information corresponding to each image, light panel position frame information, and color status corresponding to the light panel position frame; Using the training sample data to train the selected network model, and determining the model parameters of the network model; A trained target detection model is formed based on the determined model parameters and the selected network model.
3. The method according to claim 2, characterized in that The area frame is defined as the minimum circumscribed rectangular frame of a group of lamp panel position frames selected to meet the conditions.
4. The method according to claim 1, characterized in that: The determining, according to the area frame information and the original marking frame corresponding to the area frame information, a second overlap degree corresponding to the area frame includes: The area box information is used as a candidate box, and the projection box information corresponding to the area box information is used as an original marking box. The intersection and union ratio of the area box information is calculated based on the candidate box and the original marking box, and the intersection and union ratio of the area box information is used as the second overlap degree corresponding to the area box.
5. The method according to claim 4, characterized in that The projection frame information is determined in the following manner: Generate an area frame corresponding to the traffic light according to the location of the traffic light in the high-precision map; The generated area frame is projected onto the current collected image to be identified, and projection frame information corresponding to the area frame corresponding to the traffic light is obtained as projection frame information corresponding to the area frame information.
6. The method according to claim 4, characterized in that The projection frame information is determined in the following manner: According to the recognition result of the captured image of the frame before the current frame in the continuously captured images, the area frame information in the captured image of the frame before the current frame that has been recognized is selected as the projection frame information corresponding to the area frame information in the prediction information of the captured image of the current frame.
7. Traffic light recognition and detection device, characterized in that: include: A detection and recognition module, used for inputting a captured image to be recognized into a pre-trained target detection model to obtain prediction information of the captured image, wherein the target detection model is trained based on a combination of at least two image features, the at least two image features include a light panel position frame for representing the position of the traffic light and an area frame for representing the area of interest, and the prediction information includes the area frame information in the captured image, the light panel position frame information and the color state corresponding to the light panel position frame; A result determination module is used to determine the recognition result of the traffic light in the collected image according to the prediction information, specifically including determining the first overlap corresponding to each light panel position frame according to the light panel position frame information and the area frame information; determining the second overlap corresponding to the area frame according to the area frame information and the original marking frame corresponding to the area frame information; filtering out valid light panel position frame information from the light panel position frame information according to the first overlap, a preset first overlap threshold, a second overlap and a preset second overlap threshold; and using the filtered valid light panel position frame information and its color status as the recognition result of the traffic light in the collected image.
8. Traffic light recognition and detection device, characterized in that: include A memory for storing executable instructions; and A processor, configured to execute executable instructions stored in a memory, wherein the executable instructions, when executed by the processor, implement the steps of the method according to any one of claims 1 to 6.
9. A mobile tool, characterized in that: include: The traffic light recognition and detection device as described in claim 8.
10. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.
11. A computer program product, comprising a computer program stored on a non-volatile computer-readable storage medium, wherein the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Indicating information identification method and device of indicating lamp, electronic equipment and storage medium
CN112149697A
Signal lamp identification model training method and device, and signal lamp identification method and device
CN113963331A