Traffic signal lamp identification method and device, storage medium and electronic equipment

By using high-definition maps and vehicle forward-looking camera sensors in traffic light recognition, combined with grouping and neural network models, the problem of low recognition accuracy in existing technologies is solved, and high-accuracy traffic light recognition in complex scenarios is achieved.

CN120612666APending Publication Date: 2025-09-09BEIJING ZHIXINGZHE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410262202.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing traffic light recognition solutions have low recognition accuracy in complex scenarios. In particular, solutions based on high-precision maps are affected in accuracy in the case of multiple light ROIs. Purely image-based solutions cannot adapt to various complex scenarios.

Method used

By obtaining high-definition maps, vehicle positioning information and forward-looking camera sensor images, traffic lights in the map are grouped according to the grouping distance threshold and projected into the image to generate regions of interest. A neural network model is then used for identification and analysis to improve recognition accuracy.

Benefits of technology

It realizes the multi-view camera complementarity of traffic lights in complex scenes, makes up for the blind spot of the front-facing camera, and improves the accuracy of recognition results and detection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612666A_ABST
    Figure CN120612666A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic signal lamp recognition scheme, and belongs to the technical field of intelligent driving, and the method comprises the steps: obtaining a high-definition map, vehicle positioning information, and a plurality of first images collected by a vehicle foresight camera sensor; grouping the traffic lights in the high-definition map according to the vehicle positioning information and a preset traffic light grouping distance threshold; projecting traffic lights in the high-definition map into the first images to generate a plurality of second images; screening out a target image from the plurality of second images; respectively generating a region of interest for each traffic signal lamp group in the target image; detecting a traffic signal lamp frame and a countdown frame in each region of interest; and analyzing the image areas in the traffic signal lamp frame and the countdown frame to obtain a traffic signal lamp recognition result. According to the traffic signal lamp identification scheme provided by the invention, the traffic signal lamp information can be accurately and comprehensively identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving technology, and in particular to a traffic light recognition method and device, a storage medium, and an electronic device. Background Art

[0002] Traffic lights are crucial hubs connecting road intersections. Only by accurately and timely perceiving traffic light signals can autonomous vehicles ensure safe, smooth, and compliant navigation. Traffic light image perception tasks for autonomous driving primarily include traffic light target detection, color recognition, countdown recognition, sub-light attribute recognition, and flashing state recognition. Currently, two common traffic light image perception solutions exist: high-precision map-based traffic light perception and purely image-based traffic light perception.

[0003] The first traffic light perception technology based on HD maps requires projecting traffic lights onto images captured by the vehicle’s forward-looking camera sensor based on the relationship between traffic lights and roads in the HD map, and then generating a ROI (region of interest) for the traffic lights in the image. Figure 1 As shown in the figure, since multiple traffic light projections are included in one ROI, it is easy to introduce more invalid areas, affecting the accuracy of traffic light recognition results.

[0004] The second purely image-based traffic light perception solution uses a front-facing camera sensor installed on the vehicle to collect images and then analyze the images to obtain traffic light recognition results. An example of a front-facing camera sensor layout in an autonomous driving system is shown in the following figure. Figure 2 As shown in the figure, the camera is rear-mounted, usually mounted on the roof, and includes two front-facing cameras (a long focal length narrow field of view camera and a short focal length wide field of view camera), a left front-facing camera (a short focal length wide field of view camera), and a right front-facing camera (a short focal length wide field of view camera). However, using only the complementary images collected by the front-facing cameras to recognize traffic lights is not applicable to various complex traffic light recognition scenarios. Figure 2Taking the front-view camera sensor layout in as an example, the front-view telephoto camera is an H30 camera, the front-view short-focus camera is an H120 camera, the right-side front-view short-focus camera is an H100 camera, and the left-side front-view short-focus camera is an H100 camera. The blind spot of the H30 camera is approximately 0-25m, and the blind spot of the H120 camera is approximately 0-5m. In scenarios where the distance between the traffic light and the stop line is less than 5m, the short-focus H120 camera may not be able to see the traffic light in the short-focus camera image. In some left-turn waiting area scenarios, since the road only provides traffic lights ahead of the straight road, the short-focus H120 image may also easily fail to see the traffic light before and after the vehicle enters the stop line of the left-turn waiting area.

[0005] It can be seen that the existing traffic light recognition solutions all have the problem of low recognition accuracy. Summary of the Invention

[0006] The purpose of the embodiments of the present invention is to provide a traffic light recognition method and device, a storage medium, and an electronic device, which can solve the problem of low accuracy of traffic signal recognition results in the prior art.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0008] An embodiment of the present invention provides a method for identifying a traffic light, wherein the method includes:

[0009] Obtaining a high-definition map, vehicle positioning information, and multiple first images captured by a vehicle forward-looking camera sensor; wherein the high-definition map includes road structure information and traffic light information;

[0010] Grouping the traffic lights in the high-definition map according to the vehicle positioning information and a preset traffic light grouping distance threshold;

[0011] Projecting the traffic lights in the high-definition map onto each of the first images to generate a plurality of second images;

[0012] Selecting a target image from the plurality of second images;

[0013] generating a region of interest for each traffic light group in the target image;

[0014] Detecting the traffic light frame and countdown frame within each of the regions of interest;

[0015] The image areas in the traffic light frame and countdown frame are analyzed to obtain a traffic light recognition result.

[0016] Optionally, the step of grouping traffic lights in the high-definition map according to the vehicle positioning information and a preset traffic light grouping distance threshold includes:

[0017] Detecting traffic lights within a preset distance of the vehicle based on the vehicle positioning information, road structure information contained in the high-definition map, and traffic light signal information;

[0018] For each of the traffic lights, querying the distance information and structure information corresponding to the traffic light from the traffic light information;

[0019] Determine the relative distance between every two traffic lights;

[0020] Traffic lights whose relative distance is less than a preset traffic light grouping distance threshold are grouped into the same traffic light group.

[0021] Optionally, the step of projecting the traffic lights in the high-definition map onto each of the first images to generate a plurality of second images includes:

[0022] For one of the traffic lights, converting the longitude and latitude coordinates of the traffic light in the high-definition map into 3D coordinates in a geocentric coordinate system;

[0023] Converting the 3D coordinates in the geocentric coordinate system into the 3D coordinates in the vehicle coordinate system;

[0024] The 3D coordinates in the vehicle coordinate system are converted into 2D coordinates in the image coordinate system based on the intrinsic parameter matrix and the extrinsic parameter matrix of the forward-looking camera sensor that captured the first image, thereby forming a projection light frame of the traffic light and generating a second image.

[0025] Optionally, after the step of converting the 3D coordinates in the vehicle coordinate system into 2D coordinates in the image coordinate system based on the intrinsic parameter matrix and the extrinsic parameter matrix of the forward-looking camera sensor that captured the first image to form a projection light frame of the traffic light, the method further includes:

[0026] For each traffic light projected into the first image, calculating a first relative displacement difference between a projection light frame corresponding to the traffic light and an image frame captured in the first image; wherein the first relative displacement difference serves as a correction value of the image coordinate system;

[0027] The projection light frame corresponding to the traffic light is aligned with the image according to the first relative displacement difference.

[0028] Optionally, the step of selecting a target image from the plurality of second images includes:

[0029] screening valid images in the second image according to preset conditions;

[0030] In the case where there are multiple valid images, the valid image corresponding to the forward-looking camera sensor with the highest priority is selected from the multiple valid images as the target image according to the priority ranking of the forward-looking camera sensors corresponding to the valid images.

[0031] Among them, the preset conditions include: all projections of at least one group of traffic lights fall into the second image; the positions of the projection light frame of the traffic signal falling into the second image, the corrected projection light frame, and the image frame of the traffic signal in the previous frame of the second image all meet the preset boundary conditions.

[0032] Optionally, the step of generating a region of interest for each traffic light group in the target image includes:

[0033] For each traffic light group in the target image, calculating the center point coordinates of all traffic light reference frames within the traffic light group; wherein the reference frames include: a projection frame and an image frame;

[0034] Draw a rectangular area with the center point coordinate as the center, wherein the short side of the rectangular area is a preset value, and the long side of the rectangular area is the product of the preset value and the number of traffic lights included in the traffic light group;

[0035] An intersection area between the rectangular area and the target image is determined as a region of interest of the traffic light group.

[0036] Optionally, the step of analyzing the image areas in the traffic light frame and countdown frame to obtain a traffic light recognition result includes:

[0037] matching the traffic light frame with the rectified projected light frame of the traffic light in the target image;

[0038] Calculating a second relative displacement difference between the matched projection frame and the traffic light frame; wherein the second relative displacement difference is used as a correction value when correcting the projection frame;

[0039] determining the position information of the traffic light frame and the countdown frame according to the second relative displacement difference;

[0040] Classify the traffic light frame based on the pre-trained neural network model to determine the target type to which the traffic light frame belongs;

[0041] Perform light color smoothing on the image area in the traffic light frame and identify the flashing state of the traffic light to obtain the color and flashing state of the traffic light;

[0042] Perform digital smoothing on the image area in the countdown frame to obtain the countdown number;

[0043] The traffic light identification, the distance from the vehicle to the stop line, the distance from the vehicle to the traffic light, the distance from the stop line to the traffic light, and the attributes of each sub-light of the traffic light are determined based on the position information of the traffic light frame and the road structure information contained in the high-definition map.

[0044] An embodiment of the present invention provides a traffic light recognition device, comprising:

[0045] An acquisition module, configured to acquire a high-definition map, vehicle positioning information, and a plurality of first images captured by a vehicle front-view camera sensor; wherein the high-definition map includes road structure information and traffic light information;

[0046] a grouping module, configured to group the traffic lights in the high-definition map according to the vehicle positioning information and a preset traffic light grouping distance threshold;

[0047] a projection module, configured to project the traffic lights in the high-definition map onto each of the first images to generate a plurality of second images;

[0048] a screening module, configured to screen out a target image from a plurality of second images;

[0049] a generating module, configured to generate a region of interest for each traffic light group in the target image;

[0050] A detection module, configured to detect a traffic light frame and a countdown frame within each of the regions of interest;

[0051] The result recognition module is used to analyze the image areas in the traffic light frame and countdown frame to obtain a traffic light recognition result.

[0052] Optionally, the grouping module includes:

[0053] a traffic light detection submodule, configured to detect traffic lights within a preset distance of the vehicle based on the vehicle positioning information, the road structure information contained in the high-definition map, and the traffic light signal information;

[0054] a query submodule, configured to query, for each of the traffic lights, distance information and structure information corresponding to the traffic light from the traffic light information;

[0055] A distance determination submodule is used to determine the relative distance between every two traffic lights;

[0056] The division submodule is used to divide traffic lights whose relative distance is less than a preset traffic light grouping distance threshold into the same traffic light group.

[0057] Optionally, the projection module includes:

[0058] A first conversion submodule is configured to convert, for one of the traffic lights, the longitude and latitude coordinates of the traffic light in the high-definition map into 3D coordinates in a geocentric coordinate system;

[0059] A second conversion submodule is used to convert the 3D coordinates in the geocentric coordinate system into the 3D coordinates in the vehicle coordinate system;

[0060] The third conversion submodule is used to convert the 3D coordinates in the vehicle coordinate system into 2D coordinates in the image coordinate system based on the intrinsic parameter matrix and extrinsic parameter matrix of the forward-looking camera sensor that captured the first image, thereby forming a projection light frame of the traffic light and generating a second image.

[0061] Optionally, the projection module further includes:

[0062] a first calculation submodule, configured to calculate, for each traffic light projected into the first image, a first relative displacement difference between a projection light frame corresponding to the traffic light and an image frame captured in the first image; wherein the first relative displacement difference serves as a correction value of the image coordinate system;

[0063] An alignment submodule is configured to align the projection light frame corresponding to the traffic light with the image according to the first relative displacement difference.

[0064] Optionally, the screening module includes:

[0065] A valid image screening submodule, configured to screen valid images in the second image according to preset conditions;

[0066] The selection submodule is used to select the valid image corresponding to the forward-looking camera sensor with the highest priority as the target image from the multiple valid images according to the priority sorting of the forward-looking camera sensors corresponding to the valid images when there are multiple valid images.

[0067] Among them, the preset conditions include: all projections of at least one group of traffic lights fall into the second image; the positions of the projection light frame of the traffic signal falling into the second image, the corrected projection light frame, and the image frame of the traffic signal in the previous frame of the second image all meet the preset boundary conditions.

[0068] Optionally, the generating module includes:

[0069] A coordinate calculation submodule is configured to calculate, for each traffic light group in the target image, the coordinates of the center points of all traffic light reference frames within the traffic light group; wherein the reference frames include: a projection frame and an image frame;

[0070] an area drawing submodule, configured to draw a rectangular area with the center point coordinate as the center, wherein the short side of the rectangular area is a preset value, and the long side of the rectangular area is the product of the preset value and the number of traffic lights included in the traffic light group;

[0071] The region of interest determination submodule is configured to determine an intersection area between the rectangular area and the target image as the region of interest of the traffic light group.

[0072] Optionally, the result identification module includes:

[0073] a matching submodule, configured to match the light frame of the traffic light with the rectified projected light frame of the traffic light in the target image;

[0074] A second calculation submodule is configured to calculate a second relative displacement difference between the matched projection frame and the traffic light frame; wherein the second relative displacement difference is used as a correction value when correcting the projection frame;

[0075] a position information determining submodule, configured to determine the position information of the traffic light frame and the countdown frame according to the second relative displacement difference;

[0076] A type determination submodule is used to classify the traffic light frame based on a pre-trained neural network model to determine the target type of the traffic light frame;

[0077] A processing submodule is used to perform light color smoothing processing on the image area in the traffic light frame and identify the flashing state of the traffic light to obtain the color and flashing state of the traffic light;

[0078] A digital processing submodule is used to perform digital smoothing on the image area in the countdown frame to obtain the countdown number;

[0079] The information determination submodule is used to determine the traffic light logo, the distance from the vehicle to the stop line, the distance from the vehicle to the traffic light, the distance from the stop line to the traffic light, and the attributes of each sub-light of the traffic light based on the position information of the traffic light frame and the road structure information contained in the high-definition map.

[0080] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of any one of the above-mentioned traffic light recognition methods are implemented.

[0081] An embodiment of the present invention provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of any one of the above-mentioned traffic signal light recognition methods are implemented.

[0082] The traffic light recognition solution provided by an embodiment of the present invention obtains a high-definition map, vehicle positioning information, and multiple first images captured by a vehicle's forward-looking camera sensor; groups the traffic lights in the high-definition map based on the vehicle positioning information and a preset traffic light grouping distance threshold; projects the traffic lights in the high-definition map onto each of the first images to generate multiple second images; selects a target image from the multiple second images; generates a region of interest for each traffic light group in the target image; detects the traffic light frame and countdown frame within each region of interest; and analyzes the image areas in the traffic light frame and countdown frame to obtain a traffic light recognition result. The traffic light recognition solution provided by the embodiment of the present invention, on the one hand, performs traffic light recognition based on images collected by high-definition maps and vehicle forward-looking camera sensors, which can achieve multi-view camera complementarity, make up for the problem that traffic lights cannot be perceived due to the blind spots of the forward-looking camera, and improve the accuracy of traffic light recognition results; on the other hand, traffic lights are grouped based on the spatial position relationship of traffic lights in the high-precision map, and regions of interest are generated for each group, which can improve the ability to detect and recognize traffic lights, and thus improve the accuracy of traffic light recognition results. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 It is a schematic diagram of the traffic light ROI in the image;

[0084] Figure 2 is a schematic diagram showing the layout of the forward-looking camera sensor in an autonomous driving system;

[0085] Figure 3 is a flowchart showing the steps of a traffic light recognition method according to an embodiment of the present application;

[0086] Figure 4 is a schematic diagram showing another layout of forward-looking camera sensors in an autonomous driving system;

[0087] Figure 5 is a schematic diagram showing a vehicle querying a traffic light;

[0088] Figure 6It is a schematic diagram showing a traffic light intersection scene;

[0089] Figure 7 is a schematic diagram showing images before and after projection of a traffic light;

[0090] Figure 8 Schematic diagram showing the comparison before and after projection correction;

[0091] Figure 9 1 is a schematic diagram showing a traffic light recognition process according to an embodiment of the present application;

[0092] Figure 10 It is a schematic diagram showing three principles of generating ROIs for traffic light groups;

[0093] Figure 11 It is a schematic diagram showing the classification process of traffic lights;

[0094] Figure 12 It is a schematic diagram showing a traffic light frame and a countdown frame;

[0095] Figure 13 It is a schematic diagram showing the input and output information of the traffic light perception system;

[0096] Figure 14 is a structural block diagram showing a traffic light recognition device according to an embodiment of the present application;

[0097] Figure 15 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0098] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0099] The traffic light recognition solution provided by the embodiment of the present application is described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.

[0100] As attached Figure 3 As shown, the traffic light recognition method of the embodiment of the present application includes the following steps:

[0101] Step 301: Obtain a high-definition map, vehicle positioning information, and multiple first images captured by a vehicle front-view camera sensor.

[0102] Among them, the high-definition map contains road structure information and traffic light information.

[0103] In the embodiments of this application, the vehicle is equipped with multiple forward-looking camera sensors. During autonomous driving or advanced driver assistance systems, electronic devices in the vehicle, such as the autonomous driving center console, acquire multi-dimensional information to identify traffic lights and control the vehicle based on the traffic light recognition results. The autonomous driving center console is equipped with a traffic light perception system, which performs traffic light recognition.

[0104] The installation position of the front camera sensor in the vehicle can be flexibly set by the car manufacturer, for example, Figure 2 or Figure 4 The installation layout shown in .

[0105] like Figure 4 As shown in the figure, for an L3 or higher autonomous driving vehicle, the camera sensor is referred to as the camera and is usually installed on the roof. It includes two front-facing cameras (a long focal length narrow field of view camera and a short focal length wide field of view camera), a left front-facing camera (a short focal length wide field of view camera), and a right front-facing camera (a short focal length wide field of view camera). Figure 2 As shown in the figure, for an L2+ level advanced assisted driving vehicle, the camera is front-mounted, and the front-facing camera is usually installed at the position of the driving recorder. It generally includes two front-facing cameras (a long-focal-length narrow-field-of-view camera and a short-focal-length wide-field-of-view camera), a left-front-facing camera (a short-focal-length wide-field-of-view camera), and a right-front-facing camera (a short-focal-length wide-field-of-view camera). The left and right front-facing cameras are usually installed in different locations for different car models, generally in the B-pillar, rearview mirror, and on both sides of the roof.

[0106] HD maps use a structured data format and need to include road structured information such as traffic light information, lane information, stop line information, the binding relationship between lane lines and the traffic lights and stop lines at the corresponding intersections, and the topological relationship between lane lines. Traffic light information includes the traffic light's unique identifier (signal_id), the traffic light's spatial coordinates (x, y, z), the traffic light type, and the sub-light type, where (x, y) is the longitude and latitude coordinates and z is the height of the traffic light. Stop line information includes the stop line's unique identifier (stopline_id), the stop line's spatial coordinates (x, y, z = 0), where (x, y) is the longitude and latitude coordinates and z is the height of the stop line. Lane line information includes the lane_id (lane line unique identifier), the preceding lane line's lane_id, the succeeding lane line's lane_id, the lane line's spatial coordinates (x, y, z = 0), and the signal_id and stopline_id of the traffic light to which it is bound.

[0107] In the embodiment of the present application, inertial sensors and global positioning system positioning information can be combined to provide high-precision positioning for the vehicle and obtain accurate vehicle positioning information.

[0108] Each vehicle's front-view camera sensor can capture a first image to Figure 2 Taking the vehicle front-view camera sensor layout shown in as an example, a total of four first images can be captured, namely, one image captured by the front-view telephoto H30 camera, one image captured by the front-view short-focus H120 camera, one image captured by the right-side front-view short-focus H100 camera, and one image captured by the left-side front-view short-focus H100 camera. The blind spots of the two front-view cameras are compensated by the two side front-view cameras. In addition to obtaining the first images captured by the four front-view cameras, it is also necessary to obtain the internal and external parameters of each front-view camera. In this application, the internal and external parameters of each front-view camera stored in the default system are used as an example for explanation.

[0109] In an embodiment of the present application, traffic light recognition is performed based on a high-precision map, i.e., HDMap, and multiple forward-looking camera sensors. This mainly relies on the road structure information and the absolute position of the traffic lights in the high-precision map, as well as the positioning information of the vehicle itself. It is predetermined whether there are traffic lights within a certain range (for example, 100m to the stop line, where the braking distance is set according to different road grades). The information of the traffic lights on the road that has been determined to exist is combined with the internal and external parameters of multiple forward-looking camera sensors to calculate their approximate positions in the visual image, i.e., the first image, as the projection of the traffic lights on the forward-looking camera image by the HDMap. Furthermore, based on the priority of the forward-looking camera sensors, the area where the traffic lights are located is calculated in the image with the highest priority as the region of interest, i.e., ROI. Then, the detection and classification model of the deep learning neural network is used to identify the traffic light information such as the position, color, sub-light type, countdown status, etc. of the traffic lights on the ROI image. Finally, through smooth tracking, stable state recognition of the traffic lights is achieved.

[0110] Step 302: Group the traffic lights in the high-definition map based on the vehicle positioning information and a preset traffic light grouping distance threshold.

[0111] The acquired high-definition map, vehicle positioning information, and multiple first images serve as the basic data for traffic light recognition. After obtaining this basic data, it is preprocessed to ultimately select a target image for recognition. This preprocessing of the basic data may include, but is not limited to, searching for traffic lights in the high-definition map, grouping traffic lights, converting 3D data into 2D data, and selecting a target image. The processing operations at each stage are described below.

[0112] Step 302 is the operation step of searching for traffic lights in the HD map and grouping the traffic lights. Optionally, based on vehicle positioning information and a preset traffic light grouping distance threshold, grouping the traffic lights in the HD map may include the following sub-steps:

[0113] Sub-step 1: detecting traffic lights within a preset distance of the vehicle based on the vehicle positioning information, road structure information, and traffic light signal information contained in the high-definition map;

[0114] Road structured information includes lane markings, traffic lights, stop signs, and the relationship between them. The preset distance can be set to the sensing distance of the front telephoto camera. For example, if the front telephoto camera is H30, the preset distance can be set to 300m.

[0115] If there is a traffic light within the preset distance of the vehicle, sub-steps 2 to 4 are executed; if there is no traffic light within the preset distance, the process returns to step 303 to obtain basic data for the next round of traffic light recognition.

[0116] Sub-step 2: for each traffic light, query the distance information and structure information corresponding to the traffic light from the traffic light information.

[0117] Traffic light structure information may include: global position, light type, sub-light type, etc. Distance information includes the distance from the vehicle to the stop line, the distance from the vehicle to the traffic light, and the distance from the stop line to the traffic light. Figure 5 shown.

[0118] Sub-step 3: determining the relative distance between every two traffic lights;

[0119] Sub-step 4: grouping traffic lights whose relative distance is less than a preset traffic light grouping distance threshold into the same traffic light group.

[0120] According to real-world installation specifications for suspended traffic lights, the distance between traffic lights is typically less than the width of two lanes, and the width of urban motorway lanes is typically less than 3.5 meters. Therefore, the traffic light grouping distance threshold (distance_threshold) can be preset to a value less than 7 meters, such as 6 meters, 5 meters, or 5.5 meters. It should be noted that the traffic light grouping distance threshold can be dynamically adjusted based on the number of lanes and traffic lights on the current road.

[0121] Attachment Figure 6 A schematic diagram of a traffic light intersection scene is shown, where: Figure 6 (a) There are two traffic lights on both sides of the stop line. Figure 6(b) shows two common traffic lights hanging in front of the stop line. According to the position coordinates in the map, we can calculate Figure 6 The relative distance between the two traffic lights in (a) is greater than the preset traffic light grouping distance threshold, so the two traffic lights are divided into two groups. Figure 6 The relative distance between the two traffic lights in (b) is less than the preset traffic light grouping distance threshold, so the two traffic lights are grouped together.

[0122] Step 303: Project the traffic lights in the high-definition map onto each first image to generate multiple second images.

[0123] This step is to convert 3D data into 2D data during the basic data preprocessing process. In an optional embodiment, the method of projecting traffic lights in the high-definition map onto each first image to generate multiple second images may include the following sub-steps:

[0124] Sub-step 1: for a traffic light, convert the longitude and latitude coordinates of the traffic light in the high-definition map into 3D coordinates in the geocentric coordinate system;

[0125] Sub-step 2: convert the 3D coordinates in the geocentric coordinate system into the 3D coordinates in the vehicle coordinate system;

[0126] Sub-step 3: Based on the intrinsic parameter matrix and extrinsic parameter matrix of the forward-looking camera sensor that captured the first image, the 3D coordinates in the vehicle coordinate system are converted to 2D coordinates in the image coordinate system to form a projection light frame of the traffic light, thereby generating a second image.

[0127] The image diagram of traffic light before and after projection is as shown in the attached figure. Figure 7 As shown, two projection light frames are projected on the upper left of the two traffic lights. In actual implementation, since multiple first images are collected by the front-view camera sensor, this method is needed to project the traffic lights in the high-definition map into each first image.

[0128] This optional data conversion method can project traffic light signals in the high-definition map into a 2D image to improve the comprehensiveness and reliability of traffic light information in the image.

[0129] Affected by the camera calibration parameters and the positioning error of the global positioning system, the position of the traffic light projected into the image is not completely reliable. Especially during driving, the projection is easily offset up, down, left and right due to the influence of the ground slope and vehicle bumps. Therefore, the projection position is dynamically adapted in the embodiment of the present application.

[0130] In an optional embodiment, after converting the 3D coordinates in the vehicle coordinate system to 2D coordinates in the image coordinate system based on the intrinsic parameter matrix and extrinsic parameter matrix of the forward-looking camera sensor that captured the first image to form the projection frame of the traffic light, the projection position can be dynamically adapted as follows:

[0131] For each traffic light projected into the first image, a first relative displacement difference between a projection light frame corresponding to the traffic light and an image frame captured in the first image is calculated; and based on the first relative displacement difference, the projection light frame corresponding to the traffic light is aligned with the image.

[0132] An exemplary comparison diagram of projection correction before and after is shown in the attached figure. Figure 8 As shown. The first relative displacement difference serves as a correction value of the image coordinate system. In a feasible method, after the next frame image completes the conversion from 3D data to 2D data, the relative displacement difference can be added to the projection coordinates to correct the position of the projection light. This method of dynamically adapting the projection position is particularly suitable for traffic light perception systems with high processing frequencies, and the higher the processing frequency, the higher the correction accuracy. In a preferred embodiment, the processing frequency of the traffic light perception system is greater than or equal to 10 Hz.

[0133] Step 304: Select a target image from the plurality of second images.

[0134] In this step, a target image is selected as the basis for subsequent traffic light recognition. In actual implementation, an image can be randomly selected from multiple second images as the target image; an image corresponding to a specific camera can also be selected as the target image; or valid images can be first screened from multiple second images and then the target image can be further screened from the valid images. The specific method for selecting the target image is not limited in this embodiment of the application.

[0135] In an optional embodiment, a method of selecting a target image from a plurality of second images may be as follows:

[0136] First, valid images in the second image are screened according to preset conditions;

[0137] The preset conditions include: all projections of at least one set of traffic lights fall within the second image; and the positions of the traffic light projection frames, the corrected projection frames, and the image frames of the traffic lights in the previous second image all meet preset boundary conditions. The preset boundary conditions can be set as follows: the pixel values ​​from the minimum coordinate value to the upper left boundary of the image must all be greater than a preset pixel value, and the pixel values ​​from the maximum coordinate value to the lower right boundary of the image must all be greater than a preset pixel value. The preset pixel value can be set to 10, 15, or 20, for example.

[0138] Secondly, when there are multiple valid images, the valid image corresponding to the front-view camera sensor with the highest priority is selected as the target image from the multiple valid images according to the priority sorting of the front-view camera sensors corresponding to the valid images.

[0139] The priority order can be set to: front-facing telephoto camera > front-facing short-focus camera > right-facing short-focus camera > left-facing short-focus camera.

[0140] Step 305: Generate a region of interest for each traffic light group in the target image.

[0141] Because high-precision maps only provide a relatively accurate projection light frame, its size and position differ significantly from the actual light frame in the image. Therefore, it is necessary to use the position of the projection light bounding box in the image to calculate a larger ROI so that the actual traffic light is included in the ROI. Therefore, the ROI is used to find the precise bounding box of the traffic light. In this embodiment of the application, to more accurately find the traffic light bounding box, an ROI is generated for each traffic light group.

[0142] An optional method of generating a region of interest for each traffic light group in the target image may include the following sub-steps:

[0143] Sub-step 1: for each traffic light group in the target image, calculate the coordinates of the center points of all traffic light reference frames in the traffic light group;

[0144] Among them, the reference frame includes: projection frame and image frame.

[0145] Sub-step 2: Draw a rectangular area with the center point coordinates as the center.

[0146] Among them, the short side of the rectangular area is a preset value, and the long side of the rectangular area is the product of the preset value and the number of traffic lights included in the traffic light group; in a preferred embodiment, the preset value range is set to 256 to 300 pixels, and of course it can also be set to other values ​​such as 400 pixels, 350 pixels, etc.

[0147] Sub-step three: determine the intersection area of ​​the rectangular area and the target image as the region of interest for the traffic light grouping.

[0148] The above sub-steps 1 to 3 are the process of generating a ROI for a traffic light group. In actual implementation, if the target image contains multiple traffic light groups, the process is repeated to generate a corresponding ROI for each group.

[0149] Step 306: Detect the traffic light frame and countdown frame in each area of ​​interest.

[0150] After generating a traffic light ROI to further narrow the traffic light recognition area, traffic light frame and countdown frame detection can be performed within the ROI. In actual implementation, a pre-trained convolutional neural network (CNN) model is used to detect the traffic light frame and countdown frame within each ROI. The CNN model takes as input the target image and information about each ROI. Its output includes traffic light frame information (left, top, right, bottom, and detection confidence) and countdown frame information (left, top, right, bottom, and detection confidence).

[0151] Step 307: Analyze the image areas in the traffic light frame and countdown frame to obtain a traffic light recognition result.

[0152] An optional method of analyzing the image areas in the traffic light frame and countdown frame to obtain a traffic light recognition result may include the following sub-steps:

[0153] Sub-step 1: matching the traffic light frame with the rectified projected light frame of the traffic light in the target image;

[0154] After obtaining the traffic light frame, it needs to be matched one by one with the rectified high-definition map projection light frame. The nearest neighbor matching method can be used, that is, the image position distance between the traffic light frame and the projection light frame is calculated, and the minimum distance is selected to indicate the same traffic light, confirming that the two match.

[0155] Sub-step 2: calculating a second relative displacement difference between the matched projection frame and the traffic light frame;

[0156] The second relative displacement difference is used as a correction value when correcting the projection frame; the second relative displacement difference can be used as a correction value when correcting the traffic light projection frame in the next frame image and generating a region of interest for the traffic light group.

[0157] Sub-step three: determining the position information of the traffic light frame and the countdown frame based on the second relative displacement difference;

[0158] Sub-step 4: classifying the traffic light frame according to the pre-trained neural network model to determine the target type to which the traffic light frame belongs;

[0159] Traffic light frame types include, but are not limited to, horizontal, vertical, and square lights. The pre-trained neural network model further includes models for horizontal, vertical, and square light color recognition, as well as a countdown recognition model. These models can identify the traffic light type, color, and flashing state within the traffic light frame, as well as the countdown value within the countdown frame. The recognition results require further smoothing to ensure more accurate and reliable results.

[0160] Sub-step 5: performing light color smoothing processing on the image area in the traffic light frame and identifying the flashing state of the traffic light to obtain the color and flashing state of the traffic light;

[0161] Sub-step six: performing digital smoothing on the image area in the countdown frame to obtain the countdown number;

[0162] Sub-step seven: Based on the traffic light frame position information and the road structure information contained in the high-definition map, determine the traffic light logo, the distance from the vehicle to the stop line, the distance from the vehicle to the traffic light, the distance from the stop line to the traffic light, and the attributes of each sub-light of the traffic light.

[0163] The traffic light recognition method provided in the embodiment of the present application is executed based on the traffic light perception system. When performing traffic light recognition, it is only necessary to input multiple first images captured by the forward-looking camera, the vehicle's positioning information and the high-definition map to obtain the traffic light logo, traffic light frame position information, traffic light color, flashing status, countdown frame position information, countdown numbers, attributes of each sub-light of the traffic light, the distance from the vehicle to the stop line, the distance from the vehicle to the traffic light, the distance from the stop line to the traffic light and other information.

[0164] The traffic light recognition method provided in the embodiment of the present application obtains a high-definition map, vehicle positioning information, and multiple first images captured by a vehicle's forward-looking camera sensor; groups the traffic lights in the high-definition map based on the vehicle positioning information and a preset traffic light grouping distance threshold; projects the traffic lights in the high-definition map onto each of the first images to generate multiple second images; selects a target image from the multiple second images; generates a region of interest for each traffic light group in the target image; detects the traffic light frame and countdown frame in each region of interest; and analyzes the image areas in the traffic light frame and countdown frame to obtain a traffic light recognition result. The traffic light recognition method provided by the embodiment of the present invention, on the one hand, performs traffic light recognition based on images collected by high-definition maps and vehicle forward-looking camera sensors, which can realize multi-view camera complementarity, make up for the problem that traffic lights cannot be perceived due to the blind spot of the forward-looking camera, and improve the accuracy of traffic light recognition results; on the other hand, traffic lights are grouped based on the spatial position relationship of traffic lights in the high-precision map, and an area of ​​interest is generated for each group, which can improve the ability to detect and recognize traffic lights, and thus improve the accuracy of traffic light recognition results.

[0165] The traffic light recognition method of the embodiment of the present application is described below with a specific example.

[0166] This specific example involves a traffic light recognition method with multi-camera field of view complementarity. In this method, the two side forward-looking cameras of the vehicle are rationally used to compensate for the blind spots of the front-looking camera. The traffic light perception process based on high-precision maps is improved to achieve stable and accurate recognition of traffic lights in autonomous driving.

[0167] In this specific example, Figure 2 The traffic light recognition process is shown in the figure. Figure 9 As shown in Figure 1, it includes four stages: information input, preprocessing, recognition, post-processing, and output of traffic light information. Each stage is described in detail below.

[0168] Phase 1: Information input.

[0169] The input information includes a high-precision map (HDMap), the vehicle's positioning information, and four images captured by four forward-looking cameras (the first images). These four images contain the camera's internal and external parameters. In this example, the camera sensor is referred to as "camera" and the traffic light is referred to as "traffic light."

[0170] High-precision maps use a structured data format and need to include road structured information such as traffic light information, lane information, stop line information, the binding relationship between lane lines and the traffic lights and stop lines at the corresponding traffic intersections, and the topological relationship between lane lines. Traffic light information includes the traffic light unique identifier signal_id, traffic light spatial coordinates (x, y, z), traffic light type, sub-light type, etc., where (x, y) is the longitude and latitude coordinates and z is the height of the traffic light. Stop line information includes stopline_id (stop line unique identifier) ​​and stop line spatial coordinates (x, y, z = 0), where (x, y) is the longitude and latitude coordinates and z is the height of the stop line. Lane line information includes lane_id (lane line unique identifier), the previous lane line lane_id, the next lane line lane_id, the lane line spatial coordinates (x, y, z = 0), and the signal_id and stopline_id of the traffic light to which it is bound.

[0171] The vehicle positioning information is the vehicle positioning information: it combines the inertial sensor and the global positioning system positioning information to provide high-precision positioning information for the vehicle.

[0172] Four front-view images: the front long-focus image is from the H30 camera, the front short-focus image is from the H120 camera, the right front short-focus image is from the H100 camera, and the left front short-focus image is from the H100 camera, as well as the internal and external parameters of each camera.

[0173] The second stage: preprocessing, including four sub-processes: finding traffic lights in the HD map, grouping traffic lights, 3D->2D, and selecting available images.

[0174] Find traffic lights in HD maps:

[0175] Based on the positioning information given by the vehicle and the road structured information such as lane lines, traffic lights, stop lines and their relationships given by the high-precision map, the vehicle can stably query whether there are traffic lights within the preset distance. If a traffic light is found, all the traffic light structure information found is obtained, which can specifically include global position, light type, sub-light type, etc. and distance information. Among them, the distance information includes the distance from the vehicle to the stop line, the distance from the vehicle to the traffic light, and the distance from the stop line to the traffic light. The schematic diagram of the vehicle querying traffic lights is shown in the attached figure. Figure 5 shown.

[0176] In this specific example, since the telephoto camera H30 can sense traffic lights within a range of 300m, the preset distance can be set to 300m.

[0177] Traffic light grouping:

[0178] According to the installation specifications of suspended traffic lights in the real world, the distance between traffic lights is usually less than the width of two lanes, and the width of urban motor vehicle lanes is usually less than 3.5m. Therefore, the traffic light grouping distance threshold distance_threshold can be preset as a value less than 7m, and its setting can be dynamically adjusted according to the number of lane lines and traffic lights on the current road. The traffic light coordinates in the high-definition map correspond to the position information in the real world, and the relative position distance between traffic lights can be calculated through their coordinate positions. As Figure 6 shown, the relative distance between each traffic light can be calculated based on the traffic light coordinate information, denoted as distance_light2light_i_j. Combining the grouping distance threshold, the following formula can be obtained:

[0179] If distance_light2light_i_j≥distance_threshold, they can be grouped;

[0180] If distance_light2light_i_j<distance_threshold, there is no need to group.

[0181] Figure 6 Two schematic diagrams of traffic light intersection scenarios are given, Figure 6 (a) shows that two traffic lights are distributed on both sides of the stop line, Figure 6 (b) shows two common suspended traffic signal lights directly in front of the stop line. According to the position coordinates in the high-precision map, it is calculated that

[0182] {left_distance_light2light_1_2}, {right_distance_light2light_1_2}, if left_distanc e_light2light_1_2>distance_threshold, the two traffic lights are divided into two groups, denoted as group_id_0 and group_id_1; if right_distance_light2light_1_2<distance_threshold, the two traffic lights are divided into the same group, denoted as group_id_0.

[0183] 3D->2D:

[0184] For a traffic light signal_id, its map latitude and longitude coordinates are 3D (x, y, z). The 3D coordinates must first be converted to the geocentric 3D coordinate system (gx, gy, gz), then converted to the ego vehicle coordinate system (lx, ly, lz), and finally converted from the ego vehicle coordinate system to the image coordinate system (2D (ix, iy), where z = gz = lz, representing the height of the traffic light in the real world). The camera's intrinsic and extrinsic parameter matrices are required to convert the ego vehicle coordinate system to the image coordinate system.

[0185] Through coordinate system transformation, the 3D latitude and longitude coordinates of the traffic lights in the map can be projected one by one onto four cameras: the front view long focus H30 camera image, the front view short focus H120 camera image, the right front view short focus H100 camera image, and the left front view short focus H100 camera image. The projection effect can be referred to the attached Figure 7 .

[0186] Affected by camera calibration parameters and global positioning system positioning errors, the position of traffic lights projected into the image is not completely reliable. Especially during driving, it is easily affected by the ground slope and vehicle bumps, and the projection is prone to shifting up, down, left, and right. Therefore, this example proposes a method for dynamically adapting the projection position. That is, after completing the matching of the image frame of the traffic light in the image with the projection frame of the traffic light in the high-precision map, the relative displacement difference (off_x, off_y) between the image frame and the projection frame is calculated. After the next frame of the image completes the 3D->2D transformation, the relative displacement difference is added to the projection coordinates to correct the projection frame position. The schematic diagram before and after the projection correction is shown in the attached figure. Figure 8 shown.

[0187] Selecting an available image means selecting a target image:

[0188] In this sub-process, it is necessary to first filter out valid images, and then further filter out a usable image, namely the target image, from the valid images according to the priority of the camera that captured the image.

[0189] The current image is valid if the following conditions are met:

[0190] All the projection points of at least one set of traffic lights fall into the image;

[0191] The projection frame, correction projection frame, and previous frame detection frame placed on the image should have a pixel value from the minimum coordinate value to the upper left edge of the image greater than 10 pixels, and a pixel value from the maximum coordinate value to the lower right edge of the image greater than 10 pixels.

[0192] The image selection priority is as follows:

[0193] Front-facing telephoto H30 camera > Front-facing short-focus H120 camera > Right-facing short-focus H100 camera > Left-facing short-focus H100 camera

[0194] A single valid image can be selected according to the current camera priority for subsequent traffic light recognition.

[0195] Among them, the right front-view short-focus H100 camera is used to compensate for the blind spot of the right traffic light and the blind spot of the left-turn waiting area, and the left front-view short-focus H100 camera is used to compensate for the blind spot of the left traffic light.

[0196] The third stage: recognition, including processes such as generating ROIs, traffic light detection, matching with map lights, and traffic light classification.

[0197] Generate ROIs:

[0198] Because the HD map only provides a relatively accurate projection light frame, its size and position differ significantly from the actual light frame in the image. Therefore, it is necessary to use the position of the projection light bounding box in the image to calculate a larger region of interest (ROI) so that the actual traffic light is included in the ROI. This ROI is used to find the precise bounding box of the traffic light.

[0199] Based on the number of traffic light groups in the preprocessing process, a ROI is generated for all traffic lights in each group. There are three ways to generate ROIs:

[0200] Method 1: Reference Figure 10 (a) Generate ROIs with the map projection frame as the center. When querying traffic lights from the HD map, since the projection frame has not been processed in any way, ROIs can only be generated using the most basic projection frame.

[0201] Method 2: Reference Figure 10 (b) ROIs are generated using the previous frame's detection frame as the center reference. After continuous image detection, the detection results of the previous frame's traffic light frame can be stably cached. During high-frequency image processing, the position of traffic lights in the previous and subsequent frames changes relatively little. Using the previous frame's detected traffic light frame as the center reference to generate ROIs ensures that the traffic light in the current image is relatively centered, helping to improve the accuracy and recall of the traffic light detection model.

[0202] Method 3: Reference Figure 10 (c) Generate ROIs with the map projection correction frame as the center as a reference. After continuous detection of images, a stable projection correction frame can be obtained. The corrected projection frame is more accurate than the traffic light position in the image compared to method 1. When the traffic light detection fails using method 2, the current method can be used to generate ROIs for the next frame of the image.

[0203] A feasible process for generating ROIs is as follows:

[0204] Step 1: Calculate the center point position (cx, xy) of the reference frames of all traffic lights in the same group; the reference frames include: projection frame and image frame.

[0205] Step 2: Draw a rectangular area with (cx, xy) as the center, the short side of which is the preset width (usually 256 to 300 pixels), and the long side of which is the number of traffic lights multiplied by the preset width;

[0206] Step 3: The intersection of the rectangular area in step 2 and the image area is taken, and the intersection area is the ROI;

[0207] If the image contains multiple traffic light groups, repeat steps 1-3 to generate ROIs for each of the multiple traffic light groups.

[0208] Traffic light detection:

[0209] Traffic light target detection uses a CNN network model to detect traffic light frames and countdown frames within ROIs. Its input is a valid image and ROIs, and its output includes light frame information (left, top, right, bottom, detection confidence) and countdown frame information (left, top, right, bottom, detection confidence).

[0210] Matching with map lights:

[0211] After obtaining the traffic light detection frame target, it needs to be matched one by one with the map projection frame. The matching method usually adopts the nearest neighbor method, that is, the image position distance between the detection light frame and the projection light frame is calculated, and the minimum distance is selected to identify the same signal_id.

[0212] After completing the matching of the detection light and the map light, the relative displacement difference (off_x, off_y) between the target detection light frame and the projection light frame can be calculated for use in correcting the projection frame and generating ROIs in the next frame.

[0213] Traffic light classification:

[0214] As attached Figure 11As shown in the traffic light classification process diagram in , traffic light classification uses a CNN network model, which includes four classification and recognition models: horizontal light color recognition model, vertical light color recognition model, square light color recognition model, and countdown recognition model. Its input is a valid image, a detected traffic light frame, and a detected countdown frame. Its output results are the light colors red, yellow, green, black, and a confidence score. The countdown includes the actual corresponding numbers 0-99 and the confidence score.

[0215] Traffic light frame and countdown frame diagram as shown Figure 12 As shown, Figure 12 (a) is a schematic diagram of a horizontal light. Figure 12 (b) is a schematic diagram of a vertical lamp. Figure 12 (c) is a schematic diagram of a square lamp.

[0216] The fourth stage: post-processing, including light color smoothing, digital smoothing, flickering state recognition and other processes.

[0217] Light color smoothing: By caching the color status of multiple frames of signal_id, the current color status is smoothed. Usually, the color status of 10 frames of signal_id is cached. When 1 / 2 of the colors in the cache are the same as the colors of the signal_id corresponding to the current frame, the current color status is output. Otherwise, the previous color status in the cache is output.

[0218] Digital smoothing: By caching the countdown digital states of multiple frames of signal_id, the current digital state is smoothed. Usually, the digital states of 10 frames of signal_id are cached. When 1 / 2 of the numbers in the cache are the same as the numbers of the signal_id corresponding to the current frame, the current digital state is output. Otherwise, the previous countdown digital state in the cache is output.

[0219] Flashing state recognition: 10 frames of smoothed light color state information are cached to determine whether there is a red light flashing state (black-red, red-black), a short yellow light flashing state (black-yellow, yellow-black within 3 seconds), a long yellow light flashing state (black-yellow, yellow-black for more than 3 seconds), a green light flashing state (black-green, green-black), or a long black light state (continuous black light for more than 5 seconds).

[0220] It should be noted that the number of cached image frames can be flexibly set by those skilled in the art and is not limited to 10 frames, but can also be 8 frames, 12 frames, etc.

[0221] Output traffic light information:

[0222] like Figure 13As shown in the schematic diagram of the input and output information of the traffic light perception system, after the image information, i.e., four images, the positioning information, i.e., the vehicle positioning information, and the high-precision map are input into the traffic light perception system, the traffic light perception system ultimately outputs information including the traffic light signal_id, light frame position information, light color, countdown frame position information, countdown, sub-light attributes, flashing status, the distance from the vehicle to the stop line, the distance from the vehicle to the traffic light, the distance from the stop line to the traffic light, etc.

[0223] The traffic light recognition method provided in this specific example, on the one hand, recognizes traffic lights based on images collected by high-definition maps and vehicle front-view camera sensors, which can achieve multi-view camera complementarity, make up for the problem that traffic lights cannot be perceived due to the blind spot of the front-view camera, and improve the accuracy of traffic light recognition results; on the other hand, traffic lights are grouped based on the spatial position relationship of traffic lights in the high-precision map, and regions of interest are generated for each group, which can improve the ability to detect and recognize traffic lights, and thus improve the accuracy of traffic light recognition results; on the other hand, a method of dynamically adapting the projection position is provided, which can improve the accuracy of matching the image frame of the traffic light with the projection light frame of the traffic light in the high-precision map, so that the traffic lights are included in the finally determined regions of interest.

[0224] Figure 14 A structural block diagram of a traffic light recognition device for implementing an embodiment of the present application.

[0225] The traffic light recognition device provided in the embodiment of the present application includes the following functional modules:

[0226] An acquisition module 801 is configured to acquire a high-definition map, vehicle positioning information, and a plurality of first images captured by a vehicle front-view camera sensor; wherein the high-definition map includes road structure information and traffic light information;

[0227] A grouping module 802 is configured to group the traffic lights in the high-definition map according to the vehicle positioning information and a preset traffic light grouping distance threshold;

[0228] A projection module 803 is configured to project the traffic lights in the high-definition map onto each of the first images to generate a plurality of second images;

[0229] A screening module 804 is configured to screen out a target image from the plurality of second images;

[0230] A generating module 805 is configured to generate a region of interest for each traffic light group in the target image;

[0231] A detection module 806 is configured to detect the traffic light frame and countdown frame within each of the regions of interest;

[0232] The result recognition module 807 is configured to analyze the image areas in the traffic light frame and countdown frame to obtain a traffic light recognition result.

[0233] Optionally, the grouping module includes:

[0234] a traffic light detection submodule, configured to detect traffic lights within a preset distance of the vehicle based on the vehicle positioning information, the road structure information contained in the high-definition map, and the traffic light signal information;

[0235] a query submodule, configured to query, for each of the traffic lights, distance information and structure information corresponding to the traffic light from the traffic light information;

[0236] A distance determination submodule is used to determine the relative distance between every two traffic lights;

[0237] The division submodule is used to divide traffic lights whose relative distance is less than a preset traffic light grouping distance threshold into the same traffic light group.

[0238] Optionally, the projection module includes:

[0239] A first conversion submodule is configured to convert, for one of the traffic lights, the longitude and latitude coordinates of the traffic light in the high-definition map into 3D coordinates in a geocentric coordinate system;

[0240] A second conversion submodule is used to convert the 3D coordinates in the geocentric coordinate system into the 3D coordinates in the vehicle coordinate system;

[0241] The third conversion submodule is used to convert the 3D coordinates in the vehicle coordinate system into 2D coordinates in the image coordinate system based on the intrinsic parameter matrix and extrinsic parameter matrix of the forward-looking camera sensor that captured the first image, thereby forming a projection light frame of the traffic light and generating a second image.

[0242] Optionally, the projection module further includes:

[0243] a first calculation submodule, configured to calculate, for each traffic light projected into the first image, a first relative displacement difference between a projection light frame corresponding to the traffic light and an image frame captured in the first image; wherein the first relative displacement difference serves as a correction value of the image coordinate system;

[0244] An alignment submodule is configured to align the projection light frame corresponding to the traffic light with the image according to the first relative displacement difference.

[0245] Optionally, the screening module includes:

[0246] A valid image screening submodule, configured to screen valid images in the second image according to preset conditions;

[0247] The selection submodule is used to select the valid image corresponding to the forward-looking camera sensor with the highest priority as the target image from the multiple valid images according to the priority sorting of the forward-looking camera sensors corresponding to the valid images when there are multiple valid images.

[0248] Among them, the preset conditions include: all projections of at least one group of traffic lights fall into the second image; the positions of the projection light frame of the traffic signal falling into the second image, the corrected projection light frame, and the image frame of the traffic signal in the previous frame of the second image all meet the preset boundary conditions.

[0249] Optionally, the generating module includes:

[0250] A coordinate calculation submodule is configured to calculate, for each traffic light group in the target image, the coordinates of the center points of all traffic light reference frames within the traffic light group; wherein the reference frames include: a projection frame and an image frame;

[0251] an area drawing submodule, configured to draw a rectangular area with the center point coordinate as the center, wherein the short side of the rectangular area is a preset value, and the long side of the rectangular area is the product of the preset value and the number of traffic lights included in the traffic light group;

[0252] The region of interest determination submodule is configured to determine an intersection area between the rectangular area and the target image as the region of interest of the traffic light group.

[0253] Optionally, the result identification module includes:

[0254] a matching submodule, configured to match the light frame of the traffic light with the rectified projected light frame of the traffic light in the target image;

[0255] A second calculation submodule is configured to calculate a second relative displacement difference between the matched projection frame and the traffic light frame; wherein the second relative displacement difference is used as a correction value when correcting the projection frame;

[0256] a position information determining submodule, configured to determine the position information of the traffic light frame and the countdown frame according to the second relative displacement difference;

[0257] A type determination submodule is used to classify the traffic light frame based on a pre-trained neural network model to determine the target type of the traffic light frame;

[0258] A processing submodule is used to perform light color smoothing processing on the image area in the traffic light frame and identify the flashing state of the traffic light to obtain the color and flashing state of the traffic light;

[0259] A digital processing submodule is used to perform digital smoothing on the image area in the countdown frame to obtain the countdown number;

[0260] The information determination submodule is used to determine the traffic light logo, the distance from the vehicle to the stop line, the distance from the vehicle to the traffic light, the distance from the stop line to the traffic light, and the attributes of each sub-light of the traffic light based on the position information of the traffic light frame and the road structure information contained in the high-definition map.

[0261] The traffic light recognition device provided in the embodiment of the present application, on the one hand, performs traffic light recognition based on images collected by a high-definition map and a vehicle's forward-looking camera sensor, thereby realizing multi-view camera complementarity, compensating for the problem of being unable to perceive traffic lights due to the blind spot of the forward-looking camera, and improving the accuracy of traffic light recognition results; on the other hand, traffic lights are grouped based on their spatial positional relationships in a high-precision map, and regions of interest are generated for each group, thereby improving the ability to detect and recognize traffic lights, and thereby improving the accuracy of traffic light recognition results.

[0262] In the embodiment of the present application Figure 14 The traffic light recognition device shown can be installed in a mobile device or a server. The mobile device or server equipped with the device can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0263] The embodiments of the present application provide Figure 14 The traffic light recognition device shown can realize Figure 3 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0264] Optionally, refer to Figure 15 It is shown that an embodiment of the present application also provides an electronic device 900, including a processor 901, a memory 902, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the various processes executed by the above-mentioned traffic light recognition device are implemented and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0265] It should be noted that the electronic device in the embodiment of the present application includes the server described above.

[0266] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0267] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0268] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A traffic light recognition method, characterized in that: include: Obtaining a high-definition map, vehicle positioning information, and multiple first images captured by a vehicle forward-looking camera sensor; wherein the high-definition map includes road structure information and traffic light information; Grouping the traffic lights in the high-definition map according to the vehicle positioning information and a preset traffic light grouping distance threshold; Projecting the traffic lights in the high-definition map onto each of the first images to generate a plurality of second images; Selecting a target image from the plurality of second images; generating a region of interest for each traffic light group in the target image; Detecting the traffic light frame and countdown frame within each of the regions of interest; The image areas in the traffic light frame and countdown frame are analyzed to obtain a traffic light recognition result.

2. The method according to claim 1, characterized in that The step of grouping the traffic lights in the high-definition map according to the vehicle positioning information and a preset traffic light grouping distance threshold comprises: Detecting traffic lights within a preset distance of the vehicle based on the vehicle positioning information, road structure information contained in the high-definition map, and traffic light signal information; For each of the traffic lights, querying the distance information and structure information corresponding to the traffic light from the traffic light information; Determine the relative distance between every two traffic lights; Traffic lights whose relative distance is less than a preset traffic light grouping distance threshold are grouped into the same traffic light group.

3. The method according to claim 1, characterized in that The step of projecting the traffic lights in the high-definition map onto each of the first images to generate a plurality of second images comprises: For one of the traffic lights, converting the longitude and latitude coordinates of the traffic light in the high-definition map into 3D coordinates in a geocentric coordinate system; Converting the 3D coordinates in the geocentric coordinate system into the 3D coordinates in the vehicle coordinate system; The 3D coordinates in the vehicle coordinate system are converted into 2D coordinates in the image coordinate system based on the intrinsic parameter matrix and the extrinsic parameter matrix of the forward-looking camera sensor that captured the first image, thereby forming a projection light frame of the traffic light and generating a second image.

4. The method according to claim 1, wherein After the step of converting the 3D coordinates in the vehicle coordinate system into 2D coordinates in the image coordinate system based on the intrinsic parameter matrix and the extrinsic parameter matrix of the forward-looking camera sensor that captured the first image to form a projection light frame of the traffic light, the method further includes: For each traffic light projected into the first image, calculating a first relative displacement difference between a projection light frame corresponding to the traffic light and an image frame captured in the first image; wherein the first relative displacement difference serves as a correction value of the image coordinate system; The projection light frame corresponding to the traffic light is aligned with the image according to the first relative displacement difference.

5. The method according to claim 1, wherein The step of selecting a target image from the plurality of second images comprises: screening valid images in the second image according to preset conditions; In the case where there are multiple valid images, the valid image corresponding to the forward-looking camera sensor with the highest priority is selected from the multiple valid images as the target image according to the priority ranking of the forward-looking camera sensors corresponding to the valid images. Among them, the preset conditions include: all projections of at least one group of traffic lights fall into the second image; the positions of the projection light frame of the traffic signal falling into the second image, the corrected projection light frame, and the image frame of the traffic signal in the previous frame of the second image all meet the preset boundary conditions.

6. The method according to claim 1, characterized in that The step of generating a region of interest for each traffic light group in the target image comprises: For each traffic light group in the target image, calculating the center point coordinates of all traffic light reference frames within the traffic light group; wherein the reference frames include: a projection frame and an image frame; Draw a rectangular area with the center point coordinate as the center, wherein the short side of the rectangular area is a preset value, and the long side of the rectangular area is the product of the preset value and the number of traffic lights included in the traffic light group; An intersection area between the rectangular area and the target image is determined as a region of interest of the traffic light group.

7. The method according to claim 1, characterized in that The step of analyzing the image areas in the traffic light frame and countdown frame to obtain a traffic light recognition result includes: matching the traffic light frame with the rectified projected light frame of the traffic light in the target image; Calculating a second relative displacement difference between the matched projection frame and the traffic light frame; wherein the second relative displacement difference is used as a correction value when correcting the projection frame; determining the position information of the traffic light frame and the countdown frame according to the second relative displacement difference; Classify the traffic light frame based on the pre-trained neural network model to determine the target type to which the traffic light frame belongs; Perform light color smoothing on the image area in the traffic light frame and identify the flashing state of the traffic light to obtain the color and flashing state of the traffic light; Perform digital smoothing on the image area in the countdown frame to obtain the countdown number; The traffic light identification, the distance from the vehicle to the stop line, the distance from the vehicle to the traffic light, the distance from the stop line to the traffic light, and the attributes of each sub-light of the traffic light are determined based on the position information of the traffic light frame and the road structure information contained in the high-definition map.

8. A traffic light recognition device, characterized in that: include: An acquisition module, configured to acquire a high-definition map, vehicle positioning information, and a plurality of first images captured by a vehicle front-view camera sensor; wherein the high-definition map includes road structure information and traffic light information; a grouping module, configured to group the traffic lights in the high-definition map according to the vehicle positioning information and a preset traffic light grouping distance threshold; a projection module, configured to project the traffic lights in the high-definition map onto each of the first images to generate a plurality of second images; a screening module, configured to screen out a target image from a plurality of second images; a generating module, configured to generate a region of interest for each traffic light group in the target image; A detection module, configured to detect a traffic light frame and a countdown frame within each of the regions of interest; The result recognition module is used to analyze the image areas in the traffic light frame and countdown frame to obtain a traffic light recognition result.

9. An electronic device, characterized in that: The electronic device includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. The program or instruction is executed by the processor to execute the steps of any one of the traffic light recognition methods in claims 1-7.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of any one of the traffic signal light recognition methods in claims 1-7 are implemented.