Signal lamp positioning tracking method and device, electronic equipment and storage medium

By combining the traffic light features from multiple image acquisition devices and historical databases, and utilizing algorithms such as 3D projection constraint matching, the problem of recognition accuracy and reliability caused by changes in traffic light position was solved, achieving efficient positioning and tracking of traffic lights and improving the accuracy of traffic light recognition during vehicle driving.

CN122135337APending Publication Date: 2026-06-02DONGFENG MOTOR GRP

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DONGFENG MOTOR GRP
Filing Date
2026-01-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, traffic light recognition methods based on two-dimensional image analysis are difficult to accurately determine the spatial position of traffic lights, resulting in poor accuracy and reliability of the traffic light positions identified by vehicles during driving. Furthermore, three-dimensional reconstruction technology has high computational complexity and poor real-time performance, making it unable to adapt to changes in traffic light positions.

Method used

By acquiring traffic light images from multiple image acquisition devices on the target vehicle, image perception processing is performed to determine the current two-dimensional features of the traffic lights. Then, using three-dimensional projection constraint matching algorithms, epipolar geometry constraint algorithms, and two-dimensional spatial constraint algorithms, combined with traffic light features from the historical database, matching and projection processing is performed to determine the current three-dimensional position features of the traffic lights.

Benefits of technology

It enables the tracking of traffic lights that change position, improving the accuracy and reliability of traffic light position recognition during vehicle operation, reducing reliance on high-precision maps, and reducing detection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135337A_ABST
    Figure CN122135337A_ABST
Patent Text Reader

Abstract

The application discloses a positioning tracking method and device of a signal lamp, electronic equipment and a storage medium, and relates to the technical field of vehicle detection. The method comprises the following steps: performing image sensing processing on a signal lamp image of a current frame to obtain current two-dimensional features of at least one target signal lamp; determining matching point pairs of the target signal lamps based on current two-dimensional position features in the current two-dimensional features and historical position features of each historical signal lamp; performing projection processing on the matching point pairs based on the current two-dimensional position features of the target signal lamps and poses of a plurality of image acquisition devices to determine current three-dimensional position features of the target signal lamps; and determining current three-dimensional features of the target signal lamps based on current shape features, current state features and the current three-dimensional position features of the target signal lamps. The application can track the target signal lamps with changed positions, and improves the accuracy and reliability of recognizing the positions of the signal lamps in the driving process of the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle detection technology, and in particular to a method, device, electronic device and storage medium for locating and tracking traffic lights. Background Technology

[0002] Currently, with the rapid development of autonomous driving technology, accurately identifying and understanding traffic light information has become crucial to ensuring driving safety. Most mainstream traffic light recognition methods are based on analyzing two-dimensional images captured by cameras.

[0003] However, while two-dimensional image analysis is simple and easy to use, it is difficult to accurately determine the spatial position of traffic lights based solely on two-dimensional images, which can easily lead to errors in distance judgment. In recent years, some advanced solutions have attempted to introduce three-dimensional reconstruction technology to improve recognition accuracy through multi-angle observation and sensor fusion. Compared with the recognition accuracy of two-dimensional images, three-dimensional reconstruction technology can better restore the three-dimensional spatial information of traffic lights. However, directly using three-dimensional reconstruction technology to locate traffic lights requires high computational complexity, has poor real-time performance, and cannot adapt to changes in traffic light position, meaning it cannot track traffic lights that have changed positions. Consequently, the accuracy and reliability of the traffic light positions identified by vehicles during driving remain poor. Summary of the Invention

[0004] This application provides a traffic light-based positioning and tracking method, device, electronic device, and storage medium. The embodiments provided in this application solve the technical problems of existing technologies, such as high computational complexity, poor real-time performance, and inability to adapt to changes in traffic light positions, i.e., the inability to track traffic lights in changing positions, resulting in poor accuracy and reliability of traffic light positions identified by vehicles during driving. The embodiments provided in this application can track target traffic lights in changing positions, improving the accuracy and reliability of traffic light positions identified by vehicles during driving.

[0005] In a first aspect, this application provides a method for locating and tracking traffic lights, the method comprising: Acquire the traffic light images of the current frame captured by multiple image acquisition devices on the target vehicle; Image perception processing is performed on the traffic light image of the current frame to obtain the current two-dimensional features of at least one target traffic light; wherein, the current two-dimensional features include current two-dimensional position features, current shape features, and current state features; Based on the current two-dimensional position features of the at least one target traffic light and the historical position features of each historical traffic light in the historical database, the at least one target traffic light is matched with each of the historical traffic lights to determine the matching point pairs of each target traffic light; Based on the current two-dimensional position features of the target traffic light and the poses of the multiple image acquisition devices, the matching point pairs are projected to determine the current three-dimensional position features of the target traffic light. Based on the current shape features, the current state features, and the current three-dimensional position features of the target traffic light, the current three-dimensional features of the target traffic light are determined to complete the positioning and tracking of the target traffic light.

[0006] In one feasible implementation, the historical location features of each historical traffic light include historical three-dimensional location features. The step of matching the at least one target traffic light with each of the historical traffic lights based on the current two-dimensional features of the at least one target traffic light and the historical two-dimensional features of each historical traffic light in the historical database to determine the matching point pairs of each target traffic light in the current frame includes: The historical three-dimensional location features of each historical traffic light are obtained from the historical database; wherein, the historical database includes a historical frame database and a historical map database; Based on the three-dimensional projection constraint matching algorithm, the historical three-dimensional position features of each historical traffic light are projected onto the plane of the traffic light image in the current frame to determine the three-dimensional matching score between the at least one target traffic light and each historical traffic light. Based on the three-dimensional matching score and the preset matching algorithm, at least one target traffic light is matched with each of the historical traffic lights to determine the matching point pairs of each target traffic light.

[0007] In one feasible implementation, the three-dimensional projection constraint matching algorithm projects the historical three-dimensional position features of each historical traffic light onto the plane of the traffic light image in the current frame, and determines the three-dimensional matching score between the at least one target traffic light and each historical traffic light, including: Based on the three-dimensional projection constraint matching algorithm, the historical three-dimensional position features, and the camera parameters corresponding to the image acquisition device, the projection position of each historical traffic light on the plane of the traffic light image in the current frame is determined. Determine the distance difference between the projected position and the center point of the two-dimensional position feature of each of the target traffic lights; Based on the distance difference, a three-dimensional matching score is determined between the at least one target traffic light and each of the historical traffic lights, wherein the smaller the distance difference, the higher the matching score.

[0008] In one feasible implementation, the historical location features of each historical traffic light include historical two-dimensional location features. The step of matching the at least one target traffic light with each of the historical traffic lights based on the current two-dimensional features of the at least one target traffic light and the historical two-dimensional features of each historical traffic light in the historical database, to determine the matching point pairs of each target traffic light in the current frame, includes: The historical two-dimensional location features of each historical traffic light are obtained from the historical database; wherein, the historical database includes a historical frame database and a historical map database; Based on the epipolar geometry constraint algorithm, the current two-dimensional features of the at least one target traffic light, and the historical two-dimensional features of each historical traffic light, the epipolar distance from each of the historical traffic lights to the at least one target traffic light is determined. Based on the epipolar distance and the preset distance threshold, the at least one target traffic light is matched with each of the historical traffic lights to determine the matching point pairs of each of the target traffic lights in the current frame.

[0009] In one feasible implementation, determining the epipolar distance from each historical traffic light to at least one target traffic light based on the epipolar geometry constraint algorithm, the current two-dimensional features of the at least one target traffic light, and the historical two-dimensional position features of each historical traffic light includes: The relative motion data of the target vehicle is determined based on the camera parameters of multiple image acquisition devices; Based on the relative motion data, the camera parameters corresponding to the image acquisition device, and the epipolar geometry constraint algorithm, the fundamental matrix of the current frame is determined. Based on the base matrix, each of the historical two-dimensional position features, and the current position two-dimensional features of at least one target traffic light, the epipolar line of the current frame is determined. The distance from the two-dimensional position center point of at least one target signal light in the current frame to the polar line is determined as the polar distance.

[0010] In one feasible implementation, the step of matching the at least one target traffic light with each of the historical traffic lights based on the current two-dimensional features of the at least one target traffic light and the historical two-dimensional features of each historical traffic light in the historical database, to determine the matching point pairs of each of the target traffic lights in the current frame, includes: Based on the two-dimensional spatial constraint algorithm and the epipolar distance, the two-dimensional matching score of the historical position features of each historical traffic light in the traffic light image of the current frame is determined; Based on the two-dimensional matching score and the preset matching algorithm, the matching point pairs of each target traffic light in the current frame are determined.

[0011] In one feasible implementation, the preset matching algorithm includes a weighted bipartite graph matching algorithm, and the step of determining the matching point pairs of each target traffic light in the current frame based on the two-dimensional matching score and the preset matching algorithm includes: Determine the two-dimensional matching score matrix based on the two-dimensional matching score; Using the current set of traffic lights as the first vertex set, the historical set of traffic lights as the second vertex set, and the two-dimensional matching score matrix as the edge weights, a weighted bipartite graph matching algorithm is executed to determine the optimal matching pair. The optimal matching pair is determined as the matching point pair of each target traffic light in the current frame.

[0012] In one feasible implementation, the step of projecting the matching point pair based on the current two-dimensional position features of the target traffic light and the poses of the plurality of image acquisition devices to determine the current three-dimensional position features of the target traffic light includes: Based on a preset outlier filtering algorithm, the matching point pairs are filtered to determine candidate matching point pairs; Based on the poses of multiple image acquisition devices associated with the candidate matching point pairs and the current two-dimensional position features of the target traffic light under the multiple poses, a system of projection equations is constructed. Based on the candidate matching point pairs and the least squares method, the projection equations are calculated to determine the current three-dimensional position features of the target traffic light.

[0013] In one feasible implementation, determining the current three-dimensional features of the target traffic light based on the current shape features, the current state features, and the current three-dimensional position features of the target traffic light includes: Obtain the current shape features and current state features of each target traffic light in multiple consecutive frames within a preset time period; Determine the shape frequency of the target traffic light in each shape, and the state frequency of the target traffic light in each state; Based on a preset voting algorithm, the current shape feature and the current state feature are updated respectively. The current shape feature whose shape frequency exceeds a preset shape frequency threshold is determined as the target shape feature, and the current state feature whose state frequency exceeds a preset state frequency threshold is determined as the target state feature. Based on the target shape features, the target state features, and the current three-dimensional position features of the target traffic light, the current three-dimensional features of the target traffic light are determined.

[0014] In a second aspect, this application provides a positioning and tracking device for a traffic light, the traffic light positioning and tracking device comprising: The acquisition module is used to acquire the traffic light image of the current frame captured by multiple image acquisition devices on the target vehicle; The processing module is used to perform image perception processing on the traffic light image of the current frame to obtain the current two-dimensional features of at least one target traffic light; wherein, the current two-dimensional features include current two-dimensional position features, current shape features, and current state features; The first determining module is used to match the at least one target traffic light with each of the historical traffic lights based on the current two-dimensional position features of the at least one target traffic light and the historical position features of each historical traffic light in the historical database, and to determine the matching point pairs of each of the target traffic lights. The second determining module is used to perform projection processing on the matching point pair based on the current two-dimensional position features of the target traffic light and the poses of the multiple image acquisition devices to determine the current three-dimensional position features of the target traffic light. The third determining module is used to determine the current three-dimensional features of the target traffic light based on the current shape features, the current state features, and the current three-dimensional position features of the target traffic light, so as to complete the positioning and tracking of the target traffic light.

[0015] The traffic light positioning and tracking method, apparatus, electronic device, and storage medium provided in this application, compared with the prior art, can acquire traffic light images of the current frame collected by multiple image acquisition devices on a target vehicle, and perform image perception processing on the traffic light images of the current frame to obtain the current two-dimensional features of at least one target traffic light. Then, based on the current two-dimensional position features of at least one target traffic light and the historical position features of each historical traffic light in the historical database, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light. Finally, based on the current two-dimensional position features of the target traffic light... By analyzing the location features and poses of multiple image acquisition devices, the matching point pairs are projected to determine the current three-dimensional location features of the target traffic light. Finally, based on the current shape features, current state features, and the current three-dimensional location features of the target traffic light, the current three-dimensional features of the target traffic light are determined to complete the positioning and tracking of the target traffic light. This application enables the tracking of target traffic lights with changing positions, improves the accuracy and reliability of traffic light position recognition during vehicle operation, and is more conducive to the accuracy of subsequent path planning. It achieves stable tracking of target red and green traffic lights, reduces the dependence of the vehicle system on high-precision maps, and reduces detection costs. Attached Figure Description

[0016] Figure 1 A flowchart of a traffic light positioning and tracking method provided in an embodiment of this application is shown; Figure 2 This paper shows a structural block diagram of a traffic light positioning and tracking device provided in an embodiment of this application; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown.

[0017] Figure 2 and Figure 3 The correspondence between the figure labels and figure titles in the accompanying drawings is as follows: 200 Positioning and tracking device for traffic lights; 210 Acquisition module; 220 Processing module; 230 First determination module; 240 Second determination module; 250 Third determination module; 300 Electronic device; 310 Processor; 320 Memory; 330 Bus. Detailed Implementation

[0018] To better understand the technical solutions provided in the embodiments of this specification, the technical solutions of the embodiments of this specification will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of this specification and the specific features in the embodiments are detailed descriptions of the technical solutions of the embodiments of this specification, rather than limitations on the technical solutions of this specification. In the absence of conflict, the embodiments of this specification and the technical features in the embodiments can be combined with each other.

[0019] In this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element. The term "two or more" includes two or more cases.

[0020] First, the applicable application scenarios of this application will be introduced. The embodiments provided in this application are applicable to the field of vehicle detection technology, and in particular relate to a method, device, electronic device and storage medium for locating and tracking traffic lights.

[0021] Currently, while two-dimensional image analysis is simple and easy to use, it is difficult to accurately determine the spatial position of traffic lights based solely on two-dimensional images, which can easily lead to errors in distance judgment. In recent years, some advanced solutions have attempted to introduce three-dimensional reconstruction technology to improve recognition accuracy through multi-angle observation and sensor fusion. Compared with the recognition accuracy of two-dimensional images, three-dimensional reconstruction technology can better restore the three-dimensional spatial information of traffic lights. However, directly using three-dimensional reconstruction technology to locate traffic lights requires high computational complexity, has poor real-time performance, and cannot adapt to changes in traffic light position, meaning it cannot track traffic lights that have changed positions. Consequently, the accuracy and reliability of the traffic light positions identified by vehicles during driving remain poor.

[0022] Based on this, embodiments of this application provide a method, apparatus, electronic device, and storage medium for locating and tracking traffic lights. The embodiments provided by this application solve the technical problems of existing technologies, such as high computational complexity requirements, poor real-time performance, and inability to adapt to changes in traffic light positions, i.e., the inability to track traffic lights with changing positions, resulting in poor accuracy and reliability of traffic light positions identified by vehicles during driving. The embodiments provided by this application can track target traffic lights with changing positions, improving the accuracy and reliability of traffic light positions identified by vehicles during driving.

[0023] Figure 1 A flowchart illustrating a traffic light positioning and tracking method provided in an embodiment of this application is shown. Figure 1 As shown, the traffic light positioning and tracking method includes the following steps: S101. Obtain the signal light image of the current frame captured by multiple image acquisition devices on the target vehicle.

[0024] In this step, the implementation provided in this application acquires the traffic light image of the current frame through multiple image acquisition devices during the driving of the target vehicle. The angle of the traffic light image of the current frame acquired by different types of image acquisition devices is different.

[0025] It is understood that the specific model settings of the multiple image acquisition devices provided in the embodiments of this application can be customized and used according to different application scenarios and usage conditions. The multiple image acquisition devices in the embodiments provided in this application can be specifically set as narrow-view image acquisition devices and telephoto image acquisition devices.

[0026] In the above, the traffic light images of the current frame captured by different image acquisition devices need to be marked with different flag bits to distinguish them. The embodiment provided in this application uses a flag bit 1 to represent the traffic light image of the current frame captured by the narrow-view image acquisition device; and uses a flag bit 2 to represent the traffic light image of the current frame captured by the telephoto image acquisition device.

[0027] It should be noted that the traffic lights in the embodiments provided in this application can be customized and used according to different application scenarios and usage conditions. The traffic lights in the embodiments provided in this application can be specifically set as red lights, green lights, and yellow lights.

[0028] S102. Perform image perception processing on the traffic light image of the current frame to obtain the current two-dimensional features of at least one target traffic light; wherein, the current two-dimensional features include the current two-dimensional position features, the current shape features, and the current state features.

[0029] Among them, the two-dimensional position feature refers to the position coordinates of the target traffic light in the plane in the camera coordinate system; the current shape feature is used to characterize the shape of the target traffic light, such as a left-turning arrow or a right-turning arrow; and the current state feature is used to characterize the display state of the target traffic light, such as a constantly lit state or a non-flashing state.

[0030] In this step, after acquiring the traffic light image of the current frame, the embodiment provided in this application will perform image perception processing on the traffic light image of the current frame, identify the target traffic light features in the traffic light image, and use the target traffic light features as the current two-dimensional features.

[0031] In the above-mentioned embodiments, the image perception processing method provided in this application may be specific but not limited to: using a trained traffic light feature detection model for image perception processing.

[0032] S103. Based on the current two-dimensional position features of at least one target traffic light and the historical position features of each historical traffic light in the historical database, match at least one target traffic light with each historical traffic light to determine the matching point pairs of each target traffic light.

[0033] In this step, the embodiments provided in this application construct different types of constraints on the current two-dimensional position features of at least one target traffic light and the historical position features of each historical traffic light in the historical database, so as to match at least one target traffic light with each historical traffic light, determine the matching point pairs of each target traffic light, and different matching point pairs may come from the same traffic light captured by image acquisition devices at different locations.

[0034] It is understood that the methods for determining the matching point pairs of each target traffic light in the embodiments provided in this application include at least one of the following matching methods: a matching method based on a three-dimensional projection constraint matching algorithm, a matching method based on an epipolar geometric constraint algorithm, and a matching method based on a two-dimensional spatial constraint algorithm.

[0035] It should be noted that the historical location features of historical traffic lights can specifically come from historical frame databases or historical map databases.

[0036] S104. Based on the current two-dimensional position features of the target traffic light and the poses of multiple image acquisition devices, the matching point pairs are projected to determine the current three-dimensional position features of the target traffic light.

[0037] In this step, the embodiments provided in this application, after determining a series of matching point pairs on the tracking and matching of different image acquisition devices, will establish a three-dimensional model through the above matching points. Specifically, the matching point pairs are projected by the current two-dimensional position features of the target traffic light and the poses of multiple image acquisition devices over a period of time, so as to find the current three-dimensional position features of the target traffic light, so that the projection of the matching point under different viewpoints of each image acquisition device is consistent with the observation point.

[0038] Among them, the current three-dimensional position features can be obtained through The current three-dimensional position feature in this application is specifically the homogeneous coordinates of the current two-dimensional position feature projected onto the world coordinate system.

[0039] S105. Based on the current shape features, current state features, and the current three-dimensional position features of the target traffic light, determine the current three-dimensional features of the target traffic light to complete the positioning and tracking of the target traffic light.

[0040] In this step, the embodiments provided in this application, after determining the current three-dimensional position features of the target traffic light, bind and update the aforementioned current three-dimensional position features, current shape features, and current state features.

[0041] It is understood that the reconstruction of the current shape features and current state features in the embodiments provided in this application will adopt a preset voting algorithm to output the current shape features and current state features with a high and stable proportion within a preset time period.

[0042] Compared with the prior art, the traffic light positioning and tracking method provided in this application can acquire traffic light images of the current frame collected by multiple image acquisition devices on the target vehicle, and perform image perception processing on the traffic light images of the current frame to obtain the current two-dimensional features of at least one target traffic light. Then, based on the current two-dimensional position features of at least one target traffic light and the historical position features of each historical traffic light in the historical database, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light. Next, based on the current two-dimensional position features of the target traffic light and the poses of multiple image acquisition devices, the matching point pairs are projected to determine the current three-dimensional position features of the target traffic light. Finally, based on the current shape features, current state features, and current three-dimensional position features of the target traffic light, the current three-dimensional features of the target traffic light are determined to complete the positioning and tracking of the target traffic light. This application can track target traffic lights with changing positions, improving the accuracy and reliability of traffic light position identification during vehicle driving.

[0043] As an optional implementation in this embodiment, the historical position features of each historical traffic light include historical three-dimensional position features. Based on the current two-dimensional position features of at least one target traffic light and the historical position features of each historical traffic light in the historical database, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light, including: The historical three-dimensional position features of each historical traffic light are obtained from the historical database, which includes a historical frame database and a historical map database. Based on a three-dimensional projection constraint matching algorithm, the historical three-dimensional position features of each historical traffic light are projected onto the plane of the traffic light image in the current frame to determine the three-dimensional matching score between at least one target traffic light and each historical traffic light. Based on the three-dimensional matching score and a preset matching algorithm, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light.

[0044] It should be noted that the embodiments provided in this application can obtain the historical three-dimensional position features of each historical traffic light from the historical frame database, and then, based on the three-dimensional projection constraint matching algorithm of the previous and next frames in the three-dimensional projection constraint matching algorithm, project the historical three-dimensional position features of each historical traffic light onto the plane of the traffic light image of the current frame, determine the three-dimensional matching score between at least one target traffic light and each historical traffic light, and after determining the above three-dimensional matching score, match the three-dimensional matching score with the preset matching algorithm to match at least one target traffic light with each historical traffic light, and determine the matching point pair of each target traffic light.

[0045] Here, the preset matching algorithm in the embodiments provided in this application can be set as a weighted bipartite graph matching algorithm. The method for determining the matching point pairs of each target traffic light based on the weighted bipartite graph matching algorithm and the three-dimensional matching score is as follows: Initialize the first vertex label value of the first vertex set and the second vertex label value of the second vertex set, wherein the first vertex set is the target traffic light set of the current frame and the second vertex set is the historical traffic light set; Based on the two-dimensional matching score matrix, start searching for augmenting paths from the unmatched vertices in the first vertex set, wherein the unmatched vertices are used to represent target traffic lights that do not match the historical traffic lights; If an augmenting path is found, update the two-dimensional matching score matrix; If no augmenting path is found, determine the minimum distance from the first vertex in all unmatched first vertex sets to the second vertex in the second vertex set; Based on the minimum distance, search for augmenting paths again until all first vertices in the first vertex set are successfully matched or the preset number of iterations is reached; Determine the matching point pairs of the target traffic lights by assigning the same tracking ID to the successfully matched target traffic lights.

[0046] In the above, the first vertex set in the embodiments provided in this application is represented by X, and the second vertex set is represented by... Indicate; for each first vertex in X, let ,Right now The value of is the maximum weight of all edges from x to vertices in set Y. For each vertex y in set Y, let .

[0047] It is understood that the embodiments provided in this application will first obtain the historical three-dimensional position features reconstructed from the previous frame by the same source perception from the historical frame database, and then project the historical three-dimensional position features of the previous frame onto the plane of the traffic light image of the current frame to determine the three-dimensional matching score between at least one target traffic light and each historical traffic light.

[0048] It should be noted that the first feature of the historical three-dimensional location features is also determined based on the historical two-dimensional location features of each historical traffic light in the historical frame database or historical map database and the preset triangulation calculation.

[0049] Optionally, based on a 3D projection constraint matching algorithm, the historical 3D position features of each historical traffic light are projected onto the plane of the traffic light image in the current frame to determine the 3D matching score between at least one target traffic light and each historical traffic light, including: Based on the 3D projection constraint matching algorithm, historical 3D position features, and camera parameters corresponding to the image acquisition device, the projected position of each historical traffic light on the plane of the traffic light image in the current frame is determined; the distance difference between the projected position and the center point of the 2D position feature of each target traffic light is determined; based on the distance difference, the 3D matching score between at least one target traffic light and each historical traffic light is determined, wherein the smaller the distance difference, the higher the matching score.

[0050] Here, the image acquisition device in the embodiment provided in this application can be specifically set to two. All the following steps are described using two image acquisition devices as an example, and will not be repeated below.

[0051] It should be noted that the embodiments provided in this application first obtain the camera parameters of the two image acquisition devices, then obtain the coordinate transformation of the camera intrinsic parameter matrix of the current frame from the configuration, and then transform the historical 3D position features from the world coordinate system to the camera coordinate system and pixel system successively through the camera's intrinsic and extrinsic parameters, and then perform normalization processing. The transformation formula is as follows: ; in, The pixel coordinates of the center point used to characterize two-dimensional positional features; Used to characterize depth information in the camera coordinate system, and related to The calculation formula is: ; The embodiments provided in this application, after determining the projection position on the plane of the traffic light image in the current frame, determine the three-dimensional matching score between at least one target traffic light and each historical traffic light by calculating the distance difference between the projection position and the center point of the two-dimensional position features of each target traffic light. The smaller the distance difference, the higher the matching score. The calculation formula is as follows: ; ; in, Used to characterize the matching score; Used to characterize distance difference.

[0052] Optionally, the embodiments provided in this application can also obtain the historical three-dimensional location features of each historical traffic light from a historical map database, and then, based on the historical location features... Figure 3The 3D projection constraint matching algorithm projects historical 3D location features from the historical map database onto the plane of the traffic light image in the current frame, determines the 3D matching score between at least one target traffic light and each historical traffic light, and matches at least one target traffic light with each historical traffic light based on the 3D matching score and a preset matching algorithm to determine the matching point pairs of each target traffic light. The matching method and calculation formula are consistent with the matching method of the 3D projection constraint matching algorithm used in previous and subsequent frames.

[0053] In the above process, when the historical three-dimensional location features of historical traffic lights (i.e., historical red and green lights) are obtained from the historical frame database, short-term tracking matching is performed; when the historical three-dimensional location features of historical traffic lights are obtained from historical map data, map association matching is performed.

[0054] Short-term tracking matching includes: establishing mapping relationships between data from the same source; constructing a tracking context based on the most recent N frames of data; using a weighted bipartite graph algorithm to associate targets between frames; and passing tracking identifiers to successfully matched targets.

[0055] Historical map data association and matching includes: retrieving candidate map elements based on spatial index; performing multi-scale projection and visibility analysis; integrating location, size, and orientation features for comprehensive scoring; and selectively updating map data based on the matching results.

[0056] Here, determining the traffic light features in the historical map requires not only projecting its center point, but also projecting the eight vertices of its three-dimensional bounding box.

[0057] The embodiments provided in this application can determine whether a historical traffic light is visible in the current frame based on its projection position, and remove elements that are completely outside the image boundary.

[0058] The embodiments provided in this application verify whether the scales are consistent by comparing the proportional relationship between the 2D size of the historical traffic light projection and the size of the current frame detection box.

[0059] As an optional implementation in this embodiment, the historical position features of each historical traffic light include historical two-dimensional position features. Based on the current two-dimensional position features of at least one target traffic light and the historical position features of each historical traffic light in the historical database, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light, including: The historical two-dimensional position features of each historical traffic light are obtained from the historical database, which includes a historical frame database and a historical map database. Based on the epipolar geometric constraint algorithm, the current two-dimensional position features of at least one target traffic light, and the historical two-dimensional position features of each historical traffic light, the epipolar distance from each historical traffic light to at least one target traffic light is determined. Based on the epipolar distance and a preset distance threshold, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light in the current frame.

[0060] It should be noted that the embodiments provided in this application use traffic light images of the current frame acquired by different image acquisition devices. By utilizing extreme weighting constraints, the two-dimensional features of the current position of the target traffic light in the current frame are matched with the historical position features of each historical traffic light in the previous frame or historical map database to determine the matching point pairs of each target traffic light in the current frame. The specific matching method is as follows: It is understood that, in the embodiments provided in this application, the historical two-dimensional position features of the historical traffic lights with the same source input in the previous frame can be found from the input type of the previous frame based on the historical frame database. Then, the epipolar distance from each historical traffic light to at least one target traffic light is determined. Then, if the epipolar distance does not exceed a preset distance threshold, at least one target traffic light is matched with each historical traffic light to determine the matching point pair of each target traffic light in the current frame.

[0061] The preset distance threshold in the embodiments provided in this application can be selected and used in a custom manner according to different application scenarios and usage conditions.

[0062] As an optional implementation in this embodiment, the epipolar distance from each historical traffic light to at least one target traffic light is determined based on the epipolar geometry constraint algorithm, the current two-dimensional features of at least one target traffic light, and the historical two-dimensional position features of each historical traffic light, including: The relative motion data of the target vehicle is determined based on the camera parameters of multiple image acquisition devices; the fundamental matrix of the current frame is determined based on the relative motion data, the camera parameters corresponding to the image acquisition devices, and the epipolar geometric constraint algorithm; the epipolar line of the current frame is determined based on the fundamental matrix, each historical two-dimensional position feature, and the current position two-dimensional feature of at least one target traffic light; the distance from the two-dimensional position center point of at least one target traffic light in the current frame to the epipolar line is determined as the epipolar distance.

[0063] Here, the method for determining the fundamental matrix of the current frame based on relative motion data, camera parameters corresponding to the image acquisition device, and epipolar geometry constraint algorithm is as follows: ; in, The fundamental matrix used to characterize the current frame; Used to characterize the inverse of the camera intrinsic parameter matrix; Used to characterize the inverse of the transpose of the camera intrinsic parameter matrix.

[0064] Here, the cross product matrix (Skew-symmetric matrix) of t is defined as follows: The definition is as follows: ; Based on the hierarchical constraint, for corresponding points in the traffic light images of the current frame between two consecutive frames... and 2. Design the fundamental matrix to satisfy the following epipolar constraints: ; Then, the fundamental matrix of the current frame is applied to the two-dimensional position center point of the target traffic light in the traffic light image of the previous frame to obtain the epipolar line. The distance from the two-dimensional position center point of at least one target traffic light in the current frame to the epipolar line is calculated and determined as the epipolar distance. The maximum value of the length and width of the current target traffic light is used as the distance threshold to determine whether the current match meets this requirement. If the matching requirement is met, the same tracking ID is assigned.

[0065] As an optional implementation method in this embodiment, based on the current two-dimensional position features of at least one target traffic light and the historical position features of each historical traffic light in the historical database, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light. This includes: determining the two-dimensional matching score of the historical position features of each historical traffic light in the traffic light image of the current frame based on a two-dimensional spatial constraint algorithm and epipolar distance; and determining the matching point pairs of each target traffic light in the current frame based on the two-dimensional matching score and a preset matching algorithm.

[0066] In the above-described embodiments, after determining the epipolar distance, a two-dimensional spatial constraint algorithm and epipolar distance are used to determine a two-dimensional matching score of the historical position features of each historical traffic light in the traffic light image of the current frame. Then, based on the calculated two-dimensional matching score matrix, a weighted bipartite graph optimal matching algorithm is used to determine the final matching point pairs of each target traffic light in the current frame.

[0067] The preset matching algorithm can be customized and used according to different application scenarios and usage conditions. The preset matching algorithm in the embodiments provided in this application can be specifically set to adopt the weighted bipartite graph optimal matching algorithm.

[0068] As an optional implementation in this embodiment, the preset matching algorithm includes a weighted bipartite graph matching algorithm, which determines the matching point pairs of each target traffic light in the current frame based on the two-dimensional matching score and the preset matching algorithm, including: Determine the two-dimensional matching score matrix based on the two-dimensional matching score; Using the current set of traffic lights as the first vertex set, the historical set of traffic lights as the second vertex set, and the two-dimensional matching score matrix as the edge weight, a weighted bipartite graph matching algorithm is executed to determine the optimal matching pair; the optimal matching pair is determined as the matching point pair of each target traffic light in the current frame.

[0069] Here, the embodiment provided in this application calculates a matching score between each target traffic light in the current frame and each traffic light in the historical traffic light set to obtain multiple matching scores. The multiple matching scores are arranged in the order of the target traffic lights in the current frame and the order of the historical traffic lights to form a two-dimensional matrix, where the rows of the matrix correspond to the target traffic lights in the current frame, the columns correspond to the historical traffic lights, and the matrix elements are the corresponding matching scores. Based on the matching scores of the matrix elements and the matching results, the tracking identifier associated with the historical observation is assigned to the corresponding current target traffic light.

[0070] As an optional implementation method in this embodiment, based on the current two-dimensional position features of the target traffic light and the poses of multiple image acquisition devices, the matching point pairs are projected to determine the current three-dimensional position features of the target traffic light, including: Based on a preset outlier filtering algorithm, matching point pairs are screened to determine candidate matching point pairs. Based on the poses of multiple image acquisition devices related to the candidate matching point pairs and the current two-dimensional position features of the target traffic light under multiple poses, a system of projection equations is constructed. Based on the candidate matching point pairs and the least squares method, the system of projection equations is calculated to determine the current three-dimensional position features of the target traffic light.

[0071] It should be noted that the embodiments provided in this application will use a random sample consensus algorithm to filter and screen outliers in the determined matching point pairs to improve the stability of the current three-dimensional position feature prediction of the target traffic light and prevent incorrect matching points from affecting the results.

[0072] Understandably, this application provides a method to randomly select several candidate matching point pairs, then use a direct linear transformation algorithm to calculate the initial 3D position features, and then calculate the projection error corresponding to all the aforementioned candidate matching points: ; Where π is used to characterize the projection function (camera model); Used to characterize the pose of the i-th camera; Used to characterize the current two-dimensional location features of the observation.

[0073] In the above-described embodiments, the embodiments provided in this application employ triangulation in the Random Sample Consensus (RANSAC) algorithm to reconstruct the spatial position of the traffic light within the vehicle system. Specifically, by constructing a system of projection equations based on the poses of multiple cameras over a period of time and the current two-dimensional position features of the target traffic light under the corresponding poses, for the i-th camera, its projection model is: ; in, These are normalized image coordinates. It is the camera intrinsic parameter matrix. The camera extrinsic parameters (transformation from world coordinates to camera coordinates) are known. Based on these, a system of projection equations is constructed. Expanding the projection equations and eliminating the scale factor of the homogeneous coordinates yields two linear constraints:

[0074] in It is a projection matrix The j-th line. This is equivalent to: ; Here, the embodiment provided in this application selects the candidate matching point pair with the most interior points to calculate the projection equation system, and after solving it, optimizes all candidate matching points again, repeating the above steps for iteration until the interior points are the most. Finally, the n selected candidate matching point pairs are used, and the size of their matrix is ​​designed to be: A=2n×4. The projection equation system AP=0 is an overdetermined equation. By solving the least squares solution of this overdetermined equation, the current three-dimensional position feature of the traffic light is recovered.

[0075] As an optional implementation in this embodiment, the current three-dimensional features of the target traffic light are determined based on the current shape features, current state features, and the current three-dimensional position features of the target traffic light. Determining the current three-dimensional features of the target traffic light includes: The system acquires the current shape features and current state features of each target traffic light in multiple consecutive frames within a preset time period; determines the shape frequency of each target traffic light shape and the state frequency of each target traffic light state; updates the current shape features and current state features based on a preset voting algorithm, identifying current shape features with a shape frequency exceeding a preset shape frequency threshold as target shape features, and identifying current state features with a state frequency exceeding a preset state frequency threshold as target state features; and determines the current three-dimensional features of the target traffic light based on the target shape features, target state features, and the current three-dimensional position features of the target traffic light.

[0076] It should be noted that, after determining the current three-dimensional position features of the traffic light, the embodiments provided in this application need to bind the target shape features and target state features of the traffic light to the aforementioned current three-dimensional position features. The bound target shape features and target state features are shape features and state features that have a high and stable proportion over a period of time, output by a voting algorithm.

[0077] It is understood that the preset state frequency threshold and preset shape frequency threshold settings in the embodiments provided in this application can be customized and used according to different application scenarios and usage conditions.

[0078] The traffic light positioning and tracking method provided in this application, compared with the prior art, can acquire traffic light images of the current frame collected by multiple image acquisition devices on the target vehicle, and perform image perception processing on the traffic light images of the current frame to obtain the current two-dimensional features of at least one target traffic light. Then, based on the current two-dimensional position features of at least one target traffic light and the historical position features of each historical traffic light in the historical database, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light. Next, based on the current two-dimensional position features of the target traffic light and the poses of multiple image acquisition devices, the matching point pairs are projected to determine the current three-dimensional position features of the target traffic light. Finally, based on the current shape features, current state features, and the current three-dimensional position features of the target traffic light, the current three-dimensional features of the target traffic light are determined to complete the positioning and tracking of the target traffic light. This application can track target traffic lights with changing positions, improve the accuracy and reliability of traffic light position recognition during vehicle driving, and is more conducive to the accuracy of subsequent path planning. It achieves stable tracking of target red and green traffic lights, reduces the dependence of the vehicle system on high-precision maps, and reduces detection costs.

[0079] Please see Figure 2 , Figure 2 This is a structural block diagram of a traffic light positioning and tracking device provided in an embodiment of this application. Figure 2 As shown, the traffic light positioning and tracking device 200 includes: The acquisition module 210 is used to acquire the signal light image of the current frame collected by multiple image acquisition devices on the target vehicle.

[0080] The processing module 220 is used to perform image perception processing on the traffic light image of the current frame to obtain the current two-dimensional features of at least one target traffic light; wherein the current two-dimensional features include current two-dimensional position features, current shape features and current state features.

[0081] The first determining module 230 is used to match at least one target traffic light with each historical traffic light based on the current two-dimensional position features of at least one target traffic light and the historical position features of each historical traffic light in the historical database, and to determine the matching point pairs of each target traffic light.

[0082] The second determining module 240 is used to perform projection processing on the matching point pair based on the current two-dimensional position features of the target traffic light and the poses of multiple image acquisition devices to determine the current three-dimensional position features of the target traffic light.

[0083] The third determining module 250 is used to determine the current three-dimensional features of the target traffic light based on the current shape features, current state features, and the current three-dimensional position features of the target traffic light, so as to complete the positioning and tracking of the target traffic light.

[0084] For example, the historical position features of each historical traffic light include historical three-dimensional position features, and the first determining module 230 is specifically used for: The historical three-dimensional location features of each historical traffic light are obtained from the historical database, which includes a historical frame database and a historical map database.

[0085] Based on the 3D projection constraint matching algorithm, the historical 3D position features of each historical traffic light are projected onto the plane of the traffic light image in the current frame to determine the 3D matching score between at least one target traffic light and each historical traffic light.

[0086] Based on the three-dimensional matching score and the preset matching algorithm, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light.

[0087] For example, based on a 3D projection constraint matching algorithm, the historical 3D position features of each historical traffic light are projected onto the plane of the traffic light image in the current frame to determine the 3D matching score between at least one target traffic light and each historical traffic light, including: Based on the 3D projection constraint matching algorithm, historical 3D position features, and camera parameters corresponding to the image acquisition device, the projection position of each historical traffic light on the plane of the traffic light image in the current frame is determined.

[0088] Determine the distance difference between the projected position and the center point of the two-dimensional position features of each target signal light.

[0089] Based on the distance difference, a three-dimensional matching score is determined between at least one target traffic light and each historical traffic light, where the smaller the distance difference, the higher the matching score.

[0090] For example, the historical position features of each historical traffic light include historical two-dimensional position features. The first determining module 230 is specifically used for: The historical two-dimensional location features of each historical traffic light are obtained from the historical database, which includes a historical frame database and a historical map database.

[0091] Based on the epipolar geometry constraint algorithm, the current two-dimensional features of at least one target traffic light, and the historical two-dimensional features of each historical traffic light, the epipolar distance from each historical traffic light to at least one target traffic light is determined.

[0092] Based on the epipolar distance and a preset distance threshold, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light in the current frame.

[0093] For example, based on the epipolar geometry constraint algorithm, the current two-dimensional features of at least one target traffic light, and the historical two-dimensional position features of each historical traffic light, the epipolar distance from each historical traffic light to at least one target traffic light is determined, including: The relative motion data of the target vehicle is determined based on the camera parameters of multiple image acquisition devices.

[0094] The fundamental matrix of the current frame is determined based on relative motion data, camera parameters corresponding to the image acquisition device, and epipolar geometry constraint algorithm.

[0095] The epipolar line of the current frame is determined based on the fundamental matrix, the historical two-dimensional position features, and the current position two-dimensional features of at least one target traffic light.

[0096] The distance from the two-dimensional position center point of at least one target signal light in the current frame to the epipolar line is defined as the epipolar distance.

[0097] For example, the first determining module 230 is specifically used for: Based on the two-dimensional spatial constraint algorithm and epipolar distance, the two-dimensional matching score of the historical position features of each historical traffic light in the traffic light image of the current frame is determined.

[0098] Based on the two-dimensional matching score and the preset matching algorithm, the matching point pairs of each target traffic light in the current frame are determined.

[0099] For example, the preset matching algorithm includes a weighted bipartite graph matching algorithm, which determines the matching point pairs of each target traffic light in the current frame based on the two-dimensional matching score and the preset matching algorithm, including: Determine the two-dimensional matching score matrix based on the two-dimensional matching score; Using the current set of traffic lights as the first vertex set, the historical set of traffic lights as the second vertex set, and the two-dimensional matching score matrix as the edge weights, a weighted bipartite graph matching algorithm is executed to determine the optimal matching pair.

[0100] The optimal matching pair is determined as the matching point pair of each target signal light in the current frame.

[0101] For example, the second determining module 240 is specifically used to: filter matching point pairs based on a preset abnormal point filtering algorithm to determine candidate matching point pairs.

[0102] Based on the poses of multiple image acquisition devices associated with candidate matching point pairs and the current two-dimensional position features of the target signal light under multiple poses, a set of projection equations is constructed.

[0103] Based on candidate matching point pairs and the least squares method, the projection equations are calculated to determine the current three-dimensional position features of the target traffic light.

[0104] For example, the third determining module 250 is specifically used to: obtain the current shape features and current state features of each target traffic light in multiple consecutive frames within a preset time period.

[0105] Determine the shape frequency of the target traffic lights of each shape and the state frequency of the target traffic lights of each state.

[0106] Based on a preset voting algorithm, the current shape features and the current state features are updated respectively. The current shape features whose shape frequency exceeds a preset shape frequency threshold are determined as target shape features, and the current state features whose state frequency exceeds a preset state frequency threshold are determined as target state features.

[0107] Based on the target shape features, target state features, and the current three-dimensional position features of the target traffic light, the current three-dimensional features of the target traffic light are determined.

[0108] The traffic light positioning and tracking device 200 provided in this application embodiment, compared with the prior art, can acquire traffic light images of the current frame collected by multiple image acquisition devices on the target vehicle, and perform image perception processing on the traffic light images of the current frame to obtain the current two-dimensional features of at least one target traffic light. Then, based on the current two-dimensional position features of at least one target traffic light and the historical position features of each historical traffic light in the historical database, at least one target traffic light is matched with each historical traffic light to determine the matching point pairs of each target traffic light. Next, based on the current two-dimensional position features of the target traffic light and the poses of multiple image acquisition devices, the matching point pairs are projected to determine the current three-dimensional position features of the target traffic light. Finally, based on the current shape features, current state features, and current three-dimensional position features of the target traffic light, the current three-dimensional features of the target traffic light are determined to complete the positioning and tracking of the target traffic light. This application can track target traffic lights with changing positions, improving the accuracy and reliability of traffic light position identification during vehicle driving.

[0109] Please see Figure 3 , Figure 3 This application provides a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 3 As shown, the electronic device 300 includes a processor 310, a memory 320, and a bus 330.

[0110] Memory 320 stores machine-readable instructions executable by processor 310. When electronic device 300 is running, processor 310 and memory 320 communicate via bus 330. When the machine-readable instructions are executed by processor 310, they can perform the operations described above. Figure 1 The steps of the traffic light positioning and tracking method embodiment shown in the method embodiment can be found in the method embodiment for specific implementation, and will not be repeated here.

[0111] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described actions. Figure 1 The steps of the traffic light positioning and tracking method embodiment shown in the method embodiment can be found in the method embodiment for specific implementation, and will not be repeated here.

[0112] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0113] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0114] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-readable program code.

[0115] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0116] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0117] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes ​ The steps of the function specified in one or more boxes.

[0118] This application also provides a computer program product, which includes computer software instructions that, when executed on a processing device, cause the processing device to execute a process for determining a fault identification model.

[0119] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0120] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0121] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0122] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0123] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0124] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0125] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

[0126] Although preferred embodiments have been described in this specification, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this specification.

[0127] Obviously, those skilled in the art can make various modifications and variations to this specification without departing from its spirit and scope. Therefore, if such modifications and variations fall within the scope of the claims and their equivalents, this specification is also intended to include such modifications and variations.

Claims

1. A method for locating and tracking traffic lights, characterized in that, The method includes: Acquire the traffic light images of the current frame captured by multiple image acquisition devices on the target vehicle; Image perception processing is performed on the traffic light image of the current frame to obtain the current two-dimensional features of at least one target traffic light; wherein, the current two-dimensional features include current two-dimensional position features, current shape features, and current state features; Based on the current two-dimensional position features of the at least one target traffic light and the historical position features of each historical traffic light in the historical database, the at least one target traffic light is matched with each of the historical traffic lights to determine the matching point pairs of each target traffic light; Based on the current two-dimensional position features of the target traffic light and the poses of the multiple image acquisition devices, the matching point pairs are projected to determine the current three-dimensional position features of the target traffic light. Based on the current shape features, the current state features, and the current three-dimensional position features of the target traffic light, the current three-dimensional features of the target traffic light are determined to complete the positioning and tracking of the target traffic light.

2. The method for locating and tracking traffic lights according to claim 1, characterized in that, The historical location features of each historical traffic light include historical three-dimensional location features. The process of matching the at least one target traffic light with each of the historical traffic lights based on the current two-dimensional features of the at least one target traffic light and the historical two-dimensional features of each historical traffic light in the historical database, to determine the matching point pairs of each target traffic light in the current frame, includes: The historical three-dimensional location features of each historical traffic light are obtained from the historical database; wherein, the historical database includes a historical frame database and a historical map database; Based on the three-dimensional projection constraint matching algorithm, the historical three-dimensional position features of each historical traffic light are projected onto the plane of the traffic light image in the current frame to determine the three-dimensional matching score between the at least one target traffic light and each historical traffic light. Based on the three-dimensional matching score and the preset matching algorithm, at least one target traffic light is matched with each of the historical traffic lights to determine the matching point pairs of each target traffic light.

3. The method for locating and tracking traffic lights according to claim 2, characterized in that, The three-dimensional projection constraint matching algorithm projects the historical three-dimensional position features of each historical traffic light onto the plane of the traffic light image in the current frame, and determines the three-dimensional matching score between the at least one target traffic light and each historical traffic light, including: Based on the three-dimensional projection constraint matching algorithm, the historical three-dimensional position features, and the camera parameters corresponding to the image acquisition device, the projection position of each historical traffic light on the plane of the traffic light image in the current frame is determined. Determine the distance difference between the projected position and the center point of the two-dimensional position feature of each of the target traffic lights; Based on the distance difference, a three-dimensional matching score is determined between the at least one target traffic light and each of the historical traffic lights, wherein the smaller the distance difference, the higher the matching score.

4. The method for locating and tracking traffic lights according to claim 1, characterized in that, The historical location features of each historical traffic light include historical two-dimensional location features. The process of matching the at least one target traffic light with each of the historical traffic lights based on the current two-dimensional features of the at least one target traffic light and the historical two-dimensional features of each historical traffic light in the historical database, to determine the matching point pairs of each target traffic light in the current frame, includes: The historical two-dimensional location features of each historical traffic light are obtained from the historical database; wherein, the historical database includes a historical frame database and a historical map database; Based on the epipolar geometry constraint algorithm, the current two-dimensional features of the at least one target traffic light, and the historical two-dimensional features of each historical traffic light, the epipolar distance from each of the historical traffic lights to the at least one target traffic light is determined. Based on the epipolar distance and the preset distance threshold, the at least one target traffic light is matched with each of the historical traffic lights to determine the matching point pairs of each of the target traffic lights in the current frame.

5. The method for locating and tracking traffic lights according to claim 4, characterized in that, The determination of the epipolar distance from each historical traffic light to at least one target traffic light, based on the epipolar geometry constraint algorithm, the current two-dimensional features of the at least one target traffic light, and the historical two-dimensional position features of each historical traffic light, includes: The relative motion data of the target vehicle is determined based on the camera parameters of multiple image acquisition devices; Based on the relative motion data, the camera parameters corresponding to the image acquisition device, and the epipolar geometry constraint algorithm, the fundamental matrix of the current frame is determined. Based on the base matrix, each of the historical two-dimensional position features, and the current position two-dimensional features of at least one target traffic light, the epipolar line of the current frame is determined. The distance from the two-dimensional position center point of at least one target signal light in the current frame to the polar line is determined as the polar distance.

6. The method for locating and tracking traffic lights according to claim 5, characterized in that, The step of matching the at least one target traffic light with each of the historical traffic lights in the historical database based on the current two-dimensional features of the at least one target traffic light to determine the matching point pairs of each target traffic light in the current frame includes: Based on the two-dimensional spatial constraint algorithm and the epipolar distance, the two-dimensional matching score of the historical position features of each historical traffic light in the traffic light image of the current frame is determined; Based on the two-dimensional matching score and the preset matching algorithm, the matching point pairs of each target traffic light in the current frame are determined.

7. The method for locating and tracking traffic lights according to claim 6, characterized in that, The preset matching algorithm includes a weighted bipartite graph matching algorithm. The step of determining the matching point pairs of each target traffic light in the current frame based on the two-dimensional matching score and the preset matching algorithm includes: Determine the two-dimensional matching score matrix based on the two-dimensional matching score; Using the current set of traffic lights as the first vertex set, the historical set of traffic lights as the second vertex set, and the two-dimensional matching score matrix as the edge weights, a weighted bipartite graph matching algorithm is executed to determine the optimal matching pair. The optimal matching pair is determined as the matching point pair of each target traffic light in the current frame.

8. The method for locating and tracking traffic lights according to claim 1, characterized in that, The step of determining the current three-dimensional position features of the target traffic light by projecting the matching point pairs based on the current two-dimensional position features of the target traffic light and the poses of the multiple image acquisition devices includes: Based on a preset outlier filtering algorithm, the matching point pairs are filtered to determine candidate matching point pairs; Based on the poses of multiple image acquisition devices associated with the candidate matching point pairs and the current two-dimensional position features of the target traffic light under the multiple poses, a system of projection equations is constructed. Based on the candidate matching point pairs and the least squares method, the projection equations are calculated to determine the current three-dimensional position features of the target traffic light.

9. The method for locating and tracking traffic lights according to claim 1, characterized in that, Determining the current three-dimensional features of the target traffic light based on the current shape features, the current state features, and the current three-dimensional position features of the target traffic light includes: Obtain the current shape features and current state features of each target traffic light in multiple consecutive frames within a preset time period; Determine the shape frequency of the target traffic light in each shape, and the state frequency of the target traffic light in each state; Based on a preset voting algorithm, the current shape feature and the current state feature are updated respectively. The current shape feature whose shape frequency exceeds a preset shape frequency threshold is determined as the target shape feature, and the current state feature whose state frequency exceeds a preset state frequency threshold is determined as the target state feature. Based on the target shape features, the target state features, and the current three-dimensional position features of the target traffic light, the current three-dimensional features of the target traffic light are determined.

10. A positioning and tracking device for a traffic light, characterized in that, The positioning and tracking device for the traffic lights includes: The acquisition module is used to acquire the traffic light image of the current frame captured by multiple image acquisition devices on the target vehicle; The processing module is used to perform image perception processing on the traffic light image of the current frame to obtain the current two-dimensional features of at least one target traffic light; wherein, the current two-dimensional features include current two-dimensional position features, current shape features, and current state features; The first determining module is used to match the at least one target traffic light with each of the historical traffic lights based on the current two-dimensional position features of the at least one target traffic light and the historical position features of each historical traffic light in the historical database, and to determine the matching point pairs of each of the target traffic lights. The second determining module is used to perform projection processing on the matching point pair based on the current two-dimensional position features of the target traffic light and the poses of the multiple image acquisition devices to determine the current three-dimensional position features of the target traffic light. The third determining module is used to determine the current three-dimensional features of the target traffic light based on the current shape features, the current state features, and the current three-dimensional position features of the target traffic light, so as to complete the positioning and tracking of the target traffic light.