Method, device and electronic device for obtaining three-dimensional properties of target traffic light
By using neural network models to analyze the images of the image acquisition component in autonomous driving vehicles, and combining internal reference to calculate the three-dimensional properties of traffic lights, the problem of time-consuming and cost-effective labeling is solved, and the perception accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202211707327.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-27
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-12-27
AI Technical Summary
In the prior art, it takes a long time and is costly to obtain the labeling of three-dimensional attributes of traffic lights, which affects the perceptual accuracy and the difficulty of data analysis closed-loop.
By acquiring the images of the image acquisition component of the autonomous driving vehicle, using the target neural network model to analyze error information and angle information, combining the internal reference of the image acquisition component, calculate the depth information of the traffic light, and finally determine its three-dimensional properties.
It realizes efficient determination of the three-dimensional properties of traffic lights, reduces labeling costs, and improves perception accuracy and data closed-loop efficiency.
Smart Images

Figure CN115909284B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, in particular to technical fields such as autonomous driving and computer vision. Specifically, it relates to a method, device, and electronic device for obtaining three-dimensional properties of a target traffic light. Background Art
[0002] In autonomous driving scenarios, accurate traffic light perception is crucial for driving safety. Because the three-dimensional (3D) properties of traffic lights cannot be acquired through high-precision 3D scanning sensors like LiDAR, related technologies typically employ a joint annotation approach to obtain these properties. However, this annotation approach is time-consuming and labor-intensive, and is subject to various processing constraints, which can affect traffic light perception accuracy and complicate closed-loop data analysis. Summary of the Invention
[0003] The present disclosure provides a method, device, and electronic device for obtaining the three-dimensional attributes of a target traffic light, so as to at least solve the technical problems in the related art of time-consuming and high-cost labeling of the three-dimensional attributes of traffic lights.
[0004] According to one aspect of the present disclosure, a method for obtaining three-dimensional properties of a target traffic light is provided, comprising: obtaining an image to be identified, wherein the image to be identified is obtained by an image acquisition component on an autonomous driving vehicle, and the display content in the image to be identified includes: a target traffic light; using a target neural network model to analyze the image to be identified to obtain error information and angle information, wherein the target neural network model is used to estimate size information and orientation information of the target traffic light, the error information is used to determine the size information, and the angle information is used to determine the orientation information; based on the error information, the angle information, the size information of the image to be identified and the internal parameters of the image acquisition component, obtaining depth information of the target traffic light; and determining the three-dimensional properties of the target traffic light through the error information, the angle information and the depth information.
[0005] According to another aspect of the present disclosure, a device for obtaining three-dimensional properties of a target traffic light is provided, including: a first acquisition module, used to obtain an image to be identified, wherein the image to be identified is obtained by an image acquisition component on an autonomous driving vehicle, and the display content in the image to be identified includes: a target traffic light; an analysis module, used to analyze the image to be identified using a target neural network model to obtain error information and angle information, wherein the target neural network model is used to estimate size information and orientation information of the target traffic light, the error information is used to determine the size information, and the angle information is used to determine the orientation information; a second acquisition module, used to obtain depth information of the target traffic light based on the error information, angle information, size information of the image to be identified and internal parameters of the image acquisition component; a determination module, used to determine the three-dimensional properties of the target traffic light through the error information, angle information and depth information.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for obtaining three-dimensional properties of a target traffic light according to any embodiment of the present disclosure.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method for obtaining three-dimensional properties of a target traffic light according to any embodiment of the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When a processor executes the computer program, the method for obtaining three-dimensional attributes of a target traffic light according to any embodiment of the present disclosure is implemented.
[0009] In the present disclosure, an image to be identified containing a target traffic light is obtained, and then a target neural network model is used to analyze the image to be identified to obtain error information and angle information. Subsequently, based on the error information, angle information, size information of the image to be identified and the internal parameters of the image acquisition component, the depth information of the target traffic light is obtained. Finally, the three-dimensional properties of the target traffic light are determined through the error information, angle information and depth information, thereby achieving the purpose of efficiently determining the three-dimensional properties of the target traffic light in the image to be identified, and realizing the effect of improving the labeling efficiency of the three-dimensional properties of the traffic light and reducing the labeling cost, thereby solving the technical problems of long labeling time and high labeling cost for the three-dimensional properties of traffic lights in related technologies.
[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0012] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for obtaining three-dimensional attributes of a target traffic light according to an embodiment of the present disclosure;
[0013] Figure 2 is a flow chart of a method for obtaining three-dimensional attributes of a target traffic light according to an embodiment of the present disclosure;
[0014] Figure 3 is a schematic diagram of a marking box according to an embodiment of the present disclosure;
[0015] Figure 4 is a schematic diagram of a pinhole imaging model according to an embodiment of the present disclosure;
[0016] Figure 5 is a schematic structural diagram of a target neural network model according to an embodiment of the present disclosure;
[0017] Figure 6 is a schematic diagram of obtaining three-dimensional attributes of a target traffic light according to an embodiment of the present disclosure;
[0018] Figure 7 is a schematic diagram of another method for obtaining three-dimensional attributes of a target traffic light according to an embodiment of the present disclosure;
[0019] Figure 8 This is a structural block diagram of a device for obtaining three-dimensional attributes of a target traffic light according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0022] The methods for monocular 3D object detection in related technologies generally include the following four methods:
[0023] The first approach, based on the geometric assumption that the target object is located on the ground, converts the perspective image into a feature map in the Bird's Eye View (BEV) space, and then performs 3D object detection in the BEV perspective. However, this approach requires the target object to be located on the ground or on a plane at a uniform height. Traffic lights have a wide range of heights, so this requirement significantly limits the application of this detection method in traffic light perception scenarios.
[0024] The second method is to generate pseudo radar features based on 2D images and then perform 3D target detection on dense depth maps. However, this method has a strong dependence on depth map annotation, which is time-consuming and costly.
[0025] The third method uses 3D key points to match 3D models to achieve 3D target detection. This method requires additional labeling of the key points of the target object, while the typical 2D target dataset does not have key point labels. In addition, since the traffic lights in most scenes are facing the camera, only 4 of the 8 key points of the traffic lights are visible, so this method has pathological problems.
[0026] The fourth method is to directly generate a 3D object frame, but this method is not easy to decouple and there is a problem of interference between 3D and 2D attributes. Since the 2D attributes of traffic lights in autonomous driving are also needed to bind map annotation lights and perform subsequent recognition, the inaccuracy of the externally introduced 2D attributes will cause more false alarm problems.
[0027] Therefore, the 3D object detection method in the related art cannot quickly and accurately obtain the three-dimensional properties of the traffic light.
[0028] According to an embodiment of the present disclosure, a method for obtaining three-dimensional properties of a target traffic light is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0029] The method embodiments provided in the embodiments of the present disclosure can be executed in a mobile terminal, a computer terminal or a similar electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for acquiring three-dimensional properties of a target traffic light is shown.
[0030] like Figure 1 As shown, the computer terminal 100 includes a computing unit 101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 102 or a computer program loaded from a storage unit 108 into a random access memory (RAM) 103. Various programs and data required for the operation of the computer terminal 100 can also be stored in the RAM 103. The computing unit 101, the ROM 102, and the RAM 103 are connected to each other via a bus 104. An input / output (I / O) interface 105 is also connected to the bus 104.
[0031] Multiple components in the computer terminal 100 are connected to the I / O interface 105, including an input unit 106, such as a keyboard, a mouse, etc.; an output unit 107, such as various types of displays, speakers, etc.; a storage unit 108, such as a magnetic disk, an optical disk, etc.; and a communication unit 109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 109 allows the computer terminal 100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0032] The computing unit 101 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 101 performs the method for obtaining the three-dimensional properties of a target traffic light described herein. For example, in some embodiments, the method for obtaining the three-dimensional properties of a target traffic light can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the computer terminal 100 via the ROM 102 and / or the communication unit 109. When the computer program is loaded into the RAM 103 and executed by the computing unit 101, one or more steps of the method for obtaining the three-dimensional properties of a target traffic light described herein can be performed. Alternatively, in other embodiments, the computing unit 101 may be configured in any other appropriate manner (eg, by means of firmware) to execute the method for obtaining the three-dimensional properties of the target traffic light.
[0033] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0034] It should be noted that, in some optional embodiments, the above Figure 1 The electronic device shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the electronic device described above.
[0035] Under the above operating environment, the present disclosure provides the following Figure 2The method for obtaining the three-dimensional properties of the target traffic light shown in FIG. Figure 1 The computer terminal or similar electronic device shown is used for execution. Figure 2 This is a flow chart of a method for obtaining the three-dimensional attributes of a target traffic light according to an embodiment of the present disclosure. Figure 2 As shown, the method may include the following steps:
[0036] Step S21, obtaining an image to be recognized, wherein the image to be recognized is obtained by an image acquisition component on the autonomous driving vehicle, and the display content in the image to be recognized includes: a target traffic light;
[0037] The above-mentioned image to be recognized can be collected by a camera on an autonomous driving vehicle, and the display content in the image to be recognized includes traffic lights on the road where the vehicle is traveling.
[0038] Step S22: Analyze the image to be recognized using a target neural network model to obtain error information and angle information, wherein the target neural network model is used to estimate the size information and orientation information of the target traffic light, the error information is used to determine the size information, and the angle information is used to determine the orientation information;
[0039] Step S23, acquiring depth information of the target traffic light based on the error information, the angle information, the size information of the image to be recognized, and the internal parameters of the image acquisition component;
[0040] Step S24 , determining the three-dimensional attributes of the target traffic light through the error information, angle information, and depth information.
[0041] Specifically, the three-dimensional attributes of the target traffic light include size, orientation, and depth information. These three-dimensional attributes encompass three degrees of freedom. Size information includes the length, width, and height of the target traffic light; orientation information includes the pitch angle (around the X-axis), the yaw angle (around the Y-axis), and the roll angle (around the Z-axis); and depth information includes the world coordinates of the target traffic light, namely its X-axis, Y-axis, and Z-axis coordinates in space.
[0042] According to the above steps S21 to S24 of the present disclosure, by obtaining an image to be identified containing a target traffic light, and then using a target neural network model to analyze the image to be identified, error information and angle information are obtained, and then based on the error information, angle information, size information of the image to be identified and the internal parameters of the image acquisition component, the depth information of the target traffic light is obtained, and finally the three-dimensional attributes of the target traffic light are determined through the error information, angle information and depth information, thereby achieving the purpose of efficiently determining the three-dimensional attributes of the target traffic light in the image to be identified, and realizing the effect of improving the labeling efficiency of the three-dimensional attributes of the traffic light and reducing the labeling cost, thereby solving the technical problems of long labeling time and high labeling cost for the three-dimensional attributes of traffic lights in related technologies.
[0043] The following further introduces the method for obtaining the three-dimensional attributes of the target traffic light in the above embodiment.
[0044] As an optional implementation, the method for obtaining the three-dimensional attributes of the target traffic light further includes:
[0045] Step S31, obtaining a sample image, wherein the sample image is obtained by an image acquisition component, and the display content in the sample image includes: a sample traffic light;
[0046] Step S32: generating second annotation data based on the preset accuracy map, the original positioning parameters of the image acquisition component, and the first annotation data, wherein the preset accuracy map and the original positioning parameters are used to determine a first annotation box of the sample traffic light on the sample image, and the first annotation data is used to determine a second annotation box of the sample traffic light on the sample image;
[0047] Step S33: Iteratively train the initial neural network model using the second labeled data to obtain a target neural network model.
[0048] The above-mentioned preset precision maps are high-precision maps. High-precision maps have an absolute position accuracy of nearly 1 meter and a relative position accuracy of 10-20 cm at the centimeter level, accurately and comprehensively representing road characteristics. Furthermore, high-precision maps can record specific details of driving behavior, including typical driving behavior, optimal acceleration and braking points, road complexity, and annotation of signal reception conditions in different road sections.
[0049] The first annotation data is 2D annotation data of the sample image, and the second annotation data is traffic light annotation data with 3D attributes.
[0050] Based on the above optional implementation method, by acquiring a sample image, and then generating second annotation data based on a preset precision map, the original positioning parameters of the image acquisition component and the first annotation data, and finally using the second annotation data to iteratively train the initial neural network model, the target neural network model can be quickly obtained, effectively improving the model training efficiency.
[0051] As an optional implementation, in step S32, generating second annotation data based on the preset accuracy map, the original positioning parameters, and the first annotation data includes:
[0052] Step S321: projecting a first annotation frame onto the sample image using a preset accuracy map and original positioning parameters, and annotating the sample image using the first annotation data to obtain a second annotation frame;
[0053] Step S322: Adjust the original positioning parameters by grid search until the matching error between the first annotation frame and the second annotation frame is minimized, thereby obtaining the target positioning parameters.
[0054] Step S323 : generating second annotation data based on the target positioning parameters and the three-dimensional attributes of the annotated traffic lights, wherein the annotated traffic lights are traffic lights corresponding to the sample traffic lights annotated on the preset accuracy map.
[0055] Specifically, Figure 3 is a schematic diagram of a marking frame according to an embodiment of the present disclosure, such as Figure 3 As shown in the figure, the preset accuracy map and original positioning parameters are used to project the labeled traffic lights onto the sample image through the pinhole imaging model, thereby obtaining Figure 3 The first annotation box in the sample image is annotated on the sample image using the 2D annotation data of the sample image. Figure 3 The second callout box in . Figure 3 It can be seen that there is a certain matching error between the first and second annotation boxes. The sources of matching error include map annotation error, positioning error, and camera calibration error. Among them, the most important source of matching error is positioning error. Annotation error is the error introduced by sensor and algorithm errors during the high-precision map production process; positioning error is the position error caused by the positioning hardware and algorithm of the main vehicle; and camera calibration error is the error in the camera's internal and external parameters during the calibration process.
[0056] Because the matching error between the first batch of annotations and the second batch of annotations primarily stems from positioning error, a grid search method is used to adjust the original positioning parameters until the matching error between the first batch of annotations and the second batch is minimized, thereby obtaining the target positioning parameters. Grid search is a parameter adjustment method that optimizes the matching error by traversing a given parameter combination, thereby obtaining the target positioning parameters that minimize the matching error between the first batch of annotations and the second batch of annotations. Furthermore, the three-dimensional attributes of the traffic light are annotated based on the target positioning parameters to generate the second annotation data.
[0057] Based on the above optional implementation manner, a first annotation box is obtained by projecting a preset precision map and the original positioning parameters onto a sample image, and a second annotation box is obtained by annotating the sample image using the first annotation data. The original positioning parameters are then adjusted through a grid search method until the matching error between the first annotation box and the second annotation box is minimized, thereby obtaining the target positioning parameters. Finally, based on the target positioning parameters and the three-dimensional attributes of the annotated traffic lights, the second annotation data is generated, which can quickly obtain the training data of the target neural network model, thereby improving the analysis performance of the target neural network model.
[0058] As an optional implementation, in step S323, generating second annotation data based on the target positioning parameters and the three-dimensional attributes of the annotated traffic light includes:
[0059] Step S3231: determining a first association relationship between the labeled traffic light and a target labeled frame based on the target positioning parameters, wherein the target labeled frame is a labeled frame on the sample image that is adapted to the sample traffic light;
[0060] Step S3232: using the first association relationship, the three-dimensional attributes of the annotated traffic light are transferred to the target annotation frame to generate second annotation data.
[0061] Specifically, the target annotation box is a 2D annotation box on the sample image that matches the sample traffic light. In this case, the target annotation box and the annotated traffic light can correspond to the same traffic light in the real world. The annotated traffic light and the target annotation box are associated using the target positioning parameters, thereby obtaining a first association relationship. Furthermore, based on the first association relationship, the 3D attributes of the annotated traffic light can be assigned to the target annotation box, thereby obtaining second annotation data.
[0062] Based on the above optional implementation method, the first association relationship between the labeled traffic light and the target labeled frame is determined based on the target positioning parameters, and then the first association relationship is used to transfer the three-dimensional attributes of the labeled traffic light to the target labeled frame to generate second annotation data, which can quickly obtain the training data of the target neural network model, thereby improving the analysis performance of the target neural network model.
[0063] As an optional implementation, in step S3232, using the first association relationship, the three-dimensional attributes of the annotated traffic light are transferred to the target annotation frame, and the second annotation data is generated, including:
[0064] Step S32321: using the first association relationship, transferring the three-dimensional attributes of the annotated traffic light to the target annotation frame to obtain third annotation data;
[0065] Step S32322: determining first coordinate data of the sample traffic light in a first coordinate system based on the third labeled data, wherein the first coordinate system is a coordinate system corresponding to the image acquisition component;
[0066] Step S32323, using the first coordinate data and the internal parameters of the image acquisition component to determine the two-dimensional pixel difference of the sample traffic light on the sample image;
[0067] Step S32324: Correct the third labeled data based on the two-dimensional pixel difference to generate second labeled data.
[0068] Specifically, the third annotation data is the empirical value of the size of the 3D annotation box in the high-precision map. The first coordinate data of the sample traffic light in the camera coordinate system is determined based on the third annotation data. Furthermore, based on the pinhole imaging model, the first coordinate data and the camera's intrinsic parameters are used to determine the two-dimensional pixel difference of the sample traffic light in the sample image. The first coordinate data is the 3D coordinate of the sample traffic light in the camera coordinate system, the camera's intrinsic parameter is the camera's focal length, and the two-dimensional pixel difference includes the pixel difference in the length and width directions of the sample traffic light.
[0069] Figure 4 is a schematic diagram of a pinhole imaging model according to an embodiment of the present disclosure, which can be used to establish an association between a camera coordinate system and a pixel coordinate system. Figure 4 In the equation, P = (X w ,Y w ,Z w ) is the target coordinate in the world coordinate system, F c is the origin of the camera coordinate system, and (u, v) is the target coordinate in the pixel coordinate system. According to the pinhole imaging model, the first coordinate data and the camera's intrinsic parameters can be used to determine the two-dimensional pixel difference of the sample traffic light on the sample image according to the following formulas 1 and 2:
[0070] Δu=X / (Z·f x ) Formula 1
[0071] Δv=Y / (Z·f y ) Formula 2
[0072] Wherein, Δu is the pixel difference of the sample traffic light in the length direction, Δv is the pixel difference of the sample traffic light in the width direction, X, Y, Z are the 3D coordinates of the sample traffic light in the camera coordinate system, and f x and f y is the focal length of the camera.
[0073] Based on the above optional implementation manner, the three-dimensional attributes of the annotated traffic light are transferred to the target annotation frame using the first association relationship to obtain third annotation data, and then the first coordinate data of the sample traffic light in the first coordinate system is determined based on the third annotation data. Subsequently, the first coordinate data and the internal parameters of the image acquisition component are used to determine the two-dimensional pixel difference of the sample traffic light on the sample image. Finally, the third annotation data is corrected based on the two-dimensional pixel difference, and more accurate second annotation data can be obtained for training the target neural network model, thereby improving the analysis performance of the target neural network model.
[0074] As an optional implementation, in step S22, the target neural network model is used to analyze the image to be recognized, and the error information and angle information obtained include:
[0075] Step S221, using the target neural network model to perform feature extraction on the image to be identified to obtain an extraction result;
[0076] Step S222: Perform multi-head prediction based on the extraction results to obtain a first prediction result, a second prediction result, and a third prediction result. The first prediction result is used to predict the angular range of the target angle, the second prediction result is used to predict the offset of the target angle within the angular range, and the third prediction result is used to predict the error of the current predicted size relative to a preset size. The target angle is the angle between the orientation of the target traffic light and the driving direction of the autonomous vehicle. The preset size is determined by the target positioning parameter of the image acquisition component.
[0077] Step S223: Determine error information and angle information using the first prediction result, the second prediction result, and the third prediction result.
[0078] Specifically, Figure 5 is a schematic diagram of the structure of a target neural network model according to an embodiment of the present disclosure, such as Figure 5As shown in the figure, the target neural network model has multiple residual blocks (resblocks) and a prediction head composed of a fully connected layer. Among them, the three stacked residual modules 1, 2, and 3 can extract features from the image to be recognized and obtain extraction results. Based on the extraction results, multi-head prediction can be performed to obtain the first prediction result, the second prediction result, and the third prediction result.
[0079] The angle bin of the target angle is predicted based on the first prediction result, the angle offset of the target angle within the angle bin is predicted based on the second prediction result, and the error (dim) of the current predicted size relative to the preset size is predicted based on the third prediction result.
[0080] Since the pitch and roll angles of the target traffic light have little impact on the autonomous vehicle, the orientation information in the three-dimensional attributes can be simplified to only one degree of freedom: the yaw angle, which is the angle between the orientation of the target traffic light and the driving direction of the autonomous vehicle.
[0081] This error information represents the difference between the target traffic light's dimensional information and a preset dimension. The preset dimension is the mean value obtained by performing a K-means clustering calculation on the target traffic light's dimensional information. Because horizontal, vertical, and square traffic lights vary significantly in length, width, and height, the accuracy of the target neural network model's direct learning of the target traffic light's dimensional information is susceptible to the influence of the sample data distribution. Therefore, the learning of dimensional information is transformed into learning of error information.
[0082] Based on the above optional implementation method, a target neural network model is used to perform feature extraction on the image to be identified to obtain an extraction result, and then a multi-head prediction is performed based on the extraction result to obtain a first prediction result, a second prediction result and a third prediction result. Finally, the first prediction result, the second prediction result and the third prediction result are used to quickly determine the error information and angle information, thereby improving the labeling efficiency of the three-dimensional attributes of the target traffic light.
[0083] As an optional implementation, in step S23, obtaining depth information of the target traffic light based on the error information, the angle information, the size information of the image to be recognized, and the internal parameters of the image acquisition component includes:
[0084] Step S231, obtaining a second correlation relationship between first coordinate data of a target traffic light in a first coordinate system and second coordinate data of the target traffic light in a second coordinate system, wherein the first coordinate system is a coordinate system corresponding to the image acquisition component, and the second coordinate system is a coordinate system corresponding to the image to be recognized;
[0085] Step S232: acquiring depth information based on the second association relationship, error information, angle information, size information, and internal parameters of the image acquisition component.
[0086] Based on the above optional implementation manner, by obtaining a second association relationship between the first coordinate data of the target traffic light in the first coordinate system and the second coordinate data of the target traffic light in the second coordinate system, and then based on the second association relationship, error information, angle information, size information and internal parameters of the image acquisition component, depth information can be obtained, thereby improving the labeling efficiency of the three-dimensional attributes of the target traffic light.
[0087] As an optional implementation manner, in step S232, acquiring depth information based on the second association relationship, error information, angle information, size information, and internal parameters of the image acquisition component includes:
[0088] Step S2321: Calculate a first position of the target traffic light using the error information, angle information, and internal parameters of the image acquisition component, and determine a translation matrix corresponding to the target traffic light using the size information, wherein the first position is the three-dimensional center position of the target traffic light.
[0089] Step S2322: Calculate a second position based on the second association relationship, the first position, and the translation matrix, where the second position is a two-dimensional spatial position corresponding to the first position in the image to be identified;
[0090] Step S2323: Acquire depth information using the first position and the second position.
[0091] Specifically, Figure 6 is a schematic diagram of obtaining the three-dimensional attributes of a target traffic light according to an embodiment of the present disclosure, such as Figure 6 As shown in , since the 2D frame of the target traffic light will tightly wrap the 3D frame of the target traffic light, the 3D space center position of the target traffic light is calculated using the error information, angle information and the intrinsic parameters of the camera, that is, Figure 6 The position of point a in the middle.
[0092] The second correlation relationship between the first coordinate data of the target traffic light in the camera coordinate system and the second coordinate data in the coordinate system corresponding to the image to be recognized can be defined by Formula 3:
[0093]
[0094] Among them, X, Y, Z are the first coordinate data, x, y, z are the second coordinate data, K is the intrinsic parameter matrix, R is the rotation matrix, and T is the translation matrix. The intrinsic parameter matrix, rotation matrix, and translation matrix can be expressed as follows:
[0095]
[0096]
[0097]
[0098] Expanding the matrices in Formula 3, we can obtain:
[0099]
[0100] Simplifying Formula 7, we can get:
[0101]
[0102] Among them, m i is r i A linear combination of the first coordinate data [X, Y, Z]. Figure 6 As shown, since the 2D box of the object must tightly wrap the 3D box, there are only four vertices of the 3D box that fall on the four edges of the 2D box. Therefore, using the corresponding four equations, the T in the translation matrix corresponding to the target traffic light can be obtained by the least squares method. x 、T y 、T z .
[0103] Furthermore, based on the second association relationship, the first position and the translation matrix, the second position is calculated using Formula 9 and Formula 10:
[0104]
[0105]
[0106] Furthermore, the first position and the second position are used to obtain depth information.
[0107] Based on the above optional implementation manner, the first position of the target traffic light is calculated using the error information, angle information and the internal parameters of the image acquisition component, and the translation matrix corresponding to the target traffic light is determined using the size information. Then, based on the second association relationship, the first position and the translation matrix, the second position is calculated. Finally, the first position and the second position are used to obtain depth information, thereby improving the labeling efficiency of the three-dimensional attributes of the target traffic light.
[0108] As an optional implementation manner, in step S24, determining the three-dimensional attributes of the target traffic light using the error information, the angle information, and the depth information includes:
[0109] Step S241, determining size information based on error information, and determining orientation information based on angle information;
[0110] Step S242: Determine the size information, orientation information, and depth information as three-dimensional attributes.
[0111] Based on the above optional implementation, the size information is determined by the error information, and the orientation information is determined by the angle information, and then the size information, orientation information and depth information are determined as three-dimensional attributes, so that the three-dimensional attributes of the target traffic light can be efficiently obtained.
[0112] Figure 7 is another schematic diagram of obtaining the three-dimensional attributes of a target traffic light according to an embodiment of the present disclosure, such as Figure 7 As shown, the method includes the following steps:
[0113] Step S701: obtaining an image to be recognized, wherein the image to be recognized is obtained by an image acquisition component on the autonomous driving vehicle, and the display content in the image to be recognized includes: a target traffic light;
[0114] Step S702, using the target neural network model to perform feature extraction on the image to be identified to obtain an extraction result;
[0115] Step S703, performing multi-head prediction based on the extraction result to obtain a first prediction result, a second prediction result, and a third prediction result;
[0116] Step S704, using the first prediction result, the second prediction result, and the third prediction result, to determine error information and angle information;
[0117] Step S705: Acquire a second correlation between first coordinate data of the target traffic light in the first coordinate system and second coordinate data of the target traffic light in the second coordinate system, wherein the first coordinate system is a coordinate system corresponding to the image acquisition component, and the second coordinate system is a coordinate system corresponding to the image to be recognized;
[0118] Step S706: Calculate a first position of the target traffic light using the error information, the angle information, and the internal parameters of the image acquisition component, and determine a translation matrix corresponding to the target traffic light using the size information, wherein the first position is the three-dimensional center position of the target traffic light.
[0119] Step S707: Calculate a second position based on the second association relationship, the first position, and the translation matrix, wherein the second position is a two-dimensional spatial position corresponding to the first position in the image to be identified;
[0120] Step S708, acquiring depth information using the first position and the second position;
[0121] Step S709 : determining the three-dimensional attributes of the target traffic light through the error information, angle information, and depth information.
[0122] Based on the above steps, the present disclosure uses high-precision maps and positioning information, combined with image 2D annotation to obtain 3D information of traffic lights as training labels. Then, on this data, the target neural network model is used to learn the error information and angle information of the target traffic lights, and then combined with the size information of the image to be identified and the camera internal parameters, the depth information of the target traffic lights is calculated according to the geometric constraints. Therefore, compared with the existing algorithms, the embodiment of the present disclosure does not need to retrain the 2D model, and will not introduce weight changes to cause the index to decline. It can ensure the effects of 2D and 3D, and has a certain degree of versatility. It can obtain the three-dimensional properties of traffic lights without relying on lidar and 3D joint annotation, thereby effectively improving the perception accuracy and the efficiency of the data closed loop when perceiving traffic lights in autonomous driving scenarios.
[0123] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, or of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present disclosure.
[0125] The present disclosure also provides a device for obtaining the three-dimensional attributes of a target traffic light. This device is used to implement the above-mentioned embodiments and preferred embodiments, and the details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0126] Figure 8 is a structural block diagram of a device for obtaining three-dimensional attributes of a target traffic light according to one embodiment of the present disclosure, such as Figure 8 As shown, a device 800 for acquiring three-dimensional attributes of a target traffic light includes: a first acquisition module 801 , an analysis module 802 , a second acquisition module 803 , and a determination module 804 .
[0127] A first acquisition module 801 is configured to acquire an image to be recognized, wherein the image to be recognized is obtained by an image acquisition component on the autonomous driving vehicle, and the display content in the image to be recognized includes: a target traffic light;
[0128] An analysis module 802 is configured to analyze the image to be recognized using a target neural network model to obtain error information and angle information, wherein the target neural network model is used to estimate the size and orientation information of the target traffic light, the error information is used to determine the size information, and the angle information is used to determine the orientation information;
[0129] A second acquisition module 803 is configured to acquire depth information of a target traffic light based on the error information, the angle information, the size information of the image to be recognized, and the internal parameters of the image acquisition component;
[0130] The determination module 804 is configured to determine the three-dimensional attributes of the target traffic light through the error information, the angle information, and the depth information.
[0131] Optionally, the device 800 for obtaining the three-dimensional properties of the target traffic light also includes: a third acquisition module 805, used to obtain a sample image, wherein the sample image is obtained by the image acquisition component, and the display content in the sample image includes: a sample traffic light; a generation module 806, used to generate second annotation data based on a preset accuracy map, the original positioning parameters of the image acquisition component and the first annotation data, wherein the preset accuracy map and the original positioning parameters are used to determine the first annotation box of the sample traffic light on the sample image, and the first annotation data is used to determine the second annotation box of the sample traffic light on the sample image; a training module 807, used to use the second annotation data to iteratively train the initial neural network model to obtain a target neural network model.
[0132] Optionally, the generation module 806 is also used to: use a preset precision map and original positioning parameters to project a first annotation box on the sample image, and use the first annotation data to annotate the sample image to obtain a second annotation box; adjust the original positioning parameters through a grid search method until the matching error between the first annotation box and the second annotation box is minimized, thereby obtaining the target positioning parameters; and generate second annotation data based on the target positioning parameters and the three-dimensional properties of the annotated traffic light, wherein the annotated traffic light is the traffic light corresponding to the sample traffic light annotated on the preset precision map.
[0133] Optionally, the generation module 806 is also used to: determine a first association relationship between the labeled traffic light and the target labeling frame based on the target positioning parameters, wherein the target labeling frame is a labeling frame on the sample image that is adapted to the sample traffic light; and utilize the first association relationship to transfer the three-dimensional attributes of the labeled traffic light to the target labeling frame to generate second labeling data.
[0134] Optionally, the generation module 806 is also used to: utilize the first association relationship to transfer the three-dimensional attributes of the annotated traffic light to the target annotation frame to obtain third annotation data; determine the first coordinate data of the sample traffic light in the first coordinate system based on the third annotation data, wherein the first coordinate system is the coordinate system corresponding to the image acquisition component; use the first coordinate data and the internal parameters of the image acquisition component to determine the two-dimensional pixel difference of the sample traffic light on the sample image; correct the third annotation data based on the two-dimensional pixel difference to generate second annotation data.
[0135] Optionally, the analysis module 802 is also used to: use a target neural network model to perform feature extraction on the image to be identified to obtain an extraction result; perform multi-head prediction based on the extraction result to obtain a first prediction result, a second prediction result and a third prediction result, wherein the first prediction result is used to predict the angle range of the target angle, the second prediction result is used to predict the offset of the target angle within the angle range, and the third prediction result is used to predict the error of the current predicted size relative to the preset size, the target angle is the angle between the orientation of the target traffic light and the driving direction of the autonomous driving vehicle, and the preset size is determined by the target positioning parameter of the image acquisition component; use the first prediction result, the second prediction result and the third prediction result to determine the error information and angle information.
[0136] Optionally, the second acquisition module 803 is also used to: obtain a second association relationship between the first coordinate data of the target traffic light in the first coordinate system and the second coordinate data of the target traffic light in the second coordinate system, wherein the first coordinate system is the coordinate system corresponding to the image acquisition component, and the second coordinate system is the coordinate system corresponding to the image to be identified; obtain depth information based on the second association relationship, error information, angle information, size information and internal parameters of the image acquisition component.
[0137] Optionally, the second acquisition module 803 is also used to: calculate the first position of the target traffic light using error information, angle information and internal parameters of the image acquisition component, and determine the translation matrix corresponding to the target traffic light using size information, wherein the first position is the three-dimensional spatial center position of the target traffic light; calculate the second position based on the second association relationship, the first position and the translation matrix, wherein the second position is the two-dimensional spatial position corresponding to the first position in the image to be identified; and obtain depth information using the first position and the second position.
[0138] Optionally, the determination module 804 is further configured to: determine size information through error information, and determine orientation information through angle information; and determine the size information, orientation information, and depth information as three-dimensional attributes.
[0139] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0140] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, including a memory and at least one processor, wherein the memory stores computer instructions, and the processor is configured to execute the computer instructions to perform the steps in the above method embodiment.
[0141] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0142] Optionally, in the present disclosure, the processor may be configured to execute the following steps through a computer program:
[0143] S1, obtaining an image to be recognized, wherein the image to be recognized is obtained by an image acquisition component on the autonomous driving vehicle, and the display content of the image to be recognized includes: a target traffic light;
[0144] S2, using the target neural network model to analyze the image to be identified to obtain error information and angle information, wherein the target neural network model is used to estimate the size information and orientation information of the target traffic light, the error information is used to determine the size information, and the angle information is used to determine the orientation information;
[0145] S3, obtaining depth information of the target traffic light based on the error information, angle information, size information of the image to be recognized, and internal parameters of the image acquisition component;
[0146] S4, determining the three-dimensional attributes of the target traffic light through the error information, angle information and depth information.
[0147] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0148] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, wherein the non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are configured to execute the steps of the above method embodiment when running.
[0149] Optionally, in this embodiment, the non-transitory computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0150] S1, obtaining an image to be recognized, wherein the image to be recognized is obtained by an image acquisition component on the autonomous driving vehicle, and the display content of the image to be recognized includes: a target traffic light;
[0151] S2, using the target neural network model to analyze the image to be identified to obtain error information and angle information, wherein the target neural network model is used to estimate the size information and orientation information of the target traffic light, the error information is used to determine the size information, and the angle information is used to determine the orientation information;
[0152] S3, obtaining depth information of the target traffic light based on the error information, angle information, size information of the image to be recognized, and internal parameters of the image acquisition component;
[0153] S4, determining the three-dimensional attributes of the target traffic light through the error information, angle information and depth information.
[0154] Alternatively, in this embodiment, the non-transitory computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any suitable combination of the above. More specific examples of readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0155] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product. The program code for implementing the method embodiment of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0156] In the above embodiments of the present disclosure, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0157] In the several embodiments provided in the present disclosure, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0158] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0159] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0160] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0161] The above is only a preferred embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present disclosure. These improvements and modifications should also be regarded as within the scope of protection of the present disclosure.
Claims
1. A method for obtaining three-dimensional attributes of a target traffic light, comprising: Acquire an image to be recognized, wherein the image to be recognized is obtained by an image acquisition component on the autonomous driving vehicle, and the image to be recognized includes: a target traffic light; Analyzing the image to be identified using a target neural network model to obtain error information and angle information, wherein the target neural network model is used to estimate size information and orientation information of the target traffic light, the error information is used to determine the size information, and the angle information is used to determine the orientation information; Acquire depth information of the target traffic light based on the error information, the angle information, the size information of the image to be recognized, and the internal parameters of the image acquisition component; The three-dimensional properties of the target traffic light are determined using the error information, the angle information, and the depth information.
2. The method according to claim 1, wherein The method further comprises: Acquire a sample image, wherein the sample image is obtained by the image acquisition component, and the sample image includes: a sample traffic light; generating second annotation data based on a preset accuracy map, original positioning parameters of the image acquisition component, and first annotation data, wherein the preset accuracy map and the original positioning parameters are used to determine a first annotation box of the sample traffic light on the sample image, and the first annotation data is used to determine a second annotation box of the sample traffic light on the sample image; The initial neural network model is iteratively trained using the second labeled data to obtain the target neural network model.
3. The method according to claim 2, wherein: Generating the second annotation data based on the preset accuracy map, the original positioning parameters, and the first annotation data includes: Projecting the preset accuracy map and the original positioning parameters onto the sample image to obtain the first annotation frame, and annotating the sample image with the first annotation data to obtain the second annotation frame; Adjusting the original positioning parameters by a grid search method until the matching error between the first annotation frame and the second annotation frame is minimized, thereby obtaining target positioning parameters; The second annotation data is generated based on the target positioning parameters and the three-dimensional attributes of the annotated traffic light, wherein the annotated traffic light is a traffic light corresponding to the sample traffic light annotated on the preset accuracy map.
4. The method according to claim 3, wherein: Generating the second annotation data based on the target positioning parameters and the three-dimensional attributes of the annotated traffic light includes: Determining a first association relationship between the labeled traffic light and a target labeled frame based on the target positioning parameter, wherein the target labeled frame is a labeled frame on the sample image that is adapted to the sample traffic light; The first association relationship is used to transfer the three-dimensional attributes of the labeled traffic light to the target labeling frame to generate the second labeling data.
5. The method according to claim 4, wherein Using the first association relationship, transferring the three-dimensional attributes of the labeled traffic light to the target labeled frame to generate the second labeled data includes: Using the first association relationship, the three-dimensional attributes of the annotated traffic light are transferred to the target annotated frame to obtain third annotated data; Determining first coordinate data of the sample traffic light in a first coordinate system based on the third labeled data, wherein the first coordinate system is a coordinate system corresponding to the image acquisition component; Determine a two-dimensional pixel difference of the sample traffic light on the sample image using the first coordinate data and an internal parameter of the image acquisition component; The third labeled data is corrected based on the two-dimensional pixel difference to generate the second labeled data.
6. The method according to claim 1, wherein Analyzing the image to be recognized using the target neural network model to obtain the error information and the angle information includes: Using the target neural network model to perform feature extraction on the image to be identified to obtain an extraction result; Performing multi-head prediction based on the extraction result to obtain a first prediction result, a second prediction result, and a third prediction result, wherein the first prediction result is used to predict the angle range of the target angle, the second prediction result is used to predict the offset of the target angle within the angle range, and the third prediction result is used to predict the error of the current predicted size relative to a preset size, the target angle being the angle between the orientation of the target traffic light and the driving direction of the autonomous vehicle, and the preset size being determined by the target positioning parameter of the image acquisition component; The error information and the angle information are determined using the first prediction result, the second prediction result, and the third prediction result.
7. The method according to claim 1, wherein Acquiring the depth information of the target traffic light based on the error information, the angle information, the size information of the image to be recognized, and the internal parameter of the image acquisition component includes: Obtaining a second association relationship between first coordinate data of the target traffic light in a first coordinate system and second coordinate data of the target traffic light in a second coordinate system, wherein the first coordinate system is a coordinate system corresponding to the image acquisition component, and the second coordinate system is a coordinate system corresponding to the image to be recognized; The depth information is acquired based on the second association relationship, the error information, the angle information, the size information and the internal parameters of the image acquisition component.
8. The method according to claim 7, wherein: Acquiring the depth information based on the second association relationship, the error information, the angle information, the size information, and the internal parameter of the image acquisition component includes: Calculating a first position of the target traffic light using the error information, the angle information, and an internal parameter of the image acquisition component, and determining a translation matrix corresponding to the target traffic light using the size information, wherein the first position is a three-dimensional spatial center position of the target traffic light; Calculate a second position based on the second association relationship, the first position, and the translation matrix, wherein the second position is a two-dimensional spatial position corresponding to the first position in the image to be identified; The depth information is acquired using the first position and the second position.
9. The method according to claim 1, wherein Determining the three-dimensional attribute of the target traffic light by using the error information, the angle information, and the depth information includes: determining the size information based on the error information, and determining the orientation information based on the angle information; The size information, the orientation information, and the depth information are determined as the three-dimensional attributes.
10. A device for obtaining three-dimensional attributes of a target traffic light, comprising: A first acquisition module is configured to acquire an image to be recognized, wherein the image to be recognized is obtained by an image acquisition component on the autonomous driving vehicle, and the display content in the image to be recognized includes: a target traffic light; an analysis module, configured to analyze the image to be identified using a target neural network model to obtain error information and angle information, wherein the target neural network model is used to estimate size information and orientation information of the target traffic light, the error information is used to determine the size information, and the angle information is used to determine the orientation information; a second acquisition module, configured to acquire depth information of the target traffic light based on the error information, the angle information, the size information of the image to be recognized, and an internal parameter of the image acquisition component; A determination module is configured to determine the three-dimensional attributes of the target traffic light according to the error information, the angle information, and the depth information.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.
13. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method and device for identifying traffic lights
CN110390829A
Traffic signal lamp identification method and system, computing device and intelligent vehicle
CN111507210A