Sign detection method and apparatus, and vehicle

By combining the data of lidar and cameras in the vehicle, and cross-sensor methods are used to detect the identification card position, the problem of low accuracy of identification card position detection in the prior art is solved, and higher detection accuracy is achieved.

WO2025130617A1PCT designated stage expired Publication Date: 2025-06-26GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/136804
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-12-04
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the prior art, the position detection of the signboard is not very accurate, especially because the signboard size is small and the location is complex, which increases the detection difficulty.

Method used

The vehicle's lidar obtains the three-dimensional point cloud data of the target area, and combines the two-dimensional camera visual data obtained by the camera, and uses point cloud segmentation and visual detection frame matching methods to determine the position of the sign.

Benefits of technology

The accuracy of determining the position of the signboard is improved, and the problem of insufficient detection accuracy of a single sensor is overcome through the joint detection of lidar and vision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136804_26062025_PF_FP_ABST
    Figure CN2024136804_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a sign detection method and apparatus, and a vehicle. The method comprises: by means of a lidar of a vehicle, acquiring target point cloud data corresponding to a target area, and by means of a camera of the vehicle, acquiring camera visual data corresponding to the target area, wherein the target area comprises at least one sign; performing point cloud segmentation on the target point cloud data, so as to obtain at least one three-dimensional bounding box corresponding to the target point cloud data; acquiring at least one two-dimensional visual detection box corresponding to the camera visual data; and on the basis of the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box, obtaining the position of the at least one sign in the target area. In the present application, the position of the sign in the target area is obtained on the basis of three-dimensional point cloud data corresponding to the target area that is obtained by means of the lidar and two-dimensional camera visual data corresponding to the target area that is obtained by means of the camera, thereby improving the accuracy of determining the position of the sign.
Need to check novelty before this filing date? Find Prior Art

Description

Signage detection method, device, and vehicle

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 19, 2023, with application number 2023117608066 and application name “Signage plate detection method, device and vehicle”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of automobile technology, and more specifically, to a method and device for detecting a sign, and a vehicle. Background Art

[0003] With the advancement of science and technology and the improvement of people's living standards, the use of vehicles is becoming increasingly common, and vehicles are becoming increasingly versatile. Traffic sign detection is a key component of autonomous driving, and the demand for its accuracy is increasing. However, the accuracy of sign position detection in related technologies is limited. Summary of the Invention

[0004] In order to solve or partially solve the problems existing in the related art, the present application proposes a sign detection method, device and vehicle, which can improve the accuracy of determining the location of the sign.

[0005] In a first aspect, an embodiment of the present application provides a method for detecting a signboard, which is applied to a vehicle, and the method includes: obtaining target point cloud data corresponding to a target area through a laser radar of the vehicle, and obtaining camera vision data corresponding to the target area through a camera of the vehicle, wherein the target area includes at least one signboard; performing point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data; obtaining at least one two-dimensional visual detection box corresponding to the camera vision data; and obtaining the position of the at least one signboard included in the target area based on the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box.

[0006] In a second aspect, an embodiment of the present application further provides a signboard detection device, which is applied to a vehicle. The device includes: a target point cloud data acquisition module, a three-dimensional bounding box acquisition module, a two-dimensional visual detection box acquisition module, and a signboard position acquisition module. The target point cloud data acquisition module is used to acquire target point cloud data corresponding to a target area through the vehicle's laser radar, and acquire camera visual data corresponding to the target area through the vehicle's camera, wherein the target area includes at least one signboard; the three-dimensional bounding box acquisition module is used to perform point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data; the two-dimensional visual detection box acquisition module is used to acquire at least one two-dimensional visual detection box corresponding to the camera visual data; the signboard position acquisition module is used to obtain the position of the at least one signboard included in the target area based on the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box.

[0007] In a third aspect, embodiments of the present application further provide a vehicle comprising: one or more processors, a memory, and one or more applications. The one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to be executed to implement the method described in the first aspect above.

[0008] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which program code is stored. The program code can be called by a processor to execute the method described in the first aspect above.

[0009] The technical solution provided by the present application obtains target point cloud data corresponding to the target area through the vehicle's laser radar, and obtains camera vision data corresponding to the target area through the vehicle's camera, wherein the target area includes at least one signboard; performs point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data; obtains at least one two-dimensional vision detection box corresponding to the camera vision data; obtains the position of at least one signboard included in the target area based on the at least one three-dimensional bounding box and the at least one two-dimensional vision detection box, thereby obtaining the position of the signboard included in the target area based on the three-dimensional point cloud data corresponding to the target area obtained by the laser radar and the two-dimensional camera vision data corresponding to the target area obtained by the camera, and detects the signboard in a cross-sensor manner combining laser radar and vision, thereby improving the accuracy of determining the position of the signboard.

[0010] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other objects, features and advantages of the present application will become more apparent through a more detailed description of exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present application.

[0012] FIG1 is a schematic diagram showing a flow chart of a method for detecting a sign provided in an embodiment of the present application;

[0013] FIG2 is a schematic diagram showing a flow chart of a method for detecting a sign provided in an embodiment of the present application;

[0014] FIG3 is a schematic diagram showing a process of obtaining at least one three-dimensional bounding box according to an embodiment of the present application;

[0015] FIG4 is a schematic diagram showing a flow chart of a method for detecting a sign provided in an embodiment of the present application;

[0016] FIG5 is a schematic diagram showing a flow chart of a method for detecting a sign provided in an embodiment of the present application;

[0017] FIG6 shows a schematic flow chart of a method for detecting a sign provided in an embodiment of the present application;

[0018] FIG7 shows a schematic flow chart of a method for detecting a sign provided in an embodiment of the present application;

[0019] FIG8 shows a module block diagram of a sign detection device provided by an embodiment of the present application;

[0020] FIG9 shows a block diagram of a vehicle for executing a method for detecting a sign according to an embodiment of the present application;

[0021] FIG10 shows a storage unit for storing or carrying program codes for implementing the identification plate detection method according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although the accompanying drawings illustrate embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0023] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0024] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0025] With the development of science and technology, autonomous driving technology has emerged. Among them, sign detection is an important part of autonomous driving technology. Therefore, people have higher and higher requirements for sign detection.

[0026] Due to the small size of the sign, it is difficult to determine the location of the sign. In the related art, a solution of projecting a visual two-dimensional detection frame onto a three-dimensional scene to determine the location of the sign and a solution of determining the location of the sign based on deep learning of LiDAR have been proposed. Among them, the solution of projecting a visual two-dimensional detection frame onto a three-dimensional scene to determine the location of the sign has extremely high requirements for whether the detected scene is horizontal and the quality of the two-dimensional detection. Among them, since the semantic parts of traffic signs are all non-ground, and a prerequisite for two-dimensional distance detection is that the detected object is close to the ground, the use of visual projection method cannot obtain the true location of the sign. The pure deep learning solution based on LiDAR requires the use of a more powerful graphics processor, has high requirements for the deployed hardware and the amount of labeled data, and has a high development cost. In addition, due to the segmentation solution based on LiDAR, it is impossible to provide the category information of the sign and it is easy to cause missed detection due to under-segmentation.

[0027] Therefore, in the related art, the position detection of the sign has the problem of low accuracy.

[0028] In response to the above problems, the inventors discovered after long-term research and proposed the signboard detection method, device and vehicle provided in the embodiments of the present application. The positions of the signboards included in the target area are obtained based on the three-dimensional point cloud data corresponding to the target area obtained by the laser radar and the two-dimensional camera vision data corresponding to the target area obtained by the camera. The signboards are detected in a cross-sensor manner combining the laser radar and vision, thereby improving the accuracy of determining the position of the signboards.

[0029] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0030] Please refer to Figure 1, which shows a schematic flow chart of a sign detection method provided by an embodiment of the present application. In a specific embodiment, the sign detection method can be applied to a sign detection device 200 as shown in Figure 8 and a vehicle 100 (Figure 9) equipped with the sign detection device 200. The following will take a vehicle as an example to illustrate the specific process of this embodiment. Of course, it can be understood that the vehicle used in this embodiment may include electric vehicles, gasoline vehicles and other electronic devices with processing capabilities, which are not limited here. The following will be a detailed explanation of the process shown in Figure 1. The sign detection method may include the following steps:

[0031] Step S110: acquiring target point cloud data corresponding to a target area through the laser radar of the vehicle, and acquiring camera vision data corresponding to the target area through the camera of the vehicle, wherein the target area includes at least one identification plate.

[0032] In some embodiments, the vehicle may include a laser radar, such as a mechanical scanning type, a semi-solid type, etc. The vehicle may use the laser radar to obtain point cloud data of the environment in which the vehicle is located in real time.

[0033] In some embodiments, a vehicle may include one or more cameras, wherein the vehicle may obtain real-time camera visual data of the vehicle's environment through the cameras, wherein the camera visual data may include one or more images.

[0034] The target area may include at least one sign. The target area may be understood as the vehicle's driving environment or the illumination range of the vehicle's laser radar. The driving environment may include at least one sign. The sign may include a prohibition sign, a warning sign, an instruction sign, etc., without limitation.

[0035] In some embodiments, a vehicle may receive a control command input by a user, which may be used to instruct the vehicle to perform autonomous driving. Accordingly, the vehicle may, in response to the control command, acquire target point cloud data corresponding to a target area via its lidar and acquire camera vision data corresponding to the target area via its camera.

[0036] In some embodiments, to improve the accuracy of the vehicle's acquisition of the position and pose of a target area, including objects, in this embodiment, the vehicle can perform motion compensation on the point cloud data of the target area acquired by the LiDAR to obtain target point cloud data, thereby improving the usability of the target point cloud data. Optionally, different LiDARs included in the vehicle may result in different methods for performing motion compensation on the point cloud data acquired by the LiDAR to obtain target point cloud data.

[0037] Exemplarily, the laser radar included in the vehicle can be a solid-state laser radar. Accordingly, the vehicle can obtain the first point cloud data of the target area obtained by the laser radar in the current frame and the second point cloud data of the target area obtained by the laser radar in the previous frame of the current frame, and can perform motion compensation on the first pose of the object corresponding to the first point cloud data based on the second pose of the object corresponding to the second point cloud data to obtain target point cloud data, so as to improve the availability of the target point cloud data.

[0038] Exemplarily, the laser radar included in the vehicle can be a circling laser radar. Accordingly, the vehicle can obtain a preset number of frames of multiple point cloud data of the target area obtained by the laser radar, and can match the multiple point cloud data to obtain target point cloud data to improve the availability of the target point cloud data.

[0039] Step S120: performing point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data.

[0040] In some embodiments, after obtaining target point cloud data corresponding to a target area, the vehicle may perform point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data. The vehicle may be pre-configured with a preset segmentation algorithm, such as an algorithm based on plane fitting, an algorithm based on the characteristics of laser point cloud data, a segmentation algorithm, a recursive greedy algorithm, a Hungarian algorithm, and the like, without limitation herein.

[0041] Exemplarily, the vehicle may use the Hungarian algorithm to segment the target point cloud data; wherein, if the vehicle considers the under-segmentation of the target point cloud data, the vehicle may use a recursive greedy algorithm to segment the target point cloud data.

[0042] The vehicle may perform point cloud segmentation on the target point cloud data to obtain at least one irregularly shaped three-dimensional bounding box corresponding to the target point cloud data.

[0043] Step S130: Obtain at least one two-dimensional visual detection frame corresponding to the camera visual data.

[0044] In some embodiments, after obtaining camera visual data corresponding to a target area, the vehicle may obtain at least one two-dimensional visual detection frame corresponding to the camera visual data. The camera visual data may include at least one image. The at least one image may include at least one object. Accordingly, the vehicle may perform motion compensation on the object included in the at least one image based on the at least one image to obtain a target image. Accordingly, the vehicle may perform target detection on the at least one object included in the target image to obtain at least one two-dimensional visual detection frame.

[0045] The vehicle can process the camera visual data based on a target detection algorithm to obtain at least one two-dimensional visual detection frame corresponding to the camera visual data. The target detection algorithm may include a 2D perception algorithm, a YOLO algorithm, etc., which are not limited here. The at least one two-dimensional visual detection frame may include objects such as signboards, static obstacles, and dynamic obstacles, which are not limited here. Dynamic obstacles include but are not limited to motor vehicles and pedestrians; static obstacles include but are not limited to lane markings, parking spaces, drivable areas, and general obstacles.

[0046] Optionally, the vehicle can perform sign detection on the camera visual data and obtain at least one corresponding two-dimensional visual detection frame; wherein each two-dimensional visual detection frame can include a sign. The vehicle can obtain the geographic location of the target area and perform sign detection on the camera visual data based on the geographic location and obtain at least one corresponding two-dimensional visual detection frame. It can be understood that the shapes of signboards in different countries can be the same or different; wherein, the vehicle can determine the country where the signboard is located based on the geographic location of the target area, and then determine the shape of the signboard corresponding to the country, thereby performing sign detection on the camera visual data and obtaining at least one corresponding two-dimensional visual detection frame, thereby improving the accuracy of visual signboard detection.

[0047] The vehicle can also obtain the weather conditions of the target area and perform sign detection on the camera visual data based on the weather conditions, thereby obtaining at least one corresponding two-dimensional visual detection frame. It is understood that the reflective color of the sign in different weather conditions can be the same or different. The vehicle can determine the weather conditions of the sign based on the weather conditions of the target area, and then determine the reflective color of the sign corresponding to that weather, thereby performing sign detection on the camera visual data and obtaining at least one corresponding two-dimensional visual detection frame, thereby improving the accuracy of visual sign detection.

[0048] Step S140: Obtaining the position of the at least one sign included in the target area according to the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box.

[0049] In some embodiments, after the vehicle obtains at least one three-dimensional bounding box corresponding to the target area and at least one two-dimensional visual detection frame corresponding to the target area, it can obtain the position of at least one sign included in the target area based on the at least one three-dimensional bounding box and the at least one two-dimensional visual detection frame.

[0050] The vehicle can project at least one 3D bounding box corresponding to the target area into 2D space and associate it with the at least one 2D visual detection box. Alternatively, the vehicle can determine as the box to be verified a 3D bounding box in 2D space that overlaps with the 2D visual detection box; or it can determine as the box to be verified a 3D bounding box in 2D space that is within a preset distance range from the 2D visual detection box.

[0051] Accordingly, the vehicle can detect the attributes of the object corresponding to the two-dimensional visual detection frame corresponding to the frame to be verified. If the attribute of the object detected is a sign, the position of the frame to be verified can be obtained and the position of the sign can be determined as the position of the frame to be verified. The position of the frame to be verified can be obtained by the position of the three-dimensional bounding box corresponding to the frame to be verified.

[0052] In some embodiments, the number of three-dimensional bounding boxes corresponding to the frame to be verified may be one or more. Optionally, when the number of three-dimensional bounding boxes corresponding to the frame to be verified is one, the position of the frame to be verified may be the position of the three-dimensional bounding box; when the number of three-dimensional bounding boxes corresponding to the frame to be verified is multiple, the position of the frame to be verified may be the position of any one of the multiple three-dimensional bounding boxes, such as the position of the three-dimensional bounding box with the largest projected area in the two-dimensional space, or the position of the three-dimensional bounding box with the smallest projected area in the two-dimensional space. When the number of three-dimensional bounding boxes corresponding to the frame to be verified is multiple, the vehicle may also fuse the multiple three-dimensional bounding boxes to obtain a fused three-dimensional bounding box. Accordingly, the vehicle may determine the position of the fused three-dimensional bounding box as the position of the frame to be verified.

[0053] In some embodiments, the vehicle may also obtain the type of the sign, such as "No Passing" or "No Parking," when determining that the object's attribute is a sign. Accordingly, the vehicle may obtain the type of the sign when obtaining the location of the sign, and output the location and type of the sign to enable the vehicle to plan a driving route based on the location and type of the sign, thereby improving the rationality and safety of the vehicle's autonomous or assisted driving and enhancing the user experience.

[0054] In some implementations, referring to FIG. 2 , the identification plate detection method provided in an embodiment of the present application may further include steps S141 to S143 before step S140 .

[0055] Step S141: performing obstacle detection on the target point cloud data to obtain an obstacle frame corresponding to the target point cloud data.

[0056] In some embodiments, considering that the target point cloud data corresponding to the target area obtained by the vehicle through the laser radar may include at least one signboard included in the target area, static obstacles included in the target area (such as lane lines, parking spaces, drivable areas, general obstacles, etc.), dynamic obstacles (such as motor vehicles, pedestrians, etc.), and other objects, in order to reduce the amount of calculation for the vehicle to detect the signboard, in this embodiment, after the vehicle obtains the target point cloud data, it can perform obstacle detection on the target point cloud data to obtain an obstacle frame corresponding to the target point cloud data, and the vehicle can filter the target point cloud data based on the obstacle frame, thereby reducing the amount of point cloud data, reducing the amount of calculation for the position of the signboard in the target area detected by the vehicle, and reducing the power consumption of the vehicle in detecting the signboard.

[0057] The vehicle can detect obstacles in the target point cloud data based on deep learning. The deep learning method can include detecting obstacles in the target point cloud data based on convolutional neural networks, recurrent neural networks, etc., and obtaining obstacle boxes corresponding to the target point cloud data.

[0058] For example, the vehicle can pre-process the target point cloud data by filtering and removing outliers at the edge of the distribution. It can then perform obstacle detection on the pre-processed target point cloud data based on a deep learning model to obtain an obstacle box corresponding to the target point cloud data. Furthermore, the vehicle can perform ground point cloud segmentation on the target point cloud data and cluster the target point cloud data using a clustering algorithm to obtain multiple clusters, each of which can be used to represent an obstacle. Accordingly, the vehicle can fit a bounding box to each cluster to obtain an obstacle box corresponding to the target point cloud data.

[0059] Step S142: Obtain an obstacle fusion result corresponding to the target point cloud data according to the obstacle frame, the at least one three-dimensional bounding box, and the camera vision data.

[0060] In some embodiments, after the vehicle obtains the obstacle frame corresponding to the target point cloud data, at least one bounding box corresponding to the target point cloud data, and the camera vision data corresponding to the target area, it can obtain the obstacle fusion result corresponding to the target point cloud data based on the obstacle frame, the at least one three-dimensional bounding box, and the camera vision data.

[0061] The vehicle can project at least one 3D bounding box corresponding to the target point cloud data obtained through segmentation and the obstacle box obtained through deep learning into 2D space to obtain a 2D bounding box to be fused. Accordingly, the vehicle can fuse this 2D bounding box with the camera visual data to obtain an obstacle fusion result corresponding to the target point cloud data.

[0062] It will be appreciated that in this embodiment, the type of three-dimensional data in the target point cloud data is determined by projecting the 3D point cloud data into a two-dimensional space and fusing it with the two-dimensional camera vision data. Furthermore, the vehicle can obtain depth information of objects included in the three-dimensional point cloud data from the point cloud data. Based on this, the vehicle obtains the obstacle fusion result corresponding to the target point cloud data and subsequently filters out the point cloud data corresponding to obstacles in the target area included in the target point cloud data from the target point cloud data. This reduces the computational effort required for the vehicle to perform sign detection on the target point cloud data, as well as the amount of data corresponding to the target point cloud data. This saves the vehicle's storage space and improves the accuracy of the vehicle's sign detection.

[0063] Step S143: filtering the at least one three-dimensional bounding box according to the obstacle fusion result to obtain at least one filtered three-dimensional bounding box.

[0064] In some embodiments, after obtaining the obstacle fusion result corresponding to the target point cloud data, the vehicle can filter at least one three-dimensional bounding box corresponding to the target area according to the obstacle fusion result to obtain at least one filtered three-dimensional bounding box.

[0065] In some embodiments, the vehicle can obtain ground point cloud data included in the target point cloud data; this ground point cloud data can be understood as point cloud data returned by a laser radar (LiDAR) illuminating the ground. Accordingly, after obtaining the ground point cloud data included in the target point cloud data, the vehicle can filter the obstacle fusion results in the target point cloud data from at least one 3D bounding box in combination with the ground point cloud data, thereby obtaining at least one filtered 3D bounding box to improve the speed and accuracy of vehicle identification plate detection.

[0066] In some embodiments, the vehicle may obtain ground point cloud data included in the target point cloud data while obtaining an obstacle fusion result corresponding to the target point cloud data based on the obstacle box, at least one three-dimensional bounding box, and camera visual data. Accordingly, after obtaining the ground point cloud data included in the target point cloud data, the vehicle may combine the point cloud data in the target point cloud data that is not part of the obstacle fusion result with the ground point data based on the obstacle fusion result and the ground point cloud data to obtain preliminary screening point cloud data.

[0067] Accordingly, the vehicle can project the preliminary screening point cloud data included in the target point cloud data (including ground point cloud data and point cloud data that does not belong to the obstacle fusion result) into a two-dimensional space to obtain a two-dimensional frame to be verified, and can match the frame to be verified with at least one visual detection frame corresponding to the camera visual data to obtain the position and type of at least one signboard included in the target area.

[0068] For example, please refer to Figure 3, which shows a schematic diagram of the process for obtaining at least one three-dimensional bounding box provided by an embodiment of the present application. The vehicle can obtain point cloud data of a target area acquired in real time by a lidar, and can use the relative pose of the object included in the target area in the point cloud data of the previous frame as motion compensation for the pose of the object included in the target area in the point cloud data of the current frame to obtain the target point cloud data, thereby improving the usability of the point cloud data.

[0069] After obtaining the target point cloud data, the vehicle can perform obstacle detection on the target point cloud data based on a pre-set deep learning algorithm to obtain the obstacle frame corresponding to the target point cloud data. After obtaining the target point cloud data, the vehicle can obtain the point cloud data returned by illuminating the ground with a laser radar, that is, obtain the ground point cloud data within the target point cloud data.

[0070] The vehicle may also segment the target point cloud data using a segmentation algorithm to obtain at least one three-dimensional bounding box corresponding to the target point cloud data.

[0071] The vehicle can obtain real-time camera visual data of the target area captured by the camera, and can obtain at least one two-dimensional visual detection frame corresponding to the camera visual data. The vehicle can also project the obstacle frame obtained by deep learning and at least one three-dimensional bounding frame obtained by segmenting the target point cloud data into two-dimensional space, match and fuse them with the at least one two-dimensional visual detection frame to obtain an obstacle fusion result corresponding to the target point cloud data, and filter the at least one three-dimensional bounding frame based on the obstacle fusion result to obtain at least one filtered three-dimensional bounding frame.

[0072] In some implementations, referring to FIG. 4 , the identification plate detection method provided in an embodiment of the present application may further include steps S144 and S145 after step S140 .

[0073] Step S144: Obtain the target two-dimensional visual detection frame corresponding to each of the at least one signboard.

[0074] In some embodiments, after obtaining the position of at least one sign included in the target area, the vehicle may obtain a two-dimensional visual detection frame corresponding to each of the at least one sign, and may determine the two-dimensional visual detection frame as the target two-dimensional visual detection frame.

[0075] Step S145: Obtaining the type corresponding to each of the at least one signboard according to the camera visual data corresponding to the target two-dimensional visual detection frame.

[0076] In some embodiments, after obtaining a target two-dimensional visual detection frame corresponding to at least one sign, the vehicle can obtain the type of the at least one sign based on the camera visual data corresponding to the target two-dimensional visual detection frame. The vehicle can obtain the type of the sign included in the target two-dimensional visual detection frame by detecting the camera visual data corresponding to the target two-dimensional visual detection frame using an image recognition algorithm.

[0077] In some embodiments, after obtaining camera visual data, the vehicle may process the camera visual data based on an image recognition algorithm to obtain a camera sign perception result, that is, the type of at least one sign included in the target area.

[0078] Accordingly, after determining the position of at least one signboard included in the target area, the vehicle can obtain the type of signboard included in the pre-detected target two-dimensional visual detection frame based on the target two-dimensional visual detection frame corresponding to at least one signboard, thereby utilizing the obstacle results of visual detection and the lidar segmentation results to obtain reliable position depth information of the signboard through the fusion of directional lidar and vision, thereby reducing the cost of signboard detection and improving the accuracy of signboard detection.

[0079] A signboard detection method provided by an embodiment of the present application obtains target point cloud data corresponding to a target area through a vehicle's laser radar, and obtains camera vision data corresponding to the target area through a vehicle's camera, wherein the target area includes at least one signboard; performs point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data; obtains at least one two-dimensional vision detection box corresponding to the camera vision data; obtains the position of at least one signboard included in the target area based on the at least one three-dimensional bounding box and the at least one two-dimensional vision detection box, thereby obtaining the position of the signboard included in the target area based on the three-dimensional point cloud data corresponding to the target area obtained through the laser radar and the two-dimensional camera vision data corresponding to the target area obtained through the camera, and detects the signboard in a combined laser radar and vision manner, thereby improving the accuracy of determining the position of the signboard.

[0080] Please refer to Figure 5, which shows a flow chart of a method for detecting a sign provided by an embodiment of the present application. The method is applied to the above-mentioned vehicle. The flow chart shown in Figure 5 will be described in detail below. The method for detecting a sign may include the following steps:

[0081] Step S210: acquiring target point cloud data corresponding to a target area through the laser radar of the vehicle, and acquiring camera vision data corresponding to the target area through the camera of the vehicle, wherein the target area includes at least one identification plate.

[0082] Step S220: performing point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data.

[0083] Step S230: Obtain at least one two-dimensional visual detection frame corresponding to the camera visual data.

[0084] For a detailed description of steps S210 to S230 , please refer to the above description of steps S110 to S130 , which will not be repeated here.

[0085] Step S240: Projecting the at least one three-dimensional bounding box into a two-dimensional space to obtain a two-dimensional box corresponding to each of the at least one three-dimensional bounding box.

[0086] In some embodiments, after obtaining at least one 3D bounding box corresponding to the target point cloud data, the vehicle may project the at least one 3D bounding box into a 2D space to obtain a 2D box corresponding to each of the at least one 3D bounding boxes. The vehicle may project the at least one 3D bounding box into a camera coordinate system based on a camera coordinate system corresponding to a camera of the vehicle to obtain a 2D box corresponding to each of the at least one 3D bounding boxes.

[0087] Optionally, the vehicle can project the points included in the three-dimensional enclosing box to two-dimensional space to obtain a two-dimensional box corresponding to the three-dimensional enclosing box; the vehicle can also project the points of the bounding box corresponding to the three-dimensional enclosing box to two-dimensional space to obtain a two-dimensional box corresponding to the three-dimensional enclosing box; the vehicle can also project the vertices of the bounding box corresponding to the three-dimensional enclosing box to two-dimensional space to obtain a two-dimensional box corresponding to the three-dimensional enclosing box.

[0088] Step S250: Matching the two-dimensional frame corresponding to each of the at least one three-dimensional bounding frame with the at least one two-dimensional visual detection frame to obtain a matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding frame.

[0089] In some embodiments, after the vehicle obtains the two-dimensional frame corresponding to at least one three-dimensional enclosing frame, it matches the two-dimensional frame corresponding to the at least one three-dimensional enclosing frame with at least one two-dimensional visual detection frame to obtain a matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional enclosing frame.

[0090] Among them, the vehicle can match at least one two-dimensional visual detection frame in the camera coordinate system with the two-dimensional frame corresponding to the at least one three-dimensional enclosing frame obtained by projecting the at least one three-dimensional enclosing frame to the camera coordinate system, so as to improve the accuracy of signboard position detection by obtaining the position of at least one signboard included in the target area in a cross-sensor manner (combining the three-dimensional object detected by the lidar and the two-dimensional frame detected visually).

[0091] In some embodiments, the vehicle matches the two-dimensional frame corresponding to at least one three-dimensional bounding box with at least one two-dimensional visual detection frame. The process of obtaining the matching relationship between at least one two-dimensional visual detection frame and at least one three-dimensional bounding box may include obtaining the overlapping area of ​​the at least one two-dimensional visual detection frame and the two-dimensional frame corresponding to the at least one three-dimensional bounding box, and determining the matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding box based on the overlapping area.

[0092] Among them, the vehicle can obtain the area of ​​the overlapping area between at least one two-dimensional visual detection frame and the two-dimensional frame corresponding to at least one three-dimensional enclosing frame. If the area is greater than the overlapping area threshold, it can be determined that the two-dimensional visual detection frame with the overlapping area and the three-dimensional enclosing frame are in a matching relationship. Accordingly, the vehicle can obtain the area of ​​the overlapping area between the two-dimensional visual detection frame and the two-dimensional frame corresponding to at least one three-dimensional enclosing frame, and determine the matching relationship between the two-dimensional visual detection frame and the at least one three-dimensional enclosing frame based on the number of overlapping areas with an area greater than the overlapping area threshold.

[0093] In some embodiments, considering the small area of ​​the signboard, the vehicle can obtain the area of ​​the two-dimensional frame corresponding to each of the at least one three-dimensional enclosing frames, and match the two-dimensional frame corresponding to each of the at least one three-dimensional enclosing frames with the at least one two-dimensional visual detection frame based on a recursive matching method in order of the area of ​​the two-dimensional frames from small to large, to obtain a matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional enclosing frame.

[0094] Among them, the vehicle can also obtain the radius of the two-dimensional frame corresponding to at least one three-dimensional enclosing frame, and match the two-dimensional frame corresponding to at least one three-dimensional enclosing frame with at least one two-dimensional visual detection frame based on recursive matching in the order of the radius of the two-dimensional frame from small to large, to obtain the matching relationship between at least one two-dimensional visual detection frame and the at least one three-dimensional enclosing frame.

[0095] As an implementable method, the process of a vehicle determining the matching relationship between at least one two-dimensional visual detection frame and at least one three-dimensional bounding box based on the overlapping area may include: if it is determined based on the overlapping area that there is an area overlap between the fourth three-dimensional bounding box in the at least one three-dimensional bounding box and the fourth two-dimensional visual detection frame in the at least one two-dimensional visual detection, then it can be determined that the matching relationship between the fourth three-dimensional bounding box and the fourth two-dimensional visual detection frame is a one-to-one relationship.

[0096] As an implementable method, the process of a vehicle determining the matching relationship between at least one two-dimensional visual detection frame and at least one three-dimensional bounding box based on the overlapping area may include: if it is determined based on the overlapping area that multiple fifth three-dimensional bounding boxes in the at least one three-dimensional bounding box have an area overlap with the fifth two-dimensional visual detection frame in at least one two-dimensional visual detection, then it can be determined that the matching relationship between the multiple fifth three-dimensional bounding boxes and the fifth two-dimensional visual detection frame is a many-to-one relationship.

[0097] As an implementable method, the process of a vehicle determining the matching relationship between at least one two-dimensional visual detection frame and at least one three-dimensional bounding frame based on the overlapping area may include: if it is determined based on the overlapping area that there is no area overlap between the sixth three-dimensional bounding box in the at least one three-dimensional bounding box and the sixth two-dimensional visual detection frame of the at least one three-dimensional bounding box, it can be determined that the sixth three-dimensional bounding box does not match the sixth two-dimensional visual detection frame.

[0098] As an implementable method, the process of a vehicle determining the matching relationship between at least one two-dimensional visual detection frame and at least one three-dimensional enclosing frame based on the overlapping area may include: if it is determined based on the overlapping area that there are multiple seventh two-dimensional visual detection frames including identification plates in at least one two-dimensional visual detection frame and there is an overlapping area with the seventh three-dimensional enclosing frame in at least one three-dimensional enclosing frame, then it can be determined that the matching relationship between the seventh three-dimensional enclosing frame and the multiple seventh two-dimensional visual detection frames including identification plates is a one-to-many relationship.

[0099] Step S260: obtaining the position of the at least one sign included in the target area based on the matching relationship.

[0100] As an implementable method, after the vehicle obtains a matching relationship between at least one two-dimensional visual detection frame and at least one three-dimensional bounding frame, it can obtain the position of at least one signboard included in the target area based on the matching relationship. If the vehicle detects that the matching relationship is a one-to-one relationship, it can determine the three-dimensional bounding frame that is in a one-to-one relationship with the at least one two-dimensional visual detection frame from the at least one three-dimensional bounding frame as the first three-dimensional bounding frame, and can determine the two-dimensional visual detection frame that is in a one-to-one relationship with the first three-dimensional bounding frame as the first two-dimensional visual detection frame. If the vehicle detects that the first two-dimensional visual detection frame includes a signboard, it can obtain the first position corresponding to the first three-dimensional bounding frame, and can determine the first position as the position of the signboard included in the first two-dimensional visual detection frame.

[0101] Considering the relatively small area of ​​the sign, in this embodiment, the vehicle can determine, from the at least one three-dimensional bounding frame, a three-dimensional bounding frame that forms a one-to-one relationship with the at least one two-dimensional visual detection frame as the frame to be detected. If the vehicle detects that the area of ​​the frame to be detected is less than an area threshold, the frame to be detected can be determined as the first three-dimensional bounding frame.

[0102] As an implementable approach, if the vehicle detects that the matching relationship is many-to-one, multiple three-dimensional bounding boxes that form a many-to-one relationship with one of the at least one two-dimensional visual detection frames can be determined from the at least one three-dimensional bounding box, and the multiple three-dimensional bounding boxes can be fused to obtain a second three-dimensional bounding box. The two-dimensional visual detection frame in the at least one two-dimensional visual detection frame that forms a many-to-one relationship with the second three-dimensional bounding box can be determined as the second two-dimensional visual detection frame. If the vehicle detects that the second two-dimensional visual detection frame includes a sign, a second position corresponding to the second three-dimensional bounding box can be obtained, and the second position can be determined as the position of the sign included in the second two-dimensional visual detection frame.

[0103] The vehicle may fuse the multiple 3D bounding frames to obtain the second 3D bounding frame, which may include dilating the multiple 3D bounding frames until they are tangent to each other, and obtaining a single 3D bounding frame formed by dilating the multiple 3D bounding frames as the second 3D bounding frame. Alternatively, the vehicle may determine the second 3D bounding frame as the three-dimensional bounding frame with the largest corresponding 2D area among the multiple 3D bounding frames; or the vehicle may determine the second 3D bounding frame as the three-dimensional bounding frame with the smallest corresponding 2D area among the multiple 3D bounding frames.

[0104] As an implementable method, if the vehicle detects that the matching relationship is a one-to-many relationship, a third three-dimensional bounding box that is in a one-to-many relationship with multiple third two-dimensional visual detection frames in at least one two-dimensional visual detection frame can be determined from at least one three-dimensional bounding box, wherein each third two-dimensional visual detection frame includes a sign. Accordingly, the vehicle can segment the third three-dimensional bounding box based on the multiple third two-dimensional visual detection frames, obtain the target third three-dimensional bounding box corresponding to each of the multiple third two-dimensional visual detection frames, and obtain the third position of the target third three-dimensional bounding box, and use the third position as the position of the sign included in the multiple third two-dimensional visual detection frames. Among them, the corresponding multiple third two-dimensional visual detection frames can be understood as under-segmented objects.

[0105] In some embodiments, considering that the amount of computation required to determine the location of a signboard based on a one-to-one relationship is relatively small, in this embodiment, the vehicle may prioritize detecting the location of the signboard on a matching pair of a three-dimensional bounding box and a two-dimensional visual detection box that have a one-to-one matching relationship. If the vehicle detects that the two-dimensional visual detection frame in the matching pair includes the signboard, the position corresponding to the three-dimensional bounding box in the matching pair may be obtained, and the position may be determined as the position of the signboard corresponding to the matching pair. In addition, considering that the area of ​​the signboard is relatively small, the vehicle may obtain a two-dimensional visual detection frame that matches the three-dimensional bounding box based on the area of ​​the two-dimensional frame corresponding to the three-dimensional bounding box in ascending order, and perform signboard location detection to improve the efficiency and accuracy of signboard location detection.

[0106] In some implementations, referring to FIG. 6 , the identification plate detection method provided in an embodiment of the present application may further include steps S261 and S262 after step S260 .

[0107] Step S261: If it is detected that there is a two-dimensional visual detection frame in the at least one two-dimensional visual detection frame that does not match the at least one three-dimensional bounding frame, then the ground point cloud data corresponding to the target point cloud data is obtained.

[0108] In some embodiments, if the vehicle detects that a two-dimensional visual inspection box within at least one two-dimensional visual inspection box does not match the at least one three-dimensional bounding box, the vehicle may obtain ground point cloud data corresponding to the target point cloud data. The vehicle may perform ground point clustering on the target point cloud data to obtain the ground point cloud data.

[0109] Step S262: If it is detected that the unmatched two-dimensional visual detection frame matches the ground point cloud data, the unmatched two-dimensional visual detection frame is filtered out from the at least one two-dimensional visual detection frame to obtain at least one updated two-dimensional visual detection frame.

[0110] In some embodiments, if the vehicle detects that the unmatched two-dimensional visual detection frame matches the ground point cloud data, the unmatched two-dimensional visual detection frame can be filtered out from at least one two-dimensional visual detection frame to obtain at least one updated two-dimensional visual detection frame, so as to increase the rate of obtaining the matching relationship between at least one two-dimensional visual detection frame and at least one three-dimensional bounding box.

[0111] The vehicle can calculate the area of ​​the overlap between the unmatched 2D visual inspection frame and the ground point cloud data, and compare this area with a ground matching area threshold. If the area is greater than or equal to the ground matching area threshold, the vehicle can determine that the unmatched 2D visual inspection frame matches the ground point cloud data. Accordingly, the vehicle can determine that the unmatched 2D visual inspection frame that matches the ground point cloud data is a falsely detected 2D visual inspection frame; accordingly, the vehicle can delete the unmatched 2D visual inspection frame that matches the ground point cloud data to save vehicle storage space.

[0112] For example, please refer to Figure 7, which shows a flow chart of a sign detection method provided by an embodiment of the present application. The vehicle can obtain ground point cloud data based on the target point cloud data, which can be understood as unmarked used LiDAR ground points. The vehicle can also obtain at least one three-dimensional bounding box after LiDAR segmentation. The vehicle can also fuse the obstacle box obtained from the deep learning target point cloud data, at least one three-dimensional bounding box, and the camera visual data to obtain an obstacle fusion result. The vehicle can also perform traffic sign detection on the camera visual data to obtain visual traffic signs in the target area.

[0113] Among them, the vehicle can combine the point cloud data that does not belong to the obstacle fusion result in the target point cloud data determined after deep learning with the ground point cloud data corresponding to the target point cloud data and the obstacle fusion result corresponding to the target area to obtain preliminary screening point cloud data, and obtain at least one three-dimensional bounding box based on the preliminary screening point cloud data.

[0114] If the vehicle does not obtain an obstacle frame based on a deep learning algorithm, it can also project at least one 3D bounding box corresponding to the target point cloud data into 2D space. Accordingly, the vehicle can also match the 2D box corresponding to each of the at least one 3D bounding box with at least one 2D visual detection box to obtain a matching relationship between the at least one 2D visual detection box and the at least one 3D bounding box.

[0115] The vehicle can perform sign position detection on a matching pair of a three-dimensional bounding box and a two-dimensional visual detection box that have a one-to-one matching relationship. If the vehicle detects that the two-dimensional visual detection box in the matching pair includes a sign based on visual traffic signs in the target area, the vehicle can obtain the position corresponding to the three-dimensional bounding box in the matching pair and determine that position as the position of the sign corresponding to the matching pair.

[0116] Among them, the vehicle can recursively match the three-dimensional enclosing box with at least one two-dimensional visual detection frame in order from small to large radius based on the radius of the two-dimensional frame corresponding to the three-dimensional enclosing box to obtain a matching relationship, and obtain the position of at least one signboard included in the target area based on the matching relationship.

[0117] If the vehicle detects a one-to-many matching relationship, it can determine, from the at least one 3D bounding box, a third 3D bounding box that forms a one-to-many relationship with multiple third 2D visual detection boxes in the at least one 2D visual detection box, where each third 2D visual detection box includes a sign. Accordingly, the vehicle can segment the third 3D bounding box based on the multiple third 2D visual detection boxes to obtain target third 3D bounding boxes corresponding to each of the multiple third 2D visual detection boxes. The vehicle can also obtain a third position of the target third 3D bounding box and use the third position as the position of the sign included in the multiple third 2D visual detection boxes to process the under-segmented object.

[0118] The vehicle may also project ground point cloud data into a two-dimensional space and obtain at least one two-dimensional visual detection frame that does not match at least one three-dimensional bounding frame. The vehicle may cluster the ground points of the lidar in the two-dimensional space and match the clustered ground points with the unmatched two-dimensional visual detection frames. If the unmatched two-dimensional visual detection frame is detected to match the clustered ground points, the unmatched two-dimensional visual detection frame may be determined to be a falsely detected visual frame. The vehicle may filter out the unmatched two-dimensional visual detection frame to save storage space of the vehicle, increase the rate of determining the match between at least one three-dimensional bounding frame and the updated two-dimensional visual detection frame, and increase the rate of determining the location of the sign.

[0119] Among them, the vehicle can obtain the target two-dimensional visual detection frame corresponding to at least one sign, and can obtain the type corresponding to at least one sign based on the camera visual data corresponding to the target two-dimensional visual detection frame.

[0120] Among them, the vehicle can use the target point cloud data, camera vision data, and obstacle fusion results, and use the internal and external parameters of the camera and lidar in the vehicle coordinate system to obtain the position and type of traffic signs within the illumination range of the lidar sensor through projection and multi-level matching. Therefore, in the cross-sensor sign detection method that combines lidar and vision, the position of the sign is obtained through a multi-level filtering method with less computational complexity and less dependence on the scene, thereby increasing the accuracy of sign detection and reducing the false detection rate of the sign.

[0121] The signboard detection method provided in an embodiment of the present application, compared with the signboard detection method shown in Figure 1, can also project at least one three-dimensional bounding box into a two-dimensional space to obtain a two-dimensional frame corresponding to each of the at least one three-dimensional bounding boxes; match the two-dimensional frame corresponding to each of the at least one three-dimensional bounding boxes with at least one two-dimensional visual detection frame to obtain a matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding box; obtain the position of at least one signboard included in the target area based on the matching relationship, and detect the position of the signboard with less scene dependence and less calculation by performing multi-level matching and filtering on the two-dimensional visual detection frame and the three-dimensional bounding box, thereby improving the accuracy of signboard detection and reducing the power consumption and cost of signboard detection.

[0122] Please refer to Figure 8, which shows a sign detection device provided by an embodiment of the present application, which is applied to the above-mentioned vehicle. The sign detection device 200 includes: a target point cloud data acquisition module 210, a three-dimensional bounding box acquisition module 220, a two-dimensional visual detection frame acquisition module 230, and a sign position acquisition module 240, wherein:

[0123] The target point cloud data acquisition module 210 is used to acquire target point cloud data corresponding to a target area through the vehicle's laser radar, and to acquire camera vision data corresponding to the target area through the vehicle's camera, wherein the target area includes at least one identification plate.

[0124] The three-dimensional bounding box obtaining module 220 is configured to perform point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data.

[0125] The two-dimensional visual detection frame acquisition module 230 is configured to acquire at least one two-dimensional visual detection frame corresponding to the camera visual data.

[0126] The sign position obtaining module 240 is configured to obtain the position of the at least one sign included in the target area according to the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box.

[0127] Furthermore, the sign position obtaining module 240 may include: a three-dimensional bounding frame projection unit, a three-dimensional bounding frame and two-dimensional visual detection frame matching unit, and a sign position obtaining sub-unit, wherein:

[0128] The three-dimensional bounding frame projection unit is configured to project the at least one three-dimensional bounding frame into a two-dimensional space to obtain a two-dimensional frame corresponding to each of the at least one three-dimensional bounding frame.

[0129] The three-dimensional bounding box and two-dimensional visual detection frame matching unit is used to match the two-dimensional frame corresponding to each of the at least one three-dimensional bounding box with the at least one two-dimensional visual detection frame to obtain a matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding box.

[0130] The identification plate position obtaining subunit is configured to obtain the position of the at least one identification plate included in the target area based on the matching relationship.

[0131] Furthermore, the sign position obtaining subunit may include: a one-to-one relationship three-dimensional bounding box and two-dimensional visual detection box determining unit and a first position obtaining unit, wherein:

[0132] A one-to-one relationship three-dimensional bounding box and two-dimensional visual detection frame determination unit is used to determine, from the at least one three-dimensional bounding box, a three-dimensional bounding box that is in a one-to-one relationship with the at least one two-dimensional visual detection frame as a first three-dimensional bounding box if the matching relationship is a one-to-one relationship, and to determine the two-dimensional visual detection frame that is in a one-to-one relationship with the first three-dimensional bounding box as a first two-dimensional visual detection frame.

[0133] The first position obtaining unit is configured to obtain a first position corresponding to the first three-dimensional bounding frame if a signboard is detected in the first two-dimensional visual detection frame, and determine the first position as the position of the signboard included in the first two-dimensional visual detection frame.

[0134] Furthermore, the one-to-one relationship three-dimensional bounding box and two-dimensional visual detection box determining unit may include: a to-be-detected frame determining unit and a first three-dimensional bounding box determining unit, wherein:

[0135] The to-be-detected frame determining unit is configured to determine, from the at least one three-dimensional bounding frame, a three-dimensional bounding frame that is in a one-to-one relationship with the at least one two-dimensional visual detection frame as the to-be-detected frame.

[0136] The first three-dimensional bounding box determining unit is configured to determine the frame to be detected as the first three-dimensional bounding box if the area of ​​the frame to be detected is smaller than an area threshold.

[0137] Furthermore, the sign position obtaining subunit may include: a multiple three-dimensional bounding frame determining unit in a many-to-one relationship, a multiple three-dimensional bounding frame fusion unit, a multiple-to-one second two-dimensional visual detection frame determining unit, and a second position obtaining unit, wherein:

[0138] A multiple three-dimensional bounding box determination unit in a many-to-one relationship is used to determine, from the at least one three-dimensional bounding box, a multiple three-dimensional bounding box in a many-to-one relationship with one of the at least one two-dimensional visual detection frames if the matching relationship is a many-to-one relationship.

[0139] The multiple three-dimensional bounding frame fusion unit is used to fuse the multiple three-dimensional bounding frames to obtain a second three-dimensional bounding frame.

[0140] The many-to-one relationship second two-dimensional visual detection frame determining unit is used to determine the two-dimensional visual detection frame in the at least one two-dimensional visual detection frame that has a many-to-one relationship with the second three-dimensional bounding frame as the second two-dimensional visual detection frame.

[0141] The second position obtaining unit is configured to obtain a second position corresponding to the second three-dimensional bounding box if it is detected that the second two-dimensional visual detection frame includes a sign, and determine the second position as the position of the sign included in the second two-dimensional visual detection frame.

[0142] Furthermore, the sign position obtaining subunit may include: a third three-dimensional bounding box determining unit, a third three-dimensional bounding box segmenting unit, and a third position obtaining unit in a one-to-many relationship, wherein:

[0143] A one-to-many relationship third three-dimensional bounding box determination unit is used to determine, from the at least one three-dimensional bounding box, a one-to-many relationship with multiple third two-dimensional visual detection frames in the at least one two-dimensional visual detection frame if the matching relationship is a one-to-many relationship, wherein each of the third two-dimensional visual detection frames includes an identification plate.

[0144] The third three-dimensional bounding box segmentation unit is configured to segment the third three-dimensional bounding box based on the multiple third two-dimensional visual detection frames to obtain target third three-dimensional bounding boxes corresponding to each of the multiple third two-dimensional visual detection frames.

[0145] The third position obtaining unit is configured to obtain a third position of the target third three-dimensional bounding frame, and use the third position as the position of the sign included in the plurality of third two-dimensional visual detection frames.

[0146] Furthermore, after obtaining the position of the at least one sign included in the target area according to the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box, the sign detection device 200 may further include: a ground point cloud data acquisition unit and at least one two-dimensional visual detection box updating unit, wherein:

[0147] The ground point cloud data acquisition unit is configured to acquire ground point cloud data corresponding to the target point cloud data if it is detected that there is a two-dimensional visual detection frame in the at least one two-dimensional visual detection frame that does not match the at least one three-dimensional bounding frame.

[0148] At least one two-dimensional visual detection frame updating unit is used to filter out the unmatched two-dimensional visual detection frame from the at least one two-dimensional visual detection frame if it is detected that the unmatched two-dimensional visual detection frame matches the ground point cloud data, so as to obtain the at least one updated two-dimensional visual detection frame.

[0149] Furthermore, the three-dimensional bounding box and two-dimensional visual detection box matching unit may include: an overlapping area acquisition unit and a matching relationship determination unit, wherein:

[0150] The overlapping area acquisition unit is used to acquire the overlapping area of ​​the two-dimensional frames corresponding to the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding frame.

[0151] A matching relationship determining unit is configured to determine a matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding frame based on the overlapping area.

[0152] Furthermore, the matching relationship determination unit may include: a one-to-one relationship determination unit, and / or a many-to-one relationship determination unit, and / or a mismatch determination unit, and / or a one-to-many relationship determination unit, wherein:

[0153] A one-to-one relationship determination unit is used to determine that the matching relationship between the fourth three-dimensional bounding box and the fourth two-dimensional visual detection frame is a one-to-one relationship if, based on the overlapping area, it is determined that the fourth three-dimensional bounding box in the at least one three-dimensional bounding box has an area overlap with the fourth two-dimensional visual detection frame in the at least one two-dimensional visual detection.

[0154] A many-to-one relationship determination unit is used to determine that the matching relationship between the multiple fifth three-dimensional bounding boxes and the fifth two-dimensional visual detection frame is a many-to-one relationship if, based on the overlapping area, it is determined that there is an area overlap between the multiple fifth three-dimensional bounding boxes in the at least one three-dimensional bounding box and the fifth two-dimensional visual detection frame in the at least one two-dimensional visual detection.

[0155] A mismatch determination unit is used to determine that the sixth three-dimensional bounding box does not match the sixth two-dimensional visual detection frame if, based on the overlapping area, it is determined that there is no area overlap between the sixth three-dimensional bounding box in the at least one three-dimensional bounding box and the sixth two-dimensional visual detection frame of the at least one three-dimensional bounding box.

[0156] A one-to-many relationship determination unit is used to determine that the matching relationship between the seventh three-dimensional enclosing frame and the multiple seventh two-dimensional visual detection frames including the identification plate is a one-to-many relationship if it is determined based on the overlapping area that there are multiple seventh two-dimensional visual detection frames including the identification plate in the at least one two-dimensional visual detection frame and there is an overlapping area with the seventh three-dimensional enclosing frame in the at least one three-dimensional enclosing frame.

[0157] Furthermore, before obtaining the position of the sign included in the target area based on the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box, the sign detection device 200 may further include: an obstacle frame obtaining unit, an obstacle fusion result obtaining unit, and at least one three-dimensional bounding box filtering unit, wherein:

[0158] The obstacle frame obtaining unit is used to perform obstacle detection on the target point cloud data to obtain an obstacle frame corresponding to the target point cloud data.

[0159] An obstacle fusion result obtaining unit is configured to obtain an obstacle fusion result corresponding to the target point cloud data according to the obstacle frame, the at least one three-dimensional bounding box, and the camera vision data.

[0160] At least one three-dimensional bounding box filtering unit is configured to filter the at least one three-dimensional bounding box according to the obstacle fusion result to obtain at least one filtered three-dimensional bounding box.

[0161] Furthermore, after obtaining the position of the at least one sign included in the target area based on the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box, the sign detection device 200 may further include: a target two-dimensional visual detection box acquisition unit and a sign type acquisition unit, wherein:

[0162] The target two-dimensional visual detection frame acquisition unit is used to acquire the target two-dimensional visual detection frame corresponding to each of the at least one signboard.

[0163] The sign type acquisition unit is used to obtain the type corresponding to each of the at least one sign according to the camera visual data corresponding to the target two-dimensional visual detection frame.

[0164] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0165] In several embodiments provided in this application, the coupling between modules may be electrical, mechanical or other forms of coupling.

[0166] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.

[0167] Please refer to Figure 9, which shows a block diagram of a vehicle structure provided by an embodiment of the present application. The vehicle 100 can be an electronic device with processing capabilities, such as an electric vehicle or a gasoline vehicle. The vehicle 100 in the present application may include one or more of the following components: a processor 110, a memory 120, and one or more application programs, wherein the one or more application programs may be stored in the memory 120 and configured to be executed by the one or more processors 110, and the one or more programs are configured to execute the method described in the aforementioned method embodiment.

[0168] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the vehicle 100. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, as well as accesses data stored in the memory 120, to perform various functions and process data within the vehicle 100. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing displayed content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may instead be implemented via a separate communications chip.

[0169] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the various method embodiments described below, and the like. The data storage area may also store data created by the vehicle 100 during use (such as a phone book, audio and video data, and chat history data).

[0170] Please refer to Figure 10, which shows a block diagram of a computer-readable storage medium provided in an embodiment of the present application. The computer-readable storage medium 300 stores program code 310, which can be called by a processor to execute the method described in the above method embodiment.

[0171] The computer-readable storage medium 300 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium 300 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 300 has storage space for program code 310 for executing any of the method steps in the above method. These program codes 310 can be read from or written to one or more computer program products. The program code 310 can be compressed, for example, in a suitable form.

[0172] In summary, the signboard detection method, device, and vehicle provided in the embodiments of the present application obtain target point cloud data corresponding to the target area through the vehicle's laser radar, and obtain camera vision data corresponding to the target area through the vehicle's camera, wherein the target area includes at least one signboard; perform point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data; obtain at least one two-dimensional vision detection box corresponding to the camera vision data; obtain the position of at least one signboard included in the target area based on the at least one three-dimensional bounding box and the at least one two-dimensional vision detection box, thereby obtaining the position of the signboard included in the target area based on the three-dimensional point cloud data corresponding to the target area obtained by the laser radar and the two-dimensional camera vision data corresponding to the target area obtained by the camera, and detect the signboard in a combined laser radar and vision manner, thereby improving the accuracy of determining the position of the signboard.

[0173] The embodiments of the present application have been described above. The above description is illustrative and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for detecting a sign, characterized in that: Applied to a vehicle, the method comprises: Acquire target point cloud data corresponding to a target area through a laser radar of the vehicle, and acquire camera vision data corresponding to the target area through a camera of the vehicle, wherein the target area includes at least one identification plate; Performing point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data; Acquire at least one two-dimensional visual detection frame corresponding to the camera visual data; The position of the at least one identification plate included in the target area is obtained according to the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box.

2. The method according to claim 1, characterized in that The obtaining the position of the at least one sign included in the target area according to the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box includes: Projecting the at least one three-dimensional bounding box into a two-dimensional space to obtain a two-dimensional box corresponding to each of the at least one three-dimensional bounding box; Matching the two-dimensional frame corresponding to each of the at least one three-dimensional bounding frame with the at least one two-dimensional visual detection frame to obtain a matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding frame; The position of the at least one identification plate included in the target area is obtained based on the matching relationship.

3. The method according to claim 2, characterized in that The obtaining the position of the at least one identification plate included in the target area based on the matching relationship includes: If the matching relationship is a one-to-one relationship, determining a three-dimensional bounding box in a one-to-one relationship with the at least one two-dimensional visual detection box from the at least one three-dimensional bounding box as a first three-dimensional bounding box, and determining a two-dimensional visual detection box in a one-to-one relationship with the first three-dimensional bounding box as a first two-dimensional visual detection box; If it is detected that the first two-dimensional visual detection frame includes a sign, a first position corresponding to the first three-dimensional bounding frame is obtained, and the first position is determined as the position of the sign included in the first two-dimensional visual detection frame.

4. The method according to claim 3, characterized in that The determining, from the at least one three-dimensional bounding box, a three-dimensional bounding box in a one-to-one relationship with the at least one two-dimensional visual detection box as a first three-dimensional bounding box includes: Determine, from the at least one three-dimensional bounding box, a three-dimensional bounding box that is in a one-to-one relationship with the at least one two-dimensional visual detection box as a to-be-detected box; If the area of ​​the to-be-detected frame is smaller than the area threshold, the to-be-detected frame is determined as the first three-dimensional bounding frame.

5. The method according to claim 2, characterized in that: The obtaining the position of the at least one identification plate included in the target area based on the matching relationship includes: If the matching relationship is a many-to-one relationship, determining a plurality of three-dimensional bounding boxes from the at least one three-dimensional bounding box that form a many-to-one relationship with one of the at least one two-dimensional visual detection boxes; Fusing the multiple three-dimensional bounding boxes to obtain a second three-dimensional bounding box; Determine a two-dimensional visual detection frame in the at least one two-dimensional visual detection frame that is in a many-to-one relationship with the second three-dimensional bounding frame as a second two-dimensional visual detection frame; If it is detected that the second two-dimensional visual detection frame includes a sign, a second position corresponding to the second three-dimensional bounding frame is obtained, and the second position is determined as the position of the sign included in the second two-dimensional visual detection frame.

6. The method according to claim 2, characterized in that The obtaining the position of the at least one identification plate included in the target area based on the matching relationship includes: If the matching relationship is a one-to-many relationship, determining a third three-dimensional enclosing frame from the at least one three-dimensional enclosing frame that is in a one-to-many relationship with a plurality of third two-dimensional visual detection frames in the at least one two-dimensional visual detection frame, wherein each of the third two-dimensional visual detection frames includes a sign; Segmenting the third three-dimensional bounding box based on the multiple third two-dimensional visual detection boxes to obtain target third three-dimensional bounding boxes corresponding to each of the multiple third two-dimensional visual detection boxes; A third position of the target third three-dimensional bounding box is obtained, and the third position is used as the position of the sign included in the multiple third two-dimensional visual detection boxes.

7. The method according to claim 2, characterized in that After obtaining the position of the at least one sign included in the target area according to the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box, the method further includes: If it is detected that there is a two-dimensional visual detection frame in the at least one two-dimensional visual detection frame that does not match the at least one three-dimensional bounding frame, acquiring ground point cloud data corresponding to the target point cloud data; If it is detected that the unmatched two-dimensional visual detection frame matches the ground point cloud data, the unmatched two-dimensional visual detection frame is filtered out from the at least one two-dimensional visual detection frame to obtain the updated at least one two-dimensional visual detection frame.

8. The method according to claim 2, characterized in that: The matching of the two-dimensional frame corresponding to each of the at least one three-dimensional bounding frame with the at least one two-dimensional visual detection frame to obtain a matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding frame includes: Acquire an overlapping area of ​​a two-dimensional frame corresponding to each of the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding frame; A matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding frame is determined based on the overlapping area.

9. The method according to claim 8, characterized in that The determining, based on the overlapping area, a matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding frame includes: If it is determined based on the overlapping area that a fourth three-dimensional bounding box in the at least one three-dimensional bounding box overlaps with a fourth two-dimensional visual detection box in the at least one two-dimensional visual detection, then it is determined that the matching relationship between the fourth three-dimensional bounding box and the fourth two-dimensional visual detection box is a one-to-one relationship; and / or If it is determined based on the overlapping area that a plurality of fifth three-dimensional bounding boxes in the at least one three-dimensional bounding box overlap with a fifth two-dimensional visual detection box in the at least one two-dimensional visual detection, then it is determined that the matching relationship between the plurality of fifth three-dimensional bounding boxes and the fifth two-dimensional visual detection box is a many-to-one relationship; and / or If it is determined based on the overlapping area that there is no area overlap between the sixth three-dimensional bounding box in the at least one three-dimensional bounding box and the sixth two-dimensional visual detection box of the at least one three-dimensional bounding box, then it is determined that the sixth three-dimensional bounding box does not match the sixth two-dimensional visual detection box; and / or If it is determined based on the overlapping area that there are multiple seventh two-dimensional visual detection frames including identification plates in the at least one two-dimensional visual detection frame and there is an overlapping area with the seventh three-dimensional enclosing frame in the at least one three-dimensional enclosing frame, then it is determined that the matching relationship between the seventh three-dimensional enclosing frame and the multiple seventh two-dimensional visual detection frames including the identification plates is a one-to-many relationship.

10. The method according to any one of claims 1 to 9, characterized in that: Before obtaining the position of the sign included in the target area according to the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box, the method further includes: Performing obstacle detection on the target point cloud data to obtain an obstacle frame corresponding to the target point cloud data; Obtaining an obstacle fusion result corresponding to the target point cloud data according to the obstacle frame, the at least one three-dimensional bounding box and the camera visual data; The at least one three-dimensional bounding box is filtered according to the obstacle fusion result to obtain at least one filtered three-dimensional bounding box.

11. The method according to any one of claims 1 to 9, characterized in that: After obtaining the position of the at least one sign included in the target area according to the at least one three-dimensional bounding box and the at least one two-dimensional visual detection box, the method further includes: Obtain a target two-dimensional visual detection frame corresponding to each of the at least one signboard; The type corresponding to each of the at least one identification plate is obtained according to the camera vision data corresponding to the target two-dimensional visual detection frame.

12. A sign detection device, characterized in that: Applied to a vehicle, the device comprises: A target point cloud data acquisition module, used to acquire target point cloud data corresponding to a target area through a laser radar of the vehicle, and to acquire camera visual data corresponding to the target area through a camera of the vehicle, wherein the target area includes at least one identification plate; A three-dimensional bounding box obtaining module, used to perform point cloud segmentation on the target point cloud data to obtain at least one three-dimensional bounding box corresponding to the target point cloud data; A two-dimensional visual detection frame acquisition module, used to acquire at least one two-dimensional visual detection frame corresponding to the camera visual data; The sign position obtaining module is used to obtain the position of the at least one sign included in the target area according to the at least one three-dimensional enclosing frame and the at least one two-dimensional visual detection frame.

13. The device according to claim 12, characterized in that The identification plate position obtaining module comprises: A three-dimensional bounding frame projection unit, used to project the at least one three-dimensional bounding frame into a two-dimensional space to obtain a two-dimensional frame corresponding to each of the at least one three-dimensional bounding frame; a three-dimensional bounding frame and two-dimensional visual detection frame matching unit, configured to match the two-dimensional frame corresponding to each of the at least one three-dimensional bounding frame with the at least one two-dimensional visual detection frame to obtain a matching relationship between the at least one two-dimensional visual detection frame and the at least one three-dimensional bounding frame; The identification plate position obtaining subunit is used to obtain the position of the at least one identification plate included in the target area based on the matching relationship.

14. A vehicle, characterized in that: include: one or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1-11.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Target prediction method based on three-dimensional laser radar and vision fusion

    CN115205391A

  • Target identification method based on fusion of image information and laser radar point cloud information

    CN116229408A

  • Radar and camera external parameter calibration method, electronic equipment and storage medium

    CN116721162A

  • Signboard detection method and device and vehicle

    CN117671644A