Road traffic sign marking identification and positioning method and system for automatic driving

Through the joint calibration and fusion positioning technology of camera and lidar, the identification and positioning accuracy and cost of traffic sign markings in autonomous driving are solved, and high-precision and low-cost traffic sign markings are realized, which improves the perception ability of autonomous driving and high-precision map production efficiency.

CN120356184APending Publication Date: 2025-07-22BEIJING EGOVA TECH

Patent Information

Application Number
CN202510479343.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, in autonomous driving, traffic sign marking recognition based on optical cameras lacks depth information and is difficult to achieve centimeter-level accuracy. The method based on lidar is costly and lacks visual information. The hardware cost and computing power requirements of multi-sensor multi-modal fusion technology are high, which cannot effectively meet the real-time perception needs.

Method used

The joint calibration technology of camera and lidar is adopted, combined with AI data sets and identification models, and the camera and lidar are integrated to identify and locate traffic signs and markings, including traffic lights, crosswalks and lane lines, and point cloud data acquisition is used to reduce hardware costs.

Benefits of technology

It realizes high-precision identification and positioning of traffic signs and markings, with the recognition accuracy rate not less than 85%, the positioning accuracy is higher than 10cm, the hardware cost is low, and the inference time is shortened, which shortens the production cycle of high-precision maps and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356184A_ABST
    Figure CN120356184A_ABST
Patent Text Reader

Abstract

The invention discloses a road traffic sign and marking line identification and positioning method and system for automatic driving, and relates to the technical field of automatic driving. Comprising the following steps: acquiring internal parameters of a camera, external parameters between the camera and a laser radar, an AI data set, a traffic sign marking identification model, and fusion positioning of the camera and the laser radar; on the basis that the AI data set is made, when the automatic driving unmanned vehicle runs on a road, a camera is used for capturing a front video stream picture, a traffic sign marking line in the video stream picture is recognized through a traffic sign marking line recognition model, and coordinates of a detection frame of the traffic sign marking line in an image plane coordinate system are obtained; a plurality of laser radars are used for collecting original point cloud data, and the original point cloud data are analyzed through fusion positioning of a camera and the laser radars, so that a positioning result is obtained. The identification accuracy and the positioning precision are improved, and the cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and particularly to a method and system for identifying and positioning road traffic signs and markings for autonomous driving. Background Art

[0002] When an intelligent connected vehicle drives autonomously on an open road, the surrounding perception information sources are divided into two categories: one is the prior information provided based on high-precision maps or vehicle networking platforms, and the other is the real-time information collected and processed by in-vehicle sensors. The two complement each other to form the perception information for autonomous driving. Quickly identifying and positioning the road traffic signs and markings around the vehicle is one of the methods for multi-source data collection of high-precision maps, a prerequisite for autonomous driving to perform environmental perception, path planning, and decision-making, and also one of the information sources for the future vehicle-road-cloud integrated data base platform.

[0003] Currently, there are three commonly used recognition technologies: one is to use a visual recognition algorithm based on an optical camera to identify the position information of traffic signs and markings in an image; the second is to use a 3D detection algorithm based on lidar to identify the object model of traffic signs and markings; the third is a multi-sensor multi-modal fusion technology that fuses the visual information of an optical camera and the point cloud information of lidar. There are two commonly used technologies for positioning the identified traffic signs and markings: one is to use SLAM (Simultaneous Localization and Mapping) to obtain the corresponding local coordinates and convert them into world coordinates; the other is to use GNSS (Global Navigation Satellite System) to obtain the corresponding world coordinates. Through recognition and positioning, road information that can be used for high-precision map production and autonomous driving environmental perception is generated.

[0004] However, the technology based on an optical camera lacks depth information and it is difficult to obtain position information with centimeter-level accuracy. It is often used to assist in manually creating high-precision maps, increasing the production cycle and cost, and cannot be directly used for real-time perception in autonomous driving. The technology based on radar has a high hardware cost and lacks rich visual information. The technology based on multi-sensor multi-modal fusion is in the exploratory research stage, with high hardware costs and high requirements for vehicle-side computing power. Summary of the Invention

[0005] The purpose of this application is to provide a method and system for identifying and positioning road traffic signs and markings for autonomous driving, which improves the recognition accuracy and positioning precision and reduces the cost.

[0006] To achieve the above object, the present application provides a method for identifying and positioning road traffic signs and markings for autonomous driving, including the following steps: S1: Obtain the internal parameters of the camera, the external parameters between the camera and the lidar, the AI dataset, the traffic sign and marking recognition model, and the fusion positioning of the camera and the lidar; S2: On the basis of making the AI dataset, when the autonomous driving vehicle travels on the road, use the camera to capture the video stream picture in front, and use the traffic sign and marking recognition model to identify the traffic signs and markings in the video stream picture, and obtain the coordinates of the detection frame of the traffic signs and markings in the image plane coordinate system; wherein, the traffic signs and markings include: traffic lights, crosswalks, and lane lines; S3: Use multiple lidars to collect the original point cloud data, and use the fusion positioning of the camera and the lidar to analyze the original point cloud data, so as to obtain the positioning results, wherein the positioning results include: traffic light positioning results, crosswalk positioning results, and lane line positioning results.

[0007] As above, wherein, the sub-steps of obtaining the internal parameters of the camera, the external parameters between the camera and the lidar, the AI dataset, the traffic sign and marking recognition model, and the fusion positioning of the camera and the lidar are as follows: S11: Use the joint calibration technology of the camera and the radar to measure the internal parameters of the camera and the external parameters between the camera and the lidar; S12: Construct the AI dataset and the traffic sign and marking recognition model, wherein the AI dataset includes: sample data and weight data; the traffic sign and marking recognition model includes: traffic light inference model, crosswalk inference model, and lane line inference model; S13: Locate the traffic signs and markings to obtain the fusion positioning of the camera and the lidar, wherein the fusion positioning of the camera and the lidar includes: traffic light positioning strategy, crosswalk positioning strategy, and lane line positioning strategy.

[0008] As above, wherein, the sub-steps of using the joint calibration technology of the camera and the radar to measure the internal parameters of the camera and the external parameters between the camera and the lidar are as follows: S111: Set the center of the rear axle of the vehicle body of the unmanned vehicle as the origin of the vehicle body coordinate system, and by measuring the data between the installation position of the global navigation satellite system positioning antenna and the center of the rear axle of the vehicle body, the mapping matrix between the vehicle body coordinate system and the inertial navigation measurement unit can be obtained; S112: Measure the data between the installation position of the lidar and the center of the rear axle of the vehicle body, and the mapping matrix between the lidar coordinate system and the vehicle body coordinate system can be obtained; S113: Perform the internal parameter calibration of the camera to obtain the internal parameters of the camera; S114: Perform the joint calibration of the camera and the lidar to obtain the external parameters between the camera and the lidar.

[0009] As above, wherein, the expression of the internal parameters of the camera is: Among them, fx, fy, u0, and v0 are the internal parameters of the camera; fx is the equivalent focal length of the camera in the x direction, fy is the equivalent focal length of the camera in the y direction; (u0, v0) is the coordinates of the principal image point; (u, v) are the coordinates of the image plane coordinate system corresponding to the point q in the three-dimensional space of the world coordinate system; (xc, yc, zc) are the coordinates of the point q in the camera coordinate system.

[0010] As described above, among them, the sub-steps for jointly calibrating the camera and the lidar to obtain the external parameters between the camera and the lidar are as follows: S1141: Establish a projection matrix for converting the lidar coordinate system to the image plane coordinate system; S1142: Directly measure the external parameters between the lidar, the vehicle body center, and the inertial navigation unit, and establish a mapping matrix between the lidar coordinate system and the world coordinate system; S1143: After converting the original point cloud data of multiple lidars to the world coordinate system through the mapping matrix between the lidar coordinate system and the world coordinate system, then perform joint calibration of the camera and the lidar to obtain the external parameters between the camera and the lidar.

[0011] As described above, among them, the sub-steps for constructing the AI dataset and the traffic sign and marking recognition model are as follows: S121: Identify the traffic signs and markings to obtain the traffic light image dataset, the labeled traffic light dataset, and the traffic light inference model; S122: Identify the traffic signs and markings to obtain the crosswalk image dataset, the labeled crosswalk dataset, and the crosswalk inference model; S123: Identify the traffic signs and markings to obtain the lane line dataset, the labeled lane line dataset, and the lane line inference model; S124: Use the traffic light image dataset, the crosswalk image dataset, and the lane line dataset as sample data, and use the labeled traffic light dataset, the labeled crosswalk dataset, and the labeled lane line dataset as weight data; The AI dataset is composed of the sample data and the weight data; S125: Use the traffic light inference model, the crosswalk inference model, and the lane line inference model as the traffic sign and marking recognition model.

[0012] As described above, the sub-steps of identifying traffic signs and markings to obtain the traffic light image dataset, the labeled traffic light dataset, and the traffic light inference model are as follows: R1: Collect the traffic light image dataset; R2: Perform labeling processing on the traffic light image dataset to obtain the labeled traffic light dataset, where the labeling content of the labeling processing includes at least: traffic light detection frame, traffic light type, and traffic light driving direction; R3: Divide the labeled traffic light dataset into a training set and a validation set according to a preset ratio, use the training set to train the traffic light recognition model to obtain the trained traffic light model, and use the validation set and the evaluation model to analyze the accuracy of the trained traffic light model. If the accuracy of the trained traffic light model meets the requirements, use the trained traffic light model as the traffic light inference model for model inference; if the accuracy of the trained traffic light model does not meet the requirements, re-obtain the trained traffic light model.

[0013] As described above, the positioning algorithm process of the traffic light positioning strategy is as follows: E1: Obtain the traffic light recognition result and the laser point cloud data at the time of imaging; E2: Input the laser point cloud data at the time of imaging, the internal parameter matrix and the external parameter matrix of the camera, and project it into the image coordinate system by using the calculation formula for projecting the three-dimensional laser point cloud in the lidar coordinate system into the image plane coordinate points of the image plane coordinate system to obtain the traffic light image coordinates; E3: Traverse the laser point cloud data; E4: Use the traffic light image coordinates to judge whether the laser point cloud data is within the camera imaging range. If so, use this laser point cloud data as the filtered laser point cloud data and execute E5; if not, end the process; E5: Judge whether the filtered laser point cloud data is within the detection frame range. If so, use this filtered laser point cloud data as the extracted laser point cloud data and execute E6; if not, end the process; E6: Mark the extracted laser point cloud data as the recognized point cloud; E7: Filter the recognized point cloud to obtain the filtered point cloud; E8: Cluster the filtered point cloud to obtain the clustered point cloud; E9: Obtain the geometric center coordinates of the clustered point cloud as the traffic light positioning result and end the process; where the calculation formula for projecting the three-dimensional laser point cloud in the lidar coordinate system into the image plane coordinate points of the image plane coordinate system by the projection matrix is: Among them, is the internal parameter matrix of the camera; is the external parameter matrix; (u, v) is the coordinate of the image plane coordinate system corresponding to the point q in the three-dimensional space of the world coordinate system; fx is the equivalent focal length of the camera in the x direction; fy is the equivalent focal length of the camera in the y direction; (u0, v0) is the image principal point coordinate; Tk is the translation vector from the lidar coordinate system to the camera coordinate system; Rk is the rotation matrix from the lidar coordinate system to the camera coordinate system; is a 3×3 zero matrix; X is the value of any point in the laser point cloud data in the X-axis direction of the lidar coordinate system; Y is the value of any point in the laser point cloud data in the Y-axis direction of the lidar coordinate system; Z is the value of any point in the laser point cloud data in the Z-axis direction of the lidar coordinate system.

[0014] The present application also provides a road traffic sign and marking recognition and positioning system for autonomous driving, including: an unmanned vehicle, at least one camera, at least one lidar, at least one set of inertial navigation measurement units, and a recognition and positioning unit; at least one camera, at least one lidar, at least one set of inertial navigation measurement units, and the recognition and positioning unit are all mounted on the unmanned vehicle; wherein, the camera: is used to collect image data, wherein the image data includes: video stream pictures; the lidar: is used to collect original point cloud data; the inertial navigation measurement unit: is used to measure the coordinates of the phase center of the global navigation satellite system positioning antenna in the world coordinate system and the azimuth angle of the vehicle attitude in real time; the recognition and positioning unit: is used to execute the above-mentioned road traffic sign and marking recognition and positioning method for autonomous driving.

[0015] As described above, wherein the lidar includes: at least one 16-line mechanical lidar and at least one solid-state lidar; wherein, the 16-line mechanical lidar: is used to collect the laser point cloud data of traffic signs and markings within the first line-of-sight range, the first horizontal field of view range, and the first vertical field of view range; the solid-state lidar: is used to collect the laser point cloud data of traffic signs and markings within the second line-of-sight range, the second horizontal field of view range, and the second vertical field of view range.

[0016] The beneficial effects achieved by the present application are as follows:

[0017] (1) The road traffic sign and marking recognition and positioning method and system for autonomous driving of the present application make full use of the cameras, lidars, and inertial navigation integrated in the intelligent networked unmanned vehicle to achieve recognition and positioning, without the need to additionally mount sensor devices.

[0018] (2) The road traffic sign and marking recognition and positioning system for autonomous driving of the present application has the advantage of low hardware cost. Specifically, the main hardware cost in the system comes from the mechanical lidar. The cost of the 16-line mechanical lidar or solid-state lidar adopted in the present application is much lower than that of other mechanical lidars with more than 32 lines.

[0019] (3) The road traffic sign and marking recognition and positioning method and system for autonomous driving of the present application improve the recognition accuracy and recognition accuracy, and the recognition accuracy of various common traffic signs and markings on the road is not less than 85%.

[0020] (4) The road traffic sign and marking recognition and positioning method and system for autonomous driving of the present application have high positioning accuracy, and the positioning accuracy of the center point of the detection frame is higher than 10 cm.

[0021] (5) The road traffic sign and marking recognition and positioning method and system for autonomous driving of the present application shorten the inference time. In the Jetson AGX Orin 32GB processor, the inference time is less than 100 ms.

[0022] (6) The road traffic sign and marking recognition and positioning method and system for autonomous driving of the present application improve the productivity of high-precision map making. The generated high-precision map can be directly used, or it can quickly generate usable data results, greatly shortening the production cycle of the high-precision map and reducing the cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0024] Figure 1 It is a flowchart of an embodiment of the road traffic sign and marking recognition and positioning method for autonomous driving;

[0025] Figure 2 It is a schematic diagram of the road traffic sign and marking recognition and positioning technology for autonomous driving;

[0026] Figure 3 It is a schematic diagram of four coordinate systems;

[0027] Figure 4 It is a diagram of the coordinate system conversion relationship. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0029] This application provides a road traffic sign and marking recognition and positioning system for autonomous driving, including: an autonomous vehicle, at least one camera, at least one lidar, at least one set of inertial navigation measurement units, and a recognition and positioning unit; at least one camera, at least one lidar, at least one set of inertial navigation measurement units, and the recognition and positioning unit are all mounted on the autonomous vehicle.

[0030] Among them, the camera: is used to collect image data, where the image data includes: video stream pictures.

[0031] Specifically, the image data collected by the camera is the data source for recognizing traffic signs and markings.

[0032] The lidar: is used to collect raw point cloud data.

[0033] Specifically, the raw point cloud data collected by the lidar is the data source for positioning traffic signs and markings.

[0034] The inertial navigation measurement unit: is used to measure in real time the coordinates of the phase center of the GNSS (Global Navigation Satellite System) positioning antenna in the world coordinate system and the azimuth angle of the vehicle attitude.

[0035] Specifically, it is set that the center of the rear axle of the vehicle body of the autonomous vehicle is the origin of the vehicle body coordinate system, and the spatial relationship between the phase center of the GNSS (Global Navigation Satellite System) positioning antenna and the center of the rear axle of the vehicle body is rigid. By directly measuring the installation position of the GNSS (Global Navigation Satellite System) positioning antenna to the center of the rear axle of the vehicle body, the mapping matrix between the vehicle body coordinate system and the inertial navigation measurement unit can be obtained.

[0036] The recognition and positioning unit: is used to execute the following road traffic sign and marking recognition and positioning method for autonomous driving.

[0037] Furthermore, the lidar includes: at least one 16-line mechanical lidar and at least one solid-state lidar.

[0038] Specifically, the cost of the 16-line mechanical lidar or the solid-state lidar is much lower than that of other mechanical lidars with more than 32 lines, which can reduce the hardware cost.

[0039] Among them, the 16-line mechanical lidar: is used to collect the laser point cloud data of traffic signs and markings within the first line-of-sight range, the first horizontal field of view range, and the first vertical field of view range.

[0040] Specifically, the specific values of the first line-of-sight range, the first horizontal field of view range, and the first vertical field of view range are set according to the actual situation. This application preferably uses: the first line-of-sight range is 0 m to 150 m, the first horizontal field of view range is 0 degrees to 360 degrees, and the first vertical field of view range is -15 degrees to 15 degrees.

[0041] Solid-state lidar: used to collect the laser point cloud data of traffic signs and markings within the second line-of-sight range, the second horizontal field-of-view range, and the second vertical field-of-view range.

[0042] Specifically, the specific values of the second line-of-sight range, the second horizontal field-of-view range, and the second vertical field-of-view range are set according to the actual situation. In this application, preferably: the second line-of-sight range is 0 m to 40 m, the second horizontal field-of-view range is 0° to 180°, and the second vertical field-of-view range is 0° to 52°.

[0043] Furthermore, the specific number of cameras is set according to the actual situation. In this application, preferably one camera is used.

[0044] Furthermore, the camera is an optical recognition camera, but not limited to an optical recognition camera. In this application, preferably an optical recognition camera is used.

[0045] Furthermore, the specific number of 16-line mechanical lidars is set according to the actual situation. In this application, preferably 1 or 2 lidars are used.

[0046] Furthermore, the specific number of solid-state lidars is set according to the actual situation. In this application, preferably [0,4] lidars are used.

[0047] Furthermore, the specific number of inertial navigation measurement units (i.e., inertial navigation units) is set according to the actual situation. In this application, preferably one set of units is used.

[0048] As Figures 1-4 shown, this application provides a method for identifying and positioning traffic signs and markings for autonomous driving, including the following steps:

[0049] S1: Obtain the internal parameters of the camera, the external parameters between the camera and the lidar, the AI dataset, the traffic sign and marking recognition model, and the fusion positioning of the camera and the lidar.

[0050] Furthermore, the sub-steps of obtaining the internal parameters of the camera, the external parameters between the camera and the lidar, the AI dataset, the traffic sign and marking recognition model, and the fusion positioning of the camera and the lidar are as follows:

[0051] S11: Use the joint calibration technology of the camera and the lidar to measure the internal parameters of the camera and the external parameters between the camera and the lidar.

[0052] Specifically, the final positioning result of the recognition target is the world coordinate system, which can be directly used for high-precision map making and autonomous driving perception. The purpose of measuring the internal parameters of the camera and the external parameters between the camera and the lidar is to calculate the mapping matrix (i.e., transformation matrix) between the world coordinate system, the vehicle body coordinate system, the image plane coordinate system, and the lidar coordinate system. Among them, asFigure 3 As shown in the figure, the world coordinate system has its origin at the center of the earth. The X-axis points to the Greenwich Observatory in the direction of the 0-degree meridian, the Z-axis points to the average center of the north axis of the earth's pole, and the Y-axis forms a right-handed rectangular coordinate system with the X-axis and Z-axis. The vehicle body coordinate system is the coordinate system where the intelligent connected autonomous vehicle is located. The image plane coordinate system is the coordinate system of the image after camera imaging with the upper left corner of the photo as the origin. The lidar coordinate system is the coordinate system of the original point cloud data of the lidar with the optical center of the lidar as the origin.

[0053] Furthermore, the sub-steps of using the joint calibration technology of the camera and the radar to measure the internal parameters of the camera and the external parameters between the camera and the lidar are as follows:

[0054] S111: Set the center of the rear axle of the vehicle body of the autonomous vehicle as the origin of the vehicle body coordinate system. By measuring the data between the installation position of the global navigation satellite system positioning antenna and the center of the rear axle of the vehicle body, the mapping matrix between the vehicle body coordinate system and the inertial navigation measurement unit can be obtained.

[0055] Specifically, set the center of the rear axle of the vehicle body of the autonomous vehicle as the origin of the vehicle body coordinate system. The spatial relationship between the phase center of the global navigation satellite system (GNSS) and the center of the rear axle of the vehicle body is rigid. By directly measuring the data between the installation position of the global navigation satellite system (GNSS) positioning antenna and the center of the rear axle of the vehicle body, the mapping matrix between the vehicle body coordinate system and the inertial navigation measurement unit can be obtained.

[0056] S112: Measure the data between the installation position of the lidar and the center of the rear axle of the vehicle body, and the mapping matrix between the lidar coordinate system and the vehicle body coordinate system can be obtained.

[0057] Specifically, the original point cloud data collected by the lidar is the coordinate relative to the optical center in the lidar coordinate system. Therefore, it is necessary to first convert the original point cloud data into the vehicle body coordinate system. The installation position of the lidar is fixed, and the spatial relationship between the installation position of the lidar and the center of the rear axle of the autonomous vehicle is also rigid. By directly measuring the data between the installation position of the lidar and the center of the rear axle of the vehicle body, the mapping matrix between the lidar coordinate system and the vehicle body coordinate system can be obtained.

[0058] S113: Perform the internal parameter calibration of the camera to obtain the internal parameters of the camera.

[0059] Furthermore, the Zhang Zhengyou calibration method is used to perform the internal parameter calibration of the camera to obtain the internal parameters of the camera, but it is not limited to the Zhang Zhengyou calibration method.

[0060] Furthermore, the expression of the internal parameters of the camera is:

[0061]

[0062] Among them, fx, fy, u0, and v0 are the internal parameters of the camera; fx is the equivalent focal length of the camera in the x direction; fy is the equivalent focal length of the camera in the y direction; (u0, v0) is the image principal point coordinates; (u, v) is the coordinate of the image plane coordinate system corresponding to the point q in the three-dimensional space of the world coordinate system (i.e., the image coordinate);

[0063] (xc, yc, zc) is the coordinate of the point q in the camera coordinate system.

[0064] Specifically, the distorted image is restored to an undistorted image through the internal parameters of the camera.

[0065] S114: Perform the joint calibration of the camera and the lidar to obtain the external parameters between the camera and the lidar.

[0066] Furthermore, the sub-steps for performing the joint calibration of the camera and the lidar to obtain the external parameters between the camera and the lidar are as follows:

[0067] S1141: Establish a projection matrix for converting the lidar coordinate system (i.e., the three-dimensional coordinate system of the lidar point cloud) to the image plane coordinate system (i.e., the two-dimensional coordinate system of the image).

[0068] Furthermore, the calculation formula for projecting the three-dimensional lidar point cloud in the lidar coordinate system into the image plane coordinate points of the image plane coordinate system (i.e., the image plane coordinate points in the picture) through the projection matrix is:

[0069]

[0070] Among them, is the internal parameter matrix of the camera; is the transformation matrix from the lidar coordinate system to the camera coordinate system (i.e., the external parameter matrix); (u, v) is the coordinate of the image plane coordinate system corresponding to the point q in the three-dimensional space of the world coordinate system; fx is the equivalent focal length of the camera in the x direction; fy is the equivalent focal length of the camera in the y direction; (u0, v0) is the image principal point coordinates; Tk is the translation vector from the lidar coordinate system to the camera coordinate system; Rk is the rotation matrix from the lidar coordinate system to the camera coordinate system; is a 3×3 zero matrix; X is the value of any point in the lidar point cloud data in the X axis direction of the lidar coordinate system; Y is the value of any point in the lidar point cloud data in the Y axis direction of the lidar coordinate system; Z is the value of any point in the lidar point cloud data in the Z axis direction of the lidar coordinate system.

[0071] Specifically, the calculation formula for projecting the three-dimensional lidar point cloud into the image plane coordinate points in the picture is used to convert the three-dimensional coordinates into two-dimensional coordinates, and there is no precision loss in the dimension reduction process.

[0072] S1142: Directly measure the extrinsic parameters between the lidar, the vehicle body center, and the inertial navigation unit, and establish the mapping matrix between the lidar coordinate system and the world coordinate system.

[0073] S1143: After converting the original point cloud data of multiple lidars to the world coordinate system through the mapping matrix between the lidar coordinate system and the world coordinate system, perform the joint calibration of the camera and the lidar to obtain the extrinsic parameters between the camera and the lidar.

[0074] Furthermore, the expression for the conversion relationship between the camera coordinate system and the lidar coordinate system is:

[0075]

[0076] where, (Xc, Yc, Zc) is the coordinate of the point q in the three-dimensional space in the world coordinate system in the camera coordinate system; (Xt, Yt, Zt) is the coordinate of the point q in the three-dimensional space in the world coordinate system in the lidar coordinate system; Tk is the translation vector from the lidar coordinate system to the camera coordinate system, Tk = [tx ty tz] T , tx is the translation vector in the x-axis direction; ty is the translation vector in the y-axis direction; tz is the translation vector in the z-axis direction; T is the transpose; Rk is the rotation matrix from the lidar coordinate system to the camera coordinate system, Rk = Rx·Ry·Rz, Rx is the rotation matrix obtained by the lidar coordinate system rotating counterclockwise around the x-axis of the camera coordinate system; Ry is the rotation matrix obtained by the lidar coordinate system rotating counterclockwise around the y-axis of the camera coordinate system; Rz is the rotation matrix obtained by the lidar coordinate system rotating counterclockwise around the z-axis of the camera coordinate system; Tk and Rk are the extrinsic parameters between the camera and the lidar that need to be measured through the joint calibration of the camera and the lidar.

[0077] Specifically, in the common imaging space of the camera and the lidar, a three-dimensional lidar point cloud corresponds to a pixel point in a picture. Therefore, preferably: After converting the original point cloud data (i.e., point cloud coordinates) of multiple lidars to the world coordinate system, perform the joint calibration of the camera and the lidar to obtain the extrinsic parameters between the camera and the lidar, that is: the translation and rotation parameters between the camera coordinate system and the lidar coordinate system. Among them, as Figure 3 shown, the camera coordinate system is a three-dimensional coordinate system established with the camera center as the origin.

[0078] S12: Construct an AI (Artificial Intelligence) dataset and a traffic sign and marking recognition model. Among them, the AI dataset includes: sample data and weight data; the traffic sign and marking recognition model includes: a traffic light inference model, a crosswalk inference model, and a lane line inference model.

[0079] Furthermore, the sub-steps for constructing the AI dataset and the traffic sign and marking recognition model are as follows:

[0080] S121: Identify traffic signs and markings to obtain a traffic light image dataset, a labeled traffic light dataset, and a traffic light inference model.

[0081] Furthermore, the sub - steps for identifying traffic signs and markings to obtain a traffic light image dataset, a labeled traffic light dataset, and a traffic light inference model are as follows:

[0082] R1: Collect a traffic light image dataset.

[0083] Specifically, the traffic light image dataset is an image dataset containing traffic lights.

[0084] As an embodiment, use a camera to collect multiple image data (such as video stream images) containing traffic lights at different distances, different angles, and different weather conditions on the road as the traffic light image dataset, but not limited to a camera.

[0085] R2: Perform annotation processing on the traffic light image dataset to obtain a labeled traffic light dataset. Among them, the annotation content of the annotation processing at least includes: traffic light detection box, traffic light type, and traffic light driving direction.

[0086] Specifically, use the method of manual annotation to perform annotation processing on the traffic light image dataset to obtain a labeled dataset, but not limited to using the method of manual annotation.

[0087] Among them, the traffic light types include: red light, green light, yellow light, and unknown.

[0088] Among them, the traffic light driving directions include: straight - through signal light, left - turn signal light, and right - turn signal light.

[0089] R3: Divide the labeled traffic light dataset into a training set and a validation set according to a preset ratio. Use the training set to train the traffic light recognition model to obtain a trained traffic light model. Use the validation set and the evaluation model to analyze the accuracy of the trained traffic light model. If the accuracy of the trained traffic light model meets the requirements, then use the trained traffic light model as the traffic light inference model for model inference; if the accuracy of the trained traffic light model does not meet the requirements, then re - obtain the trained traffic light model.

[0090] Specifically, the specific value of the preset ratio is set according to the actual situation. This application preferably uses: training set: validation set = 8:2, that is, 80% of the labeled traffic light dataset is used for training, and 20% of the labeled traffic light dataset is used for model validation.

[0091] R4: Use the traffic light inference model to perform traffic light recognition on the images collected by the camera in real - time to obtain a traffic light recognition result.

[0092] Further, deploy the recognition algorithm in TensorRT (TensorRT: TensorRT is a high-performance deep learning inference optimizer) to call GPU resources to accelerate the inference process of the traffic light inference model for identifying traffic lights in the images captured by the camera in real time, and obtain the traffic light recognition result.

[0093] Further, select the single-stage object detection algorithm YOLOv5 based on deep learning to identify traffic lights, but not limited to selecting the single-stage object detection algorithm YOLOv5 based on deep learning to identify traffic lights.

[0094] Specifically, YOLOv5 is an object detection model based on deep learning, which is widely used in real-time object recognition tasks and can meet the requirements of real-time detection for autonomous driving.

[0095] This application selects the YOLOv5 object detection algorithm and deploys it in TensorRT, and utilizes the computing power of the GPU to ensure the real-time performance of recognition.

[0096] S122: Identify traffic signs and markings to obtain a crosswalk image dataset, an annotated crosswalk dataset, and a crosswalk inference model.

[0097] Further, the sub-steps of identifying traffic signs and markings to obtain a crosswalk image dataset, an annotated crosswalk dataset, and a crosswalk inference model are as follows:

[0098] T1: Collect a crosswalk image dataset.

[0099] Specifically, the crosswalk image dataset is an image dataset containing crosswalks.

[0100] As an embodiment, use a camera to collect multiple image data containing crosswalks at different distances, different angles, and different weather conditions on the road as the crosswalk image dataset, but not limited to a camera.

[0101] T2: Perform annotation processing on the crosswalk image dataset to obtain an annotated crosswalk dataset, where the annotation content of the annotation processing includes at least: a crosswalk detection box.

[0102] Specifically, use the method of manual annotation to perform annotation processing on the crosswalk image dataset to obtain an annotated crosswalk dataset, but not limited to using the method of manual annotation.

[0103] T3: Divide the annotated crosswalk dataset into a training set and a validation set according to a preset ratio. Use the training set to train the crosswalk recognition model to obtain the trained crosswalk model. Use the validation set and the evaluation model to analyze the accuracy of the trained crosswalk model. If the accuracy of the trained crosswalk model meets the requirements, use the trained crosswalk model as the crosswalk inference model for model inference; if the accuracy of the trained crosswalk model does not meet the requirements, re-obtain the trained crosswalk model.

[0104] Specifically, the specific value of the preset ratio is set according to the actual situation. This application preferably uses: training set: validation set = 8:2, that is, 80% of the annotated crosswalk dataset is used for training, and 20% of the annotated crosswalk dataset is used for model validation.

[0105] T4: Use the crosswalk inference model to perform crosswalk recognition on the images collected by the camera in real time to obtain the crosswalk recognition result.

[0106] Furthermore, deploy the recognition algorithm in TensorRT (Tensor Inference Optimizer) to call GPU resources to accelerate the inference process of the crosswalk inference model for performing crosswalk recognition on the images collected by the camera in real time to obtain the crosswalk recognition result.

[0107] Furthermore, select the single-stage object detection algorithm YOLOv5 based on deep learning to identify crosswalks, but it is not limited to selecting the single-stage object detection algorithm YOLOv5 based on deep learning to identify crosswalks.

[0108] Specifically, when selecting the single-stage object detection algorithm YOLOv5 based on deep learning to identify crosswalks, only the border of the crosswalk (i.e., the detection box) needs to be detected, and it is not necessary to label attributes like traffic light recognition to be used as the final crosswalk recognition result.

[0109] S123: Perform recognition on traffic signs and markings to obtain a lane line dataset, an annotated lane line dataset, and a lane line inference model.

[0110] Furthermore, the sub-steps for performing recognition on traffic signs and markings to obtain a lane line dataset, an annotated lane line dataset, and a lane line inference model are as follows:

[0111] U1: Collect the lane line dataset.

[0112] Furthermore, the lane line dataset is the BDD100K (Berkeley DeepDrive 100K) dataset, but it is not limited to the BDD100K dataset. This application preferably uses: the BDD100K dataset. The BDD100K dataset contains more samples and scenarios.

[0113] U2: Process the lane line dataset for annotation to obtain the annotated lane line dataset. The annotation content for the annotation process includes at least: weather conditions, scene locations, and clarity.

[0114] U3: Divide the annotated lane line dataset into a training set, a validation set, and a test set according to a preset ratio. Use the training set to train the lane line recognition model to obtain the trained lane line model. Analyze the accuracy of the trained lane line model using the validation set, the test set, and the evaluation model. If the accuracy of the trained lane line model meets the requirements, use the trained lane line model as the lane line inference model for model inference; if the accuracy of the trained lane line model does not meet the requirements, obtain the trained lane line model again.

[0115] Specifically, the specific value of the preset ratio is set according to the actual situation. This application preferably uses: training set: validation set: test set = 7:1:2, that is, 70% of the annotated lane line dataset is used for training, 10% of the annotated lane line dataset is used for model validation, and 20% of the annotated lane line dataset is used for model testing.

[0116] U4: Use the lane line inference model to perform lane line recognition on the images collected by the camera in real time to obtain the lane line recognition result.

[0117] Furthermore, deploy the recognition algorithm in TensorRT (Tensor Inference Optimizer) to call GPU resources to accelerate the inference process of the lane line inference model for performing lane line recognition on the images collected by the camera in real time to obtain the lane line recognition result.

[0118] Furthermore, select the YOLOPv2 algorithm (You Only Look Once for Panoptic Driving Perception, single-shot panoramic driving perception algorithm version 2) to recognize lane lines, but not limited to selecting the YOLOPv2 algorithm to recognize lane lines. YOLOPv2 is suitable for real-time panoramic driving perception in autonomous driving scenarios and can perform three different tasks simultaneously, namely vehicle detection, drivable area segmentation, and lane line segmentation.

[0119] S124: Use the traffic light image dataset, the crosswalk image dataset, and the lane line dataset as sample data, and use the annotated traffic light dataset, the annotated crosswalk dataset, and the annotated lane line dataset as weight data; constitute the AI dataset from the sample data and the weight data.

[0120] Specifically, the AI dataset is a data set used to train and validate an artificial intelligence model. The AI dataset at least includes: sample pictures collected by an in-vehicle camera and a weight file obtained after artificial annotation training on the sample pictures. This application makes full use of the existing dataset for data distillation, collects a large amount of sample data using the camera mounted on the driverless vehicle, and performs calibration to obtain weight data, which can ensure the recognition accuracy.

[0121] S125: Use the traffic light inference model, the crosswalk inference model, and the lane line inference model as the traffic sign and marking recognition model.

[0122] Specifically, autonomous driving requires that the AI models (i.e., the traffic light inference model, the crosswalk inference model, and the lane line inference model) have a short inference time and a high recognition accuracy.

[0123] S13: Locate the traffic signs and markings to obtain the camera and lidar fusion positioning, where the camera and lidar fusion positioning includes: the traffic light positioning strategy, the crosswalk positioning strategy, and the lane line positioning strategy.

[0124] Specifically, the result recognized in the camera picture is the image coordinate system. After the picture correction is made according to the internal parameters of the camera, there will still be different degrees of distortion and it cannot be directly used for positioning. After the three-dimensional point cloud collected by the lidar (i.e., the original point cloud data) is converted, it has the three-dimensional coordinates in the high-precision world coordinate system. Using the calibration result of the lidar and the camera, project the three-dimensional point cloud image at the time of photo imaging onto the recognition image, and calculate the two-dimensional coordinates of the original point cloud data in the image plane coordinate system. The dimensionality reduction process from three dimensions to two dimensions has a small accuracy loss and can be used to position the recognition result. The projection error mainly comes from the errors of the internal parameters and external parameters measured by the calibration. Only relying on the calibration is difficult to reduce the error at a long distance. Therefore, the error is controlled by screening the pictures during positioning.

[0125] The optical camera is based on the central imaging principle and has the characteristics that the image distortion near the image principal point is small and the image distortion near the lens is small. Based on these two characteristics, different positioning strategies are formulated for different application scenarios of high-precision map making and autonomous driving perception.

[0126] Furthermore, the positioning algorithm process of the traffic light positioning strategy is as follows:

[0127] E1: Obtain the traffic light recognition result and the laser point cloud data at the time of imaging.

[0128] Specifically, the camera outputs images at a frequency of 30 Hz, and the lidar outputs lidar point cloud data at 10 - 20 Hz. In the ROS system (Robot Operating System), the images are read at a frequency of 30 Hz to obtain the lidar point cloud data collected most recently closest to the image acquisition time. The time difference between the two is less than 100 milliseconds.

[0129] To create a high-precision map, the coordinates of the center point of the traffic light and the length and width of the traffic light are required, and a high precision requirement is imposed on the center point. Therefore, it is necessary to screen the images participating in the positioning. When the driverless vehicle is driving, it is preferable to use the images that are closer to the traffic light in the front of the lens and can be covered by both the camera and the lidar for positioning.

[0130] Apply the traffic light recognition result and the traffic light positioning result to the scenario of autonomous driving perception. Given the positions of the traffic lights on the road from the high-precision map deployed in the driverless vehicle, if there are multiple traffic lights in the image, it is necessary to select the traffic light closest to the traffic light in the high-precision map, and use the traffic light closest to the traffic light in the high-precision map to identify the traffic light recognition result (i.e., traffic signal information). Among them, the traffic light recognition result includes: straight-ahead signal light, left-turn signal light, and right-turn signal light.

[0131] Specifically, the traffic light recognition result is obtained through a traffic light inference model.

[0132] E2: Input the lidar point cloud data at the time of imaging, the internal parameter matrix and the external parameter matrix of the camera, and project it onto the image coordinate system using the calculation formula for projecting the three-dimensional lidar point cloud in the lidar coordinate system into the image plane coordinate points of the image plane coordinate system through the projection matrix to obtain the traffic light image coordinates (u, v).

[0133] E3: Traverse the lidar point cloud data.

[0134] Specifically, the lidar point cloud data is the data output by the lidar. The purpose is to convert the three-dimensional coordinates of the lidar point cloud data into the two-dimensional coordinates of the image, that is, to obtain the pixel position of the point cloud in the photo.

[0135] E4: Use the traffic light image coordinates (u, v) to determine whether the lidar point cloud data is within the camera imaging range. If so, use this lidar point cloud data as the filtered lidar point cloud data and execute E5; if not, end the process.

[0136] Specifically, given the resolution of a known camera, such as 1920×1080, it means that the value range of u is [0, 1920] and the value range of v is [0, 1080]. Only a part of the laser point cloud data captured by the lidar is within the imaging space of the camera. The purpose of using the traffic light image coordinates (u, v) to determine whether the laser point cloud data is within the camera imaging range (i.e., the value ranges of u and v) is to filter out the laser point cloud data outside the photo, reducing the subsequent calculation amount. The filtered laser point cloud data is the laser point cloud data located within the camera imaging range.

[0137] E5: Determine whether the filtered laser point cloud data is within the detection frame range. If so, use the filtered laser point cloud data as the extracted laser point cloud data and execute E6; if not, end the process.

[0138] Specifically, the detection frame is the image plane coordinate value and is part of the AI recognition result. The detection frame range is obtained through the traffic sign and marking recognition model. The purpose of determining whether the filtered laser point cloud data is within the detection frame range is to extract the laser point cloud data within the image recognition range as the extracted laser point cloud data.

[0139] E6: Mark the extracted laser point cloud data as recognized point cloud.

[0140] Specifically, the purpose of marking the extracted laser point cloud data as recognized point cloud is to distinguish whether it is the laser point cloud data within the detection frame range. Design the point cloud attributes in advance in the programming language, and only update the attribute values of the point cloud attributes to obtain the recognized point cloud. The purpose is to convert the extraction result (i.e., the extracted laser point cloud data) into information that is easy to process by the programming language.

[0141] E7: Filter the recognized point cloud to obtain the filtered point cloud.

[0142] Furthermore, use the conditional filtering method and the statistical outlier filtering method to filter the recognized point cloud to obtain the filtered point cloud, but it is not limited to using the conditional filtering method and the statistical outlier filtering method to filter the recognized point cloud to obtain the filtered point cloud.

[0143] E8: Cluster the filtered point cloud to obtain the clustered point cloud.

[0144] Furthermore, use the four-neighborhood clustering algorithm to cluster the filtered point cloud to obtain the clustered point cloud, but it is not limited to the four-neighborhood clustering algorithm.

[0145] Specifically, use the four-neighborhood clustering algorithm to cluster the laser point cloud data marked within the detection frame, extract the clustering border, which constitutes part of the recognition and positioning result.

[0146] E9: Obtain the geometric center coordinates of the clustered point cloud as the traffic light positioning result and end the process.

[0147] Specifically, the traffic light positioning result is used for making high-precision maps and traffic decision-making during autonomous driving planning. Among them, when making a high-precision map, the traffic light positioning result is directly used as the center point coordinates of the traffic lights. When used for traffic decision-making during autonomous driving planning, the positions of the traffic lights on the driving route are known, but their signal states are unknown. By identifying the positions of the traffic lights, it is determined which traffic light on the driving route it belongs to, and real-time traffic signals are provided for it.

[0148] Furthermore, the positioning algorithm process of the crosswalk positioning strategy is as follows:

[0149] W1: Obtain the crosswalk recognition result and the laser point cloud data at the time of imaging.

[0150] Specifically, the camera outputs images at a frequency of 30 hz, and the lidar outputs laser point cloud data at 10 - 20 h. The images are read at a frequency of 30 hz within the ros system (Robot Operating System), and the laser point cloud data collected most recently at a time close to the image acquisition time is obtained. The time difference between the two is less than 100 milliseconds.

[0151] Obtain the crosswalk recognition result through the crosswalk inference model.

[0152] W2: Input the laser point cloud data at the time of imaging, the internal parameter matrix and external parameter matrix of the camera, and project it into the image coordinate system using the calculation formula that projects the three-dimensional laser point cloud in the lidar coordinate system into the image plane coordinate points of the image plane coordinate system through the projection matrix to obtain the crosswalk image coordinates (u, v).

[0153] W3: Traverse the laser point cloud data.

[0154] Specifically, the laser point cloud data is the data output by the lidar. The purpose is to convert the three-dimensional coordinates of the laser point cloud data into the two-dimensional coordinates of the image, that is: obtain the pixel position of the point cloud in the photo.

[0155] W4: Use the crosswalk image coordinates (u, v) to determine whether the laser point cloud data is within the camera imaging range. If so, use this laser point cloud data as the filtered laser point cloud data and execute E5; if not, end the process.

[0156] Specifically, given the resolution of the camera, such as 1920×1080, it means that the value range of u is [0, 1920], and the value range of v is [0, 1080]. Only a part of the lidar point cloud data captured is within the imaging space of the camera. The purpose of using the pedestrian crossing image coordinates (u, v) to determine whether the lidar point cloud data is within the camera imaging range (i.e., the value ranges of u and v) is to filter out the lidar point cloud data outside the image, reducing the subsequent calculation amount. The filtered lidar point cloud data is the lidar point cloud data located within the camera imaging range.

[0157] W5: Determine whether the filtered lidar point cloud data is within the detection frame range. If so, use this filtered lidar point cloud data as the extracted lidar point cloud data and execute E6; if not, end the process.

[0158] Specifically, the detection frame is the image plane coordinate value and is part of the AI recognition result. The detection frame range is obtained through the traffic sign and marking recognition model. The purpose of determining whether the filtered lidar point cloud data is within the detection frame range is to extract the lidar point cloud data within the image recognition range as the extracted lidar point cloud data.

[0159] W6: Mark the extracted lidar point cloud data as the recognized point cloud.

[0160] Specifically, the purpose of marking the extracted lidar point cloud data as the recognized point cloud is to distinguish whether it is the lidar point cloud data within the detection frame range. Design the point cloud attributes in advance in the programming language, and only update the attribute values of the point cloud attributes to obtain the recognized point cloud. The purpose is to convert the extraction result (i.e., the extracted lidar point cloud data) into information that is easy to process by the programming language.

[0161] W7: Filter the recognized point cloud to obtain the filtered point cloud.

[0162] Furthermore, use the conditional filtering method and the statistical outlier filtering method to filter the recognized point cloud to obtain the filtered point cloud, but it is not limited to using the conditional filtering method and the statistical outlier filtering method to filter the recognized point cloud to obtain the filtered point cloud.

[0163] W8: Cluster the filtered point cloud to obtain the clustered point cloud.

[0164] Furthermore, use the four-neighborhood clustering algorithm to cluster the filtered point cloud to obtain the clustered point cloud, but it is not limited to the four-neighborhood clustering algorithm.

[0165] Specifically, use the four-neighborhood clustering algorithm to cluster the lidar point cloud data marked within the detection frame, extract the clustering border, which constitutes part of the recognition and positioning result.

[0166] W9: Obtain the geometric center coordinates of the clustered point cloud as the pedestrian crossing positioning result and end the process.

[0167] Specifically, when the crosswalk positioning result is used to create a high-precision map, the center point coordinates of the crosswalk are utilized, and in combination with the lane width, the crosswalk in the high-precision map can be obtained.

[0168] Furthermore, the positioning algorithm process of the lane line positioning strategy is as follows:

[0169] P1: Obtain the lane line detection result and the laser point cloud during imaging.

[0170] Specifically, the camera outputs images at a frequency of 30 hz, and the lidar outputs laser point cloud data at a frequency of 10 - 20 h. In the ros system (Robot Operating System), the images are read at a frequency of 30 hz to obtain the laser point cloud data collected most recently at a time close to the image acquisition time, and the time difference between the two is less than 100 milliseconds.

[0171] Obtain the lane line detection result through the lane line inference model.

[0172] P2: Input the laser point cloud data during imaging, the internal parameter matrix and external parameter matrix of the camera, and project it onto the image coordinate system using the calculation formula for projecting the three-dimensional laser point cloud in the lidar coordinate system into the image plane coordinate points in the image plane coordinate system to obtain the lane line image coordinates (u, v).

[0173] P3: Traverse the laser point cloud.

[0174] Specifically, the laser point cloud data is the data output by the lidar. The purpose is to convert the three-dimensional coordinates of the laser point cloud data into the two-dimensional coordinates of the image, that is: obtain the pixel position of the point cloud in the photo.

[0175] P4: Use the lane line image coordinates (u, v) to determine whether it is within the camera imaging range. If so, use this laser point cloud data as the filtered laser point cloud data and execute E5; if not, end the process.

[0176] Specifically, given the resolution of the camera, such as 1920×1080, it means that the value range of u is [0, 1920], and the value range of v is [0, 1080]. Only a part of the laser point cloud data captured by the lidar is within the imaging space of the camera. The purpose of using the lane line image coordinates (u, v) to determine whether the laser point cloud data is within the camera imaging range (that is: the value range of u and the value range of v) is to filter out the laser point cloud data outside the photo and reduce the subsequent calculation amount. The filtered laser point cloud data is the laser point cloud data located within the camera imaging range.

[0177] P5: Determine whether the filtered laser point cloud data is within the detection frame range. If so, use this filtered laser point cloud data as the extracted laser point cloud data and execute E6; if not, end the process.

[0178] Specifically, the detection box is the image plane coordinate value, which is a part of the AI recognition result. The range of the detection box is obtained through the traffic sign and marking recognition model. The purpose of determining whether the filtered laser point cloud data is within the range of the detection box is to extract the laser point cloud data within the image recognition range as the extracted laser point cloud data.

[0179] P6: Mark the extracted laser point cloud data as recognized point cloud.

[0180] Specifically, the purpose of marking the extracted laser point cloud data as recognized point cloud is to distinguish whether it is the laser point cloud data within the detection box range. The point cloud attributes are pre-designed in the programming language, and the recognized point cloud can be obtained by simply updating the attribute values of the point cloud attributes. The purpose is to convert the extraction result (i.e., the extracted laser point cloud data) into information that is easy to process by the programming language.

[0181] P7: Filter the recognized point cloud to obtain the filtered point cloud.

[0182] Furthermore, the conditional filtering method and the statistical outlier filtering method are used to filter the recognized point cloud to obtain the filtered point cloud, but it is not limited to using the conditional filtering method and the statistical outlier filtering method to filter the recognized point cloud to obtain the filtered point cloud.

[0183] P8: Cluster the filtered point cloud to obtain the clustered point cloud.

[0184] Furthermore, the four-neighborhood clustering algorithm is used to cluster the filtered point cloud to obtain the clustered point cloud, but it is not limited to the four-neighborhood clustering algorithm.

[0185] Specifically, the four-neighborhood clustering algorithm is used to cluster the laser point cloud data marked within the detection box, and the clustering border is extracted, which constitutes a part of the recognition and positioning result.

[0186] P9: Obtain the geometric center coordinates of the clustered point cloud as the lane line positioning result, and end the process.

[0187] Specifically, when making a high-precision map, the lane lines obtained according to the lane line positioning result can directly form the lanes in the high-precision map. When performing path planning for autonomous driving, the drivable area around the vehicle is generated in real time through the lane line positioning result.

[0188] Specifically, in the camera and lidar fusion positioning, for the laser point cloud data participating in the positioning, first, the conditional filtering and statistical outlier filtering methods are used to remove the noise points and the laser point cloud data outside the camera imaging range; then, the laser point cloud data screened within the detection box is filtered again to remove the noise points within the detection box, which can improve the clustering effect.

[0189] S2: Based on the completed AI dataset, when the autonomous driving vehicle is driving on the road, it uses the camera to capture the video stream image in front, and adopts a traffic sign and marking recognition model to identify the traffic signs and markings in the video stream image, and obtains the coordinates of the detection frame of the traffic signs and markings in the image plane coordinate system; among them, the traffic signs and markings include: traffic lights, crosswalks, and lane lines.

[0190] Specifically, the specific steps of adopting a traffic sign and marking recognition model to identify the traffic signs and markings in the video stream image include: using a traffic light inference model to perform traffic light recognition on the image captured by the camera in real time to obtain a traffic light recognition result; using a crosswalk inference model to perform crosswalk recognition on the image captured by the camera in real time to obtain a crosswalk recognition result; using a lane line inference model to perform lane line recognition on the image captured by the camera in real time to obtain a lane line recognition result.

[0191] S3: Use multiple lidars to collect raw point cloud data, and use camera-lidar fusion positioning to analyze the raw point cloud data to obtain a positioning result, where the positioning result includes: traffic light positioning result, crosswalk positioning result, and lane line positioning result.

[0192] Specifically, obtain the world coordinates of the raw point cloud data according to the mapping matrix between the lidar coordinate system and the vehicle body coordinate system, and according to the projection matrix between the lidar and the camera (that is, the transformation formula from the lidar coordinate system to the image plane coordinate system obtained according to the internal parameters and external parameters) and the detection frame coordinates in the image plane coordinate system, filter out the lidar point cloud data within the detection frame range from the raw point cloud data to obtain the lidar point cloud data within the corresponding type of traffic sign and marking range. The filtered lidar point cloud data finally undergoes filtering and clustering processing, and then the recognition and positioning can be automatically and quickly completed to obtain the corresponding positioning result.

[0193] Among them, the mapping matrix between the lidar coordinate system and the world coordinate system: The inertial navigation unit measures the world coordinates of the vehicle center, and the installation position of the lidar on the vehicle is fixed. The position and angle offset from the lidar center to the vehicle center can be obtained through the vehicle digital model or direct measurement, that is, the transformation matrix between the lidar coordinate system and the vehicle body coordinate system.

[0194] As Figure 4 shown, as an embodiment, obtain the coordinates of the detection frame of the picture (for example: video stream image) in the image plane coordinate system and convert them into the coordinates in the vehicle body coordinate system; convert the lidar point cloud data (lidar point cloud) in the lidar coordinate system into the coordinates in the vehicle body coordinate system; convert the coordinates in the vehicle body coordinate system into the coordinates in the projection coordinate system to obtain the recognition result and the positioning result, and use the recognition result and the positioning result for the inertial navigation unit and the high-precision map.

[0195] The beneficial effects achieved by this application are as follows:

[0196] (1) The method and system for identifying and positioning road traffic signs and markings for autonomous driving in this application make full use of the cameras, lidars, and inertial navigation integrated in intelligent connected unmanned vehicles to achieve identification and positioning, without the need to additionally install sensor devices.

[0197] (2) The system for identifying and positioning road traffic signs and markings for autonomous driving in this application has the advantage of low hardware costs. Specifically, the main hardware cost in the system comes from mechanical lidars. The cost of the 16-line mechanical lidar or solid-state lidar used in this application is much lower than that of other mechanical lidars with more than 32 lines.

[0198] (3) The method and system for identifying and positioning road traffic signs and markings for autonomous driving in this application improve the recognition accuracy and reliability. The recognition accuracy rate for various common road traffic signs and markings is not less than 85%.

[0199] (4) The method and system for identifying and positioning road traffic signs and markings for autonomous driving in this application have high positioning accuracy. The positioning accuracy of the center point of the detection frame is higher than 10 cm.

[0200] (5) The method and system for identifying and positioning road traffic signs and markings for autonomous driving in this application shorten the inference time. In the Jetson AGX Orin 32GB processor, the inference time is less than 100 ms.

[0201] (6) The method and system for identifying and positioning road traffic signs and markings for autonomous driving in this application improve the productivity of high-precision map production. The generated high-precision map can be directly used, or it can quickly generate usable data results, greatly shortening the production cycle of high-precision maps and reducing costs.

[0202] Although the preferred embodiments of this application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the protection scope of this application is intended to cover the preferred embodiments and all changes and modifications that fall within the scope of this application. Obviously, those skilled in the art can make various changes and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the protection of this application and its equivalent technologies, this application also intends to include these changes and variations.

Claims

1. A method for identifying and positioning road traffic signs and markings for autonomous driving, characterized in that, It includes the following steps: S1: Obtain the internal parameters of the camera, the external parameters between the camera and the lidar, the AI dataset, the traffic sign and marking recognition model, and the fusion positioning of the camera and the lidar; S2: On the basis of making the AI dataset, when the autonomous vehicle is driving on the road, use the camera to capture the video stream picture in front, and use the traffic sign and marking recognition model to identify the traffic signs and markings in the video stream picture, and obtain the coordinates of the detection frame of the traffic signs and markings in the image plane coordinate system; Among them, the traffic signs and markings include: traffic lights, crosswalks, and lane lines; S3: Use multiple lidars to collect the original point cloud data, and use the fusion positioning of the camera and the lidar to analyze the original point cloud data, so as to obtain the positioning results, where the positioning results include: traffic light positioning results, crosswalk positioning results, and lane line positioning results.

2. The method for identifying and positioning road traffic signs and markings for autonomous driving according to claim 1, wherein The sub-steps for obtaining the internal parameters of the camera, the external parameters between the camera and the lidar, the AI dataset, the traffic sign and marking recognition model, and the fusion positioning of the camera and the lidar are as follows: S11: Use the joint calibration technology of the camera and the radar to measure the internal parameters of the camera and the external parameters between the camera and the lidar; S12: Construct the AI dataset and the traffic sign and marking recognition model, where the AI dataset includes: sample data and weight data; the traffic sign and marking recognition model includes: traffic light inference model, crosswalk inference model, and lane line inference model; S13: Locate the traffic signs and markings to obtain the fusion positioning of the camera and the lidar, where the fusion positioning of the camera and the lidar includes: traffic light positioning strategy, crosswalk positioning strategy, and lane line positioning strategy.

3. The method for identifying and positioning road traffic signs and markings for autonomous driving according to claim 2, wherein, The sub-steps for using the joint calibration technology of the camera and the radar to measure the internal parameters of the camera and the external parameters between the camera and the lidar are as follows: S111: Set the center of the rear axle of the vehicle body of the autonomous vehicle as the origin of the vehicle body coordinate system. By measuring the data between the installation position of the global navigation satellite system positioning antenna and the center of the rear axle of the vehicle body, the mapping matrix between the vehicle body coordinate system and the inertial navigation measurement unit can be obtained; S112: Measure the data between the installation position of the lidar and the center of the rear axle of the vehicle body, and the mapping matrix between the lidar coordinate system and the vehicle body coordinate system can be obtained; S113: Perform the internal parameter calibration of the camera to obtain the internal parameters of the camera; S114: Perform the joint calibration of the camera and the lidar to obtain the external parameters between the camera and the lidar.

4. The method for identifying and positioning road traffic signs and markings for autonomous driving according to claim 3, wherein, The expression of the internal parameters of the camera is: Among them, fx, fy, u0, and v0 are the internal parameters of the camera; fx is the equivalent focal length of the camera in the x direction; fy is the equivalent focal length of the camera in the y direction; (u0, v0) is the image principal point coordinate; (u, v) is the coordinate of the image plane coordinate system corresponding to the point q in the three-dimensional space of the world coordinate system; (xc, yc, zc) is the coordinate of the point q in the camera coordinate system.

5. The method for identifying and positioning road traffic signs and markings for autonomous driving according to claim 3, wherein The sub-steps for performing the joint calibration of the camera and the lidar to obtain the external parameters between the camera and the lidar are as follows: S1141: Establish a projection matrix for converting the lidar coordinate system to the image plane coordinate system; S1142: Directly measure the extrinsic parameters between the lidar, the vehicle body center, and the inertial navigation unit, and establish the mapping matrix between the lidar coordinate system and the world coordinate system; S1143: After converting the original point cloud data of multiple lidars to the world coordinate system through the mapping matrix between the lidar coordinate system and the world coordinate system, perform the joint calibration of the camera and the lidar to obtain the extrinsic parameters between the camera and the lidar.

6. The method for identifying and positioning road traffic signs and markings for autonomous driving according to claim 2, wherein, The sub-steps for constructing the AI dataset and the traffic sign and marking recognition model are as follows: S121: Recognize the traffic signs and markings to obtain the traffic light image dataset, the labeled traffic light dataset, and the traffic light inference model; S122: Recognize the traffic signs and markings to obtain the crosswalk image dataset, the labeled crosswalk dataset, and the crosswalk inference model; S123: Recognize the traffic signs and markings to obtain the lane line dataset, the labeled lane line dataset, and the lane line inference model; S124: Use the traffic light image dataset, the crosswalk image dataset, and the lane line dataset as sample data, and use the labeled traffic light dataset, the labeled crosswalk dataset, and the labeled lane line dataset as weight data; form the AI dataset from the sample data and the weight data; S125: Use the traffic light inference model, the crosswalk inference model, and the lane line inference model as the traffic sign and marking recognition model.

7. The method for identifying and positioning road traffic signs and markings for autonomous driving according to claim 6, wherein The sub-steps for recognizing the traffic signs and markings to obtain the traffic light image dataset, the labeled traffic light dataset, and the traffic light inference model are as follows: R1: Collect the traffic light image dataset; R2: Perform annotation processing on the traffic light image dataset to obtain the labeled traffic light dataset, where the annotation content of the annotation processing includes at least: the traffic light detection frame, the traffic light type, and the traffic light driving direction; R3: Divide the labeled traffic light dataset into a training set and a validation set according to a preset ratio, use the training set to train the traffic light recognition model to obtain the trained traffic light model, use the validation set and the evaluation model to analyze the accuracy of the trained traffic light model. If the accuracy of the trained traffic light model meets the requirements, use the trained traffic light model as the traffic light inference model for model inference; if the accuracy of the trained traffic light model does not meet the requirements, re-obtain the trained traffic light model.

8. The method for identifying and positioning road traffic signs and markings for autonomous driving according to claim 2, wherein, The positioning algorithm process of the traffic light positioning strategy is as follows: E1: Obtain the traffic light recognition result and the laser point cloud data at the time of imaging; E2: Input the laser point cloud data at the time of imaging, the internal parameter matrix and the external parameter matrix of the camera, and project it to the image coordinate system using the calculation formula for projecting the three-dimensional laser point cloud in the lidar coordinate system into the image plane coordinate points of the image plane coordinate system through the projection matrix to obtain the traffic light image coordinates; E3: Traverse the laser point cloud data; E4: Use the traffic light image coordinates to determine whether the laser point cloud data is within the camera imaging range. If so, use the laser point cloud data as the filtered laser point cloud data and execute E5; if not, end the process; E5: Determine whether the filtered lidar point cloud data is within the detection frame. If so, use the filtered lidar point cloud data as the extracted lidar point cloud data and execute E6; if not, end the process. E6: Mark the extracted lidar point cloud data as the recognized point cloud. E7: Filter the recognized point cloud to obtain the filtered point cloud. E8: Cluster the filtered point cloud to obtain the clustered point cloud. E9: Obtain the geometric center coordinates of the clustered point cloud as the traffic light positioning result and end the process. Among them, the calculation formula for projecting the three-dimensional lidar point cloud in the lidar coordinate system into the image plane coordinate points in the image plane coordinate system through the projection matrix is: Among them, is the internal parameter matrix of the camera; is the external parameter matrix; (u, v) is the coordinate of the image plane coordinate system corresponding to the point q in the three-dimensional space under the world coordinate system; fx is the equivalent focal length of the camera in the x direction; fy is the equivalent focal length of the camera in the y direction; (u0, v0) is the principal point coordinate of the image; Tk is the translation vector from the lidar coordinate system to the camera coordinate system; Rk is the rotation matrix from the lidar coordinate system to the camera coordinate system; is a 3×3 zero matrix; X is the value of any point in the lidar point cloud data in the X-axis direction of the lidar coordinate system; Y is the value of any point in the lidar point cloud data in the Y-axis direction of the lidar coordinate system; Z is the value of any point in the lidar point cloud data in the Z-axis direction of the lidar coordinate system.

9. A road traffic sign and marking recognition and positioning system for autonomous driving, characterized in that, Including: An autonomous vehicle, at least one camera, at least one lidar, at least one set of inertial navigation measurement units, and an identification and positioning unit; at least one camera, at least one lidar, at least one set of inertial navigation measurement units, and the identification and positioning unit are all mounted on the autonomous vehicle. Among them, the camera: is used to collect image data, where the image data includes: video stream images. The lidar: is used to collect the original point cloud data. The inertial navigation measurement unit: is used to measure the coordinates of the phase center of the global navigation satellite system positioning antenna in the world coordinate system and the azimuth angle of the vehicle attitude in real time. The identification and positioning unit: is used to execute the method for identifying and positioning road traffic signs and markings for autonomous driving described in any one of claims 1-8.

10. The road traffic sign and marking recognition and positioning system for autonomous driving according to claim 9, characterized in that, The lidar includes: at least one 16-line mechanical lidar and at least one solid-state lidar. Among them, the 16-line mechanical lidar: is used to collect the lidar point cloud data of traffic signs and markings within the first line-of-sight range, the first horizontal field of view range, and the first vertical field of view range. The solid-state lidar: is used to collect the lidar point cloud data of traffic signs and markings within the second line-of-sight range, the second horizontal field of view range, and the second vertical field of view range.

Citation Information

Patent Citations

  • Cognitive map construction method fusing image and laser point cloud

    CN116310176A

  • Multi-sensor joint calibration method for automatic driving perception

    CN118259270A

Cited By

  • Road marking thickness continuous detection method and system

    CN122384687A