Ground identification detection method and vehicle
By mapping image features to the BEV space and considering positional offset, and using LiDAR point cloud features to train a model for position correction and feature fusion, the problem of inaccurate detection caused by camera extrinsic jitter is solved, and higher accuracy ground marker detection is achieved.
Patent Information
- Application Number
- CN202511411878.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-01-09
AI Technical Summary
During vehicle operation, image stitching misalignment caused by camera extrinsic jitter leads to inaccurate detection of ground markings.
By mapping image features to the BEV space, considering the positional offset on each feature map, and using LiDAR point cloud features to train an offset parameter model, position correction and feature fusion are performed to generate more accurate target bird's-eye view features for detection.
It improves the accuracy of ground marking detection, avoids detection errors caused by external parameter fluctuations, and enhances the accuracy rate of ground marking detection.
Smart Images

Figure CN121305495A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of target detection technology, and in particular to a ground marking detection method and vehicle. Background Technology
[0002] With the development of science and technology, intelligent driving technology is becoming increasingly mature, and vehicles are offering more and more intelligent driving functions, such as memory parking, valet parking, and urban intelligent driving, bringing great convenience to users. A crucial aspect of intelligent driving scenarios is the detection of ground markings (straight lane markings, left turn lane markings, right turn lane markings, non-motorized vehicle lane markings, etc.).
[0003] In related technologies, multiple cameras are installed on the vehicle body to collect images of the vehicle's surroundings. Then, the images collected by each camera are stitched together using the intrinsic and extrinsic parameters of each camera to obtain a fused image. Finally, ground markings are detected on the fused image to identify the corresponding ground markings.
[0004] However, since the above method stitches together the images captured by each camera using the intrinsic and extrinsic parameters of each camera to obtain the fused image, if the vehicle is moving unstably, the extrinsic parameters of the cameras are prone to shaking, resulting in inaccurate extrinsic parameters. Therefore, the image obtained by stitching together the images captured by each camera using the extrinsic and extrinsic parameters of each camera will have serious misalignment, which will lead to inaccurate ground marking detection. Summary of the Invention
[0005] This application provides a ground marking detection method, apparatus, vehicle, and storage medium. It maps image features to a BEV (Battery Electric Vehicle) space to obtain bird's-eye view features. During the conversion of P feature maps to the BEV space, it considers potential positional offsets at various locations on each feature map to obtain more accurate target bird's-eye view features, thereby enabling more accurate detection of ground markings. The technical solution includes the following:
[0006] Firstly, a method for detecting ground markings is provided, the method comprising: Feature extraction is performed on M target images to obtain P feature maps, where the M target images are images captured by M cameras on the vehicle. Based on the P feature maps and the calibration parameters of the M cameras, P target offset parameters are determined. The P target offset parameters are used to represent the possible positional offset of each pixel on each feature map. Based on the P feature maps, the P target offset parameters, and multiple first projection points, the target bird's-eye view features are determined. The multiple first projection points are the projection points of multiple coordinate points in the target bird's-eye view space onto the P feature maps respectively. The target bird's-eye view features are the features of the M target images in the target bird's-eye view space. Based on the features of the target bird's-eye view, ground markings are detected to obtain the ground marking information of the current driving road.
[0007] In this application, features are first extracted from M target images to obtain P feature maps. Then, based on the P feature maps and the calibration parameters of the M cameras on the vehicle, P target offset parameters are determined, that is, the possible positional offsets at each position on each feature map are first determined. Next, based on the P feature maps, the P target offset parameters, and multiple first projection points, the target bird's-eye view features are determined. This involves converting the P feature maps obtained from feature extraction into the BEV space, taking into account the possible positional offsets at each position on each feature map during the conversion process to obtain more accurate target bird's-eye view features. Subsequently, based on the target bird's-eye view features, ground marking detection is performed to obtain the ground marking information of the current driving road. This allows for more accurate ground marking detection based on the target bird's-eye view features. Compared to the problem of inaccurate detection caused by image misalignment due to external parameter jitter in the existing technology, this solution takes into account the impact of external parameter jitter on the positional offset of the feature map during the extraction of BEV spatial features. This avoids the problem of inaccurate image features caused by external parameter jitter, which leads to inaccurate ground marking detection, thereby improving the accuracy of ground marking detection.
[0008] Optionally, determining the P target offset parameters based on the P feature maps and the calibration parameters of the M cameras includes: For the i-th camera among the M cameras, feature extraction is performed on the calibration parameters of the i-th camera to obtain the calibration parameter features of the i-th camera; For the j-th feature map among the P feature maps, the j-th feature map is fused with the calibration parameter features of the target camera to obtain the j-th fused feature. The target image corresponding to the j-th feature map is acquired by the target camera among the M cameras. Based on the j-th fusion feature, the j-th target offset parameter among the P target offset parameters is determined.
[0009] In the above method, since the camera calibration parameters are used to establish the mapping relationship between spatial points in the real world and pixels in the image coordinate system, they represent a kind of imaging geometric information, while the feature map is image content information, used to reflect the relative relationship between pixels. Based on the feature after the fusion of the two, the position offset is determined, which can comprehensively measure the true mapping of spatial points in the real world in the image. Thus, the possible position offset of each pixel in the feature map can also be accurately determined, that is, the accurate target offset parameters can be determined.
[0010] Optionally, determining the j-th target offset parameter among the P target offset parameters based on the j-th fusion feature includes: The j-th fusion feature is input into the offset parameter determination model, and the offset parameter determination model outputs the j-th target offset parameter. The offset parameter determination model is trained based on the difference between real point cloud features and sample image features. The real point cloud features are obtained by feature extraction from point clouds acquired by LiDAR, and the sample image features are obtained by feature extraction from sample images acquired by camera.
[0011] In the above method, since the point cloud acquired by LiDAR can accurately reflect the vehicle's surrounding environment, and LiDAR does not suffer from extrinsic parameter jitter, the pixels in the features extracted from the point cloud acquired by LiDAR will not experience positional shifts. In other words, the real point cloud features can accurately reflect the mapping between spatial points in the real environment and pixels in the image, thus accurately reflecting the pixel's position in the image. In this case, based on the difference between the real point cloud features and the sample image features, an accurate prediction of the possible positional shift of the pixel can be obtained.
[0012] Optionally, the method further includes: For any one of the plurality of coordinate points, based on the calibration parameters of the M cameras, the coordinate point is projected onto the P feature maps respectively to obtain the P first projection points corresponding to the coordinate point.
[0013] Optionally, determining the target bird's-eye view features based on the P feature maps, P target offset parameters, and multiple first projection points includes: For the j-th first projection point among the P first projection points, the position of the j-th first projection point is corrected based on the j-th target offset parameter to obtain N second projection points; Obtain the first pixel feature on the N second projection points from the j-th feature map; Based on the P target offset parameters and the first pixel features on the N second projection points corresponding to each of the P first projection points, the bird's-eye view features of the coordinate points are determined. The bird's-eye view features of the multiple coordinate points are fused to obtain the target bird's-eye view features.
[0014] In the above method, during the process of converting P feature maps to the target bird's-eye view space, the position of the projection point on each feature map is corrected. Then, the pixel features on the new projection point obtained after the position correction are combined to determine the bird's-eye view feature of this coordinate point. Finally, the bird's-eye view features of multiple coordinate points are fused to determine a more accurate bird's-eye view feature, which can then be used to detect more accurate ground markings.
[0015] Optionally, the j-th target offset parameter includes N position offsets of each pixel on the j-th feature map, and the step of correcting the position of the j-th first projection point based on the j-th target offset parameter to obtain N second projection points includes: For any one of the N position offsets of the target pixel, add the position offset to the j-th first projection point to obtain a second projection point after correcting the position offset of the j-th first projection point, and the target pixel is the pixel corresponding to the j-th first projection point.
[0016] In the above method, N position offsets of the corresponding pixel are first obtained from the j-th target offset parameters. Then, N position offsets are added to the j-th first projection point, which is equivalent to offsetting the j-th first projection point by N positions to obtain N new positions (N second projection points). The N second projection points are the more realistic pixels of this coordinate point in the target bird's-eye view space projected onto the j-th feature map. Therefore, the above method can obtain more realistic N second projection points, thereby improving the accuracy of subsequent bird's-eye view features.
[0017] Optionally, the j-th target offset parameter includes N first weights and one second weight for each pixel on the j-th feature map, wherein the N first weights correspond one-to-one with the N second projection points, and the second weights correspond to the j-th feature map; determining the bird's-eye view features of the coordinate point based on the P target offset parameters and the first pixel features on the N second projection points corresponding to each of the P first projection points includes: For any one of the N second projection points corresponding to the j-th projection point, the first pixel feature on the second projection point is multiplied by the first weight corresponding to the second projection point to obtain the second pixel feature corresponding to the second projection point. The second pixel features corresponding to each of the N second projection points are summed to obtain the third pixel features. Multiply the third pixel feature by the second weight to obtain the fourth pixel feature of the j-th feature map; Based on the fourth pixel feature of the P feature maps, the bird's-eye view feature of the coordinate point is determined.
[0018] In the above method, by considering the possibility that a second projection point is the true offset of a pixel point during the process of fusing pixel features of different projection points on different feature maps, and the importance of a feature map to the ground marker detection task compared to other feature maps, a more accurate bird's-eye view feature of this coordinate point can be obtained by fusing.
[0019] Optionally, the target bird's-eye view space is composed of K planes, and the K planes include the plurality of coordinate points. The step of fusing the bird's-eye view features of the plurality of coordinate points to obtain the target bird's-eye view features includes: For any one of the K planes, attention encoding is performed on the bird's-eye view features of the coordinate points on the plane to obtain attention features; The attention features are fused with the bird's-eye view features of the coordinate points on the plane to obtain the reference bird's-eye view features of the plane. The target bird's-eye view features are obtained by fusing the reference bird's-eye view features of the K planes.
[0020] In the above method, attention encoding is first applied to the bird's-eye view features of each coordinate point to determine the importance of each coordinate point on a plane. Then, the bird's-eye view features of the coordinate points on that plane are further encoded based on the importance of each coordinate point, making the important features of the plane more prominent. Finally, the reference bird's-eye view features corresponding to each plane are fused, resulting in a fusion of the important features of each of the K planes, thus obtaining a bird's-eye view feature that more accurately represents the ground markings in the BEV space.
[0021] Optionally, the step of extracting features from M target images to obtain P feature maps includes: For the i-th target image among the M target images, the i-th target image is input into a multi-scale feature extraction network, and the multi-scale feature extraction network outputs Q feature maps, the Q feature maps having different sizes, and Q multiplied by M equals P.
[0022] In the above method, for the i-th target image among M target images, the multi-scale feature extraction network is used to extract features from the i-th target image to output Q feature maps of different sizes. This allows us to obtain image features at different levels in the i-th image, that is, to obtain image features extracted under different receptive fields. This enables the joint acquisition of global features and local detail features in the i-th image, thus providing a valid basis for subsequent ground marker detection.
[0023] Secondly, a ground marking detection device is provided, the device comprising: The feature extraction module is used to extract features from M target images respectively to obtain P feature maps, wherein the M target images are images captured by M cameras on the vehicle; The parameter determination module is used to determine P target offset parameters based on the P feature maps and the calibration parameters of the M cameras. The P target offset parameters are used to represent the possible positional offset of each pixel point on each feature map. The feature transformation module is used to determine the target bird's-eye view features based on the P feature maps, the P target offset parameters, and multiple first projection points. The multiple first projection points are the projection points of multiple coordinate points in the target bird's-eye view space onto the P feature maps respectively. The target bird's-eye view features are the features of the P feature maps transformed into the target bird's-eye view space. The detection module is used to detect ground markings based on the features of the target bird's-eye view to obtain ground marking information of the current driving road.
[0024] Optionally, the parameter determination module is specifically used for: For the i-th camera among the M cameras, feature extraction is performed on the calibration parameters of the i-th camera to obtain the calibration parameter features of the i-th camera; For the j-th feature map among the P feature maps, the j-th feature map is fused with the calibration parameter features of the target camera to obtain the j-th fused feature. The target image corresponding to the j-th feature map is acquired by the target camera among the M cameras. Based on the j-th fusion feature, the j-th target offset parameter among the P target offset parameters is determined.
[0025] Optionally, the parameter determination module is specifically used for: The j-th fusion feature is input into the offset parameter determination model, and the offset parameter determination model outputs the j-th target offset parameter. The offset parameter determination model is trained based on the difference between real point cloud features and sample image features. The real point cloud features are obtained by feature extraction from point clouds acquired by LiDAR, and the sample image features are obtained by feature extraction from sample images acquired by camera.
[0026] Optionally, the device further includes: The projection module is used to project any one of the multiple coordinate points onto the P feature maps based on the calibration parameters of the M cameras, thereby obtaining the P first projection points corresponding to the coordinate point.
[0027] Optionally, the feature determination module is specifically used for: For the j-th first projection point among the P first projection points, the position of the j-th first projection point is corrected based on the j-th target offset parameter to obtain N second projection points; Obtain the first pixel feature on the N second projection points from the j-th feature map; Based on the P target offset parameters and the first pixel features on the N second projection points corresponding to each of the P first projection points, the bird's-eye view features of the coordinate points are determined. The bird's-eye view features of the multiple coordinate points are fused to obtain the target bird's-eye view features.
[0028] Optionally, the j-th target offset parameter includes N position offsets of each pixel on the j-th feature map, and the feature determination module is specifically used for: For any one of the N position offsets of the target pixel, add the position offset to the j-th first projection point to obtain a second projection point after correcting the position offset of the j-th first projection point, and the target pixel is the pixel corresponding to the j-th first projection point.
[0029] Optionally, the j-th target offset parameter includes N first weights and one second weight for each pixel on the j-th feature map, wherein the N first weights correspond one-to-one with the N second projection points, and the second weights correspond to the j-th feature map; the feature determination module is specifically used for: For any one of the N second projection points corresponding to the j-th projection point, the first pixel feature on the second projection point is multiplied by the first weight corresponding to the second projection point to obtain the second pixel feature corresponding to the second projection point. The second pixel features corresponding to each of the N second projection points are summed to obtain the third pixel features. Multiply the third pixel feature by the second weight to obtain the fourth pixel feature of the j-th feature map; Based on the fourth pixel feature of the P feature maps, the bird's-eye view feature of the coordinate point is determined.
[0030] Optionally, the target bird's-eye view space consists of K planes, and the K planes include the plurality of coordinate points. The feature determination module is specifically used for: For any one of the K planes, attention encoding is performed on the bird's-eye view features of the coordinate points on the plane to obtain attention features; The attention features are fused with the bird's-eye view features of the coordinate points on the plane to obtain the reference bird's-eye view features of the plane. The target bird's-eye view features are obtained by fusing the reference bird's-eye view features of the K planes.
[0031] Optionally, the feature extraction module is specifically used for: For the i-th target image among the M target images, the i-th target image is input into a multi-scale feature extraction network, and the multi-scale feature extraction network outputs Q feature maps, the Q feature maps having different sizes, and Q multiplied by M equals P.
[0032] Thirdly, a vehicle is provided, the vehicle comprising: Memory, used to store executable program code; A processor is configured to call and run the executable program code from the memory, causing the vehicle to perform the aforementioned ground marking detection method.
[0033] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described ground marking detection method.
[0034] Fifthly, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the steps of the above-described ground marking detection method.
[0035] It is understood that the beneficial effects of the second, third, fourth, and fifth aspects mentioned above can be found in the relevant descriptions in the first aspect above, and will not be repeated here. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a schematic diagram of a scenario for a ground marking detection method provided in an embodiment of this application; Figure 2 This is a flowchart of a ground marking detection method provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a target bird's-eye view space provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a ground marking detection device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0039] It should be understood that "multiple" as mentioned in this application refers to two or more. In the description of this application, unless otherwise stated, " / " indicates "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist, for example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, to facilitate a clear description of the technical solutions of this application, the terms "first," "second," etc., are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and that "first," "second," etc., do not necessarily imply differences.
[0040] First, the terms used in the embodiments of this application will be explained.
[0041] 1. BEV (Bird Eye View) BEV (Browser Electric Vehicle) refers to the result of digitally presenting the environment from a perspective perpendicular to the ground. It is mainly generated by fusing and transforming raw data collected by multiple sensors on the vehicle through algorithms. By fusing scattered sensor data into a top view covering 360° around the vehicle, it can present the global environment around the vehicle, thereby solving the problem of large blind spots in the driver's field of vision.
[0042] 2. BEV Space BEV space refers to a spatial coordinate system that uses a "vertical, high-altitude perspective" as its core to digitally model the environment surrounding the vehicle. Essentially, it's a right-handed coordinate system with the autonomous vehicle itself as the origin. The X-axis is positive along the vehicle's direction of travel, the Y-axis is perpendicular to the X-axis and positive along the left side of the vehicle, and the Z-axis is perpendicular to the ground and positive upwards. By fusing data from multiple sensors, BEV space integrates 360° environmental information around the vehicle into a complete "top-view map." Whether it's the lane lines in front of the vehicle, the straight-ahead lane markings within the lane lines, or the lane markings in the right lane, everything can be accurately presented in BEV space, completely eliminating blind spots.
[0043] 3. Camera calibration parameters In the embodiments of this application, the calibration parameters of the camera may include camera intrinsic parameters, camera extrinsic parameters, and distortion coefficients.
[0044] Camera intrinsic parameters are parameters that describe the camera's optical and hardware characteristics. They are determined by the camera's manufacturing process and internal structure and are usually fixed. They are mainly used to establish the projection relationship between "three-dimensional coordinate points in the camera coordinate system" and "two-dimensional coordinate points on the image plane". Generally, camera intrinsic parameters are constructed using a 3×3 intrinsic parameter matrix. This matrix mainly consists of key parameters such as focal length in the X-axis direction, focal length in the Y-axis direction, principal point coordinates (coordinates of the intersection of the camera optical axis and the image plane), and tilt factor (used to correct projection distortion caused by sensor axis tilt).
[0045] Camera extrinsic parameters describe the camera's position and orientation in the world coordinate system; that is, they represent the camera's pose in the real world. They change as the camera moves or rotates. Their primary function is to transform 3D coordinate points in the world coordinate system to the camera coordinate system. Generally, camera extrinsic parameters consist of two parts: a rotation matrix and a translation vector. The rotation matrix describes the camera's orientation relative to the world coordinate system (including rotation angles around the X, Y, and Z axes), and the translation vector describes the camera's position relative to the origin of the world coordinate system (including offsets along the X, Y, and Z axes).
[0046] Distortion coefficients are used to correct image distortion caused by the construction of the optical system (the deviation between the projected position of a real-world 3D coordinate point and its ideal position). In autonomous driving scenarios, the original image needs to be corrected using distortion coefficients first. Without correction, the positions of lane lines, obstacles, ground markings, etc., will be deviated, leading to subsequent decision-making errors.
[0047] 4. Spatial attention Spatial attention is an algorithm module that simulates the "selective focusing" mechanism of human vision. Its core idea is to actively focus on spatial locations that are more important to the current task by calculating weights when processing input data (images, image features), while ignoring irrelevant or secondary regions, thereby improving the model's perceptual accuracy and efficiency. Its main purpose is to calculate a weight value for each spatial location in the feature map to determine the importance of each location, thus allowing subsequent feature extraction to focus on features at locations with higher weight values.
[0048] Before describing the ground marking detection method provided in the embodiments of this application, the application scenarios of the embodiments of this application will be explained first.
[0049] For example, Figure 1 This is a schematic diagram of a ground marking detection method provided in an embodiment of this application.
[0050] Figure 1 This illustrates a scenario of automatic vehicle parking, such as... Figure 1 As shown, vehicle 101 is driving in the lane, and vehicle 101 has its automatic parking function activated. Figure 1 The vehicle 101 shown is automatically searching for a parking space. Generally, the internal roads of a parking lot have a specific orientation; following the correct direction will allow for successful parking or finding the exit. Figure 1 As shown, when vehicle 101 is traveling in the lane, the road markings at the intersection ahead indicate a right turn, meaning that the vehicle should turn right at the intersection. In this situation, vehicle 101 needs to detect the right turn sign at the intersection so that it can subsequently follow the directions of the roads within the parking lot.
[0051] In related technologies, the detection of ground markings by vehicle 101 is based on an AVM (Around View Monitor) image. Specifically, fisheye images captured by cameras mounted on vehicle 101 are first projected onto the AVM image. Generally, coordinate transformation of the fisheye images is performed using pre-calibrated camera intrinsic and extrinsic parameters. The transformed images are then directly stitched together to obtain the AVM image. Subsequently, ground marking detection is performed on the AVM image using appropriate detection algorithms to obtain parameters such as the type and location of the ground markings.
[0052] However, if vehicle 101 is traveling on a bumpy road, it may cause slight vibrations in the camera's extrinsic parameters. Since the AVM image is obtained by directly stitching the transformed image using pre-calibrated camera extrinsic and extrinsic parameter data, and generally, the camera extrinsic parameters are factory-calibrated and fixed, in this case, when the actual camera extrinsic parameters change, the pre-calibrated camera extrinsic parameters cannot represent the actual camera extrinsic parameters. Therefore, after performing coordinate transformation on the fisheye image based on the pre-calibrated camera extrinsic and extrinsic parameters, the corresponding pixels may shift, resulting in serious misalignment in the stitched AVM image, and consequently, inaccurate ground marking detection.
[0053] Especially during the process of autonomous driving, if the AVM map shows serious misalignment in the detection of lane lines, stop lines, zebra crossings, etc., it may lead to incorrect detection of ground markings, which may cause corresponding driving safety issues.
[0054] Therefore, this application provides a ground marking detection method, which can be applied to the process of a vehicle detecting ground markings. For example, the ground marking detection method can be applied to the detection of ground markings during automatic parking, or to the detection of ground markings during autonomous driving on urban roads.
[0055] Specifically, M images can be acquired using M cameras on the vehicle. Feature extraction is then performed on each of the M images to obtain P feature maps. Based on these P feature maps and the calibration parameters of the M cameras, P target offset parameters are determined, with one target offset parameter corresponding to one feature map. Each target offset parameter represents the possible positional offset of each pixel on the corresponding feature map. After determining the P target offset parameters, target bird's-eye view features can be determined based on the P feature maps, the P target offset parameters, and multiple first projection points. This means that the possible positional offsets of pixels in the image features are considered during the projection of image features into the target bird's-eye view space, resulting in accurate target bird's-eye view features. Finally, ground marking detection is performed based on these target bird's-eye view features to obtain the ground marking information of the currently traveling road.
[0056] In this case, by performing corresponding ground marking detection based on more accurate target bird's-eye view features, compared with the problem of inaccurate detection caused by image misalignment due to extrinsic parameter jitter in existing technologies, this solution takes into account the impact of extrinsic parameter jitter on the positional offset of the feature map during the extraction of BEV spatial features. This can avoid the problem of inaccurate ground marking detection caused by inaccurate image features due to extrinsic parameter jitter, thereby improving the accuracy of ground marking detection.
[0057] The ground marking detection method provided in the embodiments of this application will be explained in detail below.
[0058] Figure 2 This is a flowchart illustrating a ground marking detection method provided in an embodiment of this application. This method can be applied to a vehicle's Domain Control Unit (DCU). See also... Figure 2 The method includes the following steps.
[0059] Step 201: Extract features from the M target images to obtain P feature maps. The M target images are images captured by the M cameras on the vehicle.
[0060] Where P is an integer greater than or equal to M.
[0061] In this embodiment, M can be set according to the number of cameras that acquire images during the ground marking detection process. For example, a vehicle may be equipped with four fisheye cameras (front, rear, left, and right), a front-view pinhole camera, and a rear-view pinhole camera. In this case, M can be 6, meaning that the M target images are images captured by the four fisheye cameras, the front-view pinhole camera, and the rear-view pinhole camera on the vehicle. Compared to the existing technology that uses only images acquired by four fisheye cameras for ground marking detection, this expands the field of view, allowing for the acquisition of images within a larger field of view. This expands the perception range of ground markings, meaning that ground markings at a distance can also be detected.
[0062] It should be understood that after the M cameras on the vehicle capture images, the captured images can be time-aligned using the time of one of the cameras as the reference, thereby obtaining M images captured at the same time. In the embodiments of this application, the aforementioned M target images are also the images captured by the M cameras at the same time.
[0063] One possible approach is that step 201 can be performed as follows: for the i-th target image among M target images, input the i-th target image into a multi-scale feature extraction network, and output Q feature maps through the multi-scale feature extraction network.
[0064] Where i is an integer greater than or equal to 1 and less than or equal to M. The i-th target image is the image acquired by the i-th camera.
[0065] The Q feature maps have different sizes, and Q multiplied by M equals P. That is, for each of the M target images, there are corresponding Q feature maps. In other words, Q feature maps of different scales can be extracted for each target image.
[0066] The multi-scale feature extraction network is a neural network used to extract features from two-dimensional images. In the embodiments of this application, the multi-scale feature extraction network can be a convolutional neural network, such as, but not limited to, pre-trained ResNet (Residual Network) and DenseNet (Densely Connected Convolutional Networks).
[0067] As an example, this multi-scale feature extraction network can consist of multiple convolutional layers, pooling layers, and activation layers. The convolutional layers are used to extract features at different depths from the target image, including global and local features. The pooling layers are used to reduce the feature dimensionality, thereby reducing computational cost. The activation layers are used to introduce nonlinear transformations, enhancing the network's ability to represent nonlinear relationships.
[0068] In the above structure, as the network depth increases, the receptive field of the convolutional layer gradually expands, and the extracted features will show a change from details to the global. Generally, the deeper the convolutional layer, the smaller the size of the image features extracted by the convolutional layer. In this case, Q feature maps of different scales can be obtained by outputting the image features extracted by Q convolutional layers.
[0069] In the above method, for the i-th target image among M target images, the multi-scale feature extraction network is used to extract features from the i-th target image to output Q feature maps of different sizes. This allows us to obtain image features at different levels in the i-th image, that is, to obtain image features extracted under different receptive fields. This enables the joint acquisition of global features and local detail features in the i-th image, thus providing a valid basis for subsequent ground marker detection.
[0070] Step 202: Based on P feature maps and M camera calibration parameters, determine P target offset parameters. The P target offset parameters are used to represent the possible positional offset of each pixel on each feature map.
[0071] Among them, P target offset parameters and P features Figure 1 One-to-one correspondence, that is, for any feature map, there is a corresponding target offset parameter, which is used to represent the possible positional offset of each pixel on the feature map.
[0072] In the above method, P target offset parameters are determined based on P feature maps and M camera calibration parameters. This allows us to know the possible positional offsets of each pixel on each feature map, thereby determining whether there are pixel offsets in the image features due to external parameter jitter and the specific offset of each pixel in each feature map when external parameter jitter occurs. Subsequently, the position of the corresponding pixel can be corrected to obtain accurate image features.
[0073] One possible approach is that step 202 can be performed as follows: for the i-th camera among the M cameras, feature extraction is performed on the calibration parameters of the i-th camera to obtain the calibration parameter features of the i-th camera; for the j-th feature map among the P feature maps, the j-th feature map is fused with the calibration parameters of the target camera to obtain the j-th fused feature; based on the j-th fused feature, the j-th target offset parameter among the P target offset parameters is determined.
[0074] The target image corresponding to the j-th feature map is acquired by the target camera among the M cameras; that is, it is the target image corresponding to the j-th feature map acquired by the target camera.
[0075] In the above steps, feature extraction is first performed on the calibration parameters of each of the M cameras to obtain the calibration parameter features of each camera. Then, for any feature map among the P feature maps, this feature map is concatenated with the calibration parameter features of the target camera to obtain a fused feature. The target camera is the camera that acquired the target image corresponding to this feature map, and thus the calibration parameter features of the target camera are the calibration parameter features extracted from the calibration parameter features of that camera. After obtaining the fused features corresponding to each feature map, the corresponding target offset parameters can be determined accordingly.
[0076] In the above method, since the camera calibration parameters are used to establish the mapping relationship between spatial points in the real world and pixels in the image coordinate system, they represent a kind of imaging geometric information, while the feature map is image content information, used to reflect the relative relationship between pixels. Based on the feature after the fusion of the two, the position offset is determined, which can comprehensively measure the true mapping of spatial points in the real world in the image. Thus, the possible position offset of each pixel in the feature map can also be accurately determined, that is, the accurate target offset parameters can be determined.
[0077] Furthermore, the P feature maps are feature maps of different scales corresponding to different cameras. The obtained target offset parameter is used to represent the possible positional offset of each pixel in the feature map of a certain scale corresponding to a camera. In the above method, by determining the target offset parameter for each scale feature map corresponding to each camera, the possible positional offset of each pixel in the feature map of each scale corresponding to each camera can be obtained. This is equivalent to describing the possible positional offset of each pixel in the feature map from different levels, and then the position of the corresponding pixel can be accurately corrected accordingly.
[0078] In the embodiments of this application, the camera calibration parameters may include camera intrinsic parameters, camera extrinsic parameters, and distortion coefficients.
[0079] In this case, the operation of extracting features from the calibration parameters of the i-th camera out of M cameras can be as follows: concatenate the camera intrinsic parameters, camera extrinsic parameters, and distortion coefficients of the i-th camera to obtain the concatenated parameters of the i-th camera; then input the concatenated parameters into the feature extraction network to obtain the calibration parameter features of the i-th camera.
[0080] In the above steps, the camera intrinsic parameters, camera extrinsic parameters, and distortion coefficients can be converted into one-dimensional arrays, and then these one-dimensional arrays can be concatenated to obtain the concatenated one-dimensional parameters (concatenated parameters). Feature extraction from the concatenated parameters then yields the calibration parameter features.
[0081] In this embodiment, the feature extraction network can be a neural network composed of multiple fully connected layers. For example, the feature extraction network can be a neural network composed of three fully connected layers. It should be understood that the fully connected layers, through a combination of weighted summation and activation functions, enable them to have nonlinear fitting capabilities. The essence of camera calibration is to establish a mapping relationship between "image pixel information" and "camera intrinsic parameters," camera extrinsic parameters, and distortion coefficients. This process is affected by factors such as lens distortion (radial and tangential), imaging noise, and ambient lighting, resulting in nonlinear coupling. In this case, using a feature extraction network composed of multiple fully connected layers to extract features of the camera calibration parameters allows for a thorough fitting of the nonlinear relationships between the parameters, thereby extracting more accurate calibration parameter features.
[0082] The operation of determining the j-th target offset parameter among P target offset parameters based on the j-th fusion feature can be as follows: input the j-th fusion feature into the offset parameter determination model, and output the j-th target offset parameter through the offset parameter determination model.
[0083] In this embodiment of the application, the offset parameter determination model can be a neural network model composed of multiple convolutional layers. For example, the offset parameter determination model can be composed of three convolutional layers with shared weights. The three convolutional layers are used to learn the feature representation of the j-th fused feature. After that, it can be passed through a fully connected layer and an activation layer to classify and obtain the j-th target offset parameter.
[0084] It is worth noting that the offset parameter determination model can be trained before determining the j-th target offset parameter through the offset parameter determination model.
[0085] It should be understood that the offset parameter can be obtained through server training to determine the model.
[0086] The offset parameter determination model can be trained based on the difference between real point cloud features and sample image features. Real point cloud features are obtained by extracting features from point clouds acquired by LiDAR, while sample image features are obtained by extracting features from sample images acquired by a camera. Furthermore, real point cloud features and sample image features should be features of the same scale.
[0087] Specifically, a training dataset is obtained, which includes multiple first training samples. Each of these first training samples includes sample data and sample labels. The original neural network model is trained based on these first training samples to obtain the offset parameter determination model. The sample data can be the sample image features of a sample image, and the sample labels are the true offset parameters corresponding to these sample image features (i.e., the possible offsets of each pixel in these sample image features). The true offset parameters corresponding to these sample image features can be obtained based on the difference between the sample image features and the true point cloud features. In other words, the input data for each of the multiple first training samples is the sample image features of a sample image, and the sample labels are the true offset parameters corresponding to these sample image features.
[0088] Because the point cloud acquired by LiDAR accurately reflects the vehicle's surrounding environment, and LiDAR does not suffer from extrinsic parameter jitter, the pixels extracted from the point cloud by LiDAR will not experience positional shifts. In other words, the true point cloud features accurately reflect the mapping between spatial points in the real environment and pixels in the image, thus accurately reflecting the pixel's position in the image. In this case, based on the difference between the true point cloud features and the sample image features, an accurate prediction of the pixel's potential positional shift can be trained.
[0089] When training the original neural network model using multiple first training samples, for each of these first training samples, the input data from that first training sample can be input into the neural network model to obtain output data. A loss function is then used to determine the loss value between the output data and the sample labels in that first training sample. The parameters in the neural network model are then adjusted based on this loss value. After adjusting the parameters of the neural network model based on each of these first training samples, the adjusted neural network model is the offset parameter determination model.
[0090] The operation of adjusting the parameters in the neural network model based on the loss value can be referred to in relevant technologies, and will not be described in detail in this embodiment.
[0091] For example, the server can use formulas This allows for the adjustment of any parameter in the neural network model. These are the adjusted parameters. W These are the parameters before adjustment. It's the learning rate. It can be preset, such as It can be 0.001, 0.000001, etc., and the embodiments of this application do not limit it to this only. dw Is the loss function about W The derivative can be obtained from the loss value.
[0092] In some embodiments, the j-th target offset parameter may include N position offsets, N first weights, and one second weight corresponding to each pixel in the j-th feature map.
[0093] In this system, N positional offsets of a pixel represent the N possible positional offsets of that pixel in the j-th feature map. N first weights correspond one-to-one with these N positional offsets, and each first weight indicates the probability that the pixel's actual positional offset is the positional offset corresponding to that weight. Second weights represent the importance of the j-th feature map compared to the other feature maps.
[0094] In this case, the true offset parameter corresponding to the feature of this sample image in the sample label can also include N possible position offsets, a first weight and a second weight corresponding to the N position offsets. This allows the offset parameter determination model to predict N possible position offsets corresponding to a feature map, a first weight corresponding to the N position offsets, and a second weight.
[0095] In the above method, it is equivalent to determining the N possible position offsets of each pixel in the j-th feature map, the first weight corresponding to the N position offsets, and the second weight corresponding to the j-th feature map, so that all possible position offsets of each pixel can be determined, and the pixel position can be accurately corrected based on all possible position offsets.
[0096] It is worth noting that the ground detection method provided in this application is based on BEV features. Therefore, after obtaining the feature maps of the target image, it is necessary to convert P feature maps to BEV space to obtain BEV features. However, target images acquired by different cameras have feature maps at multiple scales, and a certain feature at different scales may correspond to a pixel in the target image. Therefore, directly projecting the pixel in the two-dimensional feature map to BEV space would be quite difficult. In this application, a back-projection method is designed to first project the coordinate points in BEV space onto each feature map, and finally combine the projection positions on each feature map to determine the feature value of this coordinate point, making the determination process of BEV features simpler.
[0097] First, we will explain the specific implementation method of projecting coordinate points in the BEV space onto P feature maps.
[0098] Specifically, for any one of the multiple coordinate points in the target bird's-eye view space, based on the calibration parameters of M cameras, this coordinate point is projected onto P feature maps respectively to obtain the P first projection points corresponding to this coordinate point.
[0099] The target bird's-eye view space is the BEV space constructed in this application embodiment using a non-uniformly spaced sampling method. In this application embodiment, the target bird's-eye view space includes K planes, and each of the K planes includes multiple coordinate points. The K planes are obtained by sampling in the space near the ground using a non-uniformly spaced method. One possible approach is to use the plane where the rear axle center of the vehicle is located as the zero plane, and sample A planes at first height intervals between 1 meter and -1 meter above and below this zero plane. Then, sample B planes at second height intervals within a 3-meter vertical extension from this 1-meter to -1-meter height. The sum of the A planes and the B planes is the K planes.
[0100] The first height is smaller than the second height. For example, Figure 3 This is a schematic diagram of the structure of a target bird's-eye view space provided in an embodiment of this application. See also... Figure 3The zero plane is the plane where the center of the rear axle of the vehicle is located. For example, if the first height is 0.3 meters, then 6 planes can be sampled every 0.3 meters in height between 1 meter above and -1 meter on the zero plane. If the second height is 1 meter, then 6 more planes can be sampled within a height range of 3 meters above and below. The final constructed target bird's-eye view space can include 12 planes.
[0101] Since ground markings are mainly distributed on the ground around the vehicle and in the near-ground area, while there are almost no ground markings in the high-altitude area, the above method constructs a target bird's-eye view space by densely sampling in the space area with a vertical height of 1 meter to -1 meter, and by sparsely sampling in the high-altitude area. This allows the target bird's-eye view space to fully capture the features of the near-ground space, so that the subsequent bird's-eye view features can fully characterize the features of ground markings, thereby improving the accuracy of ground marking detection.
[0102] In this embodiment, each of the K planes can include multiple coordinate points, thus constructing a three-dimensional space of size (K, X, Y, D) for the target bird's-eye view. Furthermore, the spatial range (X, Y) of each plane can be set according to the vehicle's proximity sensing range. For example, if the vehicle's proximity sensing range is 32 meters forward, 15 meters to the left and right, and 0.25 meters backward, and the resolution of each pixel is set to 0.25 meters, then each plane can be set to contain 120×188 pixels. The height Z of each plane is the height of the zero plane compared to the corresponding plane. Therefore, in this case, the target bird's-eye view space is a three-dimensional space containing multiple coordinate points of size (12, 120, 188, D), where D is the feature vector dimension of each spatial point in the target bird's-eye view space.
[0103] The first projection point is the pixel on a feature map when a coordinate point in the target bird's-eye view space is projected onto that feature map.
[0104] After constructing the target bird's-eye view space, each of the multiple coordinate points in the target bird's-eye view space can be projected onto P feature maps. The following explains the specific implementation method of projecting a coordinate point onto P feature maps to obtain the P first projection points corresponding to this coordinate point.
[0105] Specifically, for the j-th feature map among P feature maps, the camera extrinsic parameters are multiplied by this coordinate point to obtain the first coordinate point, which is the three-dimensional coordinate point in the camera coordinate system. The first coordinate point is normalized to obtain the second coordinate point, which is the two-dimensional coordinate point on the image plane. The distortion of the second coordinate point is corrected using distortion coefficients to obtain the third coordinate point. The third coordinate point is multiplied by the camera intrinsic parameters to obtain a pixel coordinate, which is the pixel coordinate of this coordinate point in the image coordinate system. The position of this pixel coordinate on the j-th feature map is determined as the first projection point of this coordinate point on the j-th feature map.
[0106] It should be understood that the aforementioned camera extrinsic parameters, camera intrinsic parameters, and distortion coefficients are the camera calibration parameters used when acquiring the target image corresponding to the j-th feature map. Furthermore, the camera extrinsic parameters are an extrinsic parameter matrix, and the camera intrinsic parameters are an intrinsic parameter matrix.
[0107] Since camera intrinsics are used to map from the camera coordinate system to the image coordinate system, and camera extrinsic parameters are used to map from the world coordinate system to the camera coordinate system, the entire process of mapping coordinate points in BEV space to the image coordinate system can be achieved by using the camera intrinsics, camera extrinsic parameters, and distortion coefficients of the corresponding camera.
[0108] It is worth noting that, in this embodiment, before determining the position of the pixel coordinate on the j-th feature map as the first projection point of the coordinate point on the j-th feature map, it is first determined whether the pixel coordinate falls within the perception range of the corresponding camera. Then, if the pixel coordinate falls within the perception range of the corresponding camera, the position of the pixel coordinate on the j-th feature map is determined as the first projection point of the coordinate point on the j-th feature map.
[0109] It should be understood that cameras have a fixed field of view and imaging range; points outside this range cannot be observed. In the above method, by determining whether the pixel coordinates fall within the camera's perception range, and only if the pixel coordinates are within the camera's perception range, the position of these pixel coordinates on the j-th feature map is determined as the first projection point of this coordinate point on the j-th feature map. This ensures that all obtained first projection points are points observable within the camera's field of view, thereby improving the accurate mapping from BEV space coordinates to image feature coordinates.
[0110] One possible implementation is to determine whether the pixel coordinates are within the fixed resolution range of the corresponding camera.
[0111] Generally, a camera has a fixed imaging resolution. If the pixel coordinates fall within this imaging resolution, then the pixel coordinates are considered to fall within the camera's sensing range.
[0112] Another possible implementation is to determine whether the Z value of the first coordinate point corresponding to this pixel coordinate is greater than 0 for this pixel coordinate.
[0113] Generally, a camera can only observe points in front of it (points in the positive Z-axis direction), while points behind the camera or coinciding with the camera plane (the origin or negative Z-axis point) cannot be observed. Therefore, we can determine whether the Z-value of the first coordinate point is greater than 0. If the Z-value is less than or equal to 0, it means that this pixel coordinate does not fall within the sensor range of the corresponding camera.
[0114] Furthermore, when it is determined that a pixel coordinate does not fall within the perception range of the corresponding camera, it can be determined that the pixel coordinate is unreasonable, which means that the coordinate point of the target bird's-eye view space cannot be projected onto the camera plane. In other words, the coordinate point has no effective projection point in the feature map corresponding to the camera, so the pixel coordinate can be deleted.
[0115] By projecting each coordinate point in the target bird's-eye view space onto each feature map as described above, multiple first projection points can be obtained. Subsequently, the bird's-eye view features of multiple coordinate points in the target bird's-eye view space can be determined based on these multiple first projection points.
[0116] Step 203: Based on P feature maps, P target offset parameters and multiple first projection points, determine the target bird's-eye view features. The multiple first projection points are the projection points of multiple coordinate points in the target bird's-eye view space onto the P feature maps respectively. The target bird's-eye view features are the features of the P feature maps transformed into the target bird's-eye view space.
[0117] In the above method, the possible positional offsets of each position on each feature map can be taken into account during the process of converting P feature maps to the target bird's-eye view space, so as to obtain more accurate bird's-eye view features, and then more accurate ground markings can be detected based on this.
[0118] One possible way is that the operation in step 203 can be achieved through the following steps (1)-(4).
[0119] (1) For the j-th projection point among the P first projection points corresponding to a coordinate point, the position of the j-th first projection point is corrected based on the j-th target offset parameter to obtain N second projection points.
[0120] Where N is an integer greater than or equal to 1.
[0121] It should be understood that in the process of projecting coordinate points in the target bird's-eye view space onto image features, each coordinate point corresponds to a first projection point in each feature map. Therefore, P feature maps will yield P first projection points. When correcting the positions of the P first projection points, since a pixel position may have multiple possible positional offsets, after correcting the position of a first projection point, N second projection points will be obtained. That is, one first projection point can correspond to N second projection points.
[0122] One possible approach is to perform step (1) as follows: for any one of the N position offsets of the target pixel, add the position offset to the j-th first projection point to obtain the second projection point after correcting the position offset of the j-th first projection point.
[0123] It should be understood that the target offset parameter corresponding to the j-th feature map (the j-th target offset parameter) may include N position offsets, N first weights, and one second weight for each pixel on the j-th feature map. The N position offsets correspond one-to-one with the N first weights, and therefore the N first weights also correspond one-to-one with the N second projection points after each of the N position offsets. The second weight is used to represent the weight of the j-th feature map compared to the other feature maps in the ground marker detection task.
[0124] The target pixel is the pixel on the j-th feature map that corresponds to the j-th first projection point. It should be understood that the first projection point is the target pixel.
[0125] In this case, the N position offsets of the corresponding pixel are first obtained from the j-th target offset parameters. Then, the N position offsets are added to the j-th first projection point, which is equivalent to offsetting the j-th first projection point by N positions to obtain N new positions (N second projection points). The N second projection points are the more realistic pixels of this coordinate point in the target bird's-eye view space projected onto the j-th feature map. Thus, more realistic N second projection points can be obtained through the above method, thereby improving the accuracy of subsequent bird's-eye view features.
[0126] (2) Obtain the first pixel features on the N second projection points from the j-th feature map.
[0127] Since the coordinates of the target bird's-eye view space are projected onto the feature map in order to determine the bird's-eye view features of the corresponding position on the feature map, after determining the N second projection points corresponding to a first projection point, the first pixel features of the N projection points can be obtained from the j-th feature map.
[0128] (3) Based on the first pixel features of the first pixel on the N second projection points corresponding to each of the P target offset parameters and the P first projection points, determine the bird's-eye view features of this coordinate point.
[0129] One possible approach is that step (3) can be as follows: For any one of the N second projection points corresponding to the j-th projection point, multiply the first pixel feature of this second projection point by the first weight corresponding to this second projection point to obtain the second pixel feature corresponding to this second projection point; add the second pixel features corresponding to each of the N second projection points to obtain the third pixel feature; multiply the third pixel feature by the second weight to obtain the fourth pixel feature of the j-th feature map; and determine the bird's-eye view feature of this coordinate point based on the fourth pixel features of the P feature maps.
[0130] The third pixel feature is the pixel feature obtained by fusing the N second projection points of this coordinate point on the j-th feature map. The fourth pixel feature is the pixel feature obtained after considering the importance of the j-th feature map compared to the other feature maps in the ground label detection task.
[0131] In the specific operation described above, for the N second projection points corresponding to the first projection point of this coordinate point on the j-th feature map, the first pixel features of the N second projection points are first weighted and summed. That is, for the first pixel feature of any second projection point, this first pixel feature is multiplied by the first weight corresponding to the second projection point to obtain a second pixel feature. Then, the second pixel features corresponding to the N second projection points are added together to obtain the pixel feature after fusing the N second projection points. Since this coordinate point may not only be projected onto one feature map but also onto other feature maps, after obtaining the third pixel feature, it can be multiplied by the second weight corresponding to the j-th feature map to obtain the fourth pixel feature corresponding to the j-th feature map. Finally, the fourth pixel features of each of the P feature maps are fused to determine the bird's-eye view feature of this coordinate point.
[0132] In the above method, by considering the possibility that a second projection point is the true offset of a pixel point during the process of fusing pixel features of different projection points on different feature maps, and the importance of a feature map to the ground marker detection task compared to other feature maps, a more accurate bird's-eye view feature of this coordinate point can be obtained by fusing.
[0133] The operation of determining the bird's-eye view feature of this coordinate point based on the fourth pixel feature of P feature maps can be as follows: add the fourth pixel feature of each of the P feature maps to obtain the bird's-eye view feature of this coordinate point.
[0134] It is worth noting that by applying the above possible implementation methods to each second projection point of multiple coordinate points in the target bird's-eye view space on each feature map, bird's-eye view features of multiple coordinate points can be obtained.
[0135] (4) The bird's-eye view features of multiple coordinate points in the target bird's-eye view space are fused to obtain the target bird's-eye view features.
[0136] It should be understood that the target bird's-eye view space includes K planes, and the multiple coordinate points are also the coordinate points on the K planes. Therefore, after the above steps (1)-(3), the bird's-eye view features of multiple coordinate points on the K planes can be obtained. The ultimate goal is to obtain a BEV dense feature of size (1, X, Y, D) (that is, the target bird's-eye view feature). For example, it is to obtain a BEV dense feature of (1, 120, 188, D), which is to obtain a BEV dense feature on a plane containing a pixel resolution of 120×188 and a feature vector dimension of D for each pixel.
[0137] Therefore, after obtaining the bird's-eye view features from multiple coordinate points, feature fusion is required to obtain the target bird's-eye view features that meet the expectations.
[0138] One possible approach is to perform the following steps in step (4): For any one of the K planes, perform attention encoding on the bird's-eye view features of each coordinate point on the plane to obtain attention features; fuse the attention features with the bird's-eye view features of the coordinate points on the plane to obtain the reference bird's-eye view features of the plane; fuse the reference bird's-eye view features of the K planes to obtain the target bird's-eye view features.
[0139] The attention feature includes weights corresponding to each coordinate point. The weight of a coordinate point is used to represent the importance of that coordinate point in space for ground marker detection.
[0140] It should be understood that attention encoding can obtain the importance of each coordinate point in the corresponding plane. Specifically, in the embodiments of this application, it can obtain the importance of the features of each coordinate point in the corresponding plane for ground marking detection.
[0141] In the above method, attention encoding is first applied to the bird's-eye view features of each coordinate point to determine the importance of each coordinate point on a plane. Then, the bird's-eye view features of the coordinate points on that plane are further encoded based on the importance of each coordinate point, making the important features of the plane more prominent. Finally, the reference bird's-eye view features corresponding to each plane are fused, resulting in a fusion of the important features of each of the K planes, thus obtaining a bird's-eye view feature that more accurately represents the ground markings in the BEV space.
[0142] The operation of encoding the bird's-eye view features of each coordinate point on this plane with attention can be as follows: input the bird's-eye view features of each coordinate point on this plane into the spatial attention network, and output the attention features through the spatial attention network.
[0143] It should be understood that the BEV space consists of multiple spatial locations corresponding to the real-world physical space. These multiple spatial locations often cover an area far around the vehicle's location, but most of these locations are irrelevant to the current ground marking detection task. The relevant features are mainly those in the near-ground area.
[0144] Spatial attention networks are network structures that incorporate spatial attention mechanisms. The purpose of spatial attention mechanisms is to proactively focus on spatial locations that are more important to the current task (ground marker detection in this embodiment). Therefore, in the above approach, by using spatial attention networks to encode the bird's-eye view features of each coordinate point on the plane, attention features that accurately measure the importance of each spatial location can be obtained.
[0145] The operation of fusing the attention feature with the bird's-eye view features of each coordinate point on the plane to obtain the reference bird's-eye view features corresponding to the plane can be as follows: for any coordinate point on the plane, multiply the bird's-eye view feature of the coordinate point by the weight corresponding to the coordinate point to obtain the reference bird's-eye view feature of the coordinate point on the plane.
[0146] By performing the above weighting operation on each coordinate point on this plane, the reference bird's-eye view features corresponding to this plane can be obtained.
[0147] The operation of fusing the reference bird's-eye view features of K planes to obtain the target bird's-eye view features can be as follows: add the reference bird's-eye view features of corresponding coordinate points in the K planes to obtain the fused bird's-eye view features; and extract features from the fused bird's-eye view features through multiple convolutional layers to obtain the target bird's-eye view features.
[0148] It should be understood that the principle of ground marking detection based on bird's-eye view features in this application embodiment is to more accurately characterize ground marking features based on the structured relationship of physical space reflected by bird's-eye view features. Since the sum of the reference bird's-eye view features of K corresponding locations on the plane is essentially to achieve element-level superposition, that is, a simple aggregation of information, there may be spatial correlation or spatial misalignment after superposition, thus making it impossible to capture the correlation between spatial locations.
[0149] Convolutional layers have local receptive fields, which can automatically learn feature dependencies within their spatial neighborhood.
[0150] In the above method, after adding the features of the reference bird's-eye view at corresponding positions in K planes, multiple convolutional layers are used to extract features from the fused features. This allows multiple convolutional layers to aggregate the features of scattered points into structured regional features, thereby enhancing the structured expression of features and capturing the correlation between spatial locations.
[0151] In the above steps (1)-(4), during the process of converting P feature maps to the target bird's-eye view space, the position of the projection point on each feature map is corrected, and then the pixel features on the new projection point obtained after the position correction are combined to determine the bird's-eye view feature of this coordinate point. Finally, the bird's-eye view features of multiple coordinate points are fused to determine a more accurate bird's-eye view feature, and then a more accurate ground marking can be detected based on this.
[0152] Through the above steps 201-203, the feature maps corresponding to the M target images acquired by the M cameras can be mapped into the BEV space to obtain the target bird's-eye view features corresponding to the M target images. Then, ground marker detection can be performed based on the target bird's-eye view features, that is, continue to execute the following step 204.
[0153] Step 204: Based on the features of the target bird's-eye view, perform ground marking detection to obtain the ground marking information of the current driving road.
[0154] Ground markings are two-dimensional planar targets. The overhead view in BEV space has no perspective distortion, which can accurately restore the geometry of ground markings. In the above method, ground marking detection based on the target bird's-eye view features can obtain more accurate ground marking information.
[0155] One possible approach is to input the target bird's-eye view features into a ground sign detection model, process the target bird's-eye view features through the ground sign detection model, and output the ground sign category and ground sign location.
[0156] It should be understood that straight lane markings, left turn lane markings, stop lines, lane lines, etc., are all different categories of ground markings. In the above implementation method, the specific ground marking category can be output through the ground marking detection model.
[0157] Ground marker location refers to the position coordinates of a ground marker in an image. In some embodiments, the ground marker location can be the position coordinates of multiple corner points (key points) corresponding to the ground marker in the image.
[0158] Specifically, the operation of processing the target bird's-eye view features through the ground sign detection model to output the ground sign category and ground sign location can be as follows: encoding the target bird's-eye view features through the ground sign detection model to obtain reference encoded features; generating a heat map corresponding to the target bird's-eye view features based on the reference encoded features; and outputting the ground sign category and corner location based on the heat map.
[0159] In the heatmap, the response value at each location represents the probability that the location belongs to a ground marker corner. It should be understood that a higher response value in the heatmap indicates a higher probability that the location belongs to a ground marker corner, and a lower response value indicates a lower probability. In this embodiment, the probability of each location belonging to a ground marker corner can be obtained by activating the reference encoded features using a sigmoid function.
[0160] The operation of outputting ground marker categories and corner locations based on heatmaps can be as follows: identify the multiple locations with the highest response values in the heatmap as corner locations, classify the features within the corner locations using the softmax function to obtain multiple predicted categories and the confidence levels of multiple predicted categories, and identify the predicted category with the highest confidence level among the multiple predicted categories as the ground marker category.
[0161] In this embodiment of the application, the number of corner points can be preset. For example, four corner points can be set to output, which can be the four key points marked on the ground: upper left, upper right, lower left, and lower right.
[0162] Another possible approach is to perform max pooling on the heatmap to obtain the corner locations.
[0163] In the above method, by determining the multiple locations with the highest response values in the heatmap as corner locations, or by reducing the number of corner markers in the heatmap through max pooling, the corner locations corresponding to the ground markers can be obtained. This eliminates the need to perform maximum suppression on the detection boxes corresponding to the multiple corner locations in the heatmap, thus simplifying the operation process.
[0164] It is worth noting that before processing the features of the target bird's-eye view through the ground sign detection model and outputting the ground sign category and ground sign location, the ground sign detection model can be trained first.
[0165] Specifically, a target training set can be obtained, which may include multiple second training samples; the original encoder-decoder network model is trained based on these multiple second training samples to obtain a ground marker detection model.
[0166] Each of the multiple second training samples includes sample data and sample labels. The sample data can be BEV sample features corresponding to a sample image containing ground markers, and the sample labels can be the category and location of the ground markers contained in the sample data. That is, the input data for each training sample in the multiple second training samples is the BEV sample features corresponding to a sample image containing ground markers, and the sample labels are the category and location of the ground markers contained in the sample data.
[0167] The specific training process of training the original encoding and decoding network model based on the multiple second training samples to obtain the ground marker detection model is similar to the operation of training the original neural network model based on multiple first training samples to obtain the offset parameter determination model in step 202 above, and will not be repeated here.
[0168] In some embodiments, the sample data can be further segmented to obtain a boundary segmentation detection mask; then the boundary segmentation mask can be added to the sample label, that is, the sample label can contain the category and location of the boundary segmentation mask and the ground marker.
[0169] The boundary segmentation mask is a binary image with the same feature size as the BEV sample. Pixels belonging to the ground boundary line are labeled as 1, and pixels not belonging to the boundary line are labeled as 0.
[0170] During the training of the original codec network model, in the early stages of training, the model can output a binary boundary segmentation image, and the boundary segmentation loss can be calculated based on the difference between this binary segmentation image and the boundary segmentation mask in the sample label. Then, the parameters of the original codec network model can be adjusted and the BEV sample features can be optimized accordingly based on the boundary segmentation loss.
[0171] In the above approach, by adding an auxiliary task of boundary segmentation, richer supervision signals and morphological prior knowledge can be provided for model training, allowing the model to learn to recognize the outline of ground markers while learning to find ground marker corners. Ultimately, by leveraging the spatial association between boundaries and corners (corner points are the endpoints of boundary lines, and boundary lines connect two corner points), the corner detection task in model training can converge faster, thus enabling faster and more accurate localization of ground markers.
[0172] It is worth noting that the ground marking detection method provided in this application increases the field of view by adding two pinhole cameras, one in front and one behind. Furthermore, this application provides a complete process for mapping two-dimensional feature maps to the BEV space. This process can fully utilize the possible positional offsets of pixels to correct the features of a single location point, thereby reducing the impact of parameter fluctuations from cameras and other sensors on the detection performance. Finally, the use of boundary segmentation-based auxiliary loss can improve the convergence speed of model training and the stability of detection.
[0173] In this embodiment, the domain controller first extracts features from M target images to obtain P feature maps. Then, based on the P feature maps and the calibration parameters of the M cameras on the vehicle, it determines P target offset parameters, that is, it first determines the possible positional offsets at each position on each feature map. Next, based on the P feature maps, the P target offset parameters, and multiple first projection points, it determines the target bird's-eye view features. This involves converting the P feature maps obtained from feature extraction into the BEV space, taking into account the possible positional offsets at each position on each feature map during the conversion process to obtain more accurate target bird's-eye view features. Subsequently, based on the target bird's-eye view features, ground marking detection is performed to obtain the ground marking information of the current driving road. This allows for more accurate ground marking detection based on the target bird's-eye view features. Compared to the problem of inaccurate detection caused by image misalignment due to external parameter jitter in the existing technology, this solution takes into account the impact of external parameter jitter on the positional offset of the feature map during the extraction of BEV spatial features. This avoids the problem of inaccurate image features caused by external parameter jitter, which leads to inaccurate ground marking detection, thereby improving the accuracy of ground marking detection.
[0174] Figure 4 This is a schematic diagram of a ground marking detection device provided in an embodiment of this application. The ground marking detection device can be implemented as part or all of a vehicle by software, hardware, or a combination of both. The vehicle can be described below. Figure 5 The vehicle shown. See also Figure 4 The device includes: a feature extraction module 401, a parameter determination module 402, a feature conversion module 403, and a detection module 404.
[0175] The feature extraction module 401 is used to extract features from M target images respectively to obtain P feature maps. The M target images are images captured by M cameras on the vehicle. The parameter determination module 402 is used to determine P target offset parameters based on P feature maps and M camera calibration parameters. The P target offset parameters are used to represent the possible positional offset of each pixel on each feature map. The feature transformation module 403 is used to determine the target bird's-eye view features based on P feature maps, P target offset parameters and multiple first projection points. The multiple first projection points are the projection points of multiple coordinate points in the target bird's-eye view space on the P feature maps respectively. The target bird's-eye view features are the features transformed from the P feature maps to the target bird's-eye view space. The detection module 404 is used to detect ground markings based on the features of the target bird's-eye view to obtain the ground marking information of the current driving road.
[0176] Optionally, the parameter determination module 402 is specifically used for: For the i-th camera among M cameras, feature extraction is performed on the calibration parameters of the i-th camera to obtain the calibration parameter features of the i-th camera; For the j-th feature map among P feature maps, the j-th feature map is fused with the calibration parameter features of the target camera to obtain the j-th fused feature. The target image corresponding to the j-th feature map is acquired by the target camera among M cameras. Based on the j-th fusion feature, determine the j-th target offset parameter among the P target offset parameters.
[0177] Optionally, the parameter determination module 402 is specifically used for: The j-th fusion feature is input into the offset parameter determination model, and the offset parameter determination model outputs the j-th target offset parameter. The offset parameter determination model is trained based on the difference between real point cloud features and sample image features. Real point cloud features are obtained by feature extraction from point clouds acquired by LiDAR, and sample image features are obtained by feature extraction from sample images acquired by camera.
[0178] Optionally, the device further includes: The projection module is used to project any one of the multiple coordinate points onto P feature maps based on the calibration parameters of M cameras, thereby obtaining P first projection points corresponding to the coordinate point.
[0179] Optionally, the feature determination module 403 is specifically used for: For the j-th first projection point among P first projection points, the position of the j-th first projection point is corrected based on the j-th target offset parameter to obtain N second projection points; Obtain the first pixel features on the N second projection points from the j-th feature map; Based on P target offset parameters and the first pixel features on N second projection points corresponding to each of the P first projection points, the bird's-eye view features of the coordinate points are determined. The bird's-eye view features of multiple coordinate points are fused to obtain the target bird's-eye view features.
[0180] Optionally, the j-th target offset parameter includes N position offsets of each pixel on the j-th feature map, and the feature determination module 403 is specifically used for: For any one of the N position offsets of the target pixel, add the position offset to the j-th first projection point to obtain the second projection point after correcting the position offset of the j-th first projection point. The target pixel is the pixel corresponding to the j-th first projection point.
[0181] Optionally, the j-th target offset parameter includes N first weights and one second weight for each pixel on the j-th feature map, where the N first weights correspond one-to-one with the N second projection points, and the second weights correspond to the j-th feature map; the feature determination module 403 is specifically used for: For any one of the N second projection points corresponding to the j-th projection point, multiply the first pixel feature on the second projection point by the first weight corresponding to the second projection point to obtain the second pixel feature corresponding to the second projection point. The second pixel features corresponding to each of the N second projection points are summed to obtain the third pixel features. Multiply the third pixel feature by the second weight to obtain the fourth pixel feature of the j-th feature map; Based on the fourth pixel feature of P feature maps, the bird's-eye view features of the coordinate points are determined.
[0182] Optionally, the target bird's-eye view space consists of K planes, each containing multiple coordinate points. The feature determination module 403 is specifically used for: For any one of the K planes, attention encoding is performed on the bird's-eye view features of the coordinate points on the plane to obtain attention features; The attention features are fused with the bird's-eye view features of coordinate points on the plane to obtain the reference bird's-eye view features of the plane. The target bird's-eye view features are obtained by fusing the features of the reference bird's-eye view from K planes.
[0183] Optionally, the feature extraction module 401 is specifically used for: For the i-th target image among M target images, the i-th target image is input into a multi-scale feature extraction network, which outputs Q feature maps. The Q feature maps have different sizes, and Q multiplied by M equals P.
[0184] In this embodiment, features are first extracted from M target images to obtain P feature maps. Then, based on the P feature maps and the calibration parameters of the M cameras on the vehicle, P target offset parameters are determined, which means determining the possible positional offsets at each position on each feature map. Next, based on the P feature maps, the P target offset parameters, and multiple first projection points, the target bird's-eye view features are determined. This involves converting the P feature maps obtained from feature extraction into the BEV space, taking into account the possible positional offsets at each position on each feature map during the conversion process to obtain more accurate target bird's-eye view features. Subsequently, based on the target bird's-eye view features, ground marking detection is performed to obtain the ground marking information of the current driving road. This allows for more accurate ground marking detection based on the target bird's-eye view features. Compared to the problem of inaccurate detection caused by image misalignment due to external parameter jitter in the existing technology, this solution takes into account the impact of external parameter jitter on the positional offset of the feature map during the extraction of BEV spatial features. This avoids the problem of inaccurate image features caused by external parameter jitter, which leads to inaccurate ground marking detection, thereby improving the accuracy of ground marking detection.
[0185] It should be noted that the ground marking detection device provided in the above embodiments is only illustrated by the division of the above functional modules when detecting ground markings. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0186] The functional units and modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of this application.
[0187] The ground marking detection device and the ground marking detection method provided in the above embodiments belong to the same concept. The specific working process and technical effects of the units and modules in the above embodiments can be found in the method embodiments section, and will not be repeated here.
[0188] Figure 5 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application.
[0189] For example, such as Figure 5As shown, the vehicle 500 includes a memory 51 and a processor 50, wherein the memory 51 stores executable program code 52, and the processor 50 is used to call and execute the executable program code 52 to perform the above-mentioned ground marking detection method.
[0190] This embodiment can divide the vehicle into functional modules according to the above method example. For example, each function can be assigned to a separate module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0191] When each function is divided into modules corresponding to its specific function, the vehicle may include: a feature extraction module, a parameter determination module, a feature transformation module, and a detection module. It should be noted that all relevant content regarding the steps involved in the above method embodiments can be referenced from the functional descriptions of the corresponding modules, and will not be repeated here.
[0192] The vehicle provided in this embodiment is used to perform the above-described ground marking detection method, and therefore can achieve the same effect as the above-described implementation method.
[0193] When using integrated units, the vehicle may include a processing module and a storage module. The processing module is used to control and manage the vehicle's actions. The storage module is used to support the vehicle in executing corresponding program code and data.
[0194] The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits shown in conjunction with the disclosure of this application. The processor may also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module may be a memory.
[0195] This embodiment also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the above-described related method steps to implement the above-described ground marking detection method.
[0196] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement the above-described ground marking detection method.
[0197] In this embodiment, the vehicle, computer-readable storage medium, computer program product, or chip are all used to execute the method described above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the method described above, and will not be repeated here.
[0198] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0199] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are illustrative; for instance, the division of modules or units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0200] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for detecting ground markings, characterized in that, The method includes: Feature extraction is performed on M target images to obtain P feature maps, where the M target images are images captured by M cameras on the vehicle. Based on the P feature maps and the calibration parameters of the M cameras, P target offset parameters are determined. The P target offset parameters are used to represent the possible positional offset of each pixel on each feature map. Based on the P feature maps, the P target offset parameters, and multiple first projection points, target bird's-eye view features are determined. The multiple first projection points are projection points of multiple coordinate points in the target bird's-eye view space onto the P feature maps respectively. The target bird's-eye view features are the features of the P feature maps transformed into the target bird's-eye view space. Based on the features of the target bird's-eye view, ground markings are detected to obtain the ground marking information of the current driving road.
2. The method as described in claim 1, characterized in that, The determination of P target offset parameters based on the P feature maps and the calibration parameters of the M cameras includes: For the i-th camera among the M cameras, feature extraction is performed on the calibration parameters of the i-th camera to obtain the calibration parameter features of the i-th camera; For the j-th feature map among the P feature maps, the j-th feature map is fused with the calibration parameter features of the target camera to obtain the j-th fused feature. The target image corresponding to the j-th feature map is acquired by the target camera among the M cameras. Based on the j-th fusion feature, the j-th target offset parameter among the P target offset parameters is determined.
3. The method as described in claim 2, characterized in that, The step of determining the j-th target offset parameter among the P target offset parameters based on the j-th fusion feature includes: The j-th fusion feature is input into the offset parameter determination model, and the offset parameter determination model outputs the j-th target offset parameter. The offset parameter determination model is trained based on the difference between real point cloud features and sample image features. The real point cloud features are obtained by feature extraction from point clouds acquired by LiDAR, and the sample image features are obtained by feature extraction from sample images acquired by camera.
4. The method as described in claim 1, characterized in that, The method further includes: For any one of the plurality of coordinate points, based on the calibration parameters of the M cameras, the coordinate point is projected onto the P feature maps respectively to obtain the P first projection points corresponding to the coordinate point.
5. The method as described in claim 4, characterized in that, The step of determining the target bird's-eye view features based on the P feature maps, P target offset parameters, and multiple first projection points includes: For the j-th first projection point among the P first projection points, the position of the j-th first projection point is corrected based on the j-th target offset parameter to obtain N second projection points; Obtain the first pixel feature on the N second projection points from the j-th feature map; Based on the P target offset parameters and the first pixel features on the N second projection points corresponding to each of the P first projection points, the bird's-eye view features of the coordinate points are determined. The bird's-eye view features of the multiple coordinate points are fused to obtain the target bird's-eye view features.
6. The method as described in claim 5, characterized in that, The j-th target offset parameter includes N position offsets of each pixel on the j-th feature map. The step of correcting the position of the j-th first projection point based on the j-th target offset parameter to obtain N second projection points includes: For any one of the N position offsets of the target pixel, add the position offset to the j-th first projection point to obtain a second projection point after correcting the position offset of the j-th first projection point, and the target pixel is the pixel corresponding to the j-th first projection point.
7. The method as described in claim 5, characterized in that, The j-th target offset parameter includes N first weights and one second weight for each pixel on the j-th feature map, wherein the N first weights correspond one-to-one with the N second projection points, and the second weight corresponds to the j-th feature map; determining the bird's-eye view features of the coordinate point based on the P target offset parameters and the first pixel features on the N second projection points corresponding to each of the P first projection points includes: For any one of the N second projection points corresponding to the j-th projection point, the first pixel feature on the second projection point is multiplied by the first weight corresponding to the second projection point to obtain the second pixel feature corresponding to the second projection point. The second pixel features corresponding to each of the N second projection points are summed to obtain the third pixel features. Multiply the third pixel feature by the second weight to obtain the fourth pixel feature of the j-th feature map; Based on the fourth pixel feature of the P feature maps, the bird's-eye view feature of the coordinate point is determined.
8. The method as described in claim 5, characterized in that, The target bird's-eye view space is composed of K planes, and the K planes include the plurality of coordinate points. The process of fusing the bird's-eye view features of the plurality of coordinate points to obtain the target bird's-eye view features includes: For any one of the K planes, attention encoding is performed on the bird's-eye view features of the coordinate points on the plane to obtain attention features; The attention features are fused with the bird's-eye view features of the coordinate points on the plane to obtain the reference bird's-eye view features of the plane. The target bird's-eye view features are obtained by fusing the reference bird's-eye view features of the K planes.
9. The method as described in claim 1, characterized in that, The step of extracting features from M target images to obtain P feature maps includes: For the i-th target image among the M target images, the i-th target image is input into a multi-scale feature extraction network, and the multi-scale feature extraction network outputs Q feature maps, the Q feature maps having different sizes, and Q multiplied by M equals P.
10. A vehicle, characterized in that, The vehicles include: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the vehicle to perform the method as described in any one of claims 1 to 9.