Semantic grid map optimization method and system based on multi-task road surface information
By optimizing the OCC model map using computer vision models and sensor projection relationships, the problems of insufficient lane line accuracy and high obstacle recognition costs were solved, achieving efficient and lightweight semantic map optimization and improving the accuracy and reliability of autonomous driving environmental perception.
Patent Information
- Application Number
- CN202511562860.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-10-30
AI Technical Summary
Existing OCC models suffer from insufficient lane line information accuracy and high cost of obstacle semantic supplementation, resulting in inaccurate environmental perception for autonomous driving, especially in the recognition of narrow lane lines and obstacles.
By processing road environment images using computer vision models, combining the rotation and translation matrices between sensors, establishing coordinate projection relationships, supplementing lane line information and correcting the semantics of drivable areas, filtering obstacles in non-driving areas, integrating obstacle semantic associations, and generating an optimized 3D semantic raster map.
Without increasing the system burden, the quality and usability of semantic maps are significantly improved, providing a safer and more reliable environmental cognition foundation for autonomous driving systems.
Smart Images

Figure CN121033084A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of automatic driving navigation, and particularly relates to a semantic grid map optimization method and system based on multi-task road surface information. BACKGROUND
[0002] Precise environmental perception is the solid foundation of automatic driving. The most mature multi-modal fusion algorithm at present is the bird's eye view series algorithm, i.e. BEV. The OCC (Occupancy Prediction) algorithm derived from the BEV algorithm uses semantic grid to construct a dense three-dimensional occupancy map. In the task of constructing a three-dimensional semantic grid map, the OCC algorithm fully expresses the obstacles with geometric information in the automatic driving environment. With the aid of an image multi-task road surface information extraction model, the perception of lane lines and drivable areas can be well supplemented. In addition, when the category semantics of obstacles in the OCC task needs to be supplemented, the image multi-task model is selected to supplement this semantics. Then, the detection result of the image multi-task model on this semantic is associated with the corresponding position of the semantic grid result, which can avoid the high cost of repeated training of the OCC model and achieve more efficient representation.
[0003] However, due to the limitation of the computing power of the vehicle-mounted platform, the grid resolution of the semantic grid map is usually controlled at 0.5m x 0.5m x 0.5m. This accuracy cannot represent the information of narrow lane lines. The existing open-source OCC dataset defines the drivable area as flat road surface (between road edges) when manually labeled. The no-entry area, such as the emergency lane on the highway, is not distinguished. In addition, due to the complexity of the OCC task, the training cost of the OCC model is much higher than that of the conventional two-dimensional image detection and three-dimensional point cloud detection models. In the actual application process, it is difficult to avoid the situation that new obstacle semantics need to be supplemented. At this time, the cost of repeated OCC model retraining is too high. SUMMARY
[0004] To solve the above technical problems, the present application proposes a semantic grid map optimization method and system based on multi-task road surface information. The information supplement and result optimization of the three-dimensional semantic grid map are realized.
[0005] To achieve the above purpose, the present application adopts the following technical solutions: A semantic grid map optimization method based on multi-task road surface information, comprising the following steps: processing the obtained road environment image by using a computer vision model to obtain lane line segmentation results and drivable area segmentation results; establishing a coordinate projection relationship from the OCC space to the pixel coordinate system based on the rotation and translation matrix between sensors; The surface layer grid center points of the semantic grid map are projected to a pixel coordinate system based on a coordinate projection relationship, and lane line information is supplemented according to a projection result; The drivable area is semantically corrected based on the coordinate projection relationship and the drivable area segmentation result; Unknown obstacle instances in the initial map are extracted, and obstacle instances in a non-driving area are filtered out by using the corrected drivable area semantic information; the remaining unknown obstacle instances are projected to an image space by using the coordinate projection relationship again, and are semantically associated with a detection result of a computer vision model; The supplemented lane line information, the corrected drivable area semantics and the associated obstacle semantics are integrated to generate a final optimized three-dimensional semantic grid map and output.
[0006] The embodiment of the application further provides a semantic grid map optimization system based on multi-task road surface information, which comprises: A preprocessing module is configured to process a road environment image obtained by using a computer vision model to obtain lane line segmentation results and drivable area segmentation results; A mapping module is configured to establish a coordinate projection relationship from an OCC space to a pixel coordinate system based on a rotation and translation matrix between sensors; A supplement module is configured to project surface layer grid center points of a semantic grid map to a pixel coordinate system based on a coordinate projection relationship, and supplement lane line information according to a projection result; A correction module is configured to correct a drivable area based on a coordinate projection relationship and drivable area segmentation results; A semantic association module is configured to extract unknown obstacle instances in an initial map, filter out obstacle instances in a non-driving area by using corrected drivable area semantic information, project the remaining unknown obstacle instances to an image space by using a coordinate projection relationship again, and perform semantic association with a detection result of a computer vision model; An output map module is configured to integrate supplemented lane line information, corrected drivable area semantics and associated obstacle semantics, generate a final optimized three-dimensional semantic grid map and output.
[0007] The effects provided in the summary are only the effects of the embodiments, not all the effects of the application, and one of the above technical solutions has the following advantages or beneficial effects: The application provides a semantic grid map optimization method and system based on multi-task road surface information, and the method comprises the following steps: obtaining lane line segmentation results and drivable area segmentation results by processing acquired road environment images by using a computer vision model; establishing a coordinate projection relationship from an OCC space to a pixel coordinate system based on a rotation and translation matrix between sensors; projecting surface layer grid center points of a semantic grid map to the pixel coordinate system based on the coordinate projection relationship, and supplementing lane line information according to the projection results; performing semantic correction on a drivable area based on the coordinate projection relationship and the drivable area segmentation results; performing instance extraction on unknown obstacles in an initial map, and filtering out obstacle instances in a non-driving area by using the corrected drivable area semantic information; projecting the remaining unknown obstacle instances to an image space by using the coordinate projection relationship again, and performing semantic association with detection results of the computer vision model; integrating the supplemented lane line information, the corrected drivable area semantics and the associated obstacle semantics, generating a final optimized three-dimensional semantic grid map and outputting the three-dimensional semantic grid map. Based on the method, the application further provides a semantic grid map optimization system based on multi-task road surface information. The application closely integrates visual perception results and grid map structures, and constructs an efficient, automated and lightweight semantic map optimization pipeline, and the core effect is that the quality and practicability of the semantic map are significantly improved without significantly increasing the system burden, so that a safer and more reliable environment recognition basis is provided for an automatic driving system. BRIEF DESCRIPTION OF DRAWINGS
[0008] Figure 1 A semantic grid map optimization method based on multi-task road surface information is provided for the embodiment 1 of the application, and a flowchart of the method is shown in the figure; Figure 2 A schematic diagram of lane line fitting is provided for the embodiment 1 of the application; Figure 3 An image association semantic supplement schematic diagram is provided for the embodiment 1 of the application; Figure 4 A semantic grid map optimization system based on multi-task road surface information is provided for the embodiment 2 of the application, and a schematic diagram of the system is shown in the figure. DETAILED DESCRIPTION
[0009] To clearly illustrate the technical features of the present solution, the present application will be described in detail below with specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing the different structures of the present application. In order to simplify the disclosure of the present application, the components and settings of specific examples are described below. In addition, the present application can repeatedly refer to numbers and / or letters in different examples. Such repetition is for the purpose of simplification and clarity, and does not in itself indicate the relationship between the various embodiments and / or settings being discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. The present application omits the description of well-known components and processing techniques and processes to avoid unnecessarily limiting the present application.
[0010] Embodiment 1 Embodiment 1 of the present application proposes a semantic grid map optimization method based on multi-task road surface information. This method is an occupancy grid map (OCC) post-processing method, which is used to optimize the initial three-dimensional semantic grid map generated by the OCC model. The grid resolution of the initial map is 0.5m×0.5m×0.5m. The method solves the problems of missing lane line information, fuzzy drivable area definition, and high cost of unknown obstacle semantic supplement in the initial map by integrating lane line information supplement, drivable area correction, and unknown obstacle semantic association functions.
[0011] Figure 1 A flowchart of a semantic grid map optimization method based on multi-task road surface information is proposed for embodiment 1 of the present application; In step S1, the lane line segmentation result and the drivable area segmentation result are obtained by processing the acquired road environment image using a computer vision model. The specific process is as follows: A pre-trained computer vision model is deployed on a vehicle-mounted computing unit. The original image of the road environment is collected by a vehicle-mounted surround-view camera system, and the original image is input into the computer vision model. The computer vision model performs parallel processing on the input image and executes the following inference tasks: Output the probability map of each pixel belonging to different categories of lane lines, and distinguish lane lines with different IDs based on instance segmentation algorithm; Output the binary segmentation map of each pixel belonging to drivable area or non-drivable area; Extract the lane line instance segmentation result and drivable area binary segmentation result from the model output, and perform morphological post-processing on the segmentation result to eliminate noise and small cavities; The post-processed lane line segmentation result and drivable area segmentation result are output in matrix form and cached in a designated memory area for subsequent semantic grid map optimization process calls.
[0012] In step S2, a coordinate projection relationship from the OCC space to the pixel coordinate system is established based on the rotation translation matrix between the sensors; The OCC space coordinate system is ; The three-dimensional space coordinate system of the semantic grid map is generally taken as the center of the rear axle of the vehicle as the origin, The x-axis points to the forward direction of the vehicle; The y-axis points to the left side of the vehicle; The z-axis is perpendicular to the ground and points upward; The radar coordinate system is ; Wherein, represents the x-axis coordinate of the radar coordinate system, The y-axis coordinate of the radar coordinate system points to the forward direction of the vehicle; represents the x-axis coordinate of the radar coordinate system, The y-axis coordinate of the radar coordinate system points to the left side of the vehicle; represents the x-axis coordinate of the radar coordinate system, The y-axis coordinate of the radar coordinate system is perpendicular to the ground and points upward; The camera coordinate system is ; Wherein, the optical center of the camera is taken as the origin; The x-axis points to the right; The y-axis points downward; The z-axis points to the shooting direction along the optical axis.
[0013] The pixel coordinate system is : wherein, is the horizontal pixel coordinate, is the vertical pixel coordinate.
[0014] The extrinsic transformation matrix is : The extrinsic transformation matrix from the radar coordinate system to the camera coordinate system is , wherein , is a 3x3 rotation matrix, is a 3x1 translation vector with a unit of meters, representing the conversion relationship from the radar coordinate system to the camera coordinate system; represents the principal point coordinate in the camera intrinsic matrix ; The camera intrinsic matrix is : wherein, is a 3x3 matrix containing the focal length and the principal point coordinate , that is: ; The grid resolution is , and the map origin is the coordinate in the radar coordinate system .
[0015] Determine the center point of the grid in the three-dimensional coordinates of the semantic grid map: Let the index of a certain grid in the semantic grid map be , then the coordinates of the center point in the semantic grid map space are: ; wherein are the indices of the grid in the axis direction respectively; is the point in the OCC space coordinate system.
[0016] Conversion to the radar coordinate system: if the OCC space and the radar coordinate system coincide, i.e. and are the same point, and the axis systems are consistent, then ; If there is an offset, the point in the radar coordinate system can be calculated by the transformation matrix from the OCC space coordinate system to the radar coordinate system: ; and the point in the radar coordinate system is obtained .
[0017] Conversion from the radar coordinate system to the camera coordinate system: the point in the radar coordinate system is converted to the camera coordinate system by using the extrinsic transformation matrix : ; Perspective projection from the camera coordinate system to the pixel coordinate system: The point in the camera coordinate system is projected to the pixel coordinate system by using the camera intrinsic matrix .
[0018] Normalized image coordinate calculation: the point in the camera coordinate system is projected to the imaging plane of the camera, i.e. the normalized coordinates
[0019] ; Pixel coordinate conversion; the normalized coordinates are converted to pixel coordinates by using the intrinsic matrix : ; After expansion, it is: ; Through the above process, a one-to-one correspondence between any grid center point in the OCC space and the image pixel can be established, which provides a coordinate conversion basis for subsequent lane fitting, obstacle association and other algorithms.
[0020] In step S3, based on the coordinate projection relationship, the surface layer grid center points of the semantic grid map are projected to the pixel coordinate system, and lane line information is supplemented according to the projection result. Figure 2 The schematic diagram of fitting lane lines for embodiment 1 of the present application is shown in FIG. 1. The point cloud is projected to the pixel coordinate system, the center points falling on the lane line pixels are retained, and the retained center points are classified according to the numbers of the instance segmentation of the lane lines. In the classification process, a filtering rule based on the farthest point sampling is established, and the surface layer semantic grid center points are screened by using the instance lane line segmentation result, so as to ensure that the screened center points can completely represent the geometric shape of the lane lines.
[0021] For each batch of classified center points, a cubic function curve fitting is performed on the coordinates. and Before fitting, an outlier elimination operation is performed, the outliers are the center points deviating from the regular trend of the lane lines, and the outliers are determined and eliminated by calculating the Euclidean distance and angle deviation of the center points and adjacent center points. After obtaining the overhead lane lines in the OCC space, the intersection line between the lane line surface and the surface layer plane of the drivable area grid is taken as the OCC three-dimensional lane line, and the lane line information supplement is completed.
[0022] In step S4, based on the coordinate projection relationship and the drivable area segmentation result, the drivable area is semantically corrected. Figure 3 The image-related semantic supplement schematic diagram for embodiment 1 of the present application is shown in FIG. 2. Based on the coordinate projection relationship and the drivable area segmentation result, the semantics of the drivable area grid are corrected once; and according to the OCC three-dimensional lane line, the drivable area semantics on both sides of the lane line are consistency checked and corrected twice.
[0023] Based on the coordinate projection relationship and the drivable area segmentation result, the semantics of the drivable area grid are corrected once, specifically: the drivable area grid center points are projected to the pixel coordinate system; if the projected points do not fall on the drivable area segmentation pixels, the corresponding grid semantics are modified to non-drivable area, and the drivable area is corrected once.
[0024] According to the OCC three-dimensional lane line, the drivable area semantics on both sides of the lane line is consistent and corrected, specifically: the OCC three-dimensional lane line and the lane line on both sides of the area semantics output by the computer vision model, if the left side of a certain instance lane line in the image is a drivable area and the right side is a non-drivable area, then under the overhead view of the semantic grid map, a preset color mask is added to the drivable area grid on the right side of the lane line to mark it as an actual non-drivable area; the mask marking uses a preset color different from the basic color difference of the drivable area and the non-drivable area; after the correction is completed, a continuity check is performed, and if there is an isolated drivable area grid, the correction operation is re-executed.
[0025] After the drivable area correction is completed, the continuity of the corrected drivable area grid is checked, and if there is an isolated drivable area grid, it is determined that the correction is abnormal and the drivable area correction operation needs to be re-executed. Based on the secondary corrected drivable area semantics, the unknown obstacle class in the initial three-dimensional semantic grid map is extracted, single grids and sheet grids located in non-drivable areas are filtered, and unknown obstacle grid instances in the drivable area are retained.
[0026] In step S5, unknown obstacles in the initial map are extracted, and the obstacle instances in the non-driving area are filtered using the corrected drivable area semantic information; the remaining unknown obstacle instances are projected into the image space again using the coordinate projection relationship, and are associated with the detection results of the computer vision model in terms of semantics; The unknown obstacles in the initial map are extracted, specifically: the unknown obstacle class is extracted; single grids and sheet grids in the non-drivable area are filtered; and unknown obstacle grid instances in the drivable area are retained.
[0027] If new obstacle semantics need to be supplemented, only the annotation information corresponding to the new obstacle semantics to be supplemented in the image detection dataset true value (i.e. the target detection true value in the multi-task dataset) needs to be modified, and the true value data of other obstacle classes in the dataset does not need to be modified. The computer vision model is retrained with the supplemented new obstacle semantics as the training target to obtain an optimized model with new semantic detection function.
[0028] The minimum covering body of the extracted unknown obstacle grid instance is reduced in dimension and projected into the corresponding image through the projection logic of step S1, and if the grid instance is in multiple camera perspectives, multiple projections are performed.
[0029] The distance intersection over union of the projection frame obtained by calculation and the prediction frame output by the optimization model is calculated, when the distance intersection over union is greater than 0.6, it is determined that the semantics of the two are the same and the association is established, if there are multiple distance intersection over unions meeting the requirements, the highest value corresponding association relationship is taken. If all the distance intersection over unions calculated are not greater than 0.6, the corresponding unknown obstacle grid instance is not semantically associated, the "unknown obstacle" semantic state is maintained, and the spatial position information of the instance is recorded for subsequent data statistics and analysis, and the unknown obstacle semantic supplement is completed.
[0030] The distance intersection over union DIoU is an index for measuring the spatial position similarity of two boundary frames in an image, which is used to determine whether the projection frame of the unknown obstacle in the image matches the prediction frame output by the computer vision model, so that the regression becomes more stable.
[0031] ; wherein, represents the center point of the prediction frame, represents the center point of the real frame, represents the Euclidean distance between the two center points, represents the diagonal distance of the smallest closed region capable of containing the prediction frame and the real frame.
[0032] The scope of protection of the present application is not limited to the numerical values listed in embodiment 1, and persons skilled in the art can make reasonable choices according to actual conditions.
[0033] In step S6, the supplemented lane line information, the corrected drivable area semantics and the associated obstacle semantics are integrated to generate a final optimized three-dimensional semantic grid map and output.
[0034] The method of embodiment 1 of the present application adapts to the computing power limit of the vehicle-mounted platform, and all steps of the calculation process use lightweight data processing logic to avoid running delay of the vehicle-mounted system caused by excessive calculation.
[0035] Embodiment 1 of the present application proposes a semantic grid map optimization method based on multi-task road surface information, which constructs an efficient, automated and lightweight semantic map optimization pipeline by closely integrating visual perception results and grid map structure. The core effect is that without significantly increasing the system burden, the quality and practicality of the semantic map are significantly improved, thereby providing a safer and more reliable environment for the automatic driving system.
[0036] Embodiment 2 Based on the semantic grid map optimization method based on multi-task road surface information proposed in embodiment 1 of the present application, embodiment 2 of the present application further proposes a semantic grid map optimization system based on multi-task road surface information,Figure 4 A multi-task road surface information-based semantic grid map optimization system is provided for the embodiment 2 of the present application, and the system includes: A preprocessing module configured to process the acquired road environment image by using a computer vision model to obtain lane line segmentation results and drivable area segmentation results; A mapping module configured to establish a coordinate projection relationship from the OCC space to the pixel coordinate system based on a rotation and translation matrix between sensors; A supplement module configured to project the surface layer grid center points of the semantic grid map to the pixel coordinate system based on the coordinate projection relationship, and supplement the lane line information according to the projection results; A correction module configured to perform semantic correction on the drivable area based on the coordinate projection relationship and the drivable area segmentation results; A semantic association module configured to perform instance extraction on unknown obstacles in the initial map, filter out obstacle instances in the non-driving area by using the corrected drivable area semantic information, project the remaining unknown obstacle instances to the image space again by using the coordinate projection relationship, and perform semantic association with the detection results of the computer vision model; An output map module configured to integrate the supplemented lane line information, the corrected drivable area semantics, and the associated obstacle semantics, generate a final optimized three-dimensional semantic grid map, and output the map.
[0037] In the preprocessing module of the present application: a pre-trained computer vision model is deployed on a vehicle-mounted computing unit; raw images of a road environment are collected by a vehicle-mounted surround-view camera system, and the raw images are input into the computer vision model; the computer vision model performs parallel processing on the input images and performs the following inference tasks; a probability map of each pixel belonging to different categories of lane lines is output, and different ID lane lines are distinguished based on an instance segmentation algorithm; a binary segmentation map of each pixel belonging to a drivable area or a non-drivable area is output; the lane line instance segmentation results and the drivable area binary segmentation results are extracted from the model output; and morphological post-processing is performed on the segmentation results to eliminate noise and small cavities; the post-processed lane line segmentation results and drivable area segmentation results are output in matrix form and cached in a designated memory area for subsequent semantic grid map optimization process calling.
[0038] In the mapping module, the coordinate projection relationship from the OCC space to the pixel coordinate system is established, specifically: the point cloud composed of the surface layer semantic grid center points is regarded as a radar point; based on the extrinsic calibration data of the laser radar and the camera, the coordinate projection relationship from the OCC space to the pixel coordinate system is established.
[0039] In the supplement module, the supplement lane line information is specifically: projecting the grid center point to the pixel coordinate system, retaining the points falling on the lane line pixels and classifying them according to instances; performing outlier rejection and cubic function curve fitting on the classified points; calculating the intersection line of the lane line surface and the drivable area grid surface layer to generate the OCC three-dimensional lane line.
[0040] The classified points are subjected to outlier rejection, specifically: calculating the Euclidean distance and angle deviation of the center point and the adjacent points; rejecting abnormal points deviating from the regular trend of the lane line; and using the farthest point sampling rule to ensure that the screened points can completely represent the geometric shape of the lane line.
[0041] In the correction module, the drivable area is semantically corrected based on the coordinate projection relationship and the drivable area segmentation result, specifically: performing a first correction on the semantics of the drivable area grid based on the coordinate projection relationship and the drivable area segmentation result; and performing consistency checking and a second correction on the semantics of the drivable area on both sides of the lane line according to the OCC three-dimensional lane line.
[0042] The semantics of the drivable area grid is corrected once based on the coordinate projection relationship and the drivable area segmentation result, specifically: projecting the center point of the drivable area grid to the pixel coordinate system; If the projected point does not fall on the drivable area segmentation pixel, the semantics of the corresponding grid is modified to non-drivable area, realizing the first correction of the drivable area.
[0043] The semantics of the drivable area on both sides of the lane line is subjected to consistency checking and a second correction according to the OCC three-dimensional lane line, specifically: the semantics of the drivable area on both sides of the lane line output by the computer vision model and the OCC three-dimensional lane line, if the left side of an instance lane line in the image is a drivable area and the right side is a non-drivable area, then in the overhead view of the semantic grid map, a preset color mask is added to the drivable area grid on the right side of the lane line to mark it as an actual non-drivable area; the mask marking uses a preset color different from the basic color difference between the drivable area and the non-drivable area; after the correction is completed, continuity checking is performed, and if there is an isolated drivable area grid, the correction operation is re-executed.
[0044] In the correction module, unknown obstacles in the initial map are extracted as instances, specifically: unknown obstacle categories are extracted as instances; single grids and piecewise grids in the non-drivable area are filtered; and the unknown obstacle grid instances in the drivable area are retained.
[0045] Semantic association, specifically: projecting the smallest covering body of the unknown obstacle instance to the corresponding image; calculating the distance intersection ratio of the projection frame and the predicted frame output by the image multi-task model; and when the distance intersection ratio is greater than a preset value, the semantic association is established.
[0046] Embodiment 2 of the present application provides a semantic grid map optimization system based on multi-task road surface information, which constructs an efficient, automated and lightweight semantic map optimization pipeline by closely integrating visual perception results and grid map structure. The core effect is to significantly improve the quality and practicality of the semantic map without significantly increasing the system burden, thereby providing a safer and more reliable environment for the automatic driving system.
[0047] The related part of the semantic grid map optimization system based on multi-task road surface information provided in Embodiment 2 of the present application can refer to the detailed description of the corresponding part in the semantic grid map optimization method based on multi-task road surface information provided in Embodiment 1 of the present application, which will not be repeated here.
[0048] It should be noted that in this paper, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment inherent in the elements including a series of elements. Without more limitations, the element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or equipment including the element. In addition, the above technical solutions provided by the embodiments of the present application have not been described in detail, which are consistent with the implementation principles of the corresponding technical solutions in the prior art, so as not to be too repetitive.
[0049] The above describes the specific embodiments of the present application in combination with the accompanying drawings, but is not a limitation on the protection scope of the present application. For those skilled in the art, other different forms of modifications or changes can be made on the basis of the above description. Here, it is not necessary or impossible to exhaust all the embodiments. Various modifications or changes made by those skilled in the art on the basis of the technical solutions of the present application without creative labor are still within the protection scope of the present application.
Claims
1. A semantic raster map optimization method based on multi-task road surface information, characterized in that, Includes the following steps: The acquired road environment images are processed using computer vision models to obtain lane line segmentation results and drivable area segmentation results. Based on the rotation and translation matrix between sensors, a coordinate projection relationship from OCC space to the pixel coordinate system is established; Based on the coordinate projection relationship, the center point of the surface grid of the semantic raster map is projected to the pixel coordinate system, and lane line information is supplemented according to the projection result. The semantic correction of the drivable area is performed based on the coordinate projection relationship and the drivable area segmentation results. Unknown obstacles in the initial map are extracted as instances. Obstacle instances in non-driving areas are filtered out using the corrected semantic information of the drivable area. The remaining unknown obstacle instances are then projected into the image space using coordinate projection relationships and semantically associated with the detection results of the computer vision model. Integrate supplementary lane line information, corrected drivable area semantics, and associated obstacle semantics to generate and output the final optimized 3D semantic raster map.
2. The method according to claim 1, characterized in that, Based on the rotation and translation matrices between sensors, a coordinate projection relationship from OCC space to the pixel coordinate system is established, specifically as follows: The point cloud composed of the center points of the surface semantic raster is regarded as radar points; Based on the extrinsic calibration data of the LiDAR and camera, a coordinate projection relationship from the OCC space to the pixel coordinate system is established.
3. The method according to claim 1, characterized in that, Supplementing lane marking information, specifically: Project the center point of the grid onto the pixel coordinate system, retain the points that fall on the lane line pixels, and classify them by instance numbering. Outlier removal and cubic function curve fitting are performed on the classified points; Calculate the intersection of the lane line surface and the drivable area grid surface plane to generate the OCC 3D lane line.
4. The method according to claim 3, characterized in that, Outlier removal is performed on the categorized points, specifically as follows: Calculate the Euclidean distance and angular deviation between the center point and adjacent points; eliminate abnormal points that deviate from the normal direction of the lane lines; and use the farthest point sampling rule to ensure that the filtered points can completely represent the geometry of the lane lines.
5. The method according to claim 1, characterized in that, Based on coordinate projection relationships and drivable area segmentation results, semantic correction is performed on the drivable area, specifically as follows: The semantics of the drivable area raster are corrected based on the coordinate projection relationship and the drivable area segmentation results. Then, based on the OCC three-dimensional lane lines, the semantic consistency of the drivable areas on both sides of the lane lines is checked and corrected a second time.
6. The method according to claim 5, characterized in that, Based on the coordinate projection relationship and the drivable area segmentation results, the semantics of the drivable area raster are corrected, specifically as follows: Project the center point of the drivable area grid onto the pixel coordinate system; If the projection point does not fall on the segmented pixels of the drivable area, the corresponding raster semantics will be modified to non-drivable area, thus achieving one-time correction of the drivable area.
7. The method according to claim 5, characterized in that, Then, based on the OCC 3D lane lines, the semantic consistency of the drivable areas on both sides of the lane lines is checked and corrected a second time, specifically as follows: The semantics of the lane lines on both sides of the lane lines output by the OCC 3D lane lines and the computer vision model are as follows: If the left side of a lane line in an instance of the image is a drivable area and the right side is a non-drivable area, then in the top view of the semantic grid map, a preset color mask is added to the drivable area grid on the right side of the lane line to mark it as the actual non-drivable area; the mask marking uses a preset color that is different from the basic color difference between the drivable area and the non-drivable area. After the correction is completed, a continuity check is performed. If there are isolated drivable area grids, the correction operation is re-executed.
8. The method according to claim 1, characterized in that, Instance extraction of unknown obstacles in the initial map is performed as follows: Extract instances of unknown obstacle categories; filter single grids and grids in non-driving areas; retain grid instances of unknown obstacles in driving areas.
9. The method according to claim 8, characterized in that, The method also includes semantic association, specifically: The smallest bounding box of the unknown obstacle instance is projected onto the corresponding image; the distance intersection-union ratio (DIU) between the projected bounding box and the predicted bounding box output by the image multi-task model is calculated; the DIU is used to measure the spatial similarity of the two bounding boxes in the image; when the DIU is greater than a preset value, a semantic association is established.
10. A semantic raster map optimization system based on multi-task road surface information, characterized in that, include: The preprocessing module is used to process the acquired road environment images using a computer vision model to obtain lane line segmentation results and drivable area segmentation results; A mapping module is established to create a coordinate projection relationship from OCC space to the pixel coordinate system based on the rotation and translation matrices between sensors. The supplementary module is used to project the center point of the surface grid of the semantic raster map to the pixel coordinate system based on the coordinate projection relationship, and supplement lane line information according to the projection result. The correction module is used to perform semantic correction on the drivable area based on the coordinate projection relationship and the drivable area segmentation results. The semantic association module is used to extract instances of unknown obstacles in the initial map, filter out obstacle instances in non-driving areas using the corrected semantic information of the drivable area, and then project the remaining unknown obstacle instances into the image space using coordinate projection relationships, and perform semantic association with the detection results of the computer vision model. The output map module integrates supplementary lane line information, corrected drivable area semantics, and associated obstacle semantics to generate and output the final optimized 3D semantic raster map.
Citation Information
Patent Citations
Grid map generation method and system, electronic equipment and storage medium
CN112581613A
Grid map construction method, robot and machine readable storage medium
CN114779787A
Automatic driving path planning method and device based on three-dimensional space semantic information
CN118172753A
Lightweight occupancy grid prediction method and system based on large model self-labeling
CN118823139A
Plant high-definition map construction method based on edge-fog-cloud cooperation
CN120259581A