Panoramic view generation method, medium, equipment and vehicle

By combining camera and LiDAR data to construct a stereo grid map, the problem of panoramic surround view systems relying on the assumption of a flat ground is solved, generating a high-precision panoramic surround view that ensures accurate object positioning and natural stitching, supporting environmental perception for autonomous driving.

CN121095435APending Publication Date: 2025-12-09SHENZHEN DEEPROUTE AI CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511144683.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing panoramic surround view systems rely on the assumption of a flat ground, which leads to problems such as stitching distortion, object position displacement, and distance misjudgment on uneven ground, affecting the reliability of visual perception.

Method used

By acquiring image data from cameras around the vehicle and point cloud data from LiDAR, ground points are extracted and projected onto a bird's-eye view grid map to construct a 3D grid map. Combined with camera priority fusion, a panoramic view is generated, breaking the assumption of flat ground, accurately expressing terrain features, and avoiding object position offset and distance misjudgment.

Benefits of technology

It enables the accurate representation of the positional relationship between objects and vehicles on uneven ground, eliminates stitching distortion, generates high-precision panoramic surround view, and supports environmental perception for autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095435A_ABST
    Figure CN121095435A_ABST
Patent Text Reader

Abstract

The invention discloses a panoramic view generation method, a medium, equipment and a vehicle. The panoramic view generation method comprises the following steps: acquiring image data of a camera around a vehicle and point cloud data of a laser radar; extracting ground points from the point cloud data, and projecting the ground points to a preset aerial view grid map to construct a three-dimensional grid map; projecting each grid in the three-dimensional grid map into a coordinate system of each camera to obtain a color value corresponding to each grid in the image data of the camera so as to obtain a panoramic view of each camera; and according to a preset camera priority, fusing the panoramic all-around views corresponding to the cameras to generate a final panoramic all-around view.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of automatic driving, and in particular to a panoramic surround view generation method, medium, device and vehicle. BACKGROUND

[0002] The panoramic surround view system in automatic driving collects images through cameras installed around the vehicle, and fuses them into a bird's eye view panoramic image to provide the driver with a dead angle-free view of the surroundings. However, the panoramic surround view image generation in the prior art heavily relies on the assumption of flat ground, and problems such as stitching distortion, object position deviation and distance misjudgment may occur on uneven ground, affecting the reliability of visual perception. SUMMARY

[0003] The present application mainly provides a panoramic surround view generation method, medium, device and vehicle to solve the problem of the panoramic surround view system relying on the assumption of flat ground.

[0004] To solve the above technical problems, one technical solution adopted by the present application is to provide a panoramic surround view generation method, comprising: acquiring image data of cameras around the vehicle and point cloud data of a laser radar; extracting ground points from the point cloud data and projecting the ground points into a preset bird's eye grid map to construct a three-dimensional grid map; projecting each grid in the three-dimensional grid map into the coordinate system of each camera to obtain color values corresponding to each grid in the image data of each camera, thereby obtaining panoramic surround views of each camera; and fusing the panoramic surround views corresponding to each camera according to a preset camera priority to generate a final panoramic surround view.

[0005] In some embodiments, the extracting ground points from the point cloud data comprises: extracting ground points from each frame of the point cloud data; and converting the ground points in adjacent frames into the point cloud data of the current frame to extract encrypted ground points from the current frame.

[0006] In some embodiments, the projecting the ground points into a preset bird's eye grid map to construct a three-dimensional grid map comprises: constructing the bird's eye grid map with the same specification as the panoramic surround view; and projecting the encrypted ground points into the bird's eye grid map and assigning elevation values of each grid in the bird's eye grid map based on projection relationship and height values of the ground points.

[0007] In some embodiments, the assigning elevation values of each grid in the bird's eye grid map based on projection relationship and height values of the ground points comprises: assigning an average value or a median value of height values of the ground points falling into the grid as the elevation value of the corresponding grid; and deriving the elevation value of the grid without the ground points falling into it through interpolation of elevation values of neighboring grids.

[0008] In some embodiments, the projecting each grid in the stereoscopic grid map into a coordinate system of each camera to obtain a color value corresponding to each grid in image data of each camera to obtain a panoramic surround view of each camera comprises: determining coordinates of each grid in a local coordinate system based on coordinates and elevation values of each grid in the stereoscopic grid map in a global coordinate system and extrinsic parameters of the lidar; converting the coordinates of each grid in the local coordinate system into an image coordinate system of each camera to obtain a color value corresponding to each grid in the image data of each camera through parameters of the camera; and projecting the color value in the image data of each camera to the corresponding grid and setting the color value of each grid outside the effective field of view of each camera to zero to generate a panoramic surround view corresponding to each camera respectively.

[0009] In some embodiments, the fusing the panoramic surround views corresponding to each camera according to the preset camera priority to generate a final panoramic surround view comprises: obtaining a preset camera priority; and assigning the color value in each panoramic surround view of each camera to the corresponding grid in the fused panoramic surround view according to the preset camera priority to obtain the final panoramic surround view.

[0010] In some embodiments, the obtaining image data of cameras around the ego vehicle and point cloud data of lidars around the ego vehicle comprises: obtaining image data of cameras around the ego vehicle and calibrating the image data through intrinsic parameters of the cameras to correct distortion of the image data; obtaining point cloud data of lidars around the ego vehicle and filtering noise and compensating point clouds of the point cloud data; and determining poses of each camera and each lidar in a local coordinate system based on extrinsic parameters of each camera and each lidar.

[0011] To solve the above technical problems, another technical solution adopted by the present application is to provide a storage medium having program data stored thereon, the program data being executed by a processor to implement the steps of the panoramic surround view generation method as described above.

[0012] The present application also provides a vehicle-mounted device comprising a processor and a memory connected to each other, the memory storing a computer program, and the processor implementing the steps of the panoramic surround view generation method as described above when executing the computer program.

[0013] The present application also provides a vehicle comprising the storage medium as described above or the vehicle-mounted device as described above.

[0014] The beneficial effects of the present application are: different from the prior art, the present application discloses a panoramic surround view generation method, medium, equipment and vehicle. By acquiring image data of a camera around the vehicle and point cloud data of a laser radar, ground points are extracted based on the point cloud data, and the ground points are projected into a preset bird's eye grid map to construct a three-dimensional grid map. The ground points in the point cloud data are extracted and projected into the bird's eye grid map to construct a three-dimensional grid map containing elevation information, directly breaking the assumption of ground flatness, quantifying the terrain undulation features through the ground points, replacing the traditional ground flatness assumption, and accurately expressing the real terrain. Project each grid in the three-dimensional grid map into the coordinate system of each camera to obtain the color value corresponding to each grid in the image data of the camera, thereby obtaining the panoramic surround view of each camera. The image of the camera is generated on the basis of the terrain, avoiding the problems of object position deviation or distance misjudgment. Project the three-dimensional grid points in the three-dimensional grid map into the coordinate system of each camera, obtain the corresponding pixel value through the calibration parameters, realize accurate mapping from three-dimensional space to two-dimensional image, avoid position deviation caused by traditional plane projection, and the color value of each grid is directly derived from the original image of the camera, ensuring the authenticity of object texture and position. According to the preset camera priority, the panoramic surround view corresponding to each camera is fused to generate the final panoramic surround view. The effective pixels of each camera are spliced according to the preset camera priority, the color conflict in the overlapping area of the multi-camera field of view is solved, the transition at the splicing position is ensured to be natural, the splicing distortion phenomenon is eliminated, the problem that the panoramic surround system depends on the ground flatness assumption is solved, and the positional relationship between the object and the vehicle is accurately expressed in the panoramic surround view. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings, wherein: Figure 1 is a flowchart of an embodiment of the panoramic surround view generation method provided by the present application; Figure 2 is a flowchart of an embodiment of the method step 100 as shown in Figure 1 ; Figure 3 is a flowchart of an embodiment of the method step 200 as shown in Figure 1 ; Figure 4 is a flowchart of another embodiment of the method step 200 as shown in Figure 1 ; Figure 5 is an embodiment of the three-dimensional grid map generated by the method step 240 as shown in Figure 4 ; Figure 6 is a flowchart of an embodiment of the method step 240 as shown in Figure 4 Figure 7 is a flowchart of an embodiment of the method step 300 as shown in Figure 1 Figure 8 is a flowchart of an embodiment of the method step 330 as shown in Figure 7 Figure 1 is an embodiment of the panoramic surround view generated by the method step 330 as shown in Figure 9 is a flowchart of an embodiment of the method step 400 as shown in Figure 1 Figure 10 is an embodiment of the panoramic surround view generated by the method step 420 as shown in Figure 9 Figure 1 is an embodiment of the panoramic surround view generated by the method step 420 as shown in Figure 11 is a structural diagram of an embodiment of the storage medium provided by the present application; Figure 12 is a structural diagram of an embodiment of the computer device provided by the present application; Figure 13 is a structural diagram of an embodiment of the vehicle provided by the present application. DETAILED DESCRIPTION

[0016] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0017] The terms "first", "second", "third" in the embodiments of the present application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second", "third" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "multiple" is at least two, for example, two, three, etc., unless otherwise specifically limited. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.

[0018] ​​​​​Reference to“an embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase“in

[0019] Reference to“an embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase“in Figure 1 , Figure 1 is a flowchart of an embodiment of a panoramic surround view generation method provided by the application. The panoramic surround view generation method includes the following steps: 100: Obtain image data from cameras around the vehicle and point cloud data from a laser radar.

[0020] An autonomous vehicle is equipped with multiple cameras and laser radars around the vehicle body. The cameras are used to collect environmental images, and the laser radars are used to collect three-dimensional point cloud data. During the driving of the vehicle, the cameras capture two-dimensional images of the surrounding environment in real time, and the laser radars scan the surrounding space by emitting laser beams and receive reflected signals to generate point cloud data.

[0021] Optionally, the camera type can be a fisheye camera or a normal camera, and the camera is arranged around the vehicle, such as front, rear, left, and right, to ensure that the camera field of view can completely cover the 360° space around the vehicle.

[0022] The image data is a two-dimensional planar image captured by a camera sensor, which contains a pixel array and RGB color information of each pixel, reflecting the texture, color, and two-dimensional contour of objects in the environment. Image data can intuitively present the appearance, color, and texture details of objects, and is suitable for visual recognition.

[0023] Point cloud data is a set of three-dimensional coordinate points calculated by measuring the time difference or phase difference of the laser beams reflected by the laser radar, which can be accompanied by reflection intensity information, forming a three-dimensional point cloud model of the environment. Point cloud data directly provides the three-dimensional coordinates of objects, which can accurately represent the shape, position, and distance of objects and is not affected by light. Laser radars can generate millions of points per second, forming a dense point cloud that describes the ground undulations and obstacle contours in detail. The coordinates of each point implicitly indicate the distance from the sensor, which can be used for obstacle detection and path planning.

[0024] Image data is a two-dimensional plane projection, and point cloud data is a three-dimensional spatial distribution. The combination of the two can realize complementary perception of two-dimensional texture and three-dimensional structure. Image data provides rich visual semantics, and point cloud data provides accurate spatial geometry. In the generation of an Around View Monitor (AVM), the former is used for color mapping, and the latter is used for constructing a three-dimensional ground model, thereby solving the limitations of traditional methods that rely on the assumption of a flat ground.

[0025] Further, referring to Figure 2 , step 100 further comprises the following steps: 110: Obtain image data from cameras around the ego vehicle, and calibrate the image data for camera internal parameters to correct the distortion of the image data.

[0026] The environment image is collected by fisheye cameras around the vehicle body, and then the image is calibrated for internal parameters to correct the lens distortion. The image geometric accuracy is ensured, and the image distortion caused by the camera optical system is eliminated, thereby providing non-distorted two-dimensional image data for subsequent projection and stitching.

[0027] Internal parameter calibration is a process of correcting image distortion by calculating the internal parameters of the camera such as focal length, principal point coordinates, distortion coefficient, etc. through a mathematical model. It can eliminate the barrel-shaped and pillow-shaped distortion of wide-angle lenses such as fisheye cameras, restore the geometric relationship of the real scene, and ensure the shape and position accuracy of objects in the image.

[0028] 120: Obtain point cloud data from laser radars around the ego vehicle, and filter noise points and compensate point clouds.

[0029] Three-dimensional point cloud data is collected by laser radars, and then noise points are filtered, including removing outliers, dust reflection, and other noise, and point cloud compensation is performed. The accuracy and continuity of the point cloud data are improved, thereby providing a high-quality point cloud basis for subsequent ground point extraction and three-dimensional modeling.

[0030] 130: Determine the pose of each camera and each laser radar in the local coordinate system based on the external parameters of each camera and each laser radar.

[0031] Based on the pre-calibrated camera and laser radar external parameters, the accurate pose of each camera and laser radar in the local coordinate system is calculated, the spatial alignment of image data and point cloud data is realized, the coordinate conversion accuracy of the two in the same coordinate system is ensured, and a foundation is laid for the subsequent fusion of point clouds and images.

[0032] The extrinsic parameters of the camera and the lidar describe the relative position relationship between the camera and the lidar, or between the sensor and the local coordinate system. The extrinsic parameters of the camera and the lidar are used to reflect the installation position and orientation of the sensor in space, and need to be determined through a joint calibration experiment, so as to realize the spatial unification of different sensor data, for example, to project the three-dimensional point cloud of the lidar into the two-dimensional image of the camera, or vice versa, to provide a coordinate conversion basis for the fusion of point cloud and image.

[0033] The local coordinate system is a relative coordinate system with the vehicle or sensor as the origin, such as the vehicle body coordinate system, which is used to describe the pose of the camera, the lidar, and the coordinates of the spatial points. The local coordinate system is used to unify the spatial reference system of multi-sensor data, to ensure that image data, point cloud data, and grid maps are converted and fused in the same coordinate system, and to avoid projection errors caused by coordinate confusion.

[0034] 200: Extract ground points from the point cloud data and project the ground points into a preset bird's eye grid map to construct a three-dimensional grid map.

[0035] For each frame of lidar point cloud data, separate the three-dimensional points belonging to the ground from the original point cloud. Create a two-dimensional bird's eye grid map with the same spatial range and resolution as the final panoramic surround view, covering a 360° area around the vehicle. Project the ground points into the bird's eye grid map, and count the height values of the ground points falling into each grid cell to form a three-dimensional grid map.

[0036] The ground points are three-dimensional coordinate points belonging to the ground surface extracted from the lidar point cloud data through semantic segmentation or ground segmentation algorithms. The ground points are densely distributed and can reflect the actual undulations of the ground. As the core data for constructing a three-dimensional grid map, they replace the traditional assumption of flat ground and provide a real geometric basis for the panoramic surround view.

[0037] The bird's eye grid map is a two-dimensional grid formed by dividing the space around the vehicle into grids with a preset resolution. The bird's eye grid map has the same spatial range and resolution as the final panoramic surround view. As a carrier for projecting ground points, the bird's eye grid map converts three-dimensional ground points into structured two-dimensional grid data, facilitating subsequent elevation value calculation and image projection.

[0038] The three-dimensional grid map is a pseudo-three-dimensional grid model that assigns an elevation value to each grid based on the bird's eye grid map. The three-dimensional grid map is not a complete three-dimensional model, but only represents the ground undulations through grid elevation values. The three-dimensional grid map can accurately represent the true shape of the ground, breaking the dependence on flat ground in traditional panoramic surround views, providing a three-dimensional coordinate basis for subsequent projection, and avoiding image distortion caused by uneven ground.

[0039] The stereoscopic grid map accurately restores the undulating state of the ground, such as ramps, speed bumps, potholes, etc., replacing the ground flat assumption in the traditional panoramic surround view generation method, providing a real three-dimensional space reference for subsequent grid and image projection and multi-camera fusion, and fundamentally solving the problems of splicing distortion and distance misjudgment caused by uneven ground.

[0040] Further, referring to Figure 3 extracting the ground points from the point cloud data includes the following steps: 210: Extracting ground points from each frame of point cloud data.

[0041] For each frame of original point cloud data collected by the laser radar, a semantic segmentation algorithm such as a ground classification network based on deep learning, or a ground segmentation algorithm such as RANSAC plane fitting, is used to separate three-dimensional points (x, y, z) belonging to the ground from the point cloud, and filter out non-ground object points such as vehicles, pedestrians, and buildings.

[0042] By extracting ground points from each frame of point cloud data, the ground contour of the current frame is preliminarily obtained.

[0043] 220: Convert the ground points in the adjacent frame into the point cloud data of the current frame to extract the encrypted ground points from the current frame.

[0044] Using the pose estimation result of the ego vehicle, the ground points extracted from the adjacent frame are projected into the point cloud data of the current frame through coordinate conversion, merged with the ground points of the current frame, and a set of encrypted dense ground points is formed.

[0045] By encrypting the ground points, the problem of sparsity of single-frame ground points is solved, and the ground details are supplemented by multiple frames of data to ensure that the ground points are evenly distributed and the density is sufficient, providing sufficient data support for subsequent grid modeling.

[0046] Further, referring to Figure 4 projecting the ground points into the preset bird's eye grid map to construct a stereoscopic grid map includes the following steps: 230: Constructing a bird's eye grid map with the same specification as the panoramic surround view.

[0047] A two-dimensional grid map is created that is completely consistent with the spatial range of the final panoramic surround view and has the same resolution. For example, a grid map covering an area of 5m x 5m around the vehicle with a grid unit size of 0.1m x 0.1m. The bird's eye grid map forms a grid framework from a top-down perspective, and each grid corresponds to a small area in the physical space.

[0048] The ability to construct a bird's eye grid map with the same specification as the panoramic surround view provides a structured carrier for the three-dimensional coordinates of the ground points, converting discrete ground points into regular grid data, which is convenient for subsequent elevation value calculation and image projection.

[0049] 240: Project the encrypted ground points onto the bird's-eye view raster map, and assign elevation values ​​to each grid in the bird's-eye view raster map based on the projection relationship and the height values ​​of the ground points to construct a three-dimensional raster map.

[0050] The encrypted dense ground points are projected onto the bird's-eye view grid map according to their coordinates to determine the grid cell in which each ground point falls. For each grid cell, the height values ​​of all ground points falling within it are counted, and the average or median is taken as the elevation value of that grid. After all grid cells are assigned elevation values, a 2.5D grid model containing planar position and height is formed, i.e., a three-dimensional grid map, as shown in Figure 5.

[0051] See Figure 5 , Figure 5 Is it like this? Figure 4 This is an embodiment of the 3D grid map generated in step 240 of the method shown. The height variation of the 3D grid map can accurately reflect the undulations of the ground. Blank areas in the map represent grids without corresponding height values, corresponding to areas outside the effective detection range of the LiDAR or obstacles such as buildings whose height cannot be quantified in real-world scenarios.

[0052] The elevation value of a ground point is the z-coordinate of the ground point in the three-dimensional coordinate system, acquired by the lidar, representing the vertical height of that point from the origin of the local coordinate system. The elevation value of the ground point directly reflects the actual undulation of the ground and is the raw data for calculating the raster elevation value; its accuracy directly determines the precision of the 3D raster map in reproducing the ground morphology.

[0053] The elevation value of a raster is the height attribute of each raster cell in a bird's-eye view raster map, calculated statistically or by interpolation from the height values ​​of ground points falling into that raster. By replacing the simplified assumption of a flat ground in traditional panoramic view generation methods with elevation values, the raster map can accurately reflect complex terrain such as ramps and speed bumps.

[0054] This process transforms discrete ground points into a continuous 3D ground model, accurately representing the ground's undulations and providing a realistic spatial coordinate basis for subsequent projection between the raster and the image. Data sparsity is addressed through multi-frame ground point fusion, and the conversion from discrete points to a continuous model is achieved through raster elevation value calculation. Ultimately, a 3D raster map that accurately represents uneven ground is constructed, providing a crucial 3D spatial foundation for subsequent panoramic view generation.

[0055] Further, see Figure 6 Step 240 also includes the following steps: 241: Assign the average or median height value of the ground points projected onto the grid as the elevation value of the corresponding grid.

[0056] For all the encrypted ground points projected into the same grid cell, the height value is extracted, the average or median is calculated, and the value is assigned to the grid as its elevation value.

[0057] The average value calculation method is suitable for the scene where the ground points are uniformly distributed and the noise is less, and can reflect the overall height trend of the ground. The median calculation method is preferentially used in the case of a small number of outlier noise points, which can reduce the influence of extreme values on the grid elevation and improve the data robustness.

[0058] The discrete ground point cloud data is converted into structured grid attributes, so that each grid cell obtains a quantized value representing the height of the regional ground, providing basic data for subsequent stereo modeling.

[0059] 242: The elevation value of the grid without ground points is obtained by interpolation of the elevation values of the neighboring grids.

[0060] For the blank grid that does not fall into any ground point, such as the area where the ground points are missing due to occlusion or sensor blind area, the elevation value is calculated by interpolation of the elevation values of the neighboring grids. Among them, the interpolation can be performed by bilinear interpolation or Kriging interpolation algorithm.

[0061] The blank area in the grid map is eliminated to ensure that the ground elevation values are continuously distributed in space, avoiding distortion of the ground form caused by sparse data, and finally generating a complete and smooth three-dimensional ground model.

[0062] Through the combination strategy of statistical assignment and interpolation, the conversion from discrete point cloud to continuous grid is realized. The average or median reduces the influence of single point error and improves the reliability of the elevation value; the interpolation algorithm ensures that the grid map is not broken, accurately restores the micro-topography such as ground slope and depression; the three-dimensional ground information is compressed into regular grid data, providing an efficient coordinate query interface for the projection between the subsequent grid and image, supporting the accurate stitching of panoramic surround view in non-flat ground scene.

[0063] 300: Project each grid in the stereo grid map into the coordinate system of each camera to obtain the color value corresponding to each grid in the image data of the camera, thereby obtaining the panoramic surround view of each camera.

[0064] According to the pre-calibrated sensor extrinsic parameters, the three-dimensional coordinates of each grid in the stereo grid map are converted to the camera coordinate system, and the three-dimensional points in the camera coordinate system are converted to the pixel coordinates of the image plane by the perspective projection formula using the camera intrinsic matrix. According to the pixel coordinates obtained by projection, the RGB color value of the position is read from the original image collected by the camera.

[0065] The extracted color value is assigned to the corresponding grid cell in the stereo grid map, so that each grid has both three-dimensional elevation information and two-dimensional color attribute. Repeat the above steps for multiple cameras around the vehicle body to generate color grid maps for each camera view, and then stitch them into a 360° panoramic surround view through the extrinsic calibration results.

[0066] This process realizes the deep fusion of the geometric accuracy of point cloud data and the texture details of image data through the technical path of three-dimensional coordinates, two-dimensional pixels, and color mapping. The final panoramic surround view can accurately reflect the terrain height around the vehicle and present the real environmental color, providing a more reliable environmental perception basis for automatic driving parking and low-speed driving.

[0067] Further, referring to Figure 7 , step 300 further includes the following steps: 310: Based on the coordinates and elevation values of each grid in the stereo grid map in the global coordinate system, determine the coordinates of each grid in the local coordinate system through the extrinsic parameters of the laser radar.

[0068] The global coordinate system is based on the laser radar or world coordinate system and contains the three-dimensional coordinates (X global , Y global , Z global ) of each grid in the stereo grid map, where Z global is the grid elevation value The local coordinate system is a coordinate system with the vehicle or laser radar sensor as the origin, used to align with the camera coordinate system.

[0069] Use the extrinsic matrix of the laser radar to convert the coordinates of the grid in the global coordinate system to the local coordinate system, unify the coordinate reference, and make the conversion of the grid coordinates and the subsequent camera coordinate system consistent in space.

[0070] 320: Convert the coordinates of each grid in the local coordinate system to the image coordinate system of each camera through the parameters of the camera to obtain the corresponding color value of each grid in the image data of each camera.

[0071] Through the intrinsic matrix and extrinsic matrix of the camera, the coordinate conversion relationship is established. The three-dimensional coordinates (X local , Y local , Z local ) of the grid in the local coordinate system are converted to the pixel coordinates (u, v) in the image coordinate system. Specifically, first convert the local coordinates to the camera coordinate system (X cam , Y cam , Z cam ) through the camera extrinsic parameters, then use the intrinsic matrix for perspective projection to obtain the pixel coordinates (u, v). Finally, read the corresponding RGB color value from the camera original image according to the pixel coordinates (u, v).

[0072] A mapping relationship between the three-dimensional grid and the two-dimensional image pixels is established to realize accurate association between the spatial position and the color texture, and to give the grid a real environment color.

[0073] 330: Project the color values in the camera image data to the corresponding grid, and set the color values of the grid outside the effective field of view of each camera to zero to generate a panoramic surround view corresponding to each camera respectively, as shown in Figure 8 .

[0074] The color values extracted in step 320 are assigned to the corresponding grid cells in the three-dimensional grid map, so that each grid contains both elevation information and color attributes. By the field of view angle parameter of the camera, it is determined whether the grid is within the effective field of view of the camera. The color value of the grid outside the range is set to zero, i.e. black, to avoid interference from invalid areas. Repeat steps 310-330 for multiple cameras around the vehicle body to generate a color grid map corresponding to each camera, i.e. a single-camera panoramic surround view. Filter invalid data to ensure that the surround view only retains valid color information in the observable area of the camera, and lay the foundation for subsequent multi-view stitching through separate camera processing.

[0075] Refer to Figure 8 , Figure 8 For example, set fisheye cameras in front, back, left and right of the vehicle body to generate panoramic surround views corresponding to the front, rear, left and right cameras respectively to show the ground state and color around the vehicle.

[0076] By color value projection, the accurate ground shape of the laser radar contained in the three-dimensional grid is fused with the real environment texture of the camera contained in the image data, solving the problems of misalignment and distance misjudgment caused by the flat ground assumption in traditional panoramic surround view; the field of view range filtering ensures that the surround view has no redundant noise, providing high-precision environment perception basis for automatic driving parking, low-speed obstacle avoidance and other scenarios. The final panoramic surround view of each camera has both three-dimensional spatial precision and natural color performance, providing high-quality input for subsequent multi-camera surround view stitching.

[0077] 400: Fuse the panoramic surround views corresponding to each camera according to the preset camera priority to generate a final panoramic surround view.

[0078] According to the installation position and functional requirements of the camera, set the priority order, for example, the priority of the front view camera > the rear view camera > the left view camera > the right view camera. By the camera extrinsic calibration result, calculate the overlapping area of each camera surround view in advance to determine the priority range.

[0079] For the same grid in the overlapping area, compare the priority of the multiple cameras covering the grid, select the highest priority camera color value as the final grid color, and cover the color value of the low priority camera; for the independent area without overlap, directly retain the original color value of each camera, and splice according to the physical position to form a global surround view frame.

[0080] After priority fusion, boundary smoothing and color correction, a complete panoramic surround view covering the 360° environment around the vehicle is generated, which contains both the elevation information of the stereo grid and the real color texture of the camera, and can be directly used as the environment perception interface for automatic driving.

[0081] Through the combination strategy of priority rules, overlap fusion and color correction, the conflict problem of multi-camera view in the spatial overlap area is solved, ensuring that the final surround view has no redundant information, no splicing fault and consistent color, providing high-precision and intuitive environment visualization results for automatic driving parking, narrow road passing and other scenes.

[0082] Further, referring to Figure 9 , step 400 further comprises the following steps: 410: Obtain the preset camera priority.

[0083] The preset camera priority rule is read from the system configuration, which is predefined according to the vehicle driving scene, camera function positioning and safety requirement, and adjusted according to the real-time driving state, for example, setting the rear view camera as the highest priority when reversing, and improving the priority of the corresponding side view camera when turning.

[0084] The overlapping area boundary of each camera surround view is determined, and the priority rule is only effective in the overlapping area, and the original camera color value is directly retained in the non-overlapping area.

[0085] 420: According to the preset camera priority, assign the color value in each camera panoramic surround view to the corresponding grid in the fused panoramic surround view to obtain the final panoramic surround view.

[0086] Traverse each grid of the fused panoramic surround view, check whether there is a corresponding color value in each camera surround view. For multiple cameras covering the same grid, according to the priority rule obtained in step 410, select the highest priority camera color value as the final color value of the grid, and the color value of the low priority camera is automatically covered. For the grid covered by only one camera, directly use the color value of the camera without priority judgment.

[0087] For example, if a grid is covered by a front view camera with priority 1 and a right view camera with priority 4 at the same time, the color value of the front view camera with priority 1 is finally retained.

[0088] The color conflicts in the overlapping areas are eliminated by priority arbitration, ensuring that each grid in the fused panoramic surround view has only a unique and valid color value, and avoiding visual confusion caused by multi-camera data superposition.

[0089] Referring to Figure 10 , Figure 10 is Figure 8 the final panoramic surround view formed by splicing the four panoramic surround views shown in Figure 8 The images in the front, rear, left and right directions in Figure 10 are fused according to the priority of front view camera > rear view camera > left view camera > right view camera, to obtain the final panoramic surround view as shown in The final panoramic surround view contains visual information and ground relief features of the vehicle in the front, rear, left and right directions.

[0090] The arbitration strategy of the priority rule ensures that the color data of the key field of view is preferentially retained, improving the reference value of the surround view for driving decisions; without complex pixel-level fusion algorithms, efficient color value screening is achieved through preset rules, reducing the computational complexity; the finally generated panoramic surround view realizes conflict-free integration of multi-camera data, retaining the effective field of view information of each camera and ensuring the color accuracy of the key area through the priority rule, providing a clear and reliable 360° environment visualization result for the automatic driving system.

[0091] Referring to Figure 11 , Figure 11 is a structural schematic diagram of an embodiment of the storage medium provided by the present application.

[0092] The storage medium 30 stores program data 31, which, when executed by a processor, implements the steps of the panoramic surround view generation method as described in Figure 1 .

[0093] The program data 31 is stored in a storage medium 30, and includes a plurality of instructions for causing a network device (such as a router, personal computer, server, etc.) or a processor to execute all or part of the steps of the method described in various embodiments of the present application.

[0094] Optionally, the storage medium 30 can be a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc. Various media that can store program data.

[0095] Referring to Figure 12 , Figure 12 is a structural schematic diagram of an embodiment of the vehicle-mounted device provided by the present application.

[0096] The computer device 40 comprises a processor 42 and a memory 41 connected with each other, the memory 41 stores a computer program, and the processor 42 implements the steps of the panoramic surround view generation method when executing the computer program.

[0097] Referring to Figure 13 , Figure 13 A vehicle 50 is provided, comprising the on-board device 40 as described above. The vehicle 50 can complete the process of panoramic surround view generation through the control of the on-board device 40 during driving. The on-board device 40 is installed in the vehicle and is a computer type device, which is linked with various sensing modules and control modules in the vehicle to realize visual perception, sensor data processing and image display process.

[0098] Different from the prior art, the present application divides the ground into regular three-dimensional grid units based on dense ground point cloud data, each grid contains coordinates and elevation values in the global coordinate system, breaks through the traditional ground flatness assumption, accurately models the ground undulation, provides a three-dimensional geometric reference for subsequent color mapping, and improves the spatial accuracy of the panoramic surround view.

[0099] The present application also extracts color values from the camera image according to the pixel coordinates obtained by projection, and assigns the corresponding grid units in the three-dimensional grid image; the color value of the grid outside the camera field of view is set to zero. The real environment texture of the camera, such as road markings and green belts, is assigned to the grid, realizing data fusion of three-dimensional geometry and two-dimensional texture, and solving the problem of texture missing in traditional panoramic surround view. The present application also eliminates color conflicts caused by overlapping of multiple camera fields of view by presetting camera priority rules, ensures that each grid in the fused surround view has only one valid color value, avoids visual redundancy, and improves the consistency and reliability of the surround view. The present application also simplifies the complexity of multi-view stitching through camera processing, provides standardized input for subsequent global surround view fusion, and facilitates fault isolation.

[0100] Through the combination of the above technical means, the present application realizes the deep fusion of the laser radar with geometric precision advantage and the camera with texture detail advantage, and finally generates a panoramic surround view with high-precision ground morphology and real environment color, effectively supporting the environmental perception needs of automatic driving in low-speed scenarios, and avoiding the problems of stitching distortion, object projection position offset and distance misjudgment in the prior art.

[0101] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the storage medium embodiment and the computer device embodiment are basically similar to the method embodiment, so the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0102] The application is operable with numerous general purpose or special purpose car-mounted computing system environments or configurations. Examples include personal computers, hand-held or laptop devices, tablet devices, multiprocessor systems, microprocessor-based systems, network PCs, minicomputers, distributed computing environments that include any of the above systems or devices, and the like.

[0103] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the above-described device embodiments are only illustrative, for example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed.

[0104] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.

[0105] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can be physically present alone, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0106] The above is only an embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for generating a panoramic view, characterized in that, include: Acquire image data from cameras around the vehicle and point cloud data from LiDAR; Ground points are extracted from the point cloud data and projected onto a preset bird's-eye view grid map to construct a three-dimensional grid map; Each grid in the stereo grid image is projected onto the coordinate system of each camera to obtain the color value corresponding to each grid in the image data of the camera, thereby obtaining a panoramic view of each camera. The panoramic views corresponding to each camera are merged according to the preset camera priority to generate the final panoramic view.

2. The panoramic view generation method according to claim 1, characterized in that, Extracting ground points from the point cloud data includes: Extract ground points from each frame of the point cloud data; The ground points in adjacent frames are converted into the point cloud data of the current frame to extract the encrypted ground points from the current frame.

3. The panoramic view generation method according to claim 2, characterized in that, The step of projecting the ground points onto a preset bird's-eye view grid map to construct a three-dimensional grid map includes: Construct the bird's-eye view grid map of the same specifications as the panoramic panoramic view; The encrypted ground points are projected onto the bird's-eye view grid map, and the elevation values ​​of each grid in the bird's-eye view grid map are assigned based on the projection relationship and the height values ​​of the ground points to construct a three-dimensional grid map.

4. The panoramic view generation method according to claim 3, characterized in that, The elevation values ​​of each grid cell in the bird's-eye view grid map are assigned based on the projection relationship and the height values ​​of the ground points, including: The average or median value of the height of the ground points whose projection falls into the grid is assigned as the elevation value of the corresponding grid. The elevation values ​​of the grid cells into which no ground point falls are obtained by interpolating the elevation values ​​of the neighboring grid cells.

5. The panoramic view generation method according to claim 3, characterized in that, The step of projecting each grid in the stereo grid image onto the coordinate system of each camera to obtain the color value corresponding to each grid in the image data of the camera, thereby obtaining a panoramic view of each camera, includes: Based on the coordinates and elevation values ​​of each grid in the global coordinate system in the three-dimensional grid image, the coordinates of each grid in the local coordinate system are determined by the extrinsic parameters of the lidar. The coordinates of each grid in the local coordinate system are transformed to the image coordinate system of each camera using the parameters of the camera, so as to obtain the color value of each grid in the image data of each camera; The color values ​​in the image data of each camera are projected onto the corresponding grid, and the color values ​​of the grids outside the effective field of view of each camera are set to zero, so as to generate a panoramic view corresponding to each camera.

6. The panoramic view generation method according to claim 1, characterized in that, The step of fusing the panoramic views corresponding to each camera according to a preset camera priority to generate the final panoramic view includes: Obtain the preset camera priority; According to the preset camera priority, the color values ​​in the panoramic view of each camera are assigned to the corresponding grid in the fused panoramic view to obtain the final panoramic view.

7. The panoramic view generation method according to claim 1, characterized in that, The acquisition of image data from cameras around the vehicle and point cloud data from LiDAR includes: Acquire image data from cameras around the vehicle, and perform intrinsic parameter calibration on the image data to correct the distortion of the image data; Acquire point cloud data from the lidar surrounding the vehicle, and perform noise filtering and point cloud compensation on the point cloud data; Based on the extrinsic parameters of each camera and each lidar, the pose of each camera and each lidar in the local coordinate system is determined.

8. A storage medium storing program data thereon, characterized in that, When the program data is executed by the processor, it implements the steps of the panoramic view generation method as described in any one of claims 1-7.

9. A vehicle-mounted device, characterized in that, It includes an interconnected processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the steps of the panoramic view generation method as described in claims 1-7.

10. A vehicle, characterized in that, The vehicle includes the storage medium as described in claim 8 or the on-board equipment as described in claim 9.

Citation Information

Cited By

  • Point cloud-based ground line generation method and device and intelligent driving equipment

    CN121982131A

  • A ground line generation method and device based on a point cloud and an intelligent driving device

    CN121982131B