A vehicle bottom transparent registration method, system, medium and program product

Through deep neural network extracting vehicle displacement and combining the confidence of ground reference point clouds, the problem of error and misalignment in the dynamic process of vehicle bottom transparency technology is solved, and a high accuracy and robust bottom transparent image generation is achieved.

CN119169064BActive Publication Date: 2025-05-23SHENZHEN PERCHERRY TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411264343.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-05-23
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

The existing undercarriage transparency technology is easily affected by external factors during the dynamic process of the vehicle, resulting in distortion of information such as speed and angle, resulting in estimation errors and image misalignment.

Method used

A deep neural network is used to extract vehicle displacement from pure visual information, combine vehicle speed, angle and other information to determine the ground reference point cloud, and select the displacement estimation results of the visual or sensor based on the point cloud confidence, and finally generate a transparent image of the vehicle bottom.

Benefits of technology

Effectively use visual information to improve positioning accuracy, switch positioning methods through confidence evaluation, improve system robustness, avoid image misalignment caused by speed errors, and achieve accurate and stable display of transparent images under the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119169064B_ABST
    Figure CN119169064B_ABST
Patent Text Reader

Abstract

A transparent registration method, system, medium and program product for underbody vehicle, which relates to the field of driving calculation, includes: inputting two adjacent frames of vehicle top view images into a neural network model to obtain a vehicle displacement map; determining a ground reference point cloud according to the speed information, turning angle information, vehicle displacement map and vehicle top view of the test vehicle; determining the confidence level according to the coordinate distribution of multiple points in the ground reference point cloud; determining the displacement result data in the X and Y directions when the confidence level is higher than a preset credible threshold; determining the displacement result data using a speed sensor when the confidence level is not higher than the preset credible threshold; determining multiple corresponding displacement result data according to multiple vehicle displacement maps to obtain the coordinates of the vehicle body positioning points; generating a transparent underbody vehicle image according to the coordinates of the vehicle body positioning points. The method is implemented to extract vehicle displacement from pure visual information based on a deep neural network, avoiding the image misalignment problem caused by speed error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle driving calculation, and in particular to a vehicle bottom transparent registration method, system, medium and program product. Background Art

[0002] The bottom transparent technology is a technology that installs a camera at the front of the vehicle to obtain image data in front of the vehicle, determines the front and rear position coordinates of the vehicle through an algorithm, uses historical view data to fill in the blank information under the vehicle, and displays it on a large screen inside the vehicle, so that the driver can intuitively perceive the bottom of the vehicle and the surrounding environment. This technology is of great significance in the automotive field, mainly reflected in improving driving safety, parking and handling performance, enhancing driving experience, and can also be used to assist vehicle maintenance.

[0003] The common vehicle bottom transparency technology currently uses the CAN bus to obtain vehicle speed and turning angle information, and then calculates the vehicle displacement and matches the vehicle bottom image. This solution obtains speed, steering and other data through body sensors, estimates the vehicle displacement at adjacent moments based on the kinematic model, combines the displacement information with the image stitching algorithm, and generates a transparent vehicle bottom view.

[0004] However, the implementation of related technologies is easily affected by external factors. Changes in road conditions and loads can cause distortion of information such as vehicle speed and turning angle, resulting in estimation errors. Especially during dynamic processes such as vehicle turning, acceleration and deceleration, the accuracy of the calculated displacement data decreases, resulting in obvious misalignment with the actual image. Summary of the invention

[0005] The present application provides a vehicle bottom transparent registration method, system, medium and program product, which extracts vehicle displacement from pure visual information based on a deep neural network to avoid image misalignment problems caused by speed errors.

[0006] In a first aspect, the present application provides a vehicle bottom transparent registration method, which is applied to a vehicle bottom transparent registration system, the method comprising: inputting two adjacent frames of a vehicle top view of a test vehicle into a pre-trained neural network model to obtain a vehicle displacement map; the vehicle displacement map represents the displacement of each pixel in the two frames before and after the image in the X and Y directions; determining a ground reference point cloud according to the vehicle speed information, turning angle information, vehicle displacement map and vehicle top view of the test vehicle; determining the confidence of the ground reference point cloud relative to the vehicle body position according to the coordinate distribution of multiple points in the ground reference point cloud; when the confidence is higher than a preset credible threshold, determining the displacement result data of the test vehicle in the X and Y directions based on the ground reference point cloud; when the confidence is not higher than the preset credible threshold, using the displacement data in the speed sensor as the displacement result data of the test vehicle in the X and Y directions; determining multiple corresponding displacement result data of the test vehicle according to multiple vehicle displacement maps to obtain the coordinates of the vehicle body positioning points; generating a vehicle bottom transparent image according to the coordinates of the vehicle body positioning points and the vehicle top view.

[0007] In the above embodiment, the vehicle bottom transparent registration system determines the vehicle displacement map used to represent the displacement of each pixel in the front and rear frames in the X and Y directions, and then determines the ground reference point cloud based on the vehicle speed, turning angle, displacement map and top view, and determines the confidence of the relative vehicle body position based on the point cloud coordinate distribution. When the confidence is high, the X and Y direction displacement is determined based on the point cloud, and when the confidence is low, the speed sensor data is used as the displacement. Multiple displacements are determined through multiple frame displacement maps, and the vehicle body positioning coordinates are obtained, and finally the vehicle bottom transparent image is generated, avoiding the image misalignment problem caused by speed error.

[0008] In combination with some embodiments of the first aspect, in some embodiments, a ground reference point cloud is determined based on the test vehicle's speed information, turning angle information, vehicle displacement map and vehicle top view, specifically including: determining the displacement error of the test vehicle in the vehicle displacement map based on the test vehicle's speed information and turning angle information; filtering out erroneous areas in the vehicle top view based on the displacement error, and removing vertical objects in the vehicle top view to obtain a ground reference point cloud.

[0009] In the above embodiment, the bottom transparent registration system determines the displacement error of the test vehicle in the vehicle displacement map, filters out the erroneous area in the top view of the vehicle based on the displacement error, and removes vertical objects to obtain a ground reference point cloud. The interference caused by the vehicle's own movement can be removed, and irrelevant vertical objects can be eliminated, thereby improving the accuracy and stability of the bottom transparent image.

[0010] In combination with some embodiments of the first aspect, in some embodiments, the confidence of the ground reference point cloud relative to the vehicle body position is determined based on the coordinate distribution of multiple points in the ground reference point cloud, specifically including: calculating the spatial distribution density and directional consistency of multiple points in the ground reference point cloud, and taking the spatial distribution density and directional consistency as point cloud distribution features; obtaining vehicle motion state information including vehicle speed information and turning angle information, and mapping the vehicle motion state information to a vehicle motion risk level; constructing a confidence estimation model based on the point cloud distribution characteristics, the vehicle motion risk level, and the confidence compensation coefficient corresponding to the road material and weather conditions, and outputting the confidence of the ground reference point cloud relative to the vehicle body position.

[0011] In the above embodiment, the vehicle bottom transparent registration system calculates the spatial distribution density and directional consistency of the points in the ground reference point cloud as the point cloud distribution features, then obtains the vehicle state information such as vehicle speed and turning angle and maps it into the motion risk level, integrates the point cloud features, motion risk, road material, and weather compensation coefficient, builds a confidence estimation model, and outputs the confidence of the ground point cloud relative to the vehicle body. The influence of the vehicle motion state and environmental factors on the vehicle positioning credibility is determined, and the robustness of the vehicle bottom transparent image generation can be effectively improved by dynamically evaluating the positioning reliability.

[0012] In combination with some embodiments of the first aspect, in some embodiments, before the step of inputting two adjacent frames of images of the vehicle top view of the test vehicle into a pre-trained neural network model to obtain a vehicle displacement map, the method also includes: acquiring a full-view front view of the test vehicle through a vehicle-mounted panoramic camera; performing fisheye correction and perspective transformation on the full-view front view to obtain a vehicle top view.

[0013] In the above embodiment, before using the neural network to extract the vehicle displacement, the vehicle bottom transparent registration system first obtains the full-view front view through the on-board panoramic camera, and then performs fisheye correction and perspective transformation on it to obtain the vehicle top view. It can expand the field of view, obtain more complete environmental information, and eliminate the distortion caused by the fisheye lens and imaging perspective, providing a higher quality input image.

[0014] In combination with some embodiments of the first aspect, in some embodiments, fisheye correction and perspective transformation are performed on the full-view front view to obtain a top view of the vehicle, specifically including: performing distortion correction on the full-view front view to eliminate lens profile distortion and perspective distortion; extracting texture features of the full-view front view; texture features include grayscale co-occurrence matrix, gradient histogram and local binary pattern; clustering and screening texture features to remove vehicle body areas and invalid areas to obtain a texture feature set; according to the texture quality evaluation index of the texture feature set, texture features with texture quality lower than a preset quality threshold are removed to obtain a top view of the vehicle.

[0015] In the above embodiment, when obtaining the top view of the vehicle, the transparent bottom registration system performs distortion correction on the full-view front view, eliminates lens profile distortion and perspective distortion, extracts texture features such as grayscale co-occurrence matrix, gradient histogram and local binary pattern, clusters and screens the texture features, removes the body and invalid areas, and removes low-quality texture features according to the texture quality evaluation index to obtain the top view of the vehicle. By using texture information to judge the regional validity and image quality, the body occlusion and image distortion areas can be adaptively filtered out, and high-quality environmental textures can be extracted as input, thereby improving the robustness of the system.

[0016] In combination with some embodiments of the first aspect, in some embodiments, after the step of generating a transparent image of the bottom of the vehicle based on the coordinates of the vehicle body positioning points and the vehicle top view, the method also includes: obtaining a depth map corresponding to the vehicle top view, and dividing the pixel points in the vehicle top view into ground points and non-ground points based on the depth information in the depth map; in a three-dimensional coordinate system, based on the coordinates of the vehicle body positioning points and the depth information, mapping the ground points in the area near the vehicle body in the vehicle top view to a virtual vehicle bottom plane to generate a transparent texture map of the bottom of the vehicle; combining the transparent texture map of the bottom of the vehicle with the vehicle bottom model to render and generate a three-dimensional transparent chassis image; and integrating the three-dimensional transparent chassis image with the three-dimensional model of the vehicle body and the three-dimensional model of the environment to obtain a three-dimensional transparent panoramic image of the bottom of the vehicle including the driving conditions.

[0017] In the above embodiment, the transparent underbody registration system maps the ground points near the vehicle body to the virtual underbody plane to generate a transparent texture map, and then combines it with the underbody model to render a three-dimensional transparent underbody image, which is then merged with the vehicle body and the three-dimensional model of the environment to obtain a three-dimensional transparent underbody panoramic image including the driving road conditions. The real three-dimensional environment is reconstructed using depth information, and combined with the virtual underbody model to generate an immersive transparent underbody panoramic image, so that users can intuitively see the three-dimensional spatial relationship between the underbody components and the road surface, and better perceive the interactive state of the vehicle and the environment.

[0018] In combination with some embodiments of the first aspect, in some embodiments, the three-dimensional transparent chassis image is integrated with the three-dimensional model of the vehicle body and the three-dimensional model of the environment to obtain a three-dimensional transparent panoramic view of the vehicle bottom including the driving conditions, specifically including: obtaining multiple frames of three-dimensional models of the vehicle body at different times, extracting the motion change information of the chassis components including wheels and suspension, and obtaining vehicle body posture change data; obtaining multiple frames of three-dimensional models of the environment at different times, determining the elevation and slope changes of the road surface, and obtaining driving condition risk data; integrating the vehicle body posture change data and the driving condition risk data into the three-dimensional transparent chassis image, and rendering and drawing dangerous area warning signs; correcting the relative position relationship between the chassis components and the vehicle body and the environment according to the vehicle speed information and the turning angle information, and obtaining a three-dimensional transparent panoramic view of the vehicle bottom.

[0019] In the above embodiment, the transparent underbody registration system obtains multiple frames of body models at different times, extracts the movement changes of chassis components such as wheels and suspension, obtains body posture change data, and simultaneously obtains multiple frames of environmental models, determines the elevation and slope changes of the road surface, obtains driving road condition risk data, integrates the body posture and road condition risk into the transparent underbody, renders and draws warning signs for dangerous areas, and corrects the relative position between the chassis components and the body environment according to the vehicle speed and turning angle to obtain the final three-dimensional transparent underbody panoramic image. The information dimension of the transparent underbody image is enriched, and the data such as body posture and road condition risk are visualized to indicate potential driving dangers. At the same time, the local details are dynamically corrected with the vehicle speed and turning angle, making the image more realistic and improving driving safety and driving experience.

[0020] In a second aspect, an embodiment of the present application provides a vehicle bottom transparent registration system, the vehicle bottom transparent registration system comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, the one or more processors call the computer instructions to enable the vehicle bottom transparent registration system to perform the method described in the first aspect and any possible implementation method of the first aspect.

[0021] In a third aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when the computer program product is run on a vehicle bottom transparent registration system, enables the vehicle bottom transparent registration system to perform the method described in the first aspect and any possible implementation method of the first aspect.

[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, comprising instructions. When the instructions are executed on a vehicle bottom transparent registration system, the vehicle bottom transparent registration system executes the method described in the first aspect and any possible implementation method of the first aspect.

[0023] It can be understood that the vehicle bottom transparent registration system provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the method provided in the embodiment of the present application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, which will not be repeated here.

[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0025] 1. The method of inputting two adjacent frames of vehicle top view into the pre-trained neural network to obtain the pixel-level displacement map is adopted, and the ground reference point cloud is determined by combining the vehicle speed, turning angle and other information, and the displacement estimation result based on vision or sensor is selected according to the point cloud confidence. Finally, the transparent image of the bottom of the vehicle is generated through multi-frame cumulative positioning. Therefore, the visual information is effectively used to improve the positioning accuracy, and the system robustness is improved by switching the positioning method through confidence evaluation, which avoids the image misalignment problem caused by speed error, thereby achieving the improvement of the accuracy of the registration and real-time display of the transparent image of the bottom of the vehicle.

[0026] 2. The method of using an on-board panoramic camera to obtain a full-view front view, and performing fisheye correction and perspective transformation to obtain a top view of the vehicle effectively expands the field of view of environmental perception, obtains more complete road information, and eliminates lens distortion and perspective distortion through image preprocessing, thereby achieving the generation of a high-quality top view of the vehicle. This top view serves as the input for vehicle positioning and chassis registration, which improves the perception range and imaging quality of the entire system, reduces the registration error caused by insufficient field of view and image distortion, and ensures system performance.

[0027] 3. Due to the use of depth images to segment ground points and non-ground points, combined with the vehicle bottom model to render a three-dimensional transparent chassis image, and generate a three-dimensional transparent vehicle bottom panoramic view, the real three-dimensional environment is reconstructed using depth information, and the vehicle bottom texture consistent with the real environment is generated through virtual vehicle bottom model mapping. The three-dimensional display of the vehicle and the environment is realized through three-dimensional fusion, so that the transparent image of the vehicle bottom is extended from the two-dimensional plane to the three-dimensional space. Users can observe the spatial relationship between the vehicle bottom parts and the road surface from any angle, and intuitively perceive the dynamic interaction process between the vehicle driving status and the environment, which improves driving safety and control experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flow chart of the transparent registration method for the bottom of a vehicle in an embodiment of the present application;

[0029] Figure 2 is another flow chart of the transparent registration method for the bottom of a vehicle in an embodiment of the present application;

[0030] Figure 3 It is a schematic diagram of the structure of a physical device of the vehicle bottom transparent registration system in the embodiment of the present application. DETAILED DESCRIPTION

[0031] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be used as limitations to the present application. As used in the specification of the present application, the singular expressions "one", "a kind of", "above", "the" and "this" are intended to also include plural expressions, unless there is a clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations comprising one or more of the listed items.

[0032] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, "plurality" means two or more.

[0033] For ease of understanding, the application scenarios of the embodiments of the present application are introduced below.

[0034] User A, an automotive engineer, is responsible for developing a new SUV model. In order to improve the safety and driving experience of the vehicle, User A hopes to apply the bottom transparent imaging technology to this car. With this technology, the driver can clearly see the real-time status of the bottom components and the road surface, enhancing the perception of the vehicle's driving environment. However, to achieve high-quality bottom transparent imaging, it is necessary to solve a series of technical problems such as image acquisition, stitching, and registration. In particular, during the driving process of the vehicle, due to factors such as bumps and steering, the bottom image of the vehicle will be misaligned and blurred, which seriously affects the user experience. How to improve the accuracy and real-time performance of the bottom image registration has become the main technical challenge faced by User A.

[0035] In the related art, a displacement estimation method based on a kinematic model can be used and the estimated displacement can be combined with an image stitching and registration algorithm to achieve transparent imaging of the underbody. The following describes a scenario in which the transparent underbody registration method in the related art is used.

[0036] User A first tried to use the traditional transparent underbody imaging solution. This solution uses the CAN bus to obtain parameters such as vehicle speed and steering, estimates the vehicle displacement based on the kinematic model, and then combines the displacement with the image registration algorithm to generate a transparent underbody view. User A deployed the system on an experimental vehicle and conducted actual road tests. However, he found that under complex working conditions such as vehicle bumps and sharp turns, the speed and corner signals often have large errors and delays, resulting in significant offset and distortion in the generated underbody image. Although user A tried to optimize the algorithm, it was still unable to obtain satisfactory underbody imaging effects due to the limitations of sensor accuracy and vehicle dynamics models.

[0037] The transparent underbody registration method in the embodiment of the present application is used to extract image displacement by introducing a pre-trained neural network, combined with technical means such as point cloud segmentation and three-dimensional mapping, to achieve high-precision underbody image generation, which can reduce the impact of harsh environments and extreme working conditions on the accuracy of sensor signals and improve imaging frame rate and smoothness. The following introduces the scenarios in which the transparent underbody registration method in the present application is used.

[0038] In order to overcome the above technical difficulties, user A adopted the transparent underbody imaging solution proposed in this application. This solution uses a vehicle-mounted panoramic camera to obtain real-time images of the vehicle's surrounding environment, forms a bird's-eye view through image stitching, and then applies a pre-trained neural network model to extract the displacement between images, and combines point cloud segmentation, three-dimensional mapping and other technologies to generate high-quality transparent images of the underbody. User A replaced the original system with this solution and carried out a large number of real-vehicle verifications. The results show that under various complex road conditions, the underbody images generated by the new solution are highly consistent with the actual road surface, the positional relationship between the body parts and the environment is accurate and stable, the image is coherent and clear, and the delay is greatly reduced.

[0039] It can be seen that the underbody transparent registration method in the embodiment of the present application can not only achieve stable and reliable transparent visual presentation, but also effectively solve the inherent defects of traditional solutions that are limited by the vehicle kinematic model and the accuracy of a single sensor, thereby realizing highly robust and low-latency underbody environment perception and auxiliary decision-making functions, thereby improving driving experience and driving safety.

[0040] For ease of understanding, the following describes the process of the method provided by this implementation in combination with the above scenario. Figure 1 , which is a flow chart of the transparent underbody registration method in the embodiment of the present application.

[0041] S101. Input two adjacent frames of the vehicle top view of the test vehicle into a pre-trained neural network model to obtain a vehicle displacement map.

[0042] Among them, the test vehicle represents the target vehicle that needs to be transparently imaged under the vehicle. The top view of the vehicle refers to the top view image of the vehicle's surrounding environment obtained by taking a panoramic camera on the vehicle, which is obtained by image stitching and perspective transformation. Two adjacent frames of images refer to two top views of the vehicle collected continuously in time. The pre-trained neural network model refers to a deep learning model that uses a large amount of image data for offline training to learn image displacement features, and is used to directly regress pixel-level displacement vectors from input image pairs. The vehicle displacement map represents the distribution map of pixel point displacement between two frames of images predicted by the neural network model, reflecting the relative movement of the vehicle within the time interval. The displacement amount represents the distance each pixel moves in the horizontal (X) and vertical (Y) directions.

[0043] Specifically, the transparent underbody registration system first selects two adjacent top-view frames as the input of the neural network model, and generates the corresponding displacement map through forward propagation calculation. The neural network model adopts a specific architecture and is pre-trained on massive synthetic data and real scene data, so that it has the ability to accurately estimate pixel displacement vectors from image pairs. The generated displacement map has the same size as the original image, but the value of each pixel point is changed from the original RGB color value to the horizontal and vertical displacement, representing the motion displacement of the point between the two frames.

[0044] It should be noted that the neural network model here refers to the displacement estimation model; this neural network model is used to estimate the pixel-level displacement vector from two adjacent frames of the vehicle's top view and generate a vehicle displacement map. The model uses a specific network architecture and is pre-trained on massive synthetic data and real scene data. When training the model, the input is a pair of images, and the label is the true value of the pixel-level displacement vector. The model performs end-to-end learning by minimizing the error between the predicted displacement and the true value. When in use, the input of the model is two adjacent frames of the vehicle's top view, and the output is a displacement map representing the displacement of each pixel in the previous and next frames in the X and Y directions, which has the same resolution as the original image.

[0045] In some embodiments, the image can be input into a neural network model and a displacement map can be generated in a variety of ways: Optionally, before the image is input into the model, the original image is preprocessed by dedistortion, smoothing filtering, contrast enhancement, etc. to improve the accuracy of feature extraction; the preprocessed image is then input into an optical flow model, such as FlowNetS, and the encoder extracts high-dimensional features, and the decoder upsamples layer by layer and estimates the optical flow displacement; finally, the optical flow output is converted into a displacement map consistent with the original resolution.

[0046] S102: Determine a ground reference point cloud according to the vehicle speed information, the turning angle information, the vehicle displacement map and the vehicle top view of the test vehicle.

[0047] The vehicle speed information refers to the vehicle speed data measured by the wheel speed sensor. The steering angle information refers to the vehicle steering angle data collected by the steering wheel angle sensor, which is used to characterize the course change during the vehicle movement. The ground reference point cloud refers to a set of feature points belonging to the flat road surface extracted from the vehicle's top view through a certain strategy, and each feature point is composed of its three-dimensional spatial coordinates.

[0048] Specifically, the system first combines the vehicle speed and angle data with the pre-calibrated vehicle kinematic model to calculate the approximate posture transformation matrix of the vehicle body between two adjacent frames. The matrix is ​​then applied to the vehicle displacement map to compensate for the image displacement and filter out the parallax effect caused by the vehicle body's own motion. Next, the system performs semantic segmentation on the compensated displacement map, extracts the pixel mask belonging to the ground area, overlays the mask on the original top view, and obtains a binary image through grayscale threshold segmentation, thereby extracting ground feature points. Finally, these feature points are back-projected into three-dimensional space using binocular vision or multi-view geometry reconstruction algorithms to obtain an ordered ground point cloud with the vehicle body coordinate system as a reference, namely, a ground reference point cloud.

[0049] The semantic segmentation of the displacement map requires a deep image semantic segmentation model, which is used to distinguish ground points from non-ground points in the depth map corresponding to the top view of the vehicle. The model uses typical semantic segmentation network architectures such as PSPNet and DeepLab. During offline training, a large number of depth maps and pixel-level annotations (such as roads, car bodies, buildings, etc.) are input, and end-to-end learning is performed by calculating cross-entropy loss. When in use, the model takes the depth map of the top view of the vehicle as input and directly outputs the category of each pixel (binarized into ground and non-ground). The segmentation results are post-processed by morphological filtering, connected domain analysis, etc. to obtain a fine ground area mask.

[0050] In some embodiments, the extraction of ground reference point cloud can be achieved in a variety of ways: Optionally, when semantically segmenting the displacement map, a deep learning-based image segmentation model such as PSPNet can be used, only two categories of ground and non-ground can be marked in the training samples, and a priori constraints are added to the loss function to improve the segmentation accuracy; after the segmentation results are post-processed by smoothing, connected domain analysis, etc., the maximum connected area is extracted as the ground mask. After applying the mask to the top view and binarizing it, the sub-pixel corner detection algorithm is used to extract the feature points, and the three-dimensional coordinates of the feature points are solved in combination with the principle of binocular vision epipolar geometry. Optionally, when obtaining the compensated and corrected displacement map, the original displacement map can be first gridded, and traditional feature points such as SIFT and ORB can be extracted in each grid. The rigid body transformation matrix between grids is estimated by matching feature points and using the RANSAC algorithm, all transformation matrices are uniformly optimized to align to the global coordinate system, and finally the transformed feature points are back-projected into a three-dimensional point cloud. It is understandable that the extraction of ground point clouds can also be directly perceived by three-dimensional sensors such as lidar and ToF cameras, and point cloud segmentation can be performed through methods such as Bayesian probability filtering and surface fitting, which are not limited here.

[0051] S103 . Determine the confidence of the ground reference point cloud relative to the vehicle body position according to the coordinate distribution of multiple points in the ground reference point cloud.

[0052] Among them, the confidence of the ground reference point cloud relative to the vehicle body position refers to the credibility of the overall pose estimation of the point cloud, reflecting the accuracy of point cloud matching and positioning. The coordinate distribution characteristics of the point cloud generally include point density, covariance, normal vector and other attributes that reflect the geometric shape of the point set.

[0053] Specifically, the system first calculates the spatial distribution characteristics of the point cloud as a whole, such as the point density histogram, the covariance matrix in the three directions of xyz, the normal vector angle histogram, etc. At the same time, the system obtains the current vehicle speed, acceleration, steering angular velocity and other motion state information from sensors such as the on-board IMU and wheel speed meter. According to the amplitude of the motion state, the system divides the vehicle motion into different working conditions such as straight lines, smooth turns, sharp turns, acceleration and deceleration, corresponding to different levels of motion risk. In addition, the system also needs to obtain environmental information such as road material (asphalt, cement, sand and gravel, etc.), weather conditions (sunny, rainy, foggy, etc.), and light intensity. Finally, the point cloud distribution characteristics, vehicle motion risks, and environmental factors are input into the pre-trained confidence assessment model, and the confidence value representing the accuracy of the point cloud pose estimation is obtained through model reasoning.

[0054] This confidence assessment model is used to evaluate the positioning reliability of the ground reference point cloud. The input of the model includes the spatial distribution characteristics of the point cloud (such as point density histogram, covariance matrix eigenvalue ratio, etc.), vehicle motion state (such as vehicle speed, steering level), environmental factors (such as road material, weather conditions), etc., and the output is a confidence value that characterizes the accuracy of the overall pose estimation of the point cloud. During the training phase, the model uses a large number of samples with manually annotated confidence levels for supervised learning, and predicts confidence levels in the form of regression or classification. When in use, the model evaluates the confidence of the current ground point cloud in real time. If it is higher than the threshold, the visual positioning is considered reliable, otherwise it switches to other positioning methods such as odometers.

[0055] S104: When the confidence level is higher than a preset credible threshold, the displacement result data of the test vehicle in the X and Y directions are determined based on the ground reference point cloud.

[0056] The preset credible threshold is a confidence threshold set artificially based on factors such as positioning error tolerance and safety margin. A value above this threshold means that the positioning accuracy of the ground point cloud can meet the requirements of subsequent use. The displacement result data in the X and Y directions represents the relative displacement of the vehicle in the horizontal plane estimated by point cloud registration, which is a two-dimensional displacement vector.

[0057] Specifically, the system uses the current frame point cloud as the source and the previous frame point cloud as the target, and uses point cloud registration algorithms such as ICP and NDT to estimate the rigid body transformation of the two frames of point clouds. The registration process minimizes the distance metric between the two point sets, such as point-to-point distance, point-to-surface distance, etc. After multiple rounds of iterative optimization, the rotation matrix R and translation vector t that transform the source point cloud to the target point cloud coordinate system are obtained. Projecting R and t in the X-axis and Y-axis directions of the vehicle body coordinate system respectively, the displacement of the vehicle in these two directions can be obtained. The system uses this displacement as the displacement estimation result of the vehicle in this time period for subsequent bottom vehicle image generation.

[0058] In some embodiments, the vehicle displacement can be estimated using high-confidence point clouds in a variety of ways: Optionally, before performing point cloud registration, the source point cloud and the target point cloud are first filtered and downsampled to remove outliers and improve the consistency of the point cloud density; then an octree or KD tree index structure of the point cloud is constructed to facilitate the nearest neighbor search. During the registration process, a point-to-surface ICP algorithm is used to solve the rotation matrix and translation vector based on the least squares method, and the Levenberg-Marquart algorithm is used for nonlinear iterative optimization; at the same time, a weight decay mechanism and outlier removal are introduced to improve the robustness of the registration. After obtaining the transformation matrix, the Rodriguez formula is used to convert the rotation matrix into Euler angles, and the X and Y direction components are extracted from it as displacement estimates. Optionally, the histogram-based NDT registration algorithm can be directly used to divide the target point cloud into regular grids, count the distribution density of points in each grid, and maximize the normalized mutual information of the density distribution of the source point cloud and the target point cloud, so as to obtain the optimal rigid body transformation parameters; finally, the translation amounts in the X and Y directions can be separated from the transformation matrix.

[0059] S105. When the confidence level is not higher than the preset credible threshold, the displacement data in the speed sensor is used as the displacement result data of the test vehicle in the X and Y directions.

[0060] The displacement data of the speed sensor refers to the distance traveled by the vehicle during this period of time calculated using the wheel speed, which can be further decomposed into displacement components in the X and Y directions.

[0061] When the confidence assessment model determines that the confidence of the ground point cloud is low, it means that pure visual positioning may have a large deviation, and the system will switch to the odometer displacement estimation method based on the speed sensor. Specifically, the system obtains the speed pulse signals of the left and right wheels from the CAN bus, calculates the vehicle turning radius by the wheel speed difference, and converts the distance traveled by the vehicle during this time period in combination with the tire rolling radius. According to the heading angle between the vehicle body coordinate system and the earth coordinate system, the trigonometric function is used to decompose the distance into displacement components in the X and Y directions, which are used as odometer measurements to replace visual positioning for subsequent algorithms. The system performs weighted fusion of visual displacement and odometer displacement based on the confidence level to obtain the final displacement estimate.

[0062] In some embodiments, the vehicle displacement can be estimated using speed sensor data in a variety of ways: Optionally, when calculating the wheel speed, the speed pulse signals of the left and right wheels are first filtered and smoothed to remove obviously abnormal sampling points; then the pulse value is converted into an angular velocity value according to the pulse number-speed mapping relationship in the calibration data. The angular velocity difference of the wheels on both sides is used to estimate the turning radius of the vehicle in combination with the wheelbase and small angle assumptions. Assuming that the tire is rigid and does not slip laterally, the mileage is obtained by integrating the tire linear velocity. Finally, the current heading angle is aligned with the ground point cloud coordinate direction, and the converted distance is decomposed on the X and Y axes to obtain the forward and lateral displacement components respectively. Optionally, considering that the measurement error of the speed sensor will gradually increase over time, algorithms such as Kalman filtering can be used to correct the displacement estimate. The wheel speed, heading angle, and vehicle kinematic model are used as the measurement input of the Kalman filter, and the system state vector is recursively updated using the state equation and observation equation, and the confidence weights of the measured value and the predicted value are adjusted according to the state covariance, and the optimized displacement is finally output. It is understandable that odometer and visual positioning each have their own advantages and disadvantages, and the displacement estimates of the two methods can also be complementary fused, such as using particle filtering, information fusion criteria, etc., to further improve the continuity and stability of positioning, which is not limited here.

[0063] S106 , determining a plurality of corresponding displacement result data of the test vehicle according to the plurality of vehicle displacement maps, and obtaining the coordinates of the vehicle body positioning points.

[0064] Among them, multiple vehicle displacement maps refer to a series of vehicle displacement images generated continuously over a period of time, reflecting the entire process of vehicle movement. Multiple corresponding displacement result data represent vehicle displacement estimates that match these displacement maps, which can be derived from visual positioning or odometer measurement. The vehicle body positioning point coordinates refer to the position coordinates of the vehicle in the global map or its own coordinate system, which are obtained by accumulating multiple displacement values.

[0065] After obtaining a single displacement estimate, the system needs to perform cross-frame accumulation and pose fusion to generate a complete vehicle bottom track. Specifically, the system maintains a pose sequence to store the position and orientation of the vehicle at each moment. After a new frame of displacement map is generated, the system first determines its confidence, selects a high-confidence visual displacement or a low-confidence odometer displacement, and adds it to the pose of the previous moment to obtain the initial pose of the vehicle at the current moment. Then, based on feature matching between images, the system estimates the relative motion between the current frame and the previous frame, and optimizes the initial pose using methods such as ICP and image optimization adjustment to eliminate accumulated errors. After multiple iterations, the pose sequence converges to a globally consistent state, at which point the coordinate value of each item is the positioning result of the vehicle body at that moment. The translation component in the pose sequence is extracted to obtain the XY coordinate sequence of the vehicle body positioning point.

[0066] S107: Generate a transparent image of the bottom of the vehicle according to the coordinates of the vehicle body positioning points and the top view of the vehicle.

[0067] Among them, the transparent underbody image is a top-down image that visualizes the environment under the vehicle body and surrounding road conditions. The transparent effect makes the part originally blocked by the vehicle body visible.

[0068] Specifically, the system first spatially registers the top view of the historical vehicle according to the position of the vehicle body, and aligns the image position in a unified global coordinate system. Then, the system sets a virtual imaging plane parallel to the chassis of the vehicle body, projects the registered top view onto the plane frame by frame, and crops it according to the current position of the vehicle body to obtain a series of projection images of the bottom of the vehicle. Then, the system combines these images frame by frame in a splicing and fusion manner to obtain a panoramic top view covering the entire driving path. Finally, the three-dimensional model of the vehicle body is rendered into the spliced ​​image, the transparency in the material properties is adjusted, and the road surface area is set to transparent and the non-road surface area is set to translucent to obtain the final transparent image of the bottom of the vehicle. The vehicle body is partially blurred in this image, and the road surface conditions under the vehicle can be seen through, intuitively presenting the positional relationship between the vehicle and the environment.

[0069] In some embodiments, the generation of a transparent image under the vehicle can be achieved in a variety of ways: Optionally, in order to overcome the image jitter caused by the bumps of the vehicle, a global motion model can be introduced when stitching images, and feature points can be tracked through multi-layer pyramid optical flow, and the affine transformation relationship between adjacent frames can be estimated to perform pixel-based motion compensation. Then, the SIFT feature is used for image registration, duplicate areas are removed, the best stitching path is found by triangular meshing, and a multi-band fusion algorithm is used to smoothly transition the stitching boundary. When the top view is projected onto a virtual plane, the imaging transformation is represented by a homography matrix model, and the lens distortion coefficient is considered. Optionally, in order to achieve real-time rendering, the three-dimensional model can be simplified into data structures such as voxels or octrees, and ray casting is performed on the GPU. The frame buffer object is used to store color, depth and template information at different levels, and multiple views are rendered in parallel. For pixels detected as road surfaces, the corresponding color values ​​are directly rendered to the screen; for non-road objects, a physically based rendering pipeline is used to calculate the effects of lighting, reflection, refraction, etc. before Alpha blending. It is understandable that in order to enrich the information of the underbody image, the self-positioning map and multi-sensor perception results can also be integrated, and virtual information such as tire pressure and suspension load can be superimposed in real time to assist the driver in fully controlling the vehicle status. This is not limited here.

[0070] The following is a more detailed description of the process of the method provided by this implementation. Figure 2 , is another flow chart of the transparent underbody registration method in the embodiment of the present application.

[0071] S201. Obtain a full-view front view of the test vehicle through a vehicle-mounted panoramic camera.

[0072] The test vehicle refers to the vehicle model to be tested that is equipped with a transparent imaging system under the vehicle. The on-board panoramic camera is usually composed of multiple fisheye wide-angle cameras, which can achieve a 360-degree panoramic view without blind spots. The full-view front view refers to the image information within the 180-degree range directly in front of the vehicle body.

[0073] After the underbody transparent registration system is started, the first thing to do is to obtain the original image information around the vehicle. Specifically, the system controls the acquisition trigger and exposure parameters of the forward surround camera through the CAN bus, and transmits the real-time image data stream back from the camera. The system groups the images of multiple cameras according to the installation position, and extracts the original image frames covering the 180-degree field of view in front of the vehicle. Due to the use of fisheye lens imaging, these original images have obvious field of view distortion and need to be corrected. At the same time, in order to save data bandwidth, the resolution and frame rate of the original image are usually low, and interpolation may be required. The system continuously executes the above image acquisition steps in a loop to provide data input for subsequent processing.

[0074] In some embodiments, the acquisition of the full-view front view can be achieved in a variety of ways: Optionally, four or six 200-degree wide-angle cameras are used to transmit raw image data through the Ethernet bus. Each camera has a built-in ISP chip, which can complete pre-processing such as image distortion correction and color calibration. The system calls the camera to collect images through the industrial camera SDK function and sets parameters such as trigger mode, frame rate, and resolution. Three cameras are arranged in a 120-degree range in the front of the vehicle body, and the vertical field of view covers the ground to a height of 2 meters in front. The exposure time is ensured to be consistent through hardware synchronization. After obtaining the synchronized three-way images, the pinhole camera is calibrated, and the images are stitched using SURF features to generate a 360-degree panoramic view. Optionally, an integrated panoramic camera is used to integrate multiple wide-angle lenses in the same housing, and the optical axes are distributed in a circular array on a plane. By calling the panoramic stitching interface of the camera, a seamless panorama stitched by hardware is directly output. At this time, the system only needs to set the network parameters and working mode of the camera, and obtain the real-time panoramic video stream through the RTSP protocol. The video frame rate can reach 30 frames per second and the resolution can reach 4K, which can meet the needs of subsequent processing. It is understandable that in order to further expand the perception range, an additional 360-degree panoramic camera can be installed on the roof to form a complementary upper and lower field of view with the front-view camera to achieve complete all-round stereoscopic surround vision, which is not limited here.

[0075] S202: Perform fisheye correction and perspective transformation on the full-view front view to obtain a top view of the vehicle.

[0076] Among them, fisheye correction refers to correcting the barrel distortion of fisheye lens imaging and restoring the original shape and position relationship of objects in the image. Perspective transformation maps points in the imaging plane to the target plane according to certain rules to achieve perspective conversion.

[0077] Specifically, the system first calibrates the internal parameters and distortion coefficients of the fisheye camera, and generates a mapping matrix from the distortion map to the correction map. Then, for each pixel in the original image, its corresponding position in the mapping matrix is ​​found through bilinear interpolation to obtain the corrected pixel value, and finally an undistorted panoramic image is output. Next, the system calculates the perspective transformation matrix to convert the corrected image to a virtual top-down perspective. Assuming that the ground around the vehicle body is a rectangular area parallel to the imaging plane, the system determines the pixel coordinates of the four vertices of the rectangle in the full-view image and establishes the homography mapping relationship between the original image and the top view. Through perspective transformation, the pixels in the original image within the rectangle are compressed to the top view, while the pixels outside the rectangle are cropped. The system uses pyramid layered processing to smooth the image, and performs affine transformation fine-tuning to output the final high-quality top view.

[0078] In some embodiments, the conversion from a full-view image to a top-view image can be achieved in multiple ways: Optionally, in the offline calibration stage, the system captures a checkerboard calibration board from multiple angles, extracts the pixel coordinates of the corner points, and combines the physical size of the checkerboard to solve for the internal parameters and distortion coefficients of the fisheye camera using Zhang's checkerboard calibration method. Substitute the distortion coefficients into the spiral polynomial distortion model to generate a correction lookup table. During real-time correction, use a CUDA-based GPU acceleration library to perform table lookup and bilinear interpolation on the original image in parallel, and output an undistorted image. When generating the top-view image, determine the distance between the virtual plane and the ground based on the installation height, and calculate the homography matrix for perspective transformation. Convert the matrix into OpenGL shader language, and use the vertex and fragment processing units of the GPU to perform real-time rendering and output of the image. Optionally, adopt an end-to-end method based on deep learning to directly convert the original fisheye image into a top-view image through a convolutional neural network. Collect a large number of fisheye images and their corresponding manually annotated top-view images in different scenarios offline as training samples. Build a U-Net semantic segmentation network structure, use the original image as the input, set a loss function to measure the difference between the output and the annotation, and iteratively optimize the network parameters. During deployment, input the image to be processed into the trained model to directly obtain the corresponding top-view prediction result, eliminating complex geometric transformation calculations. It can be understood that considering the relatively high resolution of fisheye images, the above GPU-based solution may face memory shortage problems. Therefore, the original image can be downsampled first, and the high-level layers of the image pyramid can be processed during the correction and perspective transformation processes, and finally restored to the original resolution to achieve a balance between real-time performance and accuracy.

[0079] In some embodiments, the underbody transparent registration system corrects the distortion of the full-view front view to eliminate lens profile distortion and perspective distortion; extracts the texture features of the full-view front view; the texture features include gray-level co-occurrence matrix, histogram of gradients, and local binary pattern; clusters and filters the texture features to remove the vehicle body area and invalid areas, obtaining a texture feature set; according to the texture quality evaluation index of the texture feature set, removes the texture features with a texture quality lower than the preset quality threshold to obtain the vehicle top-view image.

[0080] Specifically, the system uses the pre-calibrated lens parameters and distortion coefficients to build a correction model, maps each pixel in the original image to the ideal imaging plane through a lookup table, and removes barrel distortion. Then, the system determines the homography transformation matrix between the ideal imaging plane and the ground reference plane, builds the coordinate mapping relationship from the front view to the top view based on the perspective projection principle, and reprojects the corrected image to the ground perspective. Next, the system extracts the texture features of the image, including the grayscale co-occurrence matrix reflecting the statistical characteristics of the grayscale distribution of the pixel neighborhood, the gradient histogram reflecting the directional information of the edge structure, and the local binary pattern reflecting the contrast relationship between the pixel and the neighborhood. By clustering analysis to screen the texture features, the system constructs a texture dictionary for different areas such as the vehicle body and the ground to achieve preliminary segmentation of the image scene. Furthermore, the system performs morphological filtering on the segmentation results to remove too small connected domains, and introduces prior indicators such as texture saliency and entropy to remove low-quality areas such as vehicle body occlusion and image blur. Finally, a ground top view with a wide field of view and clear details is obtained, providing high-quality input for subsequent processing.

[0081] In some other embodiments, the system can also combine the structural consistency of multiple frames of images, construct local closed loop constraints to optimize the parameters of distortion correction and perspective transformation, introduce IMU posture as an additional observation, and improve the accuracy of posture estimation. For ultra-wide-angle fisheye images, correction can be achieved by using spherical projection models, longitude and latitude mapping, and other methods. For the extraction and matching of texture features, more robust feature descriptors such as SIFT, SURF, and ORB can be used, combined with the FLANN approximate nearest neighbor search algorithm to accelerate matching. When dividing the ground and obstacle areas, fine segmentation can be performed by learning the occupancy grid map, and the segmentation results can be smoothed using global optimization methods such as graph cuts and random walks. In addition, the system can also adaptively adjust the extraction threshold and clustering parameters of texture features according to weather conditions, such as reduced image contrast in rainy and foggy weather, to cope with harsh environments.

[0082] S203: input two adjacent frames of the vehicle top view of the test vehicle into a pre-trained neural network model to obtain a vehicle displacement map.

[0083] Referring to step S101 , the vehicle bottom transparent registration system generates a vehicle displacement map.

[0084] S204: Determine a ground reference point cloud according to the vehicle speed information, the turning angle information, the vehicle displacement map and the vehicle top view of the test vehicle.

[0085] Referring to step S102 , the vehicle bottom transparent registration system determines the ground reference point cloud.

[0086] In some embodiments, the underbody transparent registration system determines the displacement error of the test vehicle in the vehicle displacement map based on the speed information and turning angle information of the test vehicle; filters out the erroneous area in the vehicle's top view based on the displacement error, and removes the vertical objects in the vehicle's top view to obtain a ground reference point cloud.

[0087] After obtaining the top view of the vehicle and the preliminary segmentation results, the transparent registration system for the bottom of the vehicle needs to further optimize the environmental perception and extract the ground point cloud. The system first introduces the vehicle motion prior, and estimates the motion displacement between adjacent frame images based on the speed and steering information obtained by the speedometer and steering wheel angle sensor. Assuming that the vehicle body is a rigid body, the system back-projects the estimated displacement to the image based on the internal parameters, and solves the basic matrix under the epipolar geometry constraint to obtain the pixel-level disparity field. Superimposing the disparity field on the original segmentation result can eliminate the area that does not conform to the law of vehicle motion and preliminarily eliminate the interference of dynamic obstacles. Next, the system counts the occupancy probability distribution of the scene from multiple frames of historical information and constructs a probability grid map. The Hough transform is performed on the area near the vehicle trajectory in the map to fit the road plane, and the vertical plane is detected based on this, that is, vertical objects such as curbs and guardrails. Finally, the system projects the grid belonging to the road plane back to the top view, combines the disparity filtering results, outputs the finely segmented ground area, and converts it into a three-dimensional point cloud to complete the scene perception.

[0088] S205 . Determine the confidence of the ground reference point cloud relative to the vehicle body position according to the coordinate distribution of multiple points in the ground reference point cloud.

[0089] Referring to step S103 , the vehicle bottom transparent registration system determines the confidence of the ground reference point cloud.

[0090] In some embodiments, the underbody transparent registration system calculates the spatial distribution density and directional consistency of multiple points in the ground reference point cloud, and uses the spatial distribution density and directional consistency as point cloud distribution features; obtains vehicle motion state information including vehicle speed information and turning angle information, and maps the vehicle motion state information to a vehicle motion risk level; constructs a confidence estimation model based on the point cloud distribution features, vehicle motion risk level, and confidence compensation coefficients corresponding to road materials and weather conditions, and outputs the confidence of the ground reference point cloud relative to the vehicle body position.

[0091] Specifically, based on the ground point cloud extracted above, the transparent registration system under the vehicle needs to further evaluate the confidence of environmental perception and judge the reliability of point cloud positioning. The system first analyzes the spatial distribution characteristics of the point cloud, and statistically calculates the point density distribution histogram of different regions to reflect the richness of the point cloud data. At the same time, the system calculates the directional consistency of the point cloud normal vector, that is, the ratio of the eigenvalues ​​of the covariance matrix, which reflects the regularity of the local shape of the point cloud. Areas with high point density and consistent normal vectors often correspond to flat roads, and the positioning is more reliable. Otherwise, they may contain more outliers. The system also obtains the motion state such as vehicle speed and steering as a priori, and uses the physical model to map it to the body posture change and uncertainty to predict the severity of point cloud deformation. Finally, the system establishes a confidence estimation model, comprehensively considers the point cloud distribution characteristics, motion state risks, road material, weather and other environmental factors, and outputs a confidence value reflecting the overall positioning reliability of the point cloud. When the confidence is higher than the threshold, it indicates that the match is accurate and the point cloud can be used to directly estimate the vehicle posture. Otherwise, alternative solutions such as more robust odometers are required.

[0092] S206: When the confidence level is higher than a preset credible threshold, determine the displacement result data of the test vehicle in the X and Y directions based on the ground reference point cloud.

[0093] Referring to step S104 , the vehicle bottom transparent registration system will determine the displacement result data based on the ground reference point cloud when the confidence level is higher than a preset credibility threshold.

[0094] S207: When the confidence level is not higher than the preset credible threshold, use the displacement data in the speed sensor as the displacement result data of the test vehicle in the X and Y directions.

[0095] Referring to step S105 , the vehicle bottom transparent registration system will determine the displacement result data based on the speed sensor when the confidence level is not higher than the preset confidence threshold.

[0096] S208. Determine a plurality of corresponding displacement result data of the test vehicle according to the plurality of vehicle displacement maps, and obtain the coordinates of the vehicle body positioning points.

[0097] Referring to step S106 , the vehicle bottom transparent registration system will determine the coordinates of the vehicle body positioning points.

[0098] S209: Generate a transparent image of the bottom of the vehicle according to the coordinates of the vehicle body positioning points and the top view of the vehicle.

[0099] Referring to step S107 , the vehicle bottom transparent registration system generates a vehicle bottom transparent image.

[0100] S210: Obtain a depth map corresponding to the top view of the vehicle, and divide the pixel points in the top view of the vehicle into ground points and non-ground points according to the depth information in the depth map.

[0101] The depth map is a grayscale image with the same resolution as the top view of the vehicle, which records the distance from each pixel in the image to the optical center of the camera. The ground point refers to the pixel representing the actual road plane, and the non-ground point corresponds to the non-road area such as the vehicle body and overhead objects.

[0102] Specifically, the system first obtains a depth map that matches the top view from a binocular camera or structured light sensor. The depth value represents the distance from the pixel to the camera imaging plane. The depth map is aligned to the top view coordinate system through rigid transformation and scale normalization. Then, the system applies joint bilateral filtering to the depth map to smooth the depth noise while maintaining edge details. The system binarizes the depth map according to a preset distance threshold, extracts the ground area mask, and divides the pixels in the corresponding top view into ground points and non-ground points. The system generates a point cloud of the ground area, removes suspended objects, and provides a geometric basis for virtual scene construction.

[0103] In some embodiments, ground point segmentation can be achieved in a variety of ways: Optionally, a depth camera is used to collect a noisy original depth map, and the transformation relationship from the depth map to the top view is established by calibrating the internal and external parameters of the camera. Considering the camera installation height, pixels whose vertical distance to the ground plane is less than a certain threshold are selected as ground points, and pixels above the threshold are classified as non-ground points. Morphological filtering is used to remove isolated misclassified pixels, and finally contour extraction and polygon fitting are performed on the ground mask area to obtain a smooth ground segmentation result. Optionally, a semantic segmentation method based on deep learning is used to collect a large number of RGB images taken by on-board cameras under different road conditions and corresponding pixel-by-pixel annotations offline, and the pixel categories include roads, car bodies, buildings, etc. Use these data to train segmentation networks such as PSPNet and DeepLab, and input the top view into the trained model in real time to directly obtain pixel-level semantic labels. Furthermore, RGB images can also be used to guide the upsampling of the depth map, restore the depth details at the edges, and obtain a refined depth map with the same resolution as the top view, thereby achieving more accurate ground segmentation. It is understandable that, considering that vehicle bumps may cause depth maps to drift, the depth maps at different times can be motion compensated and aligned through visual odometry, and historical frame time consistency constraints can be introduced to further improve the robustness of ground point segmentation, which is not limited here.

[0104] S211. In a three-dimensional coordinate system, according to the coordinates of the vehicle body positioning points and the depth information, the ground points in the area near the vehicle body in the top view of the vehicle are mapped to the virtual vehicle bottom plane to generate a vehicle bottom transparent texture map.

[0105] Among them, the three-dimensional coordinate system is usually defined as a right-hand system with the center of the vehicle's rear axle as the origin, the front direction of the vehicle as the positive x-axis, the left side as the positive y-axis, and the vertical upward as the positive z-axis. The body positioning point coordinates represent the position of the vehicle origin in the global world coordinate system. The virtual vehicle bottom plane is a rectangular plane parallel to the actual vehicle bottom in three-dimensional space, and its height is equal to the height of the depth camera installed.

[0106] Specifically, the system first establishes a local vehicle body coordinate system based on the global coordinates and heading angles of the vehicle body positioning points. In this coordinate system, the system defines a rectangular plane parallel to the ground and located at the height of the depth camera as the virtual vehicle bottom plane. Then, the system selects the ground area within a certain range around the vehicle body in the top view, extracts its vertex pixel coordinates, and assigns a depth value. The three-dimensional coordinates of these vertices are obtained by back-projection of the intrinsic parameter matrix. Next, the system uses these vertices as control points to construct a triangular mesh and generate a continuous and smooth ground module. Finally, the system adjusts the vertex position of the triangular mesh to the virtual vehicle bottom plane, and maps the top view as a texture to the mesh surface to obtain a transparent texture map of the vehicle bottom that is consistent with the real environment.

[0107] S212: Combining the vehicle bottom transparent texture map with the vehicle bottom model, and rendering to generate a three-dimensional transparent chassis image.

[0108] The vehicle bottom model refers to a three-dimensional model created based on the actual vehicle chassis structure, which contains the geometry and topology information of the chassis components. The three-dimensional transparent chassis image is a three-dimensional vehicle bottom image obtained by combining the vehicle bottom model with the virtual vehicle bottom environment and setting the transparent effect.

[0109] Based on the transparent texture map of the bottom of the vehicle generated above, the system further introduces a three-dimensional bottom model to render a realistic transparent chassis image. Specifically, the system constructs a three-dimensional model corresponding to the chassis structure of the real vehicle offline, and finely models the frame, transmission, suspension and other components. Then, the system imports the bottom model into a real-time rendering engine such as Unity3D, and sets the relative position of the model in the vehicle body coordinate system. Next, the system maps the bottom transparent texture map generated in the previous step to the virtual bottom plane as the background under the bottom model. The system sets the material properties of the bottom model and adjusts the transparency of different parts. Generally, the main structure such as the frame is set to be translucent, and the details such as the suspension and wheels are set to a higher transparency so that they can blend with the texture below. The system can also update rendering parameters such as sunlight direction and ambient lighting in real time to improve the rendering realism. Finally, the system outputs a realistic three-dimensional transparent chassis picture, and the observer can intuitively see the complete bottom structure and its spatial position relationship with the road surface.

[0110] In some embodiments, the rendering generation of a three-dimensional transparent chassis can be achieved in a variety of ways: Optionally, a PBR physical rendering pipeline is used to extract the bottom mesh from the CAD model and bake to generate normal maps, metalness maps, etc. The surface of the bottom model is subdivided, and the Bullet physics engine is used for soft body simulation to improve the realism of the movement of the suspension device. Realistic rendering effects such as subsurface scattering and volumetric lighting are implemented in the shader. The transparent texture of the bottom of the car is seamlessly integrated with the three-dimensional model through bilinear texture sampling, and the depth of field effect is superimposed to enhance the sense of space. Optionally, a real-time ray tracing method is used, and the virtual bottom plane is used as the light source to perform ray projection and path integration on the bottom model with transparent attributes, so as to physically and accurately simulate the interaction between the bottom material and the ambient light. The BVH is accelerated and traversed by ray tracing hardware such as RT Core, and the intersection of each ray and the scene is calculated in parallel, and the radiant brightness of the directional light and the ambient light is accumulated. The Russian roulette algorithm is used to sample the importance of the transparent light that is ejected multiple times, and finally the color value of each pixel is converged to achieve a realistic translucent rendering effect. It is understandable that when computing resources are limited on mobile platforms, the lighting calculation can be simplified by combining pre-calculated light maps, planar reflection models and other methods, and the rendering process can be optimized through techniques such as G-Buffer off-screen rendering and multi-Pass fusion to improve the frame rate while ensuring picture quality. This is not limited here.

[0111] S213, integrating the three-dimensional transparent chassis image with the three-dimensional model of the vehicle body and the three-dimensional model of the environment to obtain a three-dimensional transparent panoramic image of the vehicle bottom including the driving road conditions.

[0112] The vehicle body 3D model and the environment 3D model correspond to the mesh models of the vehicle and the road environment, respectively, and include their geometric structure and surface texture and other attribute information. Driving conditions refer to the road conditions determined by road material, slope, curvature and other characteristics. The three-dimensional transparent underbody panoramic image is an interactive transparent chassis image with driving condition information obtained by integrating the vehicle body and environment information rendering based on step S212.

[0113] After rendering the partial transparent chassis image, the system also needs to integrate the global body and environmental information to generate a complete panoramic observation map of the bottom of the vehicle. Specifically, the system first loads the three-dimensional model of the body that matches the vehicle model, extracts the motion parameters of components such as wheels, suspension, and steering, and updates their relative positions and postures in real time. Then, the system collects multiple frames of environmental three-dimensional reconstruction data and splices them to obtain a large-scale environmental model within a certain range around the vehicle. The system analyzes the environmental model, extracts geometric and semantic features such as road surface material, slope, and curvature, and evaluates the road risk level. In the rendering pipeline, the system aligns the transparent chassis image, body model, and environmental model in a unified coordinate system and updates the lighting and shadows. Visualize the ground road condition information, render different colors according to the risk level, and add warning signs and text descriptions. The system also receives control signals such as vehicle speed and direction, roams the three-dimensional scene in real time, and fine-tunes local details according to parameters such as wheel angle and suspension height, thereby outputting a realistic, information-rich, and interactive three-dimensional panoramic view of the bottom of the vehicle.

[0114] In some embodiments, the transparent underbody registration system obtains multiple frames of three-dimensional models of the vehicle body at different times, extracts the motion change information of the chassis components including wheels and suspension, and obtains vehicle body posture change data; obtains multiple frames of three-dimensional models of the environment at different times, determines the elevation and slope changes of the road surface, and obtains driving road condition risk data; integrates the vehicle body posture change data and the driving road condition risk data into the three-dimensional transparent chassis image, and renders and draws warning signs for dangerous areas; based on the vehicle speed information and the turning angle information, corrects the relative position relationship between the chassis components and the vehicle body and the environment, and obtains a three-dimensional transparent panoramic view of the underbody.

[0115] Specifically, after generating a partial transparent image of the bottom of the vehicle, the transparent registration system for the bottom of the vehicle needs to further integrate it with the global information to provide a more informative perception of the bottom of the vehicle environment. The system first introduces the vehicle body posture information, namely the pitch, roll, yaw angular velocity of the vehicle, etc., to reflect the movement of the vehicle body relative to the ground. By obtaining sensor information such as suspension displacement and wheel speed difference, and combining physical models such as simplified spring-mass-damper models, the system can estimate the posture changes of the vehicle body in real time. At the same time, the system obtains the surrounding environment model through three-dimensional reconstruction in the initialization stage, and updates the position of the vehicle in it in real time. Using the environmental model, the system can analyze the geometric parameters such as the slope and curvature of the road ahead, and combine the vehicle parameters to evaluate the possible risks of bumps and skidding when passing through the road. In the process of rendering the bottom of the vehicle image, the system obtains the body posture and road condition analysis results in real time, converts them into parameters such as frame tilt angle and suspension elastic deformation, and dynamically adjusts the relative positions of the components of the bottom of the vehicle model accordingly to simulate the changes in the body posture. At the same time, the system maps the road condition risk level to the color attribute of the ground material, and performs red, yellow, green and other warning coloring on the next road surface. At the end of the rendering pipeline, the system also integrates the transparent underbody texture, modulates the transparency of the stained area, and finally outputs a three-dimensional panoramic image of the underbody that combines global environmental information and driving safety tips.

[0116] In some other embodiments, the system can also combine prior information such as maps, such as refined semantic data such as curbs, speed bumps, and manhole covers, to further enhance scene understanding capabilities. Vehicle posture estimation can be achieved by a combination of multiple sensors such as inertial navigation and visual odometers, and algorithms such as nonlinear filtering and graph optimization can be introduced to improve estimation accuracy. For road risk analysis, vehicle physical parameters such as center of gravity height and suspension stiffness can be included, and a virtual test environment can be built using simulation tools such as CarSim, and offline iterative optimization evaluation strategies can be implemented. For transparent chassis rendering, complex material models such as BRDF and high-fidelity rendering algorithms such as ray tracing can be introduced, while in-depth exploration of the application of XR devices in the vehicle environment provides an immersive underbody surround experience. Of course, these expanded applications place higher requirements on hardware platforms such as GPUs and optical perspective, and related perception and decision-making algorithms also face automotive-grade design challenges, which require collaborative efforts from industry, academia, research, and application.

[0117] In the embodiment of the present application, due to the use of multi-faceted fusion strategies such as visual positioning, point cloud segmentation, and three-dimensional mapping, combined with confidence assessment and adaptive weight registration and other processing methods, the impact of harsh environments on sensor accuracy is effectively reduced, the stability and robustness of vehicle bottom image generation are improved, and the system response speed and computing efficiency are improved. The environmental semantic information is integrated with the driving status to achieve reliable, real-time, and immersive vehicle bottom environment perception and human-computer interaction experience.

[0118] The following describes the transparent underbody registration system in the embodiment of the present invention from the perspective of hardware processing. Figure 3 , which is a schematic diagram of the structure of a physical device of the transparent underbody registration system in an embodiment of the present application.

[0119] It should be noted that Figure 3 The structure of the vehicle bottom transparent registration system shown is only an example and shall not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0120] like Figure 3 As shown, the vehicle bottom transparent registration system includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage part 308 to the random access memory (RAM) 303, such as executing the method described in the above embodiment. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, the ROM 302 and the RAM 303 are connected to each other through the bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.

[0121] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, a button switch, etc.; an output section 307 including a liquid crystal display (LCD) and an audio output device, an indicator light, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read therefrom is installed into the storage section 308 as needed.

[0122] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 309, and / or installed from a removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, various functions defined in the present invention are performed.

[0123] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, apparatus, or device.

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings.

[0125] Specifically, the vehicle bottom transparent registration system of this embodiment includes a processor and a memory, and a computer program is stored in the memory. When the computer program is executed by the processor, the vehicle bottom transparent registration method provided in the above embodiment is implemented.

[0126] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the vehicle bottom transparent registration system described in the above embodiment; or may exist independently without being assembled into the vehicle bottom transparent registration system. The above storage medium carries one or more computer programs, and when the above one or more computer programs are executed by a processor of the vehicle bottom transparent registration system, the vehicle bottom transparent registration system implements the vehicle bottom transparent registration method provided in the above embodiment.

[0127] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0128] As used in the above embodiments, the term "when..." may be interpreted to mean "if..." or "after..." or "in response to determining..." or "in response to detecting...", depending on the context. Similarly, the phrases "upon determining..." or "if (the stated condition or event) is detected" may be interpreted to mean "if determining..." or "in response to determining..." or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)", depending on the context.

[0129] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiments, the processes can be completed by computer programs to instruct related hardware, and the programs can be stored in computer-readable storage media. When the programs are executed, they can include the processes of the above-mentioned method embodiments. The aforementioned storage media include: ROM or random access memory RAM, magnetic disk or optical disk and other media that can store program codes.

Claims

1. A transparent registration method for vehicle bottom, characterized in that: Applied to the bottom transparent registration system of a vehicle, the method comprises: Input two adjacent frames of the vehicle top view of the test vehicle into a pre-trained neural network model to obtain a vehicle displacement map; the vehicle displacement map represents the displacement of each pixel in the two preceding and succeeding frames of the image in the X and Y directions; Determine a ground reference point cloud according to the speed information, the turning angle information, the vehicle displacement map and the vehicle top view of the test vehicle; Determining the confidence of the ground reference point cloud relative to the vehicle body position according to the coordinate distribution of multiple points in the ground reference point cloud; When the confidence level is higher than a preset credible threshold, determining displacement result data of the test vehicle in the X and Y directions based on the ground reference point cloud, so as to calculate the vehicle displacement according to visual positioning; When the confidence level is not higher than a preset credible threshold, using the displacement data in the speed sensor as the displacement result data of the test vehicle in the X and Y directions to calculate the vehicle displacement according to the odometer measurement; According to the confidence of the ground reference point cloud corresponding to the multiple vehicle displacement maps, the visual positioning or the odometer measurement is selected, and the displacement result data corresponding to the multiple different selections of the test vehicle are determined to obtain the coordinates of the vehicle body positioning point; A transparent image of the vehicle bottom is generated according to the vehicle body positioning point coordinates and the vehicle top view.

2. The method according to claim 1, characterized in that Determining the ground reference point cloud according to the speed information, the turning angle information, the vehicle displacement map and the vehicle top view of the test vehicle specifically includes: Determining a displacement error of the test vehicle in the vehicle displacement map according to the vehicle speed information and the turning angle information of the test vehicle; The erroneous area in the top view of the vehicle is filtered out according to the displacement error, and the vertical objects in the top view of the vehicle are removed to obtain a ground reference point cloud.

3. The method according to claim 1, characterized in that Determining the confidence of the ground reference point cloud relative to the vehicle body position according to the coordinate distribution of multiple points in the ground reference point cloud specifically includes: Calculating the spatial distribution density and directional consistency of a plurality of points in the ground reference point cloud, and using the spatial distribution density and the directional consistency as point cloud distribution features; Acquiring vehicle motion state information including vehicle speed information and turning angle information, and mapping the vehicle motion state information into a vehicle motion risk level; According to the point cloud distribution characteristics, the vehicle motion risk level, and the confidence compensation coefficient corresponding to the road material and weather conditions, a confidence estimation model is constructed to output the confidence of the ground reference point cloud relative to the vehicle body position.

4. The method according to claim 1, characterized in that: Before the step of inputting two adjacent frames of the vehicle top view of the test vehicle into the pre-trained neural network model to obtain the vehicle displacement map, the method further includes: Obtain a full-view front view of the test vehicle through the on-board panoramic camera; The full-view front view is subjected to fisheye correction and perspective transformation to obtain a top view of the vehicle.

5. The method according to claim 4, characterized in that The performing fisheye correction and perspective transformation on the full-view front view to obtain a top view of the vehicle specifically includes: Performing distortion correction on the full-view front view to eliminate lens profile distortion and perspective distortion; Extracting texture features of the full-view front view; the texture features include gray-level co-occurrence matrix, gradient histogram and local binary pattern; Clustering and screening the texture features, removing the vehicle body area and invalid area, and obtaining a texture feature set; According to the texture quality evaluation index of the texture feature set, texture features whose texture quality is lower than a preset quality threshold are removed to obtain a top view of the vehicle.

6. The method according to claim 1, characterized in that After the step of generating a transparent image of the bottom of the vehicle according to the coordinates of the vehicle body positioning points and the vehicle top view, the method further includes: Acquire a depth map corresponding to the top view of the vehicle, and divide pixel points in the top view of the vehicle into ground points and non-ground points according to depth information in the depth map; In a three-dimensional coordinate system, according to the coordinates of the vehicle body positioning points and the depth information, the ground points in the area near the vehicle body in the top view of the vehicle are mapped to a virtual vehicle bottom plane to generate a vehicle bottom transparent texture map; Combining the vehicle bottom transparent texture map with the vehicle bottom model to render and generate a three-dimensional transparent chassis image; The three-dimensional transparent chassis image is integrated with the three-dimensional model of the vehicle body and the three-dimensional model of the environment to obtain a three-dimensional transparent panoramic image of the vehicle bottom including the driving road conditions.

7. The method according to claim 6, characterized in that The three-dimensional transparent chassis image is integrated with the three-dimensional vehicle body model and the three-dimensional environment model to obtain a three-dimensional transparent vehicle bottom panoramic image including the driving road conditions, specifically including: Acquire multiple frames of the vehicle body three-dimensional model at different times, extract the motion change information of the chassis components including wheels and suspension, and obtain the vehicle body posture change data; Obtain multiple frames of the three-dimensional model of the environment at different times, determine the changes in road elevation and slope, and obtain driving road condition risk data; Integrate the vehicle body posture change data and the driving road condition risk data into the three-dimensional transparent chassis image, and render and draw a warning sign of a dangerous area; According to the vehicle speed information and the turning angle information, the relative position relationship between the chassis component and the vehicle body and the environment is corrected to obtain a three-dimensional transparent panoramic view of the vehicle bottom.

8. A transparent registration system for vehicle bottom, characterized in that: The vehicle bottom transparent registration system includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the vehicle bottom transparent registration system to execute the method described in any one of claims 1-7.

9. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on a vehicle bottom transparent registration system, the vehicle bottom transparent registration system is enabled to execute the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that When the computer program product runs on a vehicle bottom transparent registration system, the vehicle bottom transparent registration system is enabled to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method, system and device for displaying vehicle bottom image in panoramic image and medium

    CN111959397A

  • Vehicle control method and device, vehicle, medium and product

    CN118387126A