Omnidirectional universal AEB method and system based on pure vision
Through the omnidirectional universal AEB method based on pure vision, using multi-camera and laser point cloud data to generate panoramic images, predict the collision time between obstacles and vehicles, solving the insufficient identification of unknown obstacles by the existing AEB system, realizing omnidirectional environmental perception and collision avoidance, and improving the safety and adaptability of autonomous driving.
Patent Information
- Application Number
- CN202510238065.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-11
AI Technical Summary
Existing AEB systems rely on lidar and deep learning algorithms, and cannot effectively identify non-standard or unknown obstacles, resulting in increased collision risks and inability to cope with lateral and backward collisions.
The omnidirectional universal AEB method based on pure vision is adopted, and a panoramic image is generated using multiple cameras and laser point cloud data, and the collision time between obstacles and vehicles is predicted through a deep learning network, achieving a 360° environmental perception and collision avoidance strategy.
It reduces system costs, improves safety and adaptability, can predict potential collision risks and take real-time collision avoidance measures to adapt to various obstacles, and enhances the safety performance of autonomous driving.
Smart Images

Figure CN120298985A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of driverless technology. Specifically, it relates to an omnidirectional general AEB method and system based on pure vision. Background Art
[0002] With the continuous progress of autonomous driving technology, the AEB system has become a key technology to improve road safety and has become a standard configuration for vehicles equipped with autonomous driving and assisted driving functions. However, existing AEB systems mainly rely on sensors such as lidar, which not only increases costs but also limits the adaptability of the system in different environments. In addition, these systems usually rely on deep learning algorithms such as 3D obstacle detection and 3D occupancy grids. The results of these algorithms are limited by the pre-defined whitelist of obstacle categories, resulting in the inability to effectively handle non-standard or unknown obstacles and increasing the collision risk.
[0003] To address the growing traffic safety issues and improve driving safety, the AEB system has become an essential part of autonomous driving technology. However, most of the AEB systems equipped in existing commercial vehicles have some problems, such as the inability to handle side and rear collisions, and ignoring obstacles that cannot be recognized by the obstacle detection module.
[0004] In current AEB systems, methods based on surround-view perception technology usually convert surround-view images from multiple cameras (sometimes combined with lidar point cloud data) to the bird's-eye view (BEV) perspective for processing. In the BEV perspective, the system performs key perception tasks such as 3D obstacle detection to accurately identify and locate obstacles around the vehicle. By analyzing the positions and sizes of these obstacles, the AEB system can predict potential collision risks and take corresponding collision avoidance measures to effectively protect the safety of passengers and vehicles.
[0005] In the existing technology, a multi-camera BEV perception method and system are disclosed. In the patent application No. CN202410518815.2, an innovative multi-camera BEV perception technology is proposed, specifically for 3D object detection tasks. This technology processes multi-channel surround-view inputs through a trained perception model and can output accurate 3D object boxes, object traveling directions, and object appearance probabilities. This method significantly improves the accuracy and reliability of obstacle detection and provides strong technical support for the AEB system. However, although this 3D object detection-based AEB system has made significant progress in technology, it still has some limitations. The most important problem is that the system often cannot recognize obstacle categories outside the whitelist, which may lead to missed detections of unknown or rare obstacles. In an autonomous driving environment, such missed detections may directly cause the collision avoidance failure of the AEB system, thereby increasing the risk of collision accidents. Summary of the Invention
[0006] Aiming at the defects in the prior art, the purpose of this application is to provide an omnidirectional general AEB method based on pure vision. By using the images captured by multiple cameras distributed around the vehicle and through image processing and deep learning technologies, a comprehensive perception of the vehicle's surrounding environment is achieved. The images captured by each camera not only provide views of different angles around the vehicle, but also generate a seamless panoramic image through algorithm fusion. By predicting the scale of consecutive-frame panoramic images, points with collision risks in the 360° environment around the vehicle can be predicted and their collision times can be calculated, solving the deficiencies of existing AEB systems in obstacle detection and collision avoidance strategies, and significantly improving the safety performance of autonomous vehicles in various complex environments.
[0007] One aspect of this application provides an omnidirectional general AEB method based on pure vision, including:
[0008] Based on the scene flow information made from laser point clouds, obtain the scale change information of the corresponding pixel points of the panoramic image, and generate a panoramic image scale dataset;
[0009] Construct a deep learning network model, and train the deep learning network model through the panoramic image scale dataset to generate a panoramic image scale prediction model;
[0010] Input the panoramic image data to be measured into the panoramic image scale prediction model, obtain the predicted scale change information of the panoramic image scale data, and based on the perspective view, calculate the collision time between the obstacle and the vehicle according to the predicted scale change information, and predict the collision risk.
[0011] Further, the obtaining of the scale change information of the corresponding pixel points of the panoramic image based on the scene flow information made from laser point clouds and generating a panoramic image scale dataset includes:
[0012] Through multiple cameras and lidar around the vehicle, obtain omnidirectional images of different angles around the vehicle and lidar point cloud data, and project the lidar point cloud data onto the corresponding omnidirectional images;
[0013] During the projection process, obtain the depth value z0 of the pixel point corresponding to the lidar point cloud data in the camera coordinate system, and obtain the depth value z1 of the next-frame lidar point cloud data in the camera coordinate system according to the scene flow information;
[0014] Project the omnidirectional image onto the vehicle body coordinate system, and use cylindrical projection technology to generate a panoramic image;
[0015] Obtain each coordinate point in the panoramic image and its corresponding adjacent two groups of depth values, and calculate the scale ratio to generate a panoramic image scale dataset.
[0016] Further, the projecting the lidar point cloud data onto the corresponding surround view image includes:
[0017] Converting from the lidar coordinate system to the vehicle body coordinate system, then from the vehicle body coordinate system to the camera coordinate system, and finally projecting onto the coordinate system of the surround view image;
[0018] Obtaining the extrinsic parameters of the lidar and converting the lidar point cloud data from the lidar coordinate system to the vehicle body coordinate system of the vehicle;
[0019] According to the extrinsic parameters of the camera, converting the lidar point cloud data in the vehicle body coordinate system to the camera coordinate system of the camera;
[0020] According to the intrinsic parameters of the camera, converting the lidar point cloud data in the camera coordinate system to the image coordinate system of the surround view image to obtain the pixel positions of each point of the lidar point cloud data on the surround view image.
[0021] Further, the projecting the surround view image onto the vehicle body coordinate system and generating a panoramic image using cylindrical projection technology includes:
[0022] Converting the image data of the surround view image from the image coordinate system to the camera coordinate system of the camera through the intrinsic and extrinsic parameters of the camera;
[0023] Converting the image data in the camera coordinate system to the vehicle body coordinate system of the vehicle through the extrinsic parameters of the camera;
[0024] Projecting all three-dimensional points of the image data in the vehicle body coordinate system onto a unified panoramic image through cylindrical projection technology to generate a panoramic image.
[0025] Further, in the obtaining of the panoramic image, obtaining each coordinate point and its corresponding adjacent two sets of depth values, and calculating a scale ratio to generate a panoramic image scale dataset includes:
[0026] Taking each coordinate point (u, v) and the corresponding adjacent two sets of depth values (z0, z1) as training ground truth;
[0027] Calculating the scale ratio η between two adjacent frames, η = z1 / z0, and using it as a supervision signal;
[0028] Combining the coordinate point (u, v) and the scale ratio η to generate a panoramic image scale dataset (u, v, η).
[0029] Further, the construction of the deep learning network model, training the deep learning network model with the panoramic image scale dataset to generate a panoramic image scale prediction model, includes:
[0030] Obtain the panoramic image scale dataset, input it into the backbone network of the deep learning network model, and extract preliminary feature representations;
[0031] Input the preliminary features into the multi-scale feature extraction network of the deep learning network model to generate multi-scale feature maps;
[0032] Input the multi-scale feature maps into the perspective fusion module of the deep learning network model to generate fused feature maps;
[0033] Input the fused feature maps into the scale prediction head of the deep learning network model to predict the scale changes of each pixel in the panoramic image, and generate a panoramic image scale prediction model.
[0034] Further, the predicted scale change information is the scale change rate of the object in the panoramic image between two adjacent frames.
[0035] Further, based on the predicted scale change information, calculating the collision time between the obstacle and the vehicle based on the perspective view, includes:
[0036] Obtain the predicted scale change information and calculate the depth change rate of the object in the panoramic image data to be predicted between two consecutive frames of the camera;
[0037] Based on the perspective view, through the current depth Z t and the depth change rate η t , calculate the collision time between the object and the camera.
[0038] Further, the calculation formula of the depth change rate is
[0039] where η t is the depth change rate of the object between two consecutive frames; Z t is the position of the object from the camera at time t; Z t+1 is the position where the object moves to Z t+1 when the time progresses to t + 1;
[0040] The calculation formula of the collision time (TTC) is:
[0041]
[0042] where η tis the depth change rate of the object between two consecutive frames; Δt represents the time difference between two adjacent frames.
[0043] The second aspect of this application provides an omnidirectional general AEB system based on pure vision, including:
[0044] A dataset generation module, which is used to obtain the scale change information of the corresponding pixels of the panoramic image based on the scene flow information made from the laser point cloud, and generate a panoramic image scale dataset;
[0045] A prediction module, which is used to build a deep learning network model and train the deep learning network model through the panoramic image scale dataset to generate a panoramic image scale prediction model;
[0046] An output module, which is used to input the panoramic image data to be measured into the panoramic image scale prediction model, obtain the predicted scale change information of the panoramic image scale data, and calculate the collision time between the obstacle and the vehicle based on the perspective view according to the predicted scale change information, and predict the collision risk.
[0047] Compared with the prior art, this application has at least one of the following beneficial effects:
[0048] 1. This application proposes an omnidirectional general AEB method based on pure vision, which can detect potential collision risks in front of the vehicle, and can also monitor the collision risks from the sides and rear of the vehicle at the same time. It only uses relatively economical cameras without expensive sensors such as lidar, reducing the cost of the AEB system and improving the overall safety of the system at the same time.
[0049] 2. This application uses a general collision avoidance method based on the optical expansion principle under the perspective view. It does not depend on any specific obstacle detection algorithm, but directly uses the optical expansion principle under the perspective view to avoid obstacles. Therefore, it is not restricted by any specific detection category and shows universality for various obstacles.
[0050] 3. The omnidirectional monitoring ability of the AEB system designed in this application enables the system to predict the collision risk before the obstacle enters the vehicle's driving trajectory. It is a predictive and real-time AEB system that analyzes historical frame data to predict the future time to collision (TTC), achieving real-time collision prediction, thus greatly improving the efficiency and safety of autonomous driving and obstacle avoidance. Description of the Drawings
[0051] By reading the following detailed description of the non-restrictive embodiments with reference to the accompanying drawings, other features, purposes and advantages of this application will become more obvious:
[0052] Figure 1It is a flowchart of an omnidirectional general AEB method based on pure vision in an embodiment of the present application.
[0053] Figure 2 It is a schematic structural diagram of an around-view omnidirectional perception camera system in an embodiment of the present application.
[0054] Figure 3 It is a relationship diagram between the imaging scale and depth of an object in an embodiment of the present application. Specific implementation manners
[0055] The present application will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that those of ordinary skill in the art can make several deformations and improvements without departing from the concept of the present application. These all belong to the protection scope of the present application.
[0056] Refer to Figure 1 As shown, an omnidirectional general AEB method based on pure vision in an embodiment of the present application includes: S1. Based on the scene flow (SF) information made from laser point cloud, obtain the scale change information of corresponding pixels of the panoramic image, and generate a panoramic image scale data set; S2. Construct a deep learning network model, and train the deep learning network model through the panoramic image scale data set to generate a panoramic image scale prediction model; S3. Input the panoramic image data to be measured into the panoramic image scale prediction model to obtain the predicted scale change information of the panoramic image scale data, and based on the perspective view (PV), calculate the time to collision (TTC) between the obstacle and the vehicle, and predict the collision risk.
[0057] Based on the optical expansion (OE) principle, the present application predicts the scale change of pixels in the image under the perspective view, and calculates the collision time accordingly to achieve a more general obstacle avoidance strategy, which does not depend on any specific obstacle detection algorithm and is not limited by the whitelist of any obstacle category, thus significantly improving the adaptability and robustness of the system.
[0058] Specifically, first, scene flow information is generated using lidar point cloud data. The scene flow information describes the motion state of objects in three-dimensional space over time, including speed, direction, acceleration, attitude changes, etc. Then, the lidar point cloud data is mapped onto a panoramic image, and by calculating the scale change information corresponding to each pixel point, a panoramic image scale dataset is generated. Then, by designing and constructing a deep learning network model, the generated panoramic image scale dataset is used to train the network model. By continuously adjusting the model parameters, its prediction performance is optimized so that it can output corresponding scale change prediction information. The panoramic image data to be measured is input into the trained panoramic image scale prediction model to obtain the predicted scale change information. Then, based on this information, the time to collision between the obstacle and the vehicle is calculated, and the collision risk is predicted, enabling timely collision avoidance measures to be taken before a collision occurs, thus greatly improving the safety of the vehicle.
[0059] In some specific embodiments, based on the scene flow information made from lidar point cloud data, the scale change information of the corresponding pixel points of the panoramic image is obtained to generate a panoramic image scale dataset, including:
[0060] Through multiple cameras and lidar around the vehicle, omnidirectional images and lidar point cloud data at different angles around the vehicle are obtained, and the lidar point cloud data is projected onto the corresponding omnidirectional images. During the projection process, the depth value z0 of the pixel point corresponding to the lidar point cloud data in the camera coordinate system is obtained, and according to the scene flow information, the depth value z1 of the next frame of lidar point cloud data in the camera coordinate system is obtained. The omnidirectional image is projected onto the vehicle body coordinate system, and a panoramic image is generated using cylindrical projection technology. For each coordinate point in the panoramic image and its corresponding adjacent two sets of depth values, the scale ratio is calculated to generate a panoramic image scale dataset.
[0061] Specifically, during the execution process, by using multiple cameras and lidar around the vehicle, omnidirectional images and lidar point cloud data are obtained simultaneously, and the lidar point cloud data is projected onto the corresponding omnidirectional images to achieve the fusion of image and point cloud data. Then, during the projection process, the depth value z0 of the pixel point corresponding to the lidar point cloud data in the camera coordinate system is obtained. Then, according to the scene flow information, the depth value z1 of the next frame (all points in the two point clouds correspond one by one) of lidar point cloud data in the camera coordinate system is obtained. Subsequently, the omnidirectional image is projected onto the vehicle body coordinate system, and a panoramic image is generated using cylindrical projection technology. Finally, by calculating the scale ratio, that is, the change rate of the depth value between adjacent frames of the panoramic image, a panoramic image scale dataset is generated.
[0062] Exemplarily, such as Figure 2As shown in the figure, the panoramic omnidirectional perception system adopted in this application consists of six high-resolution cameras, which are evenly distributed at different positions of the vehicle to ensure full coverage of the surrounding environment of the vehicle body. After adjusting the positions and angles of the six cameras, the overlapping area of their fields of view can be maximized. The overlapping part of the fields of view not only provides additional data redundancy for the system and enhances the robustness of the perception system, but more importantly, these overlapping areas enable the system to realize the mutual correlation and fusion between perspectives through image processing and deep learning algorithms, so as to construct a coherent and consistent environmental model. By using the panoramic camera array and combining deep learning technology, the system can realize the omnidirectional perception of the 360° environment around the vehicle, capture all dynamic and static objects around the vehicle, and provide a solid data basis for subsequent obstacle detection and collision avoidance decision-making.
[0063] Specifically, projecting the lidar point cloud data onto the corresponding panoramic image includes: converting from the lidar coordinate system to the vehicle body coordinate system, then from the vehicle body coordinate system to the camera coordinate system, and finally projecting onto the coordinate system of the panoramic image; obtaining the extrinsic parameters of the lidar and converting the lidar point cloud data from the lidar coordinate system to the vehicle body coordinate system of the vehicle; according to the extrinsic parameters of the camera, converting the lidar point cloud data in the vehicle body coordinate system to the camera coordinate system of the camera; according to the intrinsic parameters of the camera, converting the lidar point cloud data in the camera coordinate system to the image coordinate system of the panoramic image to obtain the pixel position of each point of the lidar point cloud data on the panoramic image.
[0064] By accurately projecting the lidar point cloud data onto the corresponding panoramic image, including the process of coordinate transformation, obtaining the intrinsic and extrinsic parameters of the camera, and converting the lidar point cloud data from the lidar coordinate system to the vehicle body coordinate system; using the extrinsic parameters of the camera, converting the lidar point cloud data in the vehicle body coordinate system to the camera coordinate system of the camera; using the intrinsic parameters of the camera and through the principle of perspective projection, projecting the points in the camera coordinate system onto the image coordinate system of the panoramic image, and finally the pixel position of each point of the lidar point cloud data on the panoramic image can be obtained.
[0065] Among them, the extrinsic parameters of the lidar represent the translational and rotational relationships of the lidar relative to the vehicle body; the extrinsic parameters of the camera represent the translational and rotational relationships of the camera relative to the vehicle body; the intrinsic parameters of the camera represent the transformation relationship from the camera coordinate system to the image coordinate system.
[0066] When projecting the lidar point cloud onto the image, the following several coordinate systems are involved:
[0067] 1. The lidar coordinate system, with the origin located at the lidar; 2. The vehicle body coordinate system, with the origin located at the center point of the rear axle of the vehicle; 3. The camera coordinate system, with the origin located at the camera; 4. The image coordinate system, a 2D plane coordinate system representing the imaging plane of the image, and the origin is usually located at the upper left of the image.
[0068] The specific parameters involved are as follows: 1. The extrinsic parameters of the lidar, which specifically include a rotation matrix R and a translation vector T; the rotation matrix R is a 3x3 matrix representing the rotation relationship of the lidar coordinate system relative to the vehicle body coordinate system; the translation vector T represents the translation relationship of the lidar coordinate system relative to the vehicle body coordinate system; 2. The extrinsic parameters of the camera, similar to those of the lidar, include a rotation matrix R and a translation vector T, representing the rotation and translation relationships between the camera coordinate system and the vehicle body coordinate system; 3. The intrinsic parameters of the camera: include the focal length (which determines the spatial scale in the image), the principal point (representing the camera imaging center), and the distortion coefficients (used to correct the distortion of the camera lens).
[0069] Specifically, project the surround-view image onto the vehicle body coordinate system and generate a panoramic image using cylindrical projection technology, including: converting the image data of the surround-view image from the image coordinate system to the camera coordinate system of the camera through the intrinsic and extrinsic parameters of the camera; converting the image data in the camera coordinate system to the vehicle body coordinate system through the extrinsic parameters of the camera; projecting all three-dimensional points of the image data in the vehicle body coordinate system onto a unified panoramic image through cylindrical projection technology to generate a panoramic image.
[0070] Through the intrinsic and extrinsic parameters of the camera, project the image data onto the camera coordinate system, and then, using the extrinsic parameters of the camera and the relative position relationship between the vehicle body and the camera, further project it onto the vehicle body coordinate system. Using the existing cylindrical projection technology, project all three-dimensional points in the vehicle body coordinate system onto a virtual cylinder, which surrounds the vehicle body and unfolds into a two-dimensional plane. Project all three-dimensional points onto a unified panoramic image. Thanks to the extrinsic parameters of the camera, the relative positions of the image data of each camera in the vehicle body coordinate system are accurately calibrated, and the image data of different cameras can be seamlessly stitched together to ensure the seamless stitching of the final panoramic image, ultimately forming a complete 360-degree panoramic view.
[0071] This application accurately utilizes the intrinsic and extrinsic parameters of the camera and cylindrical projection technology to achieve the three-dimensional point projection from image data to the vehicle body coordinate system and generate a seamlessly stitched 360-degree panoramic view, improving the accuracy and integrity of environmental perception, helping the system better understand and respond to complex traffic scenarios, and enhancing the safety and reliability of autonomous driving.
[0072] Specifically, when obtaining the panoramic image, for each coordinate point and its corresponding adjacent two sets of depth values, calculate the scale ratio to generate a panoramic image scale dataset, including: taking each coordinate point (u, v) and the corresponding adjacent two sets of depth values (z0, z1) as the training ground truth; calculating the scale ratio η between two adjacent frames, η = z1 / z0, and using it as the supervision signal; converting the data into the set data format (u, v, η) to generate a panoramic image scale dataset.
[0073] In the panoramic view, each coordinate point (u, v) and the two adjacent sets of depth values (z0, z1) corresponding to the point are recorded as the training true value, reflecting the distance change between the object in the scene and the camera. Finally, each coordinate point (u, v) and the corresponding scale ratio η are sorted according to the set data format (u, v, η) to form a panoramic image scale data set, that is, according to each coordinate point (u, v) and the two adjacent sets of depth values (z0, z1) corresponding to the point, the scale ratio η of two adjacent frames is calculated. The coordinate point (u, v) and the scale ratio η are combined to generate the true value (u, v, η) of the panoramic image scale dataset. In addition to the true value, the panoramic image scale dataset also includes surround images collected by the vehicle-mounted camera. The data format is (u, v, η), where η = z1 / z0 represents the scale ratio of two adjacent frames. This ratio will be used as a supervisory signal for the training of the deep learning network to ensure that the model can accurately predict scale changes. It not only contains the spatial position information of the image, but also incorporates the deep scale change information, providing rich training data, enhancing the spatial perception ability of image data and the scale invariant feature expression ability, and helping to improve the robustness and generalization ability of the algorithm in complex scenes.
[0074] In some specific embodiments, a deep learning network model is constructed, and the deep learning network model is trained with a panoramic image scale dataset to generate a panoramic image scale prediction model, including: obtaining a panoramic image scale dataset, inputting it into a backbone network of the deep learning network model, and extracting a preliminary feature representation; inputting the preliminary features into a multi-scale feature extraction network of the deep learning network model to generate a multi-scale feature map; inputting the multi-scale feature map into a perspective fusion module of the deep learning network model to generate a fused feature map; and inputting the fused feature map into a scale prediction head of the deep learning network model to predict the scale change of each pixel in the panoramic image, and generating a panoramic image scale prediction model.
[0075] Among them, the preliminary feature representation includes information such as edges, textures, shapes, and object boundaries.
[0076] In a specific implementation, the model training process includes the following steps: 1. Data set preparation and preprocessing. Preprocessing: includes uniformly resizing the image to ensure the consistency of the input data size; in addition, data enhancement of the image, such as rotation, scaling, flipping, etc., is performed to increase the diversity and robustness of the training data.
[0077] 2. Define the model structure. The model mainly includes the following parts: a. Backbone network. ResNet is selected as the backbone network to extract preliminary feature representations from panoramic images. b. Multi-scale feature extraction network. On the feature map output by the backbone network, convolution kernels of different scales are used to extract multi-scale feature maps, which helps the network understand objects of different sizes and scene details. c. Perspective fusion module. Information from different perspectives is fused to enable panoramic understanding. Fusion is performed through the attention module, combining features from different perspectives to generate a fused feature map. d. Scale prediction head. The fused feature map is input into the scale prediction head to predict the scale change value of each pixel. The output result is compared with the true value, the loss is calculated and used for back propagation.
[0078] 3. Model training: including defining loss functions, such as L1 loss function Use optimizers such as Adam and SGD for optimization. The learning rate gradually decays according to training.
[0079] Training process: The training data is fed into the network in batches for forward propagation, and the predicted output of each batch is calculated. The loss between the predicted results and the true value is calculated. The gradient is calculated and the network weights are updated through the back-propagation algorithm. The training process lasts for multiple iterations, and each iteration updates the network parameters to reduce the loss and improve the model accuracy.
[0080] 4. Model validation and evaluation: After each iteration, use the validation set to validate the model and evaluate the performance of the model on an unseen dataset. Adjust the model hyperparameters based on the results of the validation set to avoid overfitting or underfitting.
[0081] This application designs a deep learning network model to accurately estimate the scale information in panoramic images. The model consists of the following key components:
[0082] 1. Backbone network: As the foundation of the model, the backbone network is responsible for extracting preliminary feature representations from the input surround image (panoramic image scale dataset, including surround images and ground truth).
[0083] 2. Multi-scale Feature Extraction Network: This network further processes the features from the backbone network to generate multi-scale feature maps to capture different levels of visual details.
[0084] 3. Viewpoint Fusion Module: Considering the image information captured by the surround cameras from different perspectives, this module is responsible for integrating these multi-view features to ensure the accuracy and consistency of scale estimation. At the same time, it corresponds to the ground truth in the panoramic image space.
[0085] 4. Scale Prediction Head: Based on the fused feature maps, the Scale Prediction Head is responsible for predicting the scale change of each pixel.
[0086] Exemplarily, a specific implementation is that during the training of the model, first, the image data from 6 surround cameras are obtained, and the corresponding scene flow point cloud information is projected onto the images. Then, cylindrical projection is used to achieve image stitching to generate a 360-degree panoramic image. Then, the scale ratio of each pixel point in the panoramic image is calculated as the training ground truth for supervising the learning of the model, ensuring that the model can accurately learn the relationship between scale changes and images.
[0087] Next, the panoramic image data formed by the 6 surround images of two adjacent frames are sequentially input into the backbone network and the multi-scale feature extraction network to obtain multi-scale feature maps. These feature maps are then passed to the Viewpoint Fusion Module, which is responsible for integrating the feature maps from different perspectives to generate a comprehensive and consistent scale feature representation.
[0088] Finally, the fused feature maps are fed into the Scale Prediction Head to predict the scale change of each pixel, obtaining the final panoramic image scale prediction model to predict the scale change of each pixel.
[0089] Through the above implementation, the model can effectively process the image data from multiple perspectives and accurately estimate the scale change, which is crucial for obstacle avoidance and decision-making in the autonomous driving system. In addition, the training process of the model makes full use of the scene flow point cloud information, further improving the accuracy of scale estimation and making it more reliable and effective in practical applications.
[0090] In some specific embodiments, the predicted scale change information is the scale change rate of the objects in the panoramic image between two adjacent frames. By predicting the scale of consecutive frame panoramic images, it is possible to predict the points with collision risks in the 360° environment around the vehicle and calculate their collision times.
[0091] Through panoramic perception of the environment, this application can analyze the driving trajectory of a vehicle in real time and predict potential collision risks. It can not only identify obstacles in the current field of view but also, through advanced data fusion and prediction algorithms, foresee possible collision situations in the future. With forward-looking risk assessment capabilities, it can take timely collision avoidance measures before a collision occurs, thus greatly improving the safety of the vehicle.
[0092] In some specific embodiments, according to the predicted scale change information, based on the perspective view, calculate the time to collision between the obstacle and the vehicle, including: obtaining the predicted scale change information, and calculating the depth change rate of the object in the to-be-predicted panoramic image data between two consecutive frames of the camera; based on the perspective view, through the current depth Z of the object t and the depth change rate η t , calculate the time to collision between the object and the camera.
[0093] In the above embodiment, based on the relationship between the scale change of the pixel points in the panoramic image and the time to collision under the perspective view, by predicting the scale change in the panoramic image, the time to collision between the obstacle and the vehicle can be calculated. Through this relationship analysis, the system can make a more accurate judgment before a collision occurs and take corresponding collision avoidance measures.
[0094] Specifically, under the perspective view, the scale of the object on the imaging plane is closely related to its depth from the camera. When the object approaches the camera, its size in the image will increase, and when the object moves away from the camera, its size will decrease. This scale change, together with the object's movement speed and the depth from the camera, determines the time to collision between the object and the camera.
[0095] By introducing the scale change prediction based on the optical expansion principle under the perspective view, this application can predict the scale change of the pixel points in the image according to the input image information of consecutive frames, thereby judging the time to collision of each pixel point in the image and further predicting the collision risk. Different from the traditional obstacle detection-based method, the method of this application does not rely on any specific obstacle detection algorithm and is not limited by any white list of obstacle categories. Whether it is common vehicles, pedestrians, or rare obstacles such as fallen trees or temporary roadblocks, it can be effectively processed, thus greatly improving the versatility and adaptability of the system.
[0096] As Figure 3 shown, it shows the relationship between the imaging scale of the object and the depth. At time t, an object is located at a position Z t from the camera, and its scale on the imaging plane is s t . As time progresses to t + 1, the object moves to a position Z t+1 from the camera, and its scale becomes s t+1 .
[0097] Among them, the scale change is
[0098] Next, by obtaining the predicted scale change information, calculate the depth change rate of the object in the panoramic image data to be predicted between two consecutive frames of the camera. Among them, the depth change rate η t The relationship with the scale change is:
[0099]
[0100] Therefore, the calculation formula for the depth change rate is
[0101] In the formula, η t is the depth change rate of the object between two consecutive frames; Z t is the position of the object from the camera at time t; Z t+1 When the time progresses to t + 1, the object moves to Z t+1 position;
[0102] Finally, based on the perspective view, through the current depth Z of the object t and the depth change rate η t , calculate the collision time between the object and the camera. The collision time TTC of the object can be calculated by the following formula:
[0103]
[0104] Among them, represents the depth change rate of the object between two consecutive frames; Δt represents the time difference between two adjacent frames.
[0105] Therefore, by predicting the scale change of the pixels, the collision time TTC between the object and the camera can be calculated. This method allows us to directly estimate the TTC from the image sequence without relying on traditional obstacle detection algorithms, thus providing a new, purely vision-based collision risk assessment method for the AEB system. The method of this application is not restricted by the obstacle category whitelist, can adapt to various complex scenarios, provides more comprehensive and reliable safety guarantees for autonomous driving vehicles, solves the deficiencies of existing AEB systems in obstacle detection and collision avoidance strategies, and significantly improves the safety performance of autonomous driving vehicles in various complex environments.
[0106] In the second aspect of the present application, a pure vision-based omnidirectional general AEB system is provided, including: a dataset generation module, which is used to obtain the scale change information of the corresponding pixel points of the panoramic image based on the scene flow information made from the laser point cloud, and generate a panoramic image scale dataset; a prediction module, which is used to train the deep learning network model through the panoramic image scale dataset by constructing a deep learning network model, and generate a panoramic image scale prediction model; an output module, which is used to input the panoramic image data to be measured into the panoramic image scale prediction model, obtain the predicted scale change information of the panoramic image scale data, and based on the predicted scale change information, calculate the collision time between the obstacle and the vehicle from the perspective of perspective, and predict the collision risk.
[0107] Through the dataset generation module, a high-quality dataset is constructed based on the production of the panoramic image scale dataset of the scene flow. This dataset is based on the scene flow information and provides a basis for the training of the deep learning model. At the same time, under the perspective view, the relationship between the scale change of the pixel points in the image and the collision time is such that by predicting the scale change in the panoramic image through the prediction module, the collision time between the obstacle and the vehicle can be calculated; through this relationship analysis, more accurate judgments can be made before the collision occurs, and corresponding collision avoidance measures can be taken.
[0108] The pure vision-based omnidirectional general AEB method of the present application uses the images captured by multiple cameras distributed around the vehicle. Through image processing and deep learning technologies, it realizes a comprehensive perception of the vehicle's surrounding environment. The images captured by each camera not only provide views of different angles around the vehicle, but also generate a seamless panoramic image through algorithm fusion. By predicting the scale of consecutive frame panoramic images, it is possible to predict the points with collision risks in the 360° environment around the vehicle and calculate their collision times, solving the deficiencies of existing AEB systems in obstacle detection and collision avoidance strategies, and significantly improving the safety performance of autonomous vehicles in various complex environments.
[0109] The specific embodiments of the present application have been described above. It should be understood that the present application is not limited to the above specific implementation manners, and those skilled in the art can make various deformations or modifications within the scope of the claims, which does not affect the essence of the present application. The above preferred features can be combined arbitrarily without conflict.
Claims
1. A pure vision-based omnidirectional general AEB method, characterized in that, Including: Based on the scene flow information made from lidar point clouds, obtain the scale change information of the corresponding pixel points of the panoramic image, and generate a panoramic image scale dataset; Construct a deep learning network model, and train the deep learning network model through the panoramic image scale dataset to generate a panoramic image scale prediction model; Input the panoramic image data to be measured into the panoramic image scale prediction model to obtain the predicted scale change information of the panoramic image scale data, and based on the predicted scale change information, calculate the collision time between the obstacle and the vehicle from a perspective view to predict the collision risk.
2. The omnidirectional general AEB method based on pure vision according to claim 1, wherein The method of obtaining the scale change information of the corresponding pixel points of the panoramic image based on the scene flow information made from lidar point clouds and generating a panoramic image scale dataset includes: Through multiple cameras and lidar around the vehicle, obtain the omnidirectional images and lidar point cloud data at different angles around the vehicle, and project the lidar point cloud data onto the corresponding omnidirectional images; During the projection process, obtain the depth value z0 of the pixel point corresponding to the lidar point cloud data in the camera coordinate system, and obtain the depth value z1 of the next frame of lidar point cloud data in the camera coordinate system according to the scene flow information; Project the omnidirectional image onto the vehicle body coordinate system, and use the cylindrical projection technology to generate a panoramic image; Obtain each coordinate point and its corresponding adjacent two sets of depth values in the panoramic image, and calculate the scale ratio to generate a panoramic image scale dataset.
3. The omnidirectional general AEB method based on pure vision according to claim 2, wherein The step of projecting the lidar point cloud data onto the corresponding omnidirectional image includes: Convert from the lidar coordinate system to the vehicle body coordinate system, then from the vehicle body coordinate system to the camera coordinate system, and finally project to the coordinate system of the omnidirectional image; Obtain the extrinsic parameters of the lidar, and convert the lidar point cloud data from the lidar coordinate system to the vehicle body coordinate system of the vehicle; According to the extrinsic parameters of the camera, convert the lidar point cloud data in the vehicle body coordinate system to the camera coordinate system of the camera; According to the intrinsic parameters of the camera, convert the lidar point cloud data in the camera coordinate system to the image coordinate system of the omnidirectional image to obtain the pixel position of each point of the lidar point cloud data on the omnidirectional image.
4. The omnidirectional general AEB method based on pure vision according to claim 2, characterized in that, The step of projecting the omnidirectional image onto the vehicle body coordinate system and using the cylindrical projection technology to generate a panoramic image includes: Through the intrinsic and extrinsic parameters of the camera, convert the image data of the omnidirectional image from the image coordinate system to the camera coordinate system of the camera; Through the extrinsic parameters of the camera, convert the image data in the camera coordinate system to the vehicle body coordinate system of the vehicle; Through the cylindrical projection technology, project all three-dimensional points of the image data in the vehicle body coordinate system onto a unified panoramic image to generate a panoramic image.
5. The omnidirectional general AEB method based on pure vision according to claim 2, wherein The step of obtaining each coordinate point and its corresponding adjacent two sets of depth values in the panoramic image and calculating the scale ratio to generate a panoramic image scale dataset includes: Take each coordinate point (u, v) and the corresponding adjacent two sets of depth values (z0, z1) as the training ground truth; Calculate the scale ratio between two adjacent frames, η = z1 / z0, and use it as the supervision signal. Combine the coordinate points (u, v) and the scale ratio η to generate a panoramic image scale dataset (u, v, η).
6. The omnidirectional general AEB method based on pure vision according to claim 1, characterized in that, The construction of the deep learning network model, training the deep learning network model with the panoramic image scale dataset to generate a panoramic image scale prediction model, includes: Obtain the panoramic image scale dataset and input it into the backbone network of the deep learning network model to extract preliminary feature representations; Input the preliminary features into the multi-scale feature extraction network of the deep learning network model to generate multi-scale feature maps; Input the multi-scale feature maps into the view fusion module of the deep learning network model to generate fused feature maps; Input the fused feature maps into the scale prediction head of the deep learning network model to predict the scale changes of each pixel in the panoramic image and generate a panoramic image scale prediction model.
7. A pure vision-based omnidirectional general AEB method according to claim 1, characterized in that The predicted scale change information is the scale change rate of the objects in the panoramic image between two adjacent frames.
8. A pure vision-based omnidirectional general AEB method according to claim 1, characterized in that, Based on the predicted scale change information, calculate the time to collision between the obstacle and the vehicle from the perspective of perspective, including: Obtain the predicted scale change information and calculate the depth change rate of the objects in the panoramic image data to be predicted between two consecutive frames of the camera; Based on the perspective view, through the current depth Z of the object t and the depth change rate η t , calculate the collision time between the object and the camera.
9. The omnidirectional general AEB method based on pure vision according to claim 8, characterized in that The calculation formula for the depth change rate is as follows: where η t is the depth change rate of the object between two consecutive frames; Z t is the position of the object from the camera at time t; Z t+1 is the position where the object moves to Z t+1 when the time elapses to t + 1; The calculation formula for the time to collision (TTC) is: where η t is the depth change rate of the object between two consecutive frames; Δt represents the time difference between two adjacent frames.
10. An omnidirectional general AEB system based on pure vision, characterized in that, Include: A dataset generation module for obtaining the scale change information of the corresponding pixel points of the panoramic image based on the scene flow information made from the laser point cloud and generating a panoramic image scale dataset; A prediction module for generating a panoramic image scale prediction model by constructing a deep learning network model and training the deep learning network model with the panoramic image scale dataset; An output module for inputting the panoramic image data to be measured into the panoramic image scale prediction model, obtaining the predicted scale change information of the panoramic image scale data, and calculating the time to collision between the obstacle and the vehicle from the perspective of perspective according to the predicted scale change information to predict the collision risk.
Citation Information
Patent Citations
Multi-camera BEV sensing method and system
CN118429926A
Cited By
AEB enhancement method based on WiFi imaging and multi-modal perception fusion
CN122347791A
AEB enhancement method based on fusion of wifi imaging and multi-modal perception
CN122347791B