Multi-source fusion bird's-eye view perception target detection method, device, equipment and medium
By integrating fisheye and pinhole images in low-profile models, the problem of close distance and serious distortion of fisheye image detection is solved, and efficient environmental perception accuracy and cost reduction are achieved.
Patent Information
- Application Number
- CN202311294008.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-08
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-10-08
AI Technical Summary
In the prior art, fisheye images are used for bird's eye view perception to detect near distance, severe image distortion, uneven proportion of feature distribution, and uneven proportions, so that conventional pinhole camera images cannot be reused, resulting in high development costs and poor practicality.
By acquiring cylindrical projection parameters and image encoding features, the bird's-eye view perception network model is used to fuse fisheye and pinhole images under the same subject framework to perform feature mapping and perception tasks, including coordinate conversion, feature fusion and network training of cylindrical projection images to generate bird's-eye view plan encoding features.
It realizes the integration of fisheye and pinhole images of different installation layouts in low-profile models, reduce image distortion, optimize the proportion of feature distribution far and near, improve the accuracy of environmental perception and detection, and reduce development costs.
Smart Images

Figure CN117315424B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and specifically to a multi-source fusion bird's-eye view perception target detection method, device, equipment and medium. Background Art
[0002] With the rapid development of autonomous driving perception technology, multi-source, multi-modal sensor fusion is becoming the mainstream method for high-precision and high-accuracy autonomous vehicle perception. For example, bird's-eye view perception based on spatial mapping directly fuses information from different sources and modalities at the feature layer of each sensor data, achieving full 360-degree spatial feature coverage around the vehicle.
[0003] In current practice, vehicles with multiple high-definition pinhole cameras covering 360° around the vehicle are often equipped with higher-spec models due to their high requirements for data bandwidth and processor resources. However, low-spec models do not have the conditions to be equipped with all-around high-definition cameras. To further reduce the cost of using pinhole HD cameras, a method of using fisheye surround-view cameras that support parking to perform bird's-eye view perception is used. However, this solution only uses fisheye images, resulting in short detection distances, severe image distortion, a serious imbalance in the near-far distribution of features, and difficulty in data collection. Model training is heavily dependent on the corresponding fisheye images, and it is impossible to reuse the large number of existing conventional pinhole camera images. This leads to high development costs and poor practicality. Summary of the Invention
[0004] In view of the above-mentioned shortcomings of the prior art, the present invention provides a multi-source fusion bird's-eye view perception target detection method, device, equipment and medium to solve the above-mentioned technical problems.
[0005] The present invention provides a multi-source fusion bird's-eye view perception target detection method, which includes: obtaining cylindrical projection parameters, several frames of cylindrical projection images, image coding feature maps, and radial depth belonging probabilities corresponding to each pixel feature of the image coding feature map, wherein the cylindrical projection images include pinhole camera cylindrical projection images and fisheye camera cylindrical projection images, and the cylindrical projection parameters include cylindrical projection parameters of fisheye camera images and cylindrical projection parameters of pinhole cameras; based on the image coding feature map, the cylindrical projection parameters, the radial depth belonging probabilities, and the target vehicle coordinates The system determines the bird's-eye view plane coding features of the cylindrical projection image coding features of each frame; records the bird's-eye view plane coding features corresponding to the current frame cylindrical projection image as the current frame bird's-eye view plane coding features, and fuses the current frame bird's-eye view plane coding features with the historical frame bird's-eye view plane coding features to obtain the bird's-eye view temporal fusion features; decodes the bird's-eye view temporal fusion features to obtain the corresponding scale decoding features; connects the scale decoding features with the perception task head network to obtain a perception network model; and generalizes and trains the perception network model, and performs target detection based on the trained perception network model.
[0006] In one embodiment of the present invention, a plurality of frames of cylindrical projection images are obtained, including: collecting a plurality of camera images, the camera images including pinhole camera images and fisheye camera images; constructing a virtual cylinder based on intrinsic parameter information of the camera images, wherein the center point of the virtual cylinder is the origin of the camera optical axis, the radius of the virtual cylinder is the focal length of the camera, the central axis of the cylindrical projection image coincides with the optical axis of the camera image, and the lateral field of view angle of the cylindrical projection of the virtual cylinder is the same as the lateral field of view angle of the camera; performing coordinate distortion correction on the camera image using camera image distortion model parameters to obtain coordinates of the camera image to be converted; and using the virtual cylinder to convert the coordinates of the camera image to be converted into the cylindrical projection image of the camera.
[0007] In one embodiment of the present invention, based on the image coding feature map, the cylindrical projection parameters, the radial depth probability and the target vehicle coordinate system, the bird's-eye view plane coding features of the cylindrical projection image coding features of each frame are determined, including: mapping the image coding feature map to a three-dimensional grid feature space centered on the origin of the target vehicle coordinate system through the cylindrical projection parameters and the radial depth probability; and performing planar compression on the three-dimensional grid feature space to obtain the bird's-eye view plane coding features.
[0008] In one embodiment of the present invention, before mapping the encoded features of the cylindrical projection image of each frame to a three-dimensional grid feature space centered at the origin of the target vehicle coordinate system, the method further includes:
[0009] The cylindrical projection image coordinates are converted into the real object point coordinates in the camera coordinates by the following formula:
[0010]
[0011] Among them, the pixel coordinate x c is the direction of the cylindrical arc, coordinate y c is the height direction of the cylinder, and the depth ρ c is the radial direction of the central axis of the cylinder, f is the focal length, (X, Y, Z) is the real object point coordinate, (x c ,y c , ρ c ) is the cylindrical projection image coordinate.
[0012] In one embodiment of the present invention, mapping the encoded features of the cylindrical projection image of each frame to a three-dimensional grid feature space centered at the origin of the target vehicle coordinate system includes:
[0013] The cylindrical projection image encoding features of each frame are mapped to a three-dimensional grid feature space centered at the origin of the target vehicle coordinate system using a preset projection relationship; wherein the preset projection relationship is:
[0014]
[0015] in, Encode features for cylindrical projection images, is a three-dimensional grid feature space, C0 is the feature dimension, is the spatial image feature distribution after the probability weighting of the D equivalent cylindrical radial depth intervals corresponding to the i-th cylindrical projection image, Γ{·} represents the mapping of the encoded features of the cylindrical projection image to the Cartesian coordinate system centered on the vehicle by linear interpolation, by converting the cylindrical projection image coordinates into the real object point coordinates in camera coordinates, the equivalent cylindrical transformation parameters, and the extrinsic parameters of the camera relative to the target vehicle body.
[0016] In one embodiment of the present invention, the three-dimensional gridded feature space is planarized and compressed through a cylinder pooling operation, including: in the bird's-eye view space gridded space, all voxel feature vectors corresponding to the cylinder along the longitudinal axis height direction are added and averaged to obtain a first feature dimension; all voxel feature vectors corresponding to the cylinder along the longitudinal axis height direction are dimensionally stacked to obtain a second feature dimension, and the three-dimensional gridded feature space is planarized and compressed based on the first feature dimension and the second feature dimension.
[0017] In one embodiment of the present invention, the current frame bird's-eye view plane coding features and the historical frame bird's-eye view plane coding features are feature fused to obtain the bird's-eye view temporal fusion features, including: calculating the plane coordinate transformation relationship from the bird's-eye view plane coding features of each historical frame to the bird's-eye view plane coding features of the current frame based on the target vehicle positioning information of the historical frame bird's-eye view plane coding features and the target vehicle positioning information of the current frame bird's-eye view plane coding features; based on the plane coordinate transformation relationship, aligning the plane coordinate systems of the bird's-eye views of each historical frame to the same coordinate system to obtain the coordinate-aligned plane coding features of the historical frame bird's-eye view; setting weights for the coordinate-aligned plane coding features of the historical frame bird's-eye view, and linearly interpolating in the plane space based on the weights, fusing the coordinate-aligned plane coding features of the historical frame bird's-eye view with the current frame bird's-eye view plane coding features to obtain the bird's-eye view temporal fusion features.
[0018] In one embodiment of the present invention, the generalization training of the perception network model includes: collecting basic data information corresponding to the perception task, the basic data information including timestamp information; for different perception tasks, labeling task data samples of the basic data information to obtain corresponding labeled data, the perception tasks including: three-dimensional target detection, road passable areas and lane lines; constructing a loss function, and performing generalization training on the perception network model based on the labeled data.
[0019] In one embodiment of the present invention, a loss function is constructed, including: forming a total loss function of the perception network model based on the perception task prediction loss and the cylindrical radial depth prediction loss; wherein the perception task prediction loss includes the target detection loss and the image segmentation task loss; the target detection loss is obtained by weighted calculation of the classification loss function and the three-dimensional box regression loss function; the image segmentation task loss is calculated by comparing the predicted segmentation mask and the true segmentation mask; the depth prediction loss is obtained by calculating the cross entropy loss of simple classification or the ordered regression loss.
[0020] The present invention provides a multi-source fusion bird's-eye view perception target detection device, which includes: an acquisition module for acquiring cylindrical projection parameters, a plurality of frames of cylindrical projection images, an image coding feature map, and the radial depth probability corresponding to each pixel feature of the image coding feature map, wherein the cylindrical projection image includes a pinhole camera cylindrical projection image and a fisheye camera cylindrical projection image, and the cylindrical projection parameters include the cylindrical projection parameters of the fisheye camera image and the cylindrical projection parameters of the pinhole camera; a bird's-eye view plane coding feature generation module for generating a radial depth probability based on the image coding feature map, the cylindrical projection parameters, the radial depth probability, and the target vehicle seat. A standard system is provided for determining the bird's-eye view plane coding features of the cylindrical projection image coding features of each frame; a fusion module is used to record the bird's-eye view plane coding features corresponding to the cylindrical projection image of the current frame as the bird's-eye view plane coding features of the current frame, and to fuse the bird's-eye view plane coding features of the current frame with the bird's-eye view plane coding features of the historical frames to obtain the bird's-eye view temporal fusion features; a decoding module is used to decode the bird's-eye view temporal fusion features to obtain the corresponding scale decoding features; a training module is used to connect the scale decoding features with the perception task head network to obtain a perception network model; and, generalize the perception network model and perform target detection based on the trained perception network model.
[0021] Beneficial effects of the present invention: The multi-source fusion bird's-eye view perception target detection method, device, equipment and medium proposed in the present invention, by performing cylindrical mapping on different source images, the bird's-eye view perception network model can support the fusion of fisheye images and pinhole images with different installation layout orientations under the same main body framework, and jointly complete the bird's-eye view feature mapping and perception tasks, thereby performing target tracking detection based on the bird's-eye view plane coding features, and obtaining the tracking detection results of the target in the coding features of several cylindrical projection images; this is not only applicable to ordinary conventional pinhole camera images, but also to fisheye camera images, and ensures that when detecting close-range target objects, image distortion is reduced, and the feature distribution distance ratio can be optimized, thereby effectively improving the accuracy of environmental perception detection.
[0022] In addition, this solution can also be applied separately to 360° full fisheye images without pinhole images to complete bird's-eye view perception. It can reuse a large number of existing conventional pinhole camera images, has low development cost, and is highly practical.
[0023] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, serving to explain the principles of the present application. It is obvious that the drawings described below are merely some embodiments of the present application, and a person of ordinary skill in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0025] Figure 1 This is a flowchart of a multi-source fusion bird's-eye view perception target detection method shown in an exemplary embodiment of the present application;
[0026] Figure 2 1 is a schematic diagram of the radial depth of an equivalent cylinder of a cylindrical projection image shown in an exemplary embodiment of the present application;
[0027] Figure 3 A schematic diagram showing the orientation of a target vehicle according to an exemplary embodiment of the present application;
[0028] FIG4(a), FIG4(b), and FIG4(c) are schematic diagrams of equifocal cylindrical projection of pinhole imaging shown in an exemplary embodiment of the present application;
[0029] FIG5(a) and FIG5(b) are cylindrical projection top views of fisheye imaging and pinhole imaging with the same focal length and field of view, shown in an exemplary embodiment of the present application.
[0030] Figure 6 A schematic diagram of a camera coordinate system is shown for an exemplary embodiment of the present application;
[0031] Figure 7 This is a schematic diagram of a network model of a multi-source fusion bird's-eye view perception target detection method shown in an exemplary embodiment of the present application;
[0032] Figure 8 This is a schematic diagram of the module structure of a multi-source fusion bird's-eye view perception target detection method according to an exemplary embodiment of the present application;
[0033] Figure 9 This is a block diagram of a multi-source fusion bird's-eye view perception target detection device shown in an exemplary embodiment of the present application;
[0034] Figure 10 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0035] The following describes the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art will readily appreciate the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the various details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are intended only to illustrate the present invention and are not intended to limit the scope of protection of the present invention.
[0036] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.
[0037] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.
[0038] First, it's important to note that autonomous driving technology refers to the ability of a vehicle to autonomously perceive, make decisions, and control its driving without human intervention. This technology encompasses knowledge and technology from multiple fields, including the following key background technologies: Sensor technology: Autonomous vehicles use devices such as lidar, cameras, ultrasonic sensors, and radar to obtain information about their surroundings. The data collected by these sensors helps the vehicle identify roads, vehicles, pedestrians, and other obstacles. Perception and environmental modeling: The vehicle uses sensor data to construct a three-dimensional model of its surroundings. By identifying and analyzing roads, signs, traffic signals, obstacles, and other features, the vehicle understands its surroundings. Decision-making: Based on the perceived environment, the autonomous driving system makes decisions, such as selecting the appropriate lane, speed, and following distance, to ensure safe and efficient driving. Path planning: Autonomous vehicles calculate the optimal path to reach a destination or perform a specific task, taking into account traffic conditions, road rules, and the behavior of other vehicles. Control system: The control system is responsible for the actual operation of the vehicle, including acceleration, braking, steering, and other operations. These operations are performed based on instructions from the decision-making system. Artificial intelligence and machine learning: Autonomous driving technology extensively utilizes artificial intelligence and machine learning technologies for pattern recognition, behavior prediction, and decision-making. Deep learning algorithms can help vehicles recognize images, predict the behavior of other road users, and more. Data fusion: Autonomous driving systems must integrate data from multiple sensors to achieve comprehensive environmental awareness. This requires efficient data processing and sensor data fusion algorithms. Safety and reliability: Autonomous driving technology must operate in a variety of complex and uncertain situations, placing extremely high demands on system safety and reliability. Fault detection, redundant systems, and emergency switching are key safety measures.
[0039] The multi-source, multi-modal sensor forms supporting autonomous driving perception include various visual imaging sensors, such as telephoto cameras for long-range detection, wide-angle cameras for medium and short-range detection, and fisheye cameras for surround-view parking detection. Furthermore, various radar point cloud detection sensors, such as millimeter-wave radar, lidar, and 4D (fourth-dimensional) millimeter-wave radar, are used for spatial positioning. A representative and highly reliable solution is the fusion of visual images and lidar point clouds. Visual images possess the richest and most granular characteristics, fully capturing the color, shape, texture, and pose of objects in the surrounding scene. They are considered the optimal data source for object recognition and scene semantic representation. However, the raw information in visual images always reflects 2D (two-dimensional) planar features, lacking depth of field properties, making accurate measurement and positioning difficult. This is a core limitation of these sensors in practical applications. LiDAR and 4D millimeter-wave radar, which can detect 3D spatial depth and position, offer unparalleled advantages over vision for ranging and precise positioning of spatial objects. They are even more capable of detecting sudden intrusions or unknown obstacles than vision. BEV perception, which combines vision with 3D radar point cloud information, is arguably the optimal combination for autonomous driving environmental perception. However, the high cost of LiDAR significantly limits the mass production and deployment of this fusion method. Consequently, low-cost, multi-source fusion perception, exemplified by pure vision, has become the cost-effective solution pursued by most companies.
[0040] Purely visual multi-source fusion BEV perception leverages all onboard telephoto and wide-angle cameras to achieve 360° visual feature-level fusion perception of the vehicle's surroundings. Its core goal is to reduce sensor costs by extracting information solely from visual images. A neural network model extracts features from each image and, using depth estimation and imaging geometry, maps these visual features onto a bird's-eye view (BEV) space centered on the vehicle. This ultimately creates a comprehensive BEV feature map for various perception tasks, including detection and segmentation. Currently, the most commonly used multi-camera visual fusion BEV perception method targets pinhole camera images. This method utilizes conventional pinhole cameras positioned around the vehicle for forward, side, and rearward viewing to project BEV features. Pinhole cameras adhere to the principles of pinhole imaging geometry, resulting in simple and clear projection relationships, high image resolution and clarity, minimal distortion, and a balanced representation of objects from near and far, making them suitable for a variety of common scenarios. However, in current reality, vehicles equipped with multi-channel, high-definition pinhole cameras covering 360° around the vehicle are often higher-end models. Due to their high requirements for data bandwidth and processor resources, many low- and mid-range models are not equipped with all-around high-definition cameras. To further reduce the cost of using pinhole HD cameras, the industry has also proposed a small number of methods that use fisheye surround-view cameras that support parking for BEV perception. However, such solutions only use fisheye images, resulting in short detection distances, severe image distortion, a serious imbalance in the near-far distribution of features, and difficulty in data collection. Model training is heavily dependent on the corresponding fisheye images, and it is impossible to reuse the large number of existing conventional pinhole camera images. The development cost is high and the practicality is poor. Most of these methods are still in the research and exploration stage and are not yet universally applicable.
[0041] Pinhole cameras adhere to the principles of pinhole imaging geometry. Their projection relationships are simple and clear, resulting in high resolution and clarity, minimal distortion, and a balanced near-far representation of objects, making them suitable for a wide range of common scenarios. However, in current practice, multi-channel, high-definition pinhole cameras with 360° coverage are typically only available on higher-end vehicles. Due to their high requirements for data bandwidth and processor resources, many mid-range and low-end vehicles lack the capacity for all-around HD cameras. To further reduce the cost of pinhole HD cameras, a small number of approaches have been proposed for BEV perception using fisheye surround-view cameras with parking support. However, these approaches utilize only fisheye images, resulting in short detection ranges, severe image distortion, and severely uneven feature distribution. Furthermore, data acquisition is challenging, and model training relies heavily on the corresponding fisheye images, preventing the reuse of the large number of existing conventional pinhole camera images. This leads to high development costs and limited practicality. Most of these approaches remain at the exploratory stage and are not yet universally applicable.
[0042] Figure 1An exemplary embodiment of the present application illustrates a multi-source fusion bird's-eye view perception target detection method, which specifically includes the following steps:
[0043] Step S110, obtaining cylindrical projection parameters, several frames of cylindrical projection images, an image coding feature map, and the radial depth probability corresponding to each pixel feature of the image coding feature map, wherein the cylindrical projection images include cylindrical projection images of a pinhole camera and cylindrical projection images of a fisheye camera, and the cylindrical projection parameters include cylindrical projection parameters of the fisheye camera image and cylindrical projection parameters of the pinhole camera.
[0044] In one embodiment of the present application, the cylindrical projection image is input into the convolution backbone network for feature encoding to obtain an image encoding feature map. In this embodiment, the cylindrical projection image is a 2D visual cylindrical projection image I i (i=1, 2, ..., K), the image coding feature map is the 2D image coding feature, the convolutional backbone network is a shared 2D (Two-Dimensional) backbone network, and the 2D visual cylindrical projection image is encoded by the shared 2D backbone network to obtain the 2D image coding feature. In this embodiment, the 2D image encoding features and the camera internal and external parameters and projection cylinder parameters are respectively input into the cylindrical projection equivalent cylindrical radial depth prediction network to obtain the equivalent cylindrical radial depth probability P corresponding to each pixel feature of the 2D image. i ∈R D×H×W (i=1, 2, ..., K). Where K is a constant, H is a height value, W is a width value, D is a depth value, and C0 is a feature dimension.
[0045] In one embodiment of the present application, the cylindrical projection equivalent cylindrical radial depth prediction network, the input is the 2D image encoding feature The original camera intrinsic and extrinsic parameters and projection cylinder parameters corresponding to each image, the network structure can adopt SE (Squeeze-and-Excitation, compression and excitation) network, wherein the original camera intrinsic and extrinsic parameters and projection cylinder parameters are input into MLP (Multilayer perceptron, multilayer perceptron) network as one-dimensional signals, and the output high-dimensional vector is used as SE network weight to weight the image feature channel. Finally, the SE network outputs the equivalent cylindrical radial depth of each pixel of the cylindrical projection image, which is divided into various depth intervals according to the actual distance, and predicts the probability of each pixel belonging to each depth interval. In this embodiment, the radial depth corresponding to each pixel feature belongs to the probability P i ∈R D ×H×W(i=1, 2, ..., K). In this embodiment, the equivalent cylindrical radial depth of the cylindrical projection image refers to the length of the real object point corresponding to each pixel in the cylindrical projection image along the radial direction of the top-view projection cylinder, such as Figure 2 As shown, Figure 2 This diagram shows the equivalent cylindrical radial depth of a cylindrical projection image, an exemplary embodiment of this application. The length of the cylindrical radial depth diagram along the radial direction of the projection cylinder from a top-down perspective is equivalent to the depth of field of a perspective image. By encoding the input image features, camera parameters, and projection cylinder parameters, the equivalent cylindrical radial depth of each pixel in the cylindrical projection image is predicted. This helps understand the distance relationships between different objects in the scene and more accurately reflects the actual object locations based on the cylindrical projection image features.
[0046] In the above embodiment, the 2D convolution backbone network is the input cylindrical projection image I i (i=1, 2, ..., K) shared weights, the network can usually adopt but is not limited to a series of commonly used 2D convolutional network structures such as ResNet (residual neural network) with a certain depth level, EfficientNet, SwinTransformer, VoVNetV2, etc.
[0047] Step S120 , determining the bird's-eye view plane coding features of the cylindrical projection image coding features of each frame based on the image coding feature map, cylindrical projection parameters, radial depth belonging probability and the target vehicle coordinate system.
[0048] In one embodiment of the present application, the cylindrical projection image features Equivalent cylindrical transformation parameters and the probability P of each image pixel corresponding to the radial depth i ∈R D×H×W (i=1,2,...,K), combining the external parameters of each camera angle relative to the vehicle body coordinate system, linear interpolation sampling is performed on the merged data, and the encoded features of each cylindrical projection image are mapped to a unified BEV perspective 3D grid feature space centered on the origin of the target vehicle coordinate system. like Figure 3 As shown, Figure 3 The target vehicle orientation diagram is shown in an exemplary embodiment of this application. Using the cylinder pooling operation, the BEV spatial grid feature is flattened to obtain the BEV plane coding feature. In this way, by linearly interpolating and sampling to a unified BEV perspective, information from different perspectives can be mapped into a shared BEV space to obtain BEV planar encoding features. This feature representation can better adapt to certain perception tasks, such as target detection and obstacle recognition.
[0049] In step S130 , the bird's-eye view plane coding features corresponding to the current frame cylindrical projection image are recorded as the current frame bird's-eye view plane coding features, and the current frame bird's-eye view plane coding features and the historical frame bird's-eye view plane coding features are fused to obtain the bird's-eye view temporal fusion features.
[0050] Step S140 : decoding the bird's-eye view temporal fusion features to obtain corresponding scale decoding features.
[0051] Step S150: Connect the scaled decoding features with the perception task head network to obtain a perception network model, perform generalization training on the perception network model, and perform target detection based on the trained perception network model.
[0052] In one embodiment of the present application, a smaller secondary backbone decoding network is used to decode the temporal fusion feature of the bird's eye view, and a feature pyramid structure is used at the end to output appropriate scale features for each perception task. In this embodiment, the secondary backbone decoding network is a 2D convolutional secondary backbone network, typically employing, but not limited to, commonly used 2D convolutional network structures such as ResNet. This layer typically employs relatively shallow layers, such as ResNet-18. The secondary backbone decoding network restores useful information from the fused BEV features. By employing a feature pyramid structure within the secondary backbone decoding network, data information can be extracted from features at different scales to meet the needs of different perception tasks, enabling the target vehicle to more accurately perceive its surroundings and make precise decisions.
[0053] In one embodiment of the present application, each scale feature is connected to different perception task head networks that require scale adaptation, such as 3D target detection, road passable area segmentation, lane line extraction, etc., and finally outputs corresponding perception results such as 3D target detection, road passable area, lane line, etc.
[0054] In one embodiment of the present application, for each level of BEV scale features output by the feature pyramid FPN Usually, relatively small-scale features are connected to the 3D target detection head DtHead to improve network efficiency, while relatively large-scale features are connected to the road drivable area segmentation head RdHead and the lane line detection head LaHead to improve the efficiency of pixel segmentation calculation. DtHead, RdHead, and LaHead can flexibly use a variety of different task head networks, and the model has strong general adaptability.
[0055] exist Figure 1In the technical solution shown, by performing cylindrical mapping on different source images, the bird's-eye view perception network model can support the fusion of fisheye images and pinhole images with different installation layout orientations under the same main body framework, and jointly complete the bird's-eye view feature mapping and perception tasks, thereby performing target tracking and detection based on the bird's-eye view plane coding features, and obtaining the tracking detection results of the target in the coding features of several cylindrical projection images; this is not only applicable to ordinary conventional pinhole camera images, but also to fisheye camera images, and ensures that when detecting close-range target objects, image distortion is reduced, and the feature distribution distance ratio can be optimized, thereby effectively improving the accuracy of environmental perception detection.
[0056] In one embodiment of the present application, a plurality of frames of cylindrical projection images are obtained, including collecting a plurality of camera images, the camera images including pinhole camera images and fisheye camera images; a virtual cylinder is constructed based on the intrinsic parameter information of the camera image, the center point of the virtual cylinder is the origin of the camera optical axis, the radius of the virtual cylinder is the focal length of the camera, the central axis of the cylindrical projection image coincides with the optical axis of the camera image, the lateral field of view angle of the cylindrical projection of the virtual cylinder is the same as the lateral field of view angle of the camera, the coordinate distortion of the camera image is corrected by using the camera image distortion model parameters to obtain the coordinates of the pinhole camera image after dedistortion; and the virtual cylinder is used to convert the coordinates of the camera image to be converted into the cylindrical projection image of the camera.
[0057] In one embodiment of the present application, the camera is a pinhole camera, the camera image coordinates to be converted are the dedistorted pinhole camera image, and the cylindrical projection image is the pinhole camera cylindrical projection image. As shown in FIG4 , FIG4(a), FIG4(b), and FIG4(c) are schematic diagrams of equifocal cylindrical projection of pinhole imaging, illustrating an exemplary embodiment of the present application. First, the coordinates in the pinhole camera image are corrected using the image distortion model parameters to obtain the dedistorted pinhole camera image. Next, a virtual cylinder is constructed, with the center point of the virtual cylinder being the pinhole camera optical axis origin and the radius of the virtual cylinder being the pinhole camera focal length. The central axis of the cylindrical projection image coincides with the pinhole camera image optical axis, and the lateral field of view of the virtual cylinder cylindrical projection is the same as that of the pinhole camera. This ensures that no original image information is lost during the projection process. The vertical range can be controlled as needed and can typically be set to remove the black border area caused by the projection transformation to ensure that the projected image has no invalid areas.
[0058] The dedistorted pinhole camera image is then converted into a cylindrical projection image of the pinhole camera using the following formula:
[0059]
[0060] Among them, (x p ,y p ) is the pinhole camera image coordinate position, (xc ,y c ) is the converted cylindrical projection image coordinate position of the pinhole camera, f is the focal length of the pinhole camera, and θ is the azimuth angle of the pinhole camera image point.
[0061] In one embodiment of the present application, the camera is a fisheye camera, the camera image coordinates to be converted are perspective projection plane image coordinates, the cylindrical projection image is the cylindrical projection image of the fisheye camera, and the equifocal cylindrical projection of the fisheye imaging is shown in Figure 5. Figures 5(a) and 5(b) are top views of the cylindrical projection of the fisheye imaging and the pinhole imaging with the same focal length and field of view, shown in an exemplary embodiment of the present application. The coordinates in the fisheye camera image are corrected by the distortion polynomial model, internal parameters and distortion parameters of the fisheye camera to obtain the dedistorted fisheye camera image. According to the coordinates of the dedistorted fisheye camera image, its corresponding coordinates on the equifocal perspective projection plane are calculated to obtain the perspective projection plane image coordinates. Then, a virtual cylinder is constructed. The center point of the virtual cylinder is the origin of the fisheye camera optical axis, the radius of the virtual cylinder is the focal length of the fisheye camera, the central axis of the cylindrical projection image coincides with the optical axis of the fisheye camera image, and the lateral field of view angle of the cylindrical projection of the virtual cylinder is the same as that of the fisheye camera. The fisheye camera image is converted into the cylindrical projection image of the fisheye camera using the following formula:
[0062]
[0063] Where (x f ,y f ) is the fisheye camera image coordinate, (x w ,y w ) is the coordinate of the cylindrical projection image of the converted fisheye camera, f is the focal length of the fisheye camera, and θ is the azimuth angle of the image point of the fisheye camera.
[0064] In one embodiment of the present application, the cylindrical projection image of the fisheye camera is consistent with or similar to the cylindrical projection image of the pinhole camera in terms of field of view, so that a transformed image suitable for perception tasks or analysis tasks can be generated under the wide-angle viewing angle of the fisheye camera.
[0065] In one embodiment of the present application, the cylindrical projection image of the fisheye camera is guaranteed to be consistent or similar in field of view to the cylindrical projection image of the pinhole camera. At the same time, the horizontal and vertical field of view angles remain the same as the pinhole camera field of view to ensure the integrity of the image information. Typically, the FOV (field of view) of a fisheye camera covers an ultra-wide angle of approximately 180°, but the high-definition texture of the image is basically concentrated in the area near the central axis of the fisheye camera, and the information near the edge area is severely compressed. However, in order to minimize the blind area between each camera, the center axis of the new cylindrical projection image must be considered when setting the center axis. It is generally set at a certain angle to the center axis of the fisheye camera, and the field of view FOV is the same as the pinhole camera FOV. Usually, a new cylindrical projection image with a symmetrically opposite center direction is selected to complement the information from different perspectives to reduce blind spots.
[0066] In one embodiment of the present application, based on the image coding feature map, cylindrical projection parameters, radial depth belonging probability and target vehicle coordinate system, the bird's-eye view plane coding features of the cylindrical projection image coding features of each frame are determined, including mapping the image coding feature map to a three-dimensional grid feature space centered on the origin of the target vehicle coordinate system through the cylindrical projection parameters and the radial depth belonging probability; and performing planar compression on the three-dimensional grid feature space to obtain the bird's-eye view plane coding features. The grid features of the BEV space are processed by a cylinder pooling operation, and the 3D information of the cylinder is compressed onto the bird's-eye view plane, that is, the information of multiple data sources is fused and mapped into a unified bird's-eye view space to obtain the bird's-eye view plane coding features. This feature representation can better adapt to certain perception tasks, such as target detection, obstacle recognition, etc., that is, the fusion results of each perspective camera and depth information are obtained more intuitively, which can be used for subsequent perception tasks.
[0067] In one embodiment of the present application, before mapping the encoded features of each frame of cylindrical projection image to a three-dimensional grid feature space centered at the origin of the target vehicle coordinate system, the process further includes converting the cylindrical projection image coordinates into real object point coordinates in camera coordinates using the following formula:
[0068]
[0069] Among them, the pixel coordinate x c is the direction of the cylindrical arc, coordinate y c is the height direction of the cylinder, and the depth ρ c is the radial direction of the central axis of the cylinder, f is the focal length of the camera, (X, Y, Z) is the real object point coordinate, (x c ,y c , ρ c ) is the cylindrical projection image coordinate.
[0070] In one embodiment of the present application, Figure 6 As shown, Figure 6 A schematic diagram of a camera coordinate system is shown for an exemplary embodiment of the present application. In a cylindrical projection image, the position of each pixel is represented by coordinates in the cylindrical arc direction and the cylindrical height direction. These coordinates are calculated in the cylindrical projection transformation. In the camera coordinate system, the coordinates of the real object point can be determined by its three-dimensional position in space. The coordinates of the real object point include the depth in the radial direction of the central axis of the cylinder, which is invisible in the cylindrical projection image. Therefore, it is necessary to convert the cylindrical projection image coordinates into the coordinates of the real object point in the camera coordinate system, so as to estimate the position of the real object point in the camera coordinate system in the cylindrical projection image.
[0071] In one embodiment of the present application, the encoded features of each frame of cylindrical projection image are mapped to a three-dimensional grid feature space centered at the origin of the target vehicle coordinate system, including the following projection relationship of the encoded features of the cylindrical projection image to the three-dimensional grid feature space:
[0072]
[0073] in, Encode features for cylindrical projection images, is a three-dimensional grid feature space, C0 is the feature dimension, is the spatial image feature distribution after the probability weighting of the D equivalent cylindrical radial depth intervals corresponding to the i-th cylindrical projection image. Γ{·} represents the mapping of the cylindrical projection image encoding features to the Cartesian coordinate system centered on the vehicle by linear interpolation, by converting the cylindrical projection image coordinates into the real object point coordinates in camera coordinates, the equivalent cylindrical transformation parameters, and the extrinsic parameters of the camera relative to the target vehicle body.
[0074] In one embodiment of the present application, the 3D grid space R can be pre-set. Z×X×Y Each voxel grid unit corresponds to a different equivalent cylindrical radial depth space R of each pixel in the cylindrical projection image D×H×W The coordinate position conversion relationship in the pre-calculated and stored in the lookup table T F In the model training and inference runtime, the model is directly based on the lookup table T F Retrieve relevant position features and perform linear interpolation calculations to quickly obtain 3D mesh features from the BEV perspective
[0075] In one embodiment of the present application, a three-dimensional gridded feature space is planarized and compressed through a cylinder pooling operation, including: in the bird's-eye view gridded space, all voxel feature vectors corresponding to the cylinder along the longitudinal axis height direction are summed and averaged to obtain a first feature dimension; all voxel feature vectors corresponding to the cylinder along the longitudinal axis height direction are dimensionally stacked to obtain a second feature dimension, and the three-dimensional gridded feature space is compressed to the bird's-eye view plane based on the first feature dimension and the second feature dimension. The gridded features of the BEV space are processed through a cylinder pooling operation, and the 3D information of the cylinder is compressed to the BEV plane, that is, the information of multiple data sources is fused and mapped to a unified BEV space to obtain BEV plane encoding features. This feature representation can better adapt to certain perception tasks, such as target detection, obstacle recognition, etc.
[0076] In one embodiment of the present application, the 3D features The operation method of column pooling can be considered to be the sum and average or dimension stacking of all voxel feature vectors of the column corresponding to each grid along the Z axis height direction on the XY plane of the BEV space in the vehicle coordinate system. If it is the sum and average, the new feature dimension C1 = C0, if it is the dimension stacking, the new feature dimension C1 = C0×Z, and finally the BEV plane coding feature is obtained Here, C0 is the first feature dimension, and C1 is the second feature dimension. The gridded features of the BEV space are processed through a cylinder pooling operation, compressing the 3D information of the cylinder onto the BEV plane. This fusion of information from multiple data sources is then mapped into a unified BEV space to obtain BEV plane encoding features. This feature representation is better suited for certain perception tasks, such as object detection and obstacle recognition.
[0077] In one embodiment of the present application, the current frame bird's-eye view plane coding features and the historical frame bird's-eye view plane coding features are feature fused to obtain the bird's-eye view temporal fusion features, including calculating the plane coordinate transformation relationship from the bird's-eye view plane coding features of each historical frame to the bird's-eye view plane coding features of the current frame based on the target vehicle positioning information of the historical frame bird's-eye view plane coding features and the target vehicle positioning information of the current frame bird's-eye view plane coding features; based on the plane coordinate transformation relationship, aligning the plane coordinate systems of the bird's views of each historical frame to the same coordinate system to obtain the coordinate-aligned plane coding features of the historical frame bird's-eye view; setting weights for the coordinate-aligned plane coding features of the historical frame bird's-eye view, and linearly interpolating in the plane space based on the weights, fusing the coordinate-aligned plane coding features of the historical frame bird's-eye view with the current frame bird's-eye view plane coding features to obtain the bird's-eye view temporal fusion features. By fusing the bird's-eye view plane coding features of each historical frame with the bird's-eye view plane coding features of the current frame, not only can a more complete and accurate bird's-eye view plane coding feature be obtained, based on which the target vehicle can perceive the surrounding environment more accurately, but also richer temporal perception information can be obtained for use in subsequent tasks.
[0078] In one embodiment of the present application, the plane coding features of the bird's-eye view of each previous historical frame are The plane coding feature of the bird's-eye view of the current frame is Combined with the vehicle positioning information at the current frame time and the previous historical frame time, the plane coordinate transformation relationship of the bird's-eye view from each previous historical frame to the current frame is calculated, and the current frame bird's-eye view plane coordinate system is used as the reference coordinate system. The plane coordinate transformation of the bird's-eye view features of each previous historical frame is aligned to the current reference coordinate system. By setting weights for the plane coding features of the bird's-eye view of each historical frame and using the plane space linear interpolation method, the time series feature fusion of the bird's-eye view plane coding features is completed, and the bird's-eye view time series fusion feature is obtained. In this embodiment, weights can be assigned based on the time distance between the historical frame and the current frame, or can be designed as learnable variables. A smaller secondary backbone decoding network is then used to decode the temporal fusion features of the bird's-eye view, and a feature pyramid structure is used at its end to output appropriate scale features for each perception task. By fusing the bird's-eye view plane coding features of each historical frame with the bird's-eye view plane coding features of the current frame, not only can a more complete and accurate bird's-eye view plane coding feature be obtained, which enables the target vehicle to more accurately perceive its surrounding environment, but also richer temporal perception information can be obtained for use in subsequent tasks. Furthermore, the secondary backbone decoding network restores useful information from the fused BEV features. By adopting a feature pyramid structure in the secondary backbone decoding network, data information can be extracted from features at different scales to meet the needs of different perception tasks, allowing the target vehicle to more accurately perceive its surrounding environment and make accurate decisions.
[0079] In the above embodiment, the decoding network combined with the feature pyramid can output feature maps of appropriate scale for use in various perception tasks. These tasks may include object detection, semantic segmentation, instance segmentation, etc., providing appropriate features according to the task requirements.
[0080] In the above embodiment, the weight distribution can be manually set to determine the weight based on the time interval between the historical frame and the current frame. A neural network can also be used to learn dynamic weights, which requires considering factors such as time intervals during network training.
[0081] In one embodiment of the present application, the plane coding features of the bird's-eye view of each previous historical frame are The plane coding feature of the bird's-eye view of the current frame is Obtain the positioning information of the target vehicle at each moment corresponding to the plane coding features of the bird's-eye view of the historical frame and the plane coding features of the bird's-eye view of the current frame, and calculate the rotation and translation matrices of each historical moment relative to the current moment Then, the current frame bird's-eye view plane coordinate system is used as the reference coordinate system, and the plane coordinates of the BEV features of each previous historical frame are transformed and aligned to the current reference coordinate system to obtain the bird's-eye view features of the historical frame after coordinate alignment. Then, the bird's-eye view temporal fusion is completed by setting weights for each historical frame and using linear interpolation in plane space to obtain the bird's-eye view temporal fusion features. The weights can be manually assigned based on the temporal distance between the historical frames and the current frame, or they can be set as learnable variables and trained through the network. By fusing the bird's-eye view plane coding features of each historical frame with the bird's-eye view plane coding features of the current frame, we can not only obtain a more complete and accurate bird's-eye view plane coding feature, which enables the target vehicle to more accurately perceive its surroundings, but also obtain richer temporal perception information for subsequent tasks, allowing the target vehicle to more accurately perceive its surroundings and make accurate decisions.
[0082] In one embodiment of the present application, generalization training of a perception network model includes: collecting basic data information corresponding to perception tasks, including timestamp information; labeling task data samples of the basic data information for different perception tasks to obtain corresponding labeled data. Perception tasks include: 3D object detection, road drivable areas, and lane markings; constructing a loss function, and generalizing training the perception network model based on the labeled data. By establishing and applying this function to the perception network model and performing target detection using the perception network model, the accuracy of environmental perception detection can be effectively improved.
[0083] In one embodiment of the present application, the perception network model includes cross-view spatial 3D object detection, bird's-eye view road drivable area segmentation, and bird's-eye view lane line extraction. The annotated data includes 3D point cloud spatial object box annotation information and high-precision map road structured annotation information. After the collected data is fused and aligned using timestamp information, the basic data information is annotated for different perception tasks to obtain the annotated data to provide ground truth for the supervised training of each bird's-eye view perception network model. In this example, the annotation process involves drawing target bounding boxes, segmenting regions, and classifying labels. The perception network model is trained using the annotated data.
[0084] In one embodiment of the present application, by constructing a corresponding multi-task loss function and combining it with an effective random data augmentation method, the generalization training of the BEV perception network model is completed, and robust optimal network weights are obtained. The optimal network weights obtained from the training are provided to the perception network model for inference to optimize the perception network model so that the perception network model can accurately complete the perception task. In this way, by using a multi-task loss function, it is possible to learn and solve multiple perception tasks simultaneously. In addition, during the training process, random data augmentation methods such as translation, rotation, and scaling are applied to increase the generalization ability of the model.
[0085] In one embodiment of the present application, a loss function is constructed, including: forming a total loss function of a perception network model based on the perception task prediction loss and the cylindrical radial depth prediction loss; wherein the perception task prediction loss includes the target detection loss and the image segmentation task loss; the target detection loss is obtained by weighted calculation of the classification loss function and the three-dimensional box regression loss function; the image segmentation task loss is calculated by comparing the predicted segmentation mask and the true segmentation mask; the depth prediction loss is obtained by calculating the cross entropy loss of a simple classification or the ordered regression loss. By using a multi-task loss function, it is possible to learn and solve multiple perception tasks simultaneously, and during the training process, by applying multiple loss functions and random data augmentation methods, it is possible to train a perception network model with robustness and generalization performance for various perception tasks.
[0086] In one embodiment of the present application, the total loss function L for model training is composed of all perception task prediction losses and cylindrical radial depth prediction losses, where the target detection loss L det Generally, the classification Focal loss L cls With 3D box regression L1 loss L reg The image segmentation loss L is obtained by weighted calculation. seg It is obtained by binary cross entropy calculation, and the depth prediction loss L dep It can be constructed using either simple classification cross-entropy loss or ordered regression loss. Random data augmentation methods used to improve generalization during network training include, but are not limited to, random scale transformation, region cropping, symmetric mirroring, color transformation, and grid occlusion. This allows random modifications to input data during training to help the model better adapt to different scenarios and changes, improving the generalization performance of the perception network model. Furthermore, by using a multi-task loss function, it is possible to simultaneously learn and solve multiple perception tasks. Furthermore, by applying multiple loss functions and random data augmentation methods during training, it is possible to train a perception network model with robustness and generalization performance for use in various perception tasks.
[0087] In one embodiment of the present application, in the case where fisheye camera image sampling data is limited, a large amount of accumulated pinhole camera images are first converted into cylindrical projection images for large-scale pre-training, and then the network is fine-tuned in combination with the fisheye camera images converted into cylindrical projection images to achieve better prediction results.
[0088] In one embodiment of the present application, Figure 7 As shown, Figure 7It is a schematic diagram of the network model of a multi-source fusion bird's-eye view perception target detection method shown in an exemplary embodiment of the present application, wherein the three-dimensional grid feature space is a three-dimensional space characteristic, the bird's-eye view plane coding feature is a BEV characteristic, and the bird's-eye view temporal fusion feature is a BEV fusion feature. By collecting pinhole camera images of multiple pinholes and fisheye camera images of multiple fisheye cameras, and performing cylindrical projection conversion of the pinhole camera images with equal focal length and field of view, multiple pinhole camera cylindrical projection images and multiple fisheye camera cylindrical projection images are obtained. The cylindrical projection images are encoded through the backbone network to obtain image encoding features. The image encoding features are mapped to three-dimensional spatial features centered on the origin of the target vehicle coordinate system through the cylindrical radial depth prediction network. The three-dimensional spatial features are planarly compressed through the cylindrical pooling operation to obtain BEV features. The BEV features of each previous historical frame are aligned and fused with the BEV features of the current frame to obtain bird BEV fusion features. The multi-scale features of the BEV fusion features are decoded through the decoding network to obtain corresponding multi-scale decoding features. The decoded features of each scale are connected to the corresponding perception task head network to obtain perception network head 1, perception network head 2 and perception network head 3. In this embodiment, it includes cross-viewing space three-dimensional target detection, bird's-eye view road drivable area segmentation, and bird's-eye view lane line extraction. By performing cylindrical mapping on different source images, the bird's-eye view perception network model can support the fusion of fisheye images and pinhole images with different installation layout orientations under the same main body framework, and jointly complete the bird's-eye view feature mapping and perception tasks, thereby performing target tracking and detection based on the bird's-eye view planar coding features, and obtaining the tracking detection results of the target in the coding features of several cylindrical projection images; this not only ensures that image distortion is reduced when detecting close-range target objects, but also optimizes the feature distribution distance ratio, thereby effectively improving the accuracy of environmental perception detection.
[0089] In one embodiment of the present application, Figure 8 As shown, Figure 8A schematic diagram of the module structure of a multi-source fusion bird's-eye view perception target detection method is shown as an exemplary embodiment of the present application, wherein the image coding feature map is a 2D image coding feature, the three-dimensional grid feature space is a BEV space, and the bird's-eye view plane coding feature is a BEV characteristic. By collecting pinhole camera images and fisheye camera images, and performing cylindrical projection conversion on the pinhole camera image to obtain the pinhole camera cylindrical projection image, and performing cylindrical projection conversion on the fisheye camera image to obtain the fisheye camera cylindrical projection image, the cylindrical projection image is encoded through a shared 2D backbone network to obtain 2D image encoding features. Based on the cylindrical projection parameters and the radial depth probability, the 2D image encoding features are mapped to the BEV space centered on the origin of the target vehicle coordinate system. Through the cylindrical pooling operation, the three-dimensional grid feature space is planarized and compressed to obtain BEV features. The BEV features of each previous historical frame are aligned and fused with the BEV features of the current frame to obtain the bird's-eye view temporal fusion features. The BEV multi-scale features are decoded through the secondary backbone network to obtain the corresponding scale decoding features. The decoded features of each scale are connected to the corresponding perception task head network to obtain various perception network models, including cross-viewing space three-dimensional target detection, bird's-eye view road drivable area segmentation, and bird's-eye view lane line extraction. Cross-viewing spatial 3D object detection is trained using spatial object box annotation information from annotated 3D point clouds. Bird's-eye view road traversable area segmentation and lane line extraction are trained using structured annotation information from high-precision maps. Target detection is then performed based on the trained perception network models. A cylindrical projection equivalent cylindrical radial depth prediction network is trained using spatial coordinate information from 3D point clouds. The trained cylindrical projection equivalent cylindrical radial depth prediction network predicts cylindrical radial depth. By performing cylindrical mapping on different source images, the bird's-eye view perception network model can support the fusion of fisheye and pinhole images from different installation layouts within the same framework, jointly completing bird's-eye view feature mapping and perception tasks. This allows for target tracking and detection based on the bird's-eye view planar coding features, resulting in target tracking detection results for the coding features of multiple cylindrical projection images. This approach not only reduces image distortion when detecting close-range targets but also optimizes the near-far ratio of feature distribution, effectively improving the accuracy of environmental perception and detection.
[0090] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.
[0091] Exemplarily, the machine learning model may include a supervised model based on a neural network, such as a binary classification machine learning model. The machine learning model is trained by using a large number of historical trajectories so that the machine learning model adjusts model parameters during the training process, so that the adjusted model parameters have comprehensive predictive performance for all-round characteristics of the navigation object, such as the moving speed, moving direction, moving habits, dynamic and static state, etc.
[0092] Figure 9 This is a block diagram of a road condition refresh device shown in an exemplary embodiment of the present application. The device can be applied to Figure 2 The implementation environment shown in FIG2 is specifically configured in the smart terminal 210. The apparatus may also be applicable to other exemplary implementation environments and specifically configured in other devices. This embodiment does not limit the implementation environment to which the apparatus is applicable.
[0093] like Figure 9As shown, the exemplary multi-source fusion bird's-eye view perception target detection device 900 includes an acquisition module 901, a bird's-eye view plane coding feature generation module 902, a fusion module 903 and a training module 904. The acquisition module 901 is used to obtain cylindrical projection parameters, several frames of cylindrical projection images, image coding feature maps and radial depth belonging probability corresponding to each pixel feature of the image coding feature map. The cylindrical projection image includes a pinhole camera cylindrical projection image and a fisheye camera cylindrical projection image. The cylindrical projection parameters include cylindrical projection parameters of fisheye camera images and cylindrical projection parameters of pinhole cameras. The bird's-eye view plane coding feature generation module 902 is used to determine the bird's-eye view plane coding feature of each frame of cylindrical projection image coding feature based on the image coding feature map, cylindrical projection parameters, radial depth belonging probability and target vehicle coordinate system. Features; a fusion module 903 is used to record the bird's-eye view plane coding features corresponding to the current frame cylindrical projection image as the current frame bird's-eye view plane coding features, and fuse the current frame bird's-eye view plane coding features with the historical frame bird's-eye view plane coding features to obtain the bird's-eye view time series fusion features; a decoding module 904 is used to decode the bird's-eye view time series fusion features to obtain the corresponding scale decoding features; a training module 905 is used to connect the scale decoding features with the perception task head network to obtain a perception network model; and, generalize the perception network model and perform target detection based on the trained perception network model. By performing cylindrical mapping on different source images, the bird's-eye view perception network model can support the fusion of fisheye images and pinhole images with different installation layout orientations within the same main framework, jointly completing the bird's-eye view feature mapping and perception tasks. This allows for target tracking and detection based on the bird's-eye view planar coding features, obtaining tracking and detection results for targets within the coding features of several cylindrical projection images. This approach is applicable not only to conventional pinhole camera images but also to fisheye camera images, ensuring that image distortion is reduced when detecting close-range targets, optimizing the feature distribution ratio, and effectively improving the accuracy of environmental perception detection. Furthermore, the present invention can also be applied independently to 360° full fisheye images without pinhole images to complete bird's-eye view perception. This allows for the reuse of a large number of existing conventional pinhole camera images, resulting in low development costs and strong practicality.
[0094] In one embodiment of the present invention, an acquisition module is used to collect a plurality of camera images, including pinhole camera images and fisheye camera images; a virtual cylinder is constructed based on the intrinsic parameter information of the camera image, wherein the center point of the virtual cylinder is the origin of the camera optical axis, the radius of the virtual cylinder is the focal length of the camera, the center axis of the cylindrical projection image coincides with the optical axis of the camera image, and the lateral field of view angle of the cylindrical projection of the virtual cylinder is the same as the lateral field of view angle of the camera; coordinate distortion correction is performed on the camera image using the camera image distortion model parameters to obtain the coordinates of the camera image to be converted; and the virtual cylinder is used to convert the coordinates of the camera image to be converted into the cylindrical projection image of the camera.
[0095] In one embodiment of the present invention, a bird's-eye view planar coding feature generation module is used to map the coding features of each frame of cylindrical projection image to a three-dimensional grid feature space centered on the origin of the target vehicle coordinate system through cylindrical projection parameters and radial depth probability; through cylindrical pooling operation, the three-dimensional grid feature space is planarly compressed to obtain the bird's-eye view planar coding features.
[0096] In one embodiment of the present invention, the acquisition module is further configured to convert the cylindrical projection image coordinates into the real object point coordinates in the camera coordinates using the following formula:
[0097]
[0098] Among them, the pixel coordinate x c is the direction of the cylindrical arc, coordinate y c is the height direction of the cylinder, and the depth ρ c is the radial direction of the central axis of the cylinder, f is the focal length, (X, Y, Z) is the real object point coordinate, (x c ,y c , ρ c ) is the cylindrical projection image coordinate.
[0099] In one embodiment of the present invention, the bird's-eye view plane coded feature generation module is used to map the coded features of each frame cylindrical projection image to a three-dimensional grid feature space centered at the origin of the target vehicle coordinate system, using a preset projection relationship to map the coded features of each frame cylindrical projection image to the three-dimensional grid feature space centered at the origin of the target vehicle coordinate system; wherein the preset projection relationship is:
[0100]
[0101] in, Encode features for cylindrical projection images, is a three-dimensional grid feature space, C0 is the feature dimension, is the spatial image feature distribution after the probability weighting of the D equivalent cylindrical radial depth intervals corresponding to the i-th cylindrical projection image. Γ{·} represents the mapping of the cylindrical projection image encoding features to the Cartesian coordinate system centered on the vehicle by linear interpolation, by converting the cylindrical projection image coordinates into the real object point coordinates in camera coordinates, the equivalent cylindrical transformation parameters, and the extrinsic parameters of the camera relative to the target vehicle body.
[0102] In one embodiment of the present invention, the bird's-eye view planar coding feature generation module is also used to grid the bird's-eye view space, add and average all voxel feature vectors corresponding to the cylinder along the longitudinal axis height direction to obtain a first feature dimension; dimensionally stack all voxel feature vectors corresponding to the cylinder along the longitudinal axis height direction to obtain a second feature dimension, and planarly compress the three-dimensional gridded feature space based on the first feature dimension and the second feature dimension.
[0103] In one embodiment of the present invention, the fusion module calculates the plane coordinate transformation relationship between the bird's-eye view plane coding features of each historical frame and the bird's-eye view plane coding features of the current frame based on the target vehicle positioning information of the historical frame bird's-eye view plane coding features and the target vehicle positioning information of the current frame bird's-eye view plane coding features; based on the plane coordinate transformation relationship, the plane coordinate systems of the bird's-eye view of each historical frame are aligned to the same coordinate system to obtain the plane coding features of the bird's-eye view of the historical frame after coordinate alignment; weights are set for the plane coding features of the historical frame bird's-eye view that have completed coordinate alignment, and based on the weights, linear interpolation is performed in the plane space to fuse the plane coding features of the historical frame bird's-eye view after coordinate alignment with the plane coding features of the current frame bird's-eye view to obtain the bird's-eye view temporal fusion features.
[0104] In one embodiment of the present invention, a training module is used to collect basic data information corresponding to the perception task, the basic data information including timestamp information; for different perception tasks, the basic data information is labeled with task data samples to obtain corresponding labeled data, and the perception tasks include: three-dimensional target detection, road passable areas and lane lines; a loss function is constructed, and the perception network model is generalized and trained based on the labeled data.
[0105] In one embodiment of the present invention, the training module is configured to construct a total loss function for the perception network model based on the perception task prediction loss and the cylindrical radial depth prediction loss. The perception task prediction loss includes the target detection loss and the image segmentation task loss. The target detection loss is calculated by weighting the classification loss function with the 3D bounding box regression loss function. The image segmentation task loss is calculated by comparing the predicted segmentation mask with the true segmentation mask. The depth prediction loss is calculated using the cross entropy loss for simple classification or the ordinal regression loss.
[0106] It should be noted that the multi-source fusion bird's-eye view perception target detection device provided in the above embodiment and the multi-source fusion bird's-eye view perception target detection method provided in the above embodiment belong to the same concept, wherein the specific manner in which each module and unit performs the operation has been described in detail in the method embodiment and will not be repeated here. In actual application, the road condition refresh device provided in the above embodiment can distribute the above functions to different functional modules as needed, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above, and this is not limited here.
[0107] An embodiment of the present application also provides a device, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the electronic device implements the multi-source fusion bird's-eye view perception target detection method provided in the above-mentioned embodiments.
[0108] An embodiment of the present application also provides a medium having a computer program stored thereon. When the computer program is executed by a processor of a computer, the computer is enabled to execute the multi-source fusion bird's-eye view perception target detection method provided in the above-mentioned embodiments.
[0109] Figure 10 The following is a schematic diagram showing the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application. Figure 10 The computer system 1000 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0110] like Figure 10 As shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 1003. The CPU 1201, ROM 1002 and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0111] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk and the like; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.
[0112] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the various functions defined in the system of the present application are executed.
[0113] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0115] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0116] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When executed by a computer processor, the computer program causes the computer to perform the aforementioned multi-source fusion bird's-eye view perception target detection method. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device.
[0117] Another aspect of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the multi-source fusion bird's-eye view perception target detection method provided in each of the above embodiments.
[0118] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, any equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.
Claims
1. A multi-source fusion bird's-eye view perception target detection method, characterized in that: The multi-source fusion bird's-eye view perception target detection method includes: Obtaining cylindrical projection parameters, several frames of cylindrical projection images, an image coding feature map, and radial depth belonging probabilities corresponding to each pixel feature of the image coding feature map, wherein the cylindrical projection images include cylindrical projection images of a pinhole camera and cylindrical projection images of a fisheye camera, and the cylindrical projection parameters include cylindrical projection parameters of a fisheye camera image and cylindrical projection parameters of a pinhole camera; Determining a bird's-eye view plane coding feature of each frame of cylindrical projection image coding feature based on the image coding feature map, the cylindrical projection parameters, the radial depth probability and the target vehicle coordinate system; The bird's-eye view plane coding features corresponding to the current frame cylindrical projection image are recorded as the current frame bird's-eye view plane coding features, and the current frame bird's-eye view plane coding features and the historical frame bird's-eye view plane coding features are fused to obtain the bird's-eye view temporal fusion features; Decoding the bird's-eye view temporal fusion features to obtain corresponding scale decoding features; Connecting the scale decoding features with the perception task head network to obtain a perception network model; and generalizing the perception network model and performing target detection based on the trained perception network model; Acquiring a plurality of frames of cylindrical projection images includes: collecting a plurality of camera images, the camera images including pinhole camera images and fisheye camera images; constructing a virtual cylinder according to intrinsic parameter information of the camera images, wherein the center point of the virtual cylinder is the origin of the camera optical axis, the radius of the virtual cylinder is the focal length of the camera, the central axis of the cylindrical projection image coincides with the optical axis of the camera image, and the lateral field of view angle of the cylindrical projection of the virtual cylinder is the same as the lateral field of view angle of the camera; performing coordinate distortion correction on the camera image by using camera image distortion model parameters to obtain coordinates of the camera image to be converted; and using the virtual cylinder to convert the coordinates of the camera image to be converted into the cylindrical projection image of the camera.
2. The multi-source fusion bird's-eye view perception target detection method according to claim 1 is characterized in that: Determining a bird's-eye view plane coding feature of each frame of cylindrical projection image coding feature based on the image coding feature map, the cylindrical projection parameters, the radial depth probability, and the target vehicle coordinate system includes: Mapping the image encoding feature map to a three-dimensional grid feature space centered at the origin of the target vehicle coordinate system using the cylindrical projection parameters and the radial depth probability; The three-dimensional grid feature space is planarized and compressed to obtain bird's-eye view planar coding features.
3. The multi-source fusion bird's-eye view perception target detection method according to claim 2 is characterized in that: Before mapping the cylindrical projection image encoding features of each frame to a three-dimensional grid feature space centered at the origin of the target vehicle coordinate system, the method further includes: The cylindrical projection image coordinates are converted into the real object point coordinates in the camera coordinates by the following formula: Among them, x c is the direction of the cylindrical arc, y c is the height direction of the cylinder, ρ c is the radial direction of the central axis of the cylinder, f is the focal length, (X, Y, Z) is the real object point coordinate, (x c ,y c , ρ c ) is the cylindrical projection image coordinate.
4. The multi-source fusion bird's-eye view perception target detection method according to claim 3 is characterized in that: Mapping the cylindrical projection image encoding features of each frame to a three-dimensional grid feature space centered at the origin of the target vehicle coordinate system includes: The cylindrical projection image encoding features of each frame are mapped to a three-dimensional grid feature space centered at the origin of the target vehicle coordinate system using a preset projection relationship; wherein the preset projection relationship is: in, Encode features for cylindrical projection images, is a three-dimensional grid feature space, C0 is the feature dimension, is the spatial image feature distribution after probability weighting of the D equivalent cylindrical radial depth intervals corresponding to the i-th cylindrical projection image, Γ{·} represents the coordinates of the real object point converted from the cylindrical projection image coordinates to the camera coordinates, P i Represents the equivalent cylindrical radial depth probability corresponding to each pixel feature of the cylindrical projection image.
5. The multi-source fusion bird's-eye view perception target detection method according to claim 2 is characterized in that: The three-dimensional grid feature space is planarized and compressed by a cylinder pooling operation, including: In the bird's-eye view gridded space, all voxel feature vectors corresponding to the cylinder along the vertical axis height direction are summed and averaged to obtain a first feature dimension; All voxel feature vectors corresponding to the cylinder along the vertical axis height direction are stacked to obtain the second feature dimension; The three-dimensional gridded feature space is planarized and compressed based on the first feature dimension and the second feature dimension.
6. The multi-source fusion bird's-eye view perception target detection method according to claim 1 is characterized in that: The plane coding features of the current frame bird's-eye view and the plane coding features of the historical frame bird's-eye view are fused to obtain the bird's-eye view temporal fusion features, including: Calculate the plane coordinate transformation relationship between the plane coding features of the bird's-eye view of each historical frame and the plane coding features of the bird's-eye view of the current frame according to the target vehicle positioning information of the plane coding features of the bird's-eye view of the historical frame and the target vehicle positioning information of the plane coding features of the bird's-eye view of the current frame; Based on the plane coordinate transformation relationship, the plane coordinate systems of the bird's-eye view of each historical frame are aligned to the same coordinate system to obtain the plane coding features of the bird's-eye view of the historical frame after the coordinate alignment; Weights are set for the plane coding features of the bird's-eye view of the historical frame that has completed coordinate alignment, and linear interpolation is performed in the plane space based on the weights. The plane coding features of the bird's-eye view of the historical frame after coordinate alignment are fused with the plane coding features of the bird's-eye view of the current frame to obtain the bird's-eye view temporal fusion features.
7. The multi-source fusion bird's-eye view perception target detection method according to any one of claims 1 to 6, characterized in that: Generalization training of the perception network model includes: Collecting basic data information corresponding to the perception task, wherein the basic data information includes timestamp information; For different perception tasks, the basic data information is labeled with task data samples to obtain corresponding labeled data. The perception tasks include: 3D object detection, road passable area and lane markings; A loss function is constructed, and generalization training is performed on the perception network model based on the labeled data.
8. The multi-source fusion bird's-eye view perception target detection method according to claim 7 is characterized in that: Construct the loss function, including: The total loss function of the perception network model is formed based on the perception task prediction loss and the cylinder radial depth prediction loss; wherein the perception task prediction loss includes the target detection loss and the image segmentation task loss; the target detection loss is obtained by weighted calculation of the classification loss function and the 3D box regression loss function; the image segmentation task loss is calculated by comparing the predicted segmentation mask with the true segmentation mask; The depth prediction loss is calculated by cross entropy loss or ordinal regression loss for simple classification.
9. A multi-source fusion bird's-eye view perception target detection device, characterized in that: The multi-source fusion bird's-eye view perception target detection device includes: An acquisition module is used to acquire cylindrical projection parameters, several frames of cylindrical projection images, an image coding feature map, and the radial depth probability corresponding to each pixel feature of the image coding feature map, wherein the cylindrical projection images include pinhole camera cylindrical projection images and fisheye camera cylindrical projection images, and the cylindrical projection parameters include cylindrical projection parameters of fisheye camera images and cylindrical projection parameters of pinhole camera images; the process of acquiring several frames of cylindrical projection images includes: collecting several camera images, wherein the camera images include pinhole camera images and fisheye camera images; constructing a virtual cylinder according to the intrinsic parameter information of the camera images, wherein the center point of the virtual cylinder is the origin of the camera optical axis, the radius of the virtual cylinder is the camera focal length, the central axis of the cylindrical projection image coincides with the optical axis of the camera image, and the lateral field of view angle of the cylindrical projection of the virtual cylinder is the same as the lateral field of view angle of the camera; performing coordinate distortion correction on the camera image by using the camera image distortion model parameters to obtain the coordinates of the camera image to be converted; and using the virtual cylinder to convert the coordinates of the camera image to be converted into the cylindrical projection image of the camera; A bird's-eye view plane coding feature generation module is used to obtain the bird's-eye view plane coding features of the cylindrical projection image coding features of each frame based on the image coding feature map, the cylindrical projection parameters and the radial depth probability, and in combination with the external parameters of the view camera relative to the target vehicle coordinate system; A merging module is used to fuse the bird's-eye view plane coding features of each historical frame with the bird's-eye view plane coding features of the current frame to obtain a merged bird's-eye view time series feature; The training module is used to decode the merged bird's-eye view temporal features to obtain corresponding scale decoding features, connect each scale decoding feature to the corresponding perception task head network to obtain each perception network model, generalize the training of each bird's-eye view perception network model corresponding to each perception task, and perform target detection based on the trained bird's-eye view perception network model.
10. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the multi-source fusion bird's-eye view perception target detection method as described in any one of claims 1 to 8.
11. A computer-readable medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the multi-source fusion bird's-eye view perception target detection method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Target sensing method and device, computer equipment, storage medium and program product
CN116012805A
Image processing method and device, model training method and device, equipment and medium
CN116168271A