Training method and device of obstacle detection model and obstacle detection method and device

The fisheye camera image is projected onto a two-dimensional BEV grid through the obstacle detection model, combining multi-layer network extraction and fusion features, solving the problem of difficulty in identifying overhangs or irregular obstacles, achieving efficient obstacle detection on low-computing power platforms, and improving the safety of automatic parking.

CN120496026AActive Publication Date: 2025-08-15CHENGDU TIANFU INVO TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510570186.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify overhanging or irregular obstacles in underground parking lots, resulting in increased safety risks during automatic parking. In particular, obstacle detection methods based on two-dimensional environmental images cannot provide three-dimensional position information, and the three-dimensional data processing and calculation is large, making it difficult to deploy on low-computing platforms.

Method used

Using the obstacle detection model, sample images are obtained through the fisheye camera, feature maps are extracted using the first backbone network and the first neck network, and the visual converter projects the features to a two-dimensional BEV grid, combines the second backbone network and the second neck network for feature fusion, and the prediction head performs obstacle height prediction, which is suitable for low-computing power platforms.

Benefits of technology

It improves the detection accuracy of overhanging or irregular obstacles, reduces the amount of calculation, and is suitable for deployment on low-computing platforms to ensure the safety of automatic parking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496026A_ABST
    Figure CN120496026A_ABST
Patent Text Reader

Abstract

The invention provides a training method and device of an obstacle detection model and an obstacle detection method and device. The training method comprises the following steps: acquiring sample data, wherein the sample data comprises a sample image of a target area shot by a fisheye camera and a sample label; inputting the sample image into an obstacle detection model, and obtaining a first feature map through a first backbone network and a first neck network in sequence; the visual converter projects features representing different heights of the target area to a two-dimensional BEV grid based on the first feature map to obtain a BEV feature map; performing feature extraction and fusion on the BEV feature map through a second backbone network and a second neck network in sequence, and outputting a second feature map; inputting the second feature map into a prediction head to obtain a prediction thermodynamic diagram, prediction occupation information of the grids and obstacle prediction height corresponding to each grid; and adjusting parameters of the obstacle detection model based on the prediction thermodynamic diagram, the obstacle prediction height corresponding to each grid and the sample label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of assisted driving technology, and in particular to a training method and device for an obstacle detection model, and an obstacle detection method and device. Background Art

[0002] As a key component of autonomous driving, automatic parking technology has become a key hallmark of modern automotive intelligence. During automatic parking, the vehicle's controller must accurately identify obstacles around the vehicle to ensure safe and efficient parking.

[0003] Automated parking technology faces multiple challenges in complex environments like underground parking lots. For example, there are many difficult-to-identify obstacles in underground parking lots. These obstacles are often irregular, with small ground contact areas and large suspended areas, or completely suspended obstacles, such as gates, suspended pipes, and ground locks. These obstacles can affect the automated parking system's ability to identify and handle obstacles. Summary of the Invention

[0004] In view of this, the present disclosure provides a method and device for training an obstacle detection model, and a method and device for obstacle detection.

[0005] In a first aspect, a training method for an obstacle detection model is provided. The obstacle detection model includes: a first backbone network, a first neck network, a visual converter, a second backbone network, a second neck network, and a prediction head. The training method includes: obtaining sample data, the sample data including a sample image of a target area captured by a fisheye camera and a sample label; inputting the sample image into the obstacle detection model, sequentially passing the sample image through the first backbone network and the first neck network to obtain a first feature map, the first feature map including depth information of the target area; based on the first feature map, the visual converter projects features representing different heights of the target area onto a two-dimensional BEV grid to obtain a BEV feature map; sequentially passing the BEV feature map through the second backbone network and the second neck network to extract and fuse features, outputting a second feature map; inputting the second feature map into the prediction head to obtain a predicted heat map, predicted occupancy information of the grid, and the predicted obstacle height corresponding to each grid; and adjusting the parameters of the obstacle detection model based on the predicted heat map, the predicted obstacle height corresponding to each grid, and the sample labels.

[0006] In some embodiments, the sample image includes multiple sample images captured by multiple fisheye cameras at the same time, and the shooting fields of view of at least two fisheye cameras have overlapping areas.

[0007] In some embodiments, a sample image is input into an obstacle detection model and sequentially passes through a first backbone network and a first neck network to obtain a first feature map, including: the first backbone network extracts features from multiple sample images, and each sample image outputs a first intermediate feature map of multiple sizes; the first neck network fuses the first intermediate feature maps of multiple sizes corresponding to each sample image to obtain multiple first feature maps.

[0008] In some embodiments, the visual converter projects features representing different heights of the target area onto a two-dimensional BEV grid based on the first feature map to obtain a BEV feature map, including: obtaining a three-dimensional target area based on the first feature map, camera intrinsic parameters and depth information of the target area; projecting features of the three-dimensional target area at preset height intervals onto a two-dimensional BEV grid to obtain multiple BEV intermediate feature maps; and obtaining a BEV feature map after fusing multiple BEV intermediate feature maps.

[0009] In some embodiments, the BEV feature map is sequentially passed through the second backbone network and the second neck network for feature extraction and fusion, and a second feature map is output, including: the second backbone network extracts features from the BEV feature map and outputs second intermediate feature maps of multiple sizes; the second neck network fuses the second intermediate feature maps of multiple sizes to obtain a second feature map.

[0010] In some embodiments, the second feature map is input into the prediction head to obtain a predicted heat map, predicted occupancy information of the grid and the predicted height of the obstacle corresponding to each grid, including: the prediction head determines the predicted heat map through convolution; based on the predicted heat map, determines whether each grid is occupied; for occupied grids, performs layer-by-layer convolution at preset height intervals to find out whether different height layers have features; the height of the highest layer with features is used as the upper edge height of the obstacle in the grid; the height of the lowest layer with features is used as the lower edge height of the obstacle in the grid; based on the upper edge height and the lower edge height, determines the obstacle height corresponding to the grid.

[0011] In some embodiments, the sample labels include three-dimensional information about the obstacles. Adjusting the parameters of the obstacle detection model based on the predicted heat map, the predicted obstacle height corresponding to each grid, and the sample labels includes: obtaining a sample heat map and a sample obstacle height corresponding to each grid based on the three-dimensional obstacle information; and adjusting the parameters of the obstacle detection model based on the difference between the sample heat map and the predicted heat map, the predicted obstacle height corresponding to each grid, and the sample obstacle height.

[0012] In a second aspect, an obstacle detection method is provided, comprising: obtaining a fisheye image; inputting the fisheye image into a trained obstacle detection model to obtain grid occupancy information and the obstacle height corresponding to each grid; wherein the obstacle detection model is trained based on the method of the first aspect above.

[0013] In a third aspect, a training device for an obstacle detection model is provided. The obstacle detection model includes: a first backbone network, a first neck network, a visual converter, a second backbone network, a second neck network and a prediction head. The training device includes: a first acquisition module, configured to acquire sample data, the sample data including a sample image of a target area taken by a fisheye camera and a sample label; a first extraction module, configured to input the sample image into an obstacle detection model, and sequentially pass it through a first backbone network and a first neck network to obtain a first feature map, the first feature map including depth information of the target area; a conversion module, configured to input the first feature map into a visual converter, and the visual converter, based on the first feature map, projects features representing different heights of the target area onto a two-dimensional BEV grid to obtain a BEV feature map; a second extraction module, configured to sequentially pass the BEV feature map through a second backbone network and a second neck network to perform feature extraction and fusion respectively, and output a second feature map; a processing module, configured to input the second feature map into a prediction head to obtain a predicted heat map, predicted occupancy information of the grid and the predicted obstacle height corresponding to each grid; a training module, configured to adjust the parameters of the obstacle detection model based on the predicted heat map, the predicted obstacle height corresponding to each grid and the sample label.

[0014] In a fourth aspect, an obstacle detection device is provided, comprising: a second acquisition module configured to acquire a fisheye image; a prediction module configured to input the fisheye image into a trained obstacle detection model to obtain grid occupancy information and the obstacle height corresponding to each grid; wherein the obstacle detection model is trained based on the method of the first aspect above.

[0015] The training method of the obstacle detection model provided by the present disclosure extracts the first feature information including the depth information of the obstacle from the sample data, projects the features of different heights in the first feature information into the BEV space respectively, obtains the BEV intermediate feature maps corresponding to multiple heights, and splices the BEV intermediate feature maps corresponding to multiple heights to obtain the BEV feature map, thereby achieving the compression of the height information in the first feature information, reducing the amount of calculation in the subsequent process, and being suitable for deployment on a low computing power platform. After predicting the occupancy information of each grid, the height information of the occupied grid can be restored for the occupied grid. Based on the predicted height of the obstacle, the predicted height of the obstacle in the occupied grid can be further predicted, which will not cause the problem of missing obstacle height. Therefore, the detection accuracy of overhanging obstacles or irregular obstacles can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 FIG2 is a flow chart of a method for training an obstacle detection model according to an embodiment of the present disclosure.

[0017] Figure 2 The figure shows a flow chart of the steps of inputting a sample image into an obstacle detection model and sequentially passing the sample image through a first backbone network and a first neck network to obtain a first feature map, provided by an embodiment of the present disclosure.

[0018] Figure 3 The figure shows a flow chart of the steps of projecting features representing different heights of the target area onto a two-dimensional BEV grid based on the first feature map to obtain a BEV feature map provided by a visual converter provided by an embodiment of the present disclosure.

[0019] Figure 4 Shown is a front view of a three-dimensional target area provided by an embodiment of the present disclosure, and a schematic diagram of multiple BEV intermediate feature maps corresponding to the three-dimensional target area.

[0020] Figure 5 The figure shows a flow chart of the steps of extracting and fusing the BEV feature map through the second backbone network and the second neck network respectively, and outputting the second feature map, provided by an embodiment of the present disclosure.

[0021] Figure 6 The figure shows a flow chart of the steps of inputting the second feature map into the prediction head to obtain the predicted heat map, the predicted occupancy information of the grid and the predicted obstacle height corresponding to each grid, provided by an embodiment of the present disclosure.

[0022] Figure 7 FIG2 is a flow chart of the steps for adjusting the parameters of the obstacle detection model based on the predicted heat map, the predicted obstacle height corresponding to each grid, and the sample label, provided by an embodiment of the present disclosure.

[0023] Figure 8 FIG2 is a flow chart of an obstacle detection method provided in accordance with an embodiment of the present disclosure.

[0024] Figure 9 FIG2 is a schematic diagram of the structure of a training device for an obstacle detection model provided in one embodiment of the present disclosure.

[0025] Figure 10 FIG2 is a schematic structural diagram of an obstacle detection device provided in one embodiment of the present disclosure.

[0026] Figure 11 Shown is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0028] Common obstacles in underground parking lots include gates, charging stations, and fire hydrants. Gates typically consist of a main unit and a gate bar, with one side of the gate bar mounted on the main unit, making it an overhanging obstacle. Fire hydrants are typically fixed to walls or pillars, with water pipes connected around them. The ground contact area of the charging station's column is small, and the charging cables hang tangled on the column. Current obstacle detection methods are unable to effectively identify these irregular obstacles with small ground contact areas or completely overhanging obstacles, posing safety risks to the automated parking process.

[0029] In the related art, the Freespace detection method is a common method for detecting drivable areas. The Freespace detection method uses a deep learning network to predict the boundary contour of the drivable area through the environmental image around the vehicle body. The boundary contour is composed of the pixels of the nearest obstacle grounding point in each column on the environmental image. This method has some shortcomings. First, the boundary contour formed by this method has no 3D position information. Second, this method predicts the drivable area based on the obstacle grounding point. It cannot effectively identify irregular obstacles with a small grounding area and a large volume of the ungrounded part, or overhanging obstacles.

[0030] In addition, some vehicles use multiple fisheye cameras to capture fisheye images of the surrounding environment at different angles. In related technologies, the vehicle's controller uses an Around View Monitor (AVM) system to stitch together multiple fisheye images captured simultaneously by multiple fisheye cameras to create an AVM bird's-eye view of the vehicle's surroundings. The AVM bird's-eye view then predicts obstacles in the surrounding environment and determines the boundary outline of the drivable area. However, this method loses information about the height of obstacles, making it more difficult to identify irregular or overhanging obstacles.

[0031] If obstacle detection is performed based on a 2D environmental image, 3D obstacle data is created in 3D space. Although 3D data is richer, it requires 3D convolution for processing and prediction, significantly increasing the amount of computation and requiring higher computing power.

[0032] To address the above technical issues, the present disclosure provides a method and device for training an obstacle detection model, as well as an obstacle detection method and device, to improve the accuracy of identifying irregular or overhanging obstacles, and is applicable to low-computing platforms. To facilitate understanding of the technical solutions provided by the embodiments of the present disclosure, the following detailed description is provided with reference to the accompanying drawings.

[0033] Figure 1 The figure shows a flow chart of a method for training an obstacle detection model provided by an embodiment of the present disclosure. The obstacle detection model includes: a first backbone network, a first neck network, a visual converter, a second backbone network, a second neck network, and a prediction head.

[0034] like Figure 1 As shown, the training method of the obstacle detection model provided in the embodiment of the present disclosure includes the following steps.

[0035] S110, obtaining sample data.

[0036] The sample data includes a sample image of the target area captured by a fisheye camera and a sample label.

[0037] Fisheye cameras, with their extremely wide field of view, can fully perceive the vehicle's entire surroundings with fewer images, bringing significant advantages to autonomous driving applications. The target area is the real environment corresponding to the sample image. This area can include various types of obstacles, as well as background objects other than obstacles, such as the sky and the ground. Sample labels include 3D information about obstacles in the target area, a sample heat map, and the height of each obstacle sample corresponding to each grid cell.

[0038] In the embodiments of the present disclosure, sample data may be collected and labeled by the user, or an open source dataset may be used, such as the Occ3D-nuSences dataset.

[0039] S120: Input the sample image into the obstacle detection model, and pass it through the first backbone network and the first neck network in sequence to obtain a first feature map.

[0040] The sample image is fed into the obstacle detection model. First, the sample image passes through the first backbone network, which extracts multi-level and multi-scale features from the sample image. Next, the features extracted by the first backbone network are fed into the first neck network, which fuses the features extracted by the first backbone network to generate the first feature map.

[0041] The first feature map includes depth information of the target area. The depth information of the target area represents the distance between each feature in the target area and the fisheye camera that captured the sample image. The target area includes features of obstacles and features of background objects other than obstacles. Therefore, the first feature map includes depth information of obstacles and background objects.

[0042] S130, the visual converter projects the features representing different heights of the target area onto a two-dimensional BEV grid based on the first feature map to obtain a BEV feature map.

[0043] The visual converter is used to convert the features of the target area at different heights corresponding to the first feature map into a two-dimensional bird's-eye view (BEV).

[0044] The View Transformer uses Inverse Perspective Mapping (IPM) to project features of the target area at different heights onto a two-dimensional BEV grid.

[0045] Based on the depth information of the target area, the height information of the obstacle features and background object features in the first feature map can be determined. Based on the height information of the obstacle features and background object features in the first feature map, the first feature map can be restored to a three-dimensional feature corresponding to the target area.

[0046] In the three-dimensional features corresponding to the target area, the features representing different heights of the target area are projected onto the two-dimensional BEV grid according to the feature height coordinates, resulting in multiple BEV intermediate feature maps corresponding to each height. Next, the BEV intermediate feature maps corresponding to each height are fused to obtain a BEV feature map. Exemplarily, the above projection process can utilize the inverse perspective mapping (IPM) method.

[0047] In this process, the height information in the three-dimensional features is compressed to obtain a two-dimensional BEV feature map.

[0048] S140, the BEV feature map is sequentially passed through the second backbone network and the second neck network for feature extraction and fusion, and a second feature map is output.

[0049] The second backbone network is used to extract features from the BEV feature map, and the second neck network is used to fuse the features extracted by the second backbone network to obtain a second feature map.

[0050] S150: Input the second feature map into the prediction head to obtain a predicted heat map, predicted occupancy information of the grid, and predicted obstacle height corresponding to each grid.

[0051] The prediction heatmap is essentially a probability distribution map. The value of each position in the prediction heatmap represents the probability of that position being occupied. This occupancy probability is used to predict the grid occupancy information.

[0052] The predicted grid occupancy information includes information about whether each grid in the feature map formed by convolving the second feature map with the prediction head is occupied. For each grid, if a grid is occupied, it indicates that there is an obstacle in the grid; if a grid is unoccupied, it indicates that there is no obstacle in the grid.

[0053] For occupied grids, the prediction head can further predict the height of obstacles in the grids to obtain the predicted obstacle height corresponding to each grid.

[0054] Since the second feature map is two-dimensional, the prediction head uses lightweight two-dimensional convolution, which greatly reduces the computational complexity of the prediction process and is suitable for low-computing power platforms.

[0055] Before inputting the second feature map into the prediction head, the second feature map may be upsampled first, and the upsampled second feature map may be input into the prediction head, so as to improve the resolution of the second feature map.

[0056] S160 , adjusting the parameters of the obstacle detection model based on the predicted heat map, the predicted obstacle height corresponding to each grid, and the sample label.

[0057] There are two forms of sample labels. One form is: the three-dimensional information of the obstacles in the target area, and the sample heat map and the obstacle sample height corresponding to each grid can be calculated based on the three-dimensional information of the obstacles. The other form is: the sample heat map and the obstacle sample height corresponding to each grid are obtained directly based on the open source database. Regardless of which form of sample label is based on the above, the parameters of the obstacle detection model are adjusted based on the difference between the sample heat map and the predicted heat map, the predicted obstacle height and the obstacle sample height corresponding to each grid. The obstacle detection model is trained based on the above training method until the training termination conditions are met. Among them, the training termination conditions include the heat map loss, the obstacle height loss corresponding to each grid is less than the preset value, or the number of training iterations reaches a preset number, then the training process is stopped.

[0058] In the embodiment of the present disclosure, a first feature map including depth information of obstacles is extracted from a sample image, and the features of the target area at different heights corresponding to the first feature map are converted to a two-dimensional BEV grid to obtain a plurality of BEV intermediate feature maps corresponding to each height, and the BEV intermediate feature maps corresponding to each height are fused to obtain a BEV feature map, thereby achieving compression of the height information in the first feature map. When the prediction head uses the processed BEV feature map (the second feature map) for prediction, compared to the prediction process of using three-dimensional convolution to process three-dimensional spatial data, the embodiment of the present disclosure greatly reduces the amount of computation required for subsequent processes by compressing the three-dimensional space into a two-dimensional BEV grid, making it suitable for deployment on low-computing power platforms (such as low-computing power chips).

[0059] Because the BEV feature map is obtained by fusing multiple BEV intermediate feature maps corresponding to different heights, after predicting the occupancy information of each grid, the height information of the occupied grid can be restored. Based on the obstacle height initially predicted by the first feature map, the corresponding predicted obstacle height in the occupied grid can be further predicted, eliminating the problem of missing obstacle heights. This improves the detection accuracy of overhanging obstacles or irregular obstacles.

[0060] In some embodiments, the sample images include multiple sample images captured simultaneously by multiple fisheye cameras, and the fields of view of at least two of the fisheye cameras overlap. Due to the overlapping fields of view of the fisheye cameras, the target areas corresponding to the fisheye images captured by the fisheye cameras also overlap.

[0061] Since vehicles typically carry multiple fisheye cameras, each covering a 360-degree field of view around the vehicle, multiple sample images can be used to simulate the multiple fisheye images captured by these cameras during parking. Training the obstacle detection model based on these sample images improves its ability to process multiple fisheye images.

[0062] The following describes in detail the specific implementation of the obstacle detection model to process multiple sample images.

[0063] Figure 2 FIG. 1 is a flow chart of the steps of inputting a sample image into an obstacle detection model and sequentially passing through a first backbone network and a first neck network to obtain a first feature map according to an embodiment of the present disclosure. Figure 2 As shown, the sample image is input into the obstacle detection model, and the first feature map is obtained by sequentially passing through the first backbone network and the first neck network, which includes the following steps.

[0064] S121: The first backbone network extracts features from multiple sample images, and outputs first intermediate feature maps of multiple sizes for each sample image.

[0065] Input multiple sample images into the first backbone network. For example, four sample images may be input, and the four sample images correspond to fisheye images captured by fisheye cameras respectively located on the front, rear, left, and right sides of the vehicle.

[0066] The first backbone network comprises a series of convolutional layers, pooling layers, and activation functions. The first backbone network extracts features from each of the multiple sample images, outputting first intermediate feature maps of multiple sizes for each sample image. For example, each sample image can output three first intermediate feature maps of different sizes. The first backbone network can utilize a DLANet34 network.

[0067] Because the multiple sample images are captured by different fisheye cameras, they can be preprocessed before feature extraction. This preprocessing may include converting the multiple sample images to floating-point data, resizing them to the same size, and normalizing them based on the number of channels.

[0068] S122: The first neck network fuses the first intermediate feature maps of multiple sizes corresponding to each channel of sample images to obtain multiple first feature maps.

[0069] Next, the first intermediate feature maps of multiple sizes for each sample image are input into the first neck network. The first neck network fuses the first intermediate feature maps of multiple sizes corresponding to each sample image to obtain the first feature map for each sample image. For example, the first neck network can employ an FPN-LSS network.

[0070] In the disclosed embodiment, feature extraction and fusion operations are performed on each sample image to obtain the first feature map corresponding to each path. The first intermediate feature maps of different sizes contain semantic information and detail information at different levels. Among them, the low-level features extracted in the shallow layer usually include the edges and textures of the sample image, while the high-level features extracted in the deep layer usually include the shape and category of the obstacles in the sample image. Therefore, multiple first feature maps are obtained by the first backbone network and the first neck network, which can provide rich information for the subsequent prediction process.

[0071] Figure 3 The figure shows a flow chart of the steps of projecting the features representing different heights of the target area onto a two-dimensional BEV grid based on the first feature map to obtain a BEV feature map provided by a visual converter according to an embodiment of the present disclosure. Figure 3As shown, the visual converter projects the features representing different heights of the target area onto a two-dimensional BEV grid based on the first feature map to obtain a BEV feature map, which includes the following steps.

[0072] S131: Obtain a three-dimensional target area based on the first feature map, the camera intrinsic parameter, and the depth information of the target area.

[0073] The depth information of the target area includes the depth information of each feature in the target area, for example, the depth information of each obstacle.

[0074] Based on the camera intrinsic parameters and the depth information of the target area, the target area corresponding to the two-dimensional first feature map can be converted into a three-dimensional target area. The three-dimensional target area includes the three-dimensional features of the obstacle.

[0075] Exemplarily, the target area includes a street lamp, and the three-dimensional target area includes the three-dimensional features of the street lamp.

[0076] S132 , projecting the features of the three-dimensional target area at every preset height interval onto a two-dimensional BEV grid to obtain a plurality of BEV intermediate feature maps.

[0077] The 3D object region is segmented along the height direction at preset height intervals to produce multiple layers of 3D object subregions. Each 3D object subregion corresponds to a portion of the 3D object region within a certain height range. During the segmentation process, features within the 3D object region that fall within that height range are also segmented, and each 3D object subregion contains slices of the features that fall within that height range.

[0078] The upper section or lower section of each three-dimensional target sub-region is projected onto the two-dimensional BEV grid to obtain the BEV intermediate feature map corresponding to each three-dimensional target sub-region.

[0079] Figure 4 The figure shows a front view of a three-dimensional target area provided by an embodiment of the present disclosure, and a schematic diagram of a plurality of BEV intermediate feature maps corresponding to the three-dimensional target area. Figure 4 As shown in (a), continuing to take the street lamp as an example, the three-dimensional target area includes the features of the street lamp 40. The three-dimensional target area is divided along the dividing line (shown by the dotted line), and the distance between adjacent dividing lines is a preset height interval. In this way, the features of the street lamp 40 are also divided. The street lamp 40 includes a lamp 41 and a lamp post 42. The street lamp 40 is divided into multiple parts by the dividing lines with heights H1, H2, ..., Hn. Then, the cross-sections of the three-dimensional target area at the height positions of height H1, height H2, ..., and height Hn are projected onto the two-dimensional BEV grid to obtain the BEV intermediate feature maps corresponding to heights H1, height H2, ..., and height Hn.

[0080] Take the BEV intermediate feature maps corresponding to height H1, height H2, height H3, and height H4 as an example. Figure 4 (b) shows the BEV intermediate feature map corresponding to height H1, as shown in Figure 4 As shown in (b), the features of lamp 41 occupy two grids, and the coordinates corresponding to the two grids are (1,2) and (2,2) respectively; the features of lamp post 42 occupy one grid, and the coordinates are (3,2). Figure 4 (c) shows the BEV intermediate feature map corresponding to height H2, as shown in Figure 4 As shown in (c), the features of lamp 41 occupy two grids, and the coordinates corresponding to the two grids are (1,2) and (2,2) respectively; the features of lamp post 42 occupy one grid, and the coordinates are (3,2). Figure 4 (d) shows the BEV intermediate feature map corresponding to height H3, as shown in Figure 4 As shown in (d), the features of lamp 41 occupy a grid with coordinates (1, 2); the features of lamp post 42 occupy a grid with coordinates (3, 2). Figure 4 (e) shows the BEV intermediate feature map corresponding to height H4, as shown in Figure 4 As shown in (e), the feature of lamp post 42 occupies a grid with coordinates (3,2).

[0081] In this way, the three-dimensional target area is segmented along the height direction, and the upper or lower section of each layer is projected onto the two-dimensional BEV grid to obtain multiple BEV intermediate feature maps.

[0082] It is understandable that the preset height interval can be set by yourself. The smaller the value of the preset height interval, the more accurate the height regression of the obstacle corresponding to each grid will be. However, the amount of calculation will also increase. This disclosure does not impose any specific restrictions on this, and it can be set specifically in combination with actual needs. Figure 4 The number of grids, grid size, etc. are also merely illustrative and should not be construed as limiting the present disclosure.

[0083] The above process can be implemented using the inverse perspective mapping method, and the specific calculation method will not be repeated here. The inverse perspective mapping method requires less computing power, so more computing power can be reserved for subsequent prediction steps, making the obstacle detection model trained by the training method of the embodiment of the present disclosure suitable for deployment on low-computing power platforms.

[0084] S133: After fusing multiple BEV intermediate feature maps, a BEV feature map is obtained.

[0085] For multiple BEV intermediate feature maps, the multiple BEV intermediate feature maps are fused respectively to obtain a BEV feature map.

[0086] The above is for a sample image. Based on the first feature map, the camera intrinsic parameters and the depth information of the target area, a three-dimensional target area is obtained. The features of the three-dimensional target area at preset height intervals are projected onto the two-dimensional BEV grid to obtain multiple BEV intermediate feature maps; after fusing the multiple BEV intermediate feature maps, a BEV feature map is obtained.

[0087] For multiple sample images, a BEV feature map corresponding to each sample image can be obtained, that is, multiple BEV feature maps can be obtained. Subsequently, the multiple BEV feature maps can be fused and input into the second backbone network for subsequent steps.

[0088] Through the above method, the first feature map is restored to the three-dimensional target area, and all the features of the cross-sections at different heights in the three-dimensional target area are projected onto the two-dimensional BEV grid respectively, thereby achieving compression of the three-dimensional target area in the height dimension while retaining the features at different heights, making the subsequent prediction process more efficient and reducing the computational complexity of the prediction process. At the same time, the height prediction of obstacles is also more accurate.

[0089] Figure 5 The figure shows a flow chart of the steps of extracting and fusing the BEV feature map through the second backbone network and the second neck network respectively, and outputting the second feature map according to an embodiment of the present disclosure. Figure 5 As shown, the BEV feature map is sequentially passed through the second backbone network and the second neck network for feature extraction and fusion, and the second feature map step is output, which includes the following steps.

[0090] S141: The second backbone network extracts features from the BEV feature map and outputs second intermediate feature maps of multiple sizes.

[0091] The BEV feature map is input into the second backbone network. The second backbone network comprises a series of convolutional layers, pooling layers, and activation functions. The second backbone network extracts features from the BEV feature map and outputs second intermediate feature maps of multiple sizes. For example, two second intermediate feature maps of different sizes can be output. The second backbone network can employ a ResNet network.

[0092] If the sample images include multiple sample images captured by multiple fisheye cameras at the same time, it is necessary to fuse the BEV feature maps corresponding to the multiple sample images into an overall BEV feature map, and input the overall BEV feature map into the second backbone network.

[0093] S142: The second neck network fuses the second intermediate feature maps of multiple sizes to obtain a second feature map.

[0094] Next, the second intermediate feature maps of multiple sizes are fed into the second neck network. The second neck network fuses the second intermediate feature maps of multiple sizes to generate the second feature map. For example, the second neck network can employ an FPN-LSS network. The design features and simplicity of the FPN-LSS network offer significant advantages in terms of computational resource requirements, making it suitable for deployment on low-computing platforms.

[0095] In the disclosed embodiment, the BEV feature map is subjected to feature extraction through the second backbone network to obtain second intermediate feature maps of multiple sizes containing semantic information and detail information at different levels, which can provide rich information for the subsequent prediction process.

[0096] Figure 6 The figure shows a flow chart of the steps of inputting the second feature map into the prediction head to obtain the predicted heat map, the predicted occupancy information of the grid and the predicted obstacle height corresponding to each grid, provided by an embodiment of the present disclosure. Figure 6 As shown, the second feature map is input into the prediction head to obtain the predicted heat map, the predicted occupancy information of the grid and the predicted obstacle height corresponding to each grid, including the following steps.

[0097] S151, the prediction head determines the prediction heat map through convolution.

[0098] The second feature map is input into the prediction head, which includes a convolutional layer. The convolutional layer performs a convolution operation on the second feature map to capture the local correlation and hierarchical structure in the second feature map and predict the probability of each position being occupied, ultimately generating a predicted heat map. The value in the predicted heat map represents the probability of the position being occupied.

[0099] S152: Determine whether each grid is occupied based on the predicted heat map.

[0100] In the predicted heat map, if the probability of a location being occupied is greater than a preset probability threshold, the location is determined to be occupied; if the probability of a location being occupied is less than the preset probability threshold, the location is determined to be unoccupied. For each grid, if the proportion of occupied locations in the area included in the grid to the entire grid area is greater than the preset occupancy threshold, the grid is determined to be occupied; if the proportion of occupied locations in the area included in the grid to the entire grid area is less than the preset occupancy threshold, the grid is determined to be unoccupied.

[0101] In this way, we can determine whether each grid is occupied. If the grid is occupied, it means there is an obstacle in the grid; if the grid is not occupied, it means there is no obstacle in the grid, or the obstacle ratio is small.

[0102] S153: For the occupied grid, perform convolution layer by layer at every preset height interval to find out whether different height layers have features.

[0103] For each occupied grid, convolution is performed layer by layer at preset height intervals to find out whether there are features at the corresponding positions of the occupied grids at different height layers.

[0104] For example, the corresponding coordinates of the occupied grid are first determined. Next, the coordinates are checked to see if there are any features at different altitude levels. If there are features, it indicates that an obstacle exists at that coordinate location at that altitude level; if there are no features, it indicates that there is no obstacle at that coordinate location at that altitude level. This allows the height of the obstacle at that grid location to be determined.

[0105] S154: The height of the highest layer with features is used as the upper edge height of the obstacle in the grid.

[0106] S155: The height of the lowest layer with features is used as the lower edge height of the obstacle in the grid.

[0107] S156: Determine the obstacle height corresponding to the grid based on the upper edge height and the lower edge height.

[0108] The upper and lower heights of obstacles in the occupied grid can be determined. The height of the obstacle is determined based on the upper and lower heights. Specifically, when determining whether different altitude layers have characteristics, the highest altitude layer with characteristics is determined, and the height of the highest altitude layer is used as the upper height of the obstacle in the grid. The lowest altitude layer with characteristics is determined, and the height of the lowest altitude layer is used as the lower height of the obstacle in the grid.

[0109] Continue to refer Figure 4 In the grid with coordinates (1,2), convolution is performed on different height layers from H1 to Hn respectively, and the layer is searched layer by layer to see whether it has features. Through convolution, it can be seen that the highest layer with features is at a height of H1, and the lowest layer with features is at a height of H3. Therefore, in the (1,2) grid, the corresponding obstacle height is H1 minus H3; in the grid with coordinates (2,2), the highest layer with features is at a height of H1, and the lowest layer with features is at a height of H2. Therefore, in the (2,2) grid, the corresponding obstacle height is H1 minus H2; in the grid with coordinates (3,2), the highest layer with features is at a height of H1, and the lowest layer with features is at a height of Hn. Therefore, in the (3,2) grid, the corresponding obstacle height is H1 minus Hn.

[0110] In this disclosed embodiment, by predicting the top and bottom heights of obstacles in occupied grid cells, the complex shape of the obstacle is broken down into simple numerical representations, reducing the computational complexity of the prediction process. Furthermore, based on the bottom height of the obstacle, it is possible to determine whether the obstacle is in contact with the ground and its actual height above the ground, improving the ability to identify overhanging obstacles.

[0111] Figure 7 FIG. 1 is a flow chart of the steps for adjusting the parameters of the obstacle detection model based on the predicted heat map, the predicted obstacle height corresponding to each grid, and the sample label provided by an embodiment of the present disclosure. Figure 7 As shown in FIG, based on the predicted heat map, the predicted obstacle height corresponding to each grid, and the sample label, the parameter steps of the obstacle detection model are adjusted, including the following steps.

[0112] S161: Based on the three-dimensional information of the obstacle, a sample heat map and the obstacle sample height corresponding to each grid are obtained.

[0113] In one embodiment, the sample label includes three-dimensional information of obstacles in the target area, and the sample heat map is calculated based on the above information.

[0114] Next, the sample heat map is divided into multiple grids in the form of a two-dimensional BEV grid, and the obstacle sample height corresponding to each grid in the multiple grids is determined based on the height of the obstacle.

[0115] Exemplarily, the three-dimensional information of the obstacle is determined based on radar point cloud data of the target area, and the sample data is determined by combining the radar point cloud data with the sample image.

[0116] In another embodiment, the sample labels may also be directly based on the sample heat map obtained from the open source database and the obstacle sample height corresponding to each grid.

[0117] S162: Adjust the parameters of the obstacle detection model based on the difference between the sample heat map and the predicted heat map, the predicted obstacle height corresponding to each grid, and the obstacle sample height.

[0118] Next, a heatmap loss is determined based on the difference between the sample heatmap and the predicted heatmap output by the prediction head. For example, the difference between the sample heatmap and the predicted heatmap can be calculated based on a loss function. The loss function can be a CrossEntropyLoss function.

[0119] Based on the obstacle sample height corresponding to each grid and the predicted obstacle height corresponding to each grid output by the prediction head, a height loss is determined. For example, the difference between the predicted obstacle height and the obstacle sample height corresponding to each grid can be calculated based on a height loss function. The height loss function can adopt the SmoothL1Loss function.

[0120] Heatmap loss and height loss measure the difference between the obstacle detection model's predictions and the actual obstacle situation, based on the obstacle's location and height, respectively. Smaller heatmap loss and height loss indicate a closer match between the obstacle detection model's predictions and the actual situation, indicating better performance.

[0121] If the heat map loss and height loss do not meet the training termination conditions, the parameters of the obstacle detection model are adjusted based on the heat map loss and height loss.

[0122] The training process of the obstacle detection model continuously repeats the above steps. As the number of iterations increases, the parameters of the obstacle detection model are gradually optimized, the values of the heat map loss and height loss gradually decrease, and the prediction performance of the obstacle detection model gradually improves.

[0123] Combined with the above Figures 1 to 7 The training method embodiment of the obstacle detection model disclosed in the present invention is described in detail. Figure 8 The obstacle detection method embodiment of the present disclosure is described in detail. It should be understood that the description of the training method embodiment corresponds to the description of the obstacle detection method embodiment, so for parts not described in detail, reference can be made to the previous training method embodiment.

[0124] Figure 8 FIG. 1 is a flow chart of an obstacle detection method according to an embodiment of the present disclosure. Figure 8 As shown, the obstacle method provided in the embodiment of the present disclosure includes the following steps.

[0125] S810: Acquire a fisheye image.

[0126] Fisheye images are captured by a fisheye camera mounted on a vehicle.

[0127] There can be multiple fisheye cameras, which are arranged around the vehicle body and capture images of the surrounding environment of the vehicle body from different perspectives.

[0128] S820: Input the fisheye image into the trained obstacle detection model to obtain grid occupancy information and the obstacle height corresponding to each grid.

[0129] The obstacle detection model is trained based on the obstacle detection model training method in the above embodiment.

[0130] The obstacle detection model processes the fisheye image to obtain grid occupancy information and the corresponding obstacle height for each grid. This information can be used to determine the location and height of obstacles in the corresponding area of the fisheye image, thereby obtaining obstacle detection results.

[0131] The following combination Figure 9 The training device embodiment of the obstacle detection model disclosed in the present invention is described in detail. It should be understood that the description of the training method embodiment corresponds to the description of the training device embodiment, so that the parts not described in detail can be referred to the above training method embodiment.

[0132] Figure 9 FIG. 1 is a schematic diagram of the structure of a training device for an obstacle detection model provided by an embodiment of the present disclosure. Figure 9 As shown, the obstacle detection model training device 900 of the embodiment of the present disclosure includes: a first acquisition module 910, a first extraction module 920, a conversion module 930, a second extraction module 940, a processing module 950 and a training module 960.

[0133] Specifically, the obstacle detection model includes: a first backbone network, a first neck network, a visual converter, a second backbone network, a second neck network and a prediction head.

[0134] The first acquisition module 910 is configured to acquire sample data, which includes a sample image of the target area captured by a fisheye camera and sample labels. The first extraction module 920 is configured to input the sample image into the obstacle detection model, sequentially passing it through the first backbone network and the first neck network to obtain a first feature map, which includes depth information of the target area. The conversion module 930 is configured to input the first feature map into a visual transformer. Based on the first feature map, the visual transformer projects features representing different heights of the target area onto a two-dimensional BEV grid to obtain a BEV feature map. The second extraction module 940 is configured to sequentially pass the BEV feature map through the second backbone network and the second neck network, respectively, to perform feature extraction and fusion, and output a second feature map. The processing module 950 is configured to input the second feature map into the prediction head to obtain a predicted heatmap, predicted occupancy information for each grid, and the predicted obstacle height corresponding to each grid. The training module 960 is configured to adjust the parameters of the obstacle detection model based on the predicted heatmap, the predicted obstacle height corresponding to each grid, and the sample labels.

[0135] The following combination Figure 10The obstacle detection device embodiments of the present disclosure are described in detail. It should be understood that the descriptions of the obstacle detection model training method embodiments and obstacle detection method embodiments correspond to the descriptions of the obstacle detection device embodiments. Therefore, for portions not described in detail, reference can be made to the aforementioned obstacle detection model training method embodiments and obstacle detection method embodiments.

[0136] like Figure 10 As shown, the obstacle detection device 1000 according to the embodiment of the present disclosure includes: a second acquisition module 1010 and a prediction module 1020 .

[0137] The second acquisition module 1010 is configured to acquire a fisheye image. The prediction module 1020 is configured to input the fisheye image into a trained obstacle detection model to obtain grid occupancy information and the obstacle height corresponding to each grid. The obstacle detection model is trained based on the obstacle detection model training method described in the above embodiment.

[0138] Below, reference Figure 11 An electronic device according to an embodiment of the present disclosure is described. Figure 11 FIG. 1 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 11 As shown, electronic device 1100 includes one or more processors 1110 and memory 1120 .

[0139] The processor 1110 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1100 to perform desired functions.

[0140] Memory 1120 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and processor 1110 may execute the program instructions to implement the obstacle detection model training method, obstacle detection method, and / or other desired functions described above in various embodiments of the present disclosure.

[0141] In some embodiments, the electronic device 1100 may further include an input device 1130 and an output device 1140 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0142] The input device 1130 may include, for example, a touch screen, a microphone, a keyboard, a mouse, etc. The output device 1140 may include, for example, a display, a speaker, a communication network and a remote output device connected thereto.

[0143] Of course, to simplify, Figure 11 Only some of the components related to the present disclosure in the electronic device 1100 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 1100 may further include any other appropriate components.

[0144] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the obstacle detection model training method or obstacle detection method according to various embodiments of the present disclosure described above in this specification.

[0145] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0146] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, cause the processor to execute the steps of the obstacle detection model training method or obstacle detection method according to various embodiments of the present disclosure described above in this specification.

[0147] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, but is not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0148] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0149] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0150] It should also be noted that in the systems, devices, and methods of the present disclosure, each component or each step can be decomposed and / or recombined, and such decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.

[0151] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0152] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A method for training an obstacle detection model, the obstacle detection model comprising: A first backbone network, a first neck network, a visual converter, a second backbone network, a second neck network, and a prediction head, characterized by comprising: Acquire sample data, where the sample data includes a sample image of a target area captured by a fisheye camera and a sample label; Inputting the sample image into the obstacle detection model, and sequentially passing the sample image through the first backbone network and the first neck network to obtain a first feature map, wherein the first feature map includes depth information of the target area; The visual converter projects features representing different heights of the target area onto a two-dimensional BEV grid based on the first feature map to obtain a BEV feature map; Passing the BEV feature map through the second backbone network and the second neck network in sequence to perform feature extraction and fusion respectively, and outputting a second feature map; Inputting the second feature map into the prediction head to obtain a predicted heat map, predicted occupancy information of the grid, and predicted obstacle height corresponding to each grid; Adjust the parameters of the obstacle detection model based on the predicted heat map, the predicted obstacle height corresponding to each grid, and the sample label.

2. The method according to claim 1, characterized in that The sample images include multiple sample images captured by multiple fisheye cameras at the same time, and the shooting fields of at least two fisheye cameras have overlapping areas.

3. The method according to claim 2, characterized in that The step of inputting the sample image into the obstacle detection model and sequentially passing the sample image through the first backbone network and the first neck network to obtain a first feature map includes: The first backbone network performs feature extraction on the multiple sample images, and outputs a first intermediate feature map of multiple sizes for each sample image; The first neck network fuses the first intermediate feature maps of the multiple sizes corresponding to each channel of sample images to obtain multiple first feature maps.

4. The method according to claim 1, wherein The visual converter projects features representing different heights of the target area onto a two-dimensional BEV grid based on the first feature map to obtain a BEV feature map, including: Obtaining a three-dimensional target area based on the first feature map, the camera intrinsic parameter, and the depth information of the target area; Projecting the features of the three-dimensional target area at preset height intervals onto the two-dimensional BEV grid to obtain a plurality of BEV intermediate feature maps; After fusing the multiple BEV intermediate feature maps, a BEV feature map is obtained.

5. The method according to claim 1, wherein The step of sequentially passing the BEV feature map through the second backbone network and the second neck network to extract and fuse the features, and outputting a second feature map, comprises: The second backbone network performs feature extraction on the BEV feature map and outputs second intermediate feature maps of multiple sizes; The second neck network fuses the second intermediate feature maps of the multiple sizes to obtain a second feature map.

6. The method according to claim 1, characterized in that Inputting the second feature map into the prediction head to obtain a predicted heat map, predicted occupancy information of the grid, and predicted obstacle height corresponding to each grid, includes: The prediction head determines a prediction heat map through convolution; Determining whether each grid is occupied based on the predicted heat map; For occupied grids, convolution is performed layer by layer at preset height intervals to find out whether different height layers have features; The height of the highest layer with features is taken as the upper edge height of the obstacle in the grid; The height of the lowest layer with features is taken as the lower edge height of the obstacle in the grid; Based on the upper edge height and the lower edge height, the obstacle height corresponding to the grid is determined.

7. The method according to claim 1, characterized in that The sample label includes three-dimensional information of the obstacle, and adjusting the parameters of the obstacle detection model based on the predicted heat map, the predicted obstacle height corresponding to each grid, and the sample label includes: Based on the three-dimensional information of the obstacle, a sample heat map and the obstacle sample height corresponding to each grid are obtained; The parameters of the obstacle detection model are adjusted based on the difference between the sample heat map and the predicted heat map, the predicted obstacle height corresponding to each grid, and the obstacle sample height.

8. An obstacle detection method, characterized in that: include: Get fisheye image; Input the fisheye image into a trained obstacle detection model to obtain grid occupancy information and the obstacle height corresponding to each grid; The obstacle detection model is trained based on the method described in any one of claims 1 to 7.

9. A training device for an obstacle detection model, the obstacle detection model comprising: A first backbone network, a first neck network, a visual converter, a second backbone network, a second neck network and a prediction head, wherein the training device comprises: A first acquisition module is configured to acquire sample data, wherein the sample data includes a sample image of a target area captured by a fisheye camera and a sample label; a first extraction module, configured to input the sample image into an obstacle detection model, and sequentially pass the sample image through the first backbone network and the first neck network to obtain a first feature map, wherein the first feature map includes depth information of the target area; a conversion module configured to input the first feature map into a visual converter, and the visual converter projects features representing different heights of the target area onto a two-dimensional BEV grid based on the first feature map to obtain a BEV feature map; A second extraction module is configured to sequentially pass the BEV feature map through the second backbone network and the second neck network to perform feature extraction and fusion respectively, and output a second feature map; a processing module configured to input the second feature map into the prediction head to obtain a predicted heat map, predicted occupancy information of the grids, and predicted obstacle heights corresponding to each grid; The training module is configured to adjust the parameters of the obstacle detection model based on the predicted heat map, the predicted obstacle height corresponding to each grid, and the sample label.

10. An obstacle detection device, characterized in that: include: A second acquisition module is configured to acquire a fisheye image; a prediction module configured to input the fisheye image into a trained obstacle detection model to obtain grid occupancy information and obstacle height corresponding to each grid; The obstacle detection model is trained based on the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • 3D target detection system and method based on 4D millimeter wave radar and camera fusion

    CN117452396A

  • Systems and methods for generating a road surface semantic segmentation map from a sequence of point clouds

    US20230267615A1