Obstacle prediction methods, software products, and electronic devices
By generating planar feature fusion images to explicitly express feature point height information, the problem of in-vehicle equipment being unable to accurately predict suspended obstacles is solved, thus improving parking safety.
Patent Information
- Application Number
- CN202511240411.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-01
AI Technical Summary
Existing vehicle-mounted equipment cannot accurately predict suspended obstacles when parking, resulting in a high risk of collision and affecting parking safety.
By acquiring multiple images from different perspectives, the three-dimensional coordinate data of feature points are determined, and a planar feature fusion image is generated to explicitly express the height information of feature points for obstacle prediction.
It improves the accuracy of predicting suspended obstacles, reduces the probability of vehicles colliding with suspended obstacles, and enhances parking safety.
Smart Images

Figure CN120808311B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to an obstacle prediction method, program product, and electronic device. Background Technology
[0002] In parking scenarios, vehicles typically need onboard equipment to predict the occupancy status of obstacles in the surrounding environment (such as other vehicles, walls, pillars, etc.) to ensure parking safety.
[0003] However, current in-vehicle equipment does not adequately handle certain obstacles (such as suspended obstacles like fire extinguisher boxes on pillars or parking lot entrance / exit barriers) when predicting the surrounding environment. This results in low accuracy in predicting suspended obstacles, increasing the risk of collisions during parking and compromising parking safety, leading to a poor user experience. Summary of the Invention
[0004] This application provides an obstacle prediction method, a program product, and an electronic device.
[0005] In a first aspect, embodiments of this application provide an obstacle prediction method applied to an electronic device. The method includes: the electronic device acquiring multiple images, wherein the multiple images include images from different perspectives collected by sensors of a vehicle. The electronic device determines first feature data of feature points in each of the multiple images, wherein the first feature data includes first coordinate data of the feature points in a vehicle coordinate system, and the first coordinate data includes the height value of the feature points. For example, the first coordinate data is three-dimensional data of the feature points in the vehicle coordinate system. Based on the first feature data of each of the multiple images, the electronic device generates a planar feature fusion image corresponding to the multiple images, wherein the feature points included in the planar feature fusion image are: mapped feature points of each image in the multiple images onto a horizontal plane in the vehicle coordinate system. The electronic device can generate feature data of the planar feature fusion image based on the first feature data of each of the multiple images, wherein the feature data of the planar feature fusion image includes second feature data, the second feature data including: region coordinate data of each sub-region obtained after dividing the planar feature fusion image into multiple sub-regions, and the height range of the feature points in each sub-region in the vehicle coordinate system. Then, the electronic device can perform obstacle prediction based on the feature data of the planar feature fusion image.
[0006] In some embodiments of this application, the electronic device can map the first coordinate data of feature points in each image in the vehicle coordinate system onto the horizontal plane of the vehicle coordinate system to generate a planar feature fusion image. Then, based on the first feature data of each feature point, the feature data of each mapped feature point on the planar feature fusion image is determined. The feature data includes the second coordinate data of each mapped feature point and the height range data of each feature point on the corresponding image in the vehicle coordinate system. Therefore, the planar feature fusion image pays close attention to the height information of each feature point in the vehicle coordinate system, so that each feature point has more or more accurate height information when the electronic device performs obstacle prediction based on the planar feature fusion image. This ensures that the electronic device can more accurately predict the occupancy of suspended obstacles, thereby improving the safety of vehicle driving or parking.
[0007] In one possible implementation of the first aspect above, the feature data of the planar feature fusion image further includes third feature data, and the method further includes: generating third feature data of the planar feature fusion image based on the first feature data of each image in the multiple images, wherein the third feature data includes: second coordinate data and first encoding data of each mapped feature point in the planar feature fusion image; the second coordinate data is two-dimensional coordinate data that maps the first coordinate data of the feature points to the horizontal plane of the vehicle coordinate system, and the first encoding data is the combined data of visual information and height data of the feature points mapped to the same second coordinate data.
[0008] In some embodiments of this application, the electronic device can further merge the height data and visual information of the first coordinate data in the first feature data, thereby mapping them onto a planar feature fusion image, while retaining the three-dimensional coordinate information of the feature points. Therefore, when the electronic device performs obstacle prediction based on the planar feature fusion image, it can improve the prediction accuracy of the overall environment.
[0009] In one possible implementation of the first aspect above, it includes: generating fourth feature data of a planar feature fusion image based on first feature data of each image in multiple images, wherein the fourth feature data includes: second coordinate data and second encoding data of each mapped feature point in the planar feature fusion image; the second encoding data is fusion data of visual information of feature points mapped to the same second coordinate data.
[0010] In one possible implementation of the first aspect above, obstacle prediction based on feature data of a planar feature fusion image includes: fusing second feature data, third feature data, and fourth feature data into fifth feature data; and performing obstacle prediction based on the fifth feature data.
[0011] In the embodiments of this application, the second feature data can enhance the height information in the fifth feature data, the third feature data can enable the fifth feature data to retain more environmental information, and the fourth feature data can enable the fifth feature data to retain more accurate visual information, thereby ensuring both the accuracy of the fifth feature data in predicting the overall environment and the accuracy of the fifth feature data in predicting suspended obstacles.
[0012] In one possible implementation of the first aspect described above, the electronic device is further configured with a first model, which is trained based on the third coordinate data of the first sample in the vehicle coordinate system and the fourth coordinate data of the second sample mapped to the vehicle coordinate system in the following manner: adjusting the fourth coordinate data based on the first parameter to obtain first adjusted data; obtaining the similarity between the first adjusted data and the third coordinate data; based on the similarity between the first adjusted data and the third coordinate data, correcting the first parameter to a second parameter so that the similarity between the fourth coordinate data adjusted by the second parameter and the third coordinate data is greater than or equal to a similarity threshold; using the second parameter as the model parameter of the first model; wherein, the first sample is point cloud data collected by the test vehicle in the first environment, and the second sample is feature data of feature points in multiple image samples collected by the test vehicle in the first environment; the electronic device determines the height range of the feature points in the vehicle coordinate system corresponding to each sub-region in the following manner: inputting the first coordinate data of the feature points into the first model to adjust the height data of the first coordinate data of the feature points; determining the height range of the feature points in the vehicle coordinate system corresponding to each sub-region based on the adjusted first coordinate data of the feature points.
[0013] In some embodiments of this application, the point cloud sample data and image samples can be collected by a data acquisition vehicle. For example, sensors in the data acquisition vehicle may include cameras and LiDAR. The camera can obtain image samples of the environment surrounding the data acquisition vehicle, and the LiDAR can obtain point cloud data of the environment surrounding the data acquisition vehicle. Then, a server, or other devices, can train a first model based on the point cloud sample data and image samples. The point cloud sample data can be used as real data. The server can process the image samples to determine the height range sample data in the second feature data from the image samples. Then, the server can train the model parameters of the first model based on the height range sample data and the point cloud data until the similarity between the height range sample data adjusted by the trained second parameters and the point cloud sample data is greater than or equal to a similarity threshold, thereby completing the training of the first model. In this way, the first model can be deployed on vehicles without LiDAR. When the electronic devices on the vehicle determine the height range based on the image data, they can adjust based on the height range data of the first model, thereby improving the accuracy of the height range data and thus improving the accuracy of the second feature data.
[0014] In one possible implementation of the first aspect above, the feature points in the first sub-region mapped to multiple sub-regions form feature segments at different heights relative to the horizontal plane of the vehicle coordinate system. In this case, any feature point in a feature segment has more than a preset number of feature points within a region of height M. The height range of the first feature segment in the first sub-region in the vehicle coordinate system is taken as the height range of the feature points in the first sub-region corresponding to the horizontal plane of the vehicle coordinate system. The first feature segment is the feature segment in the first sub-region with the lowest horizontal plane relative to the vehicle coordinate system.
[0015] In some embodiments of this application, if the first sub-region includes multiple feature segments, the feature segment with the lowest height from the horizontal plane has a higher probability of colliding with the vehicle. Therefore, height range data can be determined from the first feature segment with the lowest height from the horizontal plane. This ensures that the electronic device can more accurately identify the first feature segment, thereby avoiding collisions between the vehicle and the obstacle corresponding to the first feature segment.
[0016] In one possible implementation of the first aspect above, the visual information further includes category information, which includes ground category and suspended obstacle category; corresponding to the feature point with the lowest height in the first feature segment having a higher height than the feature point of the ground category, the category of the feature point in the first feature segment is set to the suspended obstacle category.
[0017] In some embodiments of this application, classifying the first category corresponding to the suspended obstacle separately can further improve the accuracy of the electronic device in identifying the suspended obstacle, thereby ensuring the driving safety of the vehicle.
[0018] In one possible implementation of the first aspect described above, the obstacle prediction of the environment surrounding the vehicle based on the fifth feature data includes: acquiring the fifth feature data corresponding to a set of feature points within a preset distance from the vehicle; increasing the resolution of the fifth feature data corresponding to the feature point set to obtain sixth feature data; and predicting obstacles in the environment surrounding the vehicle based on the sixth feature data.
[0019] In some embodiments of this application, electronic devices often need to more accurately identify obstacles close to the vehicle. Therefore, after acquiring the fifth feature data, the electronic device can acquire a set of feature points within a preset distance from the vehicle and improve the resolution of the fifth feature of this feature point set to obtain the sixth feature data. Based on the sixth feature data, the electronic device can more accurately identify obstacles within the preset distance of the vehicle.
[0020] Secondly, this application provides an electronic device, comprising: a memory for storing instructions; and at least one processor for executing the instructions to cause the device to implement the method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable in the second aspect can be referred to the beneficial effects of the method provided in any embodiment of the first aspect, and will not be repeated here.
[0021] Thirdly, this application provides a computer program product that, when run on a device, enables the device to implement the methods provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable in this third aspect can be referenced to the beneficial effects of the methods provided in any embodiment of the first aspect, and will not be repeated here. Attached Figure Description
[0022] Figure 1 A schematic diagram of a vehicle parking process is shown;
[0023] Figure 2 According to some embodiments of this application, a flowchart of the implementation of obstacle prediction by an on-board device is shown;
[0024] Figure 3A According to some embodiments of this application, a schematic diagram of an in-vehicle device processing multi-view images is shown;
[0025] Figure 3B According to some embodiments of this application, a schematic diagram of an in-vehicle device obtaining fifth feature data is shown;
[0026] Figure 4 According to some embodiments of this application, a flowchart of an implementation for obtaining the truth value is shown;
[0027] Figure 5 According to some embodiments of this application, a schematic diagram of the structure of an electronic device 100 is shown. Detailed Implementation
[0028] The illustrative embodiments of this application include, but are not limited to, obstacle prediction methods, program products, and electronic devices.
[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0030] As shown in the background section, during the parking process, the vehicle does not specifically process the information of suspended obstacles, resulting in inaccurate predictions of such information. This makes the vehicle prone to colliding with the corresponding suspended obstacles during parking, thus affecting parking safety.
[0031] The following describes a vehicle parking scenario.
[0032] For example, Figure 1 A schematic diagram of a vehicle parking process is shown.
[0033] like Figure 1 As shown, during parking, vehicle 10 can acquire multiple images of its surrounding environment using various sensors (such as a front-view camera, side-view cameras, and rear-view cameras). The vehicle's onboard equipment (such as a vehicle-mounted system) can then use these images to predict obstacles in the surrounding environment. For example, obstacles around vehicle 10 include wall 20, wall 30, and a fire extinguisher box 21 on wall 20. It is understood that the fire extinguisher box 21 is fixed to wall 20, and its bottom does not contact the ground; that is, the fire extinguisher box 21 is a suspended obstacle.
[0034] In-vehicle equipment typically extracts image features from multiple images using a neural network model. Then, based on depth information (the distance of each pixel in the image from the vehicle 10), image features, and the intrinsic and extrinsic parameter matrices of the cameras corresponding to each viewpoint, it performs view transform operations to fuse information and construct a 3D feature volume of the vehicle's surrounding environment. This 3D feature volume includes information about objects in the environment surrounding the vehicle 10 (such as basic information like position, size, and shape, and potentially more advanced information like texture, color, and surface reflection properties). The in-vehicle equipment then uses this 3D feature volume to predict the occupancy of the environment around the vehicle 10, thereby determining the occupancy status of obstacles. It's important to understand that the information in the 3D feature volume is still an abstract, high-level feature representation, not direct obstacle occupancy information. Occupancy prediction transforms the information in the 3D feature volume into concrete, understandable occupancy states (such as whether space is occupied by objects).
[0035] However, in the process of generating a 3D feature body by fusing image features, depth information, and intrinsic / extrinsic parameter matrices from multiple images through view transformation operations, specific processing was not performed on certain special features (such as suspended obstacles). The accuracy of the information of each object in the 3D feature body is largely determined by the proportion of pixels occupied by each object in the multiple images. For example, walls 20 and 30 occupy a relatively large proportion of pixels in the multiple images, therefore, the vehicle-mounted equipment obtains more information about walls 20 and 30 from the multiple images, resulting in more accurate spatial information about walls 20 and 30. On the other hand, fire box 21 occupies fewer pixels in the multiple images, resulting in less information about fire box 21 provided by the multiple images. During the process of generating the 3D feature body from the multiple images, the vehicle-mounted equipment does not pay enough attention to fire box 21, leading to inaccurate information about fire box 21 in the 3D feature body. Furthermore, suspended obstacles like fire box 21 are more likely to collide with the vehicle 10. Therefore, if the information of the fire box 21 generated by the on-board equipment based on multiple images is inaccurate, the risk of the vehicle 10 colliding with the fire box 21 during parking is relatively high.
[0036] As mentioned earlier, when the onboard equipment in the vehicle predicts obstacles in the surrounding environment based on multiple images of the surrounding environment collected by sensors, it does not specifically process information about suspended obstacles. This results in the onboard equipment predicting insufficiently accurate information about suspended obstacles, thus posing a significant safety hazard to vehicle parking.
[0037] To address the issue of inaccurate information in predicting suspended obstacles by vehicle-mounted devices, this application proposes an obstacle prediction method. An electronic device acquires multiple images, including images from different perspectives captured by the vehicle's sensors (e.g., images captured by a camera on the vehicle). The electronic device determines first feature data for feature points (e.g., individual pixels) in each of the multiple images. This first feature data includes first coordinate data of the feature points in the vehicle coordinate system, including the height value of the feature points. For example, the first coordinate data may be three-dimensional data of the feature points in the vehicle coordinate system.
[0038] An electronic device generates a planar feature fusion image corresponding to multiple images based on the first feature data of each image in multiple images. The feature points in the planar feature fusion image are the mapped feature points of each image in the multiple images onto the horizontal plane of the vehicle coordinate system. The electronic device can also generate feature data for the planar feature fusion image based on the first feature data of each image in the multiple images. This feature data includes second feature data, which includes the region coordinates of each sub-region obtained after dividing the planar feature fusion image, and the height range of the feature points in each sub-region in the vehicle coordinate system. Then, the electronic device can perform obstacle prediction based on the feature data of the planar feature fusion image.
[0039] Through the above scheme, electronic devices can predict obstacles based on a fusion image of planar features corresponding to multiple images. The second feature data in the fusion image includes the height range of feature points in various regions on the horizontal plane of the vehicle coordinate system. Therefore, the fusion image pays close attention to the height information of each feature point in the vehicle coordinate system, so that each feature point has more or more accurate height information when the electronic device predicts obstacles based on the fusion image. This ensures that the electronic device can more accurately predict the occupancy of suspended obstacles, thereby improving the safety of vehicle driving or parking.
[0040] Below, we describe the process of obstacle prediction by electronic devices in some embodiments of this application.
[0041] For example, Figure 2 According to some embodiments of this application, a flowchart of an implementation of obstacle prediction by an on-board device is shown.
[0042] It is understood that the following processes can be executed by electronic devices, which can be in-vehicle devices or servers. In-vehicle devices can be mobile phones, vehicle-mounted systems, terminals in self-driving systems, or wireless terminals in transportation safety. The following uses an in-vehicle device as an example to introduce the obstacle prediction method in the embodiments of this application.
[0043] like Figure 2 As shown, the process includes:
[0044] S201, acquire multiple images, including images from different perspectives collected by the vehicle's sensors.
[0045] For example, in some embodiments of this application, the vehicle can acquire images of its surrounding environment from different perspectives using sensors during driving or parking. The vehicle's sensors can be, for example, cameras, and the vehicle can be equipped with cameras from different perspectives (e.g., front-view cameras, side-view cameras, rear-view cameras, etc.). The cameras can acquire images of the environment surrounding the vehicle. When the onboard equipment predicts obstacles, it can acquire multiple images captured by the cameras from various perspectives.
[0046] In some embodiments of this application, the images captured by cameras from different perspectives are, for example, sequential; that is, after the vehicle's driver assistance functions are activated, the cameras can continuously capture image data. ,in for A set of panoramic images captured in real time. The image is captured by the Nth camera, where N is the number of panoramic images (i.e., there are N cameras), and t is the current time. In some embodiments of this application, N is greater than 4; for example, N can be 4, 5, 6, 7, or 8, meaning that the vehicle is equipped with 4, 5, 6, 7, or 8 cameras.
[0047] S202, determine the first feature data of feature points of each image in multiple images, wherein the first feature data includes the first coordinate data of the feature points in the vehicle coordinate system, and the first coordinate data includes the height value of the feature points.
[0048] In some embodiments of this application, each of the multiple images also includes multiple feature points. For example, a feature point may be a pixel of the image. In other embodiments, a feature point may be a pixel corresponding to a specific target, or a pixel within a preset area on the image may be used as a feature point, or a feature point may be added to the image.
[0049] It can be understood that the coordinates of the feature points on each image are two-dimensional coordinates, that is, the width and height of the image. The first coordinate data of the feature points are three-dimensional coordinate data in the vehicle coordinate system. Therefore, it is necessary to map each feature point from the corresponding image coordinate system to the vehicle coordinate system.
[0050] In some embodiments of this application, after acquiring multiple images, the vehicle-mounted device can extract features from feature points in each image to obtain the position data and visual information of the feature points in the image. For example, in some embodiments, the visual information may include information such as the color, texture, category, and light reflection intensity corresponding to the feature point. Through the visual information of the feature points, the vehicle-mounted device can determine whether a feature point is an obstacle, enabling it to predict the occupancy of obstacles in the vehicle's surrounding environment. Furthermore, the vehicle-mounted device can also perform view transformations on the feature points of each image based on their position data, thereby converting the position data of the feature points from the image coordinate system to the vehicle coordinate system.
[0051] For example, the transformation process converts feature points from the image coordinate system (w, h) based on the parameters of the sensor that acquired the image (e.g., focal length, pixel size, image center coordinates, etc.) to the sensor coordinate system (w', h', d). Here, w represents the width of the image coordinate system, h represents the height of the image coordinate system, w' represents the width of the sensor coordinate system, h' represents the height of the sensor coordinate system, and d represents the depth of the sensor coordinate system, such as the distance from the sensor.
[0052] Then, based on the parameters of each sensor relative to the vehicle (such as the sensor's position on the vehicle and its rotation angle), the feature points are transformed from sensor coordinates to vehicle coordinates (x, y, z). Here, x represents the length in the vehicle coordinate system, y represents the width, and z represents the height. In some embodiments of this application, the visual information of the feature point is C, and the first feature data of the feature point is (x, y, z, C).
[0053] For example, Figure 3A According to some embodiments of this application, a schematic diagram of an in-vehicle device processing multi-view images is shown.
[0054] like Figure 3A As shown, after acquiring multi-view images, the in-vehicle device can input these images into the backbone network to extract image features. These features can include the location data of each feature point and visual information. The in-vehicle device then inputs these image features into the context head and depth head networks to encode the image features respectively. and depth features ,in and For image feature resolution (e.g., the width w and height h of the corresponding image coordinate system), The number of image feature channels (e.g., the visual information C corresponding to the feature points). Let be the number of discrete depths (e.g., the depth d corresponding to the sensor coordinates), and ... For example, it could be the type of visual information about the feature points, such as the color, texture, category, and light reflection intensity corresponding to the feature points. The number of discrete depths. For example, it could be a discrete value of the distance between a feature point and a sensor (e.g., a camera, hereinafter referred to as a camera). For instance, in some embodiments of this application, the discrete depth quantity... A value of 24 represents 0m-23m, with discrete depth values spaced 1m apart. In other embodiments, the number of discrete depths... Other values are also possible, such as 30, 40, 50, etc. The embodiments of this application specify the dispersion depth. No restrictions are imposed.
[0055] In some embodiments, during the encoding of depth features, the deep network can adjust the depth features based on the depth ground truth through a depth ground truth prediction module. The process of adjusting the depth features by the depth ground truth prediction module is described in detail below.
[0056] In-vehicle equipment will use image features and depth features After being input into the viewpoint conversion module, the viewpoint conversion module can convert the feature points from image coordinates to vehicle coordinates based on the internal and external parameters of each camera.
[0057] For example, an in-vehicle device can construct the feature volume of the feature points of the i-th image by taking the outer product of the depth features and the image features. For example, refer to equation (1):
[0058] (1);
[0059] in, Let the feature body be the pixel coordinates of the feature point in the i-th image. This feature body includes the position data, depth data, and number of feature channels in the i-th image. The image features of the feature points in the i-th image include the position data of the pixels in the i-th image and the feature channel data. Let be the depth data of the feature points in the i-th image.
[0060] In some embodiments of this application, the vehicle-mounted device can project the feature points of multi-view images onto the 3D space of the vehicle coordinate system using a lift-splat-shot (LSS) scheme, based on the positions of feature points on the images. For example, the vehicle-mounted device can first generate a view frustum based on the position data of each feature point in the image, and then transform the view frustum from the image coordinate system to the vehicle coordinate system. Exemplarily, the view frustum can be the 3D coordinates of the feature points at different discrete depths in the image. The view frustum is a coordinate system formed by adding depth coordinates to the image pixel coordinate system, where 3 represents the x', y', z' coordinates within the view frustum. For discrete depth values, and Let B be the image feature resolution, B be the total number of images, and N be the number of surround-view cameras. The onboard equipment can rotate the view cone to the vehicle coordinate system based on the internal and external parameters of the cameras. Proj represents the mapping relationship between the vehicle coordinate system and the pixel coordinates of the image, where This represents the inverse normalization of pixel coordinates. K represents the camera's intrinsic parameters, which may include the camera's focal length, pixel size, and image center coordinates. R represents the camera's extrinsic parameters, which may include the camera's position on the vehicle and its rotation angle. Using the camera's intrinsic parameters K, feature points can be transformed from pixel coordinates to the camera's coordinate system; using the camera's extrinsic parameters R, feature points can be transformed from the camera's coordinate system to the vehicle's coordinate system.
[0061] In some embodiments of this application, the vehicle coordinate system can be rasterized to fuse feature points projected from various viewpoints into the vehicle coordinate system. For example, the resolution of the vehicle coordinate system rasterization is W×H×L, where W can be the raster resolution in the x-direction, H can be the raster resolution in the y-direction, and L can be the raster resolution in the z-direction.
[0062] Based on the resolution of the vehicle coordinate system, the mapping relationship between the rasterized vehicle coordinate system and the pixel coordinates of the image is determined. Based on this, image feature points are sampled, and the features of multiple feature points falling on the same grid are summed to obtain the first feature data, for example, referring to equation (2):
[0063] (2);
[0064] in This is the first feature data. This represents the mapping relationship between the rasterized vehicle coordinate system and the pixel coordinates of the image. Let N be the number of surround-view cameras (N images), H, W, L be the resolution of the vehicle coordinate system grid, and C be the number of feature channels. C can be understood as the feature channel data of the feature point in the rasterized vehicle coordinate system. In some embodiments of this application, the position of the feature point in the vehicle coordinate system grid can also be used as the first coordinate data in the first feature data, and the feature channel data corresponding to the feature point can be used as the visual information of the feature point.
[0065] For example, in some embodiments of this application, the visual information also includes category information, which includes a ground category and a first category. The first category is used to indicate that the corresponding feature point is a feature point of a suspended obstacle. Feature points with height information higher than the ground category can be feature points of the first category. It is understood that in some embodiments of this application, classifying the first category corresponding to a suspended obstacle separately by the vehicle-mounted device can further improve the accuracy of the vehicle-mounted device in identifying suspended obstacles, thereby ensuring vehicle driving safety.
[0066] In some embodiments of this application, after the vehicle coordinate system is rasterized, the coordinates of multiple feature points may lie within the same raster. Therefore, the coordinates of multiple feature points falling within the same raster can be summed to ensure that there is at most one feature point in each raster. The summation process may involve summing the visual information corresponding to multiple feature points within a raster. In other embodiments, the average visual information of each feature point can be taken as the visual information of the raster's feature points. This application does not limit the merging process of multiple feature points within a raster. If the visual information includes category information, the union can be taken. For example, if a raster contains two feature points, one corresponding to the ground category and the other not, then the category of the feature point corresponding to that raster is the ground category.
[0067] S203, Based on the first feature data of each image in multiple images, generate a planar feature fusion image corresponding to multiple images, wherein the feature points included in the planar feature fusion image are the mapped feature points of each image in multiple images onto the horizontal plane of the vehicle coordinate system.
[0068] In some embodiments of this application, the first feature data of each image in a plurality of images includes the first coordinate data of feature points in each image in the vehicle coordinate system. Since the first coordinate data is in a three-dimensional coordinate system, the height data of the feature points in the vehicle coordinate system is not easily expressed explicitly. Therefore, the three-dimensional coordinates of each feature point in the vehicle coordinate system can be mapped to the horizontal plane of the vehicle coordinate system, thereby generating a planar feature fusion image corresponding to multiple images. Each mapped feature point on the planar feature fusion image can correspond to a height range, thereby explicitly expressing the height data of the feature points in each image, so that the electronic device can more accurately identify suspended obstacles based on the planar feature fusion image.
[0069] For example, the plane formed by the x and y directions of the vehicle coordinate system can be the horizontal plane of the vehicle coordinate system. That is, the coordinates of each mapped feature point on the planar feature fusion image are two-dimensional coordinates (x, y). It can be understood that each mapped feature point on the planar feature fusion image can be a point mapped from a feature point to the same two-dimensional coordinate on the planar feature fusion image. In other words, a mapped feature point can include multiple feature points at different heights (e.g., different z-coordinates).
[0070] S204, based on the first feature data of each image in multiple images, generate feature data of a planar feature fusion image, wherein the feature data of the planar feature fusion image includes second feature data, which includes: the region coordinate data of each sub-region obtained after dividing the planar feature fusion image into multiple sub-regions, and the height range of the feature points in each sub-region in the vehicle coordinate system.
[0071] In some embodiments of this application, the resolution of the vehicle coordinate system grid is (W, H, L), therefore the resolution of the planar grid on the horizontal plane of the vehicle coordinate system is (W, H), that is, the resolution of the planar grid of the planar feature fusion image is (W, H). In some embodiments of this application, a planar grid can be considered as a sub-region. The coordinate data of each planar grid can be the two-dimensional coordinate data of the center point of each planar grid in the planar feature fusion image, or the two-dimensional coordinate data of the lower left and upper right vertices of each planar grid in the planar feature fusion image. In some embodiments, the coordinate data of the sub-region can also be the position coordinates of the planar grid. For example, if the coordinates of a planar grid are (W1, H1), it means that the planar grid is in the W1th row in the x-direction and the H1th column in the y-direction.
[0072] It is understood that since each grid cell in the vehicle coordinate space includes at most one feature point, in a planar feature image, a planar grid cell also includes at most one mapped feature point. Therefore, in some embodiments, the coordinate data of the sub-region can also be the two-dimensional coordinate data of each mapped feature point in the planar feature fusion image, or the two-dimensional coordinate data of each mapped feature point in the planar feature fusion image can also be the coordinate data of each sub-region in the planar feature fusion image.
[0073] It can be understood that the coordinate data corresponding to the planar grid represents the position of each planar grid in the planar fused image. Therefore, the height range of a sub-region can be the height range of feature points in grids with different L-coordinates but the same planar grid position. For example, if the grid coordinates of n feature points in the vehicle coordinate space are (W1, H1, L1), (W1, H1, L2), (W1, H1, ...), (W1, H1, Ln), after mapping these n feature points to the planar feature fused image, they will all be mapped to the coordinates of the planar grid (W1, H1). Then, the height range of this planar grid (W1, H1) can be from L1 to Ln, where Ln is the height position of the grid corresponding to the feature point with the highest height (largest z-coordinate) among the n feature points, and L1 is the height position of the grid corresponding to the feature point with the lowest height (smallest z-coordinate) among the n feature points.
[0074] For example, in the process of determining the second feature data based on the first feature data, the vehicle-mounted equipment can merge the L-dimensional and C-dimensional dimensions of the first feature data, and then use conv2d to encode the merged data into a B×H×W×L×2 dimension (where 2 is to take into account the need to predict the maximum and minimum height values (e.g., Ln and L1) for suspended obstacles, in order to determine the height range of each two-dimensional grid).
[0075] For example, the vehicle-mounted equipment calculates the probability density of the maximum and minimum heights by performing a normalized exponential function (softmax) on the L-dimensional dimension on the corresponding height values in each sub-region, as shown in formulas (3) and (4):
[0076] (3);
[0077] (4);
[0078] in, This represents the highest probability density corresponding to the second coordinate data of each sub-region on the horizontal plane. This represents the probability density of the minimum height corresponding to the second coordinate data of each sub-region on the horizontal plane. ( ) represents a conv2d encoding. That is, the second feature data includes the second coordinate data (corresponding to the grid coordinates on the H×W plane) of the sub-region of the planar feature fusion image (corresponding to the H×W plane) and the coordinates corresponding to the second coordinate data. and , and As an example of height range, in some embodiments of this application, the dimension of the planar feature fusion image can be related to the number of pixels in the image. same.
[0079] In some embodiments, the first coordinate data is mapped to multiple feature points in the first sub-region to form multiple feature segments at different heights relative to the horizontal plane, wherein a feature point in a feature segment has more than a preset number of adjacent feature points in an adjacent region at a height of M.
[0080] In some embodiments of this application, the height range of the first feature segment with the lowest height among the multiple feature segments corresponding to the first sub-region can be used as the height range of the first sub-region.
[0081] It is understandable that if the first sub-region includes multiple feature segments, the feature segment with the lowest height from the horizontal plane has a higher probability of colliding with the vehicle. Therefore, the height range of the first feature segment can be determined by selecting the highest and lowest feature points from the first feature segment with the lowest height from the horizontal plane. This ensures that the on-board equipment can more accurately identify the feature segment, thereby avoiding collisions between the vehicle and obstacles corresponding to that feature segment.
[0082] In some embodiments of this application, the vehicle-mounted device can also adjust the height range in the second feature data through the height true value estimation module and the height true value, and the specific adjustment process is described below.
[0083] S205, obstacle prediction based on feature data of planar feature fusion image.
[0084] In some embodiments of this application, the feature data of the planar feature fusion image includes second feature data, which consists of second coordinate data of multiple sub-regions and height range data of each sub-region. The height range data explicitly expresses height information. Therefore, when the vehicle-mounted device performs obstacle prediction based on the planar feature image, it can more accurately predict whether an obstacle is a suspended obstacle based on the height information, thereby improving the prediction effect of the vehicle-mounted device on suspended obstacles.
[0085] Through the above scheme, electronic devices can more accurately predict the occupancy of suspended obstacles in the surrounding environment based on the fusion of planar features, thereby reducing the probability of vehicles colliding with suspended obstacles during parking or driving and improving vehicle safety.
[0086] In some embodiments of this application, the feature data of the planar feature fusion image further includes third feature data. For example, an electronic device can generate third feature data of the planar feature fusion image based on the first feature data of each image in multiple images. The third feature data includes: second coordinate data and first encoded data of each mapped feature point in the planar feature fusion image. In some embodiments of the application, the second coordinate data is two-dimensional coordinate data that maps the first coordinate data of the feature points to the horizontal plane of the vehicle coordinate system, i.e., the two-dimensional coordinate data of the mapped feature points. In other words, the second coordinate data can be the position coordinate data of a planar grid in the planar feature fusion image. The first encoded data is the combined data of visual information and height data of feature points mapped to the same second coordinate data.
[0087] For example, in the vehicle coordinate system, the resolution of the grid is (W, H, L), and each grid includes only one feature point. Therefore, the first coordinate data of the feature point in the vehicle coordinate system can also be (W, H, L), and the first feature data can also be (W, H, L, C). After the feature points of each image are mapped onto the planar feature fusion image, the coordinates of the corresponding mapped feature points are (W, H). In the embodiments of this application, the height coordinate L and visual information C in the first feature data of the feature points of each image corresponding to each mapped feature point on the planar feature fusion image can be merged. The merged feature can be, for example, M, that is, M is the merged data of C and L. It can be understood that since each mapped feature point can correspond to multiple feature points of images with the same two-dimensional coordinates (W, H) but different height coordinates L, L in M is actually a high-dimensional array. For example, in the vehicle coordinate system, the resolution of L determines the dimension of L in M. That is to say, the third feature data is essentially (W, H, M), where (W, H) are the coordinates of the mapped feature point on the planar feature fusion image, and M is the first encoded data.
[0088] For example, in some embodiments of this application, the L and C dimensions of the first feature data can be merged using a reshape function, and the electronic device can further encode the merged data using two-dimensional convolutional coding (conv2d) to transform the merged data to... Dimension, to obtain the third feature data For example, refer to formula (5):
[0089] (5);
[0090] in, This is the third feature data. ( ) represents a conv2d encoding. The first feature data is represented by `reshape()`, which is the reshaping function. In other words, the third feature data... It can include the second coordinate data of the mapped feature points onto the planar feature fusion image, and the corresponding number of channels. Number of channels The number of channels is obtained by combining the height dimension L and the number of feature channels C based on a reshaping function, followed by two-dimensional convolutional encoding. The corresponding channel data can serve as an example of the first encoded data. For instance, in some embodiments, the L dimension in the first feature data is 2, and the number of feature channels C is 3. By merging the L dimension and the number of feature channels C using a reshape method, a dimension of 2 × 3 = 6 can be obtained. This dimension can be the new number of channels in the third feature data. The new channel data is then encoded by conv2d as follows: The number of channels. and The resolution of each feature point mapped onto the planar feature fusion image is defined as follows. In embodiments of this application, the resolution of each feature point mapped onto the planar feature fusion image can be the same as the resolution of the planar raster, i.e., W×H can be the same as... × Similarly, in other embodiments, the resolution of each feature point mapped onto the planar feature fusion image can be other values. This application does not limit the resolution of each feature point mapped onto the planar feature fusion image.
[0091] In some embodiments of this application, the electronic device may also generate fourth feature data of a planar feature fusion image based on the first feature data of each image in multiple images. The fourth feature data includes: second coordinate data and second encoding data of each mapped feature point in the planar feature fusion image. The second encoding data is fusion data of visual information of feature points mapped to the same second coordinate data.
[0092] For example, if the coordinates of n feature points in the vehicle coordinate space grid are (W1, H1, L1), (W1, H1, L2), (W1, H1, ...), (W1, H1, Ln), after mapping these n feature points to the planar feature fusion image, they will all be mapped to the coordinates of the planar grid (W1, H1). Then, the second encoded data of this planar grid (W1, H1) is Cs, where Cs is the sum of the corresponding visual information of the n feature points. That is, the fourth feature data is (W, H, Cs).
[0093] For example, in some embodiments of this application, the corresponding data on the feature channels of feature points with the same second coordinate after mapping to the horizontal plane can be added together by summation, and then the number of channels can be encoded by conv2d. Dimension, thereby obtaining the fourth feature data For example, refer to formula (6):
[0094] (6);
[0095] in, This is the fourth feature data. ( ) represents a conv2d encoding, and SUM() is a summation function. It can be understood that the fourth feature data includes the second coordinate data and the second encoded data of the planar feature fusion image.
[0096] After determining the third and fourth feature data of the planar feature fusion image, the second, third, and fourth feature data can be fused into the fifth feature data, and the electronic device can perform obstacle prediction based on the fifth feature data.
[0097] For example, Figure 3B According to some embodiments of this application, a schematic diagram of an in-vehicle device obtaining fifth feature data is shown.
[0098] like Figure 3B As shown, the vehicle-mounted device can encode the height data of the highest and lowest points based on conv1×1 (representing a one-dimensional convolution kernel) and a normalized exponential function (softmax), thereby obtaining... and Then the vehicle-mounted equipment will respectively... and The maximum and minimum height features in the third feature data are enhanced by multiplying the third feature data with the attention mechanism. Then, the vehicle-mounted device is further encoded by conv3×3 (representing a 3D convolution kernel). The two sets of encoded features are concatenated with the fourth feature data, then concatenated with the third feature data after conv3×3 encoding, and then subjected to another conv3×3 feature extraction to obtain the final fifth feature data. For example, refer to formula (7):
[0099] (7);
[0100] in, This is the fifth feature data. It is a conv2d encoding. Let be the self-attention fusion function, where, CAT() is the concatenation function.
[0101] Reference Figure 3A In some embodiments of this application, the vehicle-mounted device may also input the fifth feature data into the time-series fusion module, the encoding module, and the channel conversion module respectively to finally obtain 3D features.
[0102] It is understandable that in-vehicle equipment can acquire multiple images at different times. Therefore, the fifth feature data determined from these images can be fused to determine the speed and motion of each feature point. The encoding module can then encode the speed and motion of the feature points to add motion features. The channel conversion module can separate the feature channels in the fifth feature data into L and C channels, thereby reconstructing the 3D features. The in-vehicle equipment can also adjust the 3D features based on 3D occupancy ground truth, enabling it to more accurately predict the occupancy of obstacles in the vehicle's surroundings and classify obstacles. For example, determining whether an obstacle is a moving obstacle, a suspended obstacle, a pedestrian, or a vehicle.
[0103] In some situations (such as when a vehicle is parked), obstacle prediction in scenes close to the vehicle requires high-resolution prediction accuracy, while obstacle prediction in scenes far from the vehicle does not require high-resolution prediction accuracy. Therefore, in some embodiments of this application, a local prediction head can be used to meet the accuracy requirements of different ranges. For example, the local prediction head extracts feature points of a local area close to the vehicle (as an example of a feature point set) from the fifth feature data (or 3D features) and upsamples them to obtain the sixth feature data corresponding to the upsampled feature points, thereby improving the resolution of feature points in the local area close to the vehicle. The onboard device predicts occupancy information based on the sixth feature data, which can further improve the accuracy of close-range vehicle prediction.
[0104] In summary, this solution fuses the third and second feature data using a self-attention mechanism, which strengthens the maximum and minimum height features in the second feature data. This makes the fused fifth feature more focused on the height of feature points, thereby improving the accuracy of the onboard equipment in predicting suspended obstacles. Furthermore, based on the true height values, the accuracy of the maximum and minimum heights in the fifth feature data can be further improved, enhancing the onboard equipment's obstacle prediction performance and ultimately preventing collisions between vehicles and suspended obstacles.
[0105] Understandable, refer to Figure 3AWhen the vehicle-mounted device acquires depth features, second feature data, and fifth feature data, it is necessary to acquire the ground truth value of depth, the ground truth value of height, and the ground truth value of 3D occupancy, respectively, so as to adjust the depth features, second feature data, and fifth feature data, thereby improving the accuracy of the corresponding data. The process of acquiring the ground truth value is described below in some embodiments of this application.
[0106] Exemplary examples, in some embodiments of this application, ground truth data can be acquired using a data acquisition vehicle. It is understood that the sensors in the data acquisition vehicle include cameras with multiple viewing angles, as well as sensors capable of acquiring point cloud data, such as LiDAR. During driving or parking, the data acquisition vehicle can acquire image data through cameras with multiple viewing angles, and can also acquire point cloud data from multiple viewing angles using LiDAR. For example, in some embodiments of this application, the data acquisition vehicle can acquire point cloud data in four directions. Thus, the data acquisition vehicle can acquire point cloud data from multiple viewing angles, thereby determining the ground truth values for depth, height, and 3D occupancy based on the point cloud data.
[0107] Below, we describe the process of determining the depth truth value, height truth value, and 3D occupancy truth value in some embodiments of this application.
[0108] For example, Figure 4 According to some embodiments of this application, a flowchart of an implementation for obtaining the truth value is shown.
[0109] For example, the execution entities of the following processes can be in-vehicle devices or electronic devices such as servers that are capable of processing point cloud data.
[0110] like Figure 4 As shown, the process includes:
[0111] S410, acquire point cloud data.
[0112] For example, in some embodiments of this application, the data acquisition vehicle is equipped not only with multiple cameras but also with LiDAR sensors in multiple directions. During driving or parking, the data acquisition vehicle can acquire image data through the cameras and point cloud data from multiple directions through the LiDAR sensors.
[0113] Servers or processors can acquire image data and point cloud data from different directions collected by LiDAR in order to process the point cloud data.
[0114] S420 determines the true depth value based on point cloud data.
[0115] In some embodiments of this application, the vehicle-mounted device can project point clouds onto the image coordinate system of the corresponding camera-captured images using the internal and external parameters of cameras from various viewpoints, and filter out points outside the image range according to the image resolution to obtain the depth value corresponding to the image pixels. In some embodiments of this application, the vehicle-mounted device presets [a certain parameter] when processing the image. Depth values are predicted using multiple depth planes. To facilitate the adjustment of feature points on the image using the point cloud, the vehicle-mounted device can also discretize the depth values of the point cloud to... This allows us to obtain the depth ground truth. In some embodiments of this application, the depth ground truth can be one-hot encoded, and then the feature points of the image can be analyzed using the one-hot encoded depth ground truth. The discrete-depth one-hot encoded data of the dimension is adjusted. The adjustment process could, for example, use cross-entropy loss (CELoss) to correct the feature points of the image. One-hot encoded data of discrete depth in dimensional form, so that feature points are in One-hot encoded data with discrete depths is more accurate; the adjustment process is described in detail below.
[0116] The process of determining the depth truth value is described below.
[0117] S421 projects point cloud data onto the image coordinate system based on the camera's internal and external parameters.
[0118] For example, the vehicle-mounted device can project the acquired point cloud data onto the coordinates of the image based on the camera's internal and external parameters.
[0119] S422 filters point clouds that exceed the image range based on the image resolution size.
[0120] For example, the point cloud data collected by the LiDAR on the vehicle is more comprehensive, while an image from one viewpoint may not be able to contain all the point cloud data. Therefore, after the on-board device projects the point cloud data onto the image coordinate system, it can filter out the point cloud that exceeds the image range.
[0121] S423 discretizes the point cloud within the image resolution to Dimensions, to obtain the truth value of depth.
[0122] For example, after the vehicle-mounted device obtains point cloud data within the image range, it can process the data according to a preset discrete depth. Discretize the point cloud data to Dimension, thereby obtaining the depth truth value.
[0123] S430 rasterizes the point cloud data, with the rasterization dimensions being W×H×L.
[0124] In some embodiments of this application, after the vehicle-mounted device acquires point cloud data, it can also rasterize the point cloud with a resolution of W×H×L, so that the point cloud data can correspond to the rasterized vehicle coordinate system.
[0125] S440 determines 3D occupancy truth based on rasterized point clouds.
[0126] For example, in order to better predict suspended obstacles, the vehicle-mounted device can set suspended obstacles as a separate category and distinguish whether the obstacle corresponding to the feature point is a suspended obstacle by using 3D occupancy ground truth.
[0127] In some embodiments of this application, the vehicle-mounted device can determine the 3D occupancy ground truth based on the rasterized point cloud data. The process of obtaining the 3D occupancy ground truth is described below.
[0128] S441, backup rasterized point cloud.
[0129] For example, the on-board device can back up rasterized point cloud data.
[0130] S442, traverse the cylinders corresponding to the W×H dimensions of the rasterized point cloud.
[0131] For example, the rasterized point cloud has dimensions of W×H×L. In the process of determining the 3D occupancy truth value, the vehicle-mounted device can traverse the rasterized point cloud with one grid corresponding to one cylinder on the horizontal plane in the W×H dimension as a unit and process each cylinder.
[0132] S443, determine whether the traversal is complete.
[0133] The onboard device can determine whether all cylinders within the W×H dimension have been traversed. For example, the onboard device can retrieve the current cylinder and determine whether the current cylinder has been processed.
[0134] If the judgment result is yes, the vehicle-mounted device can execute S444 to obtain the 3D occupancy truth value.
[0135] If the result is negative, the on-board equipment can execute S445 to determine whether the current column is a suspended obstacle.
[0136] S444, get the 3D occupancy truth value.
[0137] For example, once the vehicle-mounted device has traversed all the cylinders within the W×H dimension, it can be determined that the vehicle-mounted device has processed all the cylinders within the W×H dimension, thereby enabling the vehicle-mounted device to obtain the 3D occupancy truth value.
[0138] S445, determine whether the current column is a suspended obstacle.
[0139] For example, if the vehicle-mounted device determines that it has not yet traversed all the pillars within the W×H dimension, it can process the current pillar. For instance, the process of processing the current pillar can be to determine whether the current pillar is a suspended obstacle.
[0140] If the judgment result is yes, then execute S446 to determine whether the point cloud in the current suspended obstacle occupies more than a preset number of point clouds in the adjacent M region.
[0141] If the result is negative, then obtain the next cylinder within the W×H dimension and execute S443 to determine whether the traversal is complete.
[0142] For example, for a cylinder within the W×H dimension, if the point with the minimum height in the cylinder's point cloud is higher than the ground height (the ground height can refer to the height coordinates of a point cloud classified as ground), then it can be determined that the cylinder contains a suspended obstacle.
[0143] S446, determine whether the point cloud in the current suspended obstacle occupies more than a preset number of point clouds in the adjacent M region.
[0144] For example, in some embodiments, there may be interfering points (noise) within a column. Therefore, after determining that the current column corresponds to a suspended obstacle, it is also necessary to determine whether the point cloud within the column is noise. For example, within a window of m of the point cloud within the column, it is determined whether there are more than a threshold h of point clouds occupying the space.
[0145] If the judgment result is yes, then execute S447 to mark the current column as a suspended obstacle and complete the processing of the current column.
[0146] If the result is negative, then obtain the next cylinder within the W×H dimension and execute S443 to determine whether the traversal is complete.
[0147] It is understandable that if the point cloud within the main body has very little occupancy in the neighborhood of the adjacent m, then the point cloud within the cylinder may be noise, and the current cylinder category will not be changed.
[0148] S447, mark the current column as a suspended obstacle, and complete the processing of the current column.
[0149] For example, after the on-board device determines that there is a suspended obstacle within the current cylinder, it can mark the current cylinder as a suspended obstacle. Then, it obtains the next cylinder within the W×H dimension and executes S443 to determine whether the traversal is complete.
[0150] S450 determines the true height based on rasterized point clouds.
[0151] For example, after point cloud rasterization, it includes three dimensions: W, H, and L. Each grid cell on the horizontal plane (W×H dimension) can be treated as a cylinder. The height values of the highest and lowest point cells in each cylinder are determined as the true height values of the cylinder corresponding to the H×W dimension grid. The process of obtaining the true height values is described below.
[0152] S451 sets the array corresponding to the W×H dimension, with the default value of the array being ignore.
[0153] For example, after the point cloud is rasterized, each grid in the H×W dimension corresponds to a cylinder. The vehicle-mounted device can set an array for each grid in the H×W dimension, which is used to represent the true height value within the corresponding grid. In the initial state, the array for each grid in the H×W dimension is set to the default value "ignore".
[0154] S452, traverse the cylinders corresponding to the W×H dimensions of the rasterized point cloud.
[0155] The vehicle-mounted device can traverse the cylinders corresponding to the W×H dimension of the rasterized point cloud, thereby processing the array of each raster in the H×W dimension.
[0156] S453, determine whether the traversal is complete.
[0157] The vehicle-mounted device can determine whether all the cylinders corresponding to the grids in the W×H dimension have been traversed. For example, the vehicle-mounted device can determine whether the cylinder corresponding to the current grid has been traversed.
[0158] If the judgment result is yes, then obtain the height truth value.
[0159] If the result is negative, then execute S454 to determine whether the current column is occupied by point cloud.
[0160] S454 determines whether the current column is occupied by point cloud.
[0161] For example, the vehicle-mounted device can determine whether there is point cloud occupancy in the cylinder corresponding to the grid in the current W×H dimension.
[0162] If the judgment result is yes, then execute S456 to determine whether the point cloud in the current column is continuously occupied.
[0163] If the result is negative, execute S455 to set the array in the current grid to ignore.
[0164] S455 sets the array in the current grid to ignore.
[0165] For example, if the vehicle-mounted device determines that the column corresponding to the current grid in the W×H dimension is not occupied by point cloud, it means that there is no obstacle in the current column, and then the array corresponding to the current grid in the W×H dimension is set to ignore.
[0166] S456, determine whether the point cloud in the current column is continuously occupied.
[0167] For example, in some embodiments of this application, the point cloud within a cylinder can be continuously occupied or non-continuously occupied. Continuous occupation means that there is only one continuous point cloud segment within the cylinder. Specifically, any point within a point cloud segment has a predetermined number of adjacent point clouds within a height of M. Therefore, if point cloud occupancy exists within a cylinder, it is necessary to determine whether the point cloud occupancy is continuous, i.e., to determine whether multiple point cloud segments exist.
[0168] If the judgment result is yes, then execute S457, set the height index of the highest point and the height index of the lowest point in the current column into the array of the current grid, and complete the processing of the current column.
[0169] If the judgment result is negative, then execute S458, set the height index of the highest point and the height index of the lowest point in the point cloud segment with the lowest height in the array of the current grid, and complete the processing of the current column.
[0170] S457 sets the height index of the highest point and the height index of the lowest point in the current column into the array of the current grid, and completes the processing of the current column.
[0171] For example, if the onboard device determines that the point cloud of the current column is continuously occupied, it means that there is only one point cloud segment in the current column, and then sets the height index of the highest point in that point cloud segment to [the specified value]. And set the height index of the lowest point to and will and The value is assigned to the array of grid cells along the W×H dimension corresponding to the current cylinder. This completes the processing of the true height of the current cylinder, allowing the onboard device to process the true height of the next subject.
[0172] S458: In the point cloud segment with the lowest height, set the height index of the highest point and the height index of the lowest point in the array of the current grid, and complete the processing of the current column.
[0173] For example, if the onboard device determines that the point cloud of the current column is not continuously occupied, it means that there are multiple point cloud segments within the current column. In this case, the onboard device can set the height index of the highest point in the lowest point cloud segment to... And set the height index of the lowest point in the lowest point cloud segment to... and will and The value is assigned to the array of grid cells along the W×H dimension corresponding to the current cylinder. This completes the processing of the true height of the current cylinder, allowing the onboard device to process the true height of the next subject.
[0174] It is understandable that when there are multiple point cloud segments within a cylinder, the point cloud segment with the lower height is more likely to collide with a vehicle. Therefore, the positions of the highest and lowest points on the point cloud segment with the lowest height can be assigned to the array of the raster in the current W×H dimension.
[0175] Understandably, through the above process, the vehicle-mounted equipment can determine the ground truth values for depth, height, and 3D occupancy based on the point cloud data. Then, the vehicle-mounted equipment can process the image data based on these ground truth values.
[0176] For example, refer to Figure 3A Taking the high-level true value prediction module as an example, this paper introduces the process of adjusting the second feature data based on the true value.
[0177] For example, in the height ground truth prediction module (as an example of the first model), an adjustment parameter can be set to adjust the height range in the input second feature data.
[0178] The training process of the altitude ground truth prediction module is described below. For example, taking a server as an example, the altitude ground truth prediction module can be deployed on the server during training. The server can obtain image samples from the data acquisition vehicle and determine the fourth coordinate data mapped to the vehicle coordinate space based on the image samples. Exemplarily, the server can also determine the altitude range sample data in the second feature data from the image samples according to the process in S204 based on the fourth coordinate data.
[0179] The server can acquire point cloud data from the data collection vehicle and input the ground truth altitude determined based on the third coordinate data of the point cloud (as an example of a point cloud sample) into the altitude ground truth prediction module. The altitude ground truth prediction module can adjust the altitude range samples according to the initial adjustment parameters (as an example of the first parameter) to obtain first adjusted data. Then, the altitude ground truth prediction module determines the similarity between the first adjusted data and the ground truth altitude based on cross-entropy loss, and corrects the initial adjustment parameters in the altitude ground truth prediction module based on the similarity. The altitude ground truth prediction module then adjusts the altitude range sample data again based on the corrected adjustment parameters to improve the similarity between the altitude range sample data and the ground truth altitude. This process is repeated until the adjusted parameters, after adjusting the altitude range sample data, make the similarity between the adjusted altitude range sample data and the ground truth altitude greater than or equal to a similarity threshold. At this point, the adjustment parameters in the altitude ground truth prediction module can be considered trained successfully, and the trained adjustment parameters are deployed as the second parameter to the altitude ground truth prediction module.
[0180] The height ground truth prediction module after training and its relationship with Figure 3A The various modules can be deployed on the vehicle-mounted equipment of ordinary vehicles (vehicles without LiDAR). After the vehicle-mounted equipment on ordinary vehicles determines the second feature data based on the image, it can adjust the height range in the second feature data based on the height true value prediction module to make the second feature data more accurate.
[0181] It is understandable that the training process of the depth ground truth prediction module is similar to that of the height ground truth prediction module. For ordinary vehicles, adjusting the depth ground truth only requires inputting the depth features into the depth ground truth prediction module. Similarly, based on the training method of the height ground truth prediction module, a model for adjusting the fifth feature data can be trained using 3D occupancy ground truth. This model can be deployed in the channel conversion module, allowing the channel conversion module to adjust the fifth feature data. In other embodiments, the model for adjusting the fifth feature data trained based on 3D occupancy ground truth can also be deployed in the feature fusion module, temporal fusion module, encoding module, etc. In this way, the onboard device can determine more accurate fifth feature data (or 3D features), thereby improving the accuracy of the onboard device's prediction of suspended obstacles.
[0182] The electronic devices involved in the above embodiments are described below.
[0183] For example, Figure 5 According to some embodiments of this application, a schematic diagram of the structure of an electronic device 100 is shown.
[0184] The electronic device 100 may be an in-vehicle device as described in the foregoing embodiments, and the electronic device 100 is used to implement the obstacle prediction method provided in the foregoing embodiments.
[0185] like Figure 5 As shown, the electronic device 100 includes one or more processors 101, system memory 102, non-volatile memory (NVM) 103, communication interface 104, input / output device 105, and system control logic unit 106 for coupling the processor 101, system memory 102, non-volatile memory 103, communication interface 104, and input / output (I / O) device 105. Wherein:
[0186] Processor 101 may include one or more processing units, such as a central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), microprocessor (MCU), artificial intelligence (AI) processor, field programmable gate array (FPGA), neural network processing unit (NPU), etc. The processing module or circuit may include one or more single-core or multi-core processors. In some embodiments, the CPU may be used to optimize the neural network model to be run; for example, in some embodiments of this application, the neural network model may optimize the spatial data of a first obstacle model, and the NPU may be used to run the neural network model to be run.
[0187] System memory 102 is volatile memory, such as random-access memory (RAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc. System memory is used for temporary storage of data and / or instructions. For example, in some embodiments, system memory 102 can be used to store data provided in the foregoing embodiments, such as sensor data, image data, or video data, and can also be used to store instructions for obstacle prediction methods provided in the foregoing embodiments.
[0188] The non-volatile memory 103 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 may include any suitable non-volatile memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), compact disc (CD), digital versatile disc (DVD), solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 103 may also be a removable storage medium, such as secure digital (SD) storage. In other embodiments, the non-volatile memory 103 may be used to store instructions for the obstacle prediction methods provided in the foregoing embodiments.
[0189] Specifically, system memory 102 and non-volatile memory 103 may each include a temporary copy and a permanent copy of instruction 107. Instruction 107 may include, when executed by at least one of processors 101, causing electronic device 100 to implement the obstacle prediction method provided in the embodiments of this application.
[0190] The communication interface 104 may include a transceiver for providing a wired or wireless communication interface for the electronic device 100, thereby enabling communication with any other suitable device via one or more networks. In some embodiments, the communication interface 104 may be integrated into other components of the electronic device 100, for example, the communication interface 104 may be integrated into the processor 101. In some embodiments, the electronic device 100 may communicate with other devices through the communication interface 104, for example, the electronic device 100 may obtain relevant data from other devices through the communication interface 104.
[0191] Input / output (I / O) device 105 can be an input device such as a keyboard or mouse, and an output device such as a monitor. Users can interact with electronic device 100 through input / output (I / O) device 105.
[0192] The system control logic unit 106 may include any suitable interface controller to provide any suitable interface to other modules of the electronic device 100. For example, in some embodiments, the system control logic unit 106 may include one or more memory controllers to provide an interface to the system memory 102 and the non-volatile memory 103.
[0193] In some embodiments, at least one of the processors 101 may be packaged together with the logic of one or more controllers for the system control logic unit 106 to form a system in package (SiP). In other embodiments, at least one of the processors 101 may also be integrated on the same chip with the logic of one or more controllers for the system control logic unit 106 to form a system-on-chip (SoC).
[0194] Understandable. Figure 5 The structure of the electronic device 100 shown is merely an example. In other embodiments, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0195] It is understood that electronic device 100 can be any device configured on the vehicle, including but not limited to mobile phones, in-vehicle systems, terminals in self-driving vehicles, wireless terminals in transportation safety, terminals in smart cities, and so on.
[0196] This application also provides a program product that, when executed on an electronic device, enables the electronic device to implement the methods provided in the foregoing embodiments.
[0197] This application also provides a readable storage medium storing one or more programs, which, when executed by an electronic device, enable the electronic device to implement the methods provided in the foregoing embodiments.
[0198] Various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or combinations of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0199] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application-specific integrated circuit, or a microprocessor.
[0200] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0201] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried on or stored thereon by one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media can include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, compact disc-read-only memory (CD-ROMs), magneto-optical disks, read-only memory (ROM), random-access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagation signals. Therefore, machine-readable media includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.
[0202] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.
[0203] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.
[0204] It should be noted that in the examples and description of this invention, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0205] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made thereto without departing from the scope of this application.
Claims
1. An obstacle prediction method, applied to electronic devices, characterized in that, The method includes: Acquire multiple images, wherein the multiple images include images from different perspectives collected by the vehicle's sensors; First feature data of feature points of each image in the plurality of images is determined, wherein the first feature data includes the first coordinate data of the feature points in the vehicle coordinate system, and the first coordinate data includes the height value of the feature points; Based on the first feature data of each of the multiple images, a planar feature fusion image corresponding to the multiple images is generated, wherein the feature points included in the planar feature fusion image are: the feature points of each of the multiple images mapped to the horizontal plane of the vehicle coordinate system; Based on the first feature data of each image in the multiple images, feature data of the planar feature fusion image is generated, wherein the feature data of the planar feature fusion image includes second feature data, the second feature data including: the region coordinate data of each sub-region of the multiple sub-regions obtained after dividing the planar feature fusion image, and the height range of the feature points in each sub-region in the vehicle coordinate system; Obstacle prediction is performed based on the feature data of the fused planar feature image.
2. The method according to claim 1, characterized in that, The first feature data also includes visual information of each feature point, and the feature data of the planar feature fusion image also includes third feature data, and... The method further includes: Based on the first feature data of each image in the multiple images, the third feature data of the planar feature fusion image is generated, wherein the third feature data includes: the second coordinate data and the first encoding data of each mapped feature point in the planar feature fusion image; The second coordinate data is two-dimensional coordinate data that maps the first coordinate data of the feature point to the horizontal plane of the vehicle coordinate system, and the first encoded data is the combined data of the visual information and height data of the feature point mapped to the same second coordinate data.
3. The method according to claim 2, characterized in that, Also includes: Based on the first feature data of each image in the multiple images, the fourth feature data of the planar feature fusion image is generated, wherein the fourth feature data includes: the second coordinate data and the second encoding data of each mapped feature point in the planar feature fusion image; Each mapped feature point on the planar feature fusion image is a point on the planar feature fusion image that is mapped from the feature points of each image to the point corresponding to the same two-dimensional coordinate on the planar feature fusion image. The second encoded data is the sum of the visual information of the feature points mapped to the same second coordinate data.
4. The method according to claim 3, characterized in that, The obstacle prediction based on the feature data of the planar feature fusion image includes: The second feature data, the third feature data, and the fourth feature data are merged into the fifth feature data; Obstacle prediction is performed based on the fifth feature data.
5. The method according to claim 1, characterized in that, The electronic device is also equipped with a first model, which is trained based on the third coordinate data of the first sample in the vehicle coordinate system and the fourth coordinate data of the second sample mapped to the vehicle coordinate system in the following manner; First adjustment data is obtained by adjusting the fourth coordinate data based on the first parameter; Obtain the similarity between the first adjusted data and the third coordinate data; Based on the similarity between the first adjusted data and the third coordinate data, the first parameter is modified to a second parameter so that the similarity between the fourth coordinate data adjusted by the second parameter and the third coordinate data is greater than or equal to the similarity threshold. Use the second parameter as the model parameter of the first model; Wherein, the first sample is point cloud data collected by the test vehicle in the first environment, and the second sample is feature data of feature points in multiple image samples collected by the test vehicle in the first environment; The electronic device determines the height range in the vehicle coordinate system corresponding to the feature points in each sub-region according to the following method: The first coordinate data of the feature point is input into the first model to adjust the height data of the first coordinate data of the feature point; Based on the adjusted first coordinate data of the feature points, the height range of the feature points in each sub-region corresponding to the vehicle coordinate system is determined.
6. The method according to claim 1, characterized in that, The feature points mapped to the first sub-region of the plurality of sub-regions form feature segments at different heights relative to the horizontal plane of the vehicle coordinate system, wherein any feature point in a feature segment has more than a preset number of feature points in a region of height M. The height range of the first feature segment in the first sub-region in the vehicle coordinate system is taken as the height range of the feature point in the first sub-region corresponding to the vehicle coordinate system. The first feature segment is the feature segment in the first sub-region that is the lowest relative to the horizontal plane of the vehicle coordinate system.
7. The method according to claim 6, characterized in that, The first feature data also includes category information, which includes ground category and suspended obstacle category; If the height of the feature point corresponding to the lowest height in the first feature segment is higher than the height of the feature point of the ground category, the category of the feature point in the first feature segment is set to the suspended obstacle category.
8. The method according to claim 4, characterized in that, The obstacle prediction based on the fifth feature data includes: Obtain the fifth feature data corresponding to the set of feature points within a preset distance from the vehicle; Increase the resolution of the fifth feature data corresponding to the feature point set to obtain the sixth feature data; Based on the sixth feature data, obstacle prediction is performed on the environment surrounding the vehicle.
9. An electronic device, characterized in that, Includes: memory, used to store instructions; At least one processor is configured to execute the instructions to cause the electronic device to implement the method of any one of claims 1 to 8.
10. A computer program product, characterized in that, When the computer program product is run on the device, it causes the device to perform the method of any one of claims 1 to 8.
Citation Information
Patent Citations
Obstacle detection method and device, electronic equipment and storage medium
CN119068462A
Parking area determination method and device, equipment and storage medium
CN119540906A