Obstacle prediction method, program product, and electronic device

By generating a planar feature fusion image to explicitly express the height information of feature points, the problem of inaccurate prediction of suspended obstacles by on-board equipment is solved, thereby improving parking safety.

CN120808311AActive Publication Date: 2025-10-17CONTINENTAL SMART CORE TECH (SHANGHAI) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511240411.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2025-10-17
Estimated Expiration
2045-09-01

AI Technical Summary

Technical Problem

Existing on-board equipment has low accuracy in predicting suspended obstacles in parking scenarios, which makes it easy for vehicles to collide with suspended obstacles and affect parking safety.

Method used

By acquiring multiple images from different perspectives, the three-dimensional coordinate data of the feature points is determined, a plane feature fusion image is generated, and the height information of the feature points is explicitly expressed, thereby improving the accuracy of identifying suspended obstacles.

Benefits of technology

The prediction accuracy of suspended obstacles is improved, the probability of collision between vehicles and suspended obstacles is reduced, and parking safety is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808311A_ABST
    Figure CN120808311A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an obstacle prediction method, a program product and electronic equipment. The obstacle prediction method comprises the following steps: the electronic equipment determines first feature data of feature points on images according to a plurality of images, wherein the first feature data comprises first coordinate data of the feature points in a vehicle coordinate system and visual information of the feature points; the electronic device can project the first feature data to the horizontal plane so as to obtain second feature data, and the second feature data comprises height information of the highest feature point relative to the horizontal plane and height information of the lowest feature point relative to the horizontal plane. And then the electronic equipment performs obstacle prediction based on the second feature data, and the second feature data can highlight the height information of the feature points, so that the accuracy of identifying the suspended obstacle by the electronic equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to an obstacle prediction method, a program product and an electronic device. BACKGROUND

[0002] In a parking scenario, a vehicle usually needs a vehicle-mounted device to predict the occupancy of obstacles (such as the occupancy of other vehicles, walls, pillars and the like) around the vehicle to ensure the safety of the vehicle in parking.

[0003] However, when the vehicle-mounted device currently predicts the environment around the vehicle, it does not make corresponding processing for some specific obstacles (such as obstacles in the air, such as fire boxes on pillars, parking lot entrance and exit lifting rods and the like). The prediction accuracy of the vehicle-mounted device for the obstacles in the air around the vehicle is not high, and the vehicle is prone to collision when parking, thereby affecting the parking safety and causing a poor user experience. SUMMARY

[0004] The embodiments of the present application provide an obstacle prediction method, a program product and an electronic device.

[0005] In a first aspect, the embodiments of the present application provide an obstacle prediction method applied to an electronic device, which includes: the electronic device acquires multiple images, wherein the multiple images include images of different perspectives collected by a sensor of a vehicle. The electronic device determines first feature data of feature points of each image in the multiple images, wherein the first feature data includes first coordinate data of the feature points in a vehicle coordinate system, and the first coordinate data includes a height value of the feature points. For example, the first coordinate data is three-dimensional data of the feature points in the vehicle coordinate system. The electronic device generates a planar feature fusion image corresponding to the multiple images based on the first feature data of each image in the multiple images, wherein the planar feature fusion image includes mapping feature points of the feature points of each image in the multiple images mapped to a horizontal plane of the vehicle coordinate system. The electronic device can generate feature data of the planar feature fusion image based on the first feature data of each image in the multiple images, wherein the feature data of the planar feature fusion image includes second feature data, and the second feature data includes: region coordinate data of each sub-region in a plurality of sub-regions obtained by dividing the planar feature fusion image, and a height range of the feature points in each sub-region corresponding to the vehicle coordinate system. Then, the electronic device can perform obstacle prediction based on the feature data of the planar feature fusion image.

[0006] In some embodiments of the present application, the electronic device can map the first coordinate data of the feature points in each of the plurality of images to the horizontal plane of the vehicle coordinate system to generate a planar feature fusion image. Then, based on the first feature data of each feature point, the feature data of each mapped feature point on the planar feature fusion image is determined, which includes the second coordinate data of each mapped feature point and the height range data of each feature point on the image corresponding to each mapped feature point in the vehicle coordinate system. Therefore, the planar feature fusion image has a higher attention to the height information of each feature point in the vehicle coordinate system, so that the electronic device has more or more accurate height information of each feature point when predicting obstacles based on the planar feature fusion image, so as to ensure that the electronic device can more accurately predict the occupancy of the suspended obstacle, thereby improving the safety of the vehicle driving or parking.

[0007] In a possible implementation of the first aspect, the feature data of the planar feature fusion image further includes third feature data, and the method further includes: generating the third feature data of the planar feature fusion image based on the first feature data of each of the plurality of images, wherein the third feature data includes: the second coordinate data of each mapped feature point in the planar feature fusion image and first encoding data; the second coordinate data is two-dimensional coordinate data of the first coordinate data of the feature point mapped to the horizontal plane of the vehicle coordinate system, and the first encoding data is combined data of the visual information and the height data of the feature point mapped to the same second coordinate data.

[0008] In some embodiments of the present application, the electronic device can also combine the height data and the visual information of the first coordinate data in the first feature data to map to the planar feature fusion image, and the three-dimensional coordinate information of the feature point is retained. Therefore, the electronic device can improve the prediction accuracy of the overall environment when predicting obstacles based on the planar feature fusion image.

[0009] In a possible implementation of the first aspect, the method further includes: generating fourth feature data of the planar feature fusion image based on the first feature data of each of the plurality of images, wherein the fourth feature data includes: the second coordinate data of each mapped feature point in the planar feature fusion image and second encoding data; the second encoding data is the fusion data of the visual information of the feature point mapped to the same second coordinate data.

[0010] In a possible implementation of the first aspect, the obstacle prediction based on the feature data of the planar feature fusion image includes: fusing the second feature data, the third feature data and the fourth feature data into fifth feature data; and predicting obstacles based on the fifth feature data.

[0011] In the embodiments of the present application, the second feature data can enhance the height information in the fifth feature data, the third feature data can enable the fifth feature data to retain more environmental information, and the fourth feature data can enable the fifth feature data to retain more accurate visual information, thereby ensuring the accuracy of the overall environment prediction of the fifth feature data and the accuracy of the prediction of the suspended obstacle.

[0012] In a possible implementation of the first aspect, the electronic device is further configured with a first model, the first model being trained based on third coordinate data of a first sample in a vehicle coordinate system and fourth coordinate data of a second sample mapped into the vehicle coordinate system according to the following manner: adjusting the fourth coordinate data based on a first parameter to obtain first adjustment data; obtaining a similarity between the first adjustment data and the third coordinate data; modifying the first parameter to a second parameter based on the similarity between the first adjustment data and the third coordinate data, so that the similarity between the fourth coordinate data adjusted by the second parameter and the third coordinate data is greater than or equal to a similarity threshold; and taking the second parameter as a model parameter of the first model; and the electronic device determines the height range of the feature point in each sub-region in the vehicle coordinate system according to the following manner: inputting the first coordinate data of the feature point into the first model to adjust the height data of the first coordinate data of the feature point; and determining the height range of the feature point in each sub-region in the vehicle coordinate system based on the adjusted first coordinate data of the feature point.

[0013] In some embodiments of the present application, the point cloud sample data and the image sample can be collected by a data collection vehicle. For example, the sensors in the data collection vehicle can include a camera and a laser radar. The camera can obtain image samples of the environment around the data collection vehicle, and the laser radar can obtain point cloud data of the environment around the data collection vehicle. Then, the server or other devices such as the server can train the first model based on the point cloud sample data and the image sample. The point cloud sample data can be used as real data, and the server can process the image sample to determine the height range sample data in the second feature data from the image sample. Then, the server can train the model parameters of the first model based on the height range sample data and the point cloud data until the similarity between the height range sample data adjusted by the second parameter and the point cloud sample data is greater than or equal to a similarity threshold, thereby completing the training of the first model. In this way, the first model can be deployed on a vehicle without a laser radar. When the electronic device on the vehicle determines the height range based on the image data, the height range data can be adjusted based on the first model, thereby improving the accuracy of the height range data and the accuracy of the second feature data.

[0014] In a possible implementation of the first aspect, the feature points in the first sub-region of the plurality of sub-regions form feature segments at different heights relative to the horizontal plane of the vehicle coordinate system, and any feature point in a feature segment has more than a preset number of feature points in a region at a height M; a height range of a first feature segment in the first sub-region in the vehicle coordinate system is taken as a height range of the feature points in the first sub-region in the vehicle coordinate system, and the first feature segment is a lowest feature segment relative to the horizontal plane of the vehicle coordinate system in the first sub-region.

[0015] In some embodiments of the present application, if the first sub-region includes a plurality of feature segments, the feature segment closest to the horizontal plane has a higher probability of collision with the vehicle, and therefore, the height range data can be determined from the first feature segment closest to the horizontal plane. This ensures that the electronic device is more accurate in identifying the first feature segment, thereby avoiding collision between the vehicle and the obstacle corresponding to the first feature segment.

[0016] In a possible implementation of the first aspect, the visual information further includes category information, and the category information includes a ground category and a suspended obstacle category; the category of the feature points in the first feature segment is set to the suspended obstacle category, if the height of the feature point corresponding to the lowest feature point in the first feature segment is higher than the height of the feature point in the ground category.

[0017] In some embodiments of the present application, the electronic device separately classifies the first category corresponding to the suspended obstacle, which can further improve the accuracy of the electronic device in identifying the suspended obstacle, thereby ensuring the driving safety of the vehicle.

[0018] In a possible implementation of the first aspect, the obstacle prediction of the environment around the vehicle based on the fifth feature data includes: obtaining fifth feature data corresponding to a feature point set within a preset distance from the vehicle. The resolution of the fifth feature data corresponding to the feature point set is improved to obtain sixth feature data. The environment around the vehicle is predicted based on the sixth feature data.

[0019] In some embodiments of the present application, the electronic device needs to accurately identify the obstacle close to the vehicle, and therefore, after obtaining the fifth feature data, the electronic device can obtain a feature point set within a preset distance from the vehicle, and improve the resolution of the fifth feature of the feature point set to obtain the sixth feature data. The electronic device can more accurately identify the obstacle within the preset distance from the vehicle based on the sixth feature data.

[0020] In a second aspect, the present application provides an electronic device, comprising: a memory, configured to store instructions; and at least one processor, configured to execute the instructions to enable the device to implement the method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable by the second aspect can be referred to the beneficial effects of the method provided in any embodiment of the first aspect, which will not be repeated here.

[0021] In a third aspect, the present application provides a computer program product, which, when running on a device, enables the device to implement the method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable by the third aspect can be referred to the beneficial effects of the method provided in any embodiment of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 A schematic diagram of a vehicle parking process is shown; Figure 2 According to some embodiments of the present application, an implementation flowchart of obstacle prediction by a vehicle-mounted device is shown; Figure 3A According to some embodiments of the present application, a schematic diagram of processing multi-view images by a vehicle-mounted device is shown; Figure 3B According to some embodiments of the present application, a schematic diagram of obtaining fifth feature data by a vehicle-mounted device is shown; Figure 4 According to some embodiments of the present application, an implementation flowchart of obtaining ground truth is shown; Figure 5 According to some embodiments of the present application, a structural schematic diagram of an electronic device 100 is shown. DETAILED DESCRIPTION

[0023] The illustrative embodiments of the present application include, but are not limited to, an obstacle prediction method, a program product and an electronic device.

[0024] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] As shown in the background, in the parking process of a vehicle, there is no targeted processing of information about a suspended obstacle, so that the information about the suspended obstacle predicted by the vehicle is not accurate enough, which leads to the vehicle being prone to collision with the corresponding suspended obstacle in the parking process, thereby affecting the safety of the vehicle parking.

[0026] Next, the scenario of vehicle parking is introduced.

[0027] For example, Figure 1A schematic diagram of a vehicle parking process is shown.

[0028] like Figure 1 As shown, during parking, vehicle 10 can capture multiple images of its surroundings using multiple sensors onboard (e.g., a front-facing camera, side-facing cameras, and a rear-facing camera). Onboard equipment (e.g., a computer) on vehicle 10 can use these images to predict obstacles in the surroundings. For example, obstacles surrounding vehicle 10 include wall 1 20, wall 2 30, and a fire extinguisher 21 on wall 1 20. It should be understood that fire extinguisher 21 is fixed to wall 1 20, and its bottom is not in contact with the ground, meaning it is a suspended obstacle.

[0029] The on-board device typically extracts image features from multiple images using a neural network model. Based on the depth information (the distance of each pixel in the image from the vehicle 10), image features, and the camera's intrinsic and extrinsic parameter matrices corresponding to the images from each perspective, it fuses this information through a view transform operation to construct a 3D feature volume of the vehicle's surroundings. It is understood that the 3D feature volume includes information about objects in the vehicle's surroundings (e.g., basic information such as the object's position, size, and shape, and may also include more advanced information such as the object's texture, color, and surface reflectivity). The on-board device then performs occupancy prediction on the vehicle's surroundings based on the 3D feature volume, thereby determining the occupancy status of obstacles in the vehicle's surroundings. It is understood that the information in the 3D feature volume is still an abstract, high-level feature representation, rather than direct obstacle occupancy information. Occupancy prediction converts the information in the 3D feature volume into a concrete, understandable occupancy status (e.g., whether the space is occupied by an object).

[0030] However, in the process of generating the 3D feature body by fusing the image features, depth information, and internal and external parameter matrices of multiple images through the view transformation operation, some special features (such as the suspended obstacle) are not processed specifically, and the accuracy of the information of each object in the 3D feature body is more determined based on the proportion of the pixel points occupied by each object in multiple images. For example, the proportion of the pixel points occupied by wall one 20 and wall two 30 in multiple images is relatively large, and therefore, the information of wall one 20 and wall two 30 obtained by the vehicle-mounted device from multiple images is more, so that the spatial information of wall one 20 and wall two 30 generated is more accurate. The fire box 21 occupies fewer pixel points in multiple images, resulting in that the information of the fire box 21 that can be provided by multiple images is less, and the attention of the vehicle-mounted device to the fire box 21 in the process of generating the 3D feature body from multiple images is not enough, resulting in that the information of the fire box 21 in the 3D feature body is not accurate enough. For the suspended obstacle such as the fire box 21, it is more likely to collide with the body of the vehicle 10. Therefore, in the case that the information of the fire box 21 generated by the vehicle-mounted device from multiple images is not accurate, the risk of collision between the vehicle 10 and the fire box 21 in the parking process is relatively high.

[0031] As described above, in the process of predicting the obstacle in the environment around the vehicle by the vehicle-mounted device on the vehicle based on multiple images of the surrounding environment collected by the sensor, the information of the suspended obstacle is not processed specifically, resulting in that the information of the suspended obstacle predicted by the vehicle-mounted device is not accurate enough, thereby bringing a great safety hazard to the parking of the vehicle.

[0032] In order to solve the problem that the information of the suspended obstacle predicted by the vehicle-mounted device is not accurate enough, the present application provides an obstacle prediction method. An electronic device obtains multiple images, wherein the multiple images include images of different perspectives collected by a sensor of a vehicle (for example, images collected by a camera on the vehicle). The electronic device determines first feature data of feature points (for example, each pixel point of the image) of each image in the multiple images, wherein the first feature data includes first coordinate data of the feature points in a vehicle coordinate system, and the first coordinate data includes a height value of the feature points. For example, the first coordinate data is three-dimensional data of the feature points in the vehicle coordinate system.

[0033] The electronic device generates a planar feature fusion image corresponding to the plurality of images based on the first feature data of each image in the plurality of images, wherein the planar feature fusion image includes a mapping feature point of a feature point of each image in the plurality of images mapped to a horizontal plane of a vehicle coordinate system. The electronic device can generate feature data of the planar feature fusion image based on the first feature data of each image in the plurality of images, wherein the feature data of the planar feature fusion image includes second feature data, and the second feature data includes region coordinate data of each sub-region in a plurality of sub-regions obtained by dividing the planar feature fusion image, and a height range of a feature point in each sub-region corresponding to the vehicle coordinate system. Then, the electronic device can perform obstacle prediction based on the feature data of the planar feature fusion image.

[0034] Through the above scheme, the electronic device can perform obstacle prediction based on the planar feature fusion image corresponding to the plurality of images. The second feature data in the planar feature fusion image includes a height range of a feature point of each region on the horizontal plane of the vehicle coordinate system. Therefore, the planar feature fusion image has a higher attention degree on the height information of each feature point in the vehicle coordinate system, so that the electronic device has more or more accurate height information of each feature point when performing obstacle prediction based on the planar feature fusion image, so as to ensure that the electronic device can more accurately predict the occupancy of the suspended obstacle, and further improve the safety of the vehicle driving or parking.

[0035] Next, the process of obstacle prediction performed by the electronic device in some embodiments of the present application is introduced.

[0036] For example, Figure 2 According to some embodiments of the present application, an implementation flowchart of obstacle prediction performed by a vehicle-mounted device is shown.

[0037] It can be understood that the following flowcharts can be executed by an electronic device, which can be a vehicle-mounted device or a server. The vehicle-mounted device can be a mobile phone, a car machine, a terminal in self driving, and a wireless terminal in transportation safety. In the following, the vehicle-mounted device is taken as an example to introduce the obstacle prediction method in the embodiments of the present application.

[0038] As Figure 2 shown, the flowchart includes: S201, obtaining a plurality of images, wherein the plurality of images include images of different perspectives collected by sensors of a vehicle.

[0039] Exemplarily, in some embodiments of the present application, the vehicle can collect images of different perspectives of the environment where the vehicle is located by sensors during driving or parking. The sensors of the vehicle can be cameras, for example, and the vehicle can be equipped with cameras of different perspectives (e.g., front-view cameras, side-view cameras, rear-view cameras, etc.). The cameras can collect images of the environment around the vehicle. When predicting obstacles, the vehicle device can obtain multiple images collected by the cameras of different perspectives.

[0040] In some embodiments of the present application, the images collected by the cameras of different perspectives are, for example, time-sequenced, for example, after the function of assisting driving of the vehicle is turned on, the cameras can collect image data at all times. For example, , wherein is a set of images of surround view collected at all times, is an image collected by the Nth camera, N is the number of surround view images, i.e., there are N cameras, and t is the current time. In some embodiments of the present application, N is greater than 4, for example, N can be 4, 5, 6, 7, or 8, etc., i.e., the vehicle is equipped with 4, 5, 6, 7, or 8 cameras.

[0041] S202, determining first feature data of feature points of each image in the multiple images, wherein the first feature data comprises first coordinate data of the feature points in the vehicle coordinate system, and the first coordinate data comprises a height value of the feature points.

[0042] In some embodiments of the present application, each image in the multiple images further comprises multiple feature points, for example, the feature points can be pixel points of the image, in other embodiments, the feature points can also be pixel points corresponding to specific targets, or pixel points in a preset region on the image are taken as feature points, or feature points added on the image.

[0043] It can be understood that the coordinates of the feature points on each image are two-dimensional coordinates, i.e., corresponding to the width and height of the image, and the first coordinate data of the feature points is three-dimensional coordinate data in the vehicle coordinate system, therefore, it is necessary to map each feature point from the corresponding image coordinate system to the vehicle coordinate system.

[0044] In some embodiments of the present application, after the vehicle-mounted device obtains multiple images, it can extract the features of the feature points on each image to obtain the position data and visual information of the feature points in the image. For example, in some embodiments, the visual information can be the color, texture, category, and light reflection intensity of the feature points, etc. The vehicle-mounted device can determine whether the feature points are obstacles based on the visual information of the feature points, so as to predict the occupancy of the obstacles in the environment around the vehicle. In addition, the vehicle-mounted device can also perform view transformation on the feature points of each image based on the position data of the feature points in the image, so as to convert the position data of the feature points from the image coordinate system to the vehicle coordinate system.

[0045] For example, the conversion process is to convert the feature points from the image coordinate system (w, h) to the sensor coordinate (w', h', d) of the sensor based on the parameters (such as focal length, pixel size, image center coordinates, etc.) of the sensor for collecting the image. Wherein, w represents the width of the image coordinate system, h represents the height of the image coordinate system; w' represents the width of the sensor coordinate system, h' represents the height of the sensor coordinate system, and d represents the depth of the sensor coordinate, for example, the distance from the sensor.

[0046] Then, based on the parameters (such as the position of the sensor on the vehicle and the rotation angle, etc.) of each sensor relative to the vehicle, the feature points are converted from the sensor coordinate to the vehicle coordinate system (x, y, z). Wherein, x represents the length in the vehicle coordinate system, y represents the width in the vehicle coordinate system, and z represents the height in the vehicle coordinate system. In some embodiments of the present application, the visual information of the feature points is C, and the first feature data of the feature points is (x, y, z, C).

[0047] For example, Figure 3A According to some embodiments of the present application, a schematic diagram of a vehicle-mounted device processing multiple-view images is shown.

[0048] As Figure 3A shown, after the vehicle-mounted device obtains multiple-view images, it can input the multiple-view images into the backbone network to extract image features, which can include the position data and visual information of each feature point. The vehicle-mounted device then inputs the image features into the context head and the depth head to respectively encode the image features and depth features , wherein and is the resolution of the image features (for example, the width w and the height h of the image coordinate system), is the number of channels of the image features (for example, the visual information C of the feature points), For discrete depth quantity (e.g. depth d corresponding to sensor coordinate point), i represents the i-th image. It can be understood that the number of image feature channels For example, the kind of visual information can be feature points, for example, when the visual information includes color, texture, category and light reflection intensity corresponding to the feature points. The discrete depth quantity For example, the discrete value of the distance between the feature points and the sensor (e.g. camera, hereinafter the camera is taken as an example) can be the discrete depth quantity, for example, in some embodiments of the present application, the discrete depth quantity is 24, representing 0m-23m, and the interval between the discrete depth values is 1m. In other embodiments, the discrete depth quantity may also be other numerical values, for example, 30, 40, 50, etc. The embodiments of the present application do not limit the discrete depth quantity and are not limited.

[0049] In some embodiments, in the process of encoding the depth feature, the depth network can adjust the depth feature based on the depth ground truth through the depth ground truth estimation module. The process of adjusting the depth feature by the depth ground truth estimation module is described in detail below.

[0050] The vehicle-mounted device converts the image feature and the depth feature into the view angle conversion module, and the view angle conversion module can convert the feature points from the image coordinates to the vehicle coordinate system according to the internal parameters and external parameters of each camera.

[0051] For example, the vehicle-mounted device can construct the feature body of the feature points of the i-th image by the outer product of the depth feature of the i-th image and the image feature , for example, refer to formula (1): (1); wherein, is the feature body of the pixel coordinates of the feature points in the i-th image, and the feature body includes the position data, the depth data and the number of feature channels on the i-th image, is the image feature of the feature points on the i-th image, and the image feature includes the position data and the feature channel data of the pixel points on the i-th image, is the depth data of the feature points in the i-th image.

[0052] In some embodiments of the present application, the vehicle-mounted device can project the feature points on the multi-view images into the 3D space of the vehicle coordinate system according to a lift-splat-shot (LSS) scheme based on the positions of the feature points on the images. For example, the vehicle-mounted device can first generate a view cone according to the position data of each feature point in the image, and convert the view cone from the image coordinate system to the vehicle coordinate system. Illustratively, the view cone can be the 3D coordinates of the feature point at different discrete depths of the image , the view cone is a coordinate formed by adding the depth coordinate to the image pixel coordinate system, where 3 represents the x', y', z' coordinates in the view cone, is a discrete depth value, and is the image feature resolution, B is the total number of images, and N is the number of surround-view cameras. The vehicle-mounted device can convert the view cone to the vehicle coordinate system according to the internal and external parameters of the cameras , Proj represents the mapping relationship between the vehicle coordinate system and the pixel coordinates of the image, where represents the inverse normalization of the pixel coordinates. K is the internal parameter of the camera, which can include the focal length, pixel size, image center coordinates, etc. of the camera. R is the external parameter of the camera, which can include the position and rotation angle of the camera on the vehicle, etc. Through the internal parameter K of the camera, the feature point can be converted from the pixel coordinates to the coordinate system of the camera, and through the external parameter R of the camera, the feature point can be converted from the coordinate system of the camera to the vehicle coordinate system.

[0053] In some embodiments of the present application, the vehicle coordinate system can be rasterized in order to fuse the feature points of each view projected into the vehicle coordinate system. For example, the resolution of the rasterized vehicle coordinate system can be WxHxL, where W can be the grid resolution in the x direction, H can be the grid resolution in the y direction, and L can be the grid resolution in the z direction.

[0054] According to the resolution of the vehicle coordinate system, Proj is converted to the mapping relationship between the rasterized vehicle coordinate system and the pixel coordinates of the image , and based on this, the image feature points are sampled, the features of the multiple feature points falling into the same grid are calculated by summation, thereby obtaining the first feature data, for example, refer to equation (2): (2); wherein is the first feature data, is the mapping relationship between the rasterized vehicle coordinate system and the pixel coordinates of the image, wherein, i is the index of the image, N is the number of the surround-view cameras, i.e. N images, H, W, L are the resolutions of the grid in the vehicle coordinate system, and C is the number of the feature channels. It can be understood that C is the feature channel data of the feature points in the grid of the vehicle coordinate system. In some embodiments of the present application, the positions of the feature points in the grid of the vehicle coordinate system can also be the first coordinate data in the first feature data, and the feature channel data corresponding to the feature points can be the visual information of the feature points.

[0055] For example, in some embodiments of the present application, the visual information further includes category information, and the category information includes a ground category and a first category, wherein the first category is used to indicate that the corresponding feature point is a feature point of a suspended obstacle. The feature points with the height information higher than the ground category can be the feature points of the first category. It can be understood that in some embodiments of the present application, the vehicle-mounted device separately classifies the first category corresponding to the suspended obstacle, which can further improve the accuracy of the vehicle-mounted device in identifying the suspended obstacle, so as to ensure the driving safety of the vehicle.

[0056] In some embodiments of the present application, after the vehicle coordinate system is gridded, the coordinates of multiple feature points can be located in the same grid, and therefore, the multiple feature points falling into the same grid can be summed up to ensure that there is at most one feature point in each grid. For example, the visual information of the multiple feature points in one grid can be summed up, and in other embodiments, the average value of the visual information of the feature points can be taken as the visual information of the feature point of the grid. The present application does not limit the merging process of the multiple feature points in one grid. If the visual information includes category information, the union set can be taken. For example, two feature points are included in one grid, one of which corresponds to the ground category, and the other does not correspond to the ground category, and the category of the feature point corresponding to the grid is the ground category.

[0057] S203, generating a planar feature fusion image corresponding to the multiple images based on the first feature data of each image in the multiple images, wherein the planar feature fusion image includes the mapping feature points of the feature points of each image in the multiple images mapped onto the horizontal plane of the vehicle coordinate system.

[0058] In some embodiments of the present application, the first feature data of each of the plurality of images includes first coordinate data of feature points in each of the plurality of images in the vehicle coordinate system. Since the first coordinate data is in a three-dimensional coordinate system, the height data of the feature points in the vehicle coordinate system is not easy to express explicitly. Therefore, the three-dimensional coordinates of each feature point in the vehicle coordinate system can be mapped to the horizontal plane of the vehicle coordinate system, thereby generating a plane feature fusion image corresponding to the plurality of images. Each mapped feature point on the plane feature fusion image can correspond to a height range, thereby explicitly expressing the height data of the feature points in each image, so that the electronic device can more accurately identify the floating obstacles according to the plane feature fusion image.

[0059] For example, the plane formed by the x direction and the y direction of the vehicle coordinate system can be the horizontal plane of the vehicle coordinate system. That is, the coordinates of each mapped feature point on the plane feature fusion image are two-dimensional coordinates (x, y). It can be understood that each mapped feature point on the plane feature fusion image can be a point corresponding to the same two-dimensional coordinate pair on the plane feature fusion image. That is, one mapped feature point can include multiple feature points with different heights (such as different z coordinates).

[0060] S204, generating feature data of the plane feature fusion image based on the first feature data of each of the plurality of images, wherein the feature data of the plane feature fusion image includes second feature data, and the second feature data includes: region coordinate data of each sub-region in the plurality of sub-regions obtained by dividing the plane feature fusion image, and a height range corresponding to the feature points in each sub-region in the vehicle coordinate system.

[0061] In some embodiments of the present application, the resolution of the grid of the vehicle coordinate system is (W, H, L), so the resolution of the plane grid on the horizontal plane of the vehicle coordinate system is (W, H), that is, the resolution of the plane grid of the plane feature fusion image is (W, H). In some embodiments of the present application, one plane grid can be regarded as one sub-region. The coordinate data of each plane grid can be the two-dimensional coordinate data of the center point of each plane grid in the plane feature fusion image, or the two-dimensional coordinate data of the left bottom corner vertex and the right top corner vertex of each plane grid in the plane feature fusion image. In some embodiments, the coordinate data of the sub-region can also be the position coordinates of the plane grid. For example, the coordinates of one plane grid are (W1, H1), which means the plane grid is in the W1th row in the x direction and the H1th column in the y direction.

[0062] It can be understood that, since each grid in the vehicle coordinate space includes at most one feature point, and one plane grid in the plane feature image also includes at most one mapped feature point, in some embodiments, the coordinate data of the sub-region can also be two-dimensional coordinate data of each mapped feature point in the plane feature fusion image, or the two-dimensional coordinate data of each mapped feature point in the plane feature fusion image can also be coordinate data of each sub-region in the plane feature fusion image.

[0063] It can be understood that the coordinate data corresponding to the plane grid is the position of each plane grid on the plane fusion image, and the height range of a sub-region can be the height range of the feature points in the grid with different L coordinates at the same position of the plane grid. For example, the coordinates of n feature points in the grid of the vehicle coordinate space are (W1, H1, L1), (W1, H1, L2), (W1, H1,...), (W1, H1, Ln), and after mapping to the plane feature fusion image, the n feature points are all mapped to the coordinate of the plane grid (W1, H1). The height range of the plane grid (W1, H1) can be L1 to Ln, where Ln is the height position of the grid corresponding to the feature point with the highest height (z coordinate) among the n feature points, and L1 is the height position of the grid corresponding to the feature point with the lowest height (z coordinate) among the n feature points.

[0064] For example, in the process of determining the second feature data based on the first feature data, the vehicle-mounted device can combine the L dimension and the C dimension of the first feature data, and then use conv2d to encode the combined data into BxHxWxLx2 dimension (where 2 is to consider that the suspended obstacle needs to predict the maximum and minimum two height values (e.g., Ln and L1) for determining the height range of each two-dimensional grid.

[0065] For example, the vehicle-mounted device performs a softmax calculation on the corresponding height values in each sub-region in the L dimension to obtain the probability density of the maximum height and the minimum height, for example, referring to formulas (3) and (4): (3); (4); wherein, is the highest probability density of the second coordinate data of each sub-region on the horizontal plane, is the lowest probability density of the second coordinate data of each sub-region on the horizontal plane, is a kind of conv2d coding. That is, the second feature data includes the second coordinate data (corresponding to the coordinates of the grid on the HxW plane) of the sub-region of the plane feature fusion image (corresponding to the HxW plane) and the probability density corresponding to the second coordinate data and , and As an example of the height range, in some embodiments of the present application, the dimension of the planar feature fusion image can be the same as the pixel dimension of the image.

[0066] In some embodiments, the plurality of feature points of the first coordinate data mapped to the first sub-region form a plurality of feature segments at different heights from the horizontal plane, wherein the feature points in one feature segment have more than a preset number of adjacent feature points in the adjacent region at height M.

[0067] In some embodiments of the present application, the height range of the first feature segment with the lowest height in the plurality of feature segments corresponding to the first sub-region can be taken as the height range of the first sub-region.

[0068] It can be understood that if the first sub-region includes a plurality of feature segments, the feature segment with the lowest height from the horizontal plane has a higher probability of collision with the vehicle, and therefore the highest feature point and the lowest feature point in the first feature segment with the lowest height from the horizontal plane can be selected to determine the height range of the first feature segment. To ensure that the vehicle-mounted device is more accurate in identifying the feature segment, thereby avoiding collision between the vehicle and the obstacle corresponding to the feature segment.

[0069] In some embodiments of the present application, the vehicle-mounted device can also adjust the height range in the second feature data through the height true value estimation module. The specific adjustment process is described below.

[0070] S205, obstacle prediction based on the feature data of the planar feature fusion image.

[0071] In some embodiments of the present application, the feature data of the planar feature fusion image includes second feature data, and the second feature data includes the second coordinate data of the plurality of sub-regions and the height range data of each sub-region. The height range data can explicitly express the height information, and therefore the vehicle-mounted device can more accurately predict whether the obstacle is a suspended obstacle based on the height information when predicting the obstacle based on the planar feature image, thereby improving the prediction effect of the vehicle-mounted device on the suspended obstacle.

[0072] Through the above scheme, the electronic device can more accurately predict the occupancy of the corresponding suspended obstacle in the surrounding environment based on the planar feature fusion data, thereby reducing the probability of collision between the vehicle and the suspended obstacle during parking or driving, and improving the safety of the vehicle.

[0073] ​In some embodiments of the present application, the feature data of the planar feature fusion image further comprises third feature data, for example, the electronic device can generate the third feature data of the planar feature fusion image based on the first feature data of each of the plurality of images, wherein the third feature data comprises: second coordinate data and first encoding data of each mapping feature point in the planar feature fusion image. In some embodiments of the application, the second coordinate data is two-dimensional coordinate data of the feature point mapped to the horizontal plane of the vehicle coordinate system, that is, the two-dimensional coordinate data of the mapping feature point. That is, the second coordinate data can be the position coordinate data of the planar grid in the planar feature fusion image. The first encoding data is the combined data of the visual information and the height data of the feature point mapped to the same second coordinate data.

[0074] Exemplarily, in the vehicle coordinate system, the resolution of the grid is (W, H, L), and each grid includes only one feature point, so the first coordinate data of the feature point in the vehicle coordinate system can also be (W, H, L), and the first feature data can also be (W, H, L, C). After the feature points of each image are mapped to the planar feature fusion image, the coordinates of the corresponding mapping feature points are (W, H), and in the embodiments of the present application, the height coordinate L and the visual information C in the first feature data of the feature point of each image corresponding to each mapping feature point on the planar feature fusion image can be combined, and the combined feature can be M, that is, M is the combined data of C and L. It can be understood that since each mapping feature point can correspond to a plurality of feature points of each image with the same two-dimensional coordinates (W, H) and different height coordinates L, L in M is actually a high-dimensional array, for example, in the vehicle coordinate system, the resolution of L is how many, and the dimension of L in M is how many. That is, the third feature data is essentially (W, H, M), wherein (W, H) is the coordinate of the mapping feature point on the planar feature fusion image, and M is the first encoding data.

[0075] For example, in some embodiments of the present application, the dimensions of L and C of the first feature data can be combined by a reshape function, and the electronic device can further encode using a two-dimensional convolutional code (conv2d) to convert the combined data to dimension to obtain third feature data , for example, refer to formula (5): (5); wherein, is the third feature data, ( ) is a kind of conv2d coding, is the first feature data, reshape ( ) is reshape function. That is, the third feature data The second coordinate data of the mapping feature points of the feature points mapped onto the planar feature fusion image can be included, and the corresponding channel number The channel number is obtained by combining the height dimension L of the feature points and the feature channel number C based on the reshape function, and then performing two-dimensional convolution coding The corresponding channel data can be an example of the first encoding data. For example, in some embodiments, the L dimension in the first feature data is 2, and the feature channel number C is 3. After combining the L dimension and the feature channel number C based on the reshape, a dimension of 2*3=6 can be obtained. The dimension can be a new channel number of the third feature data. The new channel data is encoded by conv2d into the channel number. and The resolution of each feature point mapped onto the planar feature fusion image. In embodiments of the present application, the resolution of each feature point mapped onto the planar feature fusion image can be the same as the resolution of the planar grid, that is, W*H can be the same as * In other embodiments, the resolution of each feature point mapped onto the planar feature fusion image can also be other numerical values, and the present application does not limit the resolution of each feature point mapped onto the planar feature fusion image.

[0076] In some embodiments of the present application, the electronic device can also generate fourth feature data of the planar feature fusion image based on the first feature data of each image in the plurality of images, wherein the fourth feature data includes: second coordinate data of each mapping feature point in the planar feature fusion image and second encoding data, and the second encoding data is fusion data of the visual information of the feature points mapped onto the same second coordinate data.

[0077] For example, the coordinates of n feature points in the grid of the vehicle coordinate space are (W1, H1, L1), (W1, H1, L2), (W1, H1,...), (W1, H1, Ln), respectively. After the n feature points are mapped onto the planar feature fusion image, they are all mapped onto the coordinates of the planar grid (W1, H1). The second encoding data of the planar grid (W1, H1) is Cs, where Cs is the corresponding sum of the visual information of the n feature points. That is, the fourth feature data is (W, H, Cs).

[0078] For example, in some embodiments of the present application, the corresponding data of the feature channels of the feature points mapped onto the horizontal plane at the same second coordinate can be added by summation, and then the channel number is encoded into dimension by conv2d, so as to obtain the fourth feature data For example, refer to formula (6): (6); wherein, is the fourth feature data, ( ) is a conv2d encoding, and SUM( ) is a summation function. It can be understood that the fourth feature data includes the second coordinate data and the second encoding data of the planar feature fusion image.

[0079] After determining the third feature data and the fourth feature data of the planar feature fusion image, the second feature data, the third feature data and the fourth feature data can be fused into fifth feature data, and the electronic device can perform obstacle prediction based on the fifth feature data.

[0080] For example, Figure 3B According to some embodiments of the present application, a schematic diagram in which the vehicle-mounted device obtains the fifth feature data is shown.

[0081] As Figure 3B shown, the vehicle-mounted device can encode the height data of the highest point and the height data of the lowest point based on conv1x1 (representing a one-dimensional convolution kernel) and a normalized exponential function (softmax), thereby obtaining and . Then the vehicle-mounted device respectively multiplies and by the third feature data point through the attention mechanism, strengthening the features of the maximum height and the minimum height in the third feature data. Then, the vehicle-mounted device further encodes through conv3x3 (representing a 3-dimensional convolution kernel) respectively, and the encoded 2 groups of features are spliced with the fourth feature data. After encoding through conv3x3, they are spliced with the third feature data, and then extracted through a conv3x3 feature extraction, to obtain the final fifth feature data. For example, refer to formula (7): (7); wherein, is the fifth feature data, is a conv2d encoding, is a self-attention fusion function, wherein, , and CAT( ) is a splicing function.

[0082] Referring to Figure 3A , in some embodiments of the present application, the vehicle-mounted device can also input the fifth feature data into a time sequence fusion module, an encoding module and a channel conversion module respectively, to finally obtain a 3d feature.

[0083] It can be understood that the vehicle-mounted device can obtain multiple images at different times, and therefore, the fifth feature data determined according to the multiple images at different times can be fused to determine the speed and motion of each feature point and the like. Then, the encoding module can encode the speed and motion of the feature point to add the motion feature of the feature point. The channel conversion module can separate the feature channel in the fifth feature data into L and C, so as to restore the 3D feature. The vehicle-mounted device can also adjust the 3D feature based on the 3D occupancy ground truth, so that the vehicle-mounted device can more accurately predict the occupancy of the obstacles in the environment around the vehicle based on the 3D feature, and classify the obstacles. For example, determine whether the obstacle is a moving obstacle, a suspended obstacle, a pedestrian, a vehicle, and the like.

[0084] In some cases (for example, when the vehicle is parking), the obstacle prediction of the scene close to the vehicle needs high-resolution prediction accuracy, and the obstacle prediction of the scene far from the vehicle does not need high-resolution prediction accuracy, so in some embodiments of the present application, the accuracy requirements of different ranges can be met through the scheme of the local prediction head. For example, the local prediction head crops the feature points (as an example of a feature point set) of the local area close to the vehicle from the fifth feature data (or 3D feature) and performs up-sampling, so as to obtain the sixth feature data corresponding to the up-sampled feature points, thereby improving the resolution of the feature points in the local area close to the vehicle. The vehicle-mounted device predicts the occupancy information based on the sixth feature data, which can further improve the accuracy of the prediction close to the vehicle.

[0085] In summary, the scheme fuses the third feature data and the second feature data through the self-attention mechanism, which can strengthen the features of the maximum height and the minimum height in the second feature data. So that the fifth feature after fusion pays more attention to the height of the feature points, thereby improving the accuracy of the vehicle-mounted device in predicting suspended obstacles. And according to the height ground truth, the accuracy of the maximum height and the minimum height in the fifth feature data can be further improved to improve the effect of the vehicle-mounted device in predicting obstacles, thereby avoiding the collision between the vehicle and the suspended obstacle.

[0086] It can be understood that the vehicle-mounted device can obtain multiple images at different times, and therefore, the fifth feature data determined according to the multiple images at different times can be fused to determine the speed and motion of each feature point and the like. Then, the encoding module can encode the speed and motion of the feature point to add the motion feature of the feature point. The channel conversion module can separate the feature channel in the fifth feature data into L and C, so as to restore the 3D feature. The vehicle-mounted device can also adjust the 3D feature based on the 3D occupancy ground truth, so that the vehicle-mounted device can more accurately predict the occupancy of the obstacles in the environment around the vehicle based on the 3D feature, and classify the obstacles. For example, determine whether the obstacle is a moving obstacle, a suspended obstacle, a pedestrian, a vehicle, and the like. Figure 3A When the vehicle-mounted device obtains the depth feature, the second feature data and the fifth feature data, the depth ground truth, the height ground truth and the 3D occupancy ground truth are respectively obtained to adjust the depth feature, the second feature data and the fifth feature data respectively, so as to improve the accuracy of the corresponding data. In some embodiments of the present application, the process of obtaining the ground truth is introduced as follows.

[0087] Exemplarily, in some embodiments of the present application, the ground truth data can be acquired by a data collection vehicle. It can be understood that the sensors in the data collection vehicle include not only cameras with multiple perspectives, but also sensors such as lidars capable of collecting point cloud data. During the driving or parking process of the data collection vehicle, image data can be collected by cameras with multiple perspectives, and point cloud data with multiple perspectives can be collected by lidars. For example, in some embodiments of the present application, the data collection vehicle can collect point cloud data in 4 directions. In this way, the data collection vehicle can acquire point cloud data with multiple perspectives, so as to determine the depth ground truth, the height ground truth and the 3D occupancy ground truth according to the point cloud data.

[0088] Next, the process of determining the depth ground truth, the height ground truth and the 3D occupancy ground truth in some embodiments of the present application is introduced.

[0089] For example, Figure 4 According to some embodiments of the present application, an implementation flowchart for acquiring ground truth is shown.

[0090] Exemplarily, the execution subject of each of the following flows can be a vehicle-mounted device or an electronic device capable of processing point cloud data, such as a server.

[0091] As Figure 4 shown, the flow includes: S410, acquiring point cloud data.

[0092] Exemplarily, in some embodiments of the present application, not only multiple cameras are configured on the data collection vehicle, but also lidars are configured in multiple directions. During the driving or parking process of the data collection vehicle, image data can be collected by the cameras, and point cloud data in multiple directions can be collected by the lidars.

[0093] The server or the processor can acquire the image data and the point cloud data in different directions collected by the lidars, so as to process the point cloud data.

[0094] S420, determining the depth ground truth based on the point cloud data.

[0095] In some embodiments of the present application, the vehicle-mounted device can project the point cloud on the image coordinate system of the image collected by the camera with each perspective according to the internal parameters and the external parameters of the camera with each perspective, and filter the points beyond the image range according to the image resolution, to acquire the depth value corresponding to the image pixel. In some embodiments of the present application, the vehicle-mounted device is preset with a depth plane when processing the image, and for the convenience of adjusting the feature points on the image by the point cloud, the vehicle-mounted device can also discretize the depth value of the point cloud to ​WxHxL, and the depth ground truth is obtained. In some embodiments of the present application, the depth ground truth can be one-hot coded, and then the one-hot coded depth ground truth is used to adjust the one-hot coded data of the discrete depth of the feature points in the image WxHxL, and the depth ground truth is obtained. In some embodiments of the present application, the depth ground truth can be one-hot coded, and then the one-hot coded depth ground truth is used to adjust the one-hot coded data of the discrete depth of the feature points in the image WxHxL, and the depth ground truth is obtained. In some embodiments of the present application, the depth ground truth can be one-hot coded, and then the one-hot coded depth ground truth is used to adjust the one-hot coded data of the discrete depth of the feature points in the image WxHxL, and the depth ground truth is obtained. In some embodiments of the present application, the depth ground truth can be one-hot coded, and then the one-hot coded depth ground truth is used to adjust the one-hot coded data of the discrete depth of the feature points in the image

[0096] The determination process of the depth ground truth is described below.

[0097] S421, projecting the point cloud data in the image coordinate system according to the internal and external parameters of the camera.

[0098] Exemplarily, the vehicle-mounted device can project the acquired point cloud data to the coordinate of the image based on the internal and external parameters of the camera.

[0099] S422, filtering the point cloud beyond the image range according to the image resolution size.

[0100] Exemplarily, the point cloud data collected by the laser radar on the vehicle is more comprehensive, and an image of one perspective cannot cover all the point cloud data. Therefore, after the vehicle-mounted device projects the point cloud data to the image coordinate system, the point cloud beyond the image range can be filtered.

[0101] S423, discretizing the point cloud within the image resolution to WxHxL, and the depth ground truth is obtained. In some embodiments of the present application, the depth ground truth can be one-hot coded, and then the one-hot coded depth ground truth is used to adjust the one-hot coded data of the discrete depth of the feature points in the image

[0102] Exemplarily, after the vehicle-mounted device obtains the point cloud data within the image range, the point cloud data can be discretized to WxHxL, and the depth ground truth is obtained. In some embodiments of the present application, the depth ground truth can be one-hot coded, and then the one-hot coded depth ground truth is used to adjust the one-hot coded data of the discrete depth of the feature points in the image WxHxL, and the depth ground truth is obtained. In some embodiments of the present application, the depth ground truth can be one-hot coded, and then the one-hot coded depth ground truth is used to adjust the one-hot coded data of the discrete depth of the feature points in the image

[0103] S430, rasterizing the point cloud data, and the dimension of the rasterization is WxHxL.

[0104] In some embodiments of the present application, after the vehicle-mounted device acquires the point cloud data, the point cloud can also be rasterized, and the resolution of the rasterization is WxHxL. In this way, the point cloud data can correspond to the rasterized vehicle coordinate system.

[0105] S440, determining the 3D occupancy ground truth based on the rasterized point cloud.

[0106] For example, in order to better predict suspended obstacles, the vehicle-mounted device may set suspended obstacles as a separate category, and distinguish whether the obstacle corresponding to the feature point is a suspended obstacle based on the 3D occupancy true value.

[0107] In some embodiments of the present application, the vehicle-mounted device can determine the 3D occupancy true value based on the rasterized point cloud data. The following describes the process of obtaining the 3D occupancy true value.

[0108] S441, backup rasterized point cloud.

[0109] For example, the vehicle-mounted device may back up rasterized point cloud data.

[0110] S442 , traverse the cylinder corresponding to the W×H dimension of the rasterized point cloud.

[0111] For example, the dimension of the rasterized point cloud is W×H×L. In the process of determining the true value of 3D occupancy, the on-board equipment can traverse the rasterized point cloud and process each cylinder in the unit of one grid corresponding to one cylinder in the W×H dimension of the horizontal plane.

[0112] S443, determine whether the traversal is completed.

[0113] The vehicle-mounted device can determine whether all cylinders in the W×H dimension have been traversed. For example, the vehicle-mounted device can obtain the current cylinder and determine whether the current cylinder has been processed.

[0114] If the judgment result is yes, the vehicle-mounted device may execute S444 to obtain the true 3D occupancy value.

[0115] If the determination result is no, the vehicle-mounted device may execute S445 to determine whether the current column is a suspended obstacle.

[0116] S444, obtain the true 3D occupancy value.

[0117] For example, after the vehicle-mounted device traverses all cylinders in the W×H dimension, it can be determined that the vehicle-mounted device has processed all cylinders in the W×H dimension, so that the vehicle-mounted device can obtain the true 3D occupancy value.

[0118] S445: Determine whether the current column is a suspended obstacle.

[0119] For example, if the vehicle-mounted device determines that all cylinders in the W×H dimension have not been traversed, the current cylinder may be processed. For example, the process of processing the current cylinder may include determining whether the current cylinder is a suspended obstacle.

[0120] If the judgment result is yes, then S446 is executed to determine whether the point cloud in the current suspended obstacle occupies more than a preset number of point clouds in the adjacent M area.

[0121] If the result is no, the next column in the WxH dimension is obtained, and S443 is executed to determine whether the traversal is complete.

[0122] Illustratively, for a column in the WxH dimension, if the point with the minimum height in the column point cloud is higher than the ground height (the ground height can refer to the height coordinate of the point cloud classified as the ground category), it can be determined that the column contains a floating obstacle.

[0123] S446, determine whether the point cloud in the current floating obstacle has more than a preset number of point cloud occupancies in the adjacent M region.

[0124] Illustratively, in some embodiments, there can be interfering points (noise) in a column, so after determining that the current column is a floating obstacle corresponding column, it is also necessary to determine whether the point cloud in the column is noise. For example, in the window of the point cloud in the column, it is determined whether there are more than a threshold h number of point cloud occupancies in the neighborhood of m.

[0125] If the result is yes, S447 is executed to mark the category of the current column as a floating obstacle, and the processing of the current column is completed.

[0126] If the result is no, the next column in the WxH dimension is obtained, and S443 is executed to determine whether the traversal is complete.

[0127] It can be understood that if the point cloud in the subject has few point cloud occupancies in the neighborhood of m, the point cloud in the column can be noise, and the category of the current column is not changed.

[0128] S447, mark the category of the current column as a floating obstacle, and complete the processing of the current column.

[0129] Illustratively, after the vehicle-mounted device determines that there is a floating obstacle in the current column, the current column can be marked as a floating obstacle. Then the next column in the WxH dimension is obtained, and S443 is executed to determine whether the traversal is complete.

[0130] S450, determine the height true value based on the rasterized point cloud.

[0131] Illustratively, after the point cloud is rasterized, it includes three dimensions of W, H, and L, and each grid on the horizontal plane (WxH dimension) can be regarded as a column. The height value of the highest point cloud in each column and the height value of the lowest point cloud are determined as the height true value of the column corresponding to the grid in the HxW dimension. The process of obtaining the height true value is described below.

[0132] S451, set the array corresponding to the WxH dimension, and the default value of the array is ignore.

[0133] Exemplarily, after the point cloud is rasterized, each grid in the HxW dimension corresponds to a column, and the vehicle-mounted device can set an array on each grid in the HxW dimension, which is used to represent the height truth value in the corresponding grid. In the initial state, the array of each grid in the HxW dimension is set to the default value ignore.

[0134] S452, the column corresponding to the WxH dimension of the rasterized point cloud is traversed.

[0135] The vehicle-mounted device can traverse the column corresponding to the WxH dimension of the rasterized point cloud, thereby processing the array of each grid in the HxW dimension.

[0136] S453, it is judged whether the traversal is completed.

[0137] The vehicle-mounted device can judge whether the column corresponding to all grids in the WxH dimension has been traversed, for example, the vehicle-mounted device can judge whether the column corresponding to the current grid has been traversed.

[0138] If the judgment result is yes, the height truth value is obtained.

[0139] If the judgment result is no, S454 is executed to judge whether the current column is occupied by the point cloud.

[0140] S454, it is judged whether the current column is occupied by the point cloud.

[0141] Exemplarily, the vehicle-mounted device can judge whether the column corresponding to the current grid in the WxH dimension is occupied by the point cloud.

[0142] If the judgment result is yes, S456 is executed to judge whether the point cloud in the current column is continuously occupied.

[0143] If the judgment result is no, S455 is executed to set the array in the current grid to ignore.

[0144] S455, the array in the current grid is set to ignore.

[0145] Exemplarily, if the vehicle-mounted device judges that the column corresponding to the current grid in the WxH dimension is not occupied by the point cloud, it means that the current column has no obstacle, and the array corresponding to the current grid in the WxH dimension is set to ignore.

[0146] S456, it is judged whether the point cloud in the current column is continuously occupied.

[0147] Exemplarily, in some embodiments of the present application, the point cloud in the column body can be continuous occupation or non-continuous occupation. The continuous occupation means that there is only one continuous point cloud segment in the column body. In the point cloud segment, there are a preset number of adjacent point clouds in the range of height M. Therefore, if there is point cloud occupation in a column body, it is also necessary to judge whether the point cloud is continuously occupied, that is, whether there are multiple point cloud segments.

[0148] If the judgment result is yes, S457 is executed, the height index of the highest point and the height index of the lowest point in the current column body are set in the array of the current grid, and the processing of the current column body is completed.

[0149] If the judgment result is no, S458 is executed, in the point cloud segment with the lowest height, the height index of the highest point and the height index of the lowest point are set in the array of the current grid, and the processing of the current column body is completed.

[0150] S457, the height index of the highest point and the height index of the lowest point in the current column body are set in the array of the current grid, and the processing of the current column body is completed.

[0151] Exemplarily, the vehicle-mounted device judges that the point cloud of the current column body is continuously occupied, which means that there is only one point cloud segment in the current column body. The height index of the highest point in the point cloud segment is set as , the height index of the lowest point is set as , and and are assigned to the array of the grid with the dimension of WxH corresponding to the current column body. Thus, the processing of the height truth value of the current column body is completed, and the vehicle-mounted device can process the height truth value of the next column body.

[0152] S458, in the point cloud segment with the lowest height, the height index of the highest point and the height index of the lowest point are set in the array of the current grid, and the processing of the current column body is completed.

[0153] Exemplarily, the vehicle-mounted device judges that the point cloud of the current column body is not continuously occupied, which means that there are multiple point cloud segments in the current column body. The vehicle-mounted device can set the height index of the highest point in the point cloud segment with the lowest height as , set the height index of the lowest point in the point cloud segment with the lowest height as , and assign and to the array of the grid with the dimension of WxH corresponding to the current column body. Thus, the processing of the height truth value of the current column body is completed, and the vehicle-mounted device can process the height truth value of the next column body.

[0154] It can be understood that in the case of multiple point cloud segments in the column, the lower the height of the point cloud segment, the higher the possibility of collision with the vehicle, and therefore the position of the highest point and the lowest point on the lowest height point cloud segment can be assigned to the array of the grid in the current WxH dimension.

[0155] It can be understood that through the above process, the vehicle-mounted device can determine the depth ground truth, the height ground truth, and the 3D occupancy ground truth based on the point cloud data. Then the vehicle-mounted device can process the image data based on the above ground truths.

[0156] For example, with reference to Figure 3A Taking the height ground truth estimation module as an example, the process of adjusting the second feature data based on the ground truth is introduced.

[0157] For example, in the height ground truth estimation module (as an example of the first model), an adjustment parameter can be set, which is used to adjust the height range in the input second feature data.

[0158] Next, the training process of the height ground truth estimation module is introduced. For example, the height ground truth estimation module can be deployed to the server during training. The server can obtain image samples from the data collection vehicle, and determine fourth coordinate data mapped to the vehicle coordinate space according to the image samples. For example, the server can also determine height range sample data in the second feature data from the image samples according to the process of S204.

[0159] The server can obtain point cloud data from the data collection vehicle, and input the height ground truth (as an example of the point cloud sample) determined based on the third coordinate data of the point cloud data to the height ground truth estimation module. The height ground truth estimation module can adjust the height range sample according to the initial adjustment parameter (as an example of the first parameter) to obtain first adjustment data, and then determine the similarity between the first adjustment data and the height ground truth according to the cross-entropy loss, and correct the initial adjustment parameter in the height ground truth estimation module according to the similarity. Then the height ground truth estimation module adjusts the height range sample data based on the corrected adjustment parameter to improve the similarity between the height range sample data and the height ground truth. In this way, until the adjusted height range sample data and the height ground truth have a similarity greater than or equal to a similarity threshold after the adjusted height range sample data is adjusted by the corrected adjustment parameter, it can be determined that the training of the adjustment parameter in the height ground truth estimation module is completed, and the trained adjustment parameter is deployed to the height ground truth estimation module as a second parameter.

[0160] The trained height ground truth estimation module and the second parameter are deployed to the vehicle-mounted device. Figure 3AEach of the modules in the above embodiments can be deployed on the vehicle-mounted device of a common vehicle (a vehicle without a laser radar). After the vehicle-mounted device of the common vehicle determines the second feature data based on the image, the height range in the second feature data can be adjusted based on the height true value estimation module, so that the second feature data is more accurate.

[0161] It can be understood that the training process of the depth true value estimation module is similar to the training process of the height true value estimation module. When adjusting the depth true value, the common vehicle only needs to input the depth feature into the depth true value estimation module. Similarly, based on the training manner of the height true value estimation module, a model for adjusting the fifth feature data can be trained based on the 3D occupancy true value. The model can be deployed in the channel conversion module, so that the channel conversion module can adjust the fifth feature data. In some other embodiments, the model for adjusting the fifth feature data trained based on the 3D occupancy true value can also be deployed in the feature fusion module, the time sequence fusion module, the encoding module, etc. In this way, the vehicle-mounted device can determine more accurate fifth feature data (or 3D feature), so as to improve the accuracy of the vehicle-mounted device in predicting the overhanging obstacle.

[0162] The electronic device involved in the above embodiments is introduced below.

[0163] For example, Figure 5 According to some embodiments of the present application, a structural schematic diagram of an electronic device 100 is shown.

[0164] The electronic device 100 can be the vehicle-mounted device in the above embodiments, and the electronic device 100 is used to implement the obstacle prediction method provided in the above embodiments.

[0165] As Figure 5 shown, the electronic device 100 includes one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output device 105, and a system control logic unit 106 for coupling the processor 101, the system memory 102, the non-volatile memory 103, the communication interface 104, and the input / output (I / O) device 105. Among them: The processor 101 can include one or more processing units, e.g., can include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), an artificial intelligence (AI) processor, or a programmable logic device (FPGA), a neural-network processing unit (NPU), etc. The processing module or processing circuit can include one or more single-core or multi-core processors. In some embodiments, the CPU can be used to optimize the neural network model to be run, e.g., in some embodiments of the present application, the neural network model can optimize the spatial data of the first obstacle model, and the NPU can be used to run the neural network model to be run.

[0166] The system memory 102 is a volatile memory, e.g., a random-access memory (RAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory is used to temporarily store data and / or instructions, e.g., in some embodiments, the system memory 102 can be used to store the data provided by the foregoing embodiments, e.g., sensor data, image data or video data, etc., and can also be used to store the instructions of the obstacle prediction method provided by the foregoing embodiments, etc.

[0167] The non-volatile memory 103 can include one or more tangible, non-transitory computer-readable media for storage of data and / or instructions. In some embodiments, the non-volatile memory 103 can include any suitable non-volatile memory and / or any suitable non-volatile storage device, such as a flash memory, and / or a hard disk drive (HDD), compact disc (CD), digital versatile disc (DVD), a solid-state drive (SSD), and so forth. In some embodiments, the non-volatile memory 103 can also be a removable storage medium, such as a secure digital (SD) memory, and so forth. In other embodiments, the non-volatile memory 103 can be used to store instructions, and so forth, for the obstacle prediction method provided by the various embodiments.

[0168] In particular, the system memory 102 and the non-volatile memory 103 can include, respectively, a temporary copy and a permanent copy of the instructions 107. The instructions 107 can include instructions that, when executed by at least one of the processors 101, cause the electronic device 100 to implement the obstacle prediction method provided by the various embodiments.

[0169] The communication interface 104 can include a transceiver to provide a wired or wireless communication interface to the electronic device 100 to enable communication to and from any other suitable device via one or more networks. In some embodiments, the communication interface 104 can be integrated with other components of the electronic device 100, such as the communication interface 104 can be integrated with the processors 101. In some embodiments, the electronic device 100 can communicate with other devices via the communication interface 104, such as the electronic device 100 can obtain corresponding data from other devices via the communication interface 104, and so forth.

[0170] The input / output (I / O) device 105 can include input devices, such as a keyboard, a mouse, and so forth, and output devices, such as a display, and so forth, through which a user can interact with the electronic device 100.

[0171] The system control logic 106 can include any suitable interface controllers to provide any suitable interfaces to other modules of the electronic device 100. For example, in some embodiments, the system control logic 106 can include one or more memory controllers to provide an interface to connect to the system memory 102 and the non-volatile memory 103.

[0172] In some embodiments, at least one of the processors 101 can be packaged together with logic for one or more controllers of the system control logic 106 to form a system in package (SiP). In other embodiments, at least one of the processors 101 can also be integrated on the same die with logic for one or more controllers of the system control logic 106 to form a system-on-chip (SoC).

[0173] It can be appreciated that Figure 5 The structure of the electronic device 100 shown is only an example, and in other embodiments, the electronic device 100 can include more or fewer components than shown, or combine some components, or split some components, or different arrangement of components. The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0174] It can be appreciated that the electronic device 100 can be any device configured on a vehicle, including but not limited to a mobile phone, a car machine, a terminal in self driving, a wireless terminal in transportation safety, a terminal in smart city, and the like.

[0175] Embodiments of the present application further provide a program product, which, when executed on an electronic device, can enable the electronic device to implement the method provided by the foregoing embodiments.

[0176] Embodiments of the present application further provide a readable storage medium, which stores one or more programs, and the one or more programs, when executed by an electronic device, enable the electronic device to implement the method provided by the foregoing embodiments.

[0177] Embodiments of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. Embodiments of the application can be implemented as computer programs or program code executing on programmable systems comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0178] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information, which can be applied to one or more output devices, can be formatted in known manners. For purposes of this application, a processing system includes any system that has a processor, such as a digital signal processor, microcontroller, a programmable logic device, a microprocessor, or any other processing device.

[0179] The program code can be implemented in a high level of programming language or a object-oriented programming language to communicate with a processing system. In the event that it is desired, the program code can be implemented in an assembly or machine language. In fact, the mechanisms described herein are not limited in scope to any particular programming language. In any case, the language can be a compiled or interpreted language.

[0180] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried by or stored on a transitory or non-transitory machine-readable (e.g., computer-readable) medium, which can be read and executed by one or more processors. For example, the instructions can be distributed over the network or by other computer readable media. Thus, a machine-readable medium can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including without limitation floppy disks, optical disks, optical disks, compact discs-read only memory (CD-ROMs), magnetic disks, read only memory (ROM), random access memory (RAM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), magnetic or optical cards, flash memory, or a tangible, machine-readable storage used in the transmission of information over the Internet via electrical, optical, acoustical or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). Accordingly, the machine-readable media includes any type of media mechanisms that are suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0181] In the drawings, some of the structural or methodological features can be shown in particular arrangements and / or orders. However, it should be understood that such particular arrangements and / or orders can not be required. Instead, these features can be arranged in a different manner and / or order than shown in the illustrative figures, in some embodiments. Additionally, inclusion of structural or methodological features in a particular figure is not meant to imply that such features are required in all embodiments, and these features can not be included or can be combined with other features in some embodiments.

[0182] It should be noted that each unit / module mentioned in each device embodiment of the present application is a logical unit / module, and in the physical world, one logical unit / module can be a physical unit / module, or a part of a physical unit / module, or be realized in a combination of multiple physical unit / modules, and the physical realization of these logical units / modules is not the most important thing, and the combination of the functions implemented by these logical units / modules is the key to solving the technical problems proposed in the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce the units / modules that are not closely related to solving the technical problems proposed in the present application, which does not mean that the above-mentioned device embodiments do not have other units / modules.

[0183] It should be noted that in the examples and descriptions of the present application, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including one" does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0184] Although the present application has been illustrated and described with reference to certain preferred embodiments thereof, it should be understood by those skilled in the art that various changes in form and details can be made therein without departing from the scope of the present application.

Claims

1. An obstacle prediction method, applied to electronic equipment, characterized in that: The method comprises: Acquire a plurality of images, wherein the plurality of images include images acquired by a sensor of the vehicle from different perspectives; Determining first feature data of a feature point of each of the plurality of images, wherein the first feature data includes first coordinate data of the feature point in a vehicle coordinate system, and the first coordinate data includes a height value of the feature point; Based on the first feature data of each image in the plurality of images, generating a plane feature fusion image corresponding to the plurality of images, wherein the feature points included in the plane feature fusion image are: feature points mapped onto a horizontal plane of the vehicle coordinate system by mapping the feature points of each image in the plurality of images; generating feature data of the plane feature fusion image based on the first feature data of each image in the plurality of images, wherein the feature data of the plane feature fusion image includes second feature data, and the second feature data includes: region coordinate data of each subregion of a plurality of subregions obtained by dividing the plane feature fusion image, and a height range in the vehicle coordinate system corresponding to the feature point in each subregion; Obstacle prediction is performed based on the feature data of the plane feature fusion image.

2. The method according to claim 1, characterized in that The first feature data also includes visual information of each feature point, the feature data of the plane feature fusion image also includes third feature data, and, The method further comprises: generating third feature data of the plane feature fusion image based on the first feature data of each image in the plurality of images, wherein the third feature data includes: second coordinate data and first encoding data of each mapped feature point in the plane feature fusion image; The second coordinate data is two-dimensional coordinate data obtained by mapping the first coordinate data of the feature point to the horizontal plane of the vehicle coordinate system, and the first coded data is combined data of visual information and height data of the feature point mapped to the same second coordinate data.

3. The method according to claim 2, characterized in that Also includes: generating fourth feature data of the plane feature fusion image based on the first feature data of each image in the plurality of images, wherein the fourth feature data includes: second coordinate data and second encoding data of each mapped feature point in the plane feature fusion image; The second encoded data is fused data of the visual information of the feature points mapped to the same second coordinate data.

4. The method according to claim 3, characterized in that The obstacle prediction based on the feature data of the plane feature fusion image includes: fusing the second feature data, the third feature data, and the fourth feature data into fifth feature data; Obstacle prediction is performed based on the fifth feature data.

5. The method according to claim 1, wherein The electronic device is further configured with a first model, which is obtained by training according to the following method based on the third coordinate data of the first sample in the vehicle coordinate system and the fourth coordinate data of the second sample mapped into the vehicle coordinate system; adjusting the fourth coordinate data based on the first parameter to obtain first adjusted data; Obtaining a similarity between the first adjustment data and the third coordinate data; Based on the similarity between the first adjustment data and the third coordinate data, modifying the first parameter to a second parameter so that the similarity between the fourth coordinate data adjusted by the second parameter and the third coordinate data is greater than or equal to a similarity threshold; using the second parameter as a model parameter of the first model; The first sample is point cloud data collected by the test vehicle in a first environment, and the second sample is feature data of feature points in a plurality of image samples collected by the test vehicle in the first environment; The electronic device determines the height range corresponding to the feature point in each sub-area in the vehicle coordinate system according to the following method: inputting the first coordinate data of the feature point into the first model to adjust the height data of the first coordinate data of the feature point; A height range corresponding to the feature point in each sub-area in the vehicle coordinate system is determined based on the adjusted first coordinate data of the feature point.

6. The method according to claim 1, characterized in that The feature points in a first sub-region mapped to the multiple sub-regions form feature segments at different heights relative to a horizontal plane of the vehicle coordinate system in the vehicle coordinate system, wherein any feature point in one feature segment has more than a preset number of feature points within a region with a height M; The height range of the first characteristic segment in the first sub-area in the vehicle coordinate system is used as the height range of the characteristic point in the first sub-area corresponding to the vehicle coordinate system, and the first characteristic segment is the lowest characteristic segment in the first sub-area relative to the horizontal plane of the vehicle coordinate system.

7. The method according to claim 6, characterized in that The first feature data also includes category information, and the category information includes ground category and suspended obstacle category; Corresponding to the fact that the height of the lowest feature point in the first feature segment is higher than the height of the feature point of the ground category, the category of the feature point in the first feature segment is set to the suspended obstacle category.

8. The method according to claim 4, characterized in that The performing obstacle prediction based on the fifth characteristic data includes: Obtaining the fifth feature data corresponding to a feature point set within a preset distance from the vehicle; Improving the resolution of the fifth feature data corresponding to the feature point set to obtain sixth feature data; Obstacle prediction is performed on the environment surrounding the vehicle based on the sixth feature data.

9. An electronic device, characterized in that: The device comprises: a memory for storing instructions; At least one processor is configured to execute the instructions so that the electronic device implements the method according to any one of claims 1 to 8.

10. A computer program product, characterized in that When the computer program product is run on a device, the device is caused to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Software perception precision qualification test method, device, equipment and medium

    CN118132417A

  • Obstacle detection method and device, electronic equipment and storage medium

    CN119068462A

  • Parking area determination method and device, equipment and storage medium

    CN119540906A

  • Obstacle recognition and positioning method and device, electronic equipment and readable storage medium

    CN120032340A