Ground element information acquisition method and device, electronic equipment and vehicle

By combining the point cloud data and image data of the current frame with the feature fusion of the historical frame, and utilizing the self-attention and cross-attention mechanisms, the problem of obtaining accurate ground feature information is solved, and the real-time update and accuracy of the high-precision map are achieved.

CN115761680BActive Publication Date: 2025-10-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211488541.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2025-10-10
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

How to obtain more accurate ground feature information to support the construction and updating of high-precision maps, especially in the autonomous driving environment, where existing technologies find it difficult to effectively handle the impact of environmental factors on data.

Method used

By combining the point cloud data, image data of the current frame and the target features of the historical frames, adopting temporal and spatial fusion processing, using a pre-trained model to extract features, and performing feature fusion through self-attention and cross-attention mechanisms, we can finally obtain ground feature information.

Benefits of technology

It effectively avoids the impact of environmental factors on data, improves the accuracy and robustness of ground element information, and ensures the real-time update and accuracy of high-precision maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761680B_ABST
    Figure CN115761680B_ABST
Patent Text Reader

Abstract

The present disclosure provides a ground element information acquisition method and device, electronic equipment, vehicle and storage medium, relates to the field of artificial intelligence, in particular to the fields of automatic driving, intelligent transportation and the like. The specific implementation scheme is: obtaining image features of a current frame based on image data of the current frame, and obtaining point cloud features of the current frame based on point cloud data of the current frame; obtaining initialization features of the current frame based on the point cloud features of the current frame and target features of a historical frame; obtaining target features of the current frame based on the initialization features of the current frame and the image features of the current frame; and obtaining ground element information based on the target features of the current frame. The above method can improve the accuracy of the ground element information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and in particular to the fields of autonomous driving, intelligent transportation, etc. Background Art

[0002] With the advancement of computer technology, artificial intelligence fields such as intelligent transportation have also experienced rapid growth. In particular, technologies such as autonomous vehicles are gaining widespread application. Autonomous driving requires high-precision maps, which in turn require underlying ground feature information. However, obtaining more accurate ground feature information has become a challenge. Summary of the Invention

[0003] The present disclosure provides a method, device, electronic device, vehicle and storage medium for acquiring ground element information.

[0004] According to a first aspect of the present disclosure, a method for acquiring ground feature information is provided, comprising:

[0005] Obtaining image features of the current frame based on the image data of the current frame, and obtaining point cloud features of the current frame based on the point cloud data of the current frame;

[0006] Obtaining initialization features of the current frame based on the point cloud features of the current frame and the target features of the historical frames;

[0007] Obtaining target features of the current frame based on the initialization features of the current frame and the image features of the current frame;

[0008] Based on the target features of the current frame, ground element information is obtained.

[0009] According to a second aspect of the present disclosure, a device for acquiring ground element information is provided, comprising:

[0010] A data processing module, configured to obtain image features of the current frame based on the image data of the current frame, and obtain point cloud features of the current frame based on the point cloud data of the current frame;

[0011] An initialization processing module, configured to obtain initialization features of the current frame based on the point cloud features of the current frame and the target features of the historical frames;

[0012] A target feature processing module, configured to obtain a target feature of the current frame based on the initialization feature of the current frame and the image feature of the current frame;

[0013] The ground element acquisition module is used to obtain ground element information based on the target features of the current frame.

[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the ground element information acquisition method of the first aspect mentioned above.

[0018] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the aforementioned method.

[0019] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the aforementioned method when executed by a processor.

[0020] According to a sixth aspect of the present disclosure, a vehicle is provided, comprising the electronic device according to the embodiment of the first aspect described above.

[0021] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description.

[0022] The solution provided in this embodiment can obtain the point cloud features of the current frame based on the point cloud data of the current frame, and based on the point cloud features of the current frame and the target features of the historical frame, the initialization features of the current frame can be obtained. Then, based on the initialization features of the current frame and the image features of the current frame, the target features of the current frame can be obtained, and finally, the ground feature information can be obtained based on the target features of the current frame. In this way, by combining the point cloud data of the current frame, the image data of the current frame, and the target features of the historical frame for temporal and spatial fusion processing, the impact of environmental factors on the data in real situations can be effectively avoided, and the accuracy of the ground feature information can be improved while ensuring the real-time acquisition of ground feature information. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0024] Figure 1 is a flowchart of a method for acquiring ground feature information according to an embodiment of the present disclosure;

[0025] Figure 2 is a schematic diagram of a processing scenario for adjusting target features of a historical frame according to an embodiment of the present disclosure;

[0026] Figure 3 is a schematic diagram of a processing scenario for obtaining initialization features of a current frame based on adjusted target features of historical frames according to another embodiment of the present disclosure;

[0027] Figure 4 is a schematic flowchart of a method for acquiring ground feature information according to an embodiment of the present disclosure;

[0028] Figure 5 is another schematic flow chart of a method for acquiring ground feature information according to an embodiment of the present disclosure;

[0029] Figure 6 is a schematic diagram of the structure of a ground feature information acquisition device according to another embodiment of the present disclosure;

[0030] Figure 7 It is a block diagram of an electronic device used to implement the ground feature information acquisition method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0032] The first embodiment of the present disclosure provides a method for obtaining ground element information, such as Figure 1 Shown, including:

[0033] S101: obtaining an image feature of the current frame based on image data of the current frame, and obtaining a point cloud feature of the current frame based on point cloud data of the current frame;

[0034] S102: Obtaining initialization features of the current frame based on the point cloud features of the current frame and the target features of the historical frames;

[0035] S103: Obtaining target features of the current frame based on the initialization features of the current frame and the image features of the current frame;

[0036] S104: Obtaining ground element information based on the target features of the current frame.

[0037] The ground feature information acquisition method provided by the embodiment of the first aspect above can be applied to electronic devices. In one example, the electronic device can be a terminal device, such as a laptop, a smart phone, a tablet computer, a desktop computer, and the like; in another example, the electronic device can be a vehicle-mounted terminal set in a vehicle, and in another example, the electronic device can be a server cluster or a computing node of a distributed system that can transmit data with the vehicle. It should be understood that the above is only an exemplary description of the execution subject that can execute the ground feature information acquisition method provided by this embodiment. In actual processing, it may not be limited to the devices mentioned in the above examples. As long as the device or equipment can execute the ground feature information acquisition method provided by this embodiment, it is within the protection scope of this embodiment.

[0038] The ground feature information may be traffic features on the ground, such as traffic features within a drivable area on the ground. For example, the ground feature information may include at least one of the following: lane markings, stop lines, slow-down yield lines, ground arrows, text, guide lines, etc. The aforementioned ground feature information may be used when constructing or updating a high-precision map. The high-precision map can at least be used to execute tasks such as path regulation and trajectory prediction. This embodiment does not limit the processing or tasks that can be performed by the high-precision map.

[0039] It can be seen that by adopting the above scheme, the point cloud features of the current frame can be obtained based on the point cloud data of the current frame, and the initialization features of the current frame can be obtained based on the point cloud features of the current frame and the target features of the historical frame. Then, based on the initialization features of the current frame and the image features of the current frame, the target features of the current frame can be obtained, and finally the ground feature information of this time can be obtained based on the target features of the current frame. In this way, by combining the point cloud data of the current frame, the image data of the current frame, and the target features of the historical frame for temporal and spatial fusion processing, the influence of environmental factors on the data in real situations can be effectively avoided, and the accuracy and robustness of the ground feature information can be improved while ensuring the real-time acquisition of ground feature information.

[0040] In some possible implementations, the method further includes: acquiring image data of the current frame and acquiring point cloud data of the current frame in real time.

[0041] Specifically, the real-time acquisition of the image data of the current frame may be real-time acquisition of image information of the current frame acquired by an image acquisition device. There may be one or more image acquisition devices, and any one of the image acquisition devices may be any of the following: a camera, a surround-view camera, an infrared camera, a webcam, a time-of-flight (ToF) camera, or other similar devices with image acquisition capabilities.

[0042] The real-time acquisition of point cloud data of the current frame may be real-time acquisition of point cloud data of the current frame acquired by a point cloud acquisition device. The number of the point cloud acquisition devices may be one or more; any one of the point cloud acquisition devices may include a laser scanning device, such as a LiDAR (Light Detection and Ranging), or other similar devices with scanning capabilities.

[0043] It should be noted that the aforementioned image acquisition device and point cloud acquisition device may be mounted (or installed) on the same acquisition device. This acquisition device may include, but is not limited to, any of the following: a vehicle, aircraft, ship, unmanned aerial vehicle, robot, or other mobile device. For example, this acquisition device may be a map acquisition vehicle.

[0044] The acquisition device can control the aforementioned point cloud acquisition device and the aforementioned image acquisition device to respectively acquire point cloud data and image data at the same time, and can also control the acquisition range of image data to match the acquisition range of point cloud data. Here, the acquisition range of image data matches the acquisition range of point cloud data, including: the acquisition range of image data and point cloud data at least partially overlaps; that is, the acquisition range of image data and point cloud data has an overlapping range (hereinafter referred to as the overlapping range). Correspondingly, the field of view of the image acquisition device and the point cloud acquisition device at least partially overlaps, that is, the field of view of the image acquisition device and the point cloud acquisition device has an overlapping range. In a preferred example, the acquisition range of point cloud data and the acquisition range of image data can completely overlap. Of course, the acquisition range of point cloud data can also be larger than the acquisition range of image data, and include part or all of the acquisition range of image data. Or it can also be that the acquisition range of image data is larger than the acquisition range of point cloud data, and includes part or all of the acquisition range of point cloud data. This embodiment does not enumerate all possible situations.

[0045] Furthermore, the method may further include: acquiring, in real time, the pose of the acquisition device for the current frame captured by a pose-related sensor provided on the acquisition device. Furthermore, the pose of the acquisition device for the current frame may be stored in a dataset. The pose of the acquisition device for the current frame refers to the pose of the acquisition device at the time corresponding to the current frame.

[0046] That is, in addition to being provided with an image acquisition device and a point cloud acquisition device, the aforementioned acquisition device is also provided with a posture-related sensor. Among them, the posture-related sensor may include but is not limited to at least one of the following: Global Navigation Satellite System (GNSS), Inertial Measurement Unit (IMU), etc. The acquisition device may be configured to acquire the acquisition device posture of the current frame through the posture-related sensor while controlling the image acquisition device to acquire the image information of the current frame and controlling the point cloud acquisition device to acquire the point cloud data of the current frame; accordingly, the electronic device that executes the map element segmentation method provided in this embodiment can obtain the acquisition device posture of the current frame acquired by the posture-related sensor of the acquisition device in real time.

[0047] Among them, the position and posture of the acquisition device may include: the geographical location of the acquisition device, the relative angle of the acquisition device, etc. Among them, the geographical location of the acquisition device can be expressed by the longitude and latitude in the spherical coordinate system (or called the geodetic coordinate system); it should be understood that in actual application, the geographical location can also be expressed by other coordinate systems, but this embodiment does not list them all. The relative angle of the acquisition device may refer to the relative angle between the orientation of the acquisition device and the reference direction; the reference direction may be configured according to actual conditions, such as any one of due south, due north, due east, and due west, which is not limited here.

[0048] In some possible implementations, obtaining point cloud features of the current frame based on the point cloud data of the current frame may include: inputting the point cloud data of the current frame into a first model to obtain point cloud features of the current frame output by the first model. Obtaining image features of the current frame based on the image data of the current frame may include: inputting the image data of the current frame into a second model to obtain image features of the current frame output by the second model.

[0049] Among them, the first model and the second model are both pre-trained models. The first model and the second model can be obtained and saved by the aforementioned electronic device after being trained in other electronic devices; or, the first model and the second model can be trained and saved in the aforementioned electronic device.

[0050] The aforementioned first model may be a model for extracting point cloud features; specifically, the first model may include a backbone network for extracting point cloud features. Exemplarily, the backbone for extracting point cloud features may be a SCEOND (Sparsely Embedded Convolutional Detection) network. It should be understood that this is merely an example, and other networks or models may be used in actual processing to implement the function of extracting point cloud features, but this embodiment does not provide an exhaustive list.

[0051] The aforementioned processing of inputting the point cloud data of the current frame into the first model to obtain the point cloud features of the current frame output by the first model may be: inputting the point cloud data of the current frame into the first model, pressing the point cloud data of the current frame into the target space in the first model, and obtaining the point cloud features of the current frame; the size of the point cloud features of the current frame is C l ×H bev ×W bev The target space may be specifically a BEV (Bird's Eye View) space. l It can be a positive integer, indicating the number of channels of the point cloud feature of the current frame; H bev Can be a positive number, indicating the height of the point cloud feature of the current frame; W bev It can be a positive number, indicating the width of the point cloud feature of the current frame. In this embodiment, the number of channels may refer to the number of feature maps, which will not be repeated below.

[0052] The second model may be a model for extracting image features; the second model may include a backbone network for extracting image features. Exemplarily, the backbone for extracting image features may be at least one of the following: ResNet (residual network), EfficientNet (efficient network), etc.

[0053] The aforementioned process of inputting the image data of the current frame into the second model to obtain the image features of the current frame output by the second model may be: inputting the image data of the current frame into the second model, performing feature extraction on the image data of the current frame in the second model to obtain the image features of the current frame; the image features of the current frame have a size of n×C cam ×H×W. Where n is a positive integer, indicating the number of image acquisition devices; C cam is a positive integer representing the number of channels of the image feature of the current frame; H is a positive number representing the height of the image feature of the current frame; and W is a positive number representing the width of the image feature of the current frame. In one possible example, n can be 6, meaning that six image acquisition devices can be configured on the acquisition device.

[0054] In some possible implementations, obtaining the initialized target features based on the point cloud features of the current frame and the target features of the historical frames includes: adjusting the target features of the historical frames to obtain the adjusted target features of the historical frames that are aligned with the point cloud features of the current frame; and obtaining the initialized features of the current frame based on the adjusted target features of the historical frames and the point cloud features of the current frame.

[0055] The target feature of the historical frame may specifically be the target feature of the previous frame, the target feature may specifically be the BEV feature, and the initialization feature of the current frame may specifically be the initialization BEV feature of the current frame.

[0056] Adjusting the target features of the historical frame to obtain the adjusted target features of the historical frame aligned with the point cloud feature positions of the current frame may refer to: adjusting the target features of the historical frame based on the acquisition device posture of the current frame and the acquisition device posture of the historical frame to obtain the adjusted target features of the historical frame aligned with the point cloud feature positions of the current frame.

[0057] Through the above method, the point cloud features of the current frame and the target features of the historical frame can be aligned first, and then the point cloud features of the current frame after the alignment and the target features of the historical frame after the adjustment can be temporally fused. In this way, the accuracy of the initialization features of the current frame obtained after temporal fusion can be guaranteed, so that the final ground feature information is more accurate.

[0058] In some possible implementations, adjusting the target features of the historical frame to obtain the adjusted target features of the historical frame aligned with the point cloud feature positions of the current frame may specifically include: acquiring the acquisition device posture of the current frame and the acquisition device posture of the historical frame; determining posture change information based on the acquisition device posture of the current frame and the acquisition device posture of the historical frame; determining adjustment information based on the posture change information; and adjusting the target features of the historical frame based on the adjustment information to obtain the adjusted target features of the historical frame aligned with the point cloud feature positions of the current frame.

[0059] The method for obtaining the acquisition device posture of the current frame and the acquisition device posture of the historical frame can be: obtaining the acquisition device posture of the current frame and the acquisition device posture of the historical frame from the data set. The acquisition device posture of the historical frame refers to the acquisition device posture at the time corresponding to the historical frame; the acquisition device posture of the current frame refers to the acquisition device posture at the time corresponding to the current frame. The acquisition device posture and the processing method for saving the acquisition device posture in the data set have been explained in the previous embodiment and will not be repeated here.

[0060] Determining the posture change information based on the acquisition device posture of the current frame and the acquisition device posture of the historical frame may refer to: obtaining the position change information and angle change information of the acquisition device based on the acquisition device posture of the current frame and the acquisition device posture of the historical frame; and using the position change information and angle change information of the acquisition device as the posture change information.

[0061] Specifically, the acquisition device's location change information may refer to the change between the acquisition device's geographic location at the time corresponding to the current frame and the acquisition device's geographic location at the time corresponding to the historical frame. The acquisition device's location change information may include at least one of a longitude difference and a latitude difference. The longitude difference may be 0, a positive number, or a negative number; the latitude difference may be 0, a positive number, or a negative number.

[0062] The angle change information of the acquisition device may refer to the change between the relative angle of the acquisition device at the time corresponding to the current frame and the relative angle of the acquisition device at the time corresponding to the historical frame. The angle change information of the acquisition device may include an angle difference; the angle difference may be any one of 0, a positive number, or a negative number.

[0063] Determining the adjustment information based on the posture change information may include: determining a position adjustment value based on the position change information of the acquisition device, and determining an angle adjustment value based on the angle change information of the acquisition device; and using the position adjustment value and the angle adjustment value as the adjustment information.

[0064] The adjusting of the target features of the historical frame based on the adjustment information to obtain the adjusted target features of the historical frame aligned with the point cloud feature positions of the current frame may include: translating the target features of the historical frame based on the position adjustment value in the adjustment information, and rotating the target features of the historical frame based on the angle adjustment value in the adjustment information to obtain the adjusted target features of the historical frame aligned with the point cloud feature positions of the current frame.

[0065] Among them, determining the position adjustment value based on the position change information of the acquisition device may refer to: taking the negative value of the position change information of the acquisition device as the position adjustment value. For example, assuming that the geographical location of the acquisition device at the moment corresponding to the current frame is position 1, and the geographical location of the acquisition device at the moment corresponding to the historical frame is position 2, position 1 increases by a value A in longitude relative to position 2, and there is no change in latitude. If it is necessary to align the target features of the historical frame with the point cloud features of the current frame, it is necessary to reduce the target features of the historical frame by a value A in longitude, so the position adjustment value may be a negative value A. Further, translating the target features of the historical frame based on the position adjustment value in the adjustment information may be: translating the target features of the historical frame by a negative value A in longitude.

[0066] Determining the angle adjustment value based on the angle change information of the acquisition device may refer to: taking the negative value of the angle change information of the acquisition device as the angle adjustment value. For example, the relative angle of the acquisition device when performing the acquisition of the current frame is Angle 1, and the relative angle of the acquisition device when performing the acquisition of the historical frame is Angle 2, and Angle 1 increases the angle difference B relative to Angle 2. If it is necessary to align the target feature of the historical frame with the point cloud feature of the current frame, it is necessary to rotate the target feature of the historical frame in the opposite direction by the angle difference B, so the angle adjustment value may be a negative angle difference B. Further, rotating the target feature of the historical frame based on the angle adjustment value in the adjustment information may be: rotating the target feature of the historical frame by a negative angle difference B based on the current angle.

[0067] Combine Figure 2 For example, Figure 2 Part 200 in the figure illustrates the target feature 201 of the historical frame and the point cloud feature 202 of the current frame. After the aforementioned adjustment, the point cloud feature 212 of the current frame and the target feature 211 of the adjusted historical frame can be obtained as shown in part 210 in the figure. The target feature 211 of the adjusted historical frame is aligned with the point cloud feature of the current frame.

[0068] Since the acquisition device may be moving in real time, the position and posture of the acquisition device may be different at the time corresponding to different frames, that is, there are differences in geographical location and / or relative angle between the image data of the current frame and the image data of the historical frame. By adopting the above-mentioned scheme provided by this embodiment, it is possible to determine the posture change information of the acquisition device based on the posture of the acquisition device of the current frame and the posture of the acquisition device of the historical frame, and then adjust the target features of the historical frame based on the posture change information. This can ensure that the target features of the adjusted historical frame and the features at the same position in the point cloud features of the current frame can be correctly matched, thereby ensuring the accuracy of subsequent processing.

[0069] In some possible implementations, obtaining the initialization features of the current frame based on the adjusted target features of the historical frame and the point cloud features of the current frame includes:

[0070] Adjusting the point cloud features of the current frame to obtain adjusted point cloud features of the current frame; wherein the adjusted point cloud features of the current frame and the adjusted target features of the historical frame have the same number of channels;

[0071] The adjusted point cloud features of the current frame and the adjusted target features of the historical frames are fused to obtain the initialization features of the current frame.

[0072] Among them, the number of channels of the target features of the adjusted historical frame is the same as the number of channels of the target features of the aforementioned historical frame before adjustment, for example, both can be expressed as "C". Therefore, it can also be expressed here that the point cloud features of the adjusted current frame have the same number of channels as the target features of the historical frame.

[0073] The adjusting of the point cloud features of the current frame to obtain the adjusted point cloud features of the current frame may include: inputting the point cloud features of the current frame into a convolution layer, obtaining the output of the convolution layer, and using the output of the convolution layer as the adjusted point cloud features of the current frame. For example, inputting the point cloud features of the current frame into a convolution layer, and through the processing of the convolution layer, the number of channels of the point cloud features of the current frame is increased from C to 1. l The target feature of the adjusted historical frame becomes C, where C is a positive integer. The number of channels of the point cloud feature of the current frame has been exemplified above and will not be repeated here.

[0074] The number of convolution kernels of the convolution layer can be set according to actual needs. For example, if the actual requirement is that the number of output channels is C, the number of convolution kernels of the convolution layer can be set to C.

[0075] By using the above process, the point cloud features of the current frame can be adjusted to have the same number of channels as the target features of the adjusted historical frame. Then, the adjusted point cloud features of the current frame and the adjusted target features of the historical frame are fused to obtain the initialization features of the current frame. In this way, features with the same number of channels can be fused in the time series fusion process, which can ensure the accuracy of time series fusion and ultimately ensure that the final segmentation of ground feature information is more accurate.

[0076] In some possible implementations, the fusing of the adjusted point cloud features of the current frame and the target features of the adjusted historical frame to obtain the initialization features of the current frame may include: determining a first mask based on the target features of the adjusted historical frame; adding the target features of the adjusted historical frame to the point cloud features of the adjusted current frame to obtain the initial fusion features of the current frame; and normalizing the initial fusion features of the current frame based on the first mask to obtain the initialization features of the current frame.

[0077] Determining the first mask based on the adjusted target feature of the historical frame may refer to using the area covered by the adjusted target feature of the historical frame as the first mask.

[0078] Adding the target features of the adjusted historical frame to the point cloud features of the adjusted current frame to obtain the initial fusion features of the current frame may refer to adding the feature vectors at the same positions of the target features of the adjusted historical frame and the point cloud features of the adjusted current frame to obtain the initial fusion features of the current frame.

[0079] In one example, the target features of the adjusted historical frame may include the feature vectors corresponding to each pixel; similarly, the point cloud features of the adjusted current frame may also include the feature vectors corresponding to each pixel. The aforementioned same position may refer to the same pixel position. In another example, the target features of the adjusted historical frame may include the feature vectors corresponding to each grid; similarly, the point cloud features of the adjusted current frame may also include the feature vectors corresponding to each grid. The aforementioned same position may refer to the position of the same grid.

[0080] Normalizing the initial fused features of the current frame based on the first mask to obtain the initial features of the current frame may include: dividing the initial fused features of the current frame into overlapping regions and non-overlapping regions based on the first mask; dividing the feature vectors in the overlapping regions by two to obtain adjusted features of the overlapping regions; and merging the adjusted features of the overlapping regions with the features of the non-overlapping regions to obtain the initial features of the current frame. It should be understood that the aforementioned feature vectors may also be referred to as feature values ​​in some possible examples, and this embodiment does not exhaustively enumerate all possible names.

[0081] See also Figure 3 For example, Figure 3 Part 300 in the figure shows the adjusted point cloud features 302 of the current frame and the adjusted target features 301 of the historical frame in the above embodiment; Figure 3Part 310 in FIG. 3 illustrates a first mask 311 determined based on the target feature 301 of the adjusted historical frame; Figure 3 Part 320 in the figure illustrates dividing the initial fusion features of the current frame into an overlapping region 321 and a non-overlapping region 322 based on the first mask 311. Since the feature vector at the position of each pixel (or each grid) in the overlapping region 321 is obtained by adding the feature vector of the adjusted point cloud feature of the current frame and the target feature vector of the adjusted historical frame, the feature vector in the overlapping region 321 can be divided by two to obtain the feature of the adjusted overlapping region; and the non-overlapping region only has the feature vector of the adjusted point cloud feature of the current frame, so the feature vector of the non-overlapping region does not need to be adjusted; by merging the features of the adjusted overlapping region with the features of the non-overlapping region, the initialization features of the current frame can be obtained.

[0082] Because when obtaining the target features of the adjusted historical frame that are aligned with the point cloud feature positions of the current frame, the target features of the historical frame are rotated and translated, and there will be some positions in the area of ​​the point cloud features of the current frame that do not have the corresponding target features of the historical frame, that is, the feature vectors of the target features of the historical frame corresponding to these positions are 0. If the adjusted point cloud features of the current frame are directly used and added to the adjusted target features of the historical frame, the final result may be incorrect. This embodiment normalizes the initial fusion features of the current frame by adopting the above method, thereby ensuring the accuracy of the initialization features of the current frame. In addition, the processing method of temporal fusion provided by the aforementioned embodiment only needs to add the adjusted point cloud features of the current frame and the adjusted target features of the historical frame. While the operation is simple and the increase in the amount of calculation is not obvious, it can also ensure the accuracy of the initialization features of the current frame, thereby ensuring that the subsequent spatial fusion can obtain accurate results, and ultimately ensuring that the segmentation of the ground feature information obtained is more accurate.

[0083] In some possible implementations, obtaining the target features of the current frame based on the initialization features of the current frame and the image features of the current frame includes: performing self-attention processing on the initialization features of the current frame to obtain the initialization features of the current frame including attention weights; performing cross-attention processing on the initialization features of the current frame including attention weights and the image features of the current frame to obtain the target features of the current frame.

[0084] The aforementioned processing uses an attention mechanism (AM), which is a mechanism that simulates human attention in a computer vision system, enabling the neural network to focus on a part of the important input information, that is, the attention mechanism can be regarded as a dynamic selection process that adaptively weights the features according to the relevance or importance of the input. The attention mechanism involved in the present disclosure includes a self-attention mechanism and a cross-attention mechanism. Among them, the self-attention mechanism, also known as the internal attention mechanism, reduces the dependence on external information and is good at capturing the internal correlation of features, that is, the self-attention mechanism aims to model the internal correlation of the same group of input features to enhance the features and obtain more informative feature representations. The cross-attention mechanism is generally used to construct the cross-correlation between two different groups of input features, obtain the relationship between the two groups of input features, and achieve the effect of fusing information into one of the groups of input features.

[0085] The aforementioned self-attention processing of the initialization features of the current frame to obtain the target features of the current frame containing self-attention weights can mean that the self-attention calculation is performed on the initialization features of the current frame to obtain the target features of the current frame containing self-attention weights; or the initialization features of the current frame can be input into a self-attention model to obtain the target features of the current frame containing self-attention weights output by the self-attention model. The self-attention model can be trained and saved in the aforementioned electronic device; or the self-attention model can be trained in other devices, and after obtaining the trained self-attention model, the electronic device acquires and saves the trained self-attention model.

[0086] In a preferred example, the aforementioned self-attention calculation or self-attention model can specifically adopt a global attention mechanism. Correspondingly, the aforementioned attention weight can refer to the weight value or correlation between the feature vector at each pixel position in the initialization feature of the current frame and the feature vector corresponding to each of the other pixels. For example, the initialization feature of the current frame contains 100 pixels. For pixel 1, the weight value or correlation between the feature vector at the pixel 1 position and the feature vector at the pixel 2 position can be obtained, the weight value or correlation between the feature vector at the pixel 1 position and the feature vector at the pixel 3 position can be obtained, and so on, and the weight value or correlation between the feature vector at the pixel 1 position and the feature vector at the pixel 100 position can be obtained. For pixels 2 to 100, the same results as pixel 1 can also be obtained, until the weight value or correlation between the feature vector at each pixel position in the initialization feature of the current frame and the feature vector corresponding to each of the other pixels is obtained.

[0087] The aforementioned cross-attention processing can be implemented by using a cross-attention model. The cross-attention processing of the initialization feature of the current frame containing the attention weight and the image feature of the current frame to obtain the target feature of the current frame can refer to inputting the initialization feature of the current frame containing the attention weight and the image feature of the current frame into a cross-attention model to obtain the target feature of the current frame output by the cross-attention model. The target feature of the current frame can include the attention weight value or correlation between the initialization feature of the current frame containing the attention weight and the image feature of the current frame.

[0088] In a preferred example, the aforementioned cross-attention processing can adopt a deformable attention mechanism. Specifically, the initialization feature of the current frame containing the attention weight and the image feature of the current frame are input into the cross-attention model to obtain the target feature of the current frame output by the cross-attention model, which can include: inputting the initialization feature of the current frame containing the attention weight and the image feature of the current frame into the cross-attention model, in the cross-attention model, the grid coordinates on the initialization feature of the current frame containing the attention weight are back-projected onto the image feature of the current frame according to the internal and external parameters of the camera, and the pixel position corresponding to the grid coordinate on the initialization feature of the current frame containing the attention weight is found on the image feature of the current frame, and the pixel position is used as a reference point to select one or more pixels around the reference point to obtain the attention weight or correlation between the feature vector corresponding to the grid coordinate on the initialization feature of the current frame containing the attention weight and the feature vector corresponding to one or more pixels around the reference point.

[0089] Among them, inputting the initialization features of the current frame containing the attention weight and the image features of the current frame into the cross-attention model may refer to: using the initialization features of the current frame containing the attention weight as the query (Q, Query), and using the image features of the current frame as both the key (Key, K) and the value (V, Value); inputting the above Q, K, and V into the cross-attention model.

[0090] The above-mentioned processing provided in this embodiment, namely, performing self-attention processing on the initialization features of the current frame to obtain the initialization features of the current frame including attention weights, and performing cross-attention processing on the initialization features of the current frame including attention weights and the image features of the current frame to obtain the target features of the current frame, can be referred to as spatial fusion processing. When performing spatial fusion processing on the current frame, the spatial fusion processing can be performed only once or multiple times.

[0091] Exemplarily, if the spatial fusion process is performed multiple times, the number of executions can be set according to actual conditions, for example, it can be 10 times, 20 times, or more or less, and there is no limitation here. Assuming that the spatial fusion process is performed multiple times, if the spatial fusion process is performed for the first time, it is the same as the process described in the aforementioned embodiment and will not be described in detail. If the spatial fusion process is performed for the mth time, the target features of the current frame obtained for the m-1th time can be used as the initialization features of the current frame for the mth processing. After performing the aforementioned processing, the target features of the current frame for the mth processing are obtained; wherein m is an integer greater than 1. Further, if the mth processing is the last spatial fusion processing, the target features of the current frame obtained by the mth processing are the target features of the current frame finally obtained this time. In this way, by repeatedly performing the spatial fusion process, the target features of the current frame finally obtained can be made more accurate.

[0092] It should also be pointed out that, when it is determined that the target features of the current frame are finally obtained, the target features of the current frame can also be saved for processing the next frame.

[0093] By adopting the above scheme, the image features can be spatially fused with the initialized target features, thereby further fusing the image features of the current frame on the basis of realizing the temporal fusion of the target features of the current frame and the historical frames. In this way, the problems of data missing and inaccurate data caused by environmental influences in real situations can be effectively avoided, and the accurate target features of the current frame, that is, the BEV features of the current frame, can be obtained, thereby improving the accuracy and robustness of ground feature information.

[0094] In some possible implementations, obtaining ground element information based on the target features of the current frame may include: inputting the target features of the current frame into a segmentation model to obtain N categories of ground element segmentation results output by the segmentation model; N is a positive integer; and obtaining ground element information under each of the N categories based on the ground element segmentation results of the N categories.

[0095] The segmentation model may include a convolutional layer and a deconvolutional layer. The segmentation model may be a pre-trained model, such as one trained and stored in an electronic device; or the segmentation model may be trained in another device and then acquired and stored by the electronic device.

[0096] Inputting the target features of the current frame into a segmentation model to obtain N types of ground feature segmentation results output by the segmentation model may include: inputting the target features of the current frame into a segmentation model, upsampling the target features of the current frame a times in the segmentation model, and obtaining N×aH bev ×aW bevThe ground feature segmentation result. Among them, the upsampling multiple can be preset, for example, it can be a times, a is a positive number. The aforementioned N categories can refer to N categories of ground feature information. That is, after the target feature of the current frame is upsampled a times by the segmentation model, the number of channels (or the number of feature maps) is N and the size is aH bev ×aW bev The ground feature segmentation result.

[0097] The obtaining of ground element information of each of the N categories based on the ground element segmentation results of the N categories may include:

[0098] When the probability value at the j-th position in the ground feature segmentation result of the i-th category is greater than the i-th probability threshold, it is determined that the ground feature information of the i-th category exists at the position corresponding to the j-th position in the target feature of the current frame; wherein i is a positive integer less than or equal to N, and j is a positive integer;

[0099] When the probability value at the jth position in the ground element segmentation result of the i-th category is not greater than the i-th probability threshold, it is determined that in the target feature of the current frame, there is no ground element information under the i-th category at the position corresponding to the j-th position.

[0100] The i-th ground feature segmentation result is any one of the N ground feature segmentation results. Since the processing for each ground feature segmentation result is the same, each is not described in detail. The j-th position can specifically refer to the position of any pixel in the i-th ground feature segmentation result; the j-th position can correspond to the grid of the target feature in the current frame.

[0101] The method may further include: taking the ground element segmentation result of the i-th category as a bev ×aW bev ith second mask. Accordingly, the above-mentioned ground feature segmentation result based on the N categories obtains ground feature information under each of the N categories, which may specifically include: when the probability value corresponding to the jth pixel position of the i-th second mask is greater than the i-th probability threshold, determining that in the target feature of the current frame, at the grid corresponding to the j-th pixel position, there is ground feature information under the i-th category; when the probability value corresponding to the j-th pixel position of the i-th second mask is not greater than the i-th probability threshold, determining that in the target feature of the current frame, at the grid corresponding to the j-th pixel position, there is no ground feature information under the i-th category.

[0102] The i-th probability threshold can be configured according to actual conditions. In one possible example, different categories can be configured with different probability thresholds; for example, for the ground feature segmentation results of the i-th category, the i-th probability threshold can be set to threshold A, and for the ground feature segmentation results of the i+1-th category, the i+1-th probability threshold can be set to threshold B, where threshold A and threshold B are different. In another possible example, different categories can be configured with the same probability threshold; for example, for the ground feature segmentation results of each category, the probability threshold can be set to threshold A. In another possible example, different categories can be configured with the same probability threshold, and / or different categories can be configured with different probability values. That is to say, among the N categories, some categories may have the same probability threshold, while some categories may have different probability thresholds. For example, for the ground feature segmentation result of the i-th category, the i-th probability threshold can be set to threshold A, for the ground feature segmentation result of the i+1-th category, the i+1-th probability threshold can be set to threshold B, and for the ground feature segmentation result of the i+2-th category, the i+2-th probability threshold can be set to threshold B, and the threshold A and threshold B are different. It should be understood that the above is only an exemplary description, and the configuration method of any probability threshold is not exhaustively listed here.

[0103] The number of ground feature information under each category may be one or more, which is not limited in this embodiment.

[0104] It should be understood that the aforementioned ground feature information is obtained by processing the target features of the current frame. Since the aforementioned processing of obtaining the target features of the current frame is performed in real time, that is, the target features of different frames can be obtained at different times, the corresponding ground feature information corresponding to the target features of different frames can be obtained accordingly. Ultimately, the ground feature information acquisition method provided by this embodiment can be executed within a certain area for a period of time to obtain all ground feature information within that area.

[0105] It should also be noted that in one example, the current frame may refer to frames other than the first frame, and this example is applicable to the solutions provided by all the aforementioned implementation methods. In another example, the current frame is the first frame; in this example, since the target features of the historical frames are not saved locally, the point cloud features of the current frame can be directly used as the initialization features of the current frame after obtaining the image features of the current frame based on the image data of the current frame and the point cloud features of the current frame based on the point cloud data of the current frame. Then, the target features of the current frame are obtained based on the initialization features of the current frame and the image features of the current frame, and the ground feature information is obtained based on the target features of the current frame. This will not be elaborated here.

[0106] It can be seen that, by adopting the above scheme, the target features of the current frame after time sequence and spatial fusion can be processed to obtain segmentation results of multiple types of ground elements, and then ground element information under each type can be determined. Due to the time sequence and spatial features of the target features of the current frame, the problem of being unable to obtain comprehensive feature information due to occlusion and the like can be avoided, and the accuracy of the final ground element information is ensured.

[0107] In combination Figure 4 Taking the target features as BEV features, the initialization features of the current frame as initialization BEV features of the current frame, the image acquisition device as a camera, and the point cloud acquisition device as a laser radar as examples, the foregoing ground element information acquisition method is exemplarily described:

[0108] S401, real-time acquisition of image data of a current frame collected by a camera and real-time acquisition of point cloud data of the current frame collected by a laser radar.

[0109] S402, inputting the point cloud data of the current frame into a first model to obtain point cloud features of the current frame output by the first model, and inputting the image data of the current frame into a second model to obtain image features of the current frame output by the second model.

[0110] S403, performing time sequence fusion and spatial fusion. Specifically, the following steps can be included:

[0111] S4031, obtaining initialization BEV features of the current frame based on the point cloud features of the current frame and BEV features of a historical frame.

[0112] S4032, obtaining BEV features of the current frame based on the initialization BEV features of the current frame and the image features of the current frame.

[0113] S404, inputting the BEV features of the current frame into a segmentation model to obtain ground element segmentation results of N categories output by the segmentation model.

[0114] S405, obtaining ground element information under each category of the N categories based on the ground element segmentation results of the N categories.

[0115] In combination Figure 5 Similarly, taking the target features as BEV features, the initialization features of the current frame as initialization BEV features of the current frame, the image acquisition device as a camera, and the point cloud acquisition device as a laser radar as examples, another exemplarily description is performed:

[0116] S501, real-time acquisition of image data of a current frame collected by a camera and real-time acquisition of point cloud data of the current frame collected by a laser radar.

[0117] S502. Input the point cloud data of the current frame into a first model to obtain point cloud features of the current frame output by the first model; and input the image data of the current frame into a second model to obtain image features of the current frame output by the second model.

[0118] S503 : Adjust the BEV features of the historical frame to obtain the adjusted BEV features of the historical frame aligned with the point cloud feature positions of the current frame.

[0119] S504, adjusting the point cloud features of the current frame to obtain adjusted point cloud features of the current frame; wherein the adjusted point cloud features of the current frame have the same number of channels as the adjusted target features of the historical frame;

[0120] S505: Determine a first mask based on the adjusted BEV features of the historical frame, and add the adjusted BEV features of the historical frame to the adjusted point cloud features of the current frame to obtain an initial fusion feature of the current frame;

[0121] S506 : Normalize the initial fusion features of the current frame based on the first mask to obtain the initialized BEV features of the current frame.

[0122] S507 : Perform self-attention processing on the initialized BEV features of the current frame to obtain the initialized BEV features of the current frame including attention weights.

[0123] S508 , performing cross-attention processing on the initialized BEV features of the current frame including the attention weights and the image features of the current frame to obtain the BEV features of the current frame.

[0124] S509: Input the BEV feature of the current frame into a segmentation model to obtain N types of ground feature segmentation results output by the segmentation model;

[0125] S510 : Based on the ground feature segmentation results of the N categories, obtain ground feature information of each category in the N categories.

[0126] Finally, the beneficial effects of the solution provided by this embodiment are explained in conjunction with related technologies: Related ground feature segmentation methods often only involve spatial fusion, fusing information from multiple sensors (typically a surround-view camera and a lidar) in the spatial dimension. However, when certain ground features are obscured by pedestrians or vehicles, neither the camera nor the lidar can obtain accurate information; when vehicles are moving too fast, blurring the camera image; or when rain or snow causes noise errors in the lidar, these issues can all lead to inaccurate fusion results.

[0127] Furthermore, the spatial fusion approach used in related ground feature segmentation methods generally involves converting image information into a pseudo-radar point cloud based on depth relationships, or projecting lidar point cloud information onto the image. This makes temporal fusion of feature information difficult, or requires the use of 3D convolution for temporal fusion, which exponentially increases the computational effort. Furthermore, the segmentation results obtained by projecting the point cloud onto the image are in image space, not BEV space, and cannot be directly used by backend tasks, requiring additional post-processing steps.

[0128] Furthermore, the ground feature segmentation task under the bird's-eye view (BEV) perspective can obtain a BEV semantic map, which facilitates higher-level tasks in autonomous driving, such as path planning and trajectory prediction. Projecting the feature information of multiple sensors into the BEV space based on spatial fusion to unify the expression of ground features has become one of the mainstream methods in autonomous driving technology. The BEV space can easily fuse different sensors, such as the common surround-view multi-camera and lidar, and project different feature information from their own sensor coordinate systems into the BEV space. However, the ground feature segmentation method based on spatial fusion alone still has the problem of inaccurate fusion results in complex surrounding environments such as the aforementioned excessive vehicle speed causing camera blur, obstruction by pedestrians and vehicles on the road, and bad weather.

[0129] Taking into account that in actual autonomous driving operation, all information is continuous, it is often easy to obtain the information of historical frames, and it can be continuously iterated. Secondly, the unified expression of BEV space facilitates the fusion of the information of the current frame and the information of the historical frame. Based on this, the method and scheme provided in this embodiment can effectively obtain image features based on image data collected by multiple cameras and point cloud features based on point cloud data collected by lidar in BEV space, and then fuse them in space and time, and obtain semantic segmentation results, and finally obtain ground feature information. Through this online time series fusion method, the information of historical frames is used to effectively improve the adaptability to complex environments, especially motion blur, occlusion and other situations, thereby improving the segmentation accuracy, ensuring the accuracy of the results, and significantly improving the robustness of map feature segmentation. Furthermore, the ground feature information obtained by adopting the above scheme can provide higher quality input for downstream tasks of autonomous driving.

[0130] The second embodiment of the present disclosure provides a ground element information acquisition device, such as Figure 6 Shown, including:

[0131] A data processing module 601 is configured to obtain image features of the current frame based on the image data of the current frame, and obtain point cloud features of the current frame based on the point cloud data of the current frame;

[0132] The initialization processing module 602 is configured to obtain initialization features of the current frame based on the point cloud features of the current frame and target features of historical frames.

[0133] The target feature processing module 603 is configured to obtain target features of the current frame based on the initialization features of the current frame and image features of the current frame.

[0134] The ground element acquisition module 604 is configured to obtain ground element information based on the target features of the current frame.

[0135] The initialization processing module is configured to adjust the target features of the historical frames to obtain adjusted target features of the historical frames that are aligned with the positions of the point cloud features of the current frame, and obtain the initialization features of the current frame based on the adjusted target features of the historical frames and the point cloud features of the current frame.

[0136] The initialization processing module is configured to obtain a device pose of the current frame and a device pose of the historical frames, determine pose change information based on the device pose of the current frame and the device pose of the historical frames, determine adjustment information based on the pose change information, and adjust the target features of the historical frames based on the adjustment information to obtain the adjusted target features of the historical frames that are aligned with the positions of the point cloud features of the current frame.

[0137] The initialization processing module is configured to adjust the point cloud features of the current frame to obtain adjusted point cloud features of the current frame, wherein the adjusted point cloud features of the current frame have the same number of channels as the adjusted target features of the historical frames, and fuse the adjusted point cloud features of the current frame and the adjusted target features of the historical frames to obtain the initialization features of the current frame.

[0138] The initialization processing module is configured to determine a first mask based on the adjusted target features of the historical frames, add the adjusted target features of the historical frames and the adjusted point cloud features of the current frame to obtain initial fusion features of the current frame, and perform normalization processing on the initial fusion features of the current frame based on the first mask to obtain the initialization features of the current frame.

[0139] The target feature processing module is configured to perform self-attention processing on the initialization features of the current frame to obtain initialization features of the current frame containing attention weights, and perform cross-attention processing on the initialization features of the current frame containing the attention weights and the image features of the current frame to obtain the target features of the current frame.

[0140] The ground element acquisition module is used to input the target features of the current frame into a segmentation model to obtain N categories of ground element segmentation results output by the segmentation model; N is a positive integer; based on the N categories of ground element segmentation results, obtain ground element information under each of the N categories.

[0141] The ground feature information acquisition device provided in this embodiment can be set in an electronic device. The specific processing of each module in the device of this embodiment is the same as that in the above-mentioned ground feature information acquisition method, and will not be repeated here.

[0142] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0143] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a vehicle, a readable storage medium, and a computer program product.

[0144] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0145] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the electronic device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0146] Multiple components in the electronic device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0147] The computing unit 701 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above. For example, in some embodiments, the various methods described above can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the various methods described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the various methods described above by any other appropriate means (e.g., by means of firmware).

[0148] According to yet another embodiment of the present disclosure, a vehicle is provided. The vehicle includes the electronic device 700 according to the above embodiment.

[0149] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0150] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0151] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0152] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0153] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0154] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0155] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0156] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for acquiring ground element information, comprising: Obtaining image features of the current frame based on the image data of the current frame, and obtaining point cloud features of the current frame based on the point cloud data of the current frame; Obtaining initialization features of the current frame based on the point cloud features of the current frame and the target features of the historical frames; Obtaining a target feature of the current frame based on the initialization feature of the current frame and the image feature of the current frame, wherein the target feature is a BEV feature; Obtaining ground feature information based on the target features of the current frame; The initialization features of the current frame are obtained based on the point cloud features of the current frame and the target features of the historical frames, including: adjusting the target features of the historical frame to obtain the adjusted target features of the historical frame that are aligned with the point cloud features of the current frame; adjusting the point cloud features of the current frame to obtain the adjusted point cloud features of the current frame; wherein the adjusted point cloud features of the current frame have the same number of channels as the adjusted target features of the historical frame; determining a first mask based on the adjusted target features of the historical frame; adding the adjusted target features of the historical frame to the adjusted point cloud features of the current frame to obtain the initial fusion features of the current frame; and normalizing the initial fusion features of the current frame based on the first mask to obtain the initialization features of the current frame.

2. The method according to claim 1, wherein The adjusting the target feature of the historical frame to obtain the adjusted target feature of the historical frame aligned with the point cloud feature position of the current frame includes: Get the acquisition device pose of the current frame and the acquisition device pose of the historical frame; Determining posture change information based on the acquisition device posture of the current frame and the acquisition device posture of the historical frame; determining adjustment information based on the posture change information; The target features of the historical frame are adjusted based on the adjustment information to obtain the adjusted target features of the historical frame aligned with the point cloud feature positions of the current frame.

3. The method according to claim 1, wherein The obtaining of target features of the current frame based on the initialization features of the current frame and the image features of the current frame includes: Performing self-attention processing on the initialization features of the current frame to obtain the initialization features of the current frame including attention weights; Cross-attention processing is performed on the initialization features of the current frame including the attention weights and the image features of the current frame to obtain the target features of the current frame.

4. The method according to any one of claims 1 to 3, wherein: The obtaining of ground feature information based on the target feature of the current frame includes: Inputting the target features of the current frame into a segmentation model to obtain N types of ground feature segmentation results output by the segmentation model; N is a positive integer; Based on the ground feature segmentation results of the N categories, ground feature information of each category in the N categories is obtained.

5. A ground feature information acquisition device, comprising: A data processing module, configured to obtain image features of the current frame based on the image data of the current frame, and obtain point cloud features of the current frame based on the point cloud data of the current frame; An initialization processing module, configured to obtain initialization features of the current frame based on the point cloud features of the current frame and the target features of the historical frames; a target feature processing module, configured to obtain a target feature of the current frame based on the initialization feature of the current frame and the image feature of the current frame, wherein the target feature is a BEV feature; A ground element acquisition module, configured to obtain ground element information based on target features of the current frame; The initialization processing module is used to adjust the target features of the historical frame to obtain the adjusted target features of the historical frame that are aligned with the point cloud features of the current frame; adjust the point cloud features of the current frame to obtain the adjusted point cloud features of the current frame; wherein the point cloud features of the adjusted current frame have the same number of channels as the target features of the adjusted historical frame; determine a first mask based on the adjusted target features of the historical frame; add the adjusted target features of the historical frame to the adjusted point cloud features of the current frame to obtain the initial fusion features of the current frame; and normalize the initial fusion features of the current frame based on the first mask to obtain the initialization features of the current frame.

6. The device according to claim 5, wherein The initialization processing module is used to obtain the acquisition device posture of the current frame and the acquisition device posture of the historical frame; based on the acquisition device posture of the current frame and the acquisition device posture of the historical frame, determine posture change information; Based on the posture change information, adjustment information is determined; based on the adjustment information, the target feature of the historical frame is adjusted to obtain the adjusted target feature of the historical frame aligned with the point cloud feature position of the current frame.

7. The device according to claim 5, wherein The target feature processing module is used to perform self-attention processing on the initialization feature of the current frame to obtain the initialization feature of the current frame including the attention weight; Cross-attention processing is performed on the initialization features of the current frame including the attention weights and the image features of the current frame to obtain the target features of the current frame.

8. The device according to any one of claims 5 to 7, wherein: The ground feature acquisition module is used to input the target features of the current frame into the segmentation model to obtain N types of ground feature segmentation results output by the segmentation model; N is a positive integer; Based on the ground feature segmentation results of the N categories, ground feature information of each category in the N categories is obtained.

9. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.

11. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 4.

12. A vehicle comprising the electronic device according to claim 9.

Citation Information

Patent Citations

  • Laser radar 3D real-time target detection method fusing multi-frame time sequence point cloud

    CN111429514A

  • Positioning method, positioning device and electronic equipment

    CN111722245A