Multi-modal obstacle detection method and apparatus, electronic device, and storage medium

By using a multimodal obstacle detection method that combines feature information from laser point clouds, visual images, and millimeter-wave point clouds, obstacle detection boxes are generated. This solves the problem of low detection accuracy for dynamic obstacles with abnormal shapes or features in existing technologies, and improves the detection accuracy and safety of autonomous driving systems.

CN119810792BActive Publication Date: 2025-11-18BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411834276.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-11-18
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

Existing data-driven 3D object detection algorithms cannot accurately detect dynamic obstacles with unusual shapes or features, resulting in low detection accuracy of autonomous driving systems when facing unusual dynamic obstacles.

Method used

A multimodal obstacle detection method is adopted, which combines the feature information of laser point cloud, visual image and millimeter wave point cloud to generate multimodal fusion features. Then, using laser point representation features, image point representation features and millimeter wave detection boxes, obstacle detection boxes are generated to achieve accurate detection of dynamic obstacles.

Benefits of technology

It improves the detection accuracy of non-standard dynamic obstacles, reduces the collision risk caused by missed detections in autonomous driving systems, and enhances the robustness and safety of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810792B_ABST
    Figure CN119810792B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-modal obstacle detection method and device, electronic equipment and storage medium, relates to the technical field of computers, in particular to the fields of computer vision, data processing and the like, and can be used in obstacle detection, automatic driving and the like application scenarios. The specific scheme is: according to the obtained laser point cloud, the point feature corresponding to any feature point in the laser point cloud and the laser point representation feature are extracted; the current frame laser point cloud is projected to a visual image, and the image point representation feature corresponding to the feature point is determined in combination with the image feature extracted from the visual image; the multi-modal fusion feature is obtained according to the point feature, the laser point representation feature and the image point representation feature; the reserved feature point set and the millimeter wave detection box are generated based on the multi-modal fusion feature and the millimeter wave point cloud, and are projected to a grid image; and the obstacle detection box is generated according to the reserved feature point set and the millimeter wave detection box. The present scheme can improve the detection accuracy of heterogeneous dynamic obstacles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, particularly to the fields of computer vision and data processing, and can be used in application scenarios such as obstacle detection and autonomous driving. Specifically, it relates to multimodal obstacle detection methods, devices, electronic devices, and storage media. Background Technology

[0002] In the field of autonomous driving technology, existing data-driven three-dimensional (3D) target detection algorithms can only detect trained obstacles and cannot accurately detect dynamic obstacles with unusual shapes or features. Summary of the Invention

[0003] This disclosure provides a multimodal obstacle detection method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of this disclosure, a multimodal obstacle detection method is provided, comprising: extracting point features and laser point representation features corresponding to any feature point in the laser point cloud based on an acquired laser point cloud; the laser point cloud includes a current frame laser point cloud; projecting the current frame laser point cloud onto a visual image, and combining image features extracted from the visual image to determine image point representation features corresponding to the feature points; obtaining multimodal fusion features based on the point features, laser point representation features, and image point representation features; generating a set of retained feature points and a millimeter-wave detection box based on the multimodal fusion features and the millimeter-wave point cloud, and projecting the set of retained feature points and the millimeter-wave detection box onto a raster image; and generating an obstacle detection box based on the set of retained feature points and the millimeter-wave detection box on the raster image.

[0005] According to a second aspect of this disclosure, a multimodal obstacle detection device is provided, comprising: a laser feature extraction module, configured to extract point features and laser point representation features corresponding to any feature point in the acquired laser point cloud; the laser point cloud including a current frame laser point cloud; an image feature extraction module, configured to project the current frame laser point cloud onto a visual image, and combine the image features extracted from the visual image to determine the image point representation features corresponding to the feature points; a multimodal feature extraction module, configured to obtain multimodal fusion features based on the point features, laser point representation features, and image point representation features; a grid projection module, configured to generate a set of retained feature points and a millimeter-wave detection box based on the multimodal fusion features and the millimeter-wave point cloud, and project the set of retained feature points and the millimeter-wave detection box onto a grid image; and an obstacle detection module, configured to generate an obstacle detection box based on the set of retained feature points and the millimeter-wave detection box on the grid image.

[0006] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0007] At least one processor; and

[0008] The memory is communicatively connected to the at least one processor; wherein,

[0009] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.

[0010] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.

[0011] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.

[0012] According to a sixth aspect of this disclosure, an autonomous vehicle is provided, including electronic devices as described in the third aspect.

[0013] The solution disclosed herein can detect various types of dynamic obstacles and improve the detection accuracy of different types of dynamic obstacles.

[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0015] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0016] Figure 1 This is a schematic flowchart of a multimodal obstacle detection method according to an embodiment of the present disclosure;

[0017] Figure 2 This is a schematic diagram illustrating the process of obtaining multimodal fusion features based on laser point clouds and visual images according to an embodiment of this disclosure;

[0018] Figure 3 This is a schematic diagram illustrating the process of obtaining an obstacle detection box based on multimodal fusion features and millimeter-wave point clouds according to an embodiment of this disclosure;

[0019] Figure 4 This is a schematic diagram of the structure of a multimodal obstacle detection device according to an embodiment of the present disclosure;

[0020] Figure 5 This is a schematic diagram of a multimodal obstacle detection scenario according to an embodiment of the present disclosure;

[0021] Figure 6 This is a schematic diagram of the structure of an electronic device used to implement the multimodal obstacle detection method of the present disclosure embodiments. Detailed Implementation

[0022] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0023] In this document, the term "and / or" merely describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The term "at least one" in this document indicates any combination of at least two of a plurality of elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this document refer to and distinguish between multiple similar technical terms, not to restrict the order or to limit there to only two. For example, "first feature" and "second feature" refer to two categories / two features; the first feature can be one or more, and the second feature can also be one or more.

[0024] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0025] Before introducing the technical solutions of the embodiments of this disclosure, the technical terms that may be used in this disclosure will be further explained:

[0026] Multi-sensor fusion: This involves integrating data from multiple sensors to obtain more accurate and reliable information. By fusing data from different sensors, the limitations of a single sensor can be overcome, improving the robustness and accuracy of the system. For example, in autonomous driving, fusing data from LiDAR, cameras, and millimeter-wave radar allows for more accurate perception of the vehicle's surroundings, enhancing driving safety.

[0027] Obstacle detection: This uses sensors and algorithms to identify and locate objects in the environment to avoid collisions. Obstacle detection systems typically use multiple sensors, such as cameras, LiDAR, ultrasonic sensors, and millimeter-wave radar, to acquire environmental data. Then, image processing and machine learning algorithms are used to analyze the data, identify potential obstacles, and take appropriate avoidance measures. This helps improve the system's safety and efficiency.

[0028] In related technologies, with the popularization and development of deep learning and LiDAR, it has become increasingly possible to use deep learning methods for point cloud target detection. However, for some dynamic obstacles with unusual shapes or features on roads, such as tree-pulling vehicles and irregularly shaped non-motorized tricycles, current 3D target detection algorithms often have limitations and cannot detect these unusual dynamic obstacles. Existing data-driven 3D target detection algorithms can only detect obstacles on the whitelist already in the training set (such as cars, trucks, pedestrians, etc.), but cannot detect uncommon non-whitelisted obstacles (such as irregularly shaped vehicles, fallen pedestrians, pedestrians pushing carts, etc.), and cannot accurately detect dynamic obstacles with unusual shapes or features, resulting in low detection accuracy.

[0029] In order to at least partially solve one or more of the above-mentioned problems and other potential problems, this disclosure proposes a multimodal obstacle detection method that can make full use of the advantages of various sensors, effectively mitigate the collision risk caused by missed detection in autonomous vehicles, and achieve the detection of non-whitelisted dynamic obstacles by calculating the speed of each grid.

[0030] This disclosure provides a multimodal obstacle detection method. Figure 1 This is a flowchart illustrating a multimodal obstacle detection method according to an embodiment of the present disclosure. This multimodal obstacle detection method can be applied to a multimodal obstacle detection device. The multimodal obstacle detection device is located in an electronic device. The electronic device includes, but is not limited to, fixed devices and / or mobile devices. For example, fixed devices include, but are not limited to, servers, which can be cloud servers or ordinary servers. Mobile devices include, but are not limited to, mobile phones, tablets, and vehicle-mounted terminals. In some possible implementations, the multimodal obstacle detection method can also be implemented by a processor calling computer-readable instructions stored in memory. Figure 1 As shown, the multimodal obstacle detection method includes:

[0031] S101. Based on the acquired laser point cloud, extract the point features and laser point representation features corresponding to any feature point in the laser point cloud; the laser point cloud includes the laser point cloud of the current frame;

[0032] S102. Project the laser point cloud of the current frame onto the visual image, and combine the image features extracted from the visual image to determine the image point representation features corresponding to the feature points.

[0033] S103. Based on point features, laser point representation features, and image point representation features, multimodal fusion features are obtained;

[0034] S104. Based on multimodal fusion features and millimeter-wave point cloud, generate a set of retained feature points and millimeter-wave detection boxes, and project the set of retained feature points and millimeter-wave detection boxes onto the raster image;

[0035] S105. Generate obstacle detection boxes based on the set of preserved feature points on the grid map and the millimeter-wave detection boxes.

[0036] In this embodiment of the disclosure, the laser point cloud is a three-dimensional dataset acquired through laser scanning technology. It consists of a large number of points, each with its precise coordinates in space. These points collectively depict the three-dimensional shape and structure of an object or environment. Laser point clouds are typically generated by lidar devices, which acquire spatial information by emitting laser pulses and measuring the reflection time to calculate distances.

[0037] In this embodiment of the disclosure, feature points in a laser point cloud refer to points extracted from the laser point cloud data that possess uniqueness or saliency, exhibiting certain special geometric or intensity characteristics in space. Exemplarily, feature points can be determined from the laser point cloud using feature point detection algorithms such as Intrinsic Shape Signature (ISS) based on point cloud curvature and normal vectors, or 3D keypoint detection methods based on the geometric properties of local surfaces (Harris3D). Other methods in the prior art can also be used to determine feature points. The above are merely illustrative examples and are not intended to limit all possible cases for determining feature points; they are simply not exhaustive.

[0038] In this embodiment of the disclosure, point features refer to features extracted from laser point cloud data to describe the local or global attributes of feature points. These are basic attributes directly extracted from each feature point in the laser point cloud data. These features may include geometric features such as curvature and normal vectors, and / or intensity features such as reflection intensity, as well as other information that can help identify and classify points. The above is only an illustrative example and is not intended to limit all possible cases of point features; it is simply not exhaustive here.

[0039] In this embodiment of the disclosure, the laser point representation feature refers to the feature vector or descriptor used to represent and encode feature points in laser point cloud data. It is generated by extracting features from the laser point cloud data and then performing a point representation transformation on those features. In particular, the representation of each point in the laser point representation feature not only includes its own features but also global and multi-scale information.

[0040] In this embodiment of the disclosure, the current frame of laser point cloud refers to the laser point cloud data at a specific moment. In the dynamic scenario of autonomous driving, the vehicle-mounted LiDAR device continuously scans the environment, generating a series of laser point cloud frames. Each frame represents environmental information collected at different times, and "current frame" emphasizes the frame that is being processed or analyzed at the current moment among a series of continuously collected point cloud data.

[0041] In this embodiment of the disclosure, the visual image is a two-dimensional image captured by optical devices such as cameras, representing the scene within the device's field of view. In the dynamic scene of autonomous driving, the vehicle-mounted camera device continuously captures the environment within its field of view, generating a series of visual images. In this embodiment of the disclosure, the visual image refers to the current frame visual image, that is, the frame being processed or analyzed at the current moment among a series of continuously acquired visual images.

[0042] In this embodiment of the disclosure, image features refer to specific attributes or information extracted from a visual image that can describe its content. These features can be extracted using deep learning methods such as image encoders, or using existing methods such as manual feature extraction or feature descriptors. The above is only an illustrative example and is not intended to limit the scope of all methods for extracting image features. It is not an exhaustive list here.

[0043] In this embodiment of the disclosure, image point representation features refer to feature vectors or descriptors used to represent and encode feature points in a visual image, which can be used to better represent and distinguish different feature points in a high-dimensional space. Image point representation features can be a set of image features of all feature points on a visual image after projecting a laser point cloud onto the visual image. This set contains the feature information corresponding to each feature point in the image, combining geometric and visual information for further analysis and processing.

[0044] In this embodiment, multimodal fusion features refer to the combination of feature information from different sensors or data sources to form a more comprehensive and accurate representation. This integrates data from different modalities, leveraging their respective advantages and compensating for the shortcomings of a single modality. Multimodal fusion features fuse point features, laser point representation features, and image point representation features, utilizing their respective strengths. By combining the three-dimensional information of point features and laser point representation features with the two-dimensional information of image point representation features, a more comprehensive feature representation is formed. This enables the system to more fully understand the environment. Combining the advantages of LiDAR and cameras improves accuracy and robustness. For example, multimodal fusion features may include: speed information and semantic categories.

[0045] In this embodiment, the millimeter-wave point cloud is a type of three-dimensional point cloud data generated by millimeter-wave radar. Millimeter-wave radar emits millimeter-level electromagnetic waves to detect objects and receive reflected signals, thereby acquiring information about the object's distance, velocity, and angle. This information is then processed into point cloud data, i.e., a series of spatial coordinate points, constituting a three-dimensional model of the target object. Specifically, the electromagnetic waves emitted and received by millimeter-wave radar can be used to detect moving objects. By analyzing the frequency changes of the reflected signals, millimeter-wave radar can measure the velocity of an object using the Doppler effect. Moving objects cause changes in the frequency of the reflected signals; an increase in frequency indicates the object is approaching, while a decrease indicates it is moving away. In the field of autonomous driving, millimeter-wave point clouds generated by millimeter-wave radar can accurately detect and track the position and velocity of moving objects, enabling more accurate detection of all moving objects in the environment and ensuring comprehensive obstacle detection.

[0046] In this embodiment of the disclosure, the retained feature point set refers to the set of feature points selected and retained during the feature point filtering process. This reduces data complexity while retaining contributing feature information. The retained feature point set is obtained by filtering all feature points based on multimodal fusion features, which can retain all necessary feature points and remove unnecessary feature points, thereby reducing data volume and improving processing efficiency.

[0047] In this embodiment of the disclosure, the millimeter-wave detection box refers to a three-dimensional bounding box created using point cloud data generated by millimeter-wave radar to represent and track objects. Since a millimeter-wave point cloud is a collection of discrete points of an object in space, the point cloud data is first analyzed by an algorithm to identify the set of points belonging to the same object. Then, based on the identified set of points, a minimal three-dimensional bounding box is generated to enclose these points, thus obtaining the millimeter-wave detection box.

[0048] In this embodiment, the process of projecting the retained feature point set and millimeter-wave detection boxes onto a grid map involves projecting the three-dimensional feature points and detection boxes onto a two-dimensional grid map. The grid map is a plane composed of a grid, where each grid cell represents a fixed spatial region. On the grid map, the information of the feature points and detection boxes is mapped to the corresponding grid cells, forming a two-dimensional representation. This simplifies three-dimensional information into a two-dimensional grid map, reducing computational complexity. Furthermore, in the field of autonomous driving, grid maps aid in path planning and navigation, providing clear obstacle information.

[0049] In this embodiment of the disclosure, the process of generating obstacle detection boxes involves using information from a two-dimensional grid image to identify and describe obstacles. By combining information from a feature point set and millimeter-wave detection boxes, an algorithm is used to more accurately identify the position and shape of obstacles. Based on the fused information, a more precise obstacle detection box is generated on the grid image, representing the boundary of the detected obstacle. In the field of autonomous driving, generating obstacle detection boxes can improve the accuracy of obstacle detection.

[0050] In this embodiment, laser point cloud data scanned by a vehicle-mounted LiDAR device is acquired. Based on selected feature points, relevant point features are extracted from the point cloud, and then laser point representation features are extracted using nonlinear transformation and other methods. Through the above steps, rich feature information, including point features and laser point representation features, can be extracted from the laser point cloud.

[0051] Furthermore, after extracting image features from the visual image using feature extraction algorithms or neural network models, coordinate transformation is performed using the known relative positions and orientations between the camera and the LiDAR, thereby projecting the 3D point cloud data acquired by the LiDAR onto the 2D visual image. Next, for each feature point projected onto the image, a quadratic linear interpolation method is used to obtain the corresponding feature value from the image features. Through these steps, each feature point in the LiDAR point cloud is assigned a corresponding image feature, and the image features of all feature points constitute the image point representation features.

[0052] Furthermore, the semantic category and velocity information are predicted for each point in the current frame. The velocity estimation task outputs the two-dimensional (2D) velocity of each point, representing the 2D velocity in the X and Y directions in the current coordinate system. The semantic segmentation task outputs the semantic segmentation category of each point in the point cloud of the current frame, including movable objects, ground, curbs, fences, vegetation, etc.

[0053] Furthermore, based on multimodal fusion features, feature points are filtered according to preset rules, and those meeting the requirements are retained as the retained feature points. All retained feature points constitute the retained feature point set. Then, millimeter-wave point cloud data obtained from vehicle-mounted millimeter-wave radar scanning is acquired. Based on the millimeter-wave point cloud, point sets belonging to the same object are identified. Subsequently, a minimal 3D bounding box is generated to enclose these points, thus obtaining the millimeter-wave detection box. Next, the x and y coordinates of each feature point in the retained feature point set are mapped to the corresponding position on the raster image, ignoring the z coordinate. The corresponding grid cells are marked on the raster image according to the point's position, thus projecting the retained feature point set onto the raster image. Similarly, the bottom surface of the millimeter-wave detection box is projected onto the raster image. The boundary of the detection box is calculated, and the grid cells covered by the box are marked on the raster image, thus projecting the millimeter-wave detection box onto the raster image.

[0054] Finally, the set of preserved feature points is identified and analyzed on the grid map, and millimeter-wave detection boxes are used to determine the initial location and extent of potential obstacles. Then, the information from the set of preserved feature points and millimeter-wave detection boxes is combined to generate more accurate obstacle detection boxes on the grid map.

[0055] The technical solution of this disclosure can, based on acquired laser point clouds, millimeter-wave point clouds, and visual images, extract point features, laser point representation features, and image point representation features to generate multimodal fusion features, thereby generating a feature-preserving set of points and a millimeter-wave detection box, ultimately generating an obstacle detection box. This disclosure fully utilizes the advantages of various sensors, effectively mitigating the collision risk caused by missed detections in autonomous vehicles. Simultaneously, through various algorithm designs, it avoids the risk of sudden braking caused by noise-induced false detections, improving the versatility of obstacle detection, achieving accurate detection of dynamic obstacles, and enhancing the detection accuracy of heterogeneous dynamic obstacles.

[0056] In some embodiments, the laser point cloud includes not only the laser point cloud of the current frame but also the laser point cloud of historical frames. Based on the acquired laser point cloud, point features and laser point representation features corresponding to any feature point in the laser point cloud are extracted, including: after dimensional expansion of the laser point cloud of the current frame, point features corresponding to the feature points are extracted; the laser point cloud of the current frame and the laser point cloud of historical frames are projected onto a bird's-eye view planar image, and laser point representation features are extracted.

[0057] In this embodiment of the disclosure, the historical frame laser point cloud refers to the laser point cloud of the previous frame of the current frame laser point cloud.

[0058] In this embodiment of the disclosure, dimensional expansion refers to adding information of other dimensions, such as color, intensity, and normal vectors, to the original three-dimensional coordinates to enrich the features of the point cloud data. Dimensional expansion of laser point clouds allows for more accurate identification and description of feature points through the expanded dimensional information, thereby improving the analysis and processing capabilities of the point cloud data.

[0059] In this embodiment of the disclosure, projecting the current frame laser point cloud and the historical frame laser point cloud onto a bird's-eye view image means projecting the three-dimensional point cloud data in the current frame laser point cloud and the historical frame laser point cloud onto a horizontal two-dimensional plane. This can convert three-dimensional data into a two-dimensional plane, thereby simplifying the data processing and analysis process.

[0060] Thus, by expanding the dimensions of the current frame's laser point cloud, feature points can be identified and described more accurately, thereby improving the analysis and processing capabilities of point cloud data. By projecting the current frame's laser point cloud and historical frame's laser point cloud onto a bird's-eye view image, three-dimensional data can be converted into a two-dimensional plane, thereby simplifying the data processing and analysis process.

[0061] In some embodiments, after dimensional expansion of the current frame laser point cloud, point features corresponding to the feature points are extracted, including: dimensional expansion of the current frame laser point cloud using a neural network to obtain enhanced point cloud data; and extraction of point features corresponding to the feature points based on the enhanced point cloud data.

[0062] In this embodiment of the disclosure, the input point cloud data is processed by a trained neural network model to add additional dimensional information. After neural network processing, the enhanced point cloud data not only contains the original spatial coordinates but also includes expanded feature information. For example, the expanded feature information may include features such as the color, reflection intensity, and normal vector of each point. Other features may also be added according to the actual situation. The above is only an illustrative example and is not intended to limit all possible cases of dimensional expansion; it is not exhaustive here.

[0063] For example, the process of using a neural network to augment the dimensions of a current frame of laser point cloud data to obtain enhanced point cloud data may include: preprocessing the laser point cloud data of the current frame, which may include denoising, normalization, and coordinate transformation to ensure that the data is suitable for input into the neural network; selecting a suitable neural network architecture for processing laser point cloud data, such as PointNet or Deep Graph Convolutional Neural Network (DGCNN); training the neural network using a labeled dataset to enable it to learn the ability to extract and augment features from the laser point cloud, during which the network optimizes its weights to minimize prediction error; and inputting the processed point cloud data into the trained neural network, which outputs enhanced point cloud data containing the coordinates of the original points and the augmented feature dimensions. The above is merely an illustrative example and does not represent all possible scenarios for augmenting the dimensions of a current frame of laser point cloud data using a neural network; it is simply not exhaustive.

[0064] In this embodiment of the disclosure, point feature representations corresponding to feature points are extracted. From the enhanced point cloud data, feature information related to feature points is identified and extracted. The feature information may include the spatial location, normal vector, color information, etc., of the feature points. By extracting these features, the environment or objects can be better described and analyzed, improving performance in subsequent tasks.

[0065] For example, the process of extracting point features corresponding to feature points from enhanced point cloud data may include: local neighborhood analysis, where for each feature point, a local neighborhood of spherical or cubic size is defined, and nearby point data is collected within this range; and feature descriptor calculation, where point features are calculated within this local neighborhood. The above is merely an illustrative example and is not intended to limit all possible cases of extracting point features corresponding to feature points from enhanced point cloud data; it is simply not exhaustive.

[0066] Thus, by augmenting the dimensions of laser point clouds, the originally sparse and discrete point cloud data can be enriched and made more useful, thereby improving the effectiveness of analysis and application. By extracting point features from the enhanced point cloud data, the environment or objects can be better described and analyzed, improving performance in subsequent tasks.

[0067] In some embodiments, projecting the current frame laser point cloud and the historical frame laser point cloud onto a bird's-eye view image to extract laser point representation features includes: projecting the current frame laser point cloud and the historical frame laser point cloud onto a bird's-eye view image to obtain the current frame bird's-eye view projection and the historical frame bird's-eye view projection; determining the current frame bird's-eye view handcrafted features and the historical frame bird's-eye view handcrafted features based on the current frame bird's-eye view projection and the historical frame bird's-eye view handcrafted features, respectively; performing feature interaction and feature fusion on the current frame bird's-eye view handcrafted features and the historical frame bird's-eye view handcrafted features to obtain bird's-eye view features; and performing interpolation processing on the bird's-eye view features to extract laser point representation features.

[0068] In this embodiment of the disclosure, handcrafted features refer to features designed and extracted by human experts based on domain knowledge and experience during data processing. These features are used to help machine learning models understand and process data. Handcrafted features are typically statistical measures or descriptive indicators extracted from raw data that can better represent the key attributes of the data. For example, in the field of autonomous driving, handcrafted features include, but are not limited to, maximum height, minimum height, and average intensity. Other handcrafted features may also be selected according to actual conditions. The above is only an illustrative example and is not intended to limit all possible cases of handcrafted features; it is simply not exhaustive.

[0069] In this embodiment, since the 3D laser point cloud data of the current frame and historical frames are projected onto a 2D bird's-eye view plane, and this projection focuses on ground plane information while ignoring details such as height and intensity, manual feature extraction is required to supplement the corresponding details. For example, the maximum height can be obtained by finding the maximum height value of all points in the bird's-eye view projection and using this value as the maximum height; similarly, the minimum height can be obtained by finding the minimum height value of all points in the bird's-eye view projection and using this value as the minimum height; similarly, the average intensity can be obtained by calculating the average intensity of all points in the bird's-eye view projection and using this value as the average intensity. Specifically, the height and intensity values ​​of each point exist in the data of the corresponding point in the laser point cloud.

[0070] In this embodiment, the bird's-eye view feature obtained by performing feature interaction and feature fusion on the handcrafted bird's-eye view features of the current frame and historical frames can provide an environmental description that combines temporal and spatial information. For example, feature interaction can identify the relationship and changes between the current frame and historical frames by comparing and analyzing the handcrafted features of the current frame and historical frames. In the field of autonomous driving, this can help capture the movement of dynamic objects or changes in the environment. Furthermore, feature fusion can include integrating the features of the current frame and historical frames to form a unified feature representation. The feature fusion process can be performed through weighted averaging, feature concatenation, or using a deep learning model. The above are merely illustrative examples and are not intended to limit all possible cases of feature interaction and feature fusion using handcrafted features; they are simply not exhaustive.

[0071] In this embodiment of the disclosure, since the bird's-eye view features are typically extracted from sparse laser point clouds, spatial discontinuities may exist. Interpolation is used to fill these gaps, generating a more continuous and complete feature representation. This can be achieved through various methods, such as linear interpolation, bilinear interpolation, or more complex algorithms. Subsequently, the interpolated bird's-eye view features can be mapped back to the original laser point cloud, with each laser point obtaining a feature representation. The feature representations of all points are combined to generate the laser point representation features.

[0072] In this embodiment, the point clouds of the current frame and historical frames can be projected onto the bird's-eye view plane to calculate handcrafted features such as maximum height, minimum height, and average intensity, introducing spatial and historical information into the model. Then, a two-dimensional U-Net (2D-UNet) structure is used to perform sufficient cross-scale feature interaction and cross-scale feature fusion to obtain the bird's-eye view features. Next, quadratic linear interpolation is used to transform the bird's-eye view features into point representations, and the point representations of all points constitute the laser point representation features.

[0073] Thus, by projecting the laser point cloud of current and historical frames onto a bird's-eye view image, the dynamic changes in the vehicle's surrounding environment can be understood more intuitively. Feature interaction and fusion utilize time-series information, enabling more accurate identification and tracking of dynamic objects. Through interpolation and feature extraction, sparse data can be handled better, reducing the probability of false positives and false negatives.

[0074] In some embodiments, the current frame laser point cloud is projected onto a visual image, and image features extracted from the visual image are combined to determine the image point representation features corresponding to the feature points, including: extracting image features based on the visual image; projecting the current frame laser point cloud onto the visual image to obtain a point cloud fusion image; and determining the obtained image point representation features based on the point cloud fusion image and the image features.

[0075] In this embodiment of the disclosure, projecting the current frame of laser point cloud onto a visual image can combine the 3D point cloud data from the lidar with the 2D image data from the visual camera, generating a point cloud fusion image that can provide rich environmental information. For example, the process of projecting the current frame of laser point cloud onto a visual image to obtain a point cloud fusion image may include: ensuring that the coordinate systems of the lidar and the camera are aligned through a calibration process; projecting the 3D laser point cloud data onto the 2D camera image plane, using intrinsic parameters such as the camera's focal length and principal point, and extrinsic parameters such as rotation matrix and translation vector, to perform geometric transformations; and combining the projected point cloud data with the visual image to generate a point cloud fusion image, which simultaneously contains the depth information from the lidar and the color information from the camera.

[0076] In this embodiment of the disclosure, the process of determining the image point representation features based on the point cloud fusion image and image features is obtained by combining the three-dimensional depth information of each feature point extracted from the point cloud fusion image with the image information, which can better describe objects and scenes in the environment.

[0077] In this embodiment of the disclosure, an image encoder can be used to extract image features first, then the laser point cloud can be projected onto the image by calculating the pose transformation between the camera and the lidar, and then the image features corresponding to each point can be calculated by quadratic linear interpolation, and the image features corresponding to each point can be combined to form image point representation features.

[0078] Thus, by combining the depth information of the LiDAR point cloud with the color and texture information of the visual image, the system can perceive its surroundings more accurately. This fusion provides a more comprehensive scene representation, aiding in object identification and classification. Combining multimodal data reduces errors caused by the limitations of a single sensor; the depth information provided by the LiDAR can compensate for the shortcomings of cameras in situations of changing lighting and lack of texture. The point cloud fusion image obtained by fusing the image and point cloud features can more accurately detect and track dynamic and static objects, improving the performance of autonomous driving systems in complex traffic environments.

[0079] Figure 2 This illustrates the process of obtaining multimodal fusion features based on laser point clouds and visual images, such as... Figure 2 As shown, laser feature extraction is first performed. On one hand, a multilayer perceptron is used to expand the dimensionality of the current frame's laser point cloud to extract point features. On the other hand, the current frame's laser point cloud and the laser point clouds of historical frames are projected onto a bird's-eye-view (BEV) plane to calculate manual features such as maximum height, minimum height, and average intensity. Then, a 2D-UNet structure is used for feature interaction and fusion to obtain bird's-eye-view features. Next, quadratic linear interpolation is used to transform the bird's-eye-view features into BEV point representation features, which are the laser point representation features in this embodiment. Next, image feature extraction is performed. An image encoder is used to extract image features. Then, the point cloud is projected onto the image by calculating the pose transformation between the camera and radar. Next, quadratic linear interpolation is used to calculate the image features corresponding to each point, finally obtaining the image point representation features. Subsequently, multimodal feature fusion is performed, concatenating point features, laser point representation features, and image point representation features. A multilayer perceptron is then used to fully fuse these multimodal features, resulting in multimodal fused features. Finally, these multimodal fused features are fed into a neural network to perform semantic segmentation and velocity estimation for each point in the current frame, predicting its semantic category and velocity information. The velocity estimation task outputs the 2D velocity of each point, including the 2D velocity in the X and Y directions in the current coordinate system, while the semantic segmentation task outputs the semantic segmentation category of each point in the current frame's point cloud, including movable objects, ground, curbs, fences, and vegetation.

[0080] In some embodiments, based on multimodal fusion features and millimeter-wave point clouds, a set of retained feature points and millimeter-wave detection boxes are generated, and the set of retained feature points and millimeter-wave detection boxes are projected onto a raster image, including: filtering feature points according to multimodal fusion features, generating a set of retained feature points according to a first filtering result, and projecting the set of retained feature points onto a raster image; filtering millimeter-wave point clouds, generating millimeter-wave detection boxes according to a second filtering result, and projecting the millimeter-wave detection boxes onto a raster image.

[0081] In this embodiment of the disclosure, the process of filtering feature points based on multimodal fusion features involves selecting useful feature points from all feature points according to preset rules based on the multimodal fusion features. This process removes noise or irrelevant information, retaining the most representative feature points. Subsequently, the coordinates of the retained feature points are converted into raster coordinates, and then the feature points are marked on the raster based on the converted raster coordinates.

[0082] In this embodiment of the disclosure, the process of filtering millimeter-wave point clouds involves selecting useful points from all millimeter-wave points according to preset rules, which can remove noise or irrelevant information to retain the most representative feature points. Subsequently, the three-dimensional coordinates of the millimeter-wave detection box are converted into two-dimensional coordinates of the raster image. Based on the size and position of the detection box, its projection boundary on the raster image is calculated. Finally, the area of ​​the detection box is marked on the raster image.

[0083] In this way, by performing multi-level screening and fusion of feature points and millimeter-wave point clouds, false alarms and false negatives can be effectively reduced, thereby improving the overall performance of the autonomous driving system.

[0084] In some embodiments, feature points are filtered based on multimodal fusion features, a set of retained feature points is generated based on the first filtering result, and the set of retained feature points is projected onto a raster image, including: determining the semantic category of any feature point based on the multimodal fusion features; filtering based on the semantic category to extract all feature points that conform to the preset semantic category and generate a set of retained feature points; and projecting all retained feature points in the set of retained feature points onto a raster image.

[0085] In this embodiment of the disclosure, the multimodal fusion feature may include a semantic category. The process of filtering feature points based on the multimodal fusion feature can be based on the semantic category of the feature points. Feature points corresponding to fixed objects such as road surface points, fence points, roadside points, and green plant points are removed, and only feature points corresponding to movable objects are retained. The feature points corresponding to movable objects are used as the first filtering result, thereby generating a set of retained feature points.

[0086] In this embodiment of the disclosure, all feature points are traversed, and road surface points, fence points, roadside points, and green plant points are filtered according to the semantic category of the feature points. Only the movable object category is retained, and the feature points under the movable object category are used as the first filtering result, thereby generating a set of retained feature points.

[0087] In this way, semantic filtering can better understand the types of objects in the surrounding environment, more effectively detect and track important objects, optimize computing resources, reduce unnecessary data processing, and thus reduce accident risks and improve driving safety.

[0088] In some embodiments, filtering the millimeter-wave point cloud, generating a millimeter-wave detection box based on the second filtering result, and projecting the millimeter-wave detection box onto a raster image includes: filtering target points in the millimeter-wave point cloud, extracting all target points that conform to a preset dynamic / static category to obtain a dynamic target point set; clustering all dynamic target points in the dynamic target point set to obtain a clustering result; generating a millimeter-wave detection box based on the clustering result, and projecting the millimeter-wave detection box onto a raster image.

[0089] In this embodiment of the disclosure, the information of each millimeter-wave point in the millimeter-wave point cloud includes velocity information measured based on the Doppler effect, representing the velocity of an object relative to the radar, which helps identify moving objects. Further, based on the velocity information, the millimeter-wave points are divided into dynamic and static categories. Based on the classification of the millimeter-wave points, all millimeter-wave points classified as dynamic are designated as target points. All millimeter-wave points are traversed to obtain a set of dynamic target points.

[0090] In this embodiment of the disclosure, clustering algorithms can be used to extract features for clustering, such as spatial coordinates and velocity vectors, and then points belonging to the same object can be grouped into one category. The set of points can be transformed into an object, which can more accurately estimate the motion state and trajectory of the object, and help to more accurately understand and predict the environment. This helps to identify and track different moving objects.

[0091] Thus, by filtering and clustering dynamic target points, moving objects such as pedestrians and other vehicles can be identified and distinguished more accurately, improving the accuracy of target detection and reducing false alarms and false negatives. The clustering of dynamic target points and the generation of detection boxes help autonomous driving systems better understand the dynamic changes in the surrounding environment and identify potential traffic participants and obstacles.

[0092] In some embodiments, generating obstacle detection boxes based on a set of preserved feature points on a grid image and millimeter-wave detection boxes includes: determining the presence of millimeter waves in each grid cell based on the millimeter-wave detection boxes on the grid image; determining the presence of obstacles in each grid cell based on the set of preserved feature points on the grid image; traversing all grid cells in the grid image based on the presence of millimeter waves and the presence of obstacles to obtain an obstacle grid cell set; clustering the obstacle grid cell set; and generating obstacle detection boxes based on the clustering results.

[0093] In this embodiment of the disclosure, a raster diagram is a two-dimensional grid structure used to represent spatial information. The basic unit of a raster diagram is a grid, and each grid represents a small area in space.

[0094] In this embodiment of the disclosure, based on the millimeter wave detection frame on the grid map, it is possible to determine whether there is a millimeter wave in each grid. When there is a millimeter wave in the grid, it means that there is a dynamic object in the grid, which may be an obstacle.

[0095] In this embodiment of the disclosure, based on the set of retained feature points on the grid map, it is possible to determine whether there is a movable object within the grid. When there is a movable object within the grid, it indicates that there may be an obstacle within the grid.

[0096] In this embodiment of the disclosure, the process of traversing all grids in the grid diagram to obtain the obstacle grid set refers to filtering out all grids that may contain obstacles, thereby obtaining the obstacle grid set.

[0097] In this embodiment of the disclosure, clustering the obstacle grid set and generating obstacle detection boxes based on the clustering results refers to grouping the grid cells marked as obstacles in the grid image to identify and distinguish different obstacles. Clustering algorithms such as Density-Based Spatial Clustering of Applications with Noise (DBSCAN) and K-Means Clustering Algorithm (K-means) are applied to group these obstacle grid cells. The clustering algorithm groups adjacent or close grid cells into a group based on their spatial location. Through the clustering results, different obstacle groups can be identified, with each group representing an independent obstacle.

[0098] Thus, by combining the results of the two detection methods in the grid map, the existence of obstacles in each grid can be determined more accurately, thereby reducing false alarms and false negatives. The obstacle grid set can be processed quickly by the clustering algorithm to generate detection boxes, and the information of the vehicle's surrounding environment can be updated in real time.

[0099] In some embodiments, determining the presence of millimeter waves in each grid cell based on the millimeter wave detection frame on the grid image includes: determining whether any grid cell overlaps with the millimeter wave detection frame based on the millimeter wave detection frame on the grid image; when the corresponding grid cell overlaps with the millimeter wave detection frame, setting the millimeter wave presence of the corresponding grid cell to 1, otherwise setting it to 0.

[0100] In this embodiment of the disclosure, the detection frame provided by the millimeter-wave radar is a rectangular area, representing the spatial range in which the radar detects a possible object; the grid map is a gridded representation of the environment, with each grid representing a fixed spatial area. Calculations are performed to determine whether each grid overlaps with the millimeter-wave detection frame. If any part of a grid coincides with the detection frame, the grid is considered to potentially contain an obstacle.

[0101] In this embodiment of the disclosure, if any part of a grid coincides with the detection frame, the millimeter wave presence of the corresponding grid is set to 1, indicating that the grid may contain an obstacle; otherwise, it indicates that the grid does not contain an obstacle.

[0102] Thus, by determining the overlap between the grid and the millimeter-wave detection frame, the location of obstacles can be identified more accurately, reducing false alarms and missed alarms. Large amounts of data can be processed quickly using simple 0 or 1 markings, improving real-time performance and computational efficiency.

[0103] In some embodiments, determining the existence of obstacles in each grid cell based on the set of retained feature points on the grid image includes: determining whether any retained feature point exists in any grid cell based on the set of retained feature points on the grid image; determining the velocity information of any retained feature point in the corresponding grid cell based on multimodal fusion features; determining the average velocity information of all retained feature points in the corresponding grid cell based on the velocity information; determining the average velocity information of all retained feature points in the corresponding grid cell based on the velocity information; and setting the obstacle existence of the corresponding grid cell to 1 if a retained feature point exists in the corresponding grid cell and the average velocity information is greater than a preset average velocity threshold, otherwise setting it to 0.

[0104] In this embodiment of the disclosure, each grid cell is checked to see if it contains any reserved feature points. If a grid cell contains feature points, it indicates that the grid cell may contain obstacles or important environmental information.

[0105] In this embodiment of the disclosure, the multimodal fusion feature includes velocity information. Based on the velocity information in the multimodal fusion feature, the velocity is calculated for each feature point within the grid, and then the average velocity of the entire grid is calculated. The average velocity is compared with a preset threshold. If a feature point exists within a grid and its average velocity exceeds the threshold, the grid is marked as potentially containing an obstacle; otherwise, it is marked as not containing an obstacle.

[0106] In this embodiment of the disclosure, if a feature point exists in a grid and its average speed exceeds a threshold, the obstacle presence of the corresponding grid is set to 1, indicating that the grid may contain an obstacle; otherwise, it indicates that the grid does not contain an obstacle.

[0107] In this way, by quickly calculating the speed information and average speed of feature points, timely responses can be made to adapt to complex traffic environments. By setting speed thresholds, stationary objects and moving obstacles can be effectively distinguished, reducing the false alarm rate.

[0108] In some embodiments, based on the presence of millimeter waves and the presence of obstacles, all grids in the grid diagram are traversed to obtain a set of obstacle grids, including: traversing all grids in the grid diagram, extracting grids with a millimeter wave presence of 1 and an obstacle presence of 1, and generating a set of obstacle grids.

[0109] In this embodiment of the disclosure, each grid in the entire grid diagram is examined. During the traversal, grids that satisfy the condition that the millimeter wave existence is 1 and the obstacle existence is 1 are searched. All grids that satisfy the above conditions are collected to form a new set, called the obstacle grid set.

[0110] In this embodiment of the disclosure, a millimeter wave presence of 1 and an obstacle presence of 1 means that at the grid position, not only has the presence of a movable object been detected, but it has also been confirmed that the object is a moving obstacle.

[0111] In this way, dual verification improves the accuracy and reliability of obstacle detection, thereby enhancing the safety and decision-making capabilities of autonomous driving.

[0112] Figure 3 This illustrates the process of obtaining obstacle detection boxes based on multimodal fusion features and millimeter-wave point clouds, such as... Figure 3 As shown, the process first iterates through the feature points, filtering out road surface points, fence points, curb points, and vegetation points based on the semantic segmentation results of the laser point cloud. Combining this with velocity estimation results, only movable object categories are retained. Next, the point is projected onto the BEV grid map, and the average velocity, height, and number of points in each grid are calculated. Then, the millimeter-wave point cloud is preprocessed, removing static points and retaining only dynamic points. After clustering the dynamic points to generate detection boxes, these boxes are projected onto the BEV grid map, and the millimeter-wave presence in each grid is calculated. Finally, based on the grid map, all obstacle detection boxes for the current frame are output. Each grid is traversed; if the current grid has 0 points and no millimeter-wave points, it is set to an empty state. If the current grid has more than 0 points but the average velocity is less than 0.3 and no millimeter-wave points, it is set to static. If neither of these conditions is met, it is set to dynamic. Finally, all dynamic grids in the grid map are clustered, and obstacle detection boxes are generated based on the clustering information.

[0113] It should be understood that Figures 2 to 3 The schematic diagrams shown are merely illustrative and not limiting, and are scalable; those skilled in the art can use them as a basis. Figures 2 to 3 Even with various obvious changes and / or substitutions to the examples, the resulting technical solutions still fall within the scope of this disclosure.

[0114] This disclosure provides a multimodal obstacle detection device, such as... Figure 4 As shown, the device may include: a laser feature extraction module 401, used to extract point features and laser point representation features corresponding to any feature point in the acquired laser point cloud; the laser point cloud includes the current frame laser point cloud; an image feature extraction module 402, used to project the current frame laser point cloud onto a visual image, and combine the image features extracted from the visual image to determine the image point representation features corresponding to the feature points; a multimodal feature extraction module 403, used to obtain multimodal fusion features based on point features, laser point representation features, and image point representation features; a grid projection module 404, used to generate a set of retained feature points and a millimeter-wave detection box based on the multimodal fusion features and the millimeter-wave point cloud, and project the set of retained feature points and the millimeter-wave detection box onto a grid image; and an obstacle detection module 405, used to generate an obstacle detection box based on the set of retained feature points and the millimeter-wave detection box on the grid image.

[0115] In some embodiments, the laser point cloud further includes historical frame laser point clouds; the laser feature extraction module 401 includes: a first laser feature extraction submodule, used to extract point features corresponding to feature points after expanding the dimensions of the current frame laser point cloud; and a second laser feature extraction submodule, used to project the current frame laser point cloud and the historical frame laser point cloud onto a bird's-eye view planar image to extract laser point representation features.

[0116] In some embodiments, the first laser feature extraction submodule is configured to: expand the dimension of the laser point cloud in the current frame using a neural network to obtain enhanced point cloud data; and extract point features corresponding to the feature points based on the enhanced point cloud data.

[0117] In some embodiments, the second laser feature extraction submodule is configured to: project the current frame laser point cloud and the historical frame laser point cloud onto a bird's-eye view planar image to obtain the current frame bird's-eye view projection and the historical frame bird's-eye view projection; determine the current frame bird's-eye view handcrafted features and the historical frame bird's-eye view handcrafted features based on the current frame bird's-eye view projection and the historical frame bird's-eye view handcrafted features, respectively; perform feature interaction and feature fusion on the current frame bird's-eye view handcrafted features and the historical frame bird's-eye view handcrafted features to obtain bird's-eye view features; and perform interpolation processing on the bird's-eye view features to extract laser point representation features.

[0118] In some embodiments, the image feature extraction module 402 includes: an image feature extraction submodule, used to extract image features based on a visual image; a point cloud fusion submodule, used to project the laser point cloud of the current frame onto the visual image to obtain a point cloud fusion image; and an image feature determination submodule, used to determine the obtained image point representation features based on the point cloud fusion image and the image features.

[0119] In some embodiments, the grid projection module 404 includes: a point projection submodule, configured to filter feature points based on multimodal fusion features, generate a set of retained feature points based on a first filtering result, and project the set of retained feature points onto a grid image; and a frame projection submodule, configured to filter millimeter-wave point clouds, generate millimeter-wave detection frames based on a second filtering result, and project the millimeter-wave detection frames onto a grid image.

[0120] In some embodiments, the point projection submodule is configured to: determine the semantic category of any feature point based on the multimodal fusion features; filter based on the semantic category to extract all feature points that conform to the preset semantic category and generate a set of retained feature points; and project all retained feature points in the set of retained feature points onto the raster image.

[0121] In some embodiments, the frame projection submodule is used to: filter target points in the millimeter-wave point cloud, extract all target points that conform to the preset dynamic and static categories to obtain a dynamic target point set; cluster all dynamic target points in the dynamic target point set to obtain a clustering result; generate a millimeter-wave detection box based on the clustering result, and project the millimeter-wave detection box onto the raster image.

[0122] In some embodiments, the obstacle detection module 405 includes: a millimeter-wave presence determination submodule, configured to determine the presence of millimeter waves in each grid cell based on millimeter-wave detection boxes on the grid image; an obstacle presence determination submodule, configured to determine the presence of obstacles in each grid cell based on a set of preserved feature points on the grid image; a grid traversal submodule, configured to traverse all grid cells in the grid image based on the presence of millimeter waves and the presence of obstacles to obtain an obstacle grid cell set; and a detection box generation submodule, configured to cluster the obstacle grid cell set and generate obstacle detection boxes based on the clustering results.

[0123] In some embodiments, the millimeter wave presence determination submodule is used to: determine whether any grid and the millimeter wave detection frame have an overlapping area based on the millimeter wave detection frame on the grid map; when the corresponding grid and the millimeter wave detection frame have an overlapping area, set the millimeter wave presence of the corresponding grid to 1, otherwise set it to 0.

[0124] In some embodiments, the obstacle presence determination submodule is configured to: determine whether a reserved feature point exists in any grid cell based on the set of reserved feature points on the grid image; determine the velocity information of any reserved feature point in the corresponding grid cell based on the multimodal fusion features; determine the average velocity information of all reserved feature points in the corresponding grid cell based on the velocity information; and set the obstacle presence of the corresponding grid cell to 1 if a reserved feature point exists in the corresponding grid cell and the average velocity information is greater than a preset average velocity threshold, otherwise set it to 0.

[0125] In some embodiments, the grid traversal submodule is used to traverse all grids in the grid diagram, extract the grids with millimeter wave presence of 1 and obstacle presence of 1, and generate an obstacle grid set.

[0126] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0127] The multimodal obstacle detection device of this disclosure can detect various types of dynamic obstacles and improve the detection accuracy of different types of dynamic obstacles.

[0128] This disclosure provides a scenario illustration for multimodal obstacle detection, such as... Figure 5 As shown.

[0129] As previously described, the multimodal obstacle detection method provided in this disclosure is applied to electronic devices. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. This electronic device can be installed in or connected to an autonomous vehicle.

[0130] Specifically, the electronic device may perform the following operations:

[0131] Based on the acquired laser point cloud, extract the point features and laser point representation features corresponding to any feature point in the laser point cloud;

[0132] Project the laser point cloud of the current frame onto the visual image, and combine it with the image features extracted from the visual image to determine the image point representation features corresponding to the feature points;

[0133] Based on point features, laser point representation features, and image point representation features, multimodal fusion features are obtained;

[0134] Based on multimodal fusion features and millimeter-wave point clouds, a set of preserved feature points and millimeter-wave detection boxes are generated, and the set of preserved feature points and millimeter-wave detection boxes are projected onto the raster image;

[0135] Obstacle detection boxes are generated based on the set of preserved feature points on the raster image and the millimeter-wave detection boxes.

[0136] It should be understood that Figure 5 The scene diagrams shown are merely illustrative and not restrictive; those skilled in the art can interpret them based on... Figure 5 Even with various obvious changes and / or substitutions to the examples, the resulting technical solutions still fall within the scope of this disclosure.

[0137] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0138] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, a computer program product, and an autonomous vehicle.

[0139] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0140] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0141] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0142] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as multimodal obstacle detection methods. For example, in some embodiments, the multimodal obstacle detection method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the multimodal obstacle detection method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform a multimodal obstacle detection method by any other suitable means (e.g., by means of firmware).

[0143] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0144] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0145] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0146] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0147] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0148] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0149] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0150] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A multimodal obstacle detection method, comprising: Based on the acquired laser point cloud, point features and laser point representation features corresponding to any feature point in the laser point cloud are extracted; the laser point cloud includes the laser point cloud of the current frame; point features refer to features extracted from the laser point cloud data to describe the local or global attributes of feature points, which are basic attributes directly extracted from each feature point in the laser point cloud data; laser point representation features refer to feature vectors or descriptors used to represent and encode feature points in the laser point cloud data, which are generated by extracting features from the laser point cloud data and then performing point representation transformation on the features; The current frame laser point cloud is projected onto a visual image, and the image features extracted from the visual image are combined to determine the image point representation features corresponding to the feature points; Based on the point features, the laser point representation features, and the image point representation features, a multimodal fusion feature is obtained; Based on the multimodal fusion features, the semantic category of any feature point is determined, and all feature points that conform to the preset semantic category are extracted according to the semantic category. A set of retained feature points is generated, and the set of retained feature points is projected onto the raster image. A millimeter-wave detection bounding box is generated based on the millimeter-wave point cloud, and the millimeter-wave detection bounding box is projected onto the grid image; An obstacle detection box is generated based on the set of preserved feature points on the grid map and the millimeter-wave detection box.

2. The method according to claim 1, wherein, The laser point cloud also includes historical frame laser point clouds; The step of extracting point features and laser point representation features corresponding to any feature point in the acquired laser point cloud includes: After expanding the dimensions of the current frame laser point cloud, the point features corresponding to the feature points are extracted; The laser point cloud of the current frame and the laser point cloud of the historical frame are projected onto a bird's-eye view image, and the laser point representation features are extracted.

3. The method according to claim 2, wherein, After expanding the dimensionality of the current frame laser point cloud, extracting the point features corresponding to the feature points includes: Based on the current frame laser point cloud, the dimensionality is expanded using a neural network to obtain enhanced point cloud data; Based on the enhanced point cloud data, extract the point features corresponding to the feature points.

4. The method according to claim 2, wherein, The step of projecting the current frame laser point cloud and the historical frame laser point cloud onto a bird's-eye view image and extracting the laser point representation features includes: Project the current frame laser point cloud and the historical frame laser point cloud onto the bird's-eye view plane image to obtain the current frame bird's-eye view projection and the historical frame bird's-eye view projection. Based on the current frame bird's-eye view projection and the historical frame bird's-eye view projection, the current frame bird's-eye view handcrafted features and the historical frame bird's-eye view handcrafted features are determined respectively; wherein, handcrafted features refer to features designed and extracted by human experts based on domain knowledge and experience in data processing; The bird's-eye view handcrafted features of the current frame and the bird's-eye view handcrafted features of the historical frames are subjected to feature interaction and feature fusion to obtain the bird's-eye view features; The aerial view features are interpolated to extract the laser point representation features.

5. The method according to claim 1, wherein, The step of projecting the current frame laser point cloud onto a visual image and combining it with image features extracted from the visual image to determine the image point representation features corresponding to the feature points includes: Extract the image features based on the visual image; The current frame laser point cloud is projected onto the visual image to obtain a point cloud fusion image; Based on the point cloud fused image and the image features, the image point representation features are determined.

6. The method according to claim 1, wherein, The process of generating millimeter-wave detection boxes based on millimeter-wave point clouds and projecting the millimeter-wave detection boxes onto the grid image includes: The millimeter-wave point cloud is filtered, and the millimeter-wave detection box is generated based on the second filtering result. The millimeter-wave detection box is then projected onto the grid image.

7. The method according to claim 1, wherein, The step of projecting the preserved feature point set onto the raster image includes: Project all retained feature points in the set of retained feature points onto the raster image.

8. The method according to claim 6, wherein, The step of filtering the millimeter-wave point cloud, generating the millimeter-wave detection box based on the second filtering result, and projecting the millimeter-wave detection box onto the grid image includes: The target points in the millimeter-wave point cloud are filtered, and all target points that meet the preset dynamic and static categories are extracted to obtain a dynamic target point set; Cluster all dynamic target points in the set of dynamic target points to obtain the clustering results; Based on the clustering results, the millimeter-wave detection box is generated and projected onto the grid image.

9. The method according to claim 1, wherein, The step of generating an obstacle detection box based on the set of preserved feature points on the grid image and the millimeter-wave detection box includes: The presence of millimeter waves in each grid is determined based on the millimeter wave detection frame on the grid map. The presence of obstacles within each grid cell is determined based on the set of preserved feature points on the grid map. Based on the existence of the millimeter wave and the existence of the obstacle, all grids in the grid diagram are traversed to obtain the obstacle grid set; Cluster the obstacle grid set and generate the obstacle detection box based on the clustering results.

10. The method according to claim 9, wherein, The step of determining the presence of millimeter waves in each grid cell based on the millimeter wave detection frame on the grid image includes: Based on the millimeter-wave detection frame on the grid map, determine whether any grid cell overlaps with the millimeter-wave detection frame. When there is an overlapping area between the corresponding grid and the millimeter-wave detection frame, the millimeter-wave presence of the corresponding grid is set to 1; otherwise, it is set to 0.

11. The method according to claim 9, wherein, Determining the existence of obstacles within each grid cell based on the set of preserved feature points on the grid map includes: Based on the set of retained feature points on the raster image, determine whether there are any retained feature points in any of the raster cells; Based on the multimodal fusion features, determine the velocity information of any of the retained feature points within the grid. Based on the velocity information, determine the average velocity information of all the retained feature points within the corresponding grid. If the reserved feature point exists in the corresponding grid and the average speed information is greater than the preset average speed threshold, the obstacle existence of the corresponding grid is set to 1; otherwise, it is set to 0.

12. The method according to any one of claims 10 or 11, wherein, Based on the existence of the millimeter wave and the existence of the obstacle, the process iterates through all grids in the grid map to obtain the obstacle grid set, including: Traverse all grids in the grid diagram, extract the grids where the millimeter wave presence is 1 and the obstacle presence is 1, and generate the obstacle grid set.

13. A multimodal obstacle detection device, comprising: Laser feature extraction module; This is used to extract point features and laser point representation features corresponding to any feature point in the acquired laser point cloud; the laser point cloud includes the laser point cloud of the current frame; point features refer to features extracted from the laser point cloud data to describe the local or global attributes of feature points, which are basic attributes directly extracted from each feature point in the laser point cloud data; laser point representation features refer to feature vectors or descriptors used to represent and encode feature points in the laser point cloud data, which are generated by extracting features from the laser point cloud data and then performing point representation transformation on the features; The image feature extraction module is used to project the current frame laser point cloud onto a visual image, and combine the image features extracted from the visual image to determine the image point representation features corresponding to the feature points; A multimodal feature extraction module is used to obtain multimodal fusion features based on the point features, the laser point representation features, and the image point representation features; The grid projection module is used to determine the semantic category of any feature point according to the multimodal fusion features, filter according to the semantic category, extract all feature points that conform to the preset semantic category, generate a set of retained feature points, and project the set of retained feature points onto the grid image. A millimeter-wave detection bounding box is generated based on the millimeter-wave point cloud, and the millimeter-wave detection bounding box is projected onto the grid image; An obstacle detection module is used to generate an obstacle detection box based on the set of preserved feature points on the grid map and the millimeter-wave detection box.

14. The apparatus according to claim 13, wherein, The laser point cloud also includes historical frame laser point clouds; The image feature extraction module includes: The first laser feature extraction submodule is used to extract the point features corresponding to the feature points after expanding the dimensions of the current frame laser point cloud; The second laser feature extraction submodule is used to project the current frame laser point cloud and the historical frame laser point cloud onto a bird's-eye view planar image and extract the laser point representation features.

15. The apparatus according to claim 14, wherein, The first laser feature extraction submodule is used for: Based on the current frame laser point cloud, the dimensionality is expanded using a neural network to obtain enhanced point cloud data; Based on the enhanced point cloud data, extract the point features corresponding to the feature points.

16. The apparatus according to claim 14, wherein, The second laser feature extraction submodule is used for: Project the current frame laser point cloud and the historical frame laser point cloud onto the bird's-eye view plane image to obtain the current frame bird's-eye view projection and the historical frame bird's-eye view projection. Based on the current frame bird's-eye view projection and the historical frame bird's-eye view projection, the current frame bird's-eye view handcrafted features and the historical frame bird's-eye view handcrafted features are determined respectively; wherein, handcrafted features refer to features designed and extracted by human experts based on domain knowledge and experience in data processing; The bird's-eye view handcrafted features of the current frame and the bird's-eye view handcrafted features of the historical frames are subjected to feature interaction and feature fusion to obtain the bird's-eye view features; The aerial view features are interpolated to extract the laser point representation features.

17. The apparatus according to claim 13, wherein, The image feature extraction module includes: An image feature extraction submodule is used to extract the image features based on the visual image; The point cloud fusion submodule is used to project the current frame laser point cloud onto the visual image to obtain a point cloud fusion image; The image feature determination submodule is used to determine the image point representation features based on the point cloud fused image and the image features.

18. The apparatus according to claim 13, wherein, The grid projection module includes: The frame projection submodule is used to filter the millimeter-wave point cloud, generate the millimeter-wave detection frame according to the second filtering result, and project the millimeter-wave detection frame onto the grid image.

19. The apparatus according to claim 13, wherein, The grid projection module is also used for: Project all retained feature points in the set of retained feature points onto the raster image.

20. The apparatus according to claim 18, wherein, The frame projection submodule is used for: The target points in the millimeter-wave point cloud are filtered, and all target points that meet the preset dynamic and static categories are extracted to obtain a dynamic target point set; Cluster all dynamic target points in the set of dynamic target points to obtain the clustering results; Based on the clustering results, the millimeter-wave detection box is generated and projected onto the grid image.

21. The apparatus according to claim 13, wherein, The obstacle detection module includes: The millimeter wave presence determination submodule is used to determine the presence of millimeter waves in each grid cell based on the millimeter wave detection frame on the grid diagram. The obstacle presence determination submodule is used to determine the presence of obstacles in each grid cell based on the set of preserved feature points on the grid map. The grid traversal submodule is used to traverse all grids in the grid diagram based on the existence of the millimeter wave and the existence of the obstacle to obtain the obstacle grid set. The detection box generation submodule is used to cluster the obstacle grid set and generate the obstacle detection box based on the clustering results.

22. The apparatus according to claim 21, wherein, The millimeter wave presence determination submodule is used for: Based on the millimeter-wave detection frame on the grid map, determine whether any grid cell overlaps with the millimeter-wave detection frame. When there is an overlapping area between the corresponding grid and the millimeter-wave detection frame, the millimeter-wave presence of the corresponding grid is set to 1; otherwise, it is set to 0.

23. The apparatus according to claim 21, wherein, The obstacle presence determination submodule is used for: Based on the set of retained feature points on the raster image, determine whether there are any retained feature points in any of the raster cells; Based on the multimodal fusion features, determine the velocity information of any of the retained feature points within the grid. Based on the velocity information, determine the average velocity information of all the retained feature points within the corresponding grid. If the reserved feature point exists in the corresponding grid and the average speed information is greater than the preset average speed threshold, the obstacle existence of the corresponding grid is set to 1; otherwise, it is set to 0.

24. The apparatus according to claim 22 or 23, wherein, The grid traversal submodule is used to traverse all grids in the grid diagram, extract the grids where the millimeter wave existence is 1 and the obstacle existence is 1, and generate the obstacle grid set.

25. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.

26. A non-transitory computer-readable storage medium storing computer instructions, wherein, Computer instructions are used to cause the determinant to perform the method according to any one of claims 1-12.

27. A computer program product comprising a computer program stored on a storage medium, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-12.

28. An autonomous vehicle, including the electronic equipment as claimed in claim 25.

Citation Information

Patent Citations

  • Efficient obstacle detection method based on multi-feature fusion under grid coordinate system

    CN117037117A

  • Perception fusion system, electronic device and storage medium

    WO2024234659A1