Obstacle detection method and apparatus, electronic device, and storage medium
By combining semantic segmentation and projection processing of point cloud data and image data in unmanned vehicles, the problem of difficult detection of small target obstacles in open-pit mine environments is solved, and higher obstacle detection accuracy and driving safety are achieved.
Patent Information
- Application Number
- PCT/CN2025/072814
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-24
AI Technical Summary
It is difficult for autonomous vehicles to effectively perceive small target obstacles in open-pit mine environments, resulting in low driving safety and operational efficiency.
By acquiring point cloud data and image data of data acquisition equipment and image acquisition equipment, semantic segmentation and projection processing are performed, geometric parameters of obstacles are calculated and filtered to improve obstacle detection accuracy.
It improves the accuracy of detection of small target obstacles, improves the vehicle's perception of small target obstacles, and ensures driving safety and operation efficiency.
Smart Images

Figure CN2025072814_24072025_PF_FP_ABST
Abstract
Description
Obstacle detection method, device, electronic device and storage medium Technical Field
[0001] The present disclosure relates to the field of autonomous driving technology, and in particular to an obstacle detection method, device, electronic device, and storage medium. Background Art
[0002] With the development of autonomous driving technology, the application of autonomous vehicles in mining trucks has emerged. Currently, operations such as mining, transportation, and soil dumping are primarily performed by autonomous vehicles. Environmental perception technology enables autonomous vehicles to obtain real-time information about their surroundings, assisting with subsequent control decisions, obstacle avoidance, and path planning. Obstacle perception is a key capability in autonomous driving and is crucial for ensuring safe operation.
[0003] Among related technologies, the more mainstream and widely used solutions are pure vision solutions integrating multiple cameras and pure lidar solutions integrating multiple lidars. However, in open-pit mining scenarios, due to the special characteristics of open-pit mines, such as unstructured roads, unclear textures, rapidly changing terrain, strong winds and dust, and high risk of falling rocks, pure vision solutions struggle to effectively perceive this complex operating environment. Pure lidar solutions, while highly adaptable to environmental conditions, struggle to effectively perceive small obstacles such as ruts, small stones, and potholes on the roadway. This makes it impossible for autonomous vehicles to effectively avoid small obstacles, and thus, their driving safety and efficiency cannot be guaranteed. Summary of the Invention
[0004] In view of this, the embodiments of the present disclosure provide an obstacle detection method, device, electronic device and storage medium to solve the problem in the related art that vehicles are unable to effectively avoid small target obstacles, and thus cannot ensure the driving safety and operating efficiency of unmanned vehicles.
[0005] In a first aspect of an embodiment of the present disclosure, a method for obstacle detection is provided for use in a vehicle. The method comprises: acquiring first point cloud data acquired by a data acquisition device and first image data acquired by an image acquisition device, wherein the first point cloud data and the first image data have a temporal correspondence; performing semantic segmentation on the first image data to obtain a semantic segmentation result, wherein the semantic segmentation result is used to characterize whether the first image data includes semantic information of a target obstacle; if the first image data includes the target obstacle, projecting the target obstacle into the first point cloud data to obtain target point cloud data corresponding to the target obstacle; calculating geometric parameters of the target obstacle based on the target point cloud data, and filtering the first point cloud data based on the geometric parameters to obtain a detection result of the target obstacle.
[0006] According to a second aspect of the embodiments of the present disclosure, an obstacle detection device is provided for use in a vehicle. The device includes: an acquisition module configured to acquire first point cloud data acquired by a data acquisition device and first image data acquired by an image acquisition device, wherein the first point cloud data and the first image data have a temporal correspondence; a segmentation module configured to perform semantic segmentation on the first image data to obtain a semantic segmentation result, wherein the semantic segmentation result is used to indicate whether the first image data includes semantic information of a target obstacle; a conversion module configured to, if the first image data includes the target obstacle, project the target obstacle into the first point cloud data to obtain target point cloud data corresponding to the target obstacle; and a calculation module configured to calculate geometric parameters of the target obstacle based on the target point cloud data, and perform filtering processing on the first point cloud data based on the geometric parameters to obtain a detection result of the target obstacle.
[0007] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising at least one processor; and a memory for storing instructions executable by at least one processor; wherein the at least one processor is configured to execute instructions to implement the steps of the above method.
[0008] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the steps of the above method.
[0009] At least one of the above-mentioned technical solutions adopted in the embodiments of the present disclosure can achieve the following beneficial effects: first point cloud data collected by a data acquisition device and first image data collected by an image acquisition device are acquired, wherein the first point cloud data and the first image data have a temporal correspondence; semantic segmentation is performed on the first image data to obtain a semantic segmentation result, wherein the semantic segmentation result is used to indicate whether the first image data includes semantic information of a target obstacle; if the first image data includes a target obstacle, the target obstacle is projected onto the first point cloud data to obtain target point cloud data corresponding to the target obstacle; geometric parameters of the target obstacle are calculated based on the target point cloud data, and the first point cloud data is filtered based on the geometric parameters to obtain a target obstacle detection result. The target obstacle identified based on the image data can be projected onto the point cloud data, and the target obstacle detection result can be obtained by calculating the geometric parameters of the target obstacle. Therefore, the accuracy of small target obstacle detection is improved, the problem of small target obstacles being easily missed is solved, the vehicle's perception of small target obstacles is enhanced, and the driving safety and operating efficiency of the vehicle are further ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0011] FIG1 is a flow chart of an obstacle detection method provided by an exemplary embodiment of the present disclosure.
[0012] FIG2 is a flow chart of another obstacle detection method provided by an exemplary embodiment of the present disclosure.
[0013] FIG3 is a flow chart of yet another obstacle detection method provided by an exemplary embodiment of the present disclosure.
[0014] FIG4 is a schematic structural diagram of an obstacle detection device provided by an exemplary embodiment of the present disclosure.
[0015] FIG5 is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure.
[0016] FIG6 is a schematic diagram of the structure of a computer system provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0017] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0018] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0019] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0022] An obstacle detection method and device according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0023] FIG1 is a flow chart of an obstacle detection method provided by an exemplary embodiment of the present disclosure. The obstacle detection method of FIG1 can be executed by a server or electronic device in an autonomous driving system. As shown in FIG1 , the obstacle detection method includes:
[0024] S101, acquiring first point cloud data acquired by a data acquisition device and first image data acquired by an image acquisition device, wherein the first point cloud data and the first image data have a corresponding relationship in time;
[0025] S102, performing semantic segmentation on the first image data to obtain a semantic segmentation result, wherein the semantic segmentation result is used to indicate whether the first image data includes semantic information of a target obstacle;
[0026] S103: When the first image data includes a target obstacle, project the target obstacle onto the first point cloud data to obtain target point cloud data corresponding to the target obstacle;
[0027] S104 : Calculate geometric parameters of the target obstacle based on the target point cloud data, and perform filtering processing on the first point cloud data based on the geometric parameters to obtain a detection result of the target obstacle.
[0028] Specifically, taking the server in an autonomous driving system as an example, while the vehicle is driving, the server collects first point cloud data and first image data around the vehicle through the data acquisition device and image acquisition device installed on the vehicle, and performs semantic segmentation on the collected first image data to obtain a semantic segmentation result. Next, the server determines whether the first image data includes the target obstacle based on the semantic segmentation result. If the target obstacle is included in the first image data, the server converts the first image data corresponding to the target obstacle from the image acquisition device coordinate system to the data acquisition device coordinate system, that is, projects the target obstacle into the first point cloud data to obtain target point cloud data corresponding to the target obstacle. Furthermore, based on the target point cloud data, the server calculates the geometric parameters of the target obstacle and uses a filtering algorithm to perform point cloud filtering on the first point cloud data based on the geometric parameters to obtain a detection result of the target obstacle.
[0029] In the disclosed embodiments, an autonomous driving system refers to a system composed of hardware and software that can continuously perform part or all of a dynamic driving task (DDT). A dynamic driving task refers to the perception, decision-making, and execution required to complete vehicle driving, that is, all real-time operational and tactical functions when driving a road vehicle, but does not include planning functions, such as trip planning, destination and route selection, etc. For example, a dynamic driving task may include, but is not limited to, controlling the lateral movement of the vehicle, controlling the longitudinal movement of the vehicle, monitoring the driving environment and preparing responses by detecting, identifying, and classifying targets and events, executing responses, driving decisions, controlling vehicle lighting and signal devices, etc.
[0030] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, as well as big data and artificial intelligence platforms. The embodiments of the present disclosure do not limit this.
[0031] The vehicle can be an unmanned vehicle or an autonomous vehicle. An unmanned vehicle is an intelligent vehicle that is unmanned through a computer system. It uses an onboard sensor system to perceive the vehicle's surrounding environment and, based on the perceived information about the road, vehicle position, and obstacles, controls the vehicle's steering and speed, thereby enabling the vehicle to travel safely and reliably on the road. In the embodiments of the present disclosure, unmanned vehicles may include, but are not limited to, wide-body vehicles, large mining trucks, collection vehicles, and forklifts equipped with data acquisition and image acquisition equipment.
[0032] The data acquisition device can collect data from various points on the surrounding target objects by emitting and receiving signals, that is, collecting point data of the target points. The collection of point data collected by the data acquisition device within a sampling period is a frame of point cloud data. The data acquisition device may include but is not limited to laser radar, millimeter wave radar, ultrasonic radar, etc. The number of data acquisition devices can be one or more. In the embodiment of the present disclosure, the data acquisition device is a laser radar. A laser radar is a radar system that emits a laser beam to detect characteristic quantities such as the position and speed of a target. The measurement range of a laser radar is between 10 centimeters and 200 meters.
[0033] An image acquisition device can obtain an image of a target through physical means. An image acquisition device can directly capture video and process the video frames; it can also obtain a target image by shooting and process the target image. Image acquisition devices may include, but are not limited to, cameras, webcams, video cameras, scanners, video capture cards, and other devices with camera functions (e.g., mobile phones, tablets, etc.). The number of image acquisition devices may be one or more. In the embodiments of the present disclosure, the image acquisition device is a camera.
[0034] It should be noted that the type, quantity and installation location of the data acquisition device and the image acquisition device can be set according to actual needs, and the embodiments of the present disclosure do not limit this.
[0035] Point cloud refers to a collection of points that express the spatial distribution and surface characteristics of a target object in a certain spatial reference system. Point cloud data refers to a collection of vectors in a three-dimensional coordinate system. These vectors are usually expressed in the form of three-dimensional coordinates (X, Y, Z), and are generally used to represent the outer surface shape of an object. In addition to the geometric position information represented by the three-dimensional coordinates, point cloud data can also represent the color (R, G, B) information of a point or the intensity (Intensity) information of the object's reflective surface. In the embodiment of the present disclosure, point cloud data is used to characterize the three-dimensional coordinates of each point in the point cloud in the spatial reference coordinate system. In the embodiment of the present disclosure, the spatial reference coordinate system can be the laser radar coordinate system corresponding to the laser radar; the first point cloud data is a multi-frame point cloud data collected by the laser radar. In actual applications, the laser radar can be triggered to scan the driving area according to a preset scanning frequency (for example, 10Hz) to obtain a frame of point cloud data corresponding to each scanning frequency.
[0036] Image data refers to dynamic image data and / or static image data. In the embodiments of the present disclosure, dynamic image data may be a dynamic image comprising multiple frames of images, such as a video; static image data may be a static image comprising a single frame of images, such as a picture. The image data may be image data captured in real time by an image acquisition device, or may be image data pre-stored in a server, and the embodiments of the present disclosure are not limited thereto. Preferably, in the embodiments of the present disclosure, the first image data is multiple frames of image data captured by a camera.
[0037] There is a temporal correspondence between the first point cloud data and the first image data. Exemplarily, the timestamp of the first point cloud data may be the same as the timestamp of the first image data, or the difference between the timestamp of the first point cloud data and the timestamp of the first image data may be less than or equal to a preset time threshold. In the embodiment of the present disclosure, the preset time threshold may be a time threshold pre-set based on empirical data, or a time threshold obtained by adjusting the set time threshold based on actual needs, and the embodiment of the present disclosure is not limited to this. For example, the preset time threshold may be any value between 0.1 seconds and 1 second. Preferably, in the embodiment of the present disclosure, the preset time threshold is 0.3 or 0.4 seconds. By acquiring the first point cloud data and the first image data having the same timestamp or the difference between the timestamps being less than or equal to the preset time threshold, it is possible to ensure as much as possible that the acquired first point cloud data and the first image data belong to the same location and the same object.
[0038] Semantic segmentation refers to the process of segmenting an image based on its semantic meaning. In the field of images, semantics refers to the content of an image. Semantic segmentation is a pixel-level processing process, that is, the process of classifying all pixels in an image at the pixel level. In the disclosed embodiments, semantic segmentation refers to segmenting areas in a photo that belong to a large category and outputting their category information. The segmentation result can be a target obstacle area composed of pixels belonging to the target obstacle in the current frame.
[0039] Target obstacles are objects on the road or in the driving area that may affect or hinder the normal operation of the vehicle. Target obstacles can be divided into static obstacles and dynamic obstacles based on their state of motion. Static obstacles are generally stationary obstacles that do not move, such as piles of earth and rocks. Dynamic obstacles are movable obstacles that can move, such as pedestrians and vehicles. Target obstacles can be divided into small target obstacles, regular target obstacles, large target obstacles, and extra-large target obstacles based on their size. In the disclosed embodiments, small target obstacles may include but are not limited to wheel tracks, small rocks, potholes, road cracks, etc. Regular target obstacles may include but are not limited to streetlights, large rocks, signboards, reflective poles, etc. Large target obstacles may include but are not limited to piles of earth, materials, pedestrians, etc. Extra-large target obstacles may include but are not limited to excavators, bulldozers, wide-body vehicles, mining trucks, etc. In the disclosed embodiments, target obstacles are small target obstacles that are difficult to identify and may affect the safe operation of the vehicle.
[0040] The image acquisition device coordinate system is a polar coordinate system established with the image acquisition device's location as its origin. In the disclosed embodiments, the image acquisition device coordinate system is a camera coordinate system. A camera coordinate system is a three-dimensional coordinate system established with the camera as its center. Typically, the origin of the camera coordinate system is the camera's optical center, the X-axis is parallel to the image frame's X-axis, the Y-axis is parallel to the image frame's Y-axis, and the Z-axis is the camera's optical axis, perpendicular to the plane of the image frame.
[0041] The data acquisition device coordinate system is a polar coordinate system established with the location of the data acquisition device as the origin. In the disclosed embodiments, the data acquisition device coordinate system is a lidar coordinate system. A lidar coordinate system is a three-dimensional coordinate system established with the lidar as the center. Typically, the origin of the lidar coordinate system is the laser emission center of the lidar. When the lidar is set horizontally, the plane containing the X-axis and Y-axis is parallel to the ground, and the Z-axis is perpendicular to the ground.
[0042] The geometric parameters are used to characterize information such as the boundary shape, length, width, height (or depth) and center point of the target obstacle. In the embodiment of the present disclosure, the geometric parameters of the target obstacle can be calculated based on the multi-frame point cloud data collected by the lidar and the multi-frame image data collected by the camera. Specifically, after obtaining the multiple images collected by the camera, the multiple images are semantically segmented one by one to obtain the area of the obstacle in each image and its semantic information; the obstacles identified in each image are projected into the multi-frame point cloud data to obtain the point cloud data corresponding to the obstacle; the point cloud data of all obstacles are clustered based on the semantic information to obtain the target point cloud data of the target obstacle, and the geometric parameters of the target obstacle are calculated based on the target point cloud data.
[0043] Point cloud filtering refers to the process of removing invalid points from point cloud data. Usually, in the process of laser radar acquiring point cloud data, due to the influence of factors such as the product's own system, the surface of the object to be measured and the scanning environment, the point cloud data is inevitably mixed with some noise points and / or outliers. These noise points and / or outliers will have a certain impact on the subsequent point cloud segmentation, point cloud registration and other processing results. Therefore, they need to be directly removed or processed in a smooth manner. In the embodiment of the present disclosure, noise points refer to point cloud data that are useless for subsequent obstacle detection, that is, they cannot provide valuable information for subsequent obstacle detection; noise points may include but are not limited to dust, water mist, dynamic obstacles (for example, vehicles, people, flying insects), etc. Outliers refer to point cloud data that exceeds the measurement range of the laser radar. In the field of unmanned driving, common filtering algorithms may include but are not limited to Kalman filtering algorithms, average filtering algorithms, median filtering algorithms, etc. In the embodiment of the present disclosure, by filtering the first point cloud data based on geometric parameters, not only can the noise points and / or outliers introduced during the acquisition process of the point cloud data be removed, but also the point cloud points that are useless for detecting small target obstacles can be removed to form more accurate point cloud data, thereby improving the quality of the point cloud data.
[0044] The detection result of the target obstacle may include but is not limited to the category information, position (coordinate) information, size information, etc. of the target obstacle.
[0045] According to the technical solution provided by the embodiments of the present disclosure, first point cloud data collected by a data acquisition device and first image data collected by an image acquisition device are acquired, wherein the first point cloud data and the first image data have a temporal correspondence; semantic segmentation is performed on the first image data to obtain a semantic segmentation result, wherein the semantic segmentation result is used to indicate whether the first image data includes semantic information of a target obstacle; if the first image data includes the target obstacle, the target obstacle is projected into the first point cloud data to obtain target point cloud data corresponding to the target obstacle; geometric parameters of the target obstacle are calculated based on the target point cloud data, and the first point cloud data are filtered based on the geometric parameters to obtain a target obstacle detection result. The target obstacle identified based on the image data can be projected into the point cloud data, and the target obstacle detection result can be obtained by calculating the geometric parameters of the target obstacle. Therefore, the accuracy of small target obstacle detection is improved, the problem of small target obstacles being easily missed is solved, the vehicle's perception of small target obstacles is enhanced, and the driving safety and operating efficiency of the vehicle are further guaranteed.
[0046] In some embodiments, semantic segmentation is performed on the first image data to obtain a semantic segmentation result, including: using a target obstacle detection model to perform semantic reasoning on each pixel point in the first image data to segment a pixel area involved in the target obstacle from the first image data, and outputting a target obstacle category corresponding to the pixel area, wherein the target obstacle detection model is obtained by controlling a target obstacle detection model to be trained to perform a semantic segmentation training task on the sample image data based on the sample image data and annotated labels associated with the sample image data, the semantic segmentation training task being used to segment a pixel area for representing the target obstacle from the sample image data, and the annotated labels including the pixel area corresponding to the target obstacle in the sample image data annotated by a multimedia data annotation tool.
[0047] Specifically, a large amount of photo data of open-pit mines is obtained in advance as sample image data, and the obtained sample image data is annotated using a multimedia data annotation tool to obtain annotation labels associated with the sample image data. After obtaining the sample image data and the annotation labels associated with the sample image data, the target obstacle detection model to be trained can be controlled to perform a semantic segmentation training task on the sample image data to obtain the pixel area involved in the target obstacle in the sample image data. Furthermore, a loss calculation is performed between the pixel area involved in the predicted target obstacle and the pixel area involved in the target obstacle indicated by the annotation label, and the calculated loss is used to update the parameters of the target obstacle detection model to be trained. After multiple updates, a target obstacle detection model that meets the requirements is obtained.
[0048] After acquiring the first image data, the server uses a pre-trained target obstacle detection model to perform semantic reasoning on each pixel in the first image data to segment the pixel area involved in the target obstacle from the first image data and output the target obstacle category corresponding to the pixel area. In the disclosed embodiment, the target obstacle detection model is a deep learning model for performing semantic segmentation tasks on image data. The semantic segmentation training task is used to segment pixel areas used to represent target obstacles from sample image data. The annotation labels may include pixel areas corresponding to target obstacles in the sample image data annotated by a multimedia data annotation tool. Target obstacle categories may include, but are not limited to, ruts, small stones, and potholes.
[0049] According to the technical solution provided by the embodiments of the present disclosure, by using a pre-trained target obstacle detection model to perform semantic segmentation on image data, target obstacles in the image data can be quickly and accurately identified, thereby improving the recognition speed and recognition accuracy of target obstacles.
[0050] In some embodiments, the method further includes: acquiring second point cloud data collected by a data acquisition device and / or second image data collected by an image acquisition device; fusing the detection result of the target obstacle with the second point cloud data and / or the second image data to generate a vehicle control strategy, and controlling the vehicle based on the vehicle control strategy.
[0051] Specifically, during the driving of the vehicle, the server collects second point cloud data and / or second image data around the vehicle through a data acquisition device and / or an image acquisition device; after collecting the second point cloud data and / or the second image data, the server fuses the detection result of the target obstacle with the second point cloud data and / or the second image data, and generates a vehicle control strategy based on the fusion result to control the vehicle.
[0052] In the disclosed embodiments, the second point cloud data is multi-frame point cloud data collected by a lidar, and the second image data is multi-frame image data collected by a camera. Vehicle control strategies refer to operations performed to control the vehicle, such as accelerating, decelerating, avoiding obstacles, and maneuvering around them.
[0053] According to the technical solution provided in the embodiments of the present disclosure, by fusing the detection results of the target obstacle with the second point cloud data and / or the second image data, a vehicle control strategy can be generated based on the fusion results to provide an effective auxiliary decision-making basis for vehicle obstacle avoidance, thereby improving the stability of the automatic driving system, ensuring the driving safety of the vehicle, and improving the operating efficiency of the vehicle.
[0054] In some embodiments, the method also includes: sending the detection result of the target obstacle to the cloud server, so that the cloud server updates the obstacle layer data in the high-precision map based on the detection result of the target obstacle; receiving the updated high-precision map sent by the cloud server, so that the operator processes the target obstacle based on the obstacle layer data in the updated high-precision map; and sending the processing result of the target obstacle to the cloud server, so that the cloud server updates the updated high-precision map based on the processing result of the target obstacle.
[0055] Specifically, after obtaining the detection result of the target obstacle, the vehicle reports the detection result of the target obstacle to the cloud server. After receiving the detection result of the target obstacle, the cloud server updates the data of the obstacle layer in the original high-precision map based on the detection result of the target obstacle to obtain an updated high-precision map, and sends the updated high-precision map to the vehicle. After receiving the updated high-precision map, the operator can repair and remove target obstacles such as ruts, small stones, and potholes on the road or the vehicle's driving area based on the obstacle layer data in the updated high-precision map. After the repair and removal process is completed, the processing results are reported to the cloud server. After receiving the processing results sent by the vehicle, the cloud server updates the updated high-precision map to eliminate the obstacle layer data corresponding to the target obstacle in the updated high-precision map, thereby obtaining the target high-precision map.
[0056] According to the technical solution provided by the embodiment of the present disclosure, by sending the detection results of the target obstacle to the cloud server, the cloud server can timely update the obstacle layer data in the high-precision map, and enable the operator to process the target obstacle based on the updated obstacle layer data. Therefore, the operation efficiency is improved, the operation time is saved, the timeliness of the high-precision map is ensured, and the driving safety of the vehicle is further improved.
[0057] In some embodiments, the method further includes: receiving time data sent by a timing system, and performing timing on a data acquisition device and a computing platform of the vehicle respectively based on the time data; after receiving a second pulse signal sent by the timing system, triggering the data acquisition device to collect first point cloud data, and triggering the image acquisition device to collect first image data through the computing platform of the vehicle.
[0058] Specifically, in the field of autonomous driving, data from multiple sensors is required. Each sensor has its own internal clock, and each clock has different drifts. This results in inconsistent timing of the data received by the server from each sensor, leading to inaccurate identification of target obstacles. To ensure the accuracy of target obstacle detection results after multi-source data fusion, it is necessary to synchronize the data acquisition device and the image acquisition device. That is, align the timestamps of the data acquisition device and the image acquisition device to avoid errors in the fused data caused by fusing data at different times.
[0059] The server receives time data sent by a timing module of a timing system. In the disclosed embodiment, the timing system may include but is not limited to a global positioning system (GPS) and a global navigation satellite system (GNSS). The timing module may include but is not limited to a GPS timing module and a Beidou timing module. Preferably, in the disclosed embodiment, the timing module is a GPS timing module. The time data may be National Marine Electronics Association (NEMA) data, and NEMA data may have whole second time information accurate to nanosecond (ns) level.
[0060] After receiving the time data, the server parses the correct GPS time from the time data (for example, NEMA data) as the parsed time, which can also be called Universal Time Coordinated (UTC); further, the server calibrates the time of the data acquisition device and the image acquisition device based on the parsed GPS time (UTC time), that is, it synchronizes the time of the data acquisition device and the vehicle's computing platform to synchronize the time of the data acquisition device and the image acquisition device.
[0061] It should be noted that the method for implementing time synchronization is not limited to the synchronization based on GPS time as described above. For example, time synchronization can also be performed based on the Ethernet Precision Time Protocol (PTP).
[0062] In autonomous driving systems, data from various sensors can be fused to achieve better autonomous control. Considering that various sensors are installed in different locations on the vehicle, different coordinate systems may be used when collecting data. Therefore, to ensure a one-to-one correspondence between the collected point cloud data and each location on the image data, the data acquisition device and the image acquisition device need to be calibrated. That is, the data collected by the data acquisition device and the image acquisition device are projected into the vehicle coordinate system using a transformation matrix.
[0063] In the embodiments of the present disclosure, a transformation matrix is a concept in linear algebra. In linear algebra, linear transformations can be represented by transformation matrices. The most commonly used geometric transformations are linear transformations, including rotation, translation, scaling, shear, reflection, and projection. In the embodiments of the present disclosure, linear transformations including rotation and translation are used as an example for illustration. That is, the linear transformations that can be represented by the transformation matrix include rotation and translation.
[0064] The vehicle coordinate system (VCS) is a special moving coordinate system used to describe the motion of a vehicle. The origin of the vehicle coordinate system is fixed relative to the vehicle position and coincides with the center of mass of the vehicle. The center of mass of the vehicle may include but is not limited to the center of the head, the center of the front axle, the center of the rear axle, etc. Preferably, in the embodiment of the present disclosure, the center of mass of the vehicle is the center of the rear axle of the vehicle, that is, the vehicle coordinate system is established with the center of the rear axle of the vehicle as the origin. The vehicle coordinate system can be established as a left-handed system or a right-handed system, and the embodiment of the present disclosure does not limit this. For example, the vehicle coordinate system is established as a left-handed system. When the vehicle is stationary on a horizontal road, the X-axis of the vehicle coordinate system is parallel to the ground and points to the front of the vehicle, the Y-axis of the vehicle coordinate system points to the left of the driver, and the Z-axis of the vehicle coordinate system is perpendicular to the ground and points to the top of the vehicle. For example, a vehicle coordinate system is established as a right-hand system. When the vehicle is stationary on a level road, the X-axis of the vehicle coordinate system is parallel to the ground and points forward of the vehicle, the Y-axis of the vehicle coordinate system points to the driver's right, and the Z-axis of the vehicle coordinate system is perpendicular to the ground and points upward of the vehicle. It should be understood that the established vehicle coordinate system may differ at different times during the vehicle's driving process due to the vehicle's different positions.
[0065] By synchronizing the data acquisition device and image acquisition device using the timing system's timing module as the timing clock source, different acquisition devices can achieve time synchronization based on the same timing clock source, thereby avoiding inconsistencies in the clocks of different acquisition devices. Furthermore, by synchronizing the data acquisition device and image acquisition device, the first point cloud data and the first image data acquired can be obtained at the same time, thereby ensuring that the first point cloud data and the first image data are data from the same time, thereby improving the accuracy of the target obstacle detection results.
[0066] After timing the data acquisition device and the vehicle's computing platform, and upon receiving the rising edge of the pulse per second (PPS) signal sent by the GPS timing module of the timing system, the server triggers the data acquisition device to collect the first point cloud data, sets the time of the collected first point cloud data to the GPS time, that is, the time of the output first point cloud data is the calibrated time, and maintains this time reference to continuously accumulate and output the first point cloud data, thereby achieving time synchronization between the GPS timing module and the data acquisition device. The server also triggers the image acquisition device to collect the first image data, sets the time of the collected first image data to the GPS time, that is, the time of the output first image data is the calibrated time, and maintains this time reference to continuously accumulate and output the first image data, thereby achieving time synchronization between the GPS timing module and the image acquisition device.
[0067] It should be noted that the time synchronization steps of the data acquisition device and the image acquisition device can be executed simultaneously.
[0068] Although the above method achieves clock domain synchronization between the data acquisition device and the image acquisition device, data acquisition synchronization between the two devices remains unaffected by factors such as the image acquisition device's sampling frequency. To address this issue, in the disclosed embodiment, a trigger signal from the vehicle's computing platform is used as a reference to control the image acquisition device's exposure triggering, thereby achieving hard synchronization between the data acquisition device and the image acquisition device. This means that both devices initiate data acquisition simultaneously.
[0069] Specifically, the server sends the parsed GPS time to the vehicle's computing platform, that is, sends the parsed time to the vehicle's computing platform. After receiving the PPS signal sent by the GPS timing module of the timing system, the vehicle's computing platform adds the set time to the GPS time as its own calibration time, and updates the system time of the vehicle's computing platform to achieve time synchronization between the GPS timing module and the vehicle's computing platform. After receiving the rising edge of the PPS signal sent by the GPS timing module of the timing system, the vehicle's computing platform sends a trigger signal to the image acquisition device to trigger the image acquisition device to capture the first image data, thereby achieving time synchronization between the GPS timing module and the image acquisition device.
[0070] It should be noted that the time synchronization between the data acquisition module and the GPS timing module does not depend on the time synchronization of the vehicle's computing platform, while the time synchronization between the GPS timing module and the image acquisition device requires calibration of the vehicle's computing platform time first.
[0071] In addition, it should be noted that the above-mentioned preset time threshold may include the delay time of the image acquisition device (for example, a camera).
[0072] Optionally, considering that the image acquisition device has an exposure delay, in the embodiment of the present disclosure, time delay compensation is also performed on the first image data acquired by the image acquisition device, that is, the delay compensation time is added to the time corresponding to the trigger signal as the timestamp of the acquired first image data. In the embodiment of the present disclosure, the delay compensation time refers to the time of image exposure. The delay compensation time can be set according to the model of the image acquisition device. For example, the delay compensation time can be 30 milliseconds (ms). By performing delay compensation on the acquired first image data, the time synchronization error between the acquired first point cloud data and the first image data can be made less than 50 milliseconds, thereby making the accuracy of the detection result of the target obstacle reach 0.15 meters.
[0073] According to the technical solutions provided by the embodiments of the present disclosure, by unifying the time of the data acquisition device and the time of the image acquisition device to the GPS time frame, hard synchronization of the data acquisition device and the image acquisition device can be achieved. In addition, by using a GPS timing module and adopting hard synchronization to synchronize the time of the data acquisition device and the image acquisition device, microsecond-level accuracy can be achieved while reducing costs.
[0074] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present disclosure, and will not be described in detail here. In addition, the order of the sequence numbers of the steps in the above embodiments does not imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.
[0075] FIG2 is a flow chart of another obstacle detection method provided by an exemplary embodiment of the present disclosure. The obstacle detection method of FIG2 can be executed by a server or electronic device in an autonomous driving system. As shown in FIG2 , the obstacle detection method may include:
[0076] S201, receiving time data sent by the global positioning system, and parsing the time data to obtain parsed time;
[0077] S202, sending the parsing time to the data acquisition device, so that the data acquisition device synchronizes the time of the data acquisition device with the parsing time after receiving the pulse-per-second signal sent by the global positioning system;
[0078] S203, sending the parsing time to the computing platform of the vehicle, so that after detecting the pulse per second signal, the computing platform sends a trigger signal to the image acquisition device to trigger the image acquisition device to perform image acquisition;
[0079] S204, acquiring first point cloud data acquired by a data acquisition device and first image data acquired by an image acquisition device, wherein the first point cloud data and the first image data have a corresponding relationship in time;
[0080] S205, performing semantic segmentation on the first image data using the target obstacle detection model to obtain a semantic segmentation result, where the semantic segmentation result is used to indicate whether the first image data includes semantic information of the target obstacle;
[0081] S206, when the first image data includes a target obstacle, projecting the target obstacle into the first point cloud data to obtain target point cloud data corresponding to the target obstacle;
[0082] S207, calculating geometric parameters of the target obstacle based on the target point cloud data, and filtering the first point cloud data based on the geometric parameters to obtain a detection result of the target obstacle;
[0083] S208, acquiring second point cloud data acquired by the data acquisition device and / or second image data acquired by the image acquisition device;
[0084] S209, fusing the detection result of the target obstacle with the second point cloud data and / or the second image data to generate a vehicle control strategy;
[0085] S210: Control the vehicle based on the vehicle control strategy.
[0086] According to the technical solution provided by the embodiment of the present disclosure, the data acquisition device and the image acquisition device are time-synchronized based on the time data sent by the global positioning system, and the detection result of the target obstacle is obtained based on the first point cloud data and the first image data respectively collected by the time-synchronized data acquisition device and the image acquisition device. The detection result of the target obstacle is fused with the second point cloud data and / or the second image data collected by the data acquisition device and / or the image acquisition device to generate a vehicle control strategy, and the vehicle is controlled based on the vehicle control strategy, which can provide an effective auxiliary decision-making basis for vehicle obstacle avoidance. Therefore, the stability of the automatic driving system is improved, the driving safety of the vehicle is guaranteed, and the operation efficiency of the vehicle is improved.
[0087] FIG3 is a flow chart of another obstacle detection method provided by an exemplary embodiment of the present disclosure. The obstacle detection method of FIG3 can be executed by a server or electronic device in an autonomous driving system. As shown in FIG3, the obstacle detection method may include:
[0088] S301, receiving time data sent by the global positioning system, and parsing the time data to obtain parsed time;
[0089] S302, sending the parsing time to the data acquisition device, so that the data acquisition device synchronizes the time of the data acquisition device with the parsing time after receiving the pulse-per-second signal sent by the global positioning system;
[0090] S303, sending the parsing time to the computing platform of the vehicle, so that after detecting the pulse per second signal, the computing platform sends a trigger signal to the image acquisition device to trigger the image acquisition device to perform image acquisition;
[0091] S304, acquiring first point cloud data acquired by a data acquisition device and first image data acquired by an image acquisition device, wherein the first point cloud data and the first image data have a corresponding relationship in time;
[0092] S305: Perform semantic segmentation on the first image data using the target obstacle detection model to obtain a semantic segmentation result, where the semantic segmentation result is used to indicate whether the first image data includes semantic information of the target obstacle;
[0093] S306, when the first image data includes a target obstacle, projecting the target obstacle into the first point cloud data to obtain target point cloud data corresponding to the target obstacle;
[0094] S307, calculating geometric parameters of the target obstacle based on the target point cloud data, and filtering the first point cloud data based on the geometric parameters to obtain a detection result of the target obstacle;
[0095] S308: Send the detection result of the target obstacle to the cloud server, so that the cloud server updates the obstacle layer data in the high-precision map based on the detection result of the target obstacle;
[0096] S309, receiving the updated high-precision map sent by the cloud server, so that the operator can process the target obstacle based on the obstacle layer data in the updated high-precision map;
[0097] S310: Send the processing result of the target obstacle to the cloud server, so that the cloud server updates the updated high-precision map based on the processing result of the target obstacle.
[0098] According to the technical solution provided by the embodiment of the present disclosure, a data acquisition device and an image acquisition device are time-synchronized based on time data sent by a global positioning system, and a detection result of a target obstacle is obtained based on the first point cloud data and the first image data respectively collected by the time-synchronized data acquisition device and the image acquisition device. The detection result of the target obstacle is sent to a cloud server to update the obstacle layer data in the high-precision map. The updated high-precision map sent by the cloud server is received so that an operator can process the target obstacle based on the updated obstacle layer data, and the processing result of the target obstacle is sent to the cloud server to update the updated high-precision map. This can provide the operator with accurate obstacle layer data so that the operator can process the target obstacle in a timely manner. Therefore, the processing efficiency of the target obstacle is improved, the processing time of the target obstacle is saved, the timeliness of the high-precision map is ensured, and the driving safety of the vehicle is further ensured.
[0099] In the case of dividing each functional module according to its corresponding function, an embodiment of the present disclosure provides an obstacle detection device, which can be a server or a chip used in a server. Figure 4 is a schematic diagram of the structure of an obstacle detection device provided by an exemplary embodiment of the present disclosure. As shown in Figure 4, the obstacle detection device 400 includes:
[0100] An acquisition module 401 is configured to acquire first point cloud data acquired by a data acquisition device and first image data acquired by an image acquisition device, wherein the first point cloud data and the first image data have a temporal correspondence;
[0101] The segmentation module 402 is configured to perform semantic segmentation on the first image data to obtain a semantic segmentation result, wherein the semantic segmentation result is used to indicate whether the first image data includes semantic information of the target obstacle;
[0102] The conversion module 403 is configured to, when the first image data includes a target obstacle, project the target obstacle into the first point cloud data to obtain target point cloud data corresponding to the target obstacle;
[0103] The calculation module 404 is configured to calculate geometric parameters of the target obstacle based on the target point cloud data, and perform filtering processing on the first point cloud data based on the geometric parameters to obtain a detection result of the target obstacle.
[0104] According to the technical solution provided by the embodiments of the present disclosure, first point cloud data collected by a data acquisition device and first image data collected by an image acquisition device are acquired, wherein the first point cloud data and the first image data have a temporal correspondence; semantic segmentation is performed on the first image data to obtain a semantic segmentation result, wherein the semantic segmentation result is used to indicate whether the first image data includes semantic information of a target obstacle; if the first image data includes the target obstacle, the target obstacle is projected into the first point cloud data to obtain target point cloud data corresponding to the target obstacle; geometric parameters of the target obstacle are calculated based on the target point cloud data, and the first point cloud data are filtered based on the geometric parameters to obtain a target obstacle detection result. The target obstacle identified based on the image data can be projected into the point cloud data, and the target obstacle detection result can be obtained by calculating the geometric parameters of the target obstacle. Therefore, the accuracy of small target obstacle detection is improved, the problem of small target obstacles being easily missed is solved, the vehicle's perception of small target obstacles is enhanced, and the driving safety and operating efficiency of the vehicle are further guaranteed.
[0105] In some embodiments, the segmentation module 402 of FIG. 4 uses a target obstacle detection model to perform semantic reasoning on each pixel point in the first image data to segment the pixel region involved in the target obstacle from the first image data and output the target obstacle category corresponding to the pixel region. The target obstacle detection model is obtained by controlling the target obstacle detection model to be trained to perform a semantic segmentation training task on the sample image data based on the sample image data and the annotation labels associated with the sample image data. The semantic segmentation training task is used to segment the pixel region used to represent the target obstacle from the sample image data. The annotation labels include the pixel region corresponding to the target obstacle in the sample image data annotated by a multimedia data annotation tool.
[0106] In some embodiments, target obstacle categories include ruts, small rocks, and potholes.
[0107] In some embodiments, the obstacle detection device 400 also includes: a fusion module 405 and a control module 406, the acquisition module 401 acquires the second point cloud data acquired by the data acquisition device and / or the second image data acquired by the image acquisition device; the fusion module 405 is configured to fuse the detection result of the target obstacle with the second point cloud data and / or the second image data to generate a vehicle control strategy; the control module 406 is configured to control the vehicle based on the vehicle control strategy.
[0108] In some embodiments, the obstacle detection device 400 also includes: a sending module 407 and a receiving module 408, the sending module 407 is configured to send the detection result of the target obstacle to the cloud server, so that the cloud server updates the obstacle layer data in the high-precision map based on the detection result of the target obstacle; and send the processing result of the target obstacle to the cloud server, so that the cloud server updates the updated high-precision map based on the processing result of the target obstacle; the receiving module 408 is configured to receive the updated high-precision map sent by the cloud server, so that the operator processes the target obstacle based on the obstacle layer data in the updated high-precision map.
[0109] In some embodiments, the temporal correspondence between the first point cloud data and the first image data includes: the timestamp of the first point cloud data is the same as the timestamp of the first image data; or the difference between the timestamp of the first point cloud data and the timestamp of the first image data is less than or equal to a preset time threshold.
[0110] In some embodiments, the obstacle detection device 400 also includes: a timing module 409 and a trigger module 410. The timing module 409 is configured to receive time data sent by the timing system and to perform timing on the data acquisition device and the vehicle's computing platform respectively based on the time data; the trigger module 410 is configured to trigger the data acquisition device to collect the first point cloud data after receiving the second pulse signal sent by the timing system, and to trigger the image acquisition device to collect the first image data through the vehicle's computing platform.
[0111] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0112] An embodiment of the present disclosure also provides an electronic device, comprising: at least one processor; a memory for storing instructions executable by at least one processor; wherein the at least one processor is used to execute instructions to implement the corresponding steps in the above method disclosed in the embodiment of the present disclosure.
[0113] Figure 5 is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present disclosure. As shown in Figure 5, the electronic device 500 includes at least one processor 501 and a memory 502 coupled to the processor 501. The processor 501 can execute the above method disclosed in the embodiment of the present disclosure.
[0114] The processor 501 may also be referred to as a central processing unit (CPU), which may be an integrated circuit chip with signal processing capabilities. Each step in the method disclosed in the embodiment of the present disclosure may be performed by hardware integrated logic circuits in the processor 501 or by software instructions. The processor 501 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiment of the present disclosure may be directly embodied as being executed by a hardware decoding processor, or may be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in the memory 502, such as a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The processor 501 reads the information in the memory 502 and performs the steps of the method in combination with its hardware.
[0115] In addition, when various operations / processes according to the present disclosure are implemented via software and / or firmware, the programs constituting the software can be installed from a storage medium or a network to a computer system having a dedicated hardware structure, such as the computer system 600 shown in FIG6 . When the various programs are installed, the computer system can perform various functions, including those described above. FIG6 is a schematic diagram of the structure of a computer system provided by an exemplary embodiment of the present disclosure.
[0116] Computer system 600 is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic equipment can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0117] As shown in FIG6 , a computer system 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the computer system 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0118] Multiple components within computer system 600 are connected to I / O interface 605, including an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. Input unit 606 can be any type of device capable of inputting information into computer system 600. Input unit 606 can receive input numeric or character information and generate key input signals related to user settings and / or function control of an electronic device. Output unit 607 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 608 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 609 allows computer system 600 to exchange information / data with other devices over a network, such as the Internet, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0119] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units for running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above. For example, in some embodiments, the above-mentioned method disclosed in the embodiments of the present disclosure may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the computer system 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 may be used to perform the above-mentioned method disclosed in the embodiments of the present disclosure by any other appropriate means (e.g., by means of firmware).
[0120] An embodiment of the present disclosure further provides a computer-readable storage medium, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above method disclosed in the embodiment of the present disclosure.
[0121] The computer-readable storage medium in the embodiments of the present disclosure can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. The above-mentioned computer-readable storage medium can include but is not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the above. More specifically, the above-mentioned computer-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0122] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0123] The embodiments of the present disclosure further provide a computer program product, including a computer program, wherein the computer program implements the above method disclosed in the embodiments of the present disclosure when executed by a processor.
[0124] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer.
[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0126] The modules, components, or units described in the embodiments of the present disclosure may be implemented in software or hardware. The names of the modules, components, or units do not necessarily limit the modules, components, or units themselves.
[0127] The functions described above herein may be at least partially performed by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), an application specific standard product (ASSP), a system on a chip (SOC), a complex programmable logic device (CPLD), and the like.
[0128] The above descriptions are merely some embodiments of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present disclosure.
[0129] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. An obstacle detection method, applied to a vehicle, the method comprising: Obtaining first point cloud data collected by a data acquisition device and first image data collected by an image acquisition device, wherein there is a corresponding relationship in time between the first point cloud data and the first image data; Performing semantic segmentation on the first image data to obtain a semantic segmentation result, wherein the semantic segmentation result is used to characterize whether the first image data includes semantic information of a target obstacle; In the case where the first image data includes the target obstacle, projecting the target obstacle onto the first point cloud data to obtain target point cloud data corresponding to the target obstacle; Based on the target point cloud data, calculating geometric parameters of the target obstacle, and performing filtering processing on the first point cloud data based on the geometric parameters to obtain a detection result of the target obstacle.
2. The method according to claim 1, wherein The performing semantic segmentation on the first image data to obtain a semantic segmentation result includes: Using a target obstacle detection model to perform semantic inference on each pixel point in the first image data, so as to segment out a pixel region related to the target obstacle from the first image data, and outputting a target obstacle category corresponding to the pixel region, wherein the target obstacle detection model is obtained by controlling a target obstacle detection model to be trained to perform a semantic segmentation training task on the sample image data based on sample image data and annotation labels associated with the sample image data, the semantic segmentation training task is used to segment out a pixel region for characterizing the target obstacle from the sample image data, and the annotation labels include pixel regions corresponding to the target obstacle in the sample image data annotated by a multimedia data annotation tool.
3. The method according to claim 2, wherein The target obstacle category includes rut prints, small stones, and potholes.
4. The method according to claim 1, wherein The method further comprises: Obtaining second point cloud data collected by the data acquisition device and / or second image data collected by the image acquisition device; Fusing the detection result of the target obstacle with the second point cloud data and / or the second image data to generate a vehicle control strategy; Controlling the vehicle based on the vehicle control strategy.
5. The method according to claim 1, wherein, The method further comprises: Sending the detection result of the target obstacle to a cloud server, so that the cloud server updates obstacle layer data in a high-precision map based on the detection result of the target obstacle; Receiving the updated high-precision map sent by the cloud server, so that an operator processes the target obstacle based on the obstacle layer data in the updated high-precision map; Sending the processing result of the target obstacle to the cloud server, so that the cloud server updates the updated high-precision map based on the processing result of the target obstacle.
6. The method according to claim 1, wherein The fact that there is a corresponding relationship in time between the first point cloud data and the first image data includes: The time stamp of the first point cloud data is the same as the time stamp of the first image data; or The difference between the time stamp of the first point cloud data and the time stamp of the first image data is less than or equal to a preset time threshold.
7. The method according to any one of claims 1 to 6, wherein, The method further includes: Receiving time data sent by a timing system, and respectively timing the data acquisition device and the computing platform of the vehicle based on the time data; After receiving the second pulse signal sent by the timing system, triggering the data acquisition device to acquire the first point cloud data, and triggering the image acquisition device to acquire the first image data through the computing platform of the vehicle.
8. An obstacle detection device applied to a vehicle, the device includes: An acquisition module configured to acquire first point cloud data acquired by a data acquisition device and first image data acquired by an image acquisition device, wherein there is a corresponding relationship in time between the first point cloud data and the first image data; A segmentation module configured to perform semantic segmentation on the first image data to obtain a semantic segmentation result, wherein the semantic segmentation result is used to represent whether the first image data includes semantic information of a target obstacle; A conversion module configured to project the target obstacle into the first point cloud data to obtain target point cloud data corresponding to the target obstacle when the first image data includes the target obstacle; A calculation module configured to calculate geometric parameters of the target obstacle based on the target point cloud data, and perform filtering processing on the first point cloud data based on the geometric parameters to obtain a detection result of the target obstacle.
9. An electronic device, including: At least one processor; A memory for storing executable instructions of the at least one processor; Wherein, the at least one processor is used to execute the instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Stereoscopic vision data acquisition system with time synchronization and method thereof
CN113099211A
DeeplabV3 +-based instant positioning and mapping method for obstacles in unfamiliar sea area
CN114445572A
Three-dimensional target detection method and device
CN115147328A
Method and device for determining road sign of high-precision map and electronic equipment
CN115601516A
Three-dimensional target detection method and device and computer storage medium
CN115909269A
Cited By
Unmanned vehicle obstacle avoidance processing method and system in smart park
CN121325875A
Vehicle sensing data uploading and cloud fusion method and device based on intelligent driving and storage medium
CN121967487A