Drivable area detection method and apparatus, and vehicle

By inputting sensor data into the predictive model to generate drivable area information, the real-time performance and cost issues of relying on high-precision maps are resolved, enabling more accurate vehicle planning and decision-making, and improving the safety and economy of intelligent driving.

WO2026044764A1PCT designated stage Publication Date: 2026-03-05YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/116171
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-31
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing intelligent driving technologies rely on high-precision maps to detect drivable areas, which suffers from poor real-time performance and high costs. Furthermore, in the absence of road topology or in the presence of unstable topology, the detection results are inaccurate, affecting the accuracy of vehicle planning and decision-making.

Method used

Data is collected by sensors and input into the prediction model to generate drivable area information. A general obstacle detection model and a visual language model (VLM) are used for prediction to generate drivable area information that does not rely on high-precision maps. The prediction is then fused by combining grid features and map features.

Benefits of technology

It improves the accuracy of vehicle planning and decision-making, reduces the cost of intelligent driving, avoids unstable detection results due to road topology loss or instability, and improves driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024116171_05032026_PF_FP_ABST
    Figure CN2024116171_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a drivable area detection method and apparatus, and a vehicle. The method can be applied in the field of intelligent driving. The method comprises: acquiring data collected by a sensor; inputting the data into a general obstacle detection model, so as to obtain features of a plurality of grids, wherein a feature of each grid among the plurality of grids comprises one or more of the occupancy information, speed, direction and visibility of the grid; and inputting the features of the plurality of grids into a first prediction model, so as to obtain information of a drivable area for a vehicle. The present application can be applied to an intelligent vehicle or an electric vehicle, and a drivable area for a vehicle is determined by means of data collected by a sensor, so as to facilitate an improvement in the accuracy of planning and decision-making of the vehicle, and also facilitate a reduction in the cost of intelligent driving.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, devices and vehicles for detecting drivable areas Technical Field

[0001] This application relates to the field of intelligent driving, and more specifically, to a method, apparatus, and vehicle for detecting drivable areas. Background Technology

[0002] Vehicles in autonomous driving mode rely on the detection results of their drivable areas when making planning and decisions. Currently, drivable area detection depends on high-definition (HD) maps. However, HD maps have poor real-time performance; in scenarios involving road topology changes, such as road construction detours, the drivable areas obtained by the vehicle in autonomous driving mode may be incorrect, leading to inaccurate planning and decisions. Furthermore, relying on HD maps also increases the cost of autonomous driving. In addition, drivable areas can currently be generated based on road topology without relying on HD. However, if the road where the vehicle is located lacks road topology or the road topology detection results are unstable, the accuracy of the drivable area detection results will be affected, consequently impacting the accuracy of the vehicle's planning and decision-making.

[0003] Summary of the Invention

[0004] This application provides a drivable area detection method, apparatus, and vehicle. By inputting data collected by sensors into a prediction model for prediction, drivable areas can be obtained, which helps improve the accuracy of vehicle planning and decision-making, and also helps reduce the cost of intelligent driving.

[0005] In a first aspect, a drivable area detection method is provided, comprising: acquiring data collected by sensors; inputting the data into a general obstacle detection model to obtain features of multiple grids, wherein the features of each grid include one or more of the following: occupancy information, speed, direction, or visibility of each grid; and inputting the features of the multiple grids into a first prediction model to obtain information on the drivable area of ​​the vehicle.

[0006] Based on the above technical solution, by inputting the features of multiple grids output by the general obstacle detection model into the first prediction model, information about the vehicle's drivable area can be predicted. In this way, the generation of the drivable area does not rely on high-precision maps but is based on data collected by sensors, which helps improve the accuracy of vehicle planning and decision-making, and also helps reduce the cost of intelligent driving. At the same time, the generation of drivable area information does not depend on road topology, which avoids the inability to detect drivable areas due to road topology loss, and also avoids unstable detection results due to unstable road topology, thus preventing unexpected steering and degradation problems caused by the vehicle.

[0007] In conjunction with the first aspect, in some implementations of the first aspect, the first prediction model is a visual language model (VLM).

[0008] Based on the above technical solution, predicting drivable areas using the VLM model eliminates the need for extensive labeled training data, and can even be trained without labeled data, thus reducing training costs. Furthermore, because VLM possesses strong scene understanding and reasoning capabilities, inputting task prompts and features from multiple grids into VLM can improve prediction accuracy, thereby enhancing user driving safety.

[0009] In conjunction with the first aspect, in some implementations of the first aspect, the features of the multiple grids are input into the first prediction model to obtain information about the drivable area of ​​the vehicle, including: inputting the features of the multiple grids and task prompts into the VLM to obtain scene understanding information related to the drivable area.

[0010] In conjunction with the first aspect, in some implementations of the first aspect, the information on the drivable area of ​​the vehicle is information on the three-dimensional (3D) drivable area.

[0011] For example, the information of the 3D drivable area includes the three-dimensional coordinates of multiple points on the polygonal outline corresponding to the 3D drivable area in the vehicle coordinate system.

[0012] For example, the information of the 3D drivable area includes the three-dimensional coordinates of the road boundary in the vehicle coordinate system.

[0013] In conjunction with the first aspect, in some implementations of the first aspect, the features of the multiple grids are input into the first prediction model to obtain information about the drivable area of ​​the vehicle, including: inputting the features of the multiple grids and first information into the first prediction model to obtain information about the drivable area, wherein the first information includes standard definition (SD) maps and / or crowdsourced data.

[0014] Based on the above technical solution, by inputting the features and initial information of multiple grids into the first prediction model, beyond-line-of-sight information can be provided for the final drivable area. This helps to increase the range of detected drivable areas and provides more options for the planning and decision-making of the traffic control module.

[0015] In conjunction with the first aspect, in some implementations of the first aspect, before inputting the features of the multiple grids into the first prediction model to obtain information about the drivable area of ​​the vehicle, the method further includes: inputting second information into the second prediction model to obtain map features, the second information including standard SD maps and / or crowdsourced data; wherein, inputting the features of the multiple grids into the first prediction model to obtain information about the drivable area of ​​the vehicle includes: fusing the features of the multiple grids and the map features to obtain fused features; and inputting the fused features into the first prediction model to obtain information about the drivable area.

[0016] Based on the above technical solution, by performing pre-fusion of map features and features from multiple grids, and then making predictions, information on the drivable area of ​​a vehicle can be obtained.

[0017] In conjunction with the first aspect, in some implementations of the first aspect, the features of the multiple grids are input into a first prediction model to obtain information about the drivable area of ​​the vehicle, including: inputting the features of the multiple grids into the first prediction model to obtain a 2.5D grid prediction result, the 2.5D grid prediction result including the attributes of each grid in the multiple two-dimensional 2D grids and the height information of the grid corresponding to the drivable area in the multiple 2D grids; and determining the information of the drivable area based on the 2.5D grid prediction result.

[0018] Based on the above technical solution, by inputting the features of multiple grids into the first prediction model, a 2.5D grid prediction result can be obtained. Then, based on this 2.5D grid prediction result, information about the drivable area can be obtained. This reduces the convergence difficulty of the network.

[0019] In some possible implementations, the attributes of each grid cell in the 2D grid include one or more of the following: drivability attribute, boundary attribute, and road attribute.

[0020] In some possible implementations, the information of the drivable area is determined based on the 2.5D grid prediction result, including post-processing the 2.5D grid prediction result (e.g., edge extraction) to obtain the information of the drivable area.

[0021] In conjunction with the first aspect, in some implementations of the first aspect, the information of the drivable area includes the boundary attributes and / or road attributes of the drivable area.

[0022] Based on the above technical solution, when the planning and control module makes plans and decisions, it can combine the boundary attributes and / or road attributes to make the results of the decisions and plans more accurate, which helps to improve the driving safety of users.

[0023] In some possible implementations, the method also includes determining, based on information about the drivable area, whether the vehicle should maintain or exit intelligent driving mode.

[0024] In some possible ways, the vehicle can divide the area to be detected into a drivable area, a non-drivable area, and a low-traction area. The low-traction area can be further divided into snow-covered roads, water-covered roads, slippery roads, gravel roads, oil-stained roads, or grass, etc.

[0025] In some possible implementations, the boundary attribute may include soft boundaries and / or hard boundaries, where there is a risk of accident or collision if the vehicle crosses the hard boundary.

[0026] In conjunction with the first aspect, in some implementations of the first aspect, the drivable area includes a first area, the boundary attributes of which include information on soft boundaries and / or hard boundaries, wherein the soft boundary includes the boundary where the first area and the second area intersect, the road attributes of the first area indicate a first road surface type, the road attributes of the second area indicate a second road surface type, and the traffic priority of the first road surface type is higher than the traffic priority of the second road surface type; the hard boundary includes the boundary where the first area intersects with at least one of the following: curb, flower bed, fence, ditch edge, and cliff edge.

[0027] In conjunction with the first aspect, in some implementations of the first aspect, the road attributes of the drivable area include the road type of one or more polygonal regions where the drivable area is located, the road type being used to indicate whether the polygonal region is a dry road surface, a flooded road surface, a slippery road surface, a snow-covered road surface, a gravel road surface, or a grassy area.

[0028] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: sending information about the drivable area to the planning and control module.

[0029] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: controlling the vehicle's movement based on information about the drivable area.

[0030] Secondly, this application provides a method for detecting drivable areas, the method comprising: acquiring data collected by sensors; inputting the data into a first prediction model to obtain information on the three-dimensional drivable area of ​​the vehicle, wherein the information on the 3D drivable area includes the boundary attributes and height of one or more polygons in which the drivable area of ​​the vehicle is located.

[0031] Based on the above technical solution, the generation of drivable areas does not rely on high-precision maps, but rather on data collected by sensors. This helps improve the accuracy of vehicle planning and decision-making, and also helps reduce the cost of intelligent driving. Furthermore, the generation of drivable area information does not depend on dynamic targets or static road structure elements; that is, the generation of drivable areas is independent of road topology. This avoids the inability to detect drivable areas due to road topology loss, and also avoids unstable detection results due to unstable road topology, which could lead to unexpected steering and degradation issues.

[0032] Meanwhile, the information of this 3D drivable area includes the boundary attributes and height of one or more polygons, which can make the planning and decision-making results of the control module more accurate and help improve the driving safety of users.

[0033] In conjunction with the second aspect, in some implementations of the second aspect, the information of the 3D drivable area also includes the road attributes of each polygon in the one or more polygons.

[0034] Based on the above technical solution, the information of the 3D drivable area includes the road attributes of each polygon in one or more polygons. Since different road attributes correspond to different adhesion forces, this facilitates planning and decision-making by the control module. For example, the control module can control the vehicle to travel in high-adhesion areas, thereby preventing vehicle slippage and improving user safety. Alternatively, the control module can also control the vehicle to travel in low-adhesion areas. Because different road surface materials in low-adhesion areas correspond to different coefficients of friction and slip rates, the control module can limit the vehicle's driving parameters (e.g., speed or curvature of the driving trajectory) based on different road surface materials, further enhancing user safety.

[0035] In conjunction with the second aspect, in some implementations of the second aspect, the data is input into the first prediction model to obtain information about the vehicle's three-dimensional drivable area, including: inputting the data into a general obstacle detection model to obtain features of multiple grids, wherein the features of each grid include one or more of the following: occupancy information, speed, direction, or visibility; and inputting the features of the multiple grids into the first prediction model to obtain information about the 3D drivable domain.

[0036] In conjunction with the second aspect, in some implementations of the second aspect, before inputting the features of the multiple grids into the first prediction model to obtain the information of the 3D drivable domain, the method further includes: acquiring first information, the first information including information of a standard SD map and / or crowdsourced data; inputting the first information into a second prediction model to obtain map features; wherein, inputting the features of the multiple grids into the first prediction model to obtain the information of the 3D drivable domain includes: fusing the features of the multiple grids and the map features to obtain fused features; and inputting the fused features into the first prediction model to obtain the information of the 3D drivable area.

[0037] In conjunction with the second aspect, in some implementations of the second aspect, the features of the multiple grids are input into the first prediction model to obtain the information of the 3D drivable domain, including: inputting the features of the multiple grids into the first prediction model to obtain a 2.5D grid prediction result, the 2.5D grid prediction result including the attributes of each grid in the multiple two-dimensional 2D grids and the height information of the grids corresponding to the drivable area in the multiple 2D grids; and determining the information of the 3D drivable domain based on the 2.5D grid prediction result.

[0038] Thirdly, this application provides a drivable area detection device, which includes: an acquisition unit for acquiring data collected by sensors; a first prediction unit for inputting the data into a general obstacle detection model to obtain features of multiple grids, wherein the features of each grid include one or more of the following: occupancy information, speed, direction, or visibility; and a second prediction unit for inputting the features of the multiple grids into the first prediction model to obtain information about the drivable area of ​​the vehicle.

[0039] In conjunction with the third aspect, in some implementations of the third aspect, the second prediction unit is specifically used to: input the features and task cues of the multiple grids into the VLM to obtain scene understanding information related to the drivable area.

[0040] In conjunction with the third aspect, in some implementations of the third aspect, the second prediction unit is specifically used to: input the features of the multiple grids and the first information into the first prediction model to obtain information about the drivable area, wherein the first information includes standard SD maps and / or crowdsourced data.

[0041] In conjunction with the third aspect, in some implementations of the third aspect, the device further includes: a third prediction unit, used to input second information into a second prediction model to obtain map features, the second information including a standard SD map and / or crowdsourced data; a fusion unit, used to fuse the features of the multiple grids and the map features to obtain fused features; the second prediction unit is specifically used to: input the fused features into the first prediction model to obtain information about the drivable area.

[0042] In conjunction with the third aspect, in some implementations of the third aspect, the second prediction unit is specifically used to: input the features of the plurality of grids into the first prediction model to obtain a 2.5D grid prediction result, the 2.5D grid prediction result including the attributes of each grid in the plurality of two-dimensional 2D grids and the height information of the grid corresponding to the drivable area in the plurality of 2D grids; wherein, the device further includes: a determination unit, used to determine the information of the drivable area based on the 2.5D grid prediction result.

[0043] In conjunction with the third aspect, in some implementations of the third aspect, the information of the drivable area includes the boundary attributes and / or road attributes of the drivable area.

[0044] In conjunction with the third aspect, in some implementations of the third aspect, the drivable area includes a first area, the boundary attributes of which include information on soft boundaries and / or hard boundaries, wherein the soft boundary includes the boundary where the first area and the second area intersect, the road attributes of the first area indicate a first road surface type, the road attributes of the second area indicate a second road surface type, and the traffic priority of the first road surface type is higher than the traffic priority of the second road surface type; the hard boundary includes the boundary where the first area intersects with at least one of the following: curb, flower bed, fence, ditch edge, and cliff edge.

[0045] In conjunction with the third aspect, in some implementations of the third aspect, the road attributes of the drivable area include the road type of one or more polygonal regions where the drivable area is located, which is used to indicate whether the polygonal region is a dry road surface, a flooded road surface, a slippery road surface, a snow-covered road surface, a gravel road surface, or a grassy area.

[0046] In conjunction with the third aspect, in some implementations of the third aspect, the device further includes: a transmitting unit for transmitting information about the drivable area to the control module.

[0047] In conjunction with the third aspect, in some implementations of the third aspect, the device further includes a control unit for controlling the movement of the vehicle based on information about the drivable area.

[0048] Fourthly, this application provides a drivable area detection device, which includes: an acquisition unit for acquiring data collected by a sensor; and a prediction unit for inputting the data into a first prediction model to obtain information on the three-dimensional drivable area of ​​the vehicle, wherein the information on the 3D drivable area includes the boundary attributes and height of one or more polygons in which the drivable area of ​​the vehicle is located.

[0049] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the information of the 3D drivable area also includes the road attributes of each polygon in the one or more polygons.

[0050] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the prediction unit is specifically used to: input the data into a general obstacle detection model to obtain features of multiple grids, wherein the features of each grid include one or more of the following: occupancy information, speed, direction, or visibility of each grid; and input the features of the multiple grids into the first prediction model to obtain information about the 3D drivable domain.

[0051] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the acquisition unit is further configured to acquire first information, which includes information from a standard SD map and / or crowdsourced data; the prediction unit is further configured to input the first information into a second prediction model to obtain map features; wherein, the device further includes: a fusion unit configured to fuse the features of the multiple grids and the map features to obtain fused features; the prediction unit is specifically configured to: input the fused features into the first prediction model to obtain information about the 3D drivable area.

[0052] In conjunction with the fourth aspect, in some implementations of the fourth aspect, the prediction unit is specifically used to: input the features of the multiple grids into the first prediction model to obtain a 2.5D grid prediction result, the 2.5D grid prediction result including the attributes of each grid in the multiple two-dimensional 2D grids and the height information of the grids corresponding to the drivable area in the multiple 2D grids; wherein, the device further includes: a determination unit, used to determine the information of the 3D drivable domain based on the 2.5D grid prediction result.

[0053] Fifthly, this application provides a drivable area detection device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program in the memory, enabling the intelligent driving device to implement the method in any of the possible implementations of the first or second aspect described above.

[0054] Sixthly, this application provides a drivable area detection system, which includes a sensing system and the apparatus described in any one of the third to fifth aspects above.

[0055] In a seventh aspect, this application provides a vehicle that includes the device described in the third, fourth, or fifth aspects above, or the system described in the sixth aspect above.

[0056] The term "vehicle" in this application is used in a broad sense and can refer to means of transportation (such as commercial vehicles, passenger cars, motorcycles, flying cars, trains, etc.), industrial vehicles (such as forklifts, trailers, tractors, etc.), engineering vehicles (such as excavators, bulldozers, cranes, etc.), agricultural equipment (such as lawnmowers, harvesters, etc.), amusement equipment, toy vehicles, etc. The embodiments of this application do not specifically limit the type of vehicle.

[0057] Eighthly, this application provides a computer program product comprising: computer program code, which, when executed on a computer, causes the computer to perform the method in any possible implementation of the first or second aspect.

[0058] Ninthly, this application provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the method in any possible implementation of the first or second aspect.

[0059] In a tenth aspect, this application provides a chip including circuitry for performing the methods in any possible implementation of the first or second aspect described above. Attached Figure Description

[0060] Figure 1 is a functional block diagram of the vehicle provided in an embodiment of this application.

[0061] Figure 2 is a schematic block diagram of the intelligent driving system provided in an embodiment of this application.

[0062] Figure 3 is a schematic flowchart of the drivable area detection method provided in the embodiments of this application.

[0063] Figure 4 is a schematic diagram of the network architecture provided in an embodiment of this application.

[0064] Figure 5 is another schematic diagram of the network architecture provided in the embodiments of this application.

[0065] Figure 6 is another schematic diagram of the network architecture provided in an embodiment of this application.

[0066] Figure 7 is a schematic diagram of an intelligent driving scenario provided in an embodiment of this application.

[0067] Figure 8 is another schematic diagram of the network architecture provided in the embodiments of this application.

[0068] Figure 9 is another schematic diagram of the intelligent driving scenario provided in the embodiments of this application.

[0069] Figure 10 is another schematic diagram of the intelligent driving scenario provided in the embodiments of this application.

[0070] Figure 11 is another schematic flowchart of the drivable area detection method provided in the embodiments of this application.

[0071] Figure 12 is another schematic diagram of the network architecture provided in an embodiment of this application.

[0072] Figure 13 is another schematic diagram of the network architecture provided in the embodiments of this application.

[0073] Figure 14 is a schematic block diagram of the drivable area detection device provided in an embodiment of this application.

[0074] Figure 15 is a schematic block diagram of the drivable area detection device provided in an embodiment of this application.

[0075] Figure 16 is a schematic block diagram of the three-dimensional target detection device provided in an embodiment of this application.

[0076] Figure 17 is a schematic diagram of the module structure of the intelligent vehicle 200 provided in an embodiment of this application.

[0077] Figure 18 is a flowchart illustrating the three-dimensional target detection method provided in an embodiment of this application.

[0078] Figure 19A is a schematic diagram of panoramic segmentation provided in an embodiment of this application.

[0079] Figure 19B is a schematic diagram of generating 360-degree semantic point cloud data provided in an embodiment of this application.

[0080] Figure 20 is a flowchart illustrating the process of determining three-dimensional position information provided in an embodiment of this application.

[0081] Figure 21 is a schematic diagram of generating semantic point cloud data provided in an embodiment of this application.

[0082] Figure 22 is a schematic diagram of three-dimensional target detection provided in an embodiment of this application.

[0083] Figure 23 is a flowchart illustrating the training of the feature extraction network and the object detection network provided in the embodiments of this application.

[0084] Figure 24 is a schematic diagram of the module structure of the feature extraction network provided in the embodiment of this application.

[0085] Figure 25 is a schematic diagram of the module structure of the feature extraction network provided in the embodiments of this application. Detailed Implementation

[0086] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. "At least one" refers to one or more. For example, "at least one of A and B," similar to "A and / or B," describes the association relationship between related objects, indicating that three relationships can exist. For example, at least one of A and B can represent: A existing alone, A and B existing simultaneously, and B existing alone.

[0087] The prefixes such as "first" and "second" used in this application embodiment are merely for distinguishing different descriptive objects and do not limit the position, order, priority, quantity, or content of the described objects. The use of ordinal numbers and other prefixes used to distinguish descriptive objects in this application embodiment does not constitute a limitation on the described objects. The description of the described objects is given in the claims or the context of the embodiments, and should not constitute unnecessary restrictions due to the use of such prefixes. Furthermore, in the description of this embodiment, unless otherwise stated, "multiple" means two or more.

[0088] Figure 1 is a functional block diagram of a vehicle 100 provided in an embodiment of this application. The vehicle 100 may include a sensing system 110, a computing platform 120, and a display device 130. The sensing system 110 may include one or more sensors for sensing information about the environment surrounding the vehicle 100. For example, the sensing system 110 may include a positioning system, which may be a Global Positioning System (GPS), a BeiDou Navigation Satellite System, or another positioning system. As another example, the sensing system 110 may include one or more of the following: an inertial measurement unit (IMU), an accelerometer, a lidar, a millimeter-wave radar, an ultrasonic radar, and a camera device.

[0089] Some or all of the functions of vehicle 100 can be controlled by computing platform 120. Computing platform 120 may include one or more processors, such as processors 121 to 12n (n being a positive integer). A processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a central processing unit (CPU), microprocessor, graphics processing unit (GPU) (which can be understood as a type of microprocessor), or digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as a field-programmable gate array (FPGA). In reconfigurable hardware circuits, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement some or all of the functions of the aforementioned units. Furthermore, the processor can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), tensor processing unit (TPU), deep learning processing unit (DPU), etc. In addition, the computing platform 120 may also include a memory for storing instructions. Some or all of the processors 121 to 12n can call the instructions in the memory to implement the corresponding functions.

[0090] The in-cabin display devices 130 are mainly divided into two categories: the first is the in-vehicle display screen; the second is the projection display screen, such as the head-up display (HUD). An in-vehicle display screen is a physical display screen and an important component of the in-vehicle infotainment system. Multiple displays can be installed in the cabin, such as the digital instrument cluster display, the central control screen, the display screen in front of the front passenger (also known as the front-seat passenger), the display screen in front of the left rear passenger, the display screen in front of the right rear passenger, and even the car window can be used as a display screen. A head-up display, also known as a head-up display system, is mainly used to display driving information such as speed and navigation on a display device in front of the driver (such as the windshield). This reduces the driver's eye-shift time, avoids pupil changes caused by eye-shifting, and improves driving safety and comfort. Examples of HUDs include combiner-HUD (C-HUD) systems, windshield-HUD (W-HUD) systems, and augmented reality HUD (AR-HUD) systems. It should be understood that HUDs can also evolve into other types of systems as technology progresses, and this application does not limit them.

[0091] The above description of the display device 130 uses an in-vehicle display screen and a projection display screen as examples, but the embodiments of this application are not limited thereto. For example, the display device 130 can also be a light display screen or a projection screen.

[0092] Optionally, the structure of the vehicle 100 described above is merely illustrative. In actual applications, various components of the vehicle 100 may be added or removed as needed.

[0093] Vehicle 100 may include an intelligent driving system, which may include an advanced driving assistant system (ADAS) and an autonomous driving system (ADS). The intelligent driving system uses various sensors on the vehicle (including but not limited to: lidar, millimeter-wave radar, camera devices, ultrasonic sensors, global positioning system, inertial measurement unit) to acquire information from the vehicle's surroundings, and analyzes and processes the acquired information to achieve functions such as obstacle perception, target recognition, vehicle positioning, path planning, and driver monitoring / alerts, thereby improving the safety, automation, and comfort of driving the vehicle.

[0094] For example, Figure 2 shows a schematic block diagram of an intelligent driving system provided in an embodiment of this application. The intelligent driving system may include three functional modules: a perception module 210, a planning module 220, and a control module 230. The perception module 210 perceives the environment surrounding the vehicle through sensors and outputs corresponding perception data to the planning module 220. The planning module 220 obtains road element information based on the information acquired by the perception module 210. Based on the vehicle's current location and the road element information, the planning module 220 determines the physical connectivity of the vehicle from its current location to a sampling point, and plans the vehicle's trajectory to that sampling point when the vehicle is physically connected from its current location to that sampling point. The planning module 220 determines the vehicle's strategy space based on this trajectory. The planning module 220 can send this strategy space to the control module 230. The control module 230 can evaluate the strategy space in Euclidean space to make behavioral or interactive decisions for the vehicle.

[0095] The above-mentioned sensing module 210, planning module 220 and control module 230 can be located in the above-mentioned computing platform 120.

[0096] Vehicle-based driving automation systems are classified into five levels (or L0-L5) based on the degree to which they can perform dynamic driving tasks, according to the role allocation in performing these tasks and the presence or absence of an operational design domain (ODD), such as the external conditions (road, traffic, weather, lighting, etc.) defined during the system's design. Levels 0-2 represent driver assistance, where the system assists humans in performing dynamic driving tasks, but the driver remains the primary driver. Levels 3-5 represent autonomous driving, where the system performs dynamic driving tasks in place of the human under the designed operating conditions; when activated, the system becomes the primary driver. The names and definitions of each level are as follows:

[0097] Level 0 driving automation (also known as emergency assistance) systems cannot continuously perform lateral or longitudinal motion control of the vehicle during dynamic driving tasks, but they possess the ability to continuously perform partial target and event detection and response during dynamic driving tasks. Level 1 driving automation (also known as partial driver assistance) systems continuously perform lateral or longitudinal motion control of the vehicle during dynamic driving tasks under their design operating conditions, and possess the ability to perform partial target and event detection and response adapted to the performed lateral or longitudinal motion control. Level 2 driving automation (also known as combined driver assistance) systems continuously perform lateral and longitudinal motion control of the vehicle during dynamic driving tasks under their design operating conditions, and possess the ability to perform partial target and event detection and response adapted to the performed lateral and longitudinal motion control. Level 3 driving automation (also known as conditionally automated driving) systems continuously perform all dynamic driving tasks under their design operating conditions. Level 4 driving automation (also known as highly automated driving) systems continuously perform all dynamic driving tasks under their design operating conditions and automatically execute minimum risk strategies. Level 5 driving automation (also known as fully automated driving) systems continuously perform all dynamic driving tasks and automatically execute minimum-risk strategies under any drivable conditions. Typically, intelligent driving systems fall between Level 2 and Level 5; for example, ADAS is Level 2, and ADS is Level 3-Level 5.

[0098] Figure 3 shows a schematic flowchart of the drivable area detection method 300 provided in an embodiment of this application. This method 300 can be executed by the vehicle 100, or by the computing platform 120; or by a processor, chip, or circuit in the computing platform 120; or by an intelligent driving system; or by the perception module 210; or by a server. The method 300 includes:

[0099] S310 acquires data collected by the sensor.

[0100] For example, the sensor may include, but is not limited to, one or more of a camera device, lidar, and millimeter-wave radar.

[0101] S320, input the data into the general obstacle detection model to obtain the features of multiple grids. The features of each grid include one or more of the following: occupancy information, speed, direction, or visibility.

[0102] The general obstacle detection model involved in the embodiments of this application can refer to the three-dimensional target detection device 1000 shown in Figures 16-25.

[0103] For example, Figure 4 illustrates a schematic diagram of the network architecture provided in an embodiment of this application. This network architecture includes sensors, a general obstacle detection model, and a first prediction model (e.g., VLM). By inputting data collected by the sensors into the general obstacle detection model, a 3D world representation result can be obtained. This 3D world representation result includes features of multiple grids, and the features of each grid include one or more of the following: occupancy information, velocity, orientation, or visibility.

[0104] S330, input the features of the multiple grids into the first prediction model to obtain information on the drivable area of ​​the vehicle.

[0105] Optionally, the first prediction model can be a VLM.

[0106] For example, a VLM can include a visual model (e.g., an image encoder) and a language model (e.g., a text encoder). For instance, a VLM can leverage pre-trained visual and language models, connected by a graph-text input feature alignment module, allowing the language model to understand image features and engage in deeper question-answering reasoning.

[0107] For example, the language model can be a large language model (LLM), the image encoder can be a multilayer perceptron (MLP), and the image-text input feature alignment module can be a Q-Former model. The Q-Former model can be used in cross-modal dialogue systems to improve the interactivity and intelligence of dialogue by understanding and generating image-text mixed dialogue content.

[0108] For example, the first prediction model is the VLM shown in Figure 4. By inputting the 3D world representation results and task prompts into the VLM, scene understanding information related to the drivable area can be obtained.

[0109] For example, the task prompt may include one or more of the following:

[0110] (1) Where is the driving area located?

[0111] (2) Where is the road boundary located?

[0112] (3) What are boundary attributes?

[0113] (4) What is the surface material of the road?

[0114] (5) What is the suggested safety speed on the road?

[0115] For example, scene understanding information related to drivable areas includes one or more of the following:

[0116] (1) Coordinates of the drivable area (e.g., coordinates of points on the polygonal outline corresponding to the drivable area);

[0117] (2) Information on road boundaries;

[0118] (3) Boundary properties (e.g., soft boundary or hard boundary);

[0119] (4) Road surface material (e.g., dry road surface, snow-covered road surface, wet road surface);

[0120] (5) Information on safe driving speed within the drivable area (e.g., 80 km / h).

[0121] Optionally, the features of the multiple grids are input into the first prediction model to obtain information about the drivable area of ​​the vehicle, including: inputting the features of the multiple grids and first information into the first prediction model to obtain information about the drivable area, wherein the first information includes standard SD maps and / or crowdsourced data.

[0122] For example, Figure 5 shows another schematic diagram of the network architecture provided in an embodiment of this application. By inputting 3D world representation results, map information (or crowdsourced data), and task prompts into the VLM, scene understanding information related to drivable areas can be obtained.

[0123] Optionally, the map information (or crowdsourced data) includes the location and attributes of the points. Figures 4 and 5 above illustrate this using a VLM model as the first prediction model, but the embodiments of this application are not limited to this. The first prediction model can also be other models.

[0124] For example, Figure 6 illustrates another schematic diagram of the network architecture provided in an embodiment of this application. The first prediction model can be a convolutional neural network (CNN) as shown in Figure 6. By inputting the 3D world representation results into the CNN, information about the drivable area of ​​the vehicle can be obtained. Alternatively, by inputting the 3D world representation results and map information into the CNN, information about the drivable area of ​​the vehicle can be obtained.

[0125] For example, Figure 7 shows a schematic diagram of an intelligent driving scenario provided in an embodiment of this application.

[0126] As shown in Figure 7(a), vehicle 100 is on a highway. Vehicle 100 can determine the drivable area using data collected by sensors.

[0127] As shown in Figure 7(b), if vehicle 100 determines its drivable area based on data collected by sensors, vehicle 100 can, for example, input the sensor-collected data into the general obstacle detection model to obtain features of multiple grids. Vehicle 100 can then input these features into a first prediction model to obtain information about the drivable area. However, if the ramp entrance is blocked due to target occlusion, road boundary (e.g., guardrail) obstruction, or a large distance between the ramp entrance and vehicle 100, the perception module 210 determines that the ramp entrance in the drivable area is closed.

[0128] As shown in Figure 7(c), vehicle 100 can input the features of the multiple grids and information from the standard map (e.g., the location and attributes of points at ramp entrances in the standard map) into the first prediction model to obtain information about the drivable area. The resulting drivable area is larger. For example, this drivable area indicates that the ramp entrance is not closed. Thus, by combining information from the standard map, beyond-line-of-sight detection can be performed on the drivable area. Even when there is target occlusion, road boundary occlusion, or the ramp is far from vehicle 100, the perception module 210 determines that the ramp entrance is not closed within the drivable area. This effectively solves problems such as ramp closures, sharp curve closures, and unstable beyond-line-of-sight perception, thereby providing more basis for planning and decision-making.

[0129] Figure 7 above illustrates this using a ramp entrance as an example. Beyond-line-of-sight detection can also be applied to other intelligent driving scenarios. For example, during driving, vehicle 100 detects a portion of the polygonal outline of a guardrail using sensors, while another portion of the polygonal outline is obscured by other targets (e.g., other vehicles). Vehicle 100 can perform beyond-line-of-sight detection on the guardrail based on data collected by sensors and crowdsourced data (e.g., data collected by sensors from other vehicles passing through the road segment). This avoids situations where only a portion of the polygonal outline of a static obstacle is detected, or where the static obstacle is not detected at all, due to obstruction by other targets. This helps improve the accuracy of drivable area detection results, thereby enhancing user driving safety.

[0130] Optionally, before inputting the features of the multiple grids into the first prediction model to obtain information about the drivable area of ​​the vehicle, the method 300 further includes: inputting second information into the second prediction model to obtain map features, the second information including standard SD maps and / or crowdsourced data; wherein, inputting the features of the multiple grids into the first prediction model to obtain information about the drivable area of ​​the vehicle includes: fusing the features of the multiple grids and the map features to obtain fused features; and inputting the fused features into the first prediction model to obtain information about the drivable area.

[0131] For example, Figure 8 illustrates another schematic diagram of the network architecture provided in an embodiment of this application. Map features can be encoded by inputting map information (or crowdsourced data) into an encoder (e.g., an encoder). Fusion features can be obtained by fusing the map features with the features of multiple grid cells. Scene understanding information related to drivable areas can be obtained by inputting the fused features and task prompts into a VLM. In this embodiment, the map features and the features of multiple grid cells can be pre-fused to obtain fused features. Then, the fused features and task prompts are input into a VLM for prediction, thereby obtaining scene understanding information related to drivable areas.

[0132] For example, the map encoder can be a transformer.

[0133] Optionally, the features of the multiple grids are input into the first prediction model to obtain information about the drivable area of ​​the vehicle, including: inputting the features of the multiple grids into the first prediction model to obtain a 2.5D grid prediction result, the 2.5D grid prediction result including the attributes of each grid in the multiple two-dimensional 2D grids and the height information of the grids corresponding to the drivable area in the multiple 2D grids; and determining the information of the drivable area based on the 2.5D grid prediction result.

[0134] Optionally, determining the information of the drivable area based on the 2.5D grid prediction result includes: post-processing the 2.5D grid prediction result to obtain the information of the drivable area. For example, edge extraction can be performed on the 2.5D grid prediction result to obtain the information of the drivable area.

[0135] Optionally, the information about the drivable area includes the boundary attributes and / or road attributes of the drivable area.

[0136] Optionally, the road attributes of the drivable area include the road type of one or more polygonal regions where the drivable area is located. The road type is used to indicate whether the polygonal region is a dry road surface, a flooded road surface, a slippery road surface, a snow-covered road surface, a gravel road surface, or a grassy area.

[0137] Optionally, the drivable area includes a first area, the boundary attributes of which include information on soft boundaries and / or hard boundaries. The soft boundary includes the boundary where the first area and the second area intersect. The road attributes of the first area indicate a first road surface type, and the road attributes of the second area indicate a second road surface type. The traffic priority of the first road surface type is higher than that of the second road surface type. The hard boundary includes the boundary where the first area intersects with at least one of the following: curb, flower bed, fence, ditch edge, and cliff edge.

[0138] For example, Figure 9 shows a schematic diagram of an intelligent driving scenario provided in an embodiment of this application.

[0139] As shown in Figure 9, the vehicle 100 can determine the road surface, including guardrails and the dry, gravel, and snow-covered surfaces to the right of the guardrails, based on the data collected by sensors. The vehicle 100 can first divide the area to be detected into drivable, non-drivable, and low-adhesion areas based on the aforementioned network architecture. The area containing the dry surface is designated as the drivable area, while the areas containing the snow-covered and gravel surfaces are designated as low-adhesion areas. The boundary between the drivable and low-adhesion areas can be a soft boundary. The boundary between the drivable and non-drivable areas is a hard boundary.

[0140] Optionally, the vehicle can control its movement based on the boundary attributes of the drivable area. For example, if a collision risk between the vehicle and an obstacle is detected to meet preset conditions and the vehicle can avoid the collision risk by crossing the soft boundary, the vehicle can be controlled to cross the soft boundary. In this way, if the vehicle can avoid the collision risk by crossing the soft boundary, the vehicle can cross the soft boundary in an emergency, thereby helping to improve the safety of the user.

[0141] Alternatively, the vehicle can be controlled to drive in low-traction areas. Since different road surfaces in low-traction areas correspond to different coefficients of friction and slip rates, the planning and control module (such as the planning module 220 and control module 230 mentioned above) can limit the vehicle's driving parameters (such as speed or curvature of the driving trajectory) based on different road surfaces, which helps to improve the user's driving safety.

[0142] For example, Figure 10 shows a schematic diagram of an intelligent driving scenario provided in an embodiment of this application.

[0143] As shown in Figure 10, when a road collapse occurs in front of vehicle 100, the collapsed area can be designated as a non-drivable area. The boundary where the drivable area of ​​the vehicle intersects with the collapsed area is a hard boundary. When the planning module 220 obtains information about the drivable area, it can control the vehicle to perform emergency braking based on the distance between vehicle 100 and the polygonal outline 1 of the drivable area.

[0144] For example, if a number of vehicles greater than or equal to a certain threshold fall within a preset time period within a range of a first preset distance from the direction of travel of vehicle 100, it can be determined that a road collapse has occurred in front of the road.

[0145] For example, if, within a first preset distance from the vehicle along a first direction, a number of vehicles greater than or equal to a certain threshold fall within a preset time period, this includes: within the first preset distance from the vehicle along a first direction, if the rate of change of the visible portion of vehicles greater than or equal to the certain threshold is less than or equal to a threshold of 1, it can be determined that a road collapse has occurred ahead. For example, the visible portion can be the visible portion of the rear of the preceding vehicle. For instance, when vehicles 100 and 200 are calibrated to be on the same horizontal plane, the visible portion of the rear of vehicle 200 is 100%. If the horizontal plane of vehicle 100 is higher than that of vehicle 200, the visible portion of the rear of the preceding vehicle may be less than 100%.

[0146] As shown in Figure 10, the road surface in front of vehicle 100 collapses, causing vehicle 200 to fall into the collapsed area. At this time, vehicle 100 can still detect the position information of vehicle 200, but the visible portion of vehicle 200 is less than 100%. More specifically, when the visible portion of vehicle 200 is less than 100%, the entire rear portion of vehicle 200 can be recovered based on the existing visible portion. Then, the specific value (percentage) of the visible portion of vehicle 200 can be determined based on the proportion of the existing visible portion to the entire rear portion. For example, threshold 1 can be -5% / frame (i.e., the visible portion of the current frame image is reduced by 5% compared to the previous frame image), or it can be -10% / frame, or it can be other values. It can be understood that the smaller the rate of change of the visible portion, the larger the absolute value of the rate of change, that is, the greater the change in the visible portion of vehicle 200.

[0147] Optionally, the method 300 further includes sending information about the drivable area to the planning and control module.

[0148] For example, taking the perception module 210 as the execution subject of method 300, after obtaining the information of the drivable area, the perception module 210 can send the information of the drivable area to the planning module 220.

[0149] Optionally, the method 300 further includes controlling the vehicle's movement based on information about the drivable area.

[0150] Optionally, controlling the vehicle's movement based on the boundary attributes and / or road attributes of the drivable area includes: controlling the distance between the vehicle and the soft boundary to be greater than or equal to a first preset distance; or controlling the distance between the vehicle and the hard boundary to be greater than or equal to a second preset distance; wherein the second preset distance is greater than the first preset distance.

[0151] Figure 11 shows a schematic flowchart of a drivable area detection method 1100 provided in an embodiment of this application. This method 1100 can be executed by the vehicle 100, or by the computing platform 120; or by a processor, chip, or circuit in the computing platform 120; or by an intelligent driving system; or by the perception module 210; or by a server. The method 1100 includes:

[0152] S1110, acquires data collected by the sensor.

[0153] S1120, input the data into the first prediction model to obtain information on the vehicle's three-dimensional drivable area. The information on the 3D drivable area includes the boundary attributes and height of one or more polygons in which the vehicle's drivable area is located.

[0154] For example, Figure 12 shows a schematic diagram of the network architecture provided in an embodiment of this application. The first prediction model can be a drivable area prediction model.

[0155] For example, the drivable area prediction model can be trained using a training dataset, which includes sample data collected by sensors, as well as the coordinates of points on the polygonal contour of the labeled drivable area, the boundary attributes corresponding to the polygonal contour (e.g., soft or hard boundaries), the road attributes within the polygonal contour (e.g., dry road surface, wet road surface, snow-covered road surface, etc.), and the height at the polygonal contour. The drivable area prediction model can be trained based on the training dataset.

[0156] Optionally, the information of the 3D drivable area may also include the road attributes of each polygon in the one or more polygons.

[0157] Optionally, the drivable area includes a first area, the boundary attributes of which include information on soft boundaries and / or hard boundaries. The soft boundary includes the boundary where the first area and the second area intersect. The road attributes of the first area indicate a first road surface type, and the road attributes of the second area indicate a second road surface type. The traffic priority of the first road surface type is higher than that of the second road surface type. The hard boundary includes the boundary where the first area intersects with at least one of the following: curb, flower bed, fence, ditch edge, and cliff edge.

[0158] Optionally, the data is input into the first prediction model to obtain information about the vehicle's three-dimensional drivable area, including: inputting the data into a general obstacle detection model to obtain features of multiple grids, wherein the features of each grid include one or more of the following: occupancy information, speed, direction, or visibility; and inputting the features of the multiple grids into the first prediction model to obtain information about the 3D drivable area.

[0159] The process of inputting the features of multiple grids into the first prediction model to obtain information about the 3D drivable area can be referred to the implementation process of the above embodiment, and will not be repeated here.

[0160] Optionally, before inputting the features of the multiple grids into the first prediction model to obtain the information of the 3D drivable domain, the method 1100 further includes: acquiring first information, the first information including information of a standard SD map and / or crowdsourced data; inputting the first information into a second prediction model to obtain map features; wherein, inputting the features of the multiple grids into the first prediction model to obtain the information of the 3D drivable domain includes: fusing the features of the multiple grids and the map features to obtain fused features; and inputting the fused features into the first prediction model to obtain the information of the 3D drivable area.

[0161] For example, the second prediction model can be the encoder shown in Figure 8 above.

[0162] The process of inputting the first information and the features of multiple grids into the first prediction model to obtain the information of the 3D drivable area can be referred to the implementation process shown in Figure 8 above, and will not be repeated here.

[0163] Optionally, inputting the features of the multiple grids into the first prediction model to obtain the information of the 3D drivable domain includes: inputting the features of the multiple grids into the first prediction model to obtain a 2.5D grid prediction result, the 2.5D grid prediction result including the attributes of each grid in the multiple two-dimensional 2D grids and the height information of the grids corresponding to the drivable area in the multiple 2D grids; and determining the information of the 3D drivable domain based on the 2.5D grid prediction result.

[0164] For example, Figure 13 shows another schematic diagram of the network architecture provided in an embodiment of this application. By inputting the data collected by sensors into an end-to-end model, control commands (e.g., steering commands, braking commands, etc.) can be obtained. This end-to-end model includes a prediction model 1 and a control module (e.g., a planning module and a control module). By inputting the data collected by sensors into the prediction model 1, features of multiple grids can be obtained, which can be used to characterize the drivable area. The prediction model 1 can input the features of multiple grids into the planning module and the control module, thereby enabling the end-to-end model to output control commands. The prediction model 1 can also output the features of multiple grids, which, after being decoded, can yield information about the drivable area. Thus, the vehicle 100 can display this drivable area information through a display device (e.g., an instrument panel or a central control screen).

[0165] The planning and control module in the above embodiments may include the planning module 220 and the control module 230 as shown in Figure 2.

[0166] Figure 14 shows a schematic block diagram of a drivable area detection device 1400 provided in an embodiment of this application. The device 1400 includes: an acquisition unit 1410 for acquiring data collected by sensors; a first prediction unit 1420 for inputting the data into a general obstacle detection model to obtain features of multiple grids, wherein the features of each grid include one or more of the following: occupancy information, speed, direction, or visibility; and a second prediction unit 1430 for inputting the features of the multiple grids into the first prediction model to obtain information about the drivable area of ​​the vehicle.

[0167] Optionally, the second prediction unit 1430 is specifically used to: input the features of the multiple grids and the first information into the first prediction model to obtain information about the drivable area, wherein the first information includes standard SD maps and / or crowdsourced data.

[0168] Optionally, the device 1400 further includes: a third prediction unit, used to input second information into a second prediction model to obtain map features, the second information including a standard SD map and / or crowdsourced data; a fusion unit, used to fuse the features of the multiple grids and the map features to obtain fused features; the second prediction unit 1430 is specifically used to: input the fused features into the first prediction model to obtain information about the drivable area.

[0169] Optionally, the second prediction unit 1430 is specifically used to: input the features of the plurality of grids into the first prediction model to obtain a 2.5D grid prediction result, the 2.5D grid prediction result including the attributes of each grid in the plurality of two-dimensional 2D grids and the height information of the grid corresponding to the drivable area in the plurality of 2D grids; wherein, the device 1400 further includes: a determination unit, used to determine the information of the drivable area based on the 2.5D grid prediction result.

[0170] Optionally, the information about the drivable area includes the boundary attributes and / or road attributes of the drivable area.

[0171] Optionally, the drivable area includes a first area, the boundary attributes of which include information on soft boundaries and / or hard boundaries. The soft boundary includes the boundary where the first area and the second area intersect. The road attributes of the first area indicate a first road surface type, and the road attributes of the second area indicate a second road surface type. The traffic priority of the first road surface type is higher than that of the second road surface type. The hard boundary includes the boundary where the first area intersects with at least one of the following: curb, flower bed, fence, ditch edge, and cliff edge.

[0172] Optionally, the device 1400 further includes a transmitting unit for transmitting information about the drivable area to the control module.

[0173] Optionally, the device 1400 further includes a control unit for controlling the vehicle's movement based on information about the drivable area.

[0174] Figure 15 shows a schematic block diagram of a drivable area detection device 1500 provided in an embodiment of this application. The device 1500 includes: an acquisition unit 1510 for acquiring data collected by sensors; and a prediction unit 1520 for inputting the data into a first prediction model to obtain information about the three-dimensional drivable area of ​​the vehicle. The information about the 3D drivable area includes the boundary attributes and height of one or more polygons in which the drivable area of ​​the vehicle is located.

[0175] Optionally, the information of the 3D drivable area may also include the road attributes of each polygon in the one or more polygons.

[0176] Optionally, the prediction unit 1520 is specifically used to: input the data into a general obstacle detection model to obtain features of multiple grids, wherein the features of each grid include one or more of the following: occupancy information, speed, direction, or visibility of each grid; and input the features of the multiple grids into the first prediction model to obtain information about the 3D drivable domain.

[0177] Optionally, the acquisition unit 1510 is further configured to acquire first information, which includes information from a standard SD map and / or crowdsourced data; the prediction unit 1520 is further configured to input the first information into a second prediction model to obtain map features; wherein, the device further includes: a fusion unit, configured to fuse the features of the multiple grids and the map features to obtain fused features; the prediction unit 1520 is specifically configured to: input the fused features into the first prediction model to obtain information about the 3D drivable area.

[0178] Optionally, the prediction unit 1520 is specifically used to: input the features of the plurality of grids into the first prediction model to obtain a 2.5D grid prediction result, the 2.5D grid prediction result including the attributes of each grid in the plurality of two-dimensional 2D grids and the height information of the grid corresponding to the drivable area in the plurality of 2D grids; wherein, the device 1500 further includes: a determination unit, used to determine the information of the 3D drivable domain based on the 2.5D grid prediction result.

[0179] The general obstacle detection model involved in this application embodiment can also be called a three-dimensional target detection device. For example, this three-dimensional target detection device can detect general obstacles using a three-dimensional target detection method. For instance, Figure 16 shows a schematic block diagram of the three-dimensional target detection device 1000 provided in this application embodiment. As shown in Figure 16, the scene includes an image acquisition device 102, a point cloud acquisition device 104, and a device 106. The image acquisition device 102 can specifically be a camera, including but not limited to a monocular camera, a multi-view camera, and a depth camera. The point cloud acquisition device 104 can specifically be a LiDAR, including but not limited to a single-line LiDAR and a multi-line LiDAR. The device 106 is a processing device, which has a CPU and / or a GPU, used to process the images acquired by the image acquisition device 102 and the point cloud data acquired by the point cloud acquisition device 104, thereby achieving three-dimensional target detection. It should be noted that the device 106 can be a physical device or a cluster of physical devices, such as a terminal, a server, or a server cluster. Of course, the device 106 can also be a virtualized cloud device, such as at least one cloud computing device in a cloud computing cluster.

[0180] In a specific implementation, image acquisition device 102 acquires images of the target environment, and point cloud acquisition device 104 acquires point cloud data of the target environment, such as images and point cloud data of the same road segment. Then, image acquisition device 102 sends the images to device 106, and point cloud acquisition device 104 sends the point cloud data to device 106. Device 106 is equipped with a 3D target detection device 1000, which includes a communication module 1001, a semantic extraction module 1003, and a target detection module 1005. Communication module 1001 acquires the images and the point cloud data.

[0181] The semantic extraction module 1003 is used to extract semantic information from the image. This semantic information may include, for example, image regions of different objects in the image and category information of the objects, such as pedestrians, vehicles, roads, and trees. Then, the target detection module 1005 can obtain the three-dimensional position information of the target in the target environment based on the point cloud data, the image, and the semantic information of the image.

[0182] For example, the three-dimensional position information of the target may include features of multiple grids from step S320 above.

[0183] In another implementation scenario, as shown in Figure 17, device 106 can specifically be a vehicle 200. Vehicle 200 can be equipped with one or more sensors, such as LiDAR, camera, Global Navigation Satellite System (GNSS), and IMU. In vehicle 200, the 3D target detection device 1000 can be installed in the onboard computer. The LiDAR and camera can transmit the acquired data to the onboard computer, and the 3D target detection device 1000 performs 3D target detection in the target environment. In one embodiment of this application, as shown in Figure 17, multiple image acquisition devices 102 can be installed on vehicle 200. Specific installation locations can include the front, rear, and sides of vehicle 200, to capture surround-view images from multiple angles around vehicle 200. This application does not limit the number or installation location of the image acquisition devices 102 on vehicle 200. After acquiring the 3D position information of targets in the target environment, vehicle 200 can perform route planning, obstacle avoidance, and other driving decisions. Of course, the three-dimensional target detection device 1000 can also be embedded in an in-vehicle computer, LiDAR, or camera in the form of a chip. Specifically, the chip can be a multi-domain controller (MDC), and this application embodiment does not impose any limitations on this.

[0184] Of course, in other applications, device 106 may also include robots, robotic arms, virtual reality devices, etc., and this application does not impose any restrictions here.

[0185] The three-dimensional target detection method described in this application will be described in detail below with reference to the accompanying drawings. Figure 18 is a schematic flowchart of an embodiment of the three-dimensional target detection method provided in this application. Although this application provides method operation steps as shown in the following embodiments or figures, more or fewer operation steps may be included in the method based on conventional or non-inventive effort. In steps where there is no logically necessary causal relationship, the execution order of these steps is not limited to the execution order provided in the embodiments of this application. In actual three-dimensional target detection processes or when the device executes the method, it can be executed in the order shown in the embodiments or figures or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0186] Specifically, one embodiment of the three-dimensional target detection method provided in this application is shown in Figure 18. The method may include:

[0187] S1801: Acquire images and point cloud data of the target environment.

[0188] In one embodiment of this application, an acquisition vehicle can be used to collect images and point cloud data of the target environment. The acquisition vehicle may be equipped with an image acquisition device 102 and a point cloud acquisition device 104, and may also include one or more other sensors such as GNSS and IMU to record information such as the time and location of image or point cloud data acquisition. Specifically, the image acquisition device 102 is mainly used to acquire images of targets such as pedestrians, vehicles, roads, and greenery in the target environment. The images may include any format such as BMP, JPEG, PNG, and SVG. The point cloud acquisition device 104 is mainly used to collect point cloud data of the target environment. Because point cloud acquisition devices such as LiDAR can accurately reflect location information, the point cloud acquisition device can acquire information such as the width of the road surface, the height of pedestrians, the width of vehicles, the height of traffic lights, and other information. GNSS can be used to record the coordinates of the currently acquired images and point cloud data. The IMU is mainly used to record the angle and acceleration information of the acquisition vehicle. It should be noted that the image may include multiple images of the target environment from different angles acquired at the same time, such as a 360-degree surround view image acquired by multiple image acquisition devices 102 on the vehicle 200. The image may also include a panoramic image stitched together from the multiple images from different angles; this is not limited. Furthermore, the image acquisition device 102 may capture images or obtain a video stream of the target environment. When the image acquisition device 102 captures a video stream, the 3D target detection device 1000 may decode the video stream to obtain several frames of images, and then obtain an image from these frames.

[0189] In the process of 3D target detection, it is necessary to jointly detect images and point cloud data at the same time and location. Therefore, the acquisition time of the images and the point cloud data is the same.

[0190] Of course, in other embodiments, images and point cloud data of the target environment can also be acquired using other acquisition devices. For example, the acquisition device may include a roadside acquisition device, which may be equipped with an image acquisition device 102 and a point cloud acquisition device 104. The acquisition device may also include a robot, on which the image acquisition device 102 and the point cloud acquisition device 104 may be installed. This application does not limit the acquisition device.

[0191] S1803: Obtain the semantic information of the image, the semantic information including the category information corresponding to the pixels in the image.

[0192] In this embodiment, semantic information of the image can be obtained. The image may include original information such as size and color values ​​of each pixel, such as RGB values ​​and grayscale values. Based on the original information of the image, the semantic information of the image can be determined to further obtain semantically related information in the image. The speech information may include category information corresponding to each pixel in the image, such as pedestrians, vehicles, roads, trees, buildings, etc.

[0193] In one embodiment of this application, the semantic information of the image can be obtained using panoramic segmentation. Specifically, the image can be panoramically segmented to generate semantic information, which includes a panoramic segmentation image of the image. This panoramic segmentation image includes image regions of different objects obtained from the panoramic segmentation and the corresponding category information of each image region. In an example shown in Figure 19A, the image can be input into a panoramic segmentation network, which outputs a panoramic segmentation image. As shown in Figure 19A, the panoramic segmentation image can include multiple image blocks, where image blocks of the same color indicate that the target objects belong to the same category. For example, image blocks 1 and 10 belong to buildings, image blocks 3, 5, and 9 belong to vehicles, image block 2 is the sky, image block 4 is the road surface, image block 6 is a person, and image blocks 7, 8, and 11 belong to trees. The panoramic segmentation network can be trained using an image sample set, which can include target objects of common categories.

[0194] In another embodiment, when the image includes a 360-degree panoramic view of the target environment, i.e., multiple images from different angles, as shown in FIG19B, panoramic segmentation can be performed on the multiple images respectively, and semantic information corresponding to the multiple images can be obtained.

[0195] Of course, in other embodiments, any algorithm capable of determining the semantic information of the image, such as semantic segmentation or instance segmentation, can be used, and this application embodiment does not impose any limitations.

[0196] S1805: Determine the three-dimensional position information of the target in the target environment based on the point cloud data, the image, and the semantic information of the image.

[0197] In this embodiment, after determining the semantic information of the image, the three-dimensional position information of the target in the target environment can be determined based on the point cloud data, the image, and the semantic information. Here, combining at least three types of data to determine the three-dimensional position information of the target in the target environment provides a rich information foundation and improves the accuracy of three-dimensional target detection. In one embodiment of this application, as shown in FIG20, determining the three-dimensional position information of the target in the target environment may specifically include:

[0198] S501: Project the image and the semantic information of the image into the point cloud data to generate semantic point cloud data.

[0199] In this embodiment, since the image and the point cloud data are acquired using different devices, their spatial coordinate systems are different. The image can be based on the coordinate system of the image acquisition device 102, such as a camera coordinate system, while the point cloud data can be based on the coordinate system of the point cloud acquisition device 104, such as a lidar coordinate system. Therefore, the image and the point cloud data can be unified into the same coordinate system. In one embodiment of this application, the image and the semantic information can be projected onto the point cloud data, that is, the image and the semantic information can be transformed into the coordinate system corresponding to the point cloud data. After projecting the image and the semantic information onto the point cloud data, semantic point cloud data can be generated.

[0200] In a specific embodiment, during the process of projecting the image and semantic information onto the point cloud data, the calibration extrinsic parameters of the image acquisition device 102 and the point cloud acquisition device 104 can be obtained, such as three rotation parameters and three translation parameters. Based on these calibration extrinsic parameters, the coordinate transformation matrix P from the image acquisition device 102 to the point cloud acquisition device 104 is determined. Then, for the original image X... RGB The converted image and the original semantic information X mask Semantic information obtained from the conversion They can be represented as follows:

[0201] Here, proj() represents the projection operation.

[0202] Determine the image after projection and the semantic information after projection After that, you can and With point cloud data X pointThe semantic point cloud data X is generated by stitching together the image, semantic information, and point cloud data. Figure 21 shows the effect of generating the semantic point cloud data. As shown in Figure 21, the semantic point cloud data generated after fusing the image, semantic information, and point cloud data contains richer information. Figure 19B also shows the effect when the image includes a 360-degree panoramic view. As shown in Figure 19B, coordinate transformations can be performed on each image in the 360-degree panoramic view and the semantic information corresponding to each image, projected onto the point cloud data, to generate the 360-degree panoramic semantic point cloud shown in Figure 19B, thereby obtaining more image information. Point cloud data X point The semantic point cloud data X can include the location information (x, y, z) and inverse color ratio r of each observation point. After projecting the image and the semantic information onto the point cloud data, each observation point in the semantic point cloud data X can include not only location information and inverse color ratio, but also color information and category semantic information, and can be represented as:

[0203] Among them, R {} Let N represent the set of real numbers, N represent the number of observation points in the semantic point cloud data, 4 represent the information contained in the point cloud data, namely the positional information (x, y, z) and the inverse color ratio r, 3 represent the information in the image, namely the RGB values, and C represent the information contained in the image. K This represents the semantic information of the observation point's category.

[0204] S503: Extract feature information from the semantic point cloud data to generate semantic point cloud feature information.

[0205] S505: Determine the three-dimensional position information of the target in the semantic point cloud data based on the semantic point cloud feature information.

[0206] In this embodiment, the semantic point cloud data includes not only point cloud data but also image information and corresponding semantic information. Based on this, the extracted semantic point cloud feature information also includes the feature information corresponding to each of the aforementioned data. Performing semantic point cloud feature extraction and target detection in two separate steps reduces the complexity of 3D target detection. Of course, in other embodiments, the 3D position information of the target in the semantic point cloud data can be directly determined from the semantic point cloud data; this embodiment does not impose such limitations.

[0207] In one embodiment of this application, as shown in FIG22, the semantic point cloud data can be input into a semantic point cloud feature recognition network 701, which outputs semantic point cloud feature information. Then, the semantic point cloud feature information can be input into a target detection network 703, which outputs the three-dimensional position information of the target in the semantic point cloud data. As shown in FIG22, the three-dimensional position information can be expressed using a three-dimensional frame in the semantic point cloud data. The three-dimensional frame includes at least the following information: the center point coordinates (x, y, z) of the frame, the dimensions (length, width, height) of the frame, and the heading angle θ. Using the information of the three-dimensional frame, the three-dimensional position information of the target framed by the three-dimensional frame in the target environment can be determined.

[0208] In this embodiment of the application, supervised machine learning training can be used during the training of the semantic point cloud feature recognition network 701 and the object detection network 703. Therefore, the results need to be labeled in the training samples. Since semantic point cloud feature information is unlabelable, while the 3D position information of the target is labelable, the semantic point cloud feature recognition network 701 and the object detection network 703 can be jointly trained. In one embodiment of this application, as shown in Figure 23, the semantic point cloud feature recognition network 701 and the object detection network 703 can be trained according to the following steps:

[0209] S801: Acquire multiple semantic point cloud training samples, wherein the semantic point cloud training samples include point cloud data samples, image samples projected onto the point cloud data samples, and the semantic information of the image samples, and the semantic point cloud training samples are labeled with the three-dimensional position information of the target.

[0210] S803: Construct a semantic point cloud feature recognition network 701 and an object detection network 703, wherein the output of the semantic point cloud feature recognition network 701 is connected to the input of the object detection network 703;

[0211] S805: Input the multiple semantic point cloud training samples into the semantic point cloud feature recognition network 701 respectively, and output the prediction result through the target detection network 703;

[0212] S807: Based on the difference between the prediction result and the labeled three-dimensional position information of the target, the network parameters of the semantic point cloud feature recognition network 701 and the target detection network 703 are iteratively adjusted until the iteration meets the preset requirements.

[0213] In one embodiment of this application, the semantic point cloud training samples may include actually collected data or existing datasets; this application does not impose any limitations. Before inputting the semantic point cloud training samples into the semantic point cloud feature recognition network 701, the semantic point cloud training samples can be divided into multiple unit data volumes with preset sizes. These unit data volumes may include voxels, point pillars, etc., thus converting irregular semantic point cloud training samples into regular data volumes, reducing the difficulty of subsequent data processing. The prediction result may include the predicted target of the semantic point cloud data and its three-dimensional position, as well as the probability that the predicted target is present at the three-dimensional position. The iteration meeting preset requirements may include the difference between the prediction result and the labeled three-dimensional position information of the target being less than a difference threshold, such as 0.01 or 0.05. The iteration meeting preset requirements may also include the number of iterations being greater than a preset number threshold, such as 50 or 60 iterations. The semantic point cloud feature recognition network 701 and the object detection network 703 may include CNN and various network modules based on CNN, such as AlexNet, ResNet, ResNet1001 (pre-activation), Hourglass, Inception, Xception, SENet, etc., which are not limited herein.

[0214] In this embodiment, after projecting the image and semantic information onto the point cloud data, the image and the point cloud data remain independent data, each possessing its own data features. Based on this, as shown in Figure 24, the semantic point cloud feature recognition network 701 can be divided into a point cloud feature recognition sub-network 901 and an image feature recognition sub-network 903.

[0215] The point cloud feature recognition subnetwork 901 is used to extract point cloud feature information from the point cloud data;

[0216] The image feature recognition subnetwork 903 is used to extract image feature information from the image based on the image and the semantic information, and to dynamically adjust the network parameters of the point cloud feature recognition subnetwork using the image feature information.

[0217] Specifically, the point cloud data can be input into a point cloud feature recognition subnetwork 901, which then outputs point cloud feature information. Of course, before inputting into the point cloud feature recognition subnetwork 901, the irregular point cloud data can be divided into regularly sized unit data volumes, such as voxels or point cloud pillars. On the other hand, the image and its semantic information can be input into an image feature recognition subnetwork 903, which then outputs image feature information. Then, the network parameters of the point cloud feature recognition subnetwork can be dynamically adjusted using the image feature information. Thus, the point cloud feature information output by the dynamically adjusted point cloud feature recognition subnetwork 901 is the semantic point cloud feature information. In embodiments of this application, the network parameters can include any adjustable parameters in a neural network, such as convolution kernel parameters and batch normalization parameters. In a specific example, the generated semantic point cloud feature information y can be represented as: W = G(f img )

[0218] Where W represents the convolution kernel parameter, X point Represents point cloud data. G() represents the convolution operation, G() represents the convolution kernel generation module, and f img This represents the image feature information.

[0219] In other words, during the process of adjusting the network parameters of the point cloud feature recognition subnetwork 901 using the image feature information, it is necessary to use the convolution kernel generation module to generate the corresponding convolution kernel parameters. The convolution kernel generation module is a function represented by the neural network. In a specific example, the convolution kernel generation module may include: G(f img )=W1δ(W2f img )

[0220] Wherein, W1 and W2 represent the parameters in the convolution kernel generation module, and δ represents the activation function (such as ReLU). Of course, the result generated by the convolution kernel generation module may not match the convolution kernel size of the point cloud feature recognition sub-network 901. Based on this, the convolution kernel generation module can also generate parameter predictions that match the size of the convolution kernel of the point cloud feature recognition sub-network 901.

[0221] In this embodiment of the application, in the above-described method of dynamically adjusting network parameters using the image feature information, the optimization algorithm (such as the gradient of the loss function) corresponding to the point cloud feature recognition sub-network 901 can not only be used to adjust the parameters such as the convolution kernel of the point cloud feature recognition sub-network 901, but also to adjust the parameters in the convolution kernel generation function, so that the point cloud feature recognition sub-network 901 integrates the features of point cloud data and the features of the image, thereby improving the detection accuracy of the point cloud feature recognition sub-network 901.

[0222] In practical applications, neural networks often consist of multiple network layers. As shown in Figure 24, the point cloud feature recognition subnetwork 901 in this embodiment may include N network layers, where N ≥ 2. Based on this, in one embodiment of this application, the image feature recognition subnetwork 903 can be connected to each network layer, and the network parameters of each network layer can be dynamically adjusted using the image feature information output by the image feature recognition subnetwork 903. In this way, the influence of the image feature information on point cloud feature extraction can be gradually enhanced as the number of network layers increases. Similarly, the network parameters may include any adjustable parameters in the neural network, such as convolution kernel parameters and batch normalization parameters.

[0223] In one embodiment of this application, the network parameters may further include attention mechanism parameters. That is, as shown in Figure 25, an attention mechanism can be added between the point cloud feature recognition subnetwork 901 and the image feature recognition subnetwork 903. Specifically, information in the image feature information that is highly correlated with the point cloud data can be used to adjust the network parameters of the point cloud feature recognition subnetwork 901. More specifically, information in the image feature information whose correlation with the point cloud data is greater than a correlation threshold can be used as effective information for adjusting the point cloud feature recognition subnetwork. The correlation threshold can be automatically generated based on the model. In one embodiment of this application, the attention mechanism parameters can be determined based on the correlation between the image feature information and the point cloud data. In a specific example, the generated semantic point cloud feature information y can be represented as: Attention = γ * δ(X) point ,f img ) δ(X point ,f rgb )=ψ(f img ) T β(X point )

[0224] Where Attention⊙ represents the attention function, W represents the convolution kernel parameters, and X represents the convolution kernel parameters. point Represents point cloud data. This represents the convolution operation, γ represents the parameters of the attention function, and f img The image feature information is represented by δ(), which may include dot product operations.

[0225] Of course, the attention parameters are not limited to the examples above, and can also be determined using functions capable of implementing attention mechanisms; this application does not impose any limitations here. In the embodiments of this application, using attention parameters together with other network parameters to determine the semantic point cloud feature information can further enhance the amount of useful information in the semantic point cloud feature information and improve the accuracy of 3D target detection.

[0226] In this embodiment, the output results of each network layer can also be obtained to determine the contribution of the image feature information to each network layer. Based on this, in one embodiment, the output data of at least one network layer can be obtained, and the adjustment effect data corresponding to the network parameters can be determined based on the output data. The network parameters may include convolution kernel parameters, regularization parameters, attention mechanism parameters, etc. The output data may include the detection results of the target in the point cloud data, and the corresponding adjustment effect data may include the differences between the output results of each network layer and the differences between the output results and the final output results. Furthermore, users can view the output results of each network layer and intuitively understand the adjustment effect of each network layer. In this embodiment, compared to the black-box prediction of ordinary neural network layers, dynamically adjusted parameters can enhance the interpretability of network prediction.

[0227] The three-dimensional target detection method provided by this application has been described in detail above with reference to Figures 16 to 25. The three-dimensional target detection device 1000 and equipment 106 provided by this application will be described below with reference to the accompanying drawings.

[0228] Referring to the structural schematic diagram of the three-dimensional target detection device 1000 in the system architecture diagram shown in Figure 16, as shown in Figure 16, the device 1000 includes:

[0229] Communication module 1001 is used to acquire images and point cloud data of the target environment;

[0230] The semantic extraction module 1003 is used to obtain the semantic information of the image, the semantic information including the category information corresponding to the pixels in the image;

[0231] The target detection module 1005 is used to determine the three-dimensional position information of the target in the target environment based on the point cloud data, the image, and the semantic information of the image.

[0232] Optionally, in one embodiment of this application, the target detection module 1005 is specifically used for:

[0233] The image and its semantic information are projected onto the point cloud data to generate semantic point cloud data;

[0234] Extract feature information from the semantic point cloud data to generate semantic point cloud feature information;

[0235] The three-dimensional position information of the target in the semantic point cloud data is determined based on the semantic point cloud feature information.

[0236] Optionally, in one embodiment of this application, the semantic point cloud feature information is output through a semantic point cloud feature recognition network, and the three-dimensional position information is output through a target detection network.

[0237] Optionally, in one embodiment of this application, the semantic point cloud feature recognition network includes a point cloud feature recognition sub-network and an image feature recognition sub-network, wherein,

[0238] The point cloud feature recognition subnetwork is used to extract point cloud feature information from the point cloud data;

[0239] The image feature recognition subnetwork is used to extract image feature information from the image based on the image and the semantic information, and to dynamically adjust the network parameters of the point cloud feature recognition subnetwork using the image feature information.

[0240] Optionally, in one embodiment of this application, the point cloud feature recognition subnetwork includes at least one network layer, and the image feature recognition subnetwork is connected to the at least one network layer respectively, wherein,

[0241] The image feature recognition subnetwork is specifically used to extract image feature information from the image based on the image and the semantic information, and to dynamically adjust the network parameters of at least one network layer using the image feature information.

[0242] Optionally, in one embodiment of this application, the network parameters include convolution kernel parameters and / or attention mechanism parameters, wherein the attention mechanism parameters are used to use information in the image feature information that has a correlation greater than a correlation threshold with the point cloud data as effective information for adjusting the point cloud feature recognition subnetwork.

[0243] Optionally, in one embodiment of this application, it further includes:

[0244] The output module is used to acquire the output data of the at least one network layer respectively;

[0245] The effect determination module is used to determine the adjustment effect data corresponding to the network parameters based on the output data.

[0246] Optionally, in one embodiment of this application, the semantic point cloud feature recognition network and the object detection network are trained in the following manner:

[0247] Multiple semantic point cloud training samples are obtained. The semantic point cloud training samples include point cloud data samples, image samples projected onto the point cloud data samples, and the semantic information of the image samples. The semantic point cloud training samples are labeled with the three-dimensional position information of the target.

[0248] Construct a semantic point cloud feature recognition network and an object detection network, wherein the output of the semantic point cloud feature recognition network is connected to the input of the object detection network;

[0249] The multiple semantic point cloud training samples are respectively input into the semantic point cloud feature recognition network, and the target detection network outputs the prediction result;

[0250] Based on the difference between the prediction result and the labeled 3D location information of the target, the network parameters of the semantic point cloud feature recognition network and the target detection network are iteratively adjusted until the iteration meets the preset requirements.

[0251] Optionally, in one embodiment of this application, the semantic extraction module 1003 is specifically used for:

[0252] The image is subjected to panoramic segmentation to generate semantic information of the image. The semantic information includes a panoramic segmentation image of the image, which includes image regions of different objects obtained from panoramic segmentation and category information corresponding to the image regions.

[0253] Optionally, in one embodiment of this application, the image includes a panoramic image.

[0254] The three-dimensional target detection device 1000 according to the embodiments of this application can be used to execute the methods described in the embodiments of this application. The above and other operations and / or functions of each module in the three-dimensional target detection device 1000 are respectively to implement the corresponding processes of each method in FIG18, FIG20 and FIG23. For the sake of brevity, they will not be described in detail here.

[0255] It should be understood that the division of units in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units in the device can be implemented by a processor calling software; for example, the device includes a processor connected to memory, which stores instructions. The processor calls the instructions stored in memory to implement any of the above methods or to implement the functions of each unit in the device. The processor can be, for example, a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal or external to the device. Alternatively, the units in the device can be implemented as hardware circuits. The functions of some or all units can be implemented through the design of the hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all units are implemented through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby implementing the functions of some or all units. All units of the above device can be implemented entirely through processor calling software, or entirely through hardware circuits, or partially through processor calling software and the remaining parts through hardware circuits.

[0256] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above units. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.

[0257] As can be seen, each unit in the above device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0258] Furthermore, the units in the above devices can be integrated in whole or in part, or they can be implemented independently. In one implementation, these units are integrated together as a System-on-a-Chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the units in the device. The at least one processor may be of different types, such as CPU and FPGA, CPU and AI processor, CPU and GPU, etc.

[0259] This application also provides a drivable area detection device, which includes a processing unit and a storage unit. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to cause the device to perform the methods or steps described in the above embodiments.

[0260] Optionally, if the drivable area detection device is located in the vehicle, the processing unit may be one or more of the processors 121-12n shown in FIG1.

[0261] This application also provides a drivable area detection system, which includes a sensing system and a computing platform, the computing platform including the aforementioned drivable area detection device 1400 or 1500.

[0262] This application also provides a vehicle that may include the aforementioned drivable area detection device 1400 or 1500, or the aforementioned drivable area detection system.

[0263] This application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to perform the methods described in the above embodiments.

[0264] This application also provides a computer-readable medium storing program code that, when run on a computer, causes the computer to perform the methods described in the above embodiments.

[0265] This application also provides a chip, which includes a circuit for performing the methods described in the above embodiments.

[0266] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, power-on erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.

[0267] It should be understood that in the embodiments of this application, the memory may include read-only memory and random access memory, and provides instructions and data to the processor.

[0268] It should also be understood that, in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0269] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0270] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0271] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0272] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0273] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0274] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0275] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be covered. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for detecting drivable areas, characterized in that, include: Acquire data collected by the sensors; The data is input into a general obstacle detection model to obtain the features of multiple grids. The features of each grid include one or more of the following: occupancy information, speed, direction, or visibility. The features of the multiple grids are input into the first prediction model to obtain information about the drivable area of ​​the vehicle.

2. The method according to claim 1, characterized in that, The step of inputting the features of the multiple grids into the first prediction model to obtain information about the vehicle's drivable area includes: The features of the multiple grids and the first information are input into the first prediction model to obtain the information of the drivable area. The first information includes standard SD maps and / or crowdsourced data.

3. The method according to claim 1, characterized in that, Before inputting the features of the multiple grids into the first prediction model to obtain information about the vehicle's drivable area, the method further includes: The second information is input into the second prediction model to obtain map features, wherein the second information includes standard SD maps and / or crowdsourced data; The step of inputting the features of the multiple grids into the first prediction model to obtain information about the vehicle's drivable area includes: The features of the multiple grid cells and the map features are fused to obtain the fused features; The fused features are input into the first prediction model to obtain information about the drivable area.

4. The method according to any one of claims 1 to 3, characterized in that, The step of inputting the features of the multiple grids into the first prediction model to obtain information about the vehicle's drivable area includes: The features of the multiple grids are input into the first prediction model to obtain a 2.5D grid prediction result. The 2.5D grid prediction result includes the attributes of each grid in the multiple two-dimensional 2D grids and the height information of the grid corresponding to the drivable area in the multiple 2D grids. Based on the 2.5D grid prediction results, information about the drivable area is determined.

5. The method according to any one of claims 1 to 4, characterized in that, The information about the drivable area includes the boundary attributes and / or road attributes of the drivable area.

6. The method according to claim 5, characterized in that, The drivable area includes a first area, the boundary attributes of which include information about soft boundaries and / or hard boundaries. The soft boundary includes the boundary where the first region and the second region intersect. The road attributes of the first region indicate a first road surface type, and the road attributes of the second region indicate a second road surface type. The traffic priority of the first road surface type is higher than that of the second road surface type. The hard boundary includes the boundary where the first region intersects with at least one of the following: curb, flower bed, fence, ditch edge, and cliff edge.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Send the information about the drivable area to the planning and control module.

8. The method according to any one of claims 1 to 6, characterized in that, The method further includes: The vehicle is controlled to move based on information about the drivable area.

9. The method according to any one of claims 1 to 8, characterized in that, The features of the multiple grids are input into the first prediction model to obtain information about the vehicle's drivable area, including: The features and task cues of the multiple grids are input into the Visual Language Model (VLM) to obtain scene understanding information related to the drivable area.

10. A method for detecting drivable areas, characterized in that, include: Acquire data collected by the sensors; The data is input into the first prediction model to obtain information about the vehicle's three-dimensional drivable area. The information about the 3D drivable area includes the boundary attributes and height of one or more polygons in which the vehicle's drivable area is located.

11. The method according to claim 10, characterized in that, The information of the 3D drivable area also includes the road attributes of each polygon in the one or more polygons.

12. The method according to claim 10 or 11, characterized in that, The step of inputting the data into the first prediction model to obtain information about the vehicle's three-dimensional drivable area includes: The data is input into a general obstacle detection model to obtain the features of multiple grids. The features of each grid include one or more of the following: occupancy information, velocity, orientation, or visibility. The features of the multiple grids are input into the first prediction model to obtain the information of the 3D drivable domain.

13. The method according to claim 12, characterized in that, Before inputting the features of the plurality of grids into the first prediction model to obtain the information of the 3D drivable domain, the method further includes: Obtain first information, which includes information from standard SD maps and / or crowdsourced data; The first piece of information is input into the second prediction model to obtain map features; The step of inputting the features of the plurality of grids into the first prediction model to obtain the information of the 3D drivable domain includes: The features of the multiple grid cells and the map features are fused to obtain the fused features; The fused features are input into the first prediction model to obtain information about the 3D drivable area.

14. The method according to claim 12 or 13, characterized in that, The step of inputting the features of the multiple grids into the first prediction model to obtain the information of the 3D drivable domain includes: The features of the multiple grids are input into the first prediction model to obtain 2.5D grid features. The 2.5D grid features include the features of each grid in the multiple two-dimensional 2D grids and the height information of the grids corresponding to the drivable area in the multiple 2D grids. Based on the 2.5D grid features, the information of the 3D drivable domain is determined.

15. A drivable area detection device, characterized in that, include: The acquisition unit is used to acquire data collected by the sensor; The first prediction unit is used to input the data into a general obstacle detection model to obtain the features of multiple grids. The features of each grid include one or more of the following: occupancy information, speed, direction, or visibility of each grid. The second prediction unit is used to input the features of the multiple grids into the first prediction model to obtain information about the drivable area of ​​the vehicle.

16. The apparatus according to claim 15, characterized in that, The second prediction unit is specifically used for: The features of the multiple grids and the first information are input into the first prediction model to obtain the information of the drivable area. The first information includes standard SD maps and / or crowdsourced data.

17. The apparatus according to claim 15, characterized in that, The device further includes: The third prediction unit is used to input the second information into the second prediction model to obtain map features, wherein the second information includes standard SD maps and / or crowdsourced data; The fusion unit is used to fuse the features of the multiple grids and the map features to obtain the fused features; The second prediction unit is specifically used to: input the fused features into the first prediction model to obtain information about the drivable area.

18. The apparatus according to any one of claims 15 to 17, characterized in that, The second prediction unit is specifically used for: The features of the multiple grids are input into the first prediction model to obtain a 2.5D grid prediction result. The 2.5D grid prediction result includes the attributes of each grid in the multiple two-dimensional 2D grids and the height information of the grid corresponding to the drivable area in the multiple 2D grids. The device further includes: The determining unit is used to determine the information of the drivable area based on the 2.5D grid prediction results.

19. The apparatus according to any one of claims 15 to 18, characterized in that, The information about the drivable area includes the boundary attributes and / or road attributes of the drivable area.

20. The apparatus according to claim 19, characterized in that, The drivable area includes a first area, the boundary attributes of which include information about soft boundaries and / or hard boundaries. The soft boundary includes the boundary where the first region and the second region intersect. The road attributes of the first region indicate a first road surface type, and the road attributes of the second region indicate a second road surface type. The traffic priority of the first road surface type is higher than that of the second road surface type. The hard boundary includes the boundary where the first region intersects with at least one of the following: curb, flower bed, fence, ditch edge, and cliff edge.

21. The apparatus according to any one of claims 15 to 20, characterized in that, The device further includes: The sending unit is used to send information about the drivable area to the planning and control module.

22. The apparatus according to any one of claims 15 to 20, characterized in that, The device further includes: The control unit is used to control the movement of the vehicle based on information about the drivable area.

23. The apparatus according to any one of claims 15 to 22, characterized in that, The second prediction unit is specifically used for: The features and task cues of the multiple grids are input into the VLM to obtain scene understanding information related to the drivable area.

24. A drivable area detection device, characterized in that, include: The acquisition unit is used to acquire data collected by the sensor; The prediction unit is used to input the data into the first prediction model to obtain information about the three-dimensional drivable area of ​​the vehicle. The information about the 3D drivable area includes the boundary attributes and height of one or more polygons in which the drivable area of ​​the vehicle is located.

25. The apparatus according to claim 24, characterized in that, The information of the 3D drivable area also includes the road attributes of each polygon in the one or more polygons.

26. The apparatus according to claim 24 or 25, characterized in that, The prediction unit is specifically used for: The data is input into a general obstacle detection model to obtain the features of multiple grids. The features of each grid include one or more of the following: occupancy information, velocity, orientation, or visibility. The features of the multiple grids are input into the first prediction model to obtain the information of the 3D drivable domain.

27. The apparatus according to claim 26, characterized in that, The acquisition unit is further configured to acquire first information, the first information including information from standard SD maps and / or crowdsourced data; The prediction unit is also used to input the first information into the second prediction model to obtain map features; The device further includes a fusion unit, used to fuse the features of the plurality of grids and the map features to obtain fused features; The prediction unit is specifically used to: input the fused features into the first prediction model to obtain information about the 3D drivable area.

28. The apparatus according to claim 26 or 27, characterized in that, The prediction unit is specifically used for: The features of the multiple grids are input into the first prediction model to obtain a 2.5D grid prediction result. The 2.5D grid prediction result includes the attributes of each grid in the multiple two-dimensional 2D grids and the height information of the grid corresponding to the drivable area in the multiple 2D grids. The device further includes: The determining unit is used to determine the information of the 3D drivable domain based on the 2.5D grid prediction results.

29. A drivable area detection device, characterized in that, include: A processor for executing a computer program stored in memory to cause the apparatus to perform the method as described in any one of claims 1 to 14.

30. The apparatus according to claim 29, characterized in that, The device also includes the memory.

31. A drivable area detection system, characterized in that, The drivable area detection system includes a sensing system and a computing platform, wherein the computing platform includes the apparatus as described in any one of claims 15 to 30.

32. A vehicle, characterized in that, Includes the apparatus as described in any one of claims 15 to 30, or includes the system as described in claim 31.

33. A computer-readable storage medium, characterized in that, It stores instructions that, when executed by a processor, cause the processor to implement the method as described in any one of claims 1 to 14.

34. A computer program product, characterized in that, The computer program product includes computer program code that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 14.

35. A chip, characterized in that, The chip includes circuitry for performing the method as described in any one of claims 1 to 14.

Citation Information

Patent Citations

  • Trafficability identification method, system, equipment and computer readable storage medium

    CN112166446A

  • Acquisition method and device of drivable area and vehicle

    CN113420687A

  • Vehicle, method for a vehicle, and storage medium

    CN115082914A

  • Unmanned vehicle passable area prediction method and system based on multi-frame point cloud fusion

    CN116310681A

  • Automatic guided vehicle navigation method and device based on visual language model

    CN118258406A