Road element recognition method and apparatus
By acquiring and fusing geometric and semantic information of road elements and using neural network processing, the shortcomings of high-precision maps in autonomous driving are addressed, enabling accurate identification of road elements in complex environments and improving the safety and efficiency of autonomous driving.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-06-20
- Publication Date
- 2026-06-04
AI Technical Summary
Current autonomous driving technology relies on high-precision maps, which are costly and time-consuming to collect and produce, have insufficient coverage, and are difficult to guarantee data freshness. This makes it difficult for vehicles to accurately identify road elements in complex and ever-changing traffic environments, affecting driving safety and traffic efficiency.
By acquiring geometric and semantic information of road elements and using neural networks for fusion processing, the accuracy of semantic information recognition is improved, the impact of geometric misdetection on semantic information is reduced, the focus is on the semantic information recognition task, and the recognition results are optimized by combining the distance matrix.
Without relying on high-precision maps, it improves the accuracy of semantic information recognition of road elements, ensures safe passage of vehicles in complex traffic environments, reduces the risk of violations, and improves driving safety and traffic efficiency.
Smart Images

Figure CN2025102303_04062026_PF_FP_ABST
Abstract
Description
Methods and apparatus for identifying road elements
[0001] This application claims priority to Chinese Patent Application No. 202411751480.5, filed on November 29, 2024, entitled "Method and Apparatus for Identifying Road Elements", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of intelligent driving, and more specifically, to a method and apparatus for identifying road elements. Background Technology
[0003] With the rapid development of the automotive industry, many driver assistance and autonomous driving technologies have emerged, which can reduce driving stress and improve safety and traffic efficiency. Currently, most autonomous driving technologies rely on high-precision maps for navigation. However, high-precision maps have drawbacks such as high collection and production costs, long processing times, insufficient coverage, and difficulty in ensuring data freshness, making it difficult to promote autonomous driving technologies that rely on high-precision maps nationwide or globally.
[0004] Intelligent vehicles use sensors such as cameras and LiDAR to detect static elements (such as lanes, roads, traffic lights, and signs) within a certain range around the vehicle in real time. This allows them to model the surrounding environment, enabling vehicles to navigate smoothly in complex and ever-changing traffic conditions without relying on high-precision maps. However, the current sensing range of vehicle sensors is limited, especially in congested scenarios or when erroneous perception occurs. This can lead to some road elements not being accurately identified, significantly impacting driving safety and traffic efficiency.
[0005] Therefore, a road element detection scheme that can improve the accuracy of road element recognition urgently needs to be developed. Summary of the Invention
[0006] This application provides a method and apparatus for identifying road elements, which can improve the accuracy of semantic information recognition results of road elements, thereby improving the reliability of intelligent driving systems and vehicle driving safety.
[0007] Firstly, a method for identifying road elements is provided, which can be performed by a vehicle, for example, by the vehicle's computing platform, or by chips or circuits used in the vehicle. The method can also be performed by a cloud server communicating with the vehicle, or by its chips or circuits.
[0008] The method includes: acquiring a first set of data, which indicates the geometric information of M road elements in a first road; acquiring a second set of data, which indicates the semantic and geometric information of N road elements in the first road; and determining the semantic information of the road element to be identified in the first road based on the first set of data and the second set of data; wherein the road element to be identified is associated with M road elements and N road elements, and the M road elements and N road elements are associated, and M and N are both positive integers.
[0009] For intelligent vehicles, semantic information refers to information that enables vehicles to better understand driving rules, perceive road traffic conditions, and plan driving routes. Generally, assigning semantic information to the geometric features of road elements allows vehicles to plan more suitable and accurate driving paths based on those road elements. However, if semantic features are obtained simultaneously with geometric features, errors in semantic information recognition may occur in scenarios with geometric segmentation biases. In the above technical solution, the semantic information of the road element to be identified is determined based on its geometric information. This allows the recognition task to focus on semantic information recognition during road element identification, thereby improving the accuracy of semantic information recognition results. Furthermore, when errors occur in the detection results of geometric information, the impact of geometric misdetection on the semantic information recognition results can be mitigated, thus improving the accuracy of semantic information recognition results for road elements.
[0010] In conjunction with the first aspect, in certain implementations of the first aspect, determining the semantic information of the road element to be identified based on the first set of data and the second set of data includes: inputting the first set of data into a first type of neural network to obtain a first set of feature information, the first set of feature information indicating the geometric features of M road elements in the first road; inputting the second set of data into a second type of neural network to obtain a second set of feature information, the second set of feature information indicating the geometric and semantic features of N road elements in the first road; inputting the first set of feature information, the second set of feature information, and the position information of the road element to be identified into a fusion neural network to obtain fused feature information, the fused feature information indicating one or more initial semantic features of the road element to be identified; and determining the semantic information based on the fused feature information.
[0011] In the above technical solution, the first type of neural network focuses on extracting geometric features, which helps improve the accuracy of geometric features; the second type of neural network is used to extract the geometric and semantic features of road elements. The geometric features extracted by the first type of neural network can correct the accuracy of the geometric features extracted by the second type of neural network, thereby reducing the impact of geometric misdetection on the semantic recognition results. Furthermore, since the aforementioned first type of neural network focuses on extracting geometric features, the fusion neural network can focus on extracting semantic features, thereby improving the accuracy of the semantic information recognition results.
[0012] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: determining a distance matrix based on the first set of data and the location information of the road element to be identified, wherein the distance matrix indicates the correlation between each of the M road elements and the road element to be identified; determining the semantic information of the road element to be identified in the first road, including: determining the semantic information based on the first set of data, the second set of data and the distance matrix.
[0013] In the above technical solution, based on the distance matrix, when determining the semantic information of the road element to be identified, road elements with a higher correlation with the road element to be identified can be used, thereby improving the accuracy of the semantic information recognition result.
[0014] In conjunction with the first aspect, in some implementations of the first aspect, the first road is the road where the vehicle is currently located, the first set of data includes pixels of the first road element and / or point cloud data of the second road element, the M road elements include the first road element and the second road element, and obtaining the first set of data includes: obtaining first perception information collected by the vehicle's perception system, the first perception information including the image and / or point cloud data of the first road; processing the first perception information to obtain the pixels of the first road element and / or the point cloud data of the second road element.
[0015] In the above technical solution, the perception information collected in real time by the vehicle's perception system can determine the geometric features of road elements in real time. When the geometric features of the road change in reality, it can also obtain semantic information of road elements that are more consistent with reality.
[0016] In conjunction with the first aspect, in some implementations of the first aspect, the first road is the road where the vehicle is currently located, and the second set of data includes images and / or point cloud data of the first road collected by the vehicle's perception system.
[0017] In the above technical solution, the perception information collected in real time by the vehicle's perception system can determine the semantic information of road elements in real time, and when the road changes in reality, it can obtain the semantic information of road elements that is more consistent with reality.
[0018] In conjunction with the first aspect, in some implementations of the first aspect, the road element to be identified is a first lane line in a first road, and semantic information indicates the type of the first lane; the first set of data includes at least data indicating the location of a boundary of the first lane, and the second set of data includes at least one of the following: data indicating traffic signs associated with the first lane, traffic flow data associated with the first lane, or lane line data of the first lane; or, the first set of data includes data indicating the location of the boundary of at least one lane in the first road, at least one lane is associated with the first lane, and the second set of data includes at least one of the following: traffic flow data associated with each lane in the at least one lane, or lane line data of each lane in the at least one lane.
[0019] In the above technical solution, determining the semantic information of road elements based on multiple types of data helps to improve the richness and accuracy of semantic information, obtain detailed lane type-related semantic information, facilitate vehicle planning of more suitable routes, reduce the probability of violations, and improve vehicle driving safety.
[0020] In conjunction with the first aspect, in some implementations of the first aspect, the road element to be identified is a traffic sign in the first road, and the semantic information indicates the meaning of the traffic sign; the first set of data includes data indicating the location of the traffic sign, and the second set of data includes at least one of the following: data indicating the boundary of the first road, or data indicating the lane associated with the traffic sign.
[0021] In conjunction with the first aspect, in some implementations of the first aspect, the road element to be identified is an obstacle in the first road, and semantic information indicates the type of obstacle; the first set of data includes data indicating the position and / or position changes of the obstacle, and the second set of data includes at least one of the following: data indicating the boundary of the first road, or data indicating the lane associated with the obstacle.
[0022] Secondly, a control method is provided that can be executed by a vehicle, for example, by the vehicle's computing platform, or by a chip or circuitry for the vehicle.
[0023] The method includes controlling the vehicle to drive based on semantic information of the road element to be identified in any possible implementation of the first aspect.
[0024] The above technical solutions enable vehicles to plan driving routes that comply with traffic regulations without relying on high-precision maps, thus allowing them to drive smoothly in complex and ever-changing traffic environments.
[0025] Thirdly, an apparatus for identifying road elements is provided. The apparatus includes an acquisition unit and a processing unit. The acquisition unit is configured to: acquire a first set of data, which indicates the geometric information of M road elements in a first road; the acquisition unit is further configured to: acquire a second set of data, which indicates the semantic and geometric information of N road elements in the first road; the processing unit is configured to: determine the semantic information of the road element to be identified in the first road based on the first set of data and the second set of data; wherein the road element to be identified is associated with M road elements and N road elements, and the M road elements and N road elements are associated, and M and N are both positive integers.
[0026] In conjunction with the third aspect, in some implementations of the third aspect, the processing unit is used to: input a first set of data into a first type of neural network to obtain a first set of feature information, the first set of feature information indicating the geometric features of M road elements in the first road; input a second set of data into a second type of neural network to obtain a second set of feature information, the second set of feature information indicating the geometric and semantic features of N road elements in the first road; input the first set of feature information, the second set of feature information, and the position information of the road element to be identified into a fusion neural network to obtain fused feature information, the fused feature information indicating one or more initial semantic features of the road element to be identified; and determine semantic information based on the fused feature information.
[0027] In conjunction with the third aspect, in some implementations of the third aspect, the processing unit is used to: determine a distance matrix based on the first set of data and the location information of the road element to be identified, wherein the distance matrix indicates the degree of association between each of the M road elements and the road element to be identified; and determine semantic information based on the first set of data, the second set of data, and the distance matrix.
[0028] In conjunction with the third aspect, in some implementations of the third aspect, the first road is the road where the vehicle is currently located, the first set of data includes the pixels of the first road element and / or the point cloud data of the second road element, the M road elements include the first road element and the second road element, and the acquisition unit is used to: acquire the first perception information collected by the vehicle's perception system, the first perception information including the image and / or point cloud data of the first road; process the first perception information to obtain the pixels of the first road element and / or the point cloud data of the second road element.
[0029] In conjunction with the third aspect, in some implementations of the third aspect, the first road is the road where the vehicle is currently located, and the second set of data includes images and / or point cloud data of the first road collected by the vehicle's perception system.
[0030] In conjunction with the third aspect, in some implementations of the third aspect, the road element to be identified is the first lane line in the first road, and the semantic information indicates the type of the first lane; the first set of data includes at least data indicating the location of a boundary of the first lane, and the second set of data includes at least one of the following: data indicating traffic signs associated with the first lane, traffic flow data associated with the first lane, or lane line data of the first lane; or, the first set of data includes data indicating the location of the boundary of at least one lane in the first road, at least one lane is associated with the first lane, and the second set of data includes at least one of the following: traffic flow data associated with each lane in the at least one lane, or lane line data of each lane in the at least one lane.
[0031] In conjunction with the third aspect, in some implementations of the third aspect, the road element to be identified is a traffic sign in the first road, and the semantic information indicates the meaning of the traffic sign; the first set of data includes data indicating the location of the traffic sign, and the second set of data includes at least one of the following: data indicating the boundary of the first road, or data indicating the lane associated with the traffic sign.
[0032] In conjunction with the third aspect, in some implementations of the third aspect, the road element to be identified is an obstacle in the first road, and semantic information indicates the type of obstacle; the first set of data includes data indicating the position and / or position changes of the obstacle, and the second set of data includes at least one of the following: data indicating the boundary of the first road, or data indicating the lane associated with the obstacle.
[0033] Fourthly, a control device is provided, the device including a processing unit for: controlling vehicle movement based on semantic information of a road element to be identified in any possible implementation of the first aspect.
[0034] Fifthly, an apparatus for identifying road elements is provided, the apparatus including a processor for executing a computer program stored in the memory, such that the apparatus performs the method in any possible implementation of the first aspect described above.
[0035] In a sixth aspect, a control device is provided, the device including a processor for executing a computer program stored in the memory, such that the device performs the method in any possible implementation of the second aspect described above.
[0036] In conjunction with the fifth or sixth aspect, in some implementations of the fifth or sixth aspect, the device also includes a memory.
[0037] In a seventh aspect, a computer program product is provided, comprising: computer program code, which, when executed on a computer or processor, causes the computer or processor to perform the method in any possible implementation of the first or second aspect.
[0038] It should be noted that the above computer program code can be stored in whole or in part on a storage medium, which can be packaged together with the processor or packaged separately from the processor.
[0039] Eighthly, a computer-readable storage medium is provided, the computer-readable medium storing instructions that, when executed by a processor, cause the processor to implement the method in any possible implementation of the first or second aspect.
[0040] Ninthly, a chip is provided, the chip including circuitry for performing the methods in any possible implementation of the first or second aspect described above.
[0041] In a tenth aspect, a vehicle is provided that includes means as in any of the possible implementations of the third to sixth aspects, or the vehicle includes computer-readable storage as in any of the possible implementations of the eighth aspect, or the vehicle includes a chip as in any of the possible implementations of the ninth aspect, or the vehicle is loaded with computer program code as in any of the possible implementations of the seventh aspect.
[0042] In conjunction with aspect ten, in some implementations of aspect ten, the vehicle is a vehicle in a broad sense, such as a means of transportation (e.g., commercial vehicles, passenger cars, motorcycles, flying cars, trains, etc.), industrial vehicles (e.g., forklifts, trailers, tractors, etc.), engineering vehicles (e.g., excavators, bulldozers, cranes, etc.), agricultural equipment (e.g., lawnmowers, harvesters, etc.), amusement equipment, toy vehicles, etc. In practical implementation, the vehicle can also be a road vehicle, a water vehicle, an air vehicle, industrial equipment, agricultural equipment, or other intelligent driving equipment such as entertainment equipment.
[0043] For the beneficial effects not described in detail in aspects two through ten, please refer to the description in aspect one, which will not be repeated here. Attached Figure Description
[0044] Figure 1 is a functional schematic block diagram of the vehicle provided in an embodiment of this application;
[0045] Figure 2 is a schematic block diagram of a system for identifying road elements provided in an embodiment of this application;
[0046] Figure 3 is a schematic diagram of the process for identifying road elements provided in an embodiment of this application;
[0047] Figure 4 is another schematic diagram of the process for identifying road elements provided in the embodiments of this application;
[0048] Figure 5 is another schematic diagram of the process for identifying road elements provided in the embodiments of this application;
[0049] Figure 6 is another schematic diagram of the process for identifying road elements provided in the embodiments of this application;
[0050] Figure 7 is a schematic diagram of the application scenarios involved in the embodiments of this application;
[0051] Figure 8 is a schematic block diagram of a device for identifying road elements provided in an embodiment of this application;
[0052] Figure 9 is another schematic block diagram of the device for identifying road elements provided in the embodiments of this application. Detailed Implementation
[0053] As mentioned earlier, for intelligent vehicles, accurate perception of road elements such as lanes, lane lines, and road signs is crucial in highly dynamic and risk-sensitive driving scenarios, helping to guide timely and accurate decisions. The road information acquired by a vehicle's perception system contains rich geometric and semantic information. Geometric information refers to the length, width, position, shape, and orientation of road elements in Euclidean space; semantic information refers to multi-layered and multi-dimensional information that enables vehicles to better understand driving rules, perceive road conditions, and plan routes. For example, based on semantic information, vehicles can distinguish static road elements such as lanes, medians, roadside trees, and signs. However, under current technological limitations, such as the perception range of vehicle sensors, vehicles cannot accurately identify the semantic information of road elements while driving, leading to driving risks such as traffic violations.
[0054] In view of this, the embodiments of this application provide a scheme for identifying road elements, which can determine the semantic information of road elements based on their geometric information, reduce the impact of false detection of geometric information of road elements on the accuracy of semantic information of road elements, and when identifying road elements in this scheme, it is not necessary to identify both geometric and semantic information of road elements at the same time, so that the identification task can focus on identifying the semantic information of road elements, which helps to improve the accuracy of semantic information identification of road elements.
[0055] In some implementations, the road elements involved in this application can be understood as: elements constituting a road, elements indicating traffic conditions on the road, and / or elements affecting traffic conditions on the road. For example, road elements may include lanes, road signs, obstacles in the road, etc. Specifically, guide lines and / or lane markings in a lane can be understood as elements indicating traffic conditions on the road; traffic signs can be understood as elements indicating traffic conditions on the road; obstacles present in the road (such as other road users) can be understood as elements affecting traffic conditions on the road; lanes can be understood as elements constituting a road. Furthermore, different lanes may have different uses, which can be determined by text or road signs in the lanes, etc. Different types of vehicles may need to travel in their designated drivable lanes. For example, motor vehicles can travel in motor vehicle lanes but not in non-motor vehicle lanes. Therefore, lanes can also be understood as elements indicating traffic conditions on the road.
[0056] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0057] Figure 1 is a functional block diagram of a vehicle provided in an embodiment of this application. As shown in Figure 1, the vehicle 100 may include a perception system 120 and a computing platform 150. The perception system 120 may include several sensors for sensing information about the environment surrounding the vehicle 100. For example, the perception system 120 may include a positioning system, which may be a global navigation satellite system (GNSS), such as the global positioning system (GPS), BeiDou system, etc. Alternatively, the perception system 120 may also include one or more of the following: an inertial measurement unit (IMU), lidar, millimeter-wave radar, ultrasonic radar, and a camera device.
[0058] Some or all of the functions of vehicle 100 can be controlled by computing platform 150. Computing platform 150 may include processors 151 to 15n. A processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a central processing unit (CPU), microprocessor, graphics processing unit (GPU) (which can be understood as a type of microprocessor), or digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as a field-programmable gate array (FPGA). In reconfigurable hardware circuits, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement related functions. Furthermore, the processor can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), tensor processing unit (TPU), deep learning processing unit (DPU), etc. In addition, the computing platform 150 may also include a memory for storing instructions. Some or all of the processors 151 to 15n can call the instructions in the memory to implement the corresponding functions.
[0059] In this application, the computing platform 150 can determine the semantic information of road elements based on the environmental information collected by the perception system 120 and the geometric information of road elements. The roles of the perception system 120 and the computing platform 150 in the process of identifying road elements are explained in detail below with reference to Figure 2.
[0060] Figure 2 shows a schematic block diagram of the road boundary detection system architecture provided in an embodiment of this application. The system includes a perception module 210 and a detection module 220. In some implementations, the system may also include a control module 230. Exemplarily, the perception module 210 may include one or more sensors from the perception system 120 shown in Figure 1; the detection module 220 may include one or more processors from the computing platform 150 shown in Figure 1, or the detection module 220 may include one or more processors from a cloud server communicating with vehicles; the control module 230 may include one or more processors from the computing platform 150 shown in Figure 1. The functions of each module are as described in items (a) to (iii) below.
[0061] (i) The perception module 210 is used to collect environmental information around the vehicle. This environmental information can indicate the road structure around the vehicle, and the form of the environmental information can be images, laser point clouds, etc. The perception module 210 can send the collected environmental information to the detection module 220.
[0062] (II) The detection module 220 is used to determine the semantic information of road elements based on the geometric information of road elements and the environmental information collected by the perception module 210. More specifically, the detection module 220 may include a perception information processing module 221, a geometric feature extraction module 222, and a fusion processing module 223. In some implementations, the detection module 220 may also include a result verification module 224. The perception information processing module 221 processes the perception information from the perception module 210 to obtain scene description information related to the scene indicated by the perception information; the geometric feature extraction module 222 determines the geometric features such as the positional relationship of each road element based on the geometric information of the road elements; the fusion processing module 223 can fuse the scene description information and the positional relationship of each road element to obtain the semantic and geometric information of the road elements. By extracting geometric features such as the positional relationship of road elements through the geometric feature extraction module 222, the fusion processing module can focus on acquiring the semantic information of the road elements to be identified, thereby improving the accuracy of the semantic information recognition results. In addition, the information output by the perception information processing module 221 and the information output by the geometric feature extraction module 222 can be mutually verified to correct the geometric features of the road elements, thereby improving the accuracy and reliability of the road element recognition results.
[0063] More specifically, the aforementioned scene description information can indicate the location, meaning, and other information of various road elements. For example, when the scene indicated by the perception information includes lane lines, road boundaries, traffic signs, and other information, the scene description information can indicate the location of lane lines, road boundaries, and traffic signs; or the scene description information can also indicate the line type and color of lane lines, the meaning of traffic signs, etc.
[0064] In some implementations, the geometric information of road elements can be extracted from image or point cloud data processing, or it can be obtained through other means. For example, the geometric information of road elements can include a vector map. A vector map can be understood as a map composed of vector data representing the location and shape of geographic entities. Vector data can include at least one of points, lines, and polygons.
[0065] In one example, when the road element geometry information is extracted from image or point cloud data, the image or point cloud data may be the same as the perception information input to the perception information processing module, but the neural network used to extract the road element geometry information may be different from the neural network used by the perception information processing module 221; or the image or point cloud data may be different from the perception information input to the perception information processing module.
[0066] In another example, the vector map can be generated based on traffic flow data. The vector map can include at least one of road vectors, intersection vectors, and lane vectors. For example, a processor segments and clusters traffic flow data to obtain road vectors. Further, for multiple roads intersecting at the same intersection, based on traffic flow data and road vectors, the vector points connecting each road to the intersection are determined, and the vector points corresponding to multiple roads constitute the intersection vector. The road width is determined based on the traffic flow data, and the intersections of multiple sets of traffic flow data with the perpendicular lines from the roads are clustered. The number of lanes is determined based on the clustering results, and then the lane vectors are determined based on the road width and the number of lanes. It can be understood that lane vectors, road vectors, and intersection vectors constitute a vectorized map. Specifically, road vectors indicate the location and direction of a road segment, intersection vectors indicate the location and boundaries of intersections, and lane vectors indicate the roadway for various vehicles to travel within the same road width. Alternatively, lane vectors can also indicate the position of each lane within a road segment.
[0067] It should be noted that a road can correspond to one or more sets of traffic flow data. Each set of traffic flow data can be understood as data consisting of the trajectories formed by one or more vehicles traveling on that road. Each set of traffic flow data can include multiple traffic flow points, each traffic flow point indicating a coordinate in a vehicle's trajectory, the time the vehicle arrived at that coordinate, and the vehicle's orientation pose at that coordinate. In some implementations, each set of traffic flow data also includes information such as the type of vehicle forming the traffic flow. In actual implementation, each set of traffic flow data can be data collected by vehicles or roadside units (RSUs) containing the trajectory of at least one vehicle. When the traffic flow data is collected by vehicles, the trajectory of at least one vehicle includes its own trajectory and / or the trajectory of at least one other vehicle.
[0068] In some implementations, the result verification module 224 can also verify the output of the fusion processing module 223. For example, according to predetermined rules, data with the same or different information sources as the input fusion processing module 223 are processed to determine the semantic information of the road element to be identified, and then the semantic information is used to verify the output of the fusion processing module 223 to obtain the final semantic information recognition result.
[0069] (iii) The planning and control module 230 can plan a driving path for the vehicle based on the road elements predicted by the detection module 220, and control the vehicle to drive along the planned driving path.
[0070] It should be understood that the above modules are only an example, and in actual applications, these modules may be added or removed as needed. For example, in the system architecture shown in Figure 2, the sensing module 210 and the detection module 220 can be merged into one module; or, the detection module 220 and the control module 230 can be merged into one module.
[0071] The above describes the system architecture for identifying road elements provided in the embodiments of this application. The following details the process of implementing the method for identifying road elements provided in the embodiments of this application based on the system architecture shown in Figure 2.
[0072] Figure 3 shows a schematic flowchart of a method for identifying road elements, which can be executed by vehicle 100 or by a cloud server communicating with vehicle 100. Exemplarily, the method can be executed by the detection module 220 shown in Figure 2. The method 300 includes:
[0073] S310, Obtain the first set of data, which indicates the geometric information of M road elements in the first road.
[0074] For example, the first set of data may include road element geometric information in the foregoing embodiments. For instance, the first set of data may include vector map data; or it may also include images and point cloud data, wherein the image includes pixels of road elements and the point cloud data includes point clouds of road elements.
[0075] In some implementations, the first set of data can also indicate the semantic information of the M road elements.
[0076] S320, Obtain the second set of data, which indicates the semantic and geometric information of N road elements in the first road.
[0077] For example, the second set of data may include vehicle perception information such as images and point cloud data, and may also include traffic flow data. Since the first road is the road the vehicle is currently on, the second set of data includes images and / or point cloud data of the first road collected by the vehicle's perception system.
[0078] S330, based on the first set of data and the second set of data, determine the semantic information of the road element to be identified in the first road. The road element to be identified is associated with M road elements and N road elements, and the M road elements and N road elements are associated.
[0079] Where M and N are both positive integers. The road element to be identified is associated with M road elements and N road elements, which can be understood as: the M road elements and / or N road elements include the road element to be identified; or, the M road elements and / or N road elements do not include the road element to be identified, but the semantic information of the road element to be identified can be inferred based on the information of the M road elements and N road elements. The association of M road elements and N road elements can be understood as: the geometric information of the M road elements can be used to determine or correct the geometric information of the N road elements.
[0080] In some implementations, S330 can be refined as follows: inputting the first set of data into a first type of neural network to obtain the first set of feature information, which indicates the geometric features of M road elements in the first road; inputting the second set of data into a second type of neural network to obtain the second set of feature information, which indicates the geometric and semantic features of N road elements in the first road; inputting the first set of feature information, the second set of feature information, and the position information of the road element to be identified into a fusion neural network to obtain fused feature information, which indicates one or more initial semantic features of the road element to be identified; and determining the semantic information of the road element to be identified based on the fused feature information.
[0081] Geometric features can be understood as an abstract representation of the geometric information of road elements. These features can indicate the location, positional changes, or direction of the road elements. For example, when a road element is a lane, lane line, or road obstacle, geometric features can indicate its location; when a road element is a traffic sign, geometric features can indicate its location or shape. Semantic features can be understood as an abstract representation of the semantic information of road elements. These features can indicate the meaning, type, or other characteristics of the road elements. For example, when a road element is a lane, semantic features can indicate its type as a regular motor vehicle lane, non-motor vehicle lane, stop lane, bus lane, tidal flow lane, or reversible lane; when a road element is a traffic sign, semantic features can indicate its type as a warning sign, instruction sign, or prohibition sign, or the meaning of the traffic sign could be a prohibition on motor vehicle passage, a straight ahead warning, or a warning right turn signal.
[0082] For example, the first type of neural network may include at least one neural network focused on extracting geometric features. The aforementioned first type of neural network may include a convolutional neural network (CNN), or it may include a fully connected neural network. For instance, if the first set of data includes vector map data, images, and point cloud data, then at least one neural network may include a neural network for converting the vector map data into a format that can be processed by downstream neural networks (such as fused neural networks). At least one neural network may also include a neural network for extracting pixel or point cloud data of road elements from the image, such as a neural network based on algorithms like normal aligned radial feature (NARF), ORB (Oriented Fast and Rotated BRIEF), or the Harris corner detection algorithm. It should be understood that when the first type of neural network extracts geometric features, it may also extract some semantic features associated with the geometric features.
[0083] For example, the second type of neural network may include at least one neural network focused on extracting semantic features. The aforementioned second type of neural network may include CNNs and / or fully connected neural networks. For instance, if the second set of data includes images and point cloud data, then at least one neural network may include a neural network for target object recognition, such as a neural network based on algorithms like YOLO (you only look once), Residual Network (ResNet), or Visual Geometry Group (VGG).
[0084] In one example, the first and second sets of data are the same type of data. Different processing methods using a first-type neural network and a second-type neural network yield different feature information. In another example, the first and second sets of data are different types of data, and each set includes information about different road elements.
[0085] In some implementations, the fusion neural network can be a Transformer-type neural network. For example, the fusion neural network can be a neural network based on a multi-head attention mechanism. Specifically, the fusion neural network can include multiple head networks and a backbone network. Each head network can extract features associated with the road element to be identified from the feature information associated with a data source (such as a data source in the first set of data or a data source in the second set of data). For example, taking the road element to be identified as a lane as an example, one head network is used to extract the features of traffic signs around the road element to be identified, one head network is used to extract the features corresponding to the text or guide lines in the lane, and one head network is used to extract the features of the lane boundary (such as features indicating that the lane boundary is one or more of the lane lines, curbs, or fences).
[0086] In some implementations, the geometric features indicated by the second set of feature information may deviate from the actual geometric information of the road element, resulting in inaccurate semantic information of the road element obtained solely based on the second set of feature information. By fusing the first and second sets of feature information, the geometric features of the road element can be corrected, thereby improving the accuracy of the semantic information of the road element.
[0087] In some implementations, the method further includes: determining a distance matrix based on the first set of data and the location information of the road elements to be identified, wherein the distance matrix indicates the correlation between each of the M road elements and the road elements to be identified; determining the semantic information of the road elements to be identified in the first road, including: inputting the first set of feature information, the second set of feature information, and the distance matrix into a fusion neural network to obtain fused feature information.
[0088] For example, the closer the distance between two road elements, the higher the correlation between the two road elements. When determining the semantic information of the road element to be identified, road elements with a higher correlation with the road element to be identified can be used, thereby improving the accuracy of the semantic information recognition result.
[0089] In some implementations, the road element to be identified is the first lane in a first road, and semantic information indicates the type of the first lane; the first set of data includes at least data indicating the location of a boundary of the first lane, and the second set of data includes at least one of the following: data indicating traffic signs associated with the first lane, traffic flow data associated with the first lane, or lane line data of the first lane; or, the first set of data includes data indicating the location of the boundary of at least one lane in the first road, at least one lane is associated with the first lane, and the second set of data includes at least one of the following: traffic flow data associated with each lane in the at least one lane, or lane line data of each lane in the at least one lane.
[0090] For example, associating at least one lane with the first lane may include: at least one lane comprising the first lane, or at least one lane comprising one or more lanes adjacent to the first lane. In some implementations, at least one lane and the first lane support vehicles traveling in the same direction, or at least one lane and the first lane support vehicles traveling in the same direction for at least a period of time during the day.
[0091] In one example, taking the second set of data, which includes traffic flow data, as an example, this traffic flow data can be traffic flow data within a certain time period, meaning the traffic flow data corresponding to the first road can be updated in real time. The certain time period can be a period prior to the current moment, and its duration can be 24 hours, 48 hours, or other durations. More specifically, the type of each lane can be determined based on the changes in the number of vehicles traveling in each lane within different time periods, and / or the type of the first lane can be determined based on the changes in the travel directions supported by the first lane within different time periods.
[0092] For example, multiple time periods include a first time period and a second time period. When the driving direction supported by the first lane is the first direction in the first time period and the driving direction supported by the first lane is the second direction in the second time period, the first lane is determined to be a tidal flow lane.
[0093] For example, multiple time periods include a third time period and a fourth time period. If the number of vehicles traveling in the first lane during the third time period is the first number, and the number of vehicles traveling in the first lane during the fourth time period is the second number, and the difference between the first number and the second number is greater than or equal to the first number threshold, then the first lane is determined to be a time-specific dedicated lane.
[0094] For example, at least one lane also includes a second lane. If the difference between the number of vehicles traveling in the first lane and the number of vehicles traveling in the second lane during the third time period is greater than or equal to a second quantity threshold, and the difference between the number of vehicles traveling in the first lane and the number of vehicles traveling in the second lane during the fourth time period is less than the second quantity threshold, then the first lane is determined to be a time-specific dedicated lane.
[0095] Furthermore, the type of the first road can be determined based on geometric information. When the first road is a highway and the first lane is a time-limited lane, the first lane can be determined as an emergency lane; when the first road is an urban road and the first lane is a time-limited lane, the first lane can be determined as a bus lane.
[0096] In another example, taking the second set of data, which includes data on traffic signs associated with the first lane, the meaning of the traffic signs can be identified to determine the type of the first lane. For instance, if a traffic sign indicates that the first lane is a "dedicated lane," then the first lane can be identified as a non-reserved lane.
[0097] In another example, taking the second set of data, which includes the lane lines of the first lane, at least one of the lane line type, line shape, or color can be determined to identify the type of the first lane. For example, if the lane line type is a fence or curb, it can be determined that the first lane is located at the edge of the road. Furthermore, combined with the location of the first road, it can be preliminarily determined that the first lane is a non-motorized vehicle lane, a regular lane, or an emergency lane.
[0098] In some implementations, when the road element to be identified is the lane line of the first lane, semantic information indicates the type of lane line. The first set of data includes at least data indicating the location of a boundary of the first lane, and the second set of data includes at least one of the following: traffic flow data associated with the first lane, including an image of the pixels of the boundary of the first lane, or point cloud data including the point cloud of the first lane. For example, the type of lane line can be double yellow lines, single yellow lines, solid and dashed lines, solid white lines, dashed lines, yellow dashed lines, speed reduction warning lines, etc. For example, the type of lane line can be determined based on whether each traffic flow changes lanes, the direction of the traffic flow, and the location of the lane change in the traffic flow data.
[0099] In some other implementations, the road element to be identified is a traffic sign in the first road, and the semantic information indicates the meaning of the traffic sign; the first set of data includes data indicating the location of the traffic sign, and the second set of data includes at least one of the following: data indicating the boundary of the first road, or data indicating the lane associated with the traffic sign.
[0100] For example, the lane data associated with a traffic sign may include lane traffic flow data, an image containing lane pixels, or point cloud data containing lane point clouds, etc. When the traffic sign is related to the direction of vehicle travel, taking the lane data associated with the traffic sign including lane traffic flow data as an example, the direction of vehicle travel indicated by the traffic sign can be determined based on the direction of the traffic flow data at the intersection. When the traffic sign is related to the direction of vehicle travel, taking the lane data associated with the traffic sign including an image containing lane pixels and / or point cloud data containing lane point clouds as an example, the image and / or point cloud data can be processed to determine the direction of the guide lines in the lane, and then the direction of the vehicle indicated by the traffic sign can be determined based on the direction of the guide lines. For determining the meaning of other traffic signs, please refer to the description in this example, which will not be repeated here.
[0101] Understandably, in practical implementation, when the road element to be identified is a traffic sign, the second set of data can also include an image of the traffic sign to be identified. For the aforementioned scenario, the image can be processed to determine the meaning of the traffic sign.
[0102] In some other implementations, the road element to be identified is an obstacle in a first road, and semantic information indicates the type of obstacle; the first set of data includes data indicating the location and / or changes in location of the obstacle, and the second set of data includes at least one of the following: data indicating the boundary of the first road, or data indicating the lane associated with the obstacle.
[0103] For example, the data associated with the lane containing the obstacle may include traffic flow data for that lane, and the lane associated with the obstacle may include the lane where the obstacle is located. In some implementations, the obstacle can be determined to be a dynamic or static obstacle based on changes in its position. Further, based on traffic flow data, the obstacle can be determined to be a temporary obstacle such as a disabled vehicle or an obstacle used for road closure. For instance, if traffic flow data indicates that there are traffic points both in front of and behind the obstacle when the obstacle is in the lane, then the obstacle can be determined to be a temporary obstacle such as a disabled vehicle. If there is a traffic point behind the obstacle but no traffic point in front of it, then the obstacle can be determined to be an obstacle used for road closure. Here, "in front of" or "behind" the obstacle can be understood as "in front of" or "behind" relative to the vehicle's direction of travel. In a direction parallel to the vehicle's direction of travel, the direction relative to the obstacle moving towards the vehicle is considered "behind" of the obstacle, and the direction relative to the obstacle moving away from the vehicle is considered "in front" of the obstacle.
[0104] For example, the second type of neural network and the fusion neural network can be trained based on the rules involved in the foregoing examples, so that the second type of neural network and the fusion neural network output relevant feature information based on the data. In actual implementation, each neural network can also be trained based on data other than the data involved in the foregoing examples, so that it outputs relevant feature information.
[0105] The method for identifying road elements provided in this application determines the semantic information of the road elements to be identified based on the geometric information of the road elements. This allows the identification task to focus on the identification of semantic information when road element identification is performed. When errors occur in the detection results of geometric information, the impact of geometric misdetection on the identification results of semantic information can be weakened, thereby improving the accuracy of the semantic information identification results of road elements.
[0106] For example, taking a lane as an example of the road element to be identified, Figure 4 shows another schematic flowchart of road element identification provided by an embodiment of this application. As shown in Figure 4, geometrically related data 1 and geometrically related data 2 can be regarded as an example of the first set of data in method 300, encoding network 1 and encoding network 2 can be regarded as an example of the first type of neural network in method 300, radar bird's-eye view (BEV) image, traffic flow image, forward-looking or surround-view image can be regarded as an example of the second set of data in method 300, and image encoding network 1 and image encoding network 2 can be regarded as an example of the second type of neural network. The radar BEV image can be an image obtained by mapping laser point cloud data to BEV coordinates, the traffic flow image can be an image obtained by mapping traffic flow data to BEV coordinates, and the forward-looking or surround-view image can be the original image captured by the camera device. For example, geometrically related data 1 and geometrically related data 2 can include data indicating information related to fences, curbs, or lane lines in the road, such as images including pixels of fences, curbs, or lane lines, or point cloud data including fences, curbs, or lane lines. In some implementations, geometrically related data 1 and geometrically related data 2 can also be mapped to a vector BEV image, and this vector BEV image can be used as one of the inputs to the image coding network 1. Furthermore, the fusion network and the decoding network can each be part of a Transformer neural network, where the fusion network can be considered an example of the fusion neural network in method 300.
[0107] In some implementations, multiple initial semantic information can be determined based on a portion of the first or second set of data. These initial semantic information, along with the geometric information of the road element to be identified, are then input into a second type of neural network. The features output by the second type of neural network are then input into a fusion neural network to obtain fused feature information. In this implementation, each initial semantic information can be determined based on one or more rules involved in method 300. Specifically, as shown in Figure 5, at least one initial semantic information for a road element is determined based on the information of the road element to be identified and other road information. Further, this initial semantic information is input into an encoding network to extract the geometric and semantic features of the road element to be identified from the initial semantic information. The other road information may include a portion of the data from the first or second set of data. For example, it may include data indicating the association of fences, curbs, or lane lines in the road, such as images including pixels of fences, curbs, or lane lines, or point cloud data including fences, curbs, or lane lines. Furthermore, perceptual information is input into another encoding network to extract the geometric and semantic features of the road element indicated by the perceptual information. The perceived information may include image and / or point cloud data, such as images and / or point cloud data acquired by the vehicle in real time. Further, the feature information input from the two encoding networks is fed into the fusion network to obtain fused feature information. For example, the encoding network shown in Figure 5 can be a second type of neural network in method 300, and the fusion network and decoding network can each be part of a Transformer neural network. The fusion network can be considered an example of a fusion neural network in method 300.
[0108] In some implementations, multiple initial semantic information can be determined based on the first set of data or a portion of the second set of data. Then, using these initial semantic information and the semantic information output by the decoding network, the final semantic information of the road element can be determined. The specific process is shown in Figure 6. For example, taking lane 1 as the road element to be identified, the initial semantic information 1 of the road element to be identified can indicate that lane 1 is a regular lane and a temporary parking lane. Furthermore, the initial semantic information 1 of the road element to be identified indicates that lane 2 is a temporary parking lane. Therefore, the final semantic information of lane 1 can be determined based on the confidence level of each semantic information. For example, if the confidence level for "regular lane" is the highest, then the final semantic information of lane 1 can be determined as a regular lane. In this implementation, each initial semantic information can be determined based on one or more rules involved in method 300. More specific implementation methods and the meaning of each piece of information can be found in the descriptions of the corresponding parts of Figures 3, 4, and 5, and will not be repeated here.
[0109] Figure 7 shows a schematic diagram of an application scenario of an embodiment of this application. In this case, a frame of image captured by the vehicle's forward-facing camera can be as shown in Figure 7(a). The image includes pixels of lane lines 401, 402 and 403, as well as pixels of curb 404 and curb 405. In addition, the image also includes pixels of other vehicles and pedestrians. When the road element to be identified is a lane, the geometric and semantic information of the lane can be identified using an image recognition algorithm. Due to the occlusion of pedestrians 411, 412, vehicles 413, and 414, the curb of the road on the other side of the intersection cannot be detected. This results in the identification of the geometric information of the road element as shown in Figure 7(b), where lines A to E are respectively regarded as the positions of the pixels corresponding to lane lines 401 to 403 and curbs 404 and 405 in the image. Furthermore, due to the occlusion of the lane, the right boundary line of lane 421 on the other side of the intersection is not detected, leading to inaccurate identification results for the lane between lane line 402 and curb 404. For example, the semantic information of lane 421 cannot be obtained. For the aforementioned scenario, the road element detection method provided in this application embodiment can reduce the probability of semantic information detection errors of road elements caused by geometric misdetection.
[0110] In some implementations, the vehicle can control its own movement based on the road elements identified in the embodiments of this application; or, the vehicle's display device can be controlled to display the road elements identified in the embodiments of this application.
[0111] In the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions between the various embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0112] The method for identifying road elements provided by the embodiments of this application has been described in detail above with reference to Figures 1 to 7. The apparatus provided by the embodiments of this application will now be described in detail below with reference to Figures 8 and 9. It should be understood that the description of the apparatus embodiments corresponds to the description of the method embodiments; therefore, any content not described in detail can be referred to the method embodiments above, and for the sake of brevity, will not be repeated here.
[0113] Figure 8 shows a schematic block diagram of a road element identification device 2000 provided in an embodiment of this application. The device 2000 may include units for performing the aforementioned method. Furthermore, each unit in the device 2000 implements a corresponding process of the above method embodiment. The device 2000 includes an acquisition unit 2010, which can be used to implement corresponding data acquisition or transmission / reception functions. The device 2000 also includes a processing unit 2020, which can be used to implement corresponding processing functions.
[0114] Optionally, the device 2000 further includes a storage unit, which can be used to store instructions and / or data. The processing unit 2020 can read the instructions and / or data in the storage unit so that the device can perform the relevant actions in the aforementioned method embodiments.
[0115] It should be understood that the specific process of each unit performing the above-mentioned corresponding steps has been described in detail in the above method embodiments, and will not be repeated here for the sake of brevity.
[0116] It should also be understood that the device 2000 described herein is embodied in the form of a functional unit. The terms “module” or “unit” may refer to application-specific ASICs, electronic circuits, processors (e.g., shared processors, proprietary processors, or group processors) and memory for executing one or more software or firmware programs, integrated logic circuits, and / or other suitable components that support the described functions.
[0117] The apparatuses described above are capable of implementing the corresponding steps performed by the computing platform 150 in the methods described above. These functions can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the functions described above; for example, the acquisition unit 2010 can be replaced by a transceiver, and other units, such as processing units, can be replaced by a processor, used to execute the relevant processing operations in each method embodiment.
[0118] Exemplarily, the acquisition unit 2010 and processing unit 2020 can be disposed in the vehicle 100 shown in FIG. 1, or they can also be disposed in the system shown in FIG. 2. More specifically, the acquisition unit 2010 and processing unit 2020 can be disposed in the detection module 220. Exemplarily, the operations performed by the acquisition unit 2010 and processing unit 2020 can be performed by a single processor, or they can be performed by different processors. In specific implementation, the one or more processors can be processors disposed in the vehicle 100 shown in FIG. 1; or, the device 2000 can be a chip disposed in the vehicle 100.
[0119] In the specific implementation process, the units in the above device can be fully or partially integrated together, or they can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a system-on-a-chip (SoC).
[0120] Figure 9 is another schematic block diagram of the road element identification device provided in the embodiments of this application. The road element identification device 2100 shown in Figure 9 may include: a processor 2110, a transceiver 2120, and a memory 2130. The processor 2110, transceiver 2120, and memory 2130 are connected via internal interconnection paths. The memory 2130 is used to store instructions, and the processor 2110 is used to execute the instructions stored in the memory 2130 to implement the methods in the above embodiments. Optionally, the memory 2130 may be coupled to the processor 2110 via an interface or integrated with the processor 2110.
[0121] It should be noted that the transceiver 2120 mentioned above may include, but is not limited to, transceiver devices such as input / output interfaces, to realize communication between device 2100 and other devices or communication networks.
[0122] Memory 2130 can be volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM). For example, RAM can be used as an external cache. By way of example and not limitation, RAM includes various forms such as: static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0123] Transceiver 2120 uses transceiver devices, such as but not limited to transceivers, to enable communication between device 2100 and other devices or communication networks to receive / send data / information for implementing the methods in the above embodiments.
[0124] This application also provides an intelligent driving device, which includes the device 2000 or device 2100 in the above embodiments.
[0125] This application also provides a server, which includes the device 2000 or device 2100 in the above embodiments.
[0126] This application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to implement the methods described in the above embodiments of this application.
[0127] This application also provides a computer-readable storage medium storing computer instructions that, when executed on a computer, cause the computer to implement the methods described in the above embodiments of this application.
[0128] This application also provides a chip, including circuitry, for performing the methods described in the above embodiments of this application.
[0129] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0130] In the description of the embodiments of this application, unless otherwise stated, " / " means "or", for example, A / B can mean A or B; "and / or" in this document describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In this application, "at least one" means one or more, and "more" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0131] The use of prefixes such as "first" and "second" in this application embodiment is solely for distinguishing different descriptive objects and does not limit the position, order, priority, quantity, or content of the described objects. The use of ordinal numbers and other prefixes to distinguish descriptive objects in this application embodiment does not constitute a limitation on the described objects. The description of the described objects is found in the claims or the context of the embodiments, and the use of such prefixes should not constitute unnecessary restrictions.
[0132] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0133] In the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions between the various embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0134] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0135] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0136] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for identifying road elements, characterized in that, include: Obtain the first set of data, which indicates the geometric information of M road elements in the first road; Obtain a second set of data, which indicates the semantic and geometric information of N road elements in the first road; Based on the first set of data and the second set of data, determine the semantic information of the road elements to be identified in the first road; Wherein, the road element to be identified is associated with the M road elements and the N road elements, and the M road elements and the N road elements are associated, where M and N are both positive integers.
2. The method according to claim 1, characterized in that, The step of determining the semantic information of the road element to be identified based on the first set of data and the second set of data includes: The first set of data is input into a first type of neural network to obtain a first set of feature information, which indicates the geometric features of M road elements in the first road; The second set of data is input into a second type of neural network to obtain a second set of feature information, which indicates the geometric and semantic features of N road elements in the first road; The first set of feature information, the second set of feature information, and the location information of the road element to be identified are input into the fusion neural network to obtain fused feature information, which indicates one or more initial semantic features of the road element to be identified. The semantic information is determined based on the fused feature information.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Based on the first set of data and the location information of the road element to be identified, a distance matrix is determined, wherein the distance matrix indicates the correlation between each of the M road elements and the road element to be identified; The step of determining the semantic information of the road element to be identified in the first road includes: The semantic information is determined based on the first set of data, the second set of data, and the distance matrix.
4. The method according to any one of claims 1 to 3, characterized in that, The first road is the road where the vehicle is currently located. The first set of data includes pixel data of the first road element and / or point cloud data of the second road element. The M road elements include the first road element and the second road element. Obtaining the first set of data includes: Acquire first perception information collected by the vehicle's perception system, the first perception information including images and / or point cloud data of the first road; The first perceived information is processed to obtain the pixel data of the first road element and / or the point cloud data of the second road element.
5. The method according to any one of claims 1 to 4, characterized in that, The first road is the road where the vehicle is currently located, and the second set of data includes images and / or point cloud data of the first road collected by the vehicle's perception system.
6. The method according to any one of claims 1 to 5, characterized in that, The road element to be identified is the first lane line in the first road, and the semantic information indicates the type of the first lane. The first set of data includes at least data indicating the location of a boundary of the first lane, and the second set of data includes at least one of the following: data indicating a traffic sign associated with the first lane, traffic flow data associated with the first lane, or lane line data of the first lane; or, The first set of data includes data indicating the location of the boundary of at least one lane in the first road, the at least one lane being associated with the first lane, and the second set of data includes at least one of the following: traffic flow data associated with each lane in the at least one lane, or lane line data for each lane in the at least one lane.
7. The method according to any one of claims 1 to 5, characterized in that, The road element to be identified is a traffic sign in the first road, and the semantic information indicates the meaning of the traffic sign. The first set of data includes data indicating the location of the traffic sign, and the second set of data includes at least one of the following: data indicating the boundary of the first road, or data indicating the lane associated with the traffic sign.
8. The method according to any one of claims 1 to 5, characterized in that, The road element to be identified is an obstacle in the first road, and the semantic information indicates the type of the obstacle; The first set of data includes data indicating the location and / or changes in location of the obstacle, and the second set of data includes at least one of the following: data indicating the boundary of the first road, or data indicating the lane associated with the obstacle.
9. A device for identifying road elements, characterized in that, include: The acquisition unit is used to acquire a first set of data, which indicates the geometric information of M road elements in the first road; The acquisition unit is further configured to: acquire a second set of data, wherein the second set of data indicates the semantic and geometric information of N road elements in the first road; The processing unit is configured to: determine the semantic information of the road element to be identified in the first road based on the first set of data and the second set of data; Wherein, the road element to be identified is associated with the M road elements and the N road elements, and the M road elements and the N road elements are associated, where M and N are both positive integers.
10. The apparatus according to claim 9, characterized in that, The processing unit is used for: The first set of data is input into a first type of neural network to obtain a first set of feature information, which indicates the geometric features of M road elements in the first road; The second set of data is input into a second type of neural network to obtain a second set of feature information, which indicates the geometric and semantic features of N road elements in the first road; The first set of feature information, the second set of feature information, and the location information of the road element to be identified are input into the fusion neural network to obtain fused feature information, which indicates one or more initial semantic features of the road element to be identified. The semantic information is determined based on the fused feature information.
11. The apparatus according to claim 9 or 10, characterized in that, The processing unit is used for: Based on the first set of data and the location information of the road element to be identified, a distance matrix is determined, wherein the distance matrix indicates the correlation between each of the M road elements and the road element to be identified; The semantic information is determined based on the first set of data, the second set of data, and the distance matrix.
12. The apparatus according to any one of claims 9 to 11, characterized in that, The first road is the road where the vehicle is currently located. The first set of data includes pixel data of the first road element and / or point cloud data of the second road element. The M road elements include the first road element and the second road element. The acquisition unit is used for: Acquire first perception information collected by the vehicle's perception system, the first perception information including images and / or point cloud data of the first road; The first perceived information is processed to obtain the pixel data of the first road element and / or the point cloud data of the second road element.
13. The apparatus according to any one of claims 9 to 12, characterized in that, The first road is the road where the vehicle is currently located, and the second set of data includes images and / or point cloud data of the first road collected by the vehicle's perception system.
14. The apparatus according to any one of claims 9 to 13, characterized in that, The road element to be identified is the first lane line in the first road, and the semantic information indicates the type of the first lane. The first set of data includes at least data indicating the location of a boundary of the first lane, and the second set of data includes at least one of the following: data indicating a traffic sign associated with the first lane, traffic flow data associated with the first lane, or lane line data of the first lane; or, The first set of data includes data indicating the location of the boundary of at least one lane in the first road, the at least one lane being associated with the first lane, and the second set of data includes at least one of the following: traffic flow data associated with each lane in the at least one lane, or lane line data for each lane in the at least one lane.
15. The apparatus according to any one of claims 9 to 13, characterized in that, The road element to be identified is a traffic sign in the first road, and the semantic information indicates the meaning of the traffic sign. The first set of data includes data indicating the location of the traffic sign, and the second set of data includes at least one of the following: data indicating the boundary of the first road, or data indicating the lane associated with the traffic sign.
16. The apparatus according to any one of claims 9 to 13, characterized in that, The road element to be identified is an obstacle in the first road, and the semantic information indicates the type of the obstacle; The first set of data includes data indicating the location and / or changes in location of the obstacle, and the second set of data includes at least one of the following: data indicating the boundary of the first road, or data indicating the lane associated with the obstacle.
17. A device for identifying road elements, characterized in that, include: A processor for executing a computer program stored in memory to cause the apparatus to perform the method as described in any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that, It stores instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 8.
19. A chip, characterized in that, The chip includes circuitry for performing the method as described in any one of claims 1 to 8.
20. A computer program product, characterized in that, The computer program product includes: computer program code, which, when executed by a processor, implements the method as described in any one of claims 1 to 8.
21. A vehicle, characterized in that, Includes the apparatus as claimed in any one of claims 9 to 17, or the computer-readable storage medium as claimed in claim 18, or the chip as claimed in claim 19, or the vehicle is equipped with the computer program product as claimed in claim 20.