Road detection method, storage medium, program product, electronic device and vehicle
By obtaining the vehicle driving environment image and processing the structural feature maps of multiple road elements, and using detection models to identify and determine the location information of road elements, the problem of ignoring the relationship between road elements in the prior art is solved, and the accurate perception and stability of the autonomous driving system in complex environments is realized.
Patent Information
- Application Number
- CN202411642482.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-08-12
AI Technical Summary
The existing road structure perception technology mainly focuses on the detection of a single road element, ignoring the correlation between road elements, resulting in errors and instability of autonomous driving systems in complex environments.
By obtaining the vehicle driving environment image, the structural feature maps of multiple road elements are obtained, and the position information of road elements is identified and determined using the detection model, including lane lines, curbsides, etc., and the correlation relationship between different road elements is represented through the affinity field map.
Accurate perception of road elements is achieved, and the reliability and stability of autonomous driving systems are improved, especially the understanding and prediction of road structures in complex scenarios.
Smart Images

Figure CN120472416A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent driving technology, and in particular to a road detection method, storage medium, program product, electronic equipment and vehicle. Background Art
[0002] In the fields of autonomous driving and map reconstruction, road structure perception technology can accurately detect and understand the location information of road elements such as lane lines, lane centerlines, and curbs. This is crucial for vehicle path planning, obstacle avoidance, and decision-making. Current road structure perception technologies mostly focus on detecting single road elements, such as lane lines, while ignoring road elements such as curbs or lane centerlines, and the relationships between them. This single-element detection approach fails to provide comprehensive road condition information, leading to errors in the autonomous driving system's understanding and prediction of road structure, reducing its adaptability and safety in complex environments. Summary of the Invention
[0003] Embodiments of the present application provide a road detection method, storage medium, program product, electronic device, and vehicle to accurately extract relevant information about road structure.
[0004] In order to achieve the above-mentioned object, according to a first aspect of the present application, a road detection method is provided, the method comprising:
[0005] Acquire a vehicle's driving environment image;
[0006] Processing the driving environment image to obtain a structural feature map of a plurality of road elements in the driving environment image, wherein the structural feature map is used to indicate aggregated information of the road elements and association relationships between the plurality of road elements;
[0007] Position information of a plurality of road elements in the driving environment image is determined according to the structural feature graphs of the plurality of road elements.
[0008] Optionally, the processing the driving environment image to obtain a structural feature map of a plurality of road elements in the driving environment image includes:
[0009] performing feature extraction on the driving environment image to obtain two-dimensional features of the driving environment image;
[0010] Performing feature conversion on the two-dimensional features to obtain a bird's-eye view feature map of the driving environment image;
[0011] The bird's-eye view feature map is processed to obtain a structural feature map of multiple road elements in the driving environment image.
[0012] Optionally, the processing of the bird's-eye view feature map to obtain a structural feature map of a plurality of road elements in the driving environment image includes:
[0013] The bird's-eye view feature map is processed by the trained detection model to obtain a structural feature map of multiple road elements in the driving environment image.
[0014] Optionally, the detection model includes a first network layer and a second network layer, and the road element includes a first road element and a second road element.
[0015] The trained detection model is used to process the bird's-eye view feature map to obtain a structural feature map of multiple road elements in the driving environment image, including:
[0016] Inputting the bird's-eye view feature map into the first network layer to obtain a structural feature map of the first road element;
[0017] The bird's-eye view feature map is input into the second network layer to obtain a structural feature map of the second road element.
[0018] Optionally,
[0019] The structural feature map of the first road element includes a confidence map and a vector field map of the first road element;
[0020] The structural feature map of the second road element includes a confidence map and a regression field map of the second road element;
[0021] Among them, the confidence map is used to represent the position probability map of road elements, the vector field map is used to represent the coordinate offset between the feature points belonging to the same road element and its adjacent pixel points, and the regression field map is used to represent the relative position relationship between different road elements.
[0022] Optionally, the detection model is trained by the following steps:
[0023] Acquire training samples and annotate the training samples to obtain a truth matrix, wherein the truth matrix includes at least one of a first truth matrix, a second truth matrix, and a third truth matrix; wherein the first truth matrix is used to represent the probability that each pixel belongs to the first road element, the second truth matrix is used to represent relative position information of a feature point belonging to the first road element and its adjacent pixels, and the third truth matrix is used to represent orthogonal projection information of each pixel onto a dividing line;
[0024] The detection model is trained using the training samples to obtain the trained detection model.
[0025] Optionally, the acquiring of training samples and labeling of the training samples to obtain a truth matrix includes:
[0026] Acquire at least one sample element sequence from a preset road element list, where the sample element sequence is a sequence consisting of position information of a plurality of points on an initially marked road element;
[0027] Based on the position information of each feature point in the sample element sequence, drawing a first line connecting points corresponding to each feature point in the first preset matrix to obtain a marked first preset matrix;
[0028] Based on the marked first preset matrix, the first true value matrix is obtained.
[0029] Optionally, the training samples are obtained and labeled to obtain a truth matrix, including:
[0030] generating a first heat map based on a preset first road element list;
[0031] Marking the first heat map to obtain a plurality of marked pixels in the first heat map;
[0032] Calculating the relative position information of each marked pixel and its related pixel;
[0033] Based on the relative position information, the second preset matrix is updated to obtain the second true value matrix.
[0034] Optionally, the marking the first heat map to obtain a plurality of marked pixels in the first heat map includes:
[0035] Based on the position information of each feature point in the first heat map, drawing a second line connecting each of the feature points;
[0036] Based on the second line, a plurality of marked pixel points are obtained.
[0037] Optionally, the acquiring of training samples and labeling of the training samples to obtain a truth matrix includes:
[0038] Based on the preset lane list, a second heat map is generated;
[0039] Marking the second heat map to obtain a separation line in the second heat map;
[0040] Determine the orthogonal projection point of each pixel in the second heat map on the separation line;
[0041] Based on the orthogonal distance between each pixel point and its orthogonal projection point, the third preset matrix is updated to obtain the third true value matrix.
[0042] Optionally, marking the second heat map to obtain a separation line in the second heat map includes:
[0043] Based on the position information of each feature point in the second heat map, drawing a third line connecting each of the feature points;
[0044] A separation line in the second heat map is obtained based on the third line in the second heat map.
[0045] Optionally, determining the orthogonal projection point of each pixel in the second heat map on the separation line includes:
[0046] Calculate the projection coefficient of each pixel point in the second heat map on the separation line;
[0047] Based on the projection coefficient and the second preset threshold, the starting point or the end point of the separation line is updated, or
[0048] Determine the orthogonal projection point of each pixel in the second heat map on the separation line.
[0049] Optionally, determining the position information of the plurality of road elements in the driving environment image according to the structural feature graphs of the plurality of road elements includes:
[0050] obtaining at least one first element sequence based on the confidence map and the vector field map of the first road element, where the first element sequence is a sequence consisting of position information of a plurality of points on the first road element;
[0051] obtaining at least one second element sequence based on the first element sequence and the regression field map of the second road element, where the second element sequence is a sequence consisting of position information of a plurality of points on the second road element;
[0052] Based on the at least one first element sequence and the at least one second element sequence, position information of a plurality of road elements in the driving environment image is determined.
[0053] Optionally, obtaining at least one first element sequence based on the confidence map and the vector field map of the first road element includes:
[0054] obtaining at least one initial first element sequence based on the confidence map of the first road element;
[0055] The initial first element sequence is updated based on the coordinate offset in the vector field map of the first road element to obtain the first element sequence.
[0056] Optionally, obtaining at least one second element sequence based on the regression field map of the first element sequence and the second road element includes:
[0057] Determining, based on the first element sequence, a coordinate offset corresponding to each piece of the position information in a regression field map of the second road element;
[0058] The second element sequence is obtained based on the first element sequence and the coordinate offset of the position information.
[0059] Optionally, the first road element includes a lane centerline and / or a curb, and the second road element includes a lane line.
[0060] According to the second aspect of the present application, an embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored, and the computer-readable storage medium stores instructions, which, when executed by a computer, enable the computer to implement any one of the road detection methods provided in the embodiments of the present application.
[0061] According to the third aspect of the present application, an embodiment of the present application further provides a computer program product, which stores instructions, and when the instructions are executed by a computer, enables the computer to implement any one of the road detection methods provided in the embodiments of the present application.
[0062] According to a fourth aspect of the present application, an embodiment of the present application further provides an electronic device, including:
[0063] a memory having a computer program stored thereon;
[0064] A processor is used to execute the computer program in the memory to implement any one of the road detection methods provided in the embodiments of the present application.
[0065] According to the fifth aspect of the present application, an embodiment of the present application further provides a vehicle, comprising the electronic device described above, or capable of executing any one of the road detection methods provided in the embodiments of the present application.
[0066] Some embodiments of the present specification include at least the following beneficial effects: by converting a bird's-eye view feature map to a structural feature map, that is, three-dimensional bird's-eye view data, information on the orientation and distance between road elements and the management relationship between road elements can be obtained, which enables the system to perceive both the relative orientation of road elements in the environment and the distance between road elements, thereby achieving accurate perception and helping to improve the reliability and stability of the autonomous driving system.
[0067] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0069] In order to more completely understand the present application and its beneficial effects, the following description will be given in conjunction with the accompanying drawings, wherein the same drawing numbers represent the same parts in the following description.
[0070] Figure 1 is an application scenario diagram of the road detection method according to some embodiments of this specification;
[0071] Figure 2 is an exemplary flow chart of a road detection method according to some embodiments of this specification;
[0072] Figure 3 is an exemplary flow chart for determining location information of road elements according to some embodiments of this specification;
[0073] Figure 4 is an exemplary schematic diagram of a network structure of a detection model according to some embodiments of this specification;
[0074] Figure 5 is an exemplary flow chart of labeling a first truth matrix according to some embodiments of this specification;
[0075] Figure 6 is an exemplary flow chart of labeling a second truth matrix according to some embodiments of this specification;
[0076] Figure 7 is an exemplary flow chart of labeling a third truth matrix according to some embodiments of this specification;
[0077] Figure 8 is an exemplary schematic diagram of location information of road elements according to some embodiments of this specification;
[0078] Figure 9 This is a structural diagram of an electronic device according to some embodiments of this specification. DETAILED DESCRIPTION
[0079] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0080] In order to facilitate understanding of the implementation scheme provided in the embodiment of the present application, the relevant application background of the road detection method provided in the embodiment of the present application is first explained.
[0081] With the advancement of intelligent vehicle technology, road structure perception has become a crucial component in areas such as assisted driving and autonomous driving. For example, in assisted driving, features like lane keeping and lane departure warning, enabled by road structure perception, can effectively reduce driver workload and workload, thereby significantly reducing the occurrence of traffic accidents. Furthermore, accurate road structure perception is a fundamental requirement for the safe and stable operation of autonomous vehicles.
[0082] Current road structure perception technologies mainly include the following:
[0083] (1) Road element modeling based on segmentation mask and clustering: This method first generates a mask of the road structure through a segmentation network, and then uses traditional clustering methods to instantiate road element information. This method cannot achieve the association learning between the original image and the final road element recognition, and the obtained road element information lacks directional information.
[0084] (2) Rule-based lane centerline modeling. The lane centerline generation method output by this method is relatively complex and requires professional personnel to design and implement.
[0085] (3) Road structure perception based on map-assisted information. This method requires additional processing of map information, making the network structure more complex. Due to changes in the road scene, the map information may contain erroneous information.
[0086] Because road elements do not exist in isolation but rather have complex spatial and logical relationships, related technologies are unable to effectively model the interrelated information between road elements, leading to errors in the autonomous driving system's understanding and prediction of road structure. Furthermore, when dealing with complex scenarios such as intersections, multi-lane roads, or irregular road surfaces, the diverse and complex interrelationships of road elements make it difficult for related technologies to accurately identify and model them, thus affecting the reliability and stability of autonomous driving systems.
[0087] In view of this, some embodiments of this specification provide a road detection method that uses affinity field graphs to represent the relationship between different road elements, such as the connectivity between lane lines, so that the detection model can understand and predict the association relationship between road elements. Through the collaborative work of three network layers, it can comprehensively identify and locate road structural elements.
[0088] Figure 1 This is an application scenario diagram of the road detection method shown in some embodiments of this specification.
[0089] like Figure 1 As shown, the application scenario 100 may include a first device 110 , a second device 120 , a database 130 , a user terminal 140 and a data collection device 150 .
[0090] The data acquisition device 150 may include one or more sensors for sensing information about the surrounding environment. For example, the image acquisition device 260 may include a positioning system, which may be a global positioning system (GPS), a BeiDou system, or other positioning systems. For another example, the data acquisition device 150 may include one or more of an inertial measurement unit (IMU), an accelerometer, a lidar, a millimeter-wave radar, an ultrasonic radar, and a camera. Exemplarily, a plurality of cameras are arranged around the mobile device, and the plurality of cameras may be located at the front, rear, and both sides of the vehicle to capture 360-degree environmental information around the vehicle.
[0091] The data acquisition device 150 (e.g., a camera mounted on a mobile device) is used to obtain an open-source, large-scale data set (i.e., a training set) required by the user. In some embodiments, the data acquisition device 150 can store the collected data set in the database 130, and the second device 120 trains the detection model based on the data set maintained in the database 130. It should be noted that in the embodiment of the present application, the data set maintained in the database 130 needs to be annotated with multiple types of true value annotations. Specifically, each training sample used for the second device 120 must be annotated with a true value, which includes a first true value matrix, a second true value matrix, and a third true value matrix. The first true value matrix is used to represent the probability that each pixel belongs to the first road element, the second true value matrix is used to represent the relative position information of the feature point belonging to the first road element and its adjacent pixels, and the third true value matrix is used to represent the orthogonal projection information of each pixel to the dividing line. Similarly, each training sample in the database 130 can be annotated with the multiple types of true value annotations described above, and then the detection model can be trained using these training samples annotated with true values.
[0092] The first device 110 can call data, codes, etc. in the storage module, or store data, instructions, etc. in the storage module. The storage module can be placed in the first device 110, or the storage module can be an external memory relative to the first device 110.
[0093] In some embodiments, the first device 110 may obtain an image of the vehicle's driving environment; process the image to obtain a structural feature map of multiple road elements in the driving environment image; and determine position information of multiple road elements in the driving environment image based on the structural feature map of the multiple road elements. For more information about this embodiment, please refer to Figure 2 Related description.
[0094] The detection model trained by the second device 120 can be applied to different systems or devices (i.e., the first device 110). For example, the first device 110 can be integrated into various mobile devices (wheeled construction equipment, autonomous driving vehicles, assisted driving vehicles, etc.). Autonomous driving vehicles can also be cars, trucks, motorcycles, buses, ships, airplanes, helicopters, lawn mowers, recreational vehicles, amusement park vehicles, construction equipment, trams, golf carts, trains, and carts, etc.
[0095] In some embodiments, the first device 110 includes a processing module, which can be used to execute the trained detection model provided in the embodiments of the present application.
[0096] User terminal 140 refers to one or more terminal devices or software used by a user. In some embodiments, user terminal 140 may be used by one or more users, including users who directly use the service or other related users. In some embodiments, user terminal 140 may be a mobile device 140-1, a tablet computer 140-2, a laptop computer 140-3, a desktop computer 140-4, or any combination thereof, and may be any other device with input and / or output functions. The product form of first device 110 and user terminal 140 is not limited herein.
[0097] The network enables communication between components and with other components outside the system, facilitating the exchange of data and / or information. In some embodiments, the network can be any one or more of a wired network or a wireless network.
[0098] It should be noted that in some embodiments of the present application, the detection model can also be split into multiple sub-modules / sub-units or the detection model can also include other modules or sub-units to jointly implement the solution provided in the embodiments of the present application, which is not specifically limited here.
[0099] It should also be noted that the training of the detection model described in the above embodiment can be implemented on the cloud side, and the training of the detection model described in the above embodiment can also be implemented on the terminal side, that is, the second device 120 can be located on the terminal side, and the trained detection model can be used directly on the terminal device, or it can be sent by the terminal device to other devices for use. Specifically, the embodiments of the present application do not limit the device (cloud side or terminal side) on which the detection model is trained or applied.
[0100] It should be noted that due to Figure 1 The trained detection model (or Figure 1 The first device 110 can be deployed in various mobile devices to process relevant perception information (such as video or image) on the road surface captured by a camera device mounted on the mobile device and output location information of road elements. The location information of the road elements can be input into a downstream module of the mobile device for further processing. The road detection method is described below using the mobile device as an example of an autonomous driving vehicle.
[0101] It should be noted that the above description of the application scenario 100 and its modules is for convenience only and does not limit this specification to the scope of the embodiments. It is understandable that those skilled in the art, after understanding the principles of the system, may arbitrarily combine the modules or form a subsystem connected with other modules without deviating from the principles. In some embodiments, Figure 1 The first device 110, second device 120, database 130, user terminal 140, and data acquisition device 150 disclosed herein may be different modules within a system, or a single module may implement the functions of two or more of the aforementioned modules. For example, the modules may share a storage module, or each module may have its own storage module. Such variations are within the scope of protection of this specification.
[0102] Figure 2 is an exemplary flow chart of a road detection method according to some embodiments of this specification. In some embodiments, process 200 can be executed based on a first device of the road detection method. Figure 2 As shown, the process 200 includes the following steps.
[0103] Step S210: Acquire a driving environment image of the vehicle.
[0104] The driving environment image is an image and / or video of the vehicle's surroundings. The image and / or video of the vehicle's surroundings can reflect information about the vehicle itself and / or objects around the vehicle, such as dynamic object information (e.g., other vehicles, pedestrians, animals, etc.) and static object information (e.g., buildings, roads, streetlights, traffic signs, traffic lights, etc.). In some embodiments, static objects can include planar objects (e.g., lane boundaries, lane centerlines, etc.) or non-planar objects (e.g., buildings, traffic lights, etc.).
[0105] In some embodiments, the driving environment image is a sequence of multiple frames of images in a time sequence of scanning. In some embodiments, each frame of the driving environment image can be continuous in time.
[0106] In some embodiments, the first device may collect images of the driving environment through the data collection device 150. For example, the first device may capture images and / or videos of the environment in which the vehicle is located through a camera installed on the vehicle.
[0107] In some embodiments, the first device may acquire images of the driving environment over time to track the environment the vehicle is in. For example, a camera device may communicate with the first device via a network to continuously, periodically, or intermittently send images and / or videos of the vehicle's surroundings.
[0108] Step S220 : Process the driving environment image to obtain a structural feature map of multiple road elements in the driving environment image.
[0109] The structural feature graph is used to indicate the aggregated information of road elements and the relationships between multiple road elements. For example, the structural feature graph of multiple road elements can be used to describe the characteristics and location information of road elements (such as lane lines, curbs, lane centerlines, etc.).
[0110] In some embodiments, aggregate information is used to describe the location information of road elements and the relationship between each point on the road element and its adjacent points. For example, aggregate information may include location probability information and direction information of the road element. Association relationships refer to the relationships between different road elements. For example, association relationships may include the relative distances between different road elements.
[0111] Road elements refer to information related to the road on which the mobile device is traveling. For example, road elements may include the road boundary, lane centerline, left and right lane boundaries, lane direction, etc.
[0112] Structural feature maps of multiple road elements in a driving environment image can be obtained in a variety of ways. For example, various image processing methods and machine learning methods can be used to identify objects (e.g., traffic lights, road signs, roadblocks, buildings, or vehicles) in the driving environment image to obtain a structural feature map of the road elements. For example, image processing algorithms (e.g., support vector machines, K-nearest neighbors, decision trees, neural networks, etc.) can be used to analyze and process the acquired driving environment image and extract the structural feature map of the road elements in the driving environment image.
[0113] In some embodiments, the first device can perform feature extraction on the driving environment image to obtain two-dimensional features of the driving environment image; perform feature conversion on the two-dimensional features to obtain a bird's-eye view feature map of the driving environment image; and process the bird's-eye view feature map to obtain a structural feature map of multiple road elements in the driving environment image.
[0114] The two-dimensional features of the driving environment image refer to the characteristic information of the driving environment image from the perspective of the two-dimensional image. For example, the two-dimensional features of the driving environment image may include the edges, textures, color distribution, shape and outline of objects and / or road structures in the vehicle's surrounding environment.
[0115] In some embodiments, the first device may extract two-dimensional features of the driving environment image in a variety of ways based on the driving environment image. For example, the first device may use various image processing methods and machine learning methods to identify objects (e.g., road markings, roadblocks, buildings, or vehicles) in the driving environment image to obtain the two-dimensional features of the driving environment image. For example, the first device may analyze and process the acquired two-dimensional features of the driving environment image based on an image processing algorithm (e.g., a support vector machine, K-nearest neighbor, a decision tree, a neural network, etc.) and extract the two-dimensional features from the image.
[0116] In some embodiments, the first device can recognize the driving environment image through a pre-trained neural network model and extract the two-dimensional features of the driving environment image. The input of the pre-trained neural network model may include the driving environment image, and the output may include the two-dimensional features of the driving environment image. The prediction model can be trained based on a large number of labeled training samples through various feasible methods. For example, parameter updates can be performed based on the gradient descent method. The training samples may include sample images, and the training samples may be obtained based on historical data. The labels may be two-dimensional features of the sample images, which may be obtained by manual or automatic annotation.
[0117] The sample image refers to a sample driving environment image.
[0118] A bird's-eye view feature map is a three-dimensional feature map obtained by observing the environment of the geographical area where the mobile terminal is located from a bird's-eye view. For example, the bird's-eye view feature map may include height information of objects in the environment (obstacle height, uphill and downhill height information, etc.), latitude and longitude information of the objects, etc.
[0119] In some embodiments, a bird's-eye view feature map of the driving environment image can be obtained based on the two-dimensional features of the driving environment image in various ways. For example, the first device can convert the two-dimensional features of the driving environment image into a bird's-eye view space based on an inverse perspective transformation to obtain the bird's-eye view feature map of the driving environment image.
[0120] In some embodiments, the first device may process the driving environment image based on a pre-trained neural network model to obtain a bird's-eye view feature map. In some embodiments, the pre-trained neural network model may include a feature extraction layer and a feature conversion layer, wherein the head network layer in the feature extraction layer may adopt a residual network (ResNet) architecture to extract initial feature information from the input driving environment image. The neck network layer in the feature extraction layer may adopt a feature pyramid network (FPN) to fuse feature information of different scales, enhance the multi-scale expression of feature information, and obtain two-dimensional features.
[0121] In some embodiments, the first device may input the driving environment image into a feature extraction layer, which may output multiple initial feature maps. The initial feature maps may be feature data from a two-dimensional image perspective. The multiple initial feature maps may be of the same or different sizes.
[0122] The feature extraction layer is a model or algorithm for extracting image features. In some embodiments, the feature extraction layer may include a Visual Geometry Group Network (VGG), a Residual Network (ResNet), etc.
[0123] In some embodiments, the first device may fuse multiple initial feature maps of equal or different sizes to obtain two-dimensional features of the driving environment image, and then input the two-dimensional features into a feature conversion layer to obtain a bird's-eye view feature map corresponding to the driving environment image. Feature fusion refers to the process of combining initial feature maps of the surrounding environment captured by different cameras into a single image. In some embodiments, feature fusion can be achieved through various methods, such as feature concatenation, feature summation, and element-wise multiplication.
[0124] The feature conversion layer is used to convert the two-dimensional features in the two-dimensional image perspective into the bird's-eye view space. In some embodiments, the feature conversion layer may include a visual transformer (VT), an autonomous driving environment perception architecture (Lift-Splat-Shoot, LSS), a monocular 3D object detection network (Categorical Depth Distribution Network, CADDN), and an autonomous driving perception framework (BEVFormer) that integrates spatiotemporal information.
[0125] A structural feature map of multiple road elements is used to describe the characteristics and location information of road elements (such as lane lines, curbs, lane centerlines, etc.). In some embodiments, the structural feature map is used to indicate relevant information about the road elements. This information may include location probability information, direction information, and associations between road elements. Direction information may include the overall direction of travel of the road element or a vector pointing from one point on the road element to another. Associations refer to the relationships between different road elements. For example, associations may include the relative distances between different road elements.
[0126] In some embodiments, the structural feature map includes at least one of a confidence map, a vector field map, and a regression field map of road elements, and the road elements include at least one of a lane line, a lane centerline, and a curb.
[0127] A confidence map is a location probability map used to describe road elements (such as lane centerlines, curbs, etc.). For example, the pixel value of each pixel in the confidence map can represent the probability that the pixel is located at the lane centerline or curb.
[0128] A vector field map is used to represent the positional relationships between points belonging to the same road element in an image. For example, a vector field map may include multiple feature points, with the pixel value of each feature point representing a feature vector pointing to an adjacent feature point. Exemplarily, the pixel value of each feature point represents the coordinate position (e.g., pixel coordinate value) of the adjacent feature point. Exemplarily, the pixel value of each feature point represents the coordinate offset between the feature point and the adjacent feature point. The coordinate offset may include a horizontal offset from the feature point to the adjacent feature point, a vertical offset from the feature point to the adjacent feature point, and so on.
[0129] A feature vector generally represents the positional relationship between a feature point and adjacent feature points (e.g., the previous feature point and the next feature point). A feature point can be a point on a road element that represents a feature such as shape.
[0130] In some embodiments, the vector field map includes a first direction map and a second direction map, wherein the first direction map is used to represent the coordinate offset between each feature point and the corresponding previous feature point, and the second direction map is used to represent the coordinate offset between each feature point and the corresponding next feature point.
[0131] In some embodiments, the dimensions of the first directional map and the second directional map are (H, W, 2), where H and W represent height and width, respectively, and 2 represents that each pixel point in the first directional map and the second directional map contains two values, which correspond to the horizontal and vertical coordinates of the adjacent feature points in the image or the horizontal coordinate offset of the point and the adjacent pixel point, the vertical coordinate offset, etc.
[0132] The regression field map is a feature map used to determine the relative positional relationships between different road elements. For example, the regression field map can represent the coordinate offset between a pixel on the lane centerline and the lane line. The coordinate offset indicates the shortest distance between a pixel on the lane centerline and the lane line.
[0133] In some embodiments, the structural feature map includes a confidence map and a vector field map of the lane centerline, a confidence map and a vector field map of the roadside, and a confidence map and a regression field map of the lane line.
[0134] In some embodiments of this specification, confidence maps, vector field maps, and regression field maps can help the model better understand the overall structure of the road and help accurately perceive various road elements.
[0135] In some embodiments, the bird's-eye view feature map can be processed to obtain structural feature maps of multiple road elements in the driving environment image through various methods. For example, the first device can process the bird's-eye view feature map of the driving environment image using a model or algorithm to determine the structural feature map.
[0136] In some embodiments, the bird's-eye view feature map of the driving environment image may be processed using a trained detection model to obtain a structural feature map of multiple road elements in the driving environment image.
[0137] The detection model is an algorithm or model used to determine the structural feature maps of multiple road elements. In some embodiments, the detection model is a machine learning model. For example, the detection model can include any one or a combination of a convolutional neural network (CNN) model, a neural network (NN) model, or other custom model structures.
[0138] In some embodiments, the processing device may input a bird's-eye view feature map of the driving environment image into the detection model, and output a structural feature map of multiple road elements in the driving environment image.
[0139] In some embodiments, the detection model can be trained based on a large number of labeled training samples through various feasible methods. For example, parameter updates can be performed based on the gradient descent method. An exemplary training process includes: inputting multiple labeled training samples into the initial detection model, constructing a loss function based on the labels and the results of the initial detection model, and iteratively updating the parameters of the initial detection model based on the loss function through gradient descent or other methods. When the preset conditions are met, the model training is completed and a trained detection model is obtained. The preset conditions may include convergence of the loss function, the number of iterations reaching a threshold, etc.
[0140] In some embodiments, the training samples may include multiple groups of training samples, each group of training samples including at least a bird's-eye view feature map of the sample image. The training samples may be obtained based on historical data.
[0141] In some embodiments, the labels can be the sample confidence maps and sample vector field maps of lane centerlines, the sample confidence maps and sample vector field maps of curbs, and the sample confidence maps and sample regression field maps of lane lines corresponding to the sample images. For example, the labels can be obtained by annotating the actual structural feature maps in the sample images.
[0142] In some embodiments, the detection model may include a first network layer and a second network layer. For more information about the structure of the detection model, see Figure 4 Related description.
[0143] In some embodiments of this specification, the detection model can be used to learn the correlation between the bird's-eye view feature map and the structural feature map corresponding to the driving environment image, which helps to improve the accuracy of the output structural feature map.
[0144] Step S230 : determining position information of the plurality of road elements in the driving environment image according to the structural feature graphs of the plurality of road elements.
[0145] The position information of road elements refers to the information related to the position of each road element in the driving environment image. For example, the position information of multiple road elements may include multiple sequences, each sequence representing the position information of a roadside, lane line, or lane centerline.
[0146] In some embodiments, the position information of multiple road elements can be represented in various forms such as a sequence or a matrix. For example, the position information of multiple road elements can be represented in the form of a sequence {A, B, C}, where the elements A, B, and C in the sequence represent the characteristic information of each type of road element, respectively.
[0147] In some embodiments, each element in the sequence may include multiple sub-elements, with different sub-elements corresponding to different feature information. For example, element A1 may represent lane line 1, and element A1 may be represented as (A11, A12, ...), where A11 and A12 respectively represent the position information of multiple points on the lane line.
[0148] The multiple points include but are not limited to the start point, end point, and feature point of the road element.
[0149] Feature points are key points that reflect the shape characteristics of road elements.
[0150] The position information of the road element may be a sequence consisting of pixel coordinates of multiple points of the road element in the image.
[0151] The representation form of the position information of the road elements is only used as an example here, and the road elements may also be represented in other forms.
[0152] In some embodiments, the first device can determine the position information of multiple road elements in the driving environment image based on the structural feature map of the multiple road elements in various ways. For example, the first device can use various image processing methods and machine learning methods to identify objects (e.g., lane centerlines, lane markings, curbs, etc.) in the structural feature map to obtain road elements.
[0153] Figure 8 is an exemplary schematic diagram of location information of road elements according to some embodiments of this specification;
[0154] like Figure 8 As shown in FIG, when processing complex scenes, such as intersections, the position information of multiple road elements in the driving environment image can be accurately identified based on the structural feature graph of multiple road elements.
[0155] In some embodiments of the present specification, by converting a bird's-eye view feature map to a structural feature map, that is, three-dimensional bird's-eye view data, information on the orientation and distance between road elements and the management relationship between road elements can be obtained, which allows mobile devices to perceive both the relative orientation of road elements in the environment and the distance between road elements, thereby achieving accurate perception.
[0156] Figure 3 is an exemplary flow chart for determining the location information of road elements according to some embodiments of this specification. In some embodiments, process 300 can be executed by a first device based on a road detection method. Figure 3 As shown, the process 300 includes the following steps.
[0157] Step S310: Obtain at least one first element sequence based on the confidence map and the vector field map of the first road element.
[0158] The first element sequence is used to represent position information belonging to a first road element. The first road element includes a lane centerline and / or a roadside.
[0159] In some embodiments, the first element sequence may be a sequence consisting of position information of multiple points on a first road element.
[0160] In some embodiments, the first element sequence can be obtained through various methods based on the confidence map and vector field map of the first road element. For example, the position information of each point of the first road element in the confidence map can be updated based on the coordinate offset of the point in the vector field map. The updated position information of multiple points of the first road element in the confidence map can be obtained as the first element sequence.
[0161] For more information about road elements, confidence maps, vector field maps, etc., please refer to Figure 2 Related description.
[0162] In some embodiments, at least one initial first element sequence can be obtained based on the confidence map of the first road element; based on the coordinate offset in the vector field map of the first road element, the initial first element sequence is updated to obtain the first element sequence as the position information of the curb or lane centerline.
[0163] The initial first element sequence refers to a sequence consisting of pixel coordinates of a plurality of points of a first road element in the confidence map of the first road element.
[0164] In some embodiments, the first device can obtain at least one initial first element sequence based on the confidence map of the first road element through various methods. For example, the confidence map can be traversed to select pixels with pixel values above a set threshold as target pixels, and clustering can be performed based on the target pixels to obtain multiple initial first element sequences (i.e., sets of non-background pixels). The set threshold can be determined based on experimentation or experience.
[0165] In some embodiments, the first device can update the initial first element sequence in various ways based on the coordinate offset in the vector field map of the first road element to obtain the first element sequence. For example, the device can traverse each pixel in the initial first element sequence, determine the coordinate offset corresponding to the position information in the vector field map based on the pixel's position information, identify the pixel corresponding to the coordinate offset in the vector field map as the new pixel in the confidence map, and determine whether the pixel value of the new pixel in the confidence map is greater than a set threshold. If so, the pixel coordinates of the new pixel are added to the initial first element sequence.
[0166] The set threshold value may be a system preset value or a system default value.
[0167] In some embodiments, the first device may add all adjacent pixel points corresponding to each point in the initial first element sequence to the updated initial first element sequence based on a similar method as described above to form a complete curb or lane centerline.
[0168] In some embodiments, pixel points that are adjacent in position in the initial position may be regarded as adjacent pixel points.
[0169] In some embodiments, the first device may use the updated initial first element sequence as the first element sequence.
[0170] In some embodiments, the first device may determine a position sequence of all curbs or lane centerlines in the driving environment image based on the above method to form an array of curbs or lane centerlines.
[0171] In some embodiments of the present specification, through the confidence map and vector field map of the first road element, the specific position information of the curb, lane centerline and lane line can be effectively extracted from various confidence maps and organized into a structured array form, which is conducive to subsequent path planning and use by other autonomous driving modules.
[0172] Step S320: Obtain at least one second element sequence based on the first element sequence and the regression field map of the second road element.
[0173] The second element sequence is used to represent a position sequence belonging to a second road element. The second road element includes a lane line.
[0174] In some embodiments, the second element sequence may be a sequence consisting of position information of multiple points on a lane line.
[0175] In some embodiments, the second element sequence can be obtained through various methods based on the first element sequence and the regression field map of the second road element. For example, based on the position information of each point on the first road element, the coordinate offset of each point at the same position in the regression field map of the second road element can be determined. Based on the coordinate offset and the position information, the position information of each point on the lane line can be determined as the second element sequence.
[0176] In some embodiments, the coordinate offset of the position information can be determined in the regression field map of the second road element based on the first element sequence; and the second element sequence can be obtained based on the first element sequence and the coordinate offset corresponding to the position information.
[0177] In some embodiments, the position information of the corresponding points in the second road element can be updated or adjusted based on the first element sequence and the coordinate offset of each point in the first element sequence to obtain the second element sequence. For more details about this embodiment, please refer to the relevant description below.
[0178] In some embodiments, based on the first element sequence, a position sequence of at least one lane centerline can be determined, where the position sequence of the lane centerline is a sequence composed of position information of multiple points on the lane centerline; based on the position sequence of the lane centerline, the coordinate offset corresponding to each position information is determined in the regression field map of the lane line; based on the position sequence of the lane centerline and the coordinate offset of the position information, the position sequence of the lane line is obtained.
[0179] In some embodiments, position information of multiple points on a lane centerline can be determined from multiple first element sequences as a lane centerline position sequence.
[0180] In some embodiments, for each lane centerline, the first device can traverse each point on the lane centerline, and for each point on the lane centerline, take the point as the target point and search for the corresponding coordinate offset on the regression field map based on the position information of the target point.
[0181] In some embodiments, the first device can determine the position information of a point on the lane line (e.g., the left lane line or the right lane line) corresponding to the target point based on the position information of the target point and the coordinate offset of the target point in a certain direction (e.g., the left or right side of the target point), and update the position information of the corresponding point in the corresponding initial second element sequence based on the position information of the point on the lane line (e.g., the left lane line or the right lane line).
[0182] The initial second element sequence is a sequence of pixel coordinates of a plurality of points on the initially determined lane line. For example, the initial second element sequence is a sequence of pixel coordinates of a plurality of points on the lane line corresponding to a lane centerline in the confidence map of the second road element.
[0183] In some embodiments, based on the confidence map of the second road element, pixels with pixel values above a set threshold can be selected as target pixels, and clustering can be performed based on the target pixels to obtain multiple initial second element sequences. The set threshold can be determined based on experiments or experience.
[0184] In some embodiments, the first device may determine a point on the lane line corresponding to each point on the lane centerline based on a similar method as described above, and update the corresponding initial second element sequence based on the position information of the point to obtain a second element sequence.
[0185] In some embodiments, the first device may update (eg, replace, modify, or add, etc.) the position information of the corresponding point in the initial second element sequence based on the position information of the determined lane line point to obtain the second element sequence.
[0186] In some embodiments, a lane centerline may correspond to one or more lane lines. In some embodiments, one or more second element sequences corresponding to the same lane centerline may be determined based on the above method, and the first device may use the one or more second element sequences determined based on the same lane centerline as the position information of the lane line (e.g., the left lane line or the right lane line) corresponding to the lane centerline.
[0187] In some embodiments, the first device may determine the second element sequence corresponding to all lane centerlines in the driving environment image based on the above method. The lane line points may be distributed on both sides of the length direction of the lane centerline to form multiple complete lane lines.
[0188] In some embodiments of the present specification, the regression field map of the second road element can effectively extract the specific location information of the curb, lane centerline and lane line from various confidence maps, and organize it into a structured array form for use by subsequent path planning and other autonomous driving modules.
[0189] Step S330 : determining position information of a plurality of road elements in the driving environment image based on at least one first element sequence and at least one second element sequence.
[0190] In some embodiments, multiple first element sequences and multiple second element sequences may be used as position information of multiple road elements in the driving environment image.
[0191] In some embodiments of this specification, each lane centerline corresponds to at least two lane lines. By first determining the position sequence of the lane centerlines, it helps to locate the position of the corresponding lane and reduce the amount of calculation for calculating the position sequence of the lane lines.
[0192] Figure 4 This is an exemplary schematic diagram of the network structure of the detection model shown in some embodiments of this specification.
[0193] In some embodiments, as Figure 4 As shown, the road elements include a first road element and a second road element, and the detection model 420 may include a first network layer 421 and a second network layer 422. The bird's-eye view feature map 410 may be input into the first network layer 421 of the detection model to obtain a structural feature map of the first road element; and the bird's-eye view feature map 410 may be input into the second network layer 422 of the detection model to obtain a structural feature map of the second road element.
[0194] In some embodiments, as Figure 4 As shown, the structural feature map of the first road element includes a confidence map 431 of the first road element and a vector field map 432 of the first road element; the structural feature map of the second road element includes a confidence map 433 of the second road element and a regression field map 434 of the second road element.
[0195] Exemplarily, the structural feature map of the first road element includes a structural feature map of the roadside and / or a structural feature map of the lane centerline, wherein the structural feature map of the roadside includes a confidence map of the roadside and a vector field map of the roadside, and the structural feature map of the lane centerline includes a confidence map of the lane centerline and a vector field map of the lane centerline; the structural feature map of the second road element includes a structural feature map of the lane line, wherein the structural feature map of the lane line includes a confidence map of the lane line and a regression field map of the lane line.
[0196] The first network layer is a model for determining a structural feature graph of the first road element.
[0197] In some embodiments, the first network layer may be a machine learning model, for example, the first network layer may be a convolutional neural network (CNN).
[0198] In some embodiments, the input of the first network layer may include a bird's-eye view feature map of the driving environment image; and the output may include a confidence map and a vector field map of the first road element.
[0199] The second network layer is a model for determining the structural feature graph of the second road element.
[0200] In some embodiments, the second network layer may be a machine learning model, for example, the second network layer may be a convolutional neural network (CNN).
[0201] In some embodiments, the input of the second network layer may include a bird's-eye view feature map of the driving environment image; and the output may include a confidence map and a vector field map of the second road element.
[0202] In some embodiments, the detection model can be obtained by jointly training the first network layer and the second network layer based on a large number of labeled training samples. The training samples used for joint training include bird's-eye view feature maps of the sample images. The training samples can be obtained based on historical data.
[0203] An exemplary joint training process involves inputting the bird's-eye view feature maps of sample images from the training sample into the initial first and second network layers, respectively. A loss function is constructed based on the outputs and labels of the initial first and second network layers, and the parameters of the initial first and second network layers are simultaneously updated until a preset condition is met and training is complete. The preset condition can be that the loss function is less than a threshold, converges, or the training cycle reaches a threshold.
[0204] In some embodiments of this specification, the joint training of the initial first network layer and the initial second network layer is helpful in solving the problem of difficulty in obtaining labels when training the adjustment model alone, improving the training efficiency of the adjustment model, and reducing the training difficulty.
[0205] In some embodiments, the first network layer and the second network layer can be trained separately. The method of training the first network layer and the second network layer separately is similar to the method of training the detection model. For more information about training the detection model, please refer to Figure 2 Related description.
[0206] In some embodiments, the first network layer can be trained based on a large number of first training samples with the first label through various feasible methods. For example, parameter update can be performed based on a gradient descent method.
[0207] In some embodiments, the first training samples may include multiple groups of training samples, each group of first training samples including at least a bird's-eye view feature map of the sample image. The training samples may be obtained based on historical data.
[0208] In some embodiments, the first label may be a sample confidence map and a sample vector field map of the first road element corresponding to the sample image. The first label may be obtained based on automatic or manual annotation.
[0209] In some embodiments, the first label may include a first truth value matrix for the first road element and a second truth value matrix for the curb.
[0210] In some embodiments, the second network layer can be trained based on a large number of second training samples with second labels through various feasible methods. For example, parameter update can be performed based on a gradient descent method.
[0211] In some embodiments, the second training samples may include multiple groups of training samples, each group of first training samples including at least a bird's-eye view feature map of the sample image. The training samples may be obtained based on historical data.
[0212] In some embodiments, the second label may be a sample confidence map and a sample regression field map of the second road element corresponding to the sample image. The second label may be obtained based on automatic or manual annotation.
[0213] In some embodiments, the second label may include a first truth matrix of the second road element and a third truth matrix of the lane line.
[0214] For more information about the first truth matrix, the second truth matrix, and the third truth matrix, please refer to the relevant description below.
[0215] In some embodiments, training samples can be obtained and labeled to obtain a truth matrix, which includes at least one of a first truth matrix, a second truth matrix, and a third truth matrix; wherein the first truth matrix is used to represent feature points belonging to the lane centerline or curb, the second truth matrix is used to represent the relative position information of the feature points in the lane centerline or curb and their adjacent pixel points, and the third truth matrix is used to represent the orthogonal projection information of each pixel point; the training samples are used to train the detection model to obtain a trained detection model.
[0216] The training sample is at least a part of the sample set used to train the detection model.
[0217] The truth matrix is a matrix used to represent actual labels. In some embodiments, the truth matrix can be a two-dimensional matrix, where each element represents the true label of whether a pixel in the image belongs to a road element. In some embodiments, the truth matrix is a binary matrix, where 1 indicates that the pixel belongs to a road element and 0 indicates that it does not belong to a road element.
[0218] In some embodiments, in the first truth value matrix, if the pixel point corresponding to an element belongs to the lane centerline or the curb, the element is 1; otherwise, it is 0.
[0219] In some embodiments, the non-zero value in the second truth matrix represents the relative position information between the pixel point corresponding to the element and its adjacent pixel points.
[0220] In some embodiments, each element in the third truth matrix represents the orthogonal distance between the pixel point corresponding to the element and the lane line.
[0221] In some embodiments of this specification, training samples are labeled through the relationship between confidence maps, vector field maps, and regression field maps, which helps improve the accuracy of labeling, improves the quality of training samples, and helps the model learn the association relationship between road elements.
[0222] Figure 5 is an exemplary flowchart of labeling a first truth matrix according to some embodiments of this specification.
[0223] In some embodiments, process 500 may be performed by a second device based on a road detection method. Figure 5 As shown, process 500 includes the following steps.
[0224] Step S510: Obtain at least one sample element sequence from a preset road element list.
[0225] The preset road element list is a dataset obtained by manually annotating road elements in a bird's-eye view feature map of a sample image. For example, the preset road element list may be a list obtained by manually annotating lane centerlines.
[0226] In some embodiments, the preset road element list may include the location information and type (such as curb, lane centerline, lane line, etc.) of each road element marked in the bird's-eye view feature map of the sample image.
[0227] In some embodiments, the preset road element list may include multiple sample element sequences, each of which is a sequence of position information for multiple points on initially annotated road elements. A sample element sequence represents the position information (e.g., pixel coordinates in a bird's-eye view feature map of a sample image) of a series of points that make up a road element (e.g., a curb, lane centerline, or lane marking).
[0228] The feature point in the sample element sequence is used to represent a certain point of the initially marked road element.
[0229] In some embodiments, a sample element sequence of any roadside or lane centerline can be obtained from a preset road element list. For example, a sample element sequence of any lane centerline can be obtained from a preset road element list to form a lane centerline sample element sequence.
[0230] Step S520 : Based on the position information of each feature point in the sample element sequence, a first line connecting points corresponding to each feature point is drawn on the first preset matrix to obtain a marked first preset matrix.
[0231] In some embodiments, each sample element sequence may be traversed to determine the number of feature points contained in each sample element sequence. If the number of feature points in a sample element sequence is less than a threshold value (e.g., 2), the sample element sequence may be ignored or a warning may be issued. The threshold value may be determined based on experiments or experience.
[0232] The first preset matrix is a pre-set matrix. The size of the first preset matrix can be the same as the size of the first true value matrix. In some embodiments, the first preset matrix is a matrix of all zeros.
[0233] The first line is composed of a plurality of points of the first road element in the first preset matrix. The marked first preset matrix refers to the first preset matrix after the line is drawn.
[0234] In some embodiments, based on the position information of each feature point in the sample element sequence, a first line connecting the points corresponding to each feature point can be drawn on the first preset matrix in various ways. For example, for each sample element sequence with a point count greater than or equal to a point count threshold, the matching point of each feature point in the sample element sequence is determined in the first preset matrix based on the position information of each feature point in the sample element sequence, and a first line connecting each pair of adjacent matching points is drawn in the first preset matrix. The value of the first line can be determined according to predefined rules, such as a constant value or a weighting function. The width of the first line can be adjusted according to actual conditions to reflect the width or other attributes of the road element.
[0235] The matching points refer to corresponding feature points on other objects (eg, the first preset matrix) having the same position information as the feature points in the sample element sequence.
[0236] In some embodiments, based on the pixel coordinates of each matching point, a line connecting each pair of adjacent matching points can be drawn on the first preset matrix using a drawing algorithm such as Bresenham to obtain a line marked by a certain sample element sequence in the first preset matrix.
[0237] In some embodiments, drawing may also be setting the values of the matching points in the first preset matrix to preset values, which may be manually preset values, system preset values, and the like.
[0238] In some embodiments, based on the above method, a line marked in the first preset matrix for each sample element sequence can be obtained, and the first preset matrix is used as the marked first preset matrix.
[0239] Step S530: Obtain a first truth matrix based on the marked first preset matrix.
[0240] In some embodiments, the drawn first preset matrix may be used as the first true value matrix.
[0241] In some embodiments of this specification, a first truth matrix is obtained through automated labeling, which improves the accuracy of the labels of the curb or lane centerline obtained, facilitates the subsequent training of the detection model, and improves the training effect of the detection model.
[0242] Figure 6 is an exemplary flow chart of labeling the second truth matrix according to some embodiments of this specification. In some embodiments, process 600 can be executed by a second device based on the road detection method. Figure 6 As shown, process 600 includes the following steps.
[0243] Step S610: Generate a first heat map based on a preset first road element list.
[0244] The preset first road element list is a data set obtained by manually annotating the first road elements in the bird's-eye view feature map of the sample image. For example, the preset first road element list can be a list obtained by manually annotating lane centerlines.
[0245] In some embodiments, the preset first road element list may include each first element sequence and type (such as a curb, a lane centerline, etc.) marked in the bird's-eye view feature map of the sample image.
[0246] In some embodiments, the preset first road element list may include multiple sample element sequences, where a sample element sequence represents a sequence consisting of position information of a series of points constituting a lane centerline (e.g., pixel coordinates in a bird's-eye view feature map of a sample image, etc.).
[0247] Each element in the first heat map represents the intensity or probability of a certain attribute at the corresponding location in the bird's-eye view feature map of the sample image. For example, the value of each pixel in the first heat map represents the probability that the pixel belongs to the lane centerline or the curb. The larger the pixel value, the greater the probability of belonging to the centerline or curb.
[0248] In some embodiments, the first heat map may be a two-dimensional array having the same size as the second truth matrix.
[0249] In some embodiments, the first heat map can be generated based on a preset first road element list in various ways. For example, each feature point in a sample element sequence in the preset first road element list can be traversed, and the pixel value of the point corresponding to the feature point on the first heat map can be set to a specified value (e.g., 1 or other predefined value).
[0250] In order to make the heat map smoother, in some embodiments, a Gaussian kernel may be applied around the pixel points corresponding to the feature points so that the values of the nearby pixels gradually decrease.
[0251] In some embodiments, based on the above approach, heat maps may be generated for each lane centerline or roadside in the preset first road element list in different regions of the same image.
[0252] Step S620: Mark the first heat map to obtain a plurality of marked pixels in the first heat map.
[0253] The marked pixel points refer to the pixel points initially determined to be part of the first road element.
[0254] In some embodiments, the first heat map can be labeled in various ways to obtain a plurality of labeled pixels in the first heat map. For example, based on the position information of each feature point in the sample element sequence, the pixel value of the pixel corresponding to the feature point in the heat map can be changed to highlight the feature point. For example, the color of the pixel can be changed to red, blue, or other eye-catching colors to distinguish it from other data.
[0255] In some embodiments, based on the position information of each feature point in the first heat map, a second line connecting each feature point can be drawn; based on the second line, a plurality of marked pixel points can be obtained.
[0256] The second line is composed of multiple points of the first road element in the first heat map.
[0257] In some embodiments, based on the position information of each feature point in a certain sample element sequence, the corresponding feature point of each feature point in the sample element sequence in the first heat map is determined, and based on a method similar to drawing the first line, a second line connecting each pair of adjacent feature points is drawn on the first heat map.
[0258] In some embodiments, each point on the second line may be used as a marked pixel point.
[0259] In some embodiments, the second line of each sample element sequence on the first heat map may be determined based on the above method.
[0260] In some embodiments of this specification, by marking pixel points, the curb or lane centerline is more accurately identified, thereby improving the accuracy of the perception of the lane centerline or curb.
[0261] Step S630: Calculate the relative position information of each marked pixel and its related pixels.
[0262] The relative position information may include the coordinate offset of the two points in the horizontal direction, the coordinate offset in the vertical direction, etc.
[0263] In some embodiments, the relative position information of each marked pixel and its related pixel may be calculated in a variety of ways.
[0264] Related pixel points refer to other pixel points associated with the marked pixel point in the first heat map.
[0265] The association can be neighboring pixels that are adjacent in position or located in the same neighborhood.
[0266] Neighborhood pixels refer to pixels within the neighborhood of a certain pixel. Neighborhoods can include 4-neighborhoods, 8-neighborhoods, diagonal neighborhoods, etc. Correspondingly, neighborhood pixels include pixels within the 4-neighborhood, 8-neighborhood, or diagonal neighborhood of a certain pixel. For example, the coordinates of pixel P are (x, y), and the coordinates of the four pixels within the 4-neighborhood of pixel P are (x+1, y), (x-1, y), (x, y+1), and (x, y-1).
[0267] In some embodiments, the distance between each marked pixel and the pixels in its neighborhood may be calculated as relative position information.
[0268] In some embodiments, based on the marked pixel points, the previous pixel point or the next pixel point of each marked pixel point can be determined as the related pixel point of the marked pixel point; based on the marked pixel point and the related pixel points, the valid marked point can be determined; based on the valid marked point, the relative position information of the valid marked point and its related pixel points can be determined.
[0269] The preceding pixel or following pixel refers to the preceding pixel or following pixel of a marked pixel.
[0270] For each marked pixel, the index information of the feature point corresponding to the marked pixel in the sample element sequence is determined, and based on the index information of each marked pixel, the previous pixel or the next pixel of each marked pixel is determined.
[0271] The index information refers to the specific position of a point in a sequence or list. For example, the index of the first point is 0, the index of the second point is 1, and so on.
[0272] Valid marking points refer to marking pixels that meet the requirements.
[0273] In some embodiments of the present specification, by screening and marking pixel points, invalid pixel points can be ignored, and false detections due to isolated or abnormal points can be prevented, thereby improving the stability of the overall system.
[0274] In some embodiments, valid marking points may be determined in a variety of ways based on the marking pixel points and the related pixel points.
[0275] In some embodiments, if there are preceding and succeeding pixels of a marked pixel, and the minimum orthogonal distance between the marked pixel and the relevant pixel is less than a first preset threshold, the marked pixel is regarded as a valid marked point.
[0276] In some embodiments, if there are no preceding and following pixels of the marked pixel, or if the minimum orthogonal distance between the marked pixel and the related pixel is greater than or equal to a first preset threshold, the marked pixel is ignored.
[0277] In some embodiments of this specification, by eliminating invalid marking points, the amount of calculation can be reduced and the efficiency of determining the second truth matrix can be improved.
[0278] In some embodiments, the relative position information of a valid marker point and its associated pixel can be determined in various ways based on the valid marker point. For example, the relative position information of the valid marker point and the preceding pixel can be calculated as the first relative position information of the valid marker point, and the coordinate offset between the valid marker point and the following pixel can be calculated as the second relative position information.
[0279] Step S640: Based on the relative position information, update the second preset matrix to obtain a second true value matrix.
[0280] The second preset matrix is a pre-set matrix. The size of the second preset matrix can be the same as the size of the second true value matrix. In some embodiments, the second preset matrix is a matrix of all zeros.
[0281] In some embodiments, the second preset matrix may include a first submatrix and a second submatrix, wherein the first submatrix may store relative position information between the marked pixel and its preceding pixel, and the second submatrix may store relative position information between the marked pixel and its succeeding pixel.
[0282] In some embodiments, the first relative position information of all valid marking points can be stored in the corresponding first sub-matrix, and the second relative position information of all valid marking points can be stored in the corresponding second sub-matrix. The first sub-matrix and the second sub-matrix storing the relative position information of each valid marking point are used as the second truth matrix.
[0283] In some embodiments of this specification, by determining the relative position information of the effective marking point and its related pixel points, the positional relationship between each feature point in the vector field map and the corresponding next feature point can be determined, which helps to improve the accuracy of the labeled second truth matrix.
[0284] In some embodiments of this specification, a heat map can be used to highlight key locations of the center line or roadside, thereby improving detection accuracy; by calculating the relative position information between each feature point and its adjacent points, adjacent pixel points can be more accurately connected to form a continuous path or line; and obtaining an annotated vector field map using relative position information can improve the accuracy and robustness of lane line detection output by the detection model.
[0285] In some embodiments, the preset mask matrix may be updated based on the relative position information to obtain a mask matrix corresponding to the second true value matrix.
[0286] The preset mask matrix is a pre-set mask matrix. In some embodiments, the preset mask matrix can be a two-dimensional array of the same size as the second true value matrix, the preset mask matrix can be an all-0 matrix, or the preset mask matrix can be other values, depending on the application scenario.
[0287] The mask matrix corresponding to the second truth matrix is used to mark and extract the area belonging to the first road element in the bird's-eye view feature map of the sample image.
[0288] In some embodiments, the pixel value at the position corresponding to each valid marker point in the preset mask matrix may be updated (for example, set to 1) to obtain a mask matrix corresponding to the second true value matrix.
[0289] In some embodiments, based on the valid marking points corresponding to the calculated relative position information, the value of the pixel point corresponding to each valid marking point can be updated in the second marking matrix (for example, the pixel value is set to 1) to prevent repeated processing of the same pixel point.
[0290] The second labeling matrix is a preset labeling matrix corresponding to the second truth matrix.
[0291] The second tag matrix may be a two-dimensional array with the same size as the second truth matrix. In some embodiments, all elements of the second tag matrix are zero.
[0292] In some embodiments of this specification, through the mask matrix, during training, only the area where the mask is located is supervised, thereby avoiding training on invalid data and reducing unnecessary calculations in the training stage.
[0293] By updating the marking matrix, repeated processing of the same position can be avoided, thereby improving the processing speed of the second truth matrix; marking the processed points can effectively prevent repeated calculations, which helps the system be applied in large data sets or high-resolution images.
[0294] Figure 7 is an exemplary flow chart of labeling the third truth matrix according to some embodiments of this specification. In some embodiments, process 700 can be executed by a second device based on the road detection method. Figure 7 As shown, process 700 includes the following steps.
[0295] Step S710: Generate a second heat map based on the preset lane list.
[0296] The preset lane list is a data set obtained by manually marking a lane.
[0297] In some embodiments, the preset lane list may include multiple target sequences. A target sequence is a sequence consisting of the position information of multiple points on a target road element. A target sequence may include the position information of a series of feature points that constitute a lane line or lane centerline (e.g., pixel coordinates in a bird's-eye view feature map of a sample image).
[0298] The feature points in the target sequence are used to represent a certain point on the lane line or lane centerline of the initially marked lane.
[0299] Each pixel in the second heat map represents the intensity or probability of a certain attribute at the corresponding position in the second road element list. For example, the value of a pixel in the second heat map represents the probability that the feature point in the target sequence corresponding to the pixel belongs to a lane line.
[0300] In some embodiments, the second heat map may be a two-dimensional array having the same size as the third truth matrix.
[0301] In some embodiments, the second heat map can be generated based on a preset lane list in a variety of ways. For example, for a non-empty target sequence, each feature point in the target sequence can be traversed and the pixel value of the corresponding feature point on the second heat map can be set to a specified value (e.g., 1 or other predefined value).
[0302] In order to make the heat map smoother, in some embodiments, a Gaussian kernel may be applied around the pixel points corresponding to the feature points in the target sequence, so that the values of the nearby pixels gradually decrease.
[0303] In some embodiments, based on the above approach, heat maps may be generated for each lane centerline or lane line in the preset lane list in different regions of the same image.
[0304] Step S720: Mark the second heat map to obtain a separation line in the second heat map.
[0305] The separator line is a line in the second heat map that represents a lane line (eg, a left lane line or a right lane line) or a lane center line.
[0306] In some embodiments, the second heat map can be labeled in a variety of ways to obtain a separation line in the second heat map. For example, the pixel coordinates of lane lines or lane center lines can be determined based on the second heat map using a support vector machine, random forest, or deep learning model, and the separation line can be drawn on the second heat map based on the pixel coordinates of the lane lines or lane center lines.
[0307] In some embodiments, the position information of each feature point in the second heat map can be used to draw a third line connecting each feature point; and based on the third line in the second heat map, a separation line in the second heat map can be obtained.
[0308] In some embodiments, based on the position information of each feature point in the target sequence, the corresponding feature point of each feature point in the target sequence in the second heat map can be determined, and based on a method similar to drawing the first line, each pair of adjacent feature points is drawn on the second heat map to obtain a separation line in the second heat map.
[0309] In some embodiments, a separation line for each target sequence on the second heat map may be determined based on the above method.
[0310] In some embodiments of this specification, by determining the separation line in the second heat map, the relative position relationship between different road elements is more accurately determined, thereby improving the accuracy of lane line perception.
[0311] Step S730: Determine the orthogonal projection point of each pixel in the second heat map on the separation line.
[0312] In some embodiments, the orthogonal projection point of each pixel in the second heat map on the dividing line can be determined in a variety of ways. For example, the shortest distance from each pixel in the second heat map to the dividing line can be calculated as the orthogonal projection point of the pixel on the dividing line. For another example, the orthogonal projection point of each pixel on the dividing line can be determined using a search algorithm. In some embodiments, the search algorithm includes but is not limited to a greedy algorithm, a random walk algorithm, an A* algorithm, an enumeration algorithm, a depth-first search algorithm, a breadth-first search algorithm, and the like.
[0313] In some embodiments, the projection coefficient of each pixel point in the second heat map on the dividing line can be calculated; based on the projection coefficient and the second preset threshold, the starting point or end point of the dividing line is updated, or the orthogonal projection point of each pixel point in the second heat map on the dividing line is determined.
[0314] The second preset threshold may be determined based on experiments or experience.
[0315] In some embodiments, if the projection coefficient of a pixel point on the dividing line is between 0 and 1, it means that the pixel point can be projected onto the dividing line; then the orthogonal projection point of each pixel point on the dividing line in the second heat map.
[0316] In some embodiments, if the projection coefficient of a pixel on the dividing line is greater than 1 or less than 0, it means that the pixel is located on the extension line of the dividing line, and the dividing line can be updated based on the relative distance between the pixel and the starting point or end point of the dividing line. For example, if the relative distance between the pixel and the starting point of the dividing line is small, the pixel is used as the starting point of the dividing line; if the relative distance between the pixel and the end point of the dividing line is small, the pixel is used as the end point of the dividing line.
[0317] In some embodiments of this specification, by calculating projection coefficients, lane edge situations can be better handled. For example, when a projected point is not within the current line segment, the starting point or direct current of the lane line can be correctly adjusted to ensure that all points are correctly projected onto the nearest lane line.
[0318] Step S740: Based on the orthogonal distance between each pixel point and its orthogonal projection point, update the third preset matrix to obtain a third true value matrix.
[0319] The third preset matrix is a pre-set matrix. The size of the third preset matrix can be the same as the size of the third true value matrix. In some embodiments, the third preset matrix is a matrix of all zeros.
[0320] In some embodiments, the third preset matrix can be updated in a variety of ways based on the orthogonal distance between each pixel point and its orthogonal projection point to obtain an updated third true value matrix.
[0321] In some embodiments, if the orthogonal distance between each pixel point and its orthogonal projection point is less than the pixel value of the pixel point corresponding to the third marking matrix, the orthogonal distance between each pixel point and its orthogonal projection point can be stored in the corresponding third preset matrix.
[0322] The third labeling matrix refers to a preset labeling matrix corresponding to the third truth matrix.
[0323] In some embodiments, the corresponding pixel points may be marked in the third marking matrix based on the pixel points corresponding to the calculated orthogonal distances (eg, the pixel values are set to 1 or other values) to prevent repeated processing of the same pixel point.
[0324] The third tag matrix may be a two-dimensional array having the same size as the third truth matrix. In some embodiments, all elements in the third tag matrix may be set to a larger value.
[0325] In some embodiments of the present specification, repeated processing of the same position can be avoided by updating the marking matrix, thereby improving the processing speed of the third truth matrix; marking the processed points can effectively prevent repeated calculations, which helps the system to be used in large data sets or high-resolution images.
[0326] In some embodiments of the present specification, by calculating the orthogonal distance from each pixel point to the center line of the nearest lane line and its projection point, the lane to which each pixel point belongs can be determined, which is conducive to determining each lane and its corresponding lane line; the annotated regression field map can enhance the feature expression capability, making it easier for the detection model to learn the key features of the lane line, thereby improving detection accuracy.
[0327] It should be noted that the above description of the relevant processes is for illustration and purpose only and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the processes under the guidance of this specification. However, such modifications and changes are still within the scope of this specification.
[0328] Figure 9 This is a schematic diagram of the structure of an electronic device according to some embodiments of this specification. Figure 9 As shown, the electronic device 900 may include a processor 901 and a memory 902. The electronic device 900 may also include one or more of a multimedia component 903, an input / output (I / O) component 904, and a communication component 905. In this embodiment, the electronic device 900 may be a device integrated into a vehicle to implement the road detection method provided in this embodiment.
[0329] The processor 901 is used to control the overall operation of the electronic device 900 to complete all or part of the steps in the above-mentioned road detection method. The memory 902 is used to store various types of data to support the operation of the auxiliary electronic device 900. This data may include, for example, instructions for any application or method operating on the auxiliary electronic device 900, as well as application-related data, such as contact information, sent and received messages, pictures, audio, video, etc. The memory 902 can be implemented by any type of volatile or non-volatile storage module or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 903 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 902 or transmitted via the communication component 905. The audio component also includes at least one speaker for outputting audio signals. The I / O component 904 provides an interface between the processor 901 and other interface modules, which may include a keyboard, a mouse, buttons, etc. These buttons may be virtual or physical buttons. The communication component 905 is used for wired or wireless communication between the electronic device 900 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, Narrow Band Internet of Things (NB-IOT), Enhanced Machine-Type Communication (eMTC), or other 5G technologies, or a combination thereof, is not limited here. Therefore, the corresponding communication component 905 may include: a Wi-Fi module, a Bluetooth module, an NFC module, etc.
[0330] In an exemplary embodiment, the electronic device 900 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned road detection method.
[0331] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-mentioned road detection method are implemented. For example, the computer-readable storage medium may be the aforementioned memory 902 including the program instructions. The program instructions may be executed by the processor 901 of the electronic device 900 to implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0332] Alternatively, the instructions, when executed by a computer, enable execution to implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.
[0333] The present application also provides a vehicle, which is equipped with the electronic device provided by any of the above embodiments, or can execute any of the road detection methods provided by the embodiments of the present application, and the electronic device is used to execute the road detection method provided by any of the above embodiments.
[0334] In one embodiment, a vehicle can be configured for a fully or partially autonomous driving mode. For example, while in autonomous driving mode, the vehicle can control itself and, through human interaction, determine the current state of the vehicle and its surroundings, determine the possible behavior of at least one other vehicle in the surroundings, and determine a confidence level corresponding to the likelihood that the other vehicle will perform the possible behavior, and control the vehicle based on this information. While in autonomous driving mode, the vehicle can be configured to operate without human interaction.
[0335] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0336] The embodiments, implementation methods and related technical features of the present application can be combined and replaced with each other without conflict.
[0337] The above are only preferred embodiments of the present application and do not constitute any form of limitation to the present application. Although the descriptions of each embodiment in the embodiments of the present application have different focuses, for parts that are not described in detail in a certain embodiment, please refer to the relevant embodiments of other embodiments. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.
Claims
1. A road detection method, characterized in that: The method comprises: Acquire a vehicle's driving environment image; Processing the driving environment image to obtain a structural feature map of a plurality of road elements in the driving environment image, wherein the structural feature map is used to indicate aggregated information of the road elements and association relationships between the plurality of road elements; Position information of a plurality of road elements in the driving environment image is determined according to the structural feature graphs of the plurality of road elements.
2. The method according to claim 1, characterized in that The processing of the driving environment image to obtain a structural feature map of a plurality of road elements in the driving environment image includes: performing feature extraction on the driving environment image to obtain two-dimensional features of the driving environment image; Performing feature conversion on the two-dimensional features to obtain a bird's-eye view feature map of the driving environment image; The bird's-eye view feature map is processed to obtain a structural feature map of multiple road elements in the driving environment image.
3. The method according to claim 2, characterized in that The processing of the bird's-eye view feature map to obtain a structural feature map of a plurality of road elements in the driving environment image includes: The bird's-eye view feature map is processed by the trained detection model to obtain a structural feature map of multiple road elements in the driving environment image.
4. The method according to claim 3, characterized in that The detection model includes a first network layer and a second network layer, and the road element includes a first road element and a second road element. The trained detection model is used to process the bird's-eye view feature map to obtain a structural feature map of multiple road elements in the driving environment image, including: Inputting the bird's-eye view feature map into the first network layer to obtain a structural feature map of the first road element; The bird's-eye view feature map is input into the second network layer to obtain a structural feature map of the second road element.
5. The method according to claim 4, characterized in that The structural feature map of the first road element includes a confidence map and a vector field map of the first road element; The structural feature map of the second road element includes a confidence map and a regression field map of the second road element; Among them, the confidence map is used to represent the position probability map of road elements, the vector field map is used to represent the coordinate offset between the feature points belonging to the same road element and its adjacent pixel points, and the regression field map is used to represent the relative position relationship between different road elements.
6. The method according to claim 3, characterized in that The detection model is trained through the following steps: Acquire training samples and annotate the training samples to obtain a truth matrix, wherein the truth matrix includes at least one of a first truth matrix, a second truth matrix, and a third truth matrix; wherein the first truth matrix is used to represent the probability that each pixel belongs to the first road element, the second truth matrix is used to represent relative position information of a feature point belonging to the first road element and its adjacent pixels, and the third truth matrix is used to represent orthogonal projection information of each pixel onto a dividing line; The detection model is trained using the training samples to obtain the trained detection model.
7. The method according to claim 6, characterized in that The step of obtaining training samples and labeling the training samples to obtain a true value matrix includes: Acquire at least one sample element sequence from a preset road element list, where the sample element sequence is a sequence consisting of position information of a plurality of points on an initially marked road element; Based on the position information of each feature point in the sample element sequence, drawing a first line connecting points corresponding to each feature point in the first preset matrix to obtain a marked first preset matrix; Based on the marked first preset matrix, the first true value matrix is obtained.
8. The method according to claim 6, characterized in that The training samples are obtained and labeled to obtain a true value matrix, including: generating a first heat map based on a preset first road element list; Marking the first heat map to obtain a plurality of marked pixels in the first heat map; Calculating the relative position information of each marked pixel and its related pixel; Based on the relative position information, the second preset matrix is updated to obtain the second true value matrix.
9. The method according to claim 8, characterized in that The step of marking the first heat map to obtain a plurality of marked pixels in the first heat map includes: Based on the position information of each feature point in the first heat map, drawing a second line connecting each of the feature points; Based on the second line, a plurality of marked pixel points are obtained.
10. The method according to claim 6, characterized in that The step of obtaining training samples and labeling the training samples to obtain a true value matrix includes: Based on the preset lane list, a second heat map is generated; Marking the second heat map to obtain a separation line in the second heat map; Determine the orthogonal projection point of each pixel in the second heat map on the separation line; Based on the orthogonal distance between each pixel point and its orthogonal projection point, the third preset matrix is updated to obtain the third true value matrix.
11. The method according to claim 10, characterized in that The marking the second heat map to obtain a separation line in the second heat map includes: Based on the position information of each feature point in the second heat map, drawing a third line connecting each of the feature points; A separation line in the second heat map is obtained based on the third line in the second heat map.
12. The method according to claim 10, characterized in that Determining the orthogonal projection point of each pixel point in the second heat map on the separation line includes: Calculate the projection coefficient of each pixel point in the second heat map on the separation line; Based on the projection coefficient and the second preset threshold, the starting point or the end point of the separation line is updated, or Determine the orthogonal projection point of each pixel in the second heat map on the separation line.
13. The method according to claim 1, wherein The determining, based on the structural feature graphs of the multiple road elements, position information of the multiple road elements in the driving environment image includes: obtaining at least one first element sequence based on the confidence map and the vector field map of the first road element, where the first element sequence is a sequence consisting of position information of a plurality of points on the first road element; obtaining at least one second element sequence based on the first element sequence and the regression field map of the second road element, where the second element sequence is a sequence consisting of position information of a plurality of points on the second road element; Based on the at least one first element sequence and the at least one second element sequence, position information of a plurality of road elements in the driving environment image is determined.
14. The method according to claim 13, characterized in that The obtaining of at least one first element sequence based on the confidence map and the vector field map of the first road element includes: obtaining at least one initial first element sequence based on the confidence map of the first road element; The initial first element sequence is updated based on the coordinate offset in the vector field map of the first road element to obtain the first element sequence.
15. The method according to claim 13, characterized in that The obtaining of at least one second element sequence based on the first element sequence and the regression field map of the second road element comprises: Determining, based on the first element sequence, a coordinate offset corresponding to each piece of the position information in a regression field map of the second road element; The second element sequence is obtained based on the first element sequence and the coordinate offset of the position information.
16. The method according to any one of claims 4 to 15, characterized in that The first road element includes a lane centerline and / or a curb, and the second road element includes a lane line.
17. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer-readable storage medium stores instructions, which, when executed by a computer, enable the computer to implement the road detection method according to any one of claims 1 to 16.
18. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, enable the computer to implement the road detection method according to any one of claims 1 to 16.
19. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the road detection method according to any one of claims 1 to 16.
20. A vehicle, characterized in that: The electronic device comprises the electronic device according to claim 19, or can execute the road detection method according to any one of claims 1 to 16.