Road marking line detection model training, road marking line detection method and autonomous driving vehicle
By dividing road markings into sub-road markings and using the cross-attention mechanism to train the model, the problem of insufficient accuracy of road markings in high-precision maps is solved, achieving higher detection accuracy and stability of autonomous driving.
Patent Information
- Application Number
- CN202211467762.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-22
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-11-22
AI Technical Summary
In existing technologies, the accuracy of road markings in high-precision maps is poor, which affects the accuracy of autonomous driving, and the reliance on depth data and vector conversion processes leads to precision loss.
By dividing road markings into multiple sub-road markings, using map feature extraction network and detection network, and training models based on environmental perception data, the sub-road markings and their categories are directly predicted, avoiding image conversion and deep data dependence, and using a cross-attention mechanism to improve detection accuracy.
The accuracy and precision of road marking lines in high-precision maps are improved, reducing dependence on data and ensuring the normal operation of autonomous driving.
Smart Images

Figure CN115761677B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, in particular to technical fields such as artificial intelligence, autonomous driving, and high-precision maps. Specifically, the present disclosure relates to a road marking line detection model training, a road marking line detection method, and an autonomous driving vehicle. Background Art
[0002] High-precision maps can provide vehicles with rich road topology information and traffic rules, and are a fundamental component of autonomous driving technology.
[0003] To provide more road information, high-precision maps typically include road markings. Poor accuracy of road markings in HD maps can negatively impact autonomous driving. Therefore, accurately constructing road markings in HD maps has become a key technical challenge in the field. Summary of the Invention
[0004] In order to address at least one of the above-mentioned deficiencies, the present disclosure provides a road marking line detection model training, a detection method, and an autonomous driving vehicle.
[0005] According to a first aspect of the present disclosure, a method for training a road marking line detection model is provided, the method comprising:
[0006] Acquire sample environmental perception data of the sample road environment and annotation information of a sample environmental map corresponding to the sample road environment, the annotation information including at least one sub-road marking line contained in each road marking line in the sample environmental map and a marking line category of the sub-road marking line;
[0007] Input the sample environment perception data into the map feature extraction network of the road sign line detection model to obtain the sample map features;
[0008] Inputting the sample map features into the detection network of the road sign line detection model to obtain the predicted sub-road sign line and the sign line category of the predicted sub-road sign line;
[0009] The road marking line detection model is trained based on the first difference and the second difference, where the first difference is the difference between the predicted sub-road marking line and the sub-road marking line, and the second difference is the difference between the marking line category of the predicted sub-road marking line and the marking line category of the sub-road marking line.
[0010] According to a second aspect of the present disclosure, a method for detecting road marking lines is provided, the method comprising:
[0011] Acquire environmental perception data of the road environment;
[0012] Inputting the environmental perception data into the map feature extraction network of the road sign line detection model to obtain map features, the road sign line detection model is trained using the above-mentioned road sign line detection model training method;
[0013] The map features are input into a detection network of a road marking line detection model to obtain at least one sub-road marking line and a marking line category of each road marking line in an environment map of the road environment.
[0014] According to a third aspect of the present disclosure, a road marking line detection model training device is provided, the device comprising:
[0015] a data acquisition module, configured to acquire sample environmental perception data of a sample road environment and annotation information of a sample environmental map corresponding to the sample road environment, the annotation information including at least one sub-road marking line contained in each road marking line in the sample environmental map and a marking line category of the sub-road marking line;
[0016] A map feature extraction module is used to input sample environment perception data into the map feature extraction network of the road sign detection model to obtain sample map features;
[0017] A prediction module, configured to input sample map features into a detection network of a road sign detection model to obtain predicted sub-road sign lines and their road sign line categories;
[0018] A model training module is used to train a road marking line detection model based on a first difference and a second difference, wherein the first difference is the difference between the predicted sub-road marking line and the sub-road marking line, and the second difference is the difference between the marking line category of the predicted sub-road marking line and the marking line category of the sub-road marking line.
[0019] According to a fourth aspect of the present disclosure, a device for detecting road marking lines is provided, the device comprising:
[0020] A data acquisition module, used to acquire environmental perception data of the road environment;
[0021] A map feature extraction module, configured to input environmental perception data into a map feature extraction network of a road sign detection model to obtain map features. The road sign detection model is trained using the above-mentioned road sign detection model training method.
[0022] The prediction module is used to input the map features into the detection network of the road marking line detection model to obtain at least one sub-road marking line and the marking line category of each road marking line in the environmental map of the road environment.
[0023] According to a fifth aspect of the present disclosure, an electronic device is provided, including:
[0024] at least one processor; and
[0025] A memory communicatively connected to the at least one processor; wherein,
[0026] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the road sign line detection model training or detection method.
[0027] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned road sign line detection model training or detection method.
[0028] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the above-mentioned road marking line detection model training or detection method when executed by a processor.
[0029] According to an eighth aspect of the present disclosure, an autonomous driving vehicle is provided, comprising the electronic device described in the third aspect above.
[0030] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0032] Figure 1 1 is a flow chart of a road sign line detection model training method provided by an embodiment of the present disclosure;
[0033] Figure 2 is a schematic diagram of a segmentation situation of a sample environment map provided by an embodiment of the present disclosure;
[0034] Figure 3 1 is a flow chart of a method for detecting road marking lines provided by an embodiment of the present disclosure;
[0035] Figure 4 This is a flowchart of a specific implementation of a method for detecting road marking lines provided in an embodiment of the present disclosure;
[0036] Figure 5 is a structural diagram of a road sign line detection model training device provided by an embodiment of the present disclosure;
[0037] Figure 6 1 is a schematic structural diagram of a road marking line detection device provided by an embodiment of the present disclosure;
[0038] Figure 7 4 is a block diagram of an electronic device used to implement the road sign line detection model training or detection method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0039] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0040] In one implementation of the related technology, an environmental image of the road environment is extracted, road markings are identified from the environmental image, and then the road markings in the image are converted into a high-precision map. For example, the road markings in the environmental image are converted into a high-precision map from a bird's-eye view (BEV). This method of converting the road markings in the image into a high-precision map generally results in a loss of precision, affecting the accuracy of the road markings in the high-precision map. In addition, the conversion of road markings in the image into a high-precision map may require depth data, which is highly dependent on data.
[0041] Another implementation of the related technology involves learning map features from environmental images and then performing semantic segmentation on them to obtain semantic segmentation results for road markings. Because the image region corresponding to the semantic segmentation results for road markings can be quite wide, significant deviations can occur when the semantic segmentation results for road markings are vectorized. For example, the shape and orientation of the road markings after vectorization may differ significantly. Therefore, this approach can also affect the accuracy of road markings in high-precision maps.
[0042] The road marking line detection model training, road marking line detection method and autonomous driving vehicle provided in the embodiments of the present disclosure are intended to solve at least one of the above technical problems in the prior art.
[0043] Figure 1 FIG. 1 shows a flow chart of a road sign line detection model training method provided by an embodiment of the present disclosure, such as Figure 1 As shown in , the method may mainly include:
[0044] Step S110: Obtain sample environment perception data of the sample road environment and annotation information of the sample environment map corresponding to the sample road environment, the annotation information including at least one sub-road marking line contained in each road marking line in the sample environment map and the marking line category of the sub-road marking line.
[0045] Step S120: Input the sample environment perception data into the map feature extraction network of the road sign line detection model to obtain sample map features.
[0046] Step S130: Inputting the sample map features into the detection network of the road sign line detection model to obtain the predicted sub-road sign line and the road sign line category of the predicted sub-road sign line.
[0047] Step S140: training a road marking line detection model based on a first difference and a second difference, wherein the first difference is the difference between the predicted sub-road marking line and the sub-road marking line, and the second difference is the difference between the marking line category of the predicted sub-road marking line and the marking line category of the sub-road marking line.
[0048] Among them, the sample environment perception data is the environmental data around the sample road environment collected by the environmental perception equipment. The sample environment perception data can be used to construct a sample environment map corresponding to the sample road environment.
[0049] In the embodiment of the present disclosure, a road marking line is divided into a plurality of sub-road marking lines. By dividing the road marking line into a plurality of sub-road marking lines, it is convenient to describe each sub-road marking line.
[0050] As an example, a sub-road marking line can be described by a corresponding curve equation, and a road marking line can be described by the curve equations of the multiple sub-road marking lines it contains.
[0051] In the embodiment of the present disclosure, the marking line category of a sub-road marking line is the marking line type of the road marking line to which it belongs. The marking line type may include but is not limited to lane markings, crosswalks, stop lines, road boundary lines, etc.
[0052] In the disclosed embodiments, the road marking line detection model may include a map feature extraction network and a detection network. The feature extraction network is configured to learn sample map features based on sample environmental perception data. The detection network is configured to predict sub-road marking lines and their marking line categories based on the sample map features.
[0053] In this disclosed embodiment, a first difference between a predicted sub-road marking line and the sub-road marking line is determined, and a second difference between the line category of the predicted sub-road marking line and the line category of the sub-road marking line is determined. A road marking line detection model is trained based on the first and second differences. The trained road marking line detection model can be used to effectively detect sub-road marking lines and their line categories in an environment map.
[0054] The method provided by the embodiment of the present disclosure obtains sample environmental perception data of a sample road environment and annotation information of a sample environmental map corresponding to the sample road environment, wherein the annotation information includes at least one sub-road marking line contained in each road marking line in the sample environmental map and the marking line category of the sub-road marking line. The sample environmental perception data is input into the map feature extraction network of the road marking line detection model to obtain sample map features. The sample map features are input into the detection network of the road marking line detection model to obtain predicted sub-road marking lines and the marking line category of the predicted sub-road marking lines. The road marking line detection model is trained based on a first difference between the predicted sub-road marking line and the sub-road marking line, and a second difference between the marking line category of the predicted sub-road marking line and the marking line category of the sub-road marking line. The road marking line detection model trained based on this solution can be used to effectively detect sub-road marking lines of road marking lines such as lane lines in the environmental map, thereby ensuring the normal operation of autonomous driving.
[0055] In this disclosed embodiment, because map features are determined based on environmental perception characteristics and then sub-road markings in the environmental map are predicted based on these map features, there is no need to convert road markings in the environmental image into road markings in the HD map, thus avoiding the accuracy loss caused by this process. Furthermore, this solution can be implemented without relying on depth data, reducing its reliance on data.
[0056] In the disclosed embodiment, since the sub-road marking lines and their corresponding marking line categories in the environment map are directly predicted, no vector conversion is required, and the accuracy of the predicted sub-road marking lines can be guaranteed.
[0057] In the disclosed embodiment, the description of the entire road marking line is made more detailed by dividing the road marking line into multiple sub-road marking lines and describing each sub-road marking line separately. By using the sub-road marking lines and their corresponding marking line categories in the sample environment map as supervision, the trained road marking line detection model can accurately predict the sub-road marking lines, thereby improving the accuracy of the road marking lines constructed in the high-precision map.
[0058] In the disclosed embodiment, after sub-road marking lines are detected by the road marking line detection model, road marking lines can be constructed in the high-precision map based on the sub-road marking lines. Since the granularity of the sub-road marking lines is small, traffic signs can be expressed more finely in the high-precision map, and the edges of the road marking lines and the turning points of the road marking lines can be better expressed.
[0059] In the disclosed embodiment, the length of the sub-road marking lines can be relatively small, so that the sub-road marking lines cut out from the road marking lines are mostly straight lines or approximately straight lines, so that the road marking line detection model can learn quickly and reach convergence quickly.
[0060] In an optional embodiment of the present disclosure, sample map features are input into a detection network of a road sign detection model to obtain predicted sub-road signs and their categories, including:
[0061] Determine content features and key-value features based on sample map features;
[0062] Determining query features based on a preset query feature acquisition method;
[0063] Perform cross-attention processing based on content features, key-value features, and query features to obtain attention features;
[0064] A predicted sub-road marking line and a marking line category of the predicted sub-road marking line are determined based on the attention feature.
[0065] In an embodiment of the present disclosure, the detection network may include a cross-attention subnetwork and a detection head, and the cross-attention subnetwork is used to extract attention features based on the cross-attention mechanism.
[0066] Specifically, the sample map features can be linearly transformed to obtain content (Value) features and key (Key) features.
[0067] The query feature may be determined based on a preset query feature acquisition method.
[0068] Cross-attention processing is performed based on content features, key-value features, and query features to obtain attention features, which can then be input into the detection head to obtain the predicted sub-road sign line and the sign line category of the predicted sub-road sign line output by the detection head.
[0069] By adopting the cross-attention mechanism, we can better learn the global information of map features, and use the attention features obtained by cross-attention processing for subsequent detection, which has better detection effect.
[0070] In an optional embodiment of the present disclosure, cross-attention processing is performed based on content features, key-value features, and query features to obtain attention features, including:
[0071] Determine the attention weight based on query features and key-value features;
[0072] Based on the attention weight and content features, the attention features are determined.
[0073] In the embodiment of the present disclosure, the attention feature can be expressed by the following formula 1.
[0074]
[0075] Among them, Q represents the query feature, K represents the key value feature, and K T represents the transpose of the key-value feature, V represents the content feature, softmax represents the normalized exponential function, d k represents the dimension of the key-value feature K, represents the attention weight, and Attention(Q,K,V) represents the attention feature.
[0076] In an optional manner of the present disclosure, determining the query feature based on a preset query feature acquisition method includes:
[0077] Sampling the sample map features to obtain sampling features;
[0078] Perform linear transformation on the sampled features to obtain the query features.
[0079] In the disclosed embodiment, the sample map features may be sampled to obtain sampling features, that is, the feature points of the sample map features may be sampled to obtain sampling points, the coordinates of the sampling points may be normalized to obtain sampling features, and then the sampling features may be linearly transformed to obtain query features.
[0080] As an example, the number of sampled features may be large. In this case, a multi-layer perceptron can be used to perform linear transformation on each sampled feature to obtain the query feature.
[0081] In the disclosed embodiments, query features can also be obtained by random assignment. Compared to random assignment, the query features obtained by sampling sample map features and then performing a linear transformation on the sample features are performed using the information contained in the sample map features, resulting in better query performance.
[0082] In an optional embodiment of the present disclosure, determining the predicted sub-road marking line and the marking line category of the predicted sub-road marking line based on the attention feature includes:
[0083] The attention features are matched with the annotation information based on the Hungarian algorithm to obtain the target attention features that match the annotation information;
[0084] A predicted sub-road marking line and a marking line category of the predicted sub-road marking line are determined based on the target attention feature.
[0085] In the disclosed embodiment, the number of attention features is generally greater than the number of landmark information. In this case, the Hungarian algorithm can be used to determine the target attention features that match the landmark information.
[0086] Specifically, the prediction results based on each attention feature can be used to calculate the loss of the annotation information, and the cost (cos t) matrix of the Hungarian algorithm can be determined based on the loss, so as to achieve matching based on the Hungarian algorithm and obtain the target attention feature that matches the annotation information.
[0087] After determining the target attention feature, the loss corresponding to the target attention feature can be used as the loss of this training for gradient backpropagation.
[0088] In an optional manner of the present disclosure, before obtaining the annotation information of the sample environment map corresponding to the sample road environment, the method further includes:
[0089] Get the marked points in the sample environment map;
[0090] Split the sample environment map into map tiles of preset size;
[0091] Determine a set of annotation points based on the annotation points in the map block;
[0092] Perform curve fitting on the marked points in the marked point set to obtain the sub-road marking line.
[0093] In the embodiment of the present disclosure, the initial labels in the sample environment map generally exist in the form of annotation points, that is, the annotation user can annotate the road marking line category for a group of points corresponding to the road marking line, thereby obtaining the annotation points.
[0094] In the disclosed embodiment, the sample environment map may be divided into map blocks of a preset size. Each map block generally includes annotation points, and the annotation point set may be determined based on the annotation points in the map block.
[0095] As an example, the segmented map blocks may correspond to square areas with a side length of 2 meters in the actual environment.
[0096] As an example, a cubic Bezier curve fitting method may be used to fit the sub-road marking line.
[0097] In an optional manner of the present disclosure, determining a set of annotation points based on annotation points in a map block includes:
[0098] Determine the annotation points in each map block as a candidate annotation point set;
[0099] In response to the number of annotation points in the candidate annotation point set being not less than a preset value, determining the candidate annotation point set as the annotation point set;
[0100] In response to the number of annotation points in the candidate annotation point set being less than a preset value, the candidate annotation point set is used as the annotation point set to be merged, the annotation point set to be merged is merged with the target annotation point set to obtain a merged annotation point set, and the merged annotation point set is used as the candidate annotation point set, where the target annotation point set is the one with the smaller number of annotation points in the two candidate annotation point sets adjacent to the annotation point set to be merged.
[0101] In the disclosed embodiment, a map block may contain fewer annotation points, which may result in an inability to effectively fit a curve in the map block, or the fitted curve may be unable to effectively supervise the training of the model.
[0102] As an example, Figure 2 A schematic diagram of the segmentation of a sample environment map is shown in FIG. Figure 2 As shown in Figure 2 Each square in the example represents a segmented map block from the sample environment map. x represents a road marking line in the sample environment map, which is composed of multiple annotation points. Only a few annotation points at the top of the road marking line x fall within the map block a, indicating that the map block a contains few annotation points.
[0103] In the embodiment of the present disclosure, the annotation points in each map block can be determined as a candidate annotation point set. If the number of annotation points in the candidate annotation point set is not less than a preset value, it means that the number of annotation points in the candidate annotation point set is sufficient to effectively support subsequent processing. In this case, the candidate annotation point set can be determined as the annotation point set.
[0104] If the number of annotation points in the candidate annotation point set is less than a preset value, it means that the number of annotation points in the candidate annotation point set is small and cannot effectively support subsequent processing. In this case, the candidate annotation point set can be determined as the annotation point set to be merged, and the one containing the smaller number of annotation points is determined from the two candidate annotation point sets adjacent to the annotation point set to be merged as the target annotation point set. The annotation point set to be merged and the target annotation point set are merged to obtain a merged annotation point set. At this time, the merged annotation point set can be used as the candidate annotation point set, and the above steps of determining the candidate annotation point set as the annotation point set in response to the number of annotation points in the candidate annotation point set being not less than the preset value, using the candidate annotation point set as the annotation point set in response to the number of annotation points in the candidate annotation point set being less than the preset value, merging the annotation point set to be merged with the target annotation point set to obtain a merged annotation point set, and using the merged annotation point set as the candidate annotation point set are performed again until all annotation points are divided into the annotation point set.
[0105] For example, the annotation points include annotation point 1, annotation point 2, annotation point 3, ..., annotation point 23, wherein annotation point 1, annotation point 2, ..., annotation point 8 are located in map block 1, annotation point 9, annotation point 10, ..., annotation point 13 are located in map block 2, annotation point 14, annotation point 15 are located in map block 3, and annotation point 16, annotation point 17, ..., annotation point 23 are located in map block 4. The preset value is 3. In this case, map block 3 contains only two annotation points, which is less than the preset value. The candidate annotation point set consisting of annotation point 14 and annotation point 15 is the annotation point set to be merged. The two adjacent map blocks of map block 3 are map block 2 and map block 4. The number of annotation points contained in map block 2 is relatively small. The annotation points contained in map block 2 can be used as the target annotation point set and merged with the annotation point set to be merged to obtain a merged annotation point set. The merged annotation point set contains annotation points 9, annotation point 10, ..., annotation point 15. In this case, the number of annotation points contained in the merged annotation point set is greater than the preset value, and the merged annotation point set can be determined as the annotation point set.
[0106] In an optional manner of the present disclosure, obtaining a marked point in a sample environment map includes:
[0107] Get the initial annotation points in the sample environment map;
[0108] In response to the presence of an interpolation start point and an interpolation end point in the initial marked point, interpolation processing is performed between the interpolation start point and the interpolation end point to obtain an interpolation point, where the interpolation start point and the interpolation end point are adjacent, and a distance between the interpolation start point and the interpolation end point is greater than a preset distance;
[0109] The initial marking points and interpolation points are determined as marking points.
[0110] In the embodiment of the present disclosure, the sample environment map may have too sparse or too dense annotation points, which makes it difficult to divide the annotation point set and perform curve fitting.
[0111] In the case where the sample environment map may have too sparse annotation points, the initially acquired annotation points can be recorded as initial annotation points. Two adjacent annotation points with a distance greater than a preset distance between them can be found from the initial annotation points. These two adjacent annotation points can be recorded as the interpolation starting point and the interpolation end point, respectively. Interpolation processing can be performed between the interpolation starting point and the interpolation end point to obtain the interpolation point. The interpolation processing can use linear interpolation or curve interpolation. The initial annotation point and the interpolation point are determined as the annotation points.
[0112] Since the sample environment map may have too dense annotation points, the initial annotation points can be sampled, for example, an initial annotation point is extracted at a certain distance as a annotation point.
[0113] In an optional manner of the present disclosure, training a road sign line detection model based on the first difference and the second difference includes:
[0114] determining a curve loss function based on the first difference;
[0115] determining a type loss function based on the second difference;
[0116] Determine the target loss function based on the curve loss function and the type loss function;
[0117] The road sign detection model is trained based on the target loss function.
[0118] In this disclosed embodiment, the first difference is the difference between the predicted sub-road marking line and the sub-road marking line, namely, the curve difference. The second difference is the difference between the line category of the predicted sub-road marking line and the line category of the sub-road marking line, namely, the category difference. A curve loss function can be constructed based on the curve difference, and a category loss function can be constructed based on the category difference. The target loss function is then obtained by performing a weighted summation of the curve loss function and the category loss function.
[0119] In the disclosed embodiment, sub-road marking lines can be fitted as cubic Bezier curves. Curve differences can be represented by differences in curve length, curve control points, and curve interpolation points. Therefore, a curve length difference loss function, a curve control point difference loss function, and a curve interpolation point difference loss function can be constructed, respectively. The curve loss function is determined based on these three functions.
[0120] As an example, the target loss function can be expressed by the following formula 2.
[0121]
[0122] in, represents the target loss function, represents the curve length loss function, α represents the weight of the curve length loss function, represents the curve control point loss function, β represents the weight of the curve control point loss function, represents the loss function of the curve interpolation point, γ represents the weight of the loss function of the curve interpolation point, Represents the type loss function.
[0123] As an example, the detection head in the detection network can be a multilayer perceptron, and the detection results output by the detection head can specifically include the logistic regression value of the marking line category, and the normalized coordinate values of the starting point, middle point and end point of the cubic Bezier curve of the sub-road marking line.
[0124] In an optional embodiment of the present disclosure, the sample environment perception data includes any of the following:
[0125] Sample environment image data;
[0126] Sample environment image data and sample environment point cloud data.
[0127] In the embodiment of the present disclosure, the sample environment perception data may include sample environment image data collected by a camera, or may include sample environment image data collected by a camera and sample environment point cloud data collected by a lidar.
[0128] In an optional embodiment of the present disclosure, in response to the sample environment perception data including sample environment image data, the sample environment perception data is input into a map feature extraction network of a road sign line detection model to obtain sample map features, including:
[0129] extracting sample image features based on sample environment image data;
[0130] determining sample map features based on sample image features;
[0131] In response to the sample environment perception data including the sample environment image data and the sample environment point cloud data, the sample environment perception data is input into the map feature extraction network of the road sign line detection model to obtain sample map features, including:
[0132] Extracting sample image features based on the sample environment image data, and extracting sample point cloud features based on the sample environment point cloud data;
[0133] Based on the sample image features and the sample point cloud features, and based on preset device-related parameters, the sample map features are determined.
[0134] In the embodiment of the present disclosure, when the sample environment perception data is sample environment image data, the sample image features can be extracted by an image feature extractor, and then the initial sample map features are trained based on the sample image features to obtain the sample map features. The initial sample map features can be the sample map features obtained by initialization.
[0135] As an example, the image feature extractor can be implemented based on an efficient network (Efficient-net) or a residual network (Res idual Networks, ResNets).
[0136] When the sample environment perception data is sample environment image data and sample environment point cloud data, sample image features can be extracted using an image feature extractor, and sample point cloud features can be extracted using a point cloud feature extractor. Based on the sample image features and the sample point cloud features, and based on preset device-related parameters, initial sample map features are trained to obtain sample map features. The initial sample map features may be sample map features obtained through initialization.
[0137] Device-related parameters may include internal and external parameters of the camera, internal and external parameters of the lidar, internal and external parameters of the inertial measurement unit (IMU), etc.
[0138] Figure 3 A schematic diagram of a road marking line detection method provided by an embodiment of the present disclosure is shown. Figure 3 As shown in , the method may mainly include:
[0139] Step S310: Acquire environmental perception data of the road environment.
[0140] Step S320: Inputting the environmental perception data into the map feature extraction network of the road sign line detection model to obtain map features. The road sign line detection model is trained using the above-mentioned road sign line detection model training method.
[0141] Step S330: inputting the map features into the detection network of the road marking line detection model to obtain at least one sub-road marking line and the marking line category of each road marking line contained in the environmental map of the road environment.
[0142] Among them, environmental perception data is environmental data around the road environment collected by environmental perception equipment. The environmental perception data can be used to construct an environmental map corresponding to the road environment.
[0143] In the disclosed embodiment, environmental perception data can be collected in real time and an online high-precision map can be generated while the vehicle is driving.
[0144] In the disclosed embodiments, a road sign detection model may include a map feature extraction network and a detection network. The feature extraction network is configured to learn map features based on environmental perception data. The detection network is configured to predict sub-road signs and their respective road sign categories based on the map features.
[0145] In the disclosed embodiment, the road marking line detection model detects sub-road marking lines. A road marking line may be composed of multiple sub-road marking lines. By dividing a road marking line into multiple sub-road marking lines, it is easier to describe each sub-road marking line.
[0146] As an example, a sub-road marking line can be described by a corresponding curve equation, and a road marking line can be described by the curve equations of the multiple sub-road marking lines it contains.
[0147] In the embodiment of the present disclosure, the marking line category of a sub-road marking line is the marking line type of the road marking line to which it belongs. The marking line type may include but is not limited to lane markings, crosswalks, stop lines, road boundary lines, etc.
[0148] The method provided in the disclosed embodiments obtains environmental perception data of the road environment and inputs this data into the map feature extraction network of a road sign detection model to obtain map features. The road sign detection model is trained using the aforementioned road sign detection model training method. The map features are then input into the detection network of the road sign detection model to obtain at least one sub-road sign contained in each road sign in the environmental map of the road environment, as well as the road sign category of the sub-road sign. Based on this solution, the pre-trained road sign detection model can effectively detect sub-road signs of road signs, such as lane markings, in the environmental map, ensuring the normal operation of autonomous driving.
[0149] In this disclosed embodiment, because map features are determined based on environmental perception characteristics and then sub-road markings in the environmental map are predicted based on these map features, there is no need to convert road markings in the environmental image into road markings in the HD map, thus avoiding the accuracy loss caused by this process. Furthermore, this solution can be implemented without relying on depth data, reducing its reliance on data.
[0150] In the disclosed embodiment, since the sub-road marking lines and their corresponding marking line categories in the environment map are directly predicted, no vector conversion is required, and the accuracy of the predicted sub-road marking lines can be guaranteed.
[0151] In the disclosed embodiment, by dividing the road marking line into multiple sub-road marking lines and describing each sub-road marking line separately, the description of the entire road marking line is made more detailed. By using the sub-road marking lines and their corresponding marking line categories in the environmental map as supervision, the trained road marking line detection model can accurately predict the sub-road marking lines, thereby improving the accuracy of the road marking lines constructed in the high-precision map.
[0152] In the disclosed embodiment, after sub-road marking lines are detected by the road marking line detection model, road marking lines can be constructed in the high-precision map based on the sub-road marking lines. Since the granularity of the sub-road marking lines is small, traffic signs can be expressed more finely in the high-precision map, and the edges of the road marking lines and the turning points of the road marking lines can be better expressed.
[0153] In the disclosed embodiment, the length of the sub-road marking lines can be relatively small, so that the sub-road marking lines cut out from the road marking lines are mostly straight lines or approximately straight lines, so that the road marking line detection model can learn quickly and reach convergence quickly.
[0154] In an optional embodiment of the present disclosure, the environmental perception data includes any of the following:
[0155] Environmental image data;
[0156] Environmental image data and environmental point cloud data.
[0157] In an optional embodiment of the present disclosure, after obtaining at least one sub-road marking line and the marking line category of each road marking line in the environment map of the road environment, the method further includes:
[0158] Based on the sub-road marking lines and their marking line categories, the sub-road marking lines are merged to obtain a road marking line.
[0159] In the embodiment of the present disclosure, after the sub-road marking lines are predicted, the sub-road marking lines can be merged according to their marking line categories to obtain a road marking line.
[0160] Specifically, a plurality of continuous sub-road marking lines belonging to the same marking line category are merged into one road marking line.
[0161] As an example, Figure 4 A flowchart illustrating a specific implementation of a method for detecting road marking lines provided in an embodiment of the present disclosure is shown.
[0162] like Figure 4As shown in , picture is the environment image data. BEV is the initial BEV map feature.
[0163] The image feature extractor is used to extract an image feature map (FeatuerMap) from the environmental image data.
[0164] The BEV generation model is used to train the initial BEV features based on the image feature map to obtain the BEV map feature map (BEVFeatuerMap).
[0165] Transformer Decoder, used to extract attention features based on BEV map features, equivalent to the cross attention network in the embodiment of the present disclosure.
[0166] The multi-layer perceptron is used to output the detection results, which may specifically include the logistic regression value of the road marking line category (i.e., the probability in the figure), and the normalized coordinate values of the starting point, middle point, and end point of the cubic Bezier curve of the sub-road marking line.
[0167] Based on Figure 1 The same principle as shown in the method, Figure 5 FIG. 1 shows a schematic diagram of a road sign line detection model training device provided by an embodiment of the present disclosure. Figure 5 As shown, the road marking line detection model training device 50 may include:
[0168] A data acquisition module 510 is configured to acquire sample environmental perception data of a sample road environment and annotation information of a sample environmental map corresponding to the sample road environment, wherein the annotation information includes at least one sub-road marking line contained in each road marking line in the sample environmental map and a marking line category of the sub-road marking line;
[0169] A map feature extraction module 520 is used to input sample environment perception data into a map feature extraction network of a road sign detection model to obtain sample map features;
[0170] A prediction module 530 is configured to input sample map features into a detection network of a road sign detection model to obtain predicted sub-road sign lines and their road sign line categories;
[0171] The model training module 540 is used to train the road marking line detection model based on the first difference and the second difference, wherein the first difference is the difference between the predicted sub-road marking line and the sub-road marking line, and the second difference is the difference between the marking line category of the predicted sub-road marking line and the marking line category of the sub-road marking line.
[0172] The device provided by the embodiment of the present disclosure obtains sample environment perception data of a sample road environment and annotation information of a sample environment map corresponding to the sample road environment, wherein the annotation information includes at least one sub-road marking line contained in each road marking line in the sample environment map and the marking line category of the sub-road marking line, inputs the sample environment perception data into the map feature extraction network of the road marking line detection model to obtain sample map features, inputs the sample map features into the detection network of the road marking line detection model to obtain predicted sub-road marking lines and the marking line category of the predicted sub-road marking lines, and trains the road marking line detection model based on a first difference between the predicted sub-road marking line and the sub-road marking line, and a second difference between the marking line category of the predicted sub-road marking line and the marking line category of the sub-road marking line. The road marking line detection model trained based on this solution can be used to effectively detect sub-road marking lines of road marking lines such as lane lines in the environment map, thereby ensuring the normal operation of autonomous driving.
[0173] Optionally, the prediction module is specifically used to:
[0174] Determine content features and key-value features based on sample map features;
[0175] Determining query features based on a preset query feature acquisition method;
[0176] Perform cross-attention processing based on content features, key-value features, and query features to obtain attention features;
[0177] A predicted sub-road marking line and a marking line category of the predicted sub-road marking line are determined based on the attention feature.
[0178] Optionally, when determining the predicted sub-road marking line and the marking line category of the predicted sub-road marking line based on the attention feature, the prediction module is specifically configured to:
[0179] Determine the attention weight based on query features and key-value features;
[0180] Based on the attention weight and content features, the attention features are determined.
[0181] Optionally, when determining the query feature based on a preset query feature acquisition method, the prediction module is specifically configured to:
[0182] Sampling the sample map features to obtain sampling features;
[0183] Perform linear transformation on the sampled features to obtain the query features.
[0184] Optionally, when determining the predicted sub-road marking line and the marking line category of the predicted sub-road marking line based on the attention feature, the prediction module is specifically configured to:
[0185] The attention features are matched with the annotation information based on the Hungarian algorithm to obtain the target attention features that match the annotation information;
[0186] A predicted sub-road marking line and a marking line category of the predicted sub-road marking line are determined based on the target attention feature.
[0187] Optionally, the above device further includes a sub-road marking line determination module, and the sub-road marking line determination module is specifically configured to:
[0188] Before obtaining the annotation information of the sample environment map corresponding to the sample road environment, obtaining the annotation points in the sample environment map;
[0189] Split the sample environment map into map tiles of preset size;
[0190] Determine a set of annotation points based on the annotation points in the map block;
[0191] Perform curve fitting on the marked points in the marked point set to obtain the sub-road marking line.
[0192] Optionally, when determining the marked point set based on the marked points in the map block, the sub-road marking line determination module is specifically configured to:
[0193] Determine the annotation points in each map block as a candidate annotation point set;
[0194] In response to the number of annotation points in the candidate annotation point set being not less than a preset value, determining the candidate annotation point set as the annotation point set;
[0195] In response to the number of annotation points in the candidate annotation point set being less than a preset value, the candidate annotation point set is used as the annotation point set to be merged, the annotation point set to be merged is merged with the target annotation point set to obtain a merged annotation point set, and the merged annotation point set is used as the candidate annotation point set, where the target annotation point set is the one with the smaller number of annotation points in the two candidate annotation point sets adjacent to the annotation point set to be merged.
[0196] Optionally, when obtaining the marked points in the sample environment map, the sub-road marking line determination module is specifically configured to:
[0197] Get the initial annotation points in the sample environment map;
[0198] In response to the presence of an interpolation start point and an interpolation end point in the initial marked point, interpolation processing is performed between the interpolation start point and the interpolation end point to obtain an interpolation point, where the interpolation start point and the interpolation end point are adjacent, and a distance between the interpolation start point and the interpolation end point is greater than a preset distance;
[0199] The initial marking points and interpolation points are determined as marking points.
[0200] Optionally, when training the road sign line detection model based on the first difference and the second difference, the model training module is specifically configured to:
[0201] determining a curve loss function based on the first difference;
[0202] determining a type loss function based on the second difference;
[0203] Determine the target loss function based on the curve loss function and the type loss function;
[0204] The road sign detection model is trained based on the target loss function.
[0205] Optionally, the sample environmental perception data includes any of the following:
[0206] Sample environment image data;
[0207] Sample environment image data and sample environment point cloud data.
[0208] Optionally, in response to the sample environment perception data including sample environment image data, the map feature extraction module is specifically configured to:
[0209] extracting sample image features based on sample environment image data;
[0210] determining sample map features based on sample image features;
[0211] In response to the sample environment perception data including the sample environment image data and the sample environment point cloud data, the map feature extraction module is specifically configured to:
[0212] Extracting sample image features based on the sample environment image data, and extracting sample point cloud features based on the sample environment point cloud data;
[0213] Based on the sample image features and the sample point cloud features, and based on preset device-related parameters, the sample map features are determined.
[0214] It can be understood that the above modules of the road sign line detection model training device in the embodiment of the present disclosure have the function of realizing Figure 1 The functions of the corresponding steps of the road marking line detection model training method in the embodiment shown in . This function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above modules can be implemented separately or integrated into multiple modules. For the functional description of each module of the road marking line detection model training device, please refer to Figure 1 The corresponding description of the road sign line detection model training method in the embodiment shown in will not be repeated here.
[0215] Based on Figure 3 The same principle as shown in the method, Figure 6 FIG. 1 shows a schematic diagram of a road marking line detection device provided by an embodiment of the present disclosure. Figure 6 As shown, the road marking line detection device 60 may include:
[0216] A data acquisition module 610 is used to acquire environmental perception data of the road environment;
[0217] A map feature extraction module 620 is configured to input the environmental perception data into a map feature extraction network of a road sign detection model to obtain map features. The road sign detection model is trained using the above-mentioned road sign detection model training method.
[0218] The prediction module 630 is configured to input the map features into the detection network of the road marking line detection model to obtain at least one sub-road marking line and the marking line category of each road marking line in the environment map of the road environment.
[0219] The device provided in the disclosed embodiments obtains environmental perception data of the road environment and inputs this data into the map feature extraction network of a road sign detection model to obtain map features. The road sign detection model is trained using the aforementioned road sign detection model training method. The map features are then input into the detection network of the road sign detection model to obtain at least one sub-road sign contained in each road sign in the environmental map of the road environment, as well as the road sign category of the sub-road sign. Based on this solution, the pre-trained road sign detection model can effectively detect sub-road signs of road signs, such as lane markings, in the environmental map, ensuring the normal operation of autonomous driving.
[0220] Optionally, the environmental perception data includes any of the following:
[0221] Environmental image data;
[0222] Environmental image data and environmental point cloud data.
[0223] Optionally, the above device further includes a road marking line determination module, configured to:
[0224] After obtaining at least one sub-road marking line and a marking line category of each road marking line in the environment map of the road environment, the sub-road marking lines are merged based on the sub-road marking lines and the marking line categories of the sub-road marking lines to obtain a road marking line.
[0225] It is understandable that the above modules of the road marking line detection device in the embodiment of the present disclosure have the function of realizing Figure 3 The functions of the corresponding steps of the method for detecting road marking lines in the embodiment shown in FIG. This function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and each of the above modules can be implemented separately or integrated into multiple modules. For a detailed description of the functions of each module of the above road marking line detection device, please refer to Figure 3 The corresponding description of the method for detecting road marking lines in the embodiment shown in is omitted here for brevity.
[0226] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0227] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, a computer program product, and an autonomous driving vehicle.
[0228] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the road marking line detection model training or road marking line detection method provided in the embodiments of the present disclosure.
[0229] Compared with the existing technology, this electronic device obtains sample environmental perception data of a sample road environment and annotation information of a sample environmental map corresponding to the sample road environment, the annotation information including at least one sub-road marking line contained in each road marking line in the sample environmental map and the marking line category of the sub-road marking line. The sample environmental perception data is input into the map feature extraction network of the road marking line detection model to obtain sample map features. The sample map features are then input into the detection network of the road marking line detection model to obtain predicted sub-road marking lines and the marking line category of the predicted sub-road marking lines. The road marking line detection model is trained based on a first difference between the predicted sub-road marking line and the sub-road marking line, and a second difference between the marking line category of the predicted sub-road marking line and the marking line category of the sub-road marking line. The road marking line detection model trained based on this solution can be used to effectively detect sub-road marking lines of road marking lines such as lane lines in the environmental map, ensuring the normal operation of autonomous driving.
[0230] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the road marking line detection model training or road marking line detection method provided in the embodiments of the present disclosure.
[0231] Compared with the prior art, this readable storage medium obtains sample environmental perception data of a sample road environment and annotation information of a sample environmental map corresponding to the sample road environment, the annotation information including at least one sub-road marking line and the marking line category of each road marking line in the sample environmental map. The sample environmental perception data is input into the map feature extraction network of the road marking line detection model to obtain sample map features. The sample map features are then input into the detection network of the road marking line detection model to obtain predicted sub-road marking lines and the marking line category of the predicted sub-road marking lines. The road marking line detection model is trained based on a first difference between the predicted sub-road marking line and the sub-road marking line, and a second difference between the marking line category of the predicted sub-road marking line and the marking line category of the sub-road marking line. The road marking line detection model trained based on this solution can be used to effectively detect sub-road marking lines and the marking line categories of sub-road marking lines in the environmental map, ensuring the normal operation of autonomous driving.
[0232] The computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the road marking line detection model training or the road marking line detection method provided in the embodiments of the present disclosure.
[0233] Compared to the prior art, this computer program product obtains sample environmental perception data of a sample road environment and annotation information of a sample environmental map corresponding to the sample road environment, the annotation information including at least one sub-road marking line contained in each road marking line in the sample environmental map and the marking line category of the sub-road marking line. The sample environmental perception data is input into the map feature extraction network of a road marking line detection model to obtain sample map features. The sample map features are then input into the detection network of the road marking line detection model to obtain predicted sub-road marking lines and the marking line category of the predicted sub-road marking lines. The road marking line detection model is trained based on a first difference between the predicted sub-road marking line and the sub-road marking line, and a second difference between the marking line category of the predicted sub-road marking line and the marking line category of the sub-road marking line. The road marking line detection model trained based on this solution can be used to effectively detect sub-road marking lines of road marking lines, such as lane markings, in an environmental map, ensuring the normal operation of autonomous driving.
[0234] The autonomous driving vehicle includes the above-mentioned electronic equipment.
[0235] Compared to existing technologies, this autonomous driving vehicle obtains sample environmental perception data of a sample road environment and annotation information of a sample environmental map corresponding to the sample road environment. The annotation information includes at least one sub-road marking line contained in each road marking line in the sample environmental map and the marking line category of the sub-road marking line. The sample environmental perception data is input into the map feature extraction network of the road marking line detection model to obtain sample map features. The sample map features are then input into the detection network of the road marking line detection model to obtain predicted sub-road marking lines and the marking line category of the predicted sub-road marking lines. The road marking line detection model is trained based on a first difference between the predicted sub-road marking line and the sub-road marking line, and a second difference between the marking line category of the predicted sub-road marking line and the marking line category of the sub-road marking line. The road marking line detection model trained based on this solution can be used to effectively detect sub-road marking lines of road marking lines, such as lane markings, in the environmental map, ensuring the normal operation of autonomous driving.
[0236] Figure 7 A schematic block diagram of an example electronic device 70 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0237] like Figure 7 As shown, the electronic device 70 includes a computing unit 710, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 720 or a computer program loaded from a storage unit 780 into a random access memory (RAM) 730. Various programs and data required for the operation of the device 70 can also be stored in the RAM 730. The computing unit 710, the ROM 720, and the RAM 730 are connected to each other via a bus 740. An input / output (I / O) interface 750 is also connected to the bus 740.
[0238] Various components in device 70 are connected to I / O interface 750, including an input unit 760, such as a keyboard, mouse, etc.; an output unit 770, such as various types of displays, speakers, etc.; a storage unit 780, such as a magnetic disk, optical disk, etc.; and a communication unit 790, such as a network card, modem, wireless communication transceiver, etc. Communication unit 790 allows device 70 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0239] The computing unit 710 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 710 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 710 performs the road marking line detection model training or road marking line detection method provided in the embodiments of the present disclosure. For example, in some embodiments, the execution of the road marking line detection model training or road marking line detection method provided in the embodiments of the present disclosure can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 780. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 70 via the ROM 720 and / or the communication unit 790. When the computer program is loaded into the RAM 730 and executed by the computing unit 710, one or more steps of the road marking line detection model training or road marking line detection method provided in the embodiments of the present disclosure can be performed. Alternatively, in other embodiments, the computing unit 710 may be configured in any other appropriate manner (for example, by means of firmware) to execute the road marking line detection model training or road marking line detection method provided in the embodiments of the present disclosure.
[0240] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0241] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0242] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0243] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0244] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0245] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0246] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0247] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A road marking line detection model training method, comprising: Acquiring sample environmental perception data of a sample road environment and annotation information of a sample environmental map corresponding to the sample road environment, the annotation information including at least one sub-road marking line included in each road marking line in the sample environmental map and a marking line category of the sub-road marking line; Inputting the sample environment perception data into a map feature extraction network of a road sign line detection model to obtain sample map features; Inputting the sample map features into the detection network of the road sign line detection model to obtain predicted sub-road sign lines and the road sign line categories of the predicted sub-road sign lines; training the road marking line detection model based on a first difference and a second difference, wherein the first difference is a difference between the predicted sub-road marking line and the sub-road marking line, and the second difference is a difference between a marking line category of the predicted sub-road marking line and a marking line category of the sub-road marking line; The step of inputting the sample map features into the detection network of the road sign line detection model to obtain the predicted sub-road sign line and the road sign line category of the predicted sub-road sign line includes: Determining content features and key-value features based on the sample map features; Determining query features based on a preset query feature acquisition method; Performing cross-attention processing based on the content feature, the key-value feature, and the query feature to obtain an attention feature; A predicted sub-road marking line and a marking line category of the predicted sub-road marking line are determined based on the attention feature.
2. The method according to claim 1, wherein The cross-attention processing based on the content feature, the key-value feature, and the query feature to obtain the attention feature includes: Determining an attention weight based on the query feature and the key-value feature; An attention feature is determined based on the attention weight and the content feature.
3. The method according to claim 1 or 2, wherein: The determining of the query feature based on a preset query feature acquisition method includes: Sampling the sample map features to obtain sampling features; Performing linear transformation on the sampled features to obtain query features.
4. The method according to claim 1 or 2, wherein: The determining, based on the attention feature, a predicted sub-road marking line and a marking line category of the predicted sub-road marking line includes: Matching the attention feature with the annotation information based on the Hungarian algorithm to obtain a target attention feature that matches the annotation information; A predicted sub-road marking line and a marking line category of the predicted sub-road marking line are determined based on the target attention feature.
5. The method according to claim 1 or 2, before obtaining the annotation information of the sample environment map corresponding to the sample road environment, the method further comprises: Obtaining the marked points in the sample environment map; Dividing the sample environment map into map blocks of preset sizes; Determining a set of marked points based on the marked points in the map block; Curve fitting is performed on the marked points in the marked point set to obtain the sub-road marking line.
6. The method according to claim 5, wherein: The determining of the annotation point set based on the annotation points in the map block includes: Determine the annotation points in each of the map blocks as a candidate annotation point set; In response to the number of annotation points in the candidate annotation point set being not less than a preset value, determining the candidate annotation point set as the annotation point set; In response to the number of annotation points in the candidate annotation point set being less than a preset value, the candidate annotation point set is used as the annotation point set to be merged, the annotation point set to be merged is merged with a target annotation point set to obtain a merged annotation point set, and the merged annotation point set is used as the candidate annotation point set, where the target annotation point set is the one with the smaller number of annotation points in two candidate annotation point sets adjacent to the annotation point set to be merged.
7. The method according to claim 6, wherein: The obtaining of the marked points in the sample environment map includes: Obtaining initial marking points in the sample environment map; In response to the presence of an interpolation start point and an interpolation end point in the initial marked points, interpolation processing is performed between the interpolation start point and the interpolation end point to obtain an interpolation point, the interpolation start point and the interpolation end point are adjacent, and a distance between the interpolation start point and the interpolation end point is greater than a preset distance; The initial marking point and the interpolation point are determined as marking points.
8. The method according to claim 5, wherein The training of the road sign line detection model based on the first difference and the second difference includes: determining a curve loss function based on the first difference; determining a type loss function based on the second difference; Determine a target loss function based on the curve loss function and the type loss function; The road sign line detection model is trained based on the target loss function.
9. The method according to claim 1, wherein: The sample environmental perception data includes any of the following: Sample environment image data; Sample environment image data and sample environment point cloud data.
10. The method according to claim 1, wherein In response to the sample environment perception data including sample environment image data, inputting the sample environment perception data into a map feature extraction network of a road sign line detection model to obtain sample map features includes: extracting sample image features based on the sample environment image data; determining a sample map feature based on the sample image feature; In response to the sample environment perception data including sample environment image data and sample environment point cloud data, the sample environment perception data is input into a map feature extraction network of a road sign line detection model to obtain sample map features, including: extracting sample image features based on the sample environment image data, and extracting sample point cloud features based on the sample environment point cloud data; Based on the sample image features and the sample point cloud features, and based on preset device-related parameters, sample map features are determined.
11. A method for detecting road marking lines, comprising: Acquire environmental perception data of the road environment; Inputting the environmental perception data into a map feature extraction network of a road sign line detection model to obtain map features, wherein the road sign line detection model is trained by the road sign line detection model training method according to any one of claims 1 to 10; The map features are input into a detection network of the road marking line detection model to obtain at least one sub-road marking line contained in each road marking line in the environment map of the road environment and a marking line category of the sub-road marking line.
12. The method according to claim 11, wherein The environmental perception data includes any of the following: Environmental image data; Environmental image data and environmental point cloud data.
13. The method according to claim 11 or 12, further comprising, after obtaining at least one sub-road marking line contained in each road marking line in the environment map of the road environment and the marking line category of the sub-road marking line: The sub-road marking lines are merged based on the sub-road marking lines and the marking line categories of the sub-road marking lines to obtain the road marking line.
14. A road marking line detection model training device, comprising: a data acquisition module, configured to acquire sample environmental perception data of a sample road environment and annotation information of a sample environmental map corresponding to the sample road environment, the annotation information including at least one sub-road marking line included in each road marking line in the sample environmental map and a marking line category of the sub-road marking line; a map feature extraction module, configured to input the sample environment perception data into a map feature extraction network of a road sign line detection model to obtain sample map features; a prediction module, configured to input the sample map features into a detection network of the road sign line detection model to obtain a predicted sub-road sign line and a road sign line category of the predicted sub-road sign line; a model training module, configured to train the road marking line detection model based on a first difference and a second difference, wherein the first difference is a difference between the predicted sub-road marking line and the sub-road marking line, and the second difference is a difference between a marking line category of the predicted sub-road marking line and a marking line category of the sub-road marking line; The prediction module is specifically used for: Determining content features and key-value features based on the sample map features; Determining query features based on a preset query feature acquisition method; Performing cross-attention processing based on the content feature, the key-value feature, and the query feature to obtain an attention feature; A predicted sub-road marking line and a marking line category of the predicted sub-road marking line are determined based on the attention feature.
15. A road marking detection device, comprising: A data acquisition module, used to acquire environmental perception data of the road environment; a map feature extraction module, configured to input the environmental perception data into a map feature extraction network of a road sign detection model to obtain map features, wherein the road sign detection model is trained using the road sign detection model training method according to any one of claims 1 to 10; A prediction module is used to input the map features into the detection network of the road marking line detection model to obtain at least one sub-road marking line contained in each road marking line in the environmental map of the road environment and the marking line category of the sub-road marking line.
16. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-13.
18. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 13.
19. An autonomous driving vehicle comprising the electronic device according to claim 16.
Citation Information
Patent Citations
Automatic driving method and device, electronic equipment and readable storage medium
CN110794844A
Lane line detection method and device
CN111767853A