Detection result determination method and device, storage medium and vehicle

By combining the multi-source feature fusion of vehicle perimeter camera and high-precision prior map, using self-attention mechanism and convolutional neural network, the efficiency and accuracy problems of lane line detection in complex environments are solved, and efficient and real-time lane line detection is achieved.

CN120356172APending Publication Date: 2025-07-22GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510465413.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing lane line detection technology has low detection efficiency in complex environments, making it difficult to cope with road wear, lighting changes and bad weather conditions. Inappropriate fusion of multi-source information leads to low detection accuracy, poor dynamic adaptability of prior maps, and insufficient real-time and computational efficiency.

Method used

By combining the real-time image data acquired by the vehicle's perimeter camera and high-precision prior map, multi-source bird's-eye viewing features are extracted for fusion, and feature conversion and detection are used for self-attention mechanism and convolutional neural network to enhance environmental adaptability and robustness.

Benefits of technology

It improves the accuracy and robustness of lane line detection, improves the detection efficiency and real-time performance in complex environments, adapts to changes in road environments, and reduces misjudgment and misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356172A_ABST
    Figure CN120356172A_ABST
Patent Text Reader

Abstract

The invention discloses a detection result determination method and device, a storage medium and a vehicle. The method comprises the steps that image data of an area where a vehicle is located currently are obtained, and the image data are used for representing a multi-angle view of the area; a first aerial view angle feature in the image data is extracted, and the first aerial view angle feature is at least used for representing road information in the image data; determining local map data corresponding to the area from the priori map, and extracting a second bird's-eye view feature in the local map data, the second bird's-eye view feature being at least used for representing road information in the local map data; fusing the first aerial view angle feature and the second aerial view angle feature to obtain a fused feature; and determining a detection result of the lane line in the area based on the fusion feature. The technical problem of low lane line detection efficiency is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of data processing, and in particular, to a method, apparatus, storage medium, and vehicle for determining a detection result. Background Art

[0002] Currently, for the detection of lane lines, in order to meet the requirements of high precision and high robustness in autonomous driving, usually only rely on single-sensor data, such as on-vehicle cameras or lidar, to separately perform feature extraction and lane line recognition.

[0003] However, lane lines are often blurred due to road wear, light changes, rain, fog and other adverse weather conditions. The above method is only a lane line detection algorithm based on single visual perception, which is difficult to cope with extremely complex situations, and there is a technical problem of low detection efficiency of lane lines. Summary of the Invention

[0004] Embodiments of the present application provide a method, apparatus, storage medium, and vehicle for determining a detection result, aiming to improve the technical problem of low monitoring efficiency of lane lines.

[0005] According to one embodiment of the present application, a method for determining a detection result is provided, including: obtaining image data of the current area where the vehicle is located, where the image data is used to represent a multi-angle view of the area; extracting a first bird's-eye view feature from the image data, where the first bird's-eye view feature is at least used to represent road information in the image data; determining local map data corresponding to the area from a prior map, and extracting a second bird's-eye view feature from the local map data, where the second bird's-eye view feature is at least used to represent road information in the local map data; fusing the first bird's-eye view feature and the second bird's-eye view feature to obtain a fused feature; and determining a detection result of lane lines in the area based on the fused feature.

[0006] The above optional embodiments of the present application can achieve the following beneficial effects: By fusing the obtained first bird's-eye view feature and the second bird's-eye view feature extracted from the local map data in the prior map to obtain a fused feature, and predicting lane lines based on the fused feature, the understanding of road information and the adaptability to complex environments can be enhanced, thereby achieving the purpose of improving the detection accuracy and robustness of lane lines, and thus achieving the technical effect of improving the detection efficiency of lane lines, and solving the technical problem of low detection efficiency of lane lines.

[0007] Optionally, determining a detection result of lane lines in the area based on the fused feature includes: converting the fused feature to obtain a target feature, where the target feature is different from the fused feature; and detecting the target feature to obtain a detection result.

[0008] The above optional embodiments of the present application can achieve the following beneficial effects: By further converting the fusion features into more optimized target features, the expression ability of the features is enhanced specifically and the key information of the lane line detection task is focused on, so as to achieve the purpose of improving the accuracy and processing speed of lane line detection. Through the above method, features irrelevant to lane line detection (such as noise features) can be effectively filtered out, so as to extract the target features of lane lines more efficiently, solve the problem of insufficient extraction of lane line features in complex image data, and can perform lane line detection stably and with high precision in various driving environments.

[0009] Optionally, converting the fusion features to obtain target features includes: performing dimensionality reduction conversion on the fusion features to obtain a first feature sequence; injecting the position information of the vehicle into the first feature sequence to obtain a second feature sequence; performing mapping processing on the second feature sequence to obtain a feature matrix; and performing normalization processing on the feature matrix to obtain target features.

[0010] The above optional embodiments of the present application can achieve the following beneficial effects: By performing dimensionality reduction conversion on the fusion features and injecting position information, the pertinence and spatial positioning ability of the fusion features are enhanced, so as to achieve the purpose of improving the accuracy of lane line detection and the depth of environmental understanding. Through the above method, the computational complexity of feature processing can be effectively reduced, and at the same time, it is ensured that the obtained target features can carry sufficient spatial position information, so that the lane lines can be located and identified more accurately.

[0011] Optionally, detecting the target features to obtain a detection result includes: retrieving a query vector, where the query vector corresponds to each lane line in the region; and using a decoder to perform regression prediction on the query vector and the target features to obtain a detection result.

[0012] The above optional embodiments of the present application can achieve the following beneficial effects: By retrieving the query vector corresponding to the lane line and using the decoder to perform regression prediction to achieve precise decoding of the target features, the accuracy and reliability of the lane line detection result can be significantly improved. Through the above method, dedicated feature interaction and prediction can be performed for each lane line, so as to infer the monitoring results corresponding to the lane line (such as position, type, and attributes, etc.) more accurately, solve the problem of inaccurate lane line detection in multi-lane and complex road environments. This decoding strategy based on query vectors makes full use of the information in the target features, provides accurate lane line detection results for the autonomous driving system, and helps to improve the positioning accuracy and driving safety of the vehicle.

[0013] Optionally, from the prior map, local map data corresponding to the area is determined, and second bird's-eye view features in the local map data are extracted, including: obtaining the positioning information of the vehicle; determining the global coordinate position of the vehicle in the global coordinate system based on the positioning information; determining the local map data from the prior map according to the global coordinate position; performing feature encoding on the lane lines in the local map data to obtain encoded data; performing pooling processing on the encoded data to obtain the pooling result of the lane lines in the local map data; and performing transformation on the pooling result to obtain the second bird's-eye view features.

[0014] The above optional embodiment of the present application can achieve the following beneficial effects: By starting from the vehicle positioning information, local map data is determined, and feature encoding and pooling processing are performed on the lane lines to achieve efficient extraction and representation of the prior map information. Through the above method, local map data related to the current position of the vehicle can be accurately captured from the high-precision prior map, so as to more effectively utilize the prior map features to assist lane line detection, and avoid the problems of real-time updating of map information and effective utilization of map features in a dynamically changing environment.

[0015] Optionally, the first bird's-eye view feature and the second bird's-eye view feature are fused to obtain a fused feature, including: splicing the first bird's-eye view feature and the second bird's-eye view feature to obtain a spliced feature; and performing feature extraction on the spliced feature to obtain the fused feature.

[0016] The above optional embodiment of the present application can achieve the following beneficial effects: By splicing the first bird's-eye view feature (for example, the real-time BEV feature from the panoramic camera) and the second bird's-eye view feature (for example, the BEV feature from the prior map), the real-time perception information and the prior geographical information are integrated, so as to achieve the purpose of enhancing the multi-modal information fusion ability of lane line detection. Through the above method, multi-source information can be effectively merged in the channel dimension, and the uniqueness and complementarity of their respective features are retained, so as to more comprehensively and accurately describe the road environment around the vehicle, and solve the limitations and robustness problems of a single information source in a complex environment. This strategy of feature splicing and fusion provides a more complete environmental model for the autonomous driving system, which helps to improve the performance of the vehicle in road recognition and path planning.

[0017] Optionally, the method may further include: matching the detection result with the prior map to obtain the similarity between the detection result and the prior map; and outputting the similarity.

[0018] The above optional embodiments of the present application can achieve the following beneficial effects: By matching the detection results of lane lines with a prior map, calculating and outputting the similarity between the two to evaluate the accuracy and consistency of the system's detection results. Through the above method, the matching degree between the lane line detection results and the prior map can be monitored in real time, so as to trigger system calibration or manual intervention when necessary, solving the problem that the lane line detection results may not conform to the actual road conditions in a dynamically changing or map information that is not completely accurate environment, achieving the purpose of dynamically verifying the detection results and continuously optimizing the system performance. This similarity calculation mechanism provides an effective means for performance monitoring and adaptive adjustment of the autonomous driving system, helping to improve the driving safety and stability of the vehicle in an unknown or changing road environment.

[0019] According to one embodiment of the present application, there is also provided a device for determining detection results, which may include: an acquisition unit for acquiring image data of the area where the vehicle is currently located, where the image data is used to represent the multi-angle view of the area; an extraction unit for extracting the first bird's-eye view feature from the image data, where the first bird's-eye view feature is at least used to represent the road information in the image data; a first determination unit for determining, from the prior map, local map data corresponding to the area and extracting the second bird's-eye view feature in the local map data, where the second bird's-eye view feature is at least used to represent the road information in the local map data; a fusion unit for fusing the first bird's-eye view feature and the second bird's-eye view feature to obtain a fusion feature; a second determination unit for determining the detection results of the lane lines in the area based on the fusion feature.

[0020] Optionally, the second determination unit may include: a conversion module for converting the fusion feature to obtain a target feature, where the target feature is different from the fusion feature; a detection module for detecting the target feature to obtain the detection results.

[0021] Optionally, the conversion module may include: a conversion sub-module for performing dimensionality reduction conversion on the fusion feature to obtain a first feature sequence; an injection sub-module for injecting the position information of the vehicle into the first feature sequence to obtain a second feature sequence; a first processing sub-module for performing mapping processing on the second feature sequence to obtain a feature matrix; a second processing sub-module for performing normalization processing on the feature matrix to obtain the target feature.

[0022] Optionally, the detection unit may include: a retrieval module for retrieving query vectors, where the query vectors correspond to the lane lines in the area one by one; a prediction module for using a decoder to perform regression prediction on the query vectors and the target feature to obtain the detection results.

[0023] Optionally, the first determination unit may include: an acquisition module configured to acquire the positioning information of the vehicle; a first determination module configured to determine the global coordinate position of the vehicle in the global coordinate system based on the positioning information; a second determination module configured to determine local map data from the prior map according to the global coordinate position; an encoding module configured to perform feature encoding on the lane lines in the local map data to obtain encoded data; a processing module configured to perform pooling processing on the encoded data to obtain the pooling result of the lane lines in the local map data; and a conversion module configured to perform conversion on the pooling result to obtain the second bird's-eye view feature.

[0024] Optionally, the fusion unit may include: a splicing module configured to splice the first bird's-eye view feature and the second bird's-eye view feature to obtain a spliced feature; and an extraction module configured to perform feature extraction on the spliced feature to obtain a fused feature.

[0025] Optionally, the apparatus may further include: a matching unit configured to match the detection result with the prior map to obtain the similarity between the detection result and the prior map; and an output unit configured to output the similarity.

[0026] According to another aspect of the embodiments of the present application, there is provided an electronic device, including a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the above method.

[0027] According to another aspect of the embodiments of the present application, there is provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the above method when run by a processor.

[0028] According to another aspect of the embodiments of the present application, there is provided a computer program product including a computer program, which implements the above method when executed by a processor.

[0029] According to another aspect of the embodiments of the present application, there is provided a vehicle, including an in-vehicle processor and an in-vehicle memory, wherein the in-vehicle memory can be used to store a computer program; and the in-vehicle processor is configured to execute the computer program stored on the memory to implement the above method.

[0030] It should be noted that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. Description of the Drawings

[0031] Figure 1 is a flowchart of a method for determining a detection result provided by an embodiment of the present application;

[0032] Figure 2It is a flowchart of another end-to-end lane line detection based on prior map feature fusion provided by an embodiment of the present application;

[0033] Figure 3 It is a schematic diagram of an end-to-end lane line detection based on prior map feature fusion provided by an embodiment of the present application;

[0034] Figure 4 It is a structural diagram of a device for determining a detection result provided by an embodiment of the present application;

[0035] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0036] In order to make the technical problems, technical solutions and beneficial effects solved by the present application clearer, the present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0037] Currently, lane line detection is one of the key technologies in autonomous driving and Advanced Driver Assistance Systems (ADAS for short). Precise lane line detection helps the vehicle to complete self-positioning, path planning and driving control, especially in complex road environments, such as highways, urban roads, intersections, etc.

[0038] In related technologies, in order to meet the requirements of high-precision and high-robustness in autonomous driving, existing lane line detection technologies usually rely on the fusion of multi-source data such as visual sensors, such as cameras, lidar, and prior maps. This method obtains road images or point cloud data and combines deep learning models, such as Convolutional Neural Network (CNN for short), to extract the geometric and semantic information of lane lines. However, the above method has the following problems:

[0039] First, the environmental complexity and perception robustness are insufficient: in actual driving scenarios, lane lines become blurred due to road wear, lighting changes, rain, fog and other bad weather conditions. The above method is only a lane line detection algorithm based on single visual perception and is difficult to handle extremely complex situations, resulting in unstable or inaccurate lane line detection;

[0040] II. Difficulties in data fusion: Although existing methods attempt to use multi-modal data such as lidar, cameras, and prior maps for lane detection, during the process of multi-source information fusion, the heterogeneity and synchronization issues of different data sources are prominent. Improper fusion may lead to information redundancy or loss, reducing the detection accuracy of lane lines.

[0041] III. Poor dynamic adaptability of prior maps: The lane line information provided by prior maps can improve the detection accuracy. However, when the road environment changes (such as temporary construction or road marking changes), the detection method based on static prior maps cannot be updated in real time, resulting in misdetection or missed detection of lane lines.

[0042] IV. Limitations of real-time performance and computational efficiency: To achieve real-time response in autonomous driving, the lane detection system not only requires high accuracy but also fast computational capabilities. However, although some existing large deep learning models perform excellently under experimental conditions, their computational complexity is high, making it difficult to meet the requirements of real-time processing on actual vehicles.

[0043] In view of the above problems, the present application proposes a method for determining detection results. This method can be applied to the field of autonomous driving technology, especially in the perception and environment understanding module of autonomous vehicles, and is related to the technologies of lane detection and map construction. This method can help the vehicle to sense the road environment in real time, accurately detect and infer the geometric shape and semantic information of lane lines, thereby realizing the safe navigation and path planning of the vehicle. In addition, this method is also applicable to scenarios such as complex urban roads, intersections, and highways, especially in cases where high-precision lane detection is required, such as in applications like unmanned driving, vehicle autonomous positioning, and map generation and update. It should be noted that only examples are given here, and the scope of use of the method for determining detection results proposed in the present application is not specifically limited.

[0044] In this embodiment, through multi-source feature fusion and self-attention mechanism, the adaptability of the system to environmental changes is enhanced, and misjudgment or missed judgment is reduced; by combining convolutional neural network and self-attention mechanism, while ensuring high accuracy, the computational efficiency is improved, meeting the real-time requirements in autonomous driving; by combining real-time perception data (i.e., image data) with prior map information, the lane line inference result is automatically adjusted to adapt to the changes in the road environment, solving the limitations brought by the static nature of prior maps. This method is applicable to various scenarios such as highways, urban roads, and complex intersections, and shows better reliability especially in cases where high-precision lane detection requirements are relatively high.

[0045] The method for determining the above detection result provided by the embodiment of the present application achieves the following technical effects: obtaining image data of the area where the vehicle is currently located, where the image data is used to represent the multi-angle view of the area; extracting the first bird's-eye view feature in the image data, where the first bird's-eye view feature is at least used to represent the road information in the image data; determining local map data corresponding to the area from the prior map, and extracting the second bird's-eye view feature in the local map data, where the second bird's-eye view feature is at least used to represent the road information in the local map data; fusing the first bird's-eye view feature and the second bird's-eye view feature to obtain a fused feature; determining the detection result of the lane line in the area based on the fused feature, that is, in this embodiment, by fusing the second bird's-eye view feature of the local map data in the prior map with the first bird's-eye view feature of the image data collected in real time, a fused feature is obtained, and further, the fused feature can be determined to obtain the monitoring result of the lane line, thus completing the accurate detection of the lane line, and further achieving the technical effect of improving the monitoring efficiency of the lane line and solving the technical problem of low monitoring efficiency of the lane line.

[0046] Embodiment 1

[0047] The embodiment of the present application provides a method for determining a detection result, Figure 1 is a flowchart of a method for determining a detection result provided by an embodiment of the present application. Refer to Figure 1 and the method may include the following steps:

[0048] S102: Obtain image data of the area where the vehicle is currently located, where the image data is used to represent the multi-angle view of the area.

[0049] In step S102, the above area may be the location or environment where the vehicle is currently located. The above image data may be a real-time image, may be an environmental image collected in real time by a multi-camera system (such as a panoramic camera) around the vehicle, may be a bird's-eye view image, or may also be called a perspective view image, and may be used to represent the multi-angle view of the area. It should be noted that this is only an example, and there is no specific limitation on the type of device for collecting image data and the form of presentation of the image data. As long as the image data is used to represent the road conditions of the area where the vehicle is currently located, it should be within the protection scope of the present application.

[0050] Optionally, by using a panoramic camera installed on the vehicle, real-time images of the area (or environment) where the vehicle is currently located can be obtained to obtain image data. The image data can be used to determine the road information in the area. For example, it can be used to determine data such as the number of vehicles on the road, the number and type of lane lines in the road.

[0051] For example, images of the vehicle's surrounding environment can be captured by multiple cameras installed around the vehicle, and then these images can be stitched together to form a 360-degree bird's-eye view image around the vehicle. Alternatively, image data of a perspective view (PV) can be captured by a panoramic camera.

[0052] S104: Extract the first bird's-eye view feature from the image data, where the first bird's-eye view feature is at least used to characterize the road information in the image data.

[0053] In step S104, the above first bird's-eye view feature can be obtained by extracting from the image data and can be represented by F BEV (u, v). The above road information can be used to characterize the shape of the road where the vehicle is located, including straight roads, curves, fork points, etc., and the distribution of lane lines, such as one-way roads, multi-lane roads, lane-changing areas, etc., and the types of lane lines, for example, solid lines, dashed lines, etc. It should be noted that this is only an example, and the type of road information is not specifically limited here.

[0054] Optionally, the first bird's-eye view feature can be obtained through the following steps: The real-time image of the vehicle's current environment can be obtained through the panoramic camera of the vehicle to obtain image data. Further, the image feature (PV feature) of the image data can be extracted based on a deep learning model (such as ResNet-50), and the Lift-Splat-Shoot (LSS) model can be used to convert the image feature into a bird's-eye view (BEV) representation to obtain the first bird's-eye view feature. By using the first bird's-eye view feature, it can be ensured that the system can perceive the environment around the vehicle. The above first bird's-eye view feature can be used to provide environmental information such as roads and lane lines around the vehicle as the basis for subsequent processing.

[0055] For example, obtain the real-time image of the panoramic camera and preprocess the real-time data to adjust each perspective view image (I PV ) obtained from the panoramic camera to the input size suitable for ResNet-50: I′ PV = resize(I PV , size=(H, W)). I′ PV can be used to represent the adjusted image data.

[0056] Furthermore, the adjusted image data can be normalized, mapping the pixel values corresponding to the image data from (0, 255) to (0, 1), and subtracting the mean and standard deviation to meet the pre-training requirements of ResNet-50:

[0057] I″PV = I' PV - μ ImageNet / σ ImageNet

[0058] where μ ImageNet can be used to represent the mean of a dataset (e.g., ImageNet); σ ImageNet can be used to represent the standard deviation of the ImageNet dataset.

[0059] After obtaining the normalized image data, ResNet-50 can be used to extract the two-dimensional perspective features (PV features) of the image. Among them, the above ResNet-50 can be composed of multiple convolutional layers and residual blocks, and multi-scale features can be extracted through the above convolutional operations. Input the normalized image I″ PV into the convolutional layer of ResNet-50 for forward propagation:

[0060] F PV = ResNet50(I″ PV )

[0061] where F PV is the PV feature extracted through convolution and residual blocks, and can be a multi-channel two-dimensional (TwoDimensional, abbreviated as 2D) feature map. The PV feature map F PV can have a dimension of C × H' × W'. C can be the number of feature channels, and H' and W' can be the spatial dimensions after the convolutional operation.

[0062] In the Lift stage, the 2D PV features can be lifted to a three-dimensional (ThreeDimensional, abbreviated as 3D) space through the LSS model, and the depth probability of each pixel point can be predicted using the depth estimation module. Estimate the depth probability distribution P(z|F PV ) for the pixel points in each PV feature F PV . This process can be achieved through a depth prediction network:

[0063] P(z i |F PV ) = softmax(z i )

[0064] where z i represents candidate points at different depth positions, and the predicted depth distribution determines the position of the 2D feature points in the 3D space. According to the depth probability distribution P(z|F PV ), lift each 2D feature point (x PV , y PV ) to the 3D voxel space (x BEV , yBEV , z BEV ):

[0065] (x BEV , y BEV , z BEV ) = f((x PV , y PV , P(z|F PV ))

[0066] where f(·) is a lifting function from 2D to 3D.

[0067] In the Splat stage, the feature point cloud lifted to the 3D space can be projected onto the BEV (Bird's Eye View) plane. The voxel points in the 3D point cloud are projected onto the two-dimensional BEV grid plane. Assume the height of the BEV plane is 0, that is, only the information on the (x, y) plane is retained:

[0068] (u, v) = II((x BEV , y BEV )

[0069] where II(·) is a projection operation, and (u, v) are the coordinates projected onto the BEV grid.

[0070] Finally, for multiple voxel points projected onto the same BEV grid, through max pooling or weighted average operations, multiple features can be aggregated into a single feature vector, which can be the first bird's eye view feature:

[0071] F BEV (u, v) = MaxPooling(F BEV,1 , F BEV,2 , …, F BEV,n )

[0072] where F BEV,i is the voxel point feature projected onto the same (u, v) position. The final BEV feature map F BEV is a two-dimensional feature map, and each grid position (u, v) can contain the aggregated high-dimensional feature vector. The above first bird's eye view feature can reflect information such as roads, lane lines, and obstacles in the surrounding environment. where C BEV is the number of channels of the BEV feature, and H BEV , W BEV are the spatial dimensions of the BEV feature map.

[0073] S106: Determine local map data corresponding to the area from the prior map, and extract the second bird's-eye view feature in the local map data, where the second bird's-eye view feature is at least used to characterize the road information in the local map data.

[0074] In step S106, the above prior map can be a pre-constructed high-precision prior map containing detailed road network information, which can be collected and produced by a dedicated survey vehicle or data collection device before the vehicle runs, and can include geometric features of the road (such as the shape and position of lane lines), road attributes (such as lane line types, road surface types, traffic sign positions, etc.), and static elements of the road environment (such as buildings, street lights, traffic signals, etc.). The above prior map can be a vector map, where roads, lane lines, etc. can be represented as vector objects such as points, lines, polygons, etc., or a raster map, which divides the road environment into many grid cells. It should be noted that this is only an example here, and there is no specific limitation on the representation form of the prior map.

[0075] Optionally, the above local map data refers to the map information most relevant to the current operation of the vehicle extracted from the prior map according to the current position and driving direction of the vehicle, which can include road structures, lane line information, traffic signs, etc. within a certain range around the vehicle, and can be used for real-time navigation and environmental understanding of the vehicle. It can be obtained by extracting from the prior map based on the vehicle's positioning information (such as GPS, IMU, etc.), and is used to ensure that the obtained map information matches the current position and driving direction of the vehicle. It should be noted that this is only an example here, and there is no specific limitation on the acquisition method of the local map data.

[0076] Optionally, the above second bird's-eye view feature (F BEV-map ) can be the feature information extracted from the local map data, and can be represented as the road environment feature observed from the vertical overhead view angle from above. The above road environment features can include geometric information such as the position, type, and width of lane lines, as well as semantic information such as road surface markings, road attributes, and obstacle types. The above second bird's-eye view feature can be used to assist the real-time perception system, and by combining with the first bird's-eye view feature perceived in real time, it can improve the vehicle's understanding of the surrounding environment and the accuracy of lane line detection.

[0077] In this embodiment, the first bird's-eye view feature comes from the image data perceived by the vehicle in real time and can capture the lane lines and road conditions in the current environment; the second bird's-eye view feature is based on high-precision prior map information and provides prior knowledge of the road structure. Fusing these two types of features can make up for the limitations of a single data source, enable the system to perform better in the face of environmental changes or complex road scenarios, and effectively improve the accuracy and reliability of lane line detection.

[0078] Optionally, from the prior map, local map data corresponding to the current area can be determined, and the neural network model can be used to extract the local map data to obtain the second bird's-eye view feature. It should be noted that the method for obtaining the second bird's-eye view feature here is only for illustrative purposes and is not specifically limited here.

[0079] For example, the second bird's-eye view feature can be obtained through the following steps: The prior map can be captured and parsed through the positioning information, the vectorized prior map feature can be extracted using a multi-layer perceptron, combined with the vehicle's positioning information, the feature of the area where the current vehicle is located can be extracted from the prior map, and based on this feature, the local map data in the prior map can be determined. Among them, the above prior map can include information such as road structure, lane type, and color, and these data can improve the accuracy of lane line detection.

[0080] S108: Fuse the first bird's-eye view feature and the second bird's-eye view feature to obtain a fused feature.

[0081] In step S108, the above fused feature can be represented by F fusion It is expressed.

[0082] Optionally, the first bird's-eye view feature (i.e., the real-time BEV feature) and the second bird's-eye view feature are fused to obtain a fused feature containing multi-source information.

[0083] Optionally, the first bird's-eye view feature obtained through the panoramic camera is fused with the second bird's-eye view feature extracted from the vectorized prior map to obtain a fused feature.

[0084] For example, the BEV features from two sources can be spliced together to obtain a fused feature containing multi-source information, and then downsampled through a CNN to further compress the information to retain key features while reducing the data dimension to ensure computational efficiency. These fused features can more comprehensively describe the current environment and its corresponding prior map features, providing a basis for subsequent deep feature learning.

[0085] In this embodiment, the first bird's-eye view feature (the real-time BEV feature from the panoramic camera) and the second bird's-eye view feature (the BEV feature from the prior map) are fused to obtain a fused feature. This process aims to integrate real-time perception information and prior geographical information to obtain a more comprehensive and accurate environmental description, thereby improving the accuracy and robustness of lane line detection.

[0086] For example, the first bird's-eye view feature obtained in real time can be concatenated with the second bird's-eye view feature in the prior map in the channel dimension to obtain a fused feature. The fused feature not only contains the real-time perceived environmental changes but also incorporates the high-precision road information of the prior map, providing a comprehensive description of the vehicle's surrounding environment. This step ensures that the independent information of the two feature sources is retained while being able to capture the complementarity and differences between real-time perception and prior information. Further, the concatenated feature can be downsampled by a convolutional neural network (CNN) to compress the data dimension, extract important geometric and semantic information, and ensure computational efficiency. While downsampling, the CNN also performs feature learning to further enhance the expressive power of the feature.

[0087] S110: Determine the detection result of the lane lines in the region based on the fused feature.

[0088] In step S110, the fused feature can be recognized to determine the detection result of the lane lines in the region. The above detection result can include the geometric shape of the lane lines (such as position and curve parameters) and semantic types (such as information on dashed lines, solid lines, etc.). It should be noted that this is only an example, and there is no specific limitation on the type of the detection result.

[0089] Optionally, through a pre-set neural network, such as a neural network composed of a self-attention mechanism to enhance feature representation and a corresponding encoder-decoder, the fused feature is recognized to obtain the detection result. Through the above method, end-to-end monitoring can be achieved. It should be noted that this is only an example, and there is no specific limitation on the method of determining the detection result.

[0090] For example, the self-attention mechanism can be used to process the fused feature. This mechanism can capture the spatial correlation and global context information between features, which helps to more accurately infer the geometric shape and semantic information of the lane lines in complex scenarios (such as multi-lane intersections). The fused feature processed by the self-attention mechanism is fed into a query-based decoder. The initialized query vector in the decoder interacts with the fused feature to infer the specific position, shape, and type of the lane lines (such as solid lines, dashed lines, double yellow lines, etc.). The query-based decoder predicts the semantic type and geometric parameters of the lane lines through a classification and regression network respectively, and finally outputs the complete information of each lane line, including its accurate position, shape, and type.

[0091] Optionally, in the above multi-lane road scenario, the fused features are input into a self-attention mechanism, which learns the spatial relationships between lane lines and their associations with road edges, vehicles, and other obstacles. The query decoder uses these processed fused features to decode lane line information, and it can simultaneously infer the type and shape of each lane line. Even in poor lighting conditions or when lane lines are worn, with the supplement of prior map information, it can ensure the accuracy and robustness of lane line detection results. The finally output lane line detection results contain the detailed position information and types of each lane line, providing a reliable basis for lane keeping and path planning of autonomous vehicles.

[0092] In this embodiment, the first bird's-eye view feature of the environment is obtained through a vehicle surround camera to ensure that the system can perceive the surrounding environment in real time. Subsequently, combined with vehicle positioning information, the second bird's-eye view feature of the corresponding area is extracted from the prior map, and effective prior map features are extracted through a multi-layer perceptron and a max pooling layer. In the feature fusion step, the BEV feature of the surround camera (i.e., the first bird's-eye view feature) is concatenated with the vectorized prior map feature (i.e., the second bird's-eye view feature), and further downsampled by a CNN to compress information and retain key features. Then, a self-attention mechanism is used to perform deep learning on the fused multi-source BEV features to enhance the understanding of spatial associations and global information. Finally, through a lane line decoder based on query decoding, the fused features are inferred to respectively infer the geometric shape and semantic information of the lane lines, specifically including classifying and regressing and predicting the position and type of the lane lines (such as dotted lines, solid lines, etc.). At the same time, the detection results of the lane lines and the similarity between the online mapping and the prior map are output. Through the above method, the accuracy, robustness, and real-time performance of lane line detection can be improved.

[0093] In related technologies, if only radar is used for lane line detection, it is easily interfered by environmental factors (such as buildings or obstacles), resulting in inaccurate echo signals, especially in complex urban or tunnel scenarios. In addition, the synthetic aperture radar (SAR) imaging process is computationally complex and difficult to meet the real-time requirements in high-speed scenarios. To solve this problem, this application combines a vehicle surround camera with a prior map for feature fusion, reduces the dependence on radar, uses an efficient convolutional neural network (CNN) and a self-attention mechanism to improve the computational efficiency, and enhances the robustness and adaptability of the system to complex environments.

[0094] In the related art, complex data fusion of multiple cameras and lidar can also be relied on to complete the prediction of lane lines. However, in adverse weather or insufficient lighting conditions, the performance of the sensors is unstable. In addition, processing multi-sensor data requires high computing resources, affecting the real-time performance of the system and its detection ability in complex scenarios. To solve the above problems, the present application adopts multi-source feature fusion of a panoramic camera and a prior map, enhances the understanding of spatial and global information through a self-attention mechanism, simplifies the complex fusion process of multi-sensors, and at the same time maintains the real-time performance and robustness of detection.

[0095] In addition, in the related art, relying on a single in-vehicle camera for image detection is susceptible to environmental conditions such as light and weather, resulting in inaccurate detection results. Moreover, a single geometric model performs poorly in complex roads (such as sharp turns and fork intersections), affecting the detection accuracy. In this embodiment, by combining the geometric features of the prior map with the real-time image data captured by the camera, the detection ability of the system in the case of blurred or damaged lane lines is enhanced. At the same time, the fusion method of multi-source data improves the inference accuracy of geometric shapes and semantic information in complex scenarios.

[0096] The embodiment of the present application enhances the ability to extract geometric and semantic information of lane lines in complex scenarios by fusing multi-source features of a panoramic camera and a vectorized prior map and using a self-attention mechanism, reducing the detection error caused by environmental changes. By fusing the features of the panoramic camera and the prior map and combining CNN downsampling with the self-attention mechanism, the heterogeneity problem of multi-modal data is effectively solved, and the expression ability of the fused features is improved. This embodiment can automatically adjust and update the inference result of the lane line according to the real-time perceived camera information, combined with the prior map, and at the same time output the detection result of the lane line and the similarity between the online mapping and the prior map. It adapts to the dynamic changes of the road environment, avoiding the limitations brought by the static prior map. By combining an efficient CNN with the self-attention mechanism, the speed and computational efficiency of lane line detection are improved, enabling real-time lane line detection in an autonomous driving system and meeting the requirements in actual driving scenarios.

[0097] Based on the above steps S102 to S110, image data of the current area where the vehicle is located is obtained, where the image data is used to represent multi-angle views of the area; the first bird's-eye view feature in the image data is extracted, where the first bird's-eye view feature is at least used to represent road information in the image data; from the prior map, local map data corresponding to the area is determined, and the second bird's-eye view feature in the local map data is extracted, where the second bird's-eye view feature is at least used to represent road information in the local map data; the first bird's-eye view feature and the second bird's-eye view feature are fused to obtain a fused feature; based on the fused feature, the detection result of the lane lines in the area is determined. That is, in this embodiment, according to the second bird's-eye view feature of the local map data in the prior map, it is fused with the first bird's-eye view feature of the image data collected in real time to obtain a fused feature, and further, the fused feature can be determined to obtain the monitoring result of the lane lines, thus completing the accurate detection of the lane lines, and further achieving the technical effect of improving the monitoring efficiency of the lane lines and solving the technical problem of low monitoring efficiency of the lane lines.

[0098] Next, the above method of this embodiment will be further introduced.

[0099] As an optional implementation manner, in step S110, based on the fused feature, determining the detection result of the lane lines in the area includes: converting the fused feature to obtain a target feature, where the target feature is different from the fused feature; detecting the target feature to obtain a detection result.

[0100] In this embodiment, the fused feature can be converted through a self-attention mechanism to obtain a target feature. This target feature can be a feature sequence with enhanced spatial correlation and global information understanding.

[0101] Optionally, after obtaining the fused feature, a self-attention mechanism can be used to further learn the fused feature after fusion (that is, the multi-source BEV feature) to obtain a target feature. By using the self-attention mechanism, the spatial correlation between features can be effectively captured, enhancing the system's understanding and perception of global information, which helps the system better infer the geometric shape and semantic information of lane lines in complex road scenarios (such as multi-lane, intersections, etc.).

[0102] Optionally, after obtaining the target feature, a lane line decoder based on query decoding can be used to process the target feature to output a detection result.

[0103] Optionally, through a lane line decoder based on query-based decoding, inferences can be made on the fused features processed by the self-attention mechanism. This decoder can simultaneously infer the geometric shape and semantic information of the lane lines, which may include the position, shape, type (such as dashed line, solid line, etc.) of the lane lines, and perform classification and regression predictions on these features, and finally output the detection results of the lane lines. In addition, the detection results of the lane lines and the similarity between the online mapping and the prior map can also be output.

[0104] As an alternative implementation, the fused features are transformed to obtain target features, including: performing dimensionality reduction transformation on the fused features to obtain a first feature sequence; injecting the position information of the vehicle into the first feature sequence to obtain a second feature sequence; performing mapping processing on the second feature sequence to obtain a feature matrix; and performing normalization processing on the feature matrix to obtain the target features.

[0105] In this embodiment, after obtaining the fused features, the fused features (F fusion ) can be first subjected to dimensionality reduction transformation to obtain a first feature sequence. Further, the position information of the vehicle can be injected into the first feature sequence to obtain a second feature sequence (F pos ). Mapping processing can be performed on the second feature sequence to obtain a feature matrix. Normalization processing is performed on the feature matrix to obtain the target features. The above position information can be used to determine the current position of the vehicle, which can be coordinate information or longitude and latitude information, etc. It should be noted that only examples are given here, and the type of the position information is not specifically limited, as long as it is information used to determine the current position of the vehicle, it can be the position information in this application.

[0106] Optionally, after obtaining the fused features, the shape of the feature map F fusion after convolutional fusion is (B, C, H, W), where B is the batch size, C is the number of channels (i.e., the output channel number after convolutional fusion), and H and W are the height and width of the feature map respectively. Before entering the Transformer Encoder, it is usually necessary to flatten the 2D feature map into a serialized vector representation (token) so that the Transformer can process it. Therefore, F fusion can be transformed from the shape (B, C, H, W) to (B, N, C) to obtain a first feature sequence. Where N = H × W is the total number of pixels in the feature map, and the feature vector dimension at each position is C (the number of channels).

[0107] Furthermore, the sine positional encoding can be used to explicitly inject the position information of the vehicle into the input to ensure that the Transformer can understand the relative spatial relationships between the positions in these sequences. pos = F seq + P pos . Among them, F seq is the flattened feature sequence (i.e., the first feature sequence), and P pos is the positional encoding, which can be the position information or the encoding information determined based on the position information.

[0108] Optionally, after obtaining the second feature sequence, the self-attention mechanism can be used to perform a mapping process on the second feature sequence to obtain a feature matrix. Among them, the role of the above second feature sequence is to enable the features at each position to pay attention to the features at other positions in the input sequence, so as to capture the global context information. The input feature sequence F pos can be mapped to a feature matrix, and this feature matrix can include three matrices, namely: query (Q), key (K), and value (V):

[0109] Q = F pos W Q , K = F pos W K , V = F pos W V

[0110] Among them, W Q , W K , W V are learnable weight matrices. Calculate the similarity (i.e., attention score) between each position through the query matrix Q and the key matrix K, and then normalize these scores:

[0111]

[0112] Among them, d k is the dimension of the key vector, and Softmax is used to normalize the attention weights. To enhance the expression ability of the model, the Transformer can calculate multiple independent attention heads (heads) in parallel and then splice their results together:

[0113] Multi-head Attention = Concat(head1, head2, …, head h )W O

[0114] Among them, W Ois the output weight matrix, and h is the number of attention heads. The multi-head attention mechanism can capture different feature relationships in different subspaces. Transformer uses a residual connection after the self-attention layer, that is, adding the input features and the attention output:

[0115] F att = LayerNorm(F pos + Attention(Q, K, V))

[0116] After that, the output features at each position can pass through a two-layer feed-forward neural network. For example, it can be composed of two linear layers and a non-linear activation function (such as ReLU):

[0117] F FFN = ReLU(F att W1 + b1)W2 + b2

[0118] Furthermore, multiple such Transformer Encoders can be utilized to extract the global information of features multiple times and capture the long-range dependencies in the sequence. Finally, after being processed by multiple encoding layers, the Transformer Encoder outputs a new feature sequence, and this feature sequence can be the target feature. The shape of the target feature can be consistent with the input (that is, the fused feature) (B, N, C), that is, the features at each position have undergone multi-layer context information fusion and contain global information. These features can be used for further lane line perception tasks by passing these context-rich features to the downstream query decoding network to decode the lane line position and type.

[0119] As an alternative implementation, the target feature is detected to obtain a detection result, including: retrieving a query vector, where the query vector corresponds to each lane line in the region; using a decoder to perform regression prediction on the query vector and the target feature to obtain the detection result.

[0120] In this embodiment, after obtaining the target feature, a query vector can be retrieved. The query vector can be a latent representation and can be used to represent the geometric shape and semantic features of a lane line, which can be represented by and can correspond to each lane line one by one. That is, a query vector can be used to represent the information corresponding to a lane line. The above decoder can be a lane line decoder and can be used to perform regression prediction on the query vector and the target feature to obtain the detection result.

[0121] Optionally, the features obtained from multi-modal fusion Input to the lane line decoder based on query decoding. Initialize a set of query vectors Q = {q1, q2, …, q N}, where N is the number of queries. For each query vector q i , it can interact with the fused feature F final through the self-attention mechanism. The self-attention mechanism aggregates highly concerned features by calculating the attention distribution between the query vector and the features:

[0122]

[0123] where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the feature dimension. In this way, each query vector can find the most relevant features from the fused feature and learn the geometric information of the lane line. After being processed by the self-attention mechanism, each query vector q i will contain the geometric and semantic information of the lane line. Next, these query vectors are decoded through a fully connected layer to infer the geometric shape (position and curve parameters) and semantic type (such as dashed line, solid line, etc.) of the lane line respectively.

[0124] For each query vector q i , the geometric shape of the lane line is predicted through a regression model. The lane line can be represented using a polynomial or a curve equation. The shape of the lane line is controlled by multiple geometric parameters, such as the starting point, ending point, curvature, etc. Using a polynomial model to perform regression on the lane line, assuming the lane line can be represented as a cubic polynomial:

[0125] y = a3x 3 + a2x 2 + a1x + a0

[0126] where a3, a2, a1, a0 are the polynomial coefficients obtained by regression, representing the geometric shape of the lane line. Through a fully connected network, the features of each query q i are regressed to obtain the corresponding geometric parameters {a3, a2, a1, a0}, and these parameters determine the shape of the lane line:

[0127]

[0128] where f regression (·) is the regression network.

[0129] In addition to the geometric shape of the lane line, it is also necessary to predict the semantic information of the lane line, such as the lane line type (dashed line, solid line, double yellow line, etc.). This part of the task is usually completed through a classification network. Through a fully connected classifier for each query vector q iClassify to predict the semantic category of the lane lines. The classifier outputs the category probability distribution P(c i ):

[0130] P(c i ) = Softmax(W cls q i +b cls )

[0131] where W cls and b cls are the weights and biases of the classifier, and c i is the category label of the lane line. The final semantic classification output is the lane line type Select the category according to the maximum value in the probability distribution:

[0132]

[0133] Output the lane line detection result. The detection result L i of each lane line contains its geometry and semantic type representing the complete information of the lane line:

[0134]

[0135] In this embodiment, a similarity measure between the online map building and the prior map can also be introduced as an auxiliary regression task and output, and the lane line detection is combined with the map similarity regression to achieve multi-task learning during the training process, optimizing two objectives: lane line detection and the similarity regression between the online map building and the prior map. The similarity regression task is to calculate the similarity between the online map building and the prior map for each map region, combining both geometric similarity and semantic similarity. The position similarity between the online map building and the prior map:

[0136]

[0137] where can be used to represent the position information in the online map building. can be used to represent the position information in the prior map.

[0138] The curvature similarity between the online map building and the prior map:

[0139]

[0140] where κ build (x i ) can be used to represent the online map building curvature. κ map (x i ) can be used to represent the prior map curvature.

[0141] Online mapping and prior map type similarity:

[0142]

[0143] where t build,i can be used to represent the online mapping type. t build,i can be used to represent the prior map type.

[0144] Comprehensive similarity metric assisted regression loss:

[0145] ε sim = λ1ε geo + λ2ε curvature + λ3ε semantics

[0146] where the above λ1, λ2, λ2 can be calibration values.

[0147] End-to-end learning model joint optimization objective:

[0148]

[0149] For N lane lines, the final output detection result set is:

[0150]

[0151] As an optional implementation, in step S106, from the prior map, determine the local map data corresponding to the area, and extract the second bird's-eye view features in the local map data, including: obtaining the positioning information of the vehicle; based on the positioning information, determining the global coordinate position of the vehicle in the global coordinate system; according to the global coordinate position, determining the local map data from the prior map; performing feature encoding on the lane lines in the local map data to obtain encoded data; performing pooling processing on the encoded data to obtain the pooling result of the lane lines in the local map data; and performing transformation on the pooling result to obtain the second bird's-eye view features.

[0152] In this embodiment, through the positioning information of the vehicle (such as GPS, IMU, odometer, etc. information), determine the global coordinate position (x vehicle , y vehicle , θ vehicle ) of the current vehicle. Where (x vehicle , y vehicle ) can be used to represent the planar position of the vehicle in the global coordinate system, and θ vehicle can be used to represent the heading angle of the vehicle. According to the current position information (x vehicle , y vehicle) The map information of the vehicle's current location can be extracted from a pre-stored high-precision prior map to obtain local map data, which can be a vectorized map. The vectorized map can be transformed into the vehicle's ego coordinate system.

[0153] Optionally, assume that the extracted local area range is W map ×H map , then the prior map corresponding to this area (i.e., the local map data) can be expressed as:

[0154] M region = extract_region(M global ,(x vehicle ,y vehicle ),W map ,H map )

[0155] where M global can be used to represent the global prior map, and M region is the local map, which can contain important information such as lane line geometry and lane line semantics.

[0156] Furthermore, the lane lines in the local prior map M region can be feature-encoded to obtain encoded data. The input data of the local map data can be represented in the form of a tensor T map , and its shape can be [L, N, F], where L can be the number of lane lines. N can be the number of nodes in each lane line. F can be the feature dimension of each node, including road geometry information and semantic attributes.

[0157] Optionally, a multi-layer perceptron is used to encode the features of each node to obtain encoded data (F map ):

[0158] F map = mlp_embed_1(T map )

[0159] where mlp_embed_1 can be a multi-layer perceptron (MLP) network, the input feature dimension is F, and the output feature dimension is:

[0160] D = H BEV *W BEV

[0161] where the above MLP can be composed of a fully connected layer, layer normalization (LayerNorm), and ReLU activation function to enhance the feature expression ability. Finally, the shape of F map can be [L, N, D].

[0162] Perform a max pooling operation on each node of each lane line to obtain the global feature (i.e., the pooling result) of each lane line:

[0163] F maxpool = maxpool(F map , dim = 1, keepdim = True)[0]

[0164] This operation can effectively capture the overall feature information of the lane line, ensuring that important features are not affected by local noise. At this time, the shape of F maxpool is [L, N, D].

[0165] Optionally, concatenate the pooled global feature (i.e., the pooling result) with the feature of each node to enhance the model's ability to jointly represent global and local information:

[0166] F map = concat((F map , F maxpool .repeat(1, N, 1)), dim = -1)

[0167] The shape of the concatenated feature becomes [L, N, 2D], further improving the ability to describe road information. Further, perform another transformation on the fused feature F map using an MLP for non-linear mapping:

[0168] F map = mlp_embed_2(F map )

[0169] This MLP consists of a fully connected layer, layer normalization (LayerNorm), and ReLU to optimize the feature representation to make it more suitable for subsequent tasks. At this time, the shape of F map returns to [L, N, D].

[0170] Finally, perform a max pooling operation on each node of each lane line again to obtain the final global feature of each lane line:

[0171] F map = maxpool(F map , dim = 1)[0]

[0172] This step can highlight key global information while reducing noise interference. Finally, the shape of F map is [L, D]. Subsequently, along dimension D = (H BEV * W BEV) is reshaped to generate the final vectorized prior map high-dimensional feature map (i.e., the second bird's-eye view feature):

[0173] F BEV-map = F map .view(L, H bev , W bev )

[0174] where L represents the number of channels; H bev , W bev are the spatial dimensions of the feature map.

[0175] F BEV-map provides a BEV-view high-dimensional feature representation containing the key geometric and semantic information of the prior map, providing accurate environmental understanding for lane line perception in autonomous driving tasks.

[0176] As an alternative implementation, in step S108, the first bird's-eye view feature and the second bird's-eye view feature are fused to obtain a fused feature, including: splicing the first bird's-eye view feature and the second bird's-eye view feature to obtain a spliced feature; performing feature extraction on the spliced feature to obtain a fused feature.

[0177] In this embodiment, the above-mentioned spliced feature can be represented by F concate .

[0178] Optionally, the first bird's-eye view feature (F BEV-cam ) extracted by the panoramic camera is spliced with the second bird's-eye view feature (F BEV-map ) extracted from the prior map. The splicing operation can be performed on the channel dimension to retain the independent information of the two source features.

[0179] Optionally, the real-time BEV feature and the prior map BEV feature have the number of channels C1 and C2 respectively, and the spatial dimensions of the above two features are the same. Therefore, the two features can be spliced on the channel dimension to generate a spliced feature F concate :

[0180] F concate = concat(F BEV-cam , F BEV-map , axis = 1)

[0181] where the fused feature after splicing retains multimodal information (i.e., real-time perception features and prior map features).

[0182] After splicing, in order to further extract key features and reduce the dimension of the data, the fused spliced feature F can be processed through a convolutional neural network (CNN)concate Perform further feature extraction to obtain fused features. The goal of this step is to downsample the concatenated features, retain important geometric and semantic information, and at the same time reduce the computational burden. The convolution operation can be expressed as:

[0183] F fusion = CNN(F concate )

[0184] Optionally, through the convolution operation of the CNN, higher-level features can be extracted from the fused features, that is, the fused features, which contain the comprehensive information of the panoramic camera and the prior map.

[0185] As an alternative implementation, the method may further include: matching the detection result with the prior map to obtain the similarity between the detection result and the prior map; outputting the similarity.

[0186] In this embodiment, the detection result output by the lane detection algorithm can be standardized to ensure that it uses the same coordinate system and representation as the lane information in the prior map. For example, the detection result may be given in the form of pixel coordinates, lane type, and geometric parameters, while the lane information in the prior map is based on the global coordinate system and a specific lane description. Therefore, preprocessing such as coordinate transformation and unit unification needs to be performed on the detection result to ensure that they can be directly compared with the prior map information. Extract the lane line features related to the detection area from the prior map, including attributes such as the position, type (such as dotted line, solid line), width, and curvature of the lane line. These features will be used for subsequent matching calculations. By comparing the geometric position and semantic information of the lane line in the detection result with that in the prior map, the similarity between the two is calculated. Among them, the above similarity may include geometric similarity, semantic similarity, and comprehensive similarity, etc. It should be noted that the method for determining the similarity here and the type of similarity are only for illustrative purposes and are not specifically limited here.

[0187] Optionally, the calculated similarity can be an important indicator for evaluating the performance of the lane detection algorithm. In an autonomous driving system, it can also be used to trigger system correction or manual takeover. For example, when the similarity between the detection result and the prior map information is low, it may mean that there is a large difference between the current environment and the map information, and the system needs to adopt a more cautious driving strategy or switch to a safer mode, such as reducing the autonomous driving level or completely stopping autonomous driving and switching to manual driving.

[0188] Figure 2 is a flowchart of another end-to-end lane detection based on prior map feature fusion provided by an embodiment of the present application. As Figure 2 shown, the method may include the following steps:

[0189] Step S202: Obtain the real-time image of the environment through the panoramic camera, and extract the first bird's-eye view features in the image.

[0190] In this embodiment, obtain the bird's-eye view (BEV) features of the panoramic camera. Convert these features to a bird's-eye view display to extract the first bird's-eye view features.

[0191] Optionally, the camera parameters of the above panoramic camera may include the viewing angle, focal length, lens distortion coefficient, etc. These parameters can be calibrated in the system to ensure that the images obtained by the camera can be accurately mapped to the BEV features. It should be noted that the above is only an example, and the type of camera parameters is not specifically limited.

[0192] Optionally, the input size of ResNet-50 can be (H, W). It should be noted that the size here can be set according to actual needs.

[0193] Optionally, determine the size of the camera image before entering ResNet-50. This parameter will affect the computational burden and the effect of feature extraction.

[0194] Optionally, this embodiment can use the mean and standard deviation of ImageNet to normalize the input image. These parameters are predefined during the training process of ResNet-50, aiming to keep the input of the feature extraction network consistent.

[0195] In this embodiment, obtain the real-time image of the environment through the camera and convert it into BEV features, so that the road and lane line information around the vehicle can be accurately obtained.

[0196] Optionally, obtain the bird's-eye view (BEV) features of the panoramic camera to achieve the purpose of obtaining the real-time perception information of the environment around the vehicle and providing a basis for subsequent processing.

[0197] Optionally, use the multi-camera system (panoramic camera) around the vehicle to collect environmental images in real time. The images captured by the camera are usually in perspective view. Through a deep convolutional neural network (such as ResNet-50), extract the two-dimensional features (PV features) of these images. ResNet-50 obtains the multi-scale information in the image through multiple convolutional layers and residual blocks. Use the LSS model to lift the 2D PV features to 3D space and convert these 3D features into a bird's-eye view (BEV) representation. This step involves a depth estimation module, which predicts the depth probability of each pixel, thereby converting the 2D information into 3D voxels and further projecting them onto the BEV plane.

[0198] Optionally, the finally obtained first bird's-eye view feature may include the structural information of the vehicle's surrounding environment, such as lane lines, roads, and other information.

[0199] Step S204, grab and parse the prior map through the positioning information, and extract the second bird's-eye view feature in the local map data.

[0200] In this embodiment, the prior map is grabbed and parsed through the vehicle's positioning information to obtain local map data. A multi-layer perceptron and a max pooling layer can be used to extract vectorized prior map features.

[0201] Optionally, the positioning accuracy determines the accuracy of the local area extracted from the prior map. Positioning information with a large error may lead to the extraction of incorrect map areas.

[0202] Optionally, the addition of prior map information can significantly improve the accuracy of lane line detection, especially in complex or multi-lane scenarios. It provides additional geometric and semantic information (such as the type and position of lane lines), enabling the system to make references and comparisons during the detection process. Although the system can still perform lane line detection based on the information of the panoramic camera without the prior map, the detection accuracy and robustness may decrease. Therefore, the accuracy of lane line detection can be improved through the above steps.

[0203] In this embodiment, the vehicle's positioning information (such as GPS, IMU, odometer, etc.) determines the global coordinate position of the vehicle, so as to extract the map information of the area where the vehicle is located from the high-precision prior map.

[0204] Optionally, a multi-layer perceptron and a max pooling neural network are used to extract features from the vectorized prior map. These features include information such as road structure and lane lines. These prior map features, as the system's prior understanding of the vehicle's surrounding environment, enhance the accuracy of lane line detection.

[0205] Step S206, fuse the first bird's-eye view feature and the second bird's-eye view feature.

[0206] In this embodiment, the BEV features are concatenated in the channel dimension, and this operation ensures that the multi-source features do not lose their original information during the fusion process.

[0207] Optionally, when using a CNN to downsample the concatenated features, parameters such as the convolution kernel size, stride, and pooling layer will affect the resolution and computational complexity of the output features. Therefore, the downsampling ratio of the fused features can be set according to actual needs.

[0208] Optionally, by controlling the number of channels after downsampling, the balance between feature retention and computational efficiency can be achieved. Therefore, the number of channels of the fused features can be set according to actual needs.

[0209] Optionally, fusing the BEV features of the panoramic camera with the BEV features of the prior map can enhance the system's understanding ability of the environment and lane information.

[0210] In this embodiment, the real-time perception information and the prior map information are fused to generate a comprehensive feature containing multi-source information.

[0211] Optionally, the real-time BEV features extracted from the panoramic camera and the BEV features extracted from the prior map are concatenated in the channel dimension. This concatenation preserves the independence of the two features. The concatenated multi-source features are downsampled and processed by a convolutional neural network (CNN) to further compress the data dimension while extracting important geometric and semantic information. These fused features can more comprehensively describe the vehicle's surrounding environment, including the real-time perception situation and the prior road structure.

[0212] Step S208, using the self-attention mechanism to learn the fused features to capture the spatial correlation between the features.

[0213] In this embodiment, after the feature map is flattened into a sequence, the dimension N is equal to H×W, and it is necessary to ensure that the flattened sequence can preserve the spatial relationship of the original features.

[0214] Optionally, using sine position encoding is a common choice, but the encoding dimension and pattern (such as linear or non-linear) can be adjusted according to requirements.

[0215] In this embodiment, the self-attention mechanism helps to capture the long-range dependence relationship and global context information between the features, which is very important for accurately detecting the geometric shape and semantic information of lane lines in complex scenarios.

[0216] Optionally, the self-attention mechanism is used to further learn the fused features after fusion to obtain the target features.

[0217] In this embodiment, global information is captured through the self-attention mechanism to enhance the system's understanding of complex scenarios.

[0218] Optionally, the fused multi-modal features are flattened into a serialized vector representation (token), and spatial position information is injected through position encoding. The self-attention mechanism models the correlation between each feature by calculating the similarity between the features. This mechanism can effectively capture the long-range dependence relationship between the features and ensure the system's perception and understanding of the global environment.

[0219] Optionally, a Transformer Encoder is used to perform multi-layer processing on these serialized features, further enhancing the expressive power of the features and integrating global context information into each feature to provide richer input for the subsequent decoder.

[0220] Step S210, based on the lane line decoder, predict the target features to obtain the detection result.

[0221] In this embodiment, through the lane line decoder based on query decoding, classify and regressively predict the fused features processed by the self-attention mechanism to output the detection result of the lane line. Further, the similarity between the detection result and the online mapping and the prior map can be determined.

[0222] Optionally, the above detection result can be used to complete online mapping.

[0223] Optionally, each query vector represents a lane line, and the size of the number (N) of query vectors determines the number of lane lines that the system can detect simultaneously.

[0224] Optionally, the initial value and dimension (d) of the query vector need to be reasonably set to ensure that it can cover the geometric and semantic features of the lane line.

[0225] Optionally, parameters such as the number of heads and dimensions of the self-attention mechanism affect the interaction effect between the query vector and the features. Therefore, the number of heads, dimensions, etc. of the self-attention mechanism can be set according to actual needs, and no specific limitations are made here.

[0226] Optionally, by decoding the target features through the query decoder, the final lane line position and type can be obtained.

[0227] Optionally, initialize a set of query vectors, each query vector representing a potential lane line. These query vectors interact with the features enhanced by the self-attention mechanism, and the feature information related to each query is aggregated together through the attention mechanism.

[0228] Optionally, the query vector contains the geometric and semantic information of the lane line. Decode these query vectors through a fully connected network to infer the geometric shape (such as position, curvature) and type (such as solid line, dotted line, etc.) of the lane line.

[0229] Optionally, the final decoded output is the detection result containing the geometric and semantic information of the lane line and the similarity with the prior map. The system can integrate these results and output the complete information of each lane line and the similarity with the prior map.

[0230] Figure 3 It is a schematic diagram of an end-to-end lane line detection based on prior map feature fusion provided by an embodiment of the present application, asFigure 3 As shown, the panoramic camera image 301 and the vehicle-end real-time fusion positioning pose 302 can be obtained. The panoramic camera image 301 is extracted to obtain the first bird's-eye view feature 303. The prior local map 304 is determined based on the positioning pose 302. The prior local map 304 is processed by a multi-layer perceptron 305 to obtain the max-pooling map feature 306. The first bird's-eye view feature 303 and the max-pooling map feature 306 are fused to obtain a fused feature. The fused feature is downsampled by a convolutional neural network, and the fused feature is processed by a multi-head self-attention mechanism feature fusion network 307 to obtain a target feature. The target feature is processed by a query decoder 308 to obtain a detection result 309. Online mapping can be performed based on the detection result 309, and the similarity between the result of the online mapping and the prior map is regressed to obtain the similarity between the two, which can be output and displayed.

[0231] Optionally, the above prior local map may be local map data. The above max-pooling map feature 306 may be the second bird's-eye view feature 303.

[0232] In this embodiment, the vehicle's surrounding environment information is collected in real time by a panoramic camera, and at the same time, the map features of the vehicle's current area are extracted from the prior map in combination with the vehicle's positioning information. Through multi-modal feature fusion, the real-time perception features and the prior features are stitched and fused to generate richer environmental features. The self-attention mechanism is used to enhance the representation ability of these fused features, and the specific geometric and semantic information of the lane lines is inferred in the query decoder. The system outputs the similarity between the detection result and the prior map, realizing high-precision and robust lane line detection.

[0233] In this embodiment, the real-time image information collected by the vehicle's panoramic camera is fused with the geometric and semantic information in the high-precision prior map. The encoding of the map information adopts a method called vectorization. This method can capture the complex spatial structure in the scene more efficiently compared with the traditional rasterization method (such as converting the scene into a bird's-eye view image). The combination of this multi-source information is an innovation in the field of lane line detection. Compared with using only single perception information (such as only a panoramic camera or only radar), this method can perceive the vehicle's surrounding environment and road structure more accurately and comprehensively.

[0234] In the related art, usually only relying on single-sensor data, such as cameras or radars, may lose effectiveness in some environments (such as occlusion, light changes or weather effects). However, this application provides a stable reference for the road and lane structure through the information of the prior map, improving the robustness and accuracy of the system in complex environments.

[0235] In this embodiment, a complete end-to-end process is designed, from data acquisition, feature extraction, feature fusion to the final lane line decoding, and the above processes are all implemented in a unified deep learning framework. This end-to-end method simplifies the lane line detection process and avoids the error accumulation that may be caused by multi-stage processing. A similarity metric between the online map building and the prior map is introduced as an auxiliary regression task. The lane line detection is combined with the map similarity regression, and these two objectives are jointly optimized through an end-to-end learning model.

[0236] In the related art, for lane line detection, multi-stage processing is often required. For example, feature extraction is first performed, and then a separate model is used for lane line detection. Although this method is flexible, it is prone to error transmission and accumulation. In this application, through the end-to-end method, the final detection result is directly optimized, improving the overall performance of the system. By regressing the similarity between the prior map and the online map, the accuracy and consistency of the current environment model are evaluated, which is used to trigger the downstream optimized post-processing module or manual takeover.

[0237] In this embodiment, after multi-modal feature fusion, a self-attention mechanism is introduced to capture the long-range dependencies between features. This mechanism can better integrate and understand the global environmental information. Especially in complex scenarios such as multi-lane intersections, it can model the lane line geometry and semantic information more accurately.

[0238] In the related art, although it performs well in extracting local features, it has limitations in dealing with global relationships and long-range dependencies. In this application, the self-attention mechanism enhances the system's ability to integrate global information, making the lane line detection in complex scenarios more accurate and stable.

[0239] In this embodiment, a query-based decoding mechanism is introduced. By interacting the query vector with the fused features, the specific geometry and semantic information of the lane line are inferred. This query-based decoding method draws on the Detection Transformer architecture in object detection but is optimized for lane line detection. In the lane line detection step based on query decoding, the beyond-horizon features and the multi-modal fused features are input into the decoder together, so as to realize the joint decoding and prediction of long-distance lane lines. This joint decoding method can simultaneously process the lane line information at close range and long range, and predict the geometric and semantic information of the lane line within different perception ranges.

[0240] Optionally, through a query mechanism, this embodiment can simultaneously infer the type and shape of lane lines and flexibly adjust their geometric parameters in complex scenarios. This method improves the efficiency and flexibility of lane line detection while ensuring high precision. Compared with traditional decoders that only rely on close-range features, the joint decoding method of this application can better handle complex road scenarios, such as highways and multi-lane lane-changing scenarios, to achieve accurate lane line detection within the ultra-long sight range.

[0241] In this embodiment, in the multi-modal feature fusion stage, a dedicated convolutional network structure is designed to process the fusion of real-time BEV features and prior map features. By splicing and downsampling features from different sources, the CNN structure of this application can effectively compress the information dimension while retaining key geometric and semantic features.

[0242] In the related art, only simple feature splicing or superposition methods are used to complete the detection of lane lines, while this application further extracts the key features after fusion through an optimized convolutional network, ensuring computational efficiency without loss of information quality. This optimization enables the system to maintain real-time performance in scenarios with high-precision requirements.

[0243] This application adopts methods such as multi-modal information fusion, self-attention mechanism, and ultra-long sight range perception, which not only improve the accuracy and robustness of lane line detection but also greatly expand the application scenarios and adaptability of the system.

[0244] Optionally, by performing multi-modal fusion of the real-time BEV features, prior map features, and ultra-long sight range perception features of the panoramic camera, this application can detect the lane line information around the vehicle more comprehensively and accurately. The prior map provides high-precision geometric and semantic information to supplement the deficiencies of real-time perception; while the ultra-long sight range perception further expands the detection range, enabling the system to accurately perceive lane information at long distances, thus significantly improving the accuracy of lane line detection.

[0245] Optionally, through the fusion of multi-source information, this embodiment reduces the risk of a single sensor (such as a camera or radar) failing under special conditions (such as strong light, rain, snow, occlusion, etc.). Even if the data quality of a certain sensor deteriorates, the system can still rely on other feature sources (such as the prior map or ultra-long sight range perception) to maintain high detection accuracy. This robustness plays an important role when autonomous vehicles cope with different environmental conditions, ensuring that the vehicle can still operate stably in extreme weather or complex road environments.

[0246] Optionally, after introducing the beyond-visual-range long-distance perception technology in this embodiment, it is possible to detect and predict roads and lane lines in advance at distances far beyond the traditional visual perception range. This enables the autonomous driving system to react in advance when encountering curves, slopes, or complex traffic conditions ahead during high-speed driving, improving safety and the driving stability of the vehicle. When driving on highways or long straight sections, the vehicle can rely on the beyond-visual-range perception information for anticipation and path optimization, reducing the safety risks caused by temporary reactions, thereby significantly enhancing the driving experience and safety.

[0247] Optionally, this embodiment adopts an end-to-end deep learning architecture, reducing the latency caused by multi-stage processing. At the same time, through adaptive feature fusion and an optimized convolutional network structure, it is possible to reduce the computational burden while ensuring accuracy, thereby improving the real-time performance of the system. The autonomous driving system has extremely high requirements for real-time performance, and the efficient computing and fast response capabilities of this application can ensure that the vehicle makes timely decisions and controls in high-speed and complex scenarios.

[0248] Optionally, the introduction of the self-attention mechanism enables the system to capture global information in complex scenarios and can handle complex road scenarios such as multi-lane, intersections, and lane changes. Under this mechanism, the geometric and semantic information of lane lines can be more accurately modeled and inferred. Even in cases where visual perception is blurred, the system can still correctly identify the type and position of lane lines. In the urban driving environment, vehicles often encounter situations such as intersections, lane bifurcations, and multi-lane lane changes. The processing ability of this application can effectively improve the adaptability and response ability of vehicles in complex scenarios. By regressing the similarity between the prior map and the online map, the accuracy and consistency of the current environment model are evaluated, which is used to trigger the downstream optimization post-processing module or manual takeover.

[0249] Optionally, although the prior map provides a large amount of valuable information for the system, in practical applications, there are certain difficulties in data updating and accuracy maintenance of the prior map. Through the supplementation of beyond-visual-range perception and real-time perception in this application, the absolute dependence of the system on a high-precision prior map is reduced. Even when the map data is not completely accurate or updated in a timely manner, the system can still maintain a stable detection ability through other data sources. This makes the system more applicable globally, especially in areas with a low map update frequency or incomplete data, the autonomous driving system can still work properly.

[0250] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0251] Embodiment 2

[0252] According to an embodiment of the present application, a device for determining a detection result is further provided. It should be noted that the device for determining the detection result can be used to execute the method for determining the detection result in Embodiment 1.

[0253] The embodiment of the present application also provides a device 40 for determining a detection result. Please refer to Figure 4 , Figure 4 FIG. is a structural diagram of a device for determining a detection result provided by an embodiment of the present application. The device may include: an acquisition unit 402, configured to acquire image data of the area where the vehicle is currently located, where the image data is used to represent a multi-angle view of the area; an extraction unit 404, configured to extract a first bird's-eye view feature from the image data, where the first bird's-eye view feature is at least used to represent road information in the image data; a first determination unit 406, configured to determine local map data corresponding to the area from a prior map, and extract a second bird's-eye view feature in the local map data, where the second bird's-eye view feature is at least used to represent road information in the local map data; a fusion unit 408, configured to fuse the first bird's-eye view feature and the second bird's-eye view feature to obtain a fusion feature; a second determination unit 410, configured to determine a detection result of lane lines in the area based on the fusion feature.

[0254] The device for determining the detection result provided by the embodiment of the present application achieves the following technical effects: By fusing the second bird's-eye view feature of the local map data in the prior map with the first bird's-eye view feature of the image data collected in real time, a fusion feature is obtained. Further, the fusion feature can be determined to obtain a monitoring result of the lane lines, thereby completing the accurate detection of the lane lines, and further achieving the technical effect of improving the monitoring efficiency of the lane lines, and solving the technical problem of low monitoring efficiency of the lane lines.

[0255] It should be noted that the above-mentioned units can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.

[0256] Embodiment III

[0257] The embodiment of the present application also provides an electronic device 50. Please refer to Figure 5 , Figure 5 which is a structural diagram of an electronic device provided by an embodiment of the present application, including a processor 510 and a memory 520. Among them, the memory 510 is used to store a computer program; the processor 520 is used to execute the program stored on the memory 510 to implement the method for determining the detection result introduced in any embodiment of the present application.

[0258] Optionally, in this embodiment, the above-mentioned processor can be set to execute the following steps through a computer program:

[0259] Step S1, obtain image data of the area where the vehicle is currently located, where the image data is used to represent the multi-angle view of the area;

[0260] Step S2, extract the first bird's-eye view feature in the image data, where the first bird's-eye view feature is at least used to represent the road information in the image data;

[0261] Step S3, determine the local map data corresponding to the area from the prior map, and extract the second bird's-eye view feature in the local map data, where the second bird's-eye view feature is at least used to represent the road information in the local map data;

[0262] Step S4, fuse the first bird's-eye view feature and the second bird's-eye view feature to obtain a fused feature;

[0263] Step S5, determine the detection result of the lane line in the area based on the fused feature.

[0264] The above-mentioned electronic device provided by the embodiment of the present application achieves the following technical effects: when a pose adjustment instruction is obtained, based on the current pose information of the seat, the target pose information, and the abnormal pose information of the obstacles within the target range, determine the control data of the motor to be controlled corresponding to at least one component, and further control the motor to be controlled according to the control data, thereby achieving the technical effect of improving the monitoring efficiency of the lane line and solving the technical problem of low monitoring efficiency of the lane line.

[0265] Those of ordinary skill in the art can understand that Figure 5The structure shown is only schematic, and the electronic device can also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a Mobile Internet Device (MID for short). Figure 5 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device 50 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 5 in the figure, or have a different configuration from that shown Figure 5 in the figure.

[0266] Embodiment 4

[0267] The embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the method for determining the detection result introduced in any embodiment of the present application.

[0268] Optionally, in this embodiment, the above storage medium may be set to store a computer program for performing the following steps:

[0269] Step S1: Obtain image data of the area where the vehicle is currently located, where the image data is used to represent the multi-angle view of the area;

[0270] Step S2: Extract the first bird's-eye view feature in the image data, where the first bird's-eye view feature is at least used to represent the road information in the image data;

[0271] Step S3: Determine the local map data corresponding to the area from the prior map, and extract the second bird's-eye view feature in the local map data, where the second bird's-eye view feature is at least used to represent the road information in the local map data;

[0272] Step S4: Fuse the first bird's-eye view feature and the second bird's-eye view feature to obtain a fused feature;

[0273] Step S5: Based on the fused feature, determine the detection result of the lane lines in the area.

[0274] Optionally, in this embodiment, the above storage medium may include but is not limited to: various media that can store computer programs such as a USB flash drive, a read-only memory (abbreviated as ROM), a random access memory (abbreviated as RAM), a mobile hard disk, a magnetic disk, or an optical disc.

[0275] The above electronic device provided by the embodiments of the present application achieves the following technical effects: By obtaining and analyzing the current pose information of the sub-components and the abnormal pose information of the obstacles within the target area in real time, accurate control data for controlling the sub-components of the seat is obtained, thereby achieving the purpose of intelligently adjusting and optimizing the motor control strategy. In an environment where the internal space of the vehicle is limited and the positions of obstacles are variable, there may be a risk of collision when the sub-components are adjusted. Through the above method, it can be ensured that the seat can be adjusted safely and efficiently in any initial pose, thereby improving the user's comfort and usage experience, solving the technical problem of low monitoring efficiency of lane lines, and achieving the technical effect of improving the safety of seat control.

[0276] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0277] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0278] In the present application, "a plurality of" means two or more.

[0279] In the present application, unless otherwise clearly defined, the terms "installation", "connection", and "coupling" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0280] The terms "first", "second", "third", "fourth", etc. (if any) in the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence.

[0281] The term "and / or" in the present application is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the front and rear associated objects.

[0282] If there is no special instruction, all steps of the present application can be carried out in sequence or randomly. For example, the method includes steps A and B, indicating that the method may include steps A and B carried out in sequence, or may include steps B and A carried out in sequence. For example, it is mentioned that the method may further include step C, indicating that step C can be added to the method in any order. For example, the method may include steps A, B, and C, or may include steps A, C, and B, or may include steps C, A, and B, etc.

[0283] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for determining a detection result, characterized in that Including: Obtain image data of the area where the vehicle is currently located, where the image data is used to represent a multi-angle view of the area; Extract the first bird's-eye view feature from the image data, where the first bird's-eye view feature is at least used to represent road information in the image data; Determine local map data corresponding to the area from a prior map, and extract a second bird's-eye view feature from the local map data, where the second bird's-eye view feature is at least used to represent road information in the local map data; Fuse the first bird's-eye view feature and the second bird's-eye view feature to obtain a fused feature; Based on the fused feature, determine the detection result of lane lines in the area.

2. The method according to claim 1, wherein The determining the detection result of lane lines in the area based on the fused feature includes: Convert the fused feature to obtain a target feature, where the target feature is different from the fused feature; Detect the target feature to obtain the detection result.

3. The method according to claim 2, wherein The converting the fused feature to obtain a target feature includes: Perform dimensionality reduction conversion on the fused feature to obtain a first feature sequence; Inject the position information of the vehicle into the first feature sequence to obtain a second feature sequence; Perform mapping processing on the second feature sequence to obtain a feature matrix; Perform normalization processing on the feature matrix to obtain the target feature.

4. The method according to claim 2, wherein The detecting the target feature to obtain the detection result includes: Retrieve a query vector, where the query vector corresponds one-to-one to lane lines in the area; Use a decoder to perform regression prediction on the query vector and the target feature to obtain the detection result.

5. The method according to claim 1, wherein The determining local map data corresponding to the area from a prior map and extracting a second bird's-eye view feature from the local map data includes: Obtain the positioning information of the vehicle; Based on the positioning information, determine the global coordinate position of the vehicle in the global coordinate system; According to the global coordinate position, determine the local map data from the prior map; Perform feature encoding on lane lines in the local map data to obtain encoded data; Perform pooling processing on the encoded data to obtain the pooling result of lane lines in the local map data; Convert the pooling result to obtain the second bird's-eye view feature.

6. The method according to claim 1, wherein The fusing the first bird's-eye view feature and the second bird's-eye view feature to obtain a fused feature includes: Concatenate the first bird's-eye view feature and the second bird's-eye view feature to obtain a concatenated feature; Extract features from the concatenated feature to obtain the fused feature.

7. The method according to claim 1, characterized in that The method further includes: Match the detection result with the prior map to obtain the similarity between the detection result and the prior map; Output the similarity.

8. A determination device for detection results, characterized in that, Including: An acquisition unit for acquiring image data of the area where the vehicle is currently located, where the image data is used to represent a multi-angle view of the area; An extraction unit, configured to extract a first bird's-eye view feature from the image data, where the first bird's-eye view feature is at least used to characterize road information in the image data; A first determination unit, configured to determine local map data corresponding to the area from a prior map, and extract a second bird's-eye view feature from the local map data, where the second bird's-eye view feature is at least used to characterize road information in the local map data; A fusion unit, configured to fuse the first bird's-eye view feature and the second bird's-eye view feature to obtain a fused feature; A second determination unit, configured to determine a detection result of lane lines in the area based on the fused feature.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-8 is implemented.

10. A vehicle, characterized in that, Including a vehicle-mounted processor and a vehicle-mounted memory, where The vehicle-mounted memory is configured to store a computer program; The vehicle-mounted processor is configured to execute the computer program stored on the memory to implement the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • BEVLane-based lightweight lane line detection method and system

    CN122024211A

  • A lane line detection method and system based on BEV lane

    CN122024211B