Vehicle lane line detection method, electronic device, storage medium, and mobile platform

By dividing the front view into multiple front view and projecting it into multiple top view, and using different resolutions for lane line recognition, the problem of missing lane line under top view in the prior art is solved, and the recognition effect is improved.

WO2025102345A1PCT designated stage expired Publication Date: 2025-05-22SZ ZHUOYU TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/132282
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

In the top view, the existing lane line detection method has a small vertical pixel density, resulting in too few horizontal or large curvature lane line pixel points, resulting in serious missed detection and poor recognition effect.

Method used

By obtaining the front view of the mobile platform, segmenting it from near and far into multiple front views, and projecting these subfront views into multiple top views of different resolutions, where the higher the resolution of the sub-top view closer to the mobile platform, and then lane line recognition is performed.

Benefits of technology

Through multi-resolution input, lane line recognition is used to identify different field of view ranges using different resolutions, which reduces the missed detection of transverse lane lines or large curvature lane lines close to the mobile platform, and improves the recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023132282_22052025_PF_FP_ABST
    Figure CN2023132282_22052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a vehicle lane line detection method, comprising: acquiring a front view of a mobile platform; segmenting the front view into multiple sub-front views from near to far; projecting the multiple sub-front views into multiple sub-bird views of different resolutions, wherein the closer a sub-bird view is to the mobile platform, the higher the resolution of the sub-bird view; and performing vehicle lane line recognition at least on the basis of the multiple sub-bird views. The present application uses multi-resolution input to recognize different field-of-view ranges using different resolutions, solving the problem of missed detection of large-curvature vehicle lane lines and ensuring a sufficient recognition distance as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Lane line detection method, electronic device, storage medium and mobile platform Technical Field

[0001] The present application relates to the field of autonomous driving technology, and in particular to a lane line detection method, electronic equipment, storage medium, and mobile platform. Background Art

[0002] False and missed lane detections significantly impact autonomous driving control and are a long-standing problem in the field. Semantic segmentation is a mainstream method for lane detection, classifying each pixel in a sensor image to identify lane lines, road markings, and other scene information.

[0003] In the process of realizing this application, the inventors discovered that in the lane line detection method of the prior art, in order to ensure sufficient recognition distance under the top view, the vertical pixel density of the top view is often relatively small, resulting in too few pixels for the lateral / large curvature lane lines. This leads to serious missed detection of lateral or oblique lane lines (common on large curvature curves) and poor recognition effect. In addition, the real-time performance required for autonomous driving limits the input resolution size, and the single perspective (top view or front view) limits the recognition ability of the model. Most current methods recognize lane lines based on single-perspective segmentation of the front view (FrontView, FV) image or the bird's-eye view (BirdView, BV) image, but the front view has poor recognition effect for distant objects, and the top view is prone to misdetection of objects similar to lane lines such as guardrails and curbs.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a lane line detection method, electronic device, storage medium and mobile platform, which are used to solve at least one of the above technical problems.

[0006] In a first aspect, an embodiment of the present application provides a lane line detection method, comprising:

[0007] Get the front view of the mobile platform;

[0008] Segmenting the front view into multiple front views from near to far;

[0009] Projecting the multi-molecule front view into multi-molecule top views with different resolutions, wherein a sub-top view closer to the mobile platform has a higher resolution;

[0010] Lane line recognition is performed at least based on the multiple molecular top views.

[0011] In some embodiments, performing lane line recognition based at least on the multi-molecule top view includes: identifying lane line pixels in the multi-molecule top view to obtain a lane line recognition result.

[0012] In some embodiments, identifying lane line pixels in the multiple molecular top-view images to obtain lane line recognition results includes:

[0013] The multi-molecule top view is input into a deep neural network to identify lane line pixels in the multi-molecule top view to obtain a lane line recognition result.

[0014] In some embodiments, the multi-molecule top view is input into a deep neural network to identify lane line pixels in the multi-molecule top view to obtain a lane line recognition result, including:

[0015] Inputting the multi-molecule top view into a deep neural network to extract deep features and shallow features;

[0016] Lane line pixel points are identified based on the deep features and shallow features to obtain lane line recognition results.

[0017] In some embodiments, the lane line detection method further includes: obtaining a scene semantic understanding result of the front view;

[0018] The performing lane line recognition at least based on the multi-molecule top view includes: performing lane line recognition based on the scene semantic understanding result and the multi-molecule top view.

[0019] In some embodiments, the performing lane line recognition based on the scene semantic understanding result and the multi-molecule top view includes:

[0020] Projecting the scene semantic understanding result of the front view into a top-down scene semantic understanding result under a top-down perspective;

[0021] Determine a first lane line recognition result according to the multi-molecule top view;

[0022] The first lane line recognition result is filtered according to the overhead scene semantic understanding result to obtain a second lane line recognition result.

[0023] In some embodiments, the semantic understanding result of the overhead scene includes the probability that the multiple pixel points in the front view projection as the overhead scene are non-lane line pixels; the first lane line recognition result includes the probability that the multiple pixel points in the multi-molecule overhead view are lane line pixels.

[0024] In some embodiments, filtering the first lane line recognition result according to the overhead scene semantic understanding result to obtain the second lane line recognition result includes:

[0025] The lane line pixels in the multi-molecule top view are screened and determined based on the probability that the multiple pixel points in the front view projection as the top view scene are non-lane line pixels and the probability that the multiple pixel points in the multi-molecule top view are lane line pixels.

[0026] In some embodiments, screening and determining lane line pixels in the multi-molecule top view based on the probability that the plurality of pixels in the top view projection as the top view scene are non-lane line pixels and the probability that the plurality of pixels in the multi-molecule top view are lane line pixels includes:

[0027] Determining a probability that a plurality of pixel points in the front view projection as a top-down scene are non-lane line pixels;

[0028] Pixels whose probability of being non-lane line pixels is greater than a first preset probability are screened out from the multiple pixels and determined as non-lane line pixels.

[0029] In some embodiments, filtering lane line pixels based on the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels and the probability that the plurality of pixels in the molecular top-down views are lane line pixels includes:

[0030] Pixel points having a probability of being lane line pixels greater than a second preset probability are screened out from the plurality of pixel points under the multi-molecule top view and are determined as lane line pixels.

[0031] In some embodiments, filtering lane line pixels based on the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels and the probability that the plurality of pixels in the molecular top-down views are lane line pixels includes:

[0032] Determining a probability that a plurality of pixel points in the front view projection as a top-down scene are non-lane line pixels;

[0033] Filtering pixel points whose probability of being non-lane line pixel points is greater than a first preset probability from the plurality of pixel points, and determining them as a first non-lane line pixel point set;

[0034] Filtering out pixel points whose probability of being lane line pixels is less than a second preset probability from the plurality of pixel points in the multi-molecule top view, and determining them as a second non-lane line pixel point set;

[0035] Pixel points other than the first non-lane line pixel set and the second non-lane line pixel set in the multi-molecule top view are screened and determined as lane line pixel points.

[0036] In some embodiments, the lane line detection method further includes: generating a front view lane line in the front view based on the determined lane line pixel points in the multiple molecular top views.

[0037] In some embodiments, generating a front view lane line in the front view based on the determined lane line pixel points in the plurality of molecular top views includes:

[0038] Generating a plurality of lane lines in each sub-view according to the determined lane line pixel points in the plurality of sub-views;

[0039] The plurality of lane lines are projected back into the front view to generate front view lane lines.

[0040] In a second aspect, an embodiment of the present application provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the following steps are implemented:

[0041] Get the front view of the mobile platform;

[0042] Segmenting the front view into multiple front views from near to far;

[0043] Projecting the multi-molecule front view into multi-molecule top views with different resolutions, wherein a sub-top view closer to the mobile platform has a higher resolution;

[0044] Lane line recognition is performed at least based on the multiple molecular top views.

[0045] In some embodiments, performing lane line recognition at least based on the multi-molecule top view includes: identifying lane line pixels in the multi-molecule top view to obtain a lane line recognition result.

[0046] In some embodiments, identifying lane line pixels in the multiple molecular top-view images to obtain lane line recognition results includes:

[0047] The multi-molecule top view is input into a deep neural network to identify lane line pixels in the multi-molecule top view to obtain a lane line recognition result.

[0048] In some embodiments, the multi-molecule top view is input into a deep neural network to identify lane line pixels in the multi-molecule top view to obtain a lane line recognition result, including:

[0049] Inputting the multi-molecule top view into a deep neural network to extract deep features and shallow features;

[0050] Lane line pixel points are identified based on the deep features and shallow features to obtain lane line recognition results.

[0051] In some embodiments, the electronic device further comprises: acquiring a scene semantic understanding result of the front view;

[0052] The performing lane line recognition at least based on the multi-molecule top view includes: performing lane line recognition based on the scene semantic understanding result and the multi-molecule top view.

[0053] In some embodiments, the performing lane line recognition based on the scene semantic understanding result and the multi-molecule top view includes:

[0054] Projecting the scene semantic understanding result of the front view into a top-down scene semantic understanding result under a top-down perspective;

[0055] Determine a first lane line recognition result according to the multi-molecule top view;

[0056] The first lane line recognition result is filtered according to the overhead scene semantic understanding result to obtain a second lane line recognition result.

[0057] In some embodiments, the semantic understanding result of the overhead scene includes the probability that the multiple pixel points in the front view projection as the overhead scene are non-lane line pixels; the first lane line recognition result includes the probability that the multiple pixel points in the multi-molecule overhead view are lane line pixels.

[0058] In some embodiments, filtering the first lane line recognition result according to the overhead scene semantic understanding result to obtain the second lane line recognition result includes:

[0059] The lane line pixels in the multi-molecule top view are screened and determined based on the probability that the multiple pixel points in the front view projection as the top view scene are non-lane line pixels and the probability that the multiple pixel points in the multi-molecule top view are lane line pixels.

[0060] In some embodiments, screening and determining lane line pixels in the multi-molecule top view based on the probability that the plurality of pixels in the top view projection as the top view scene are non-lane line pixels and the probability that the plurality of pixels in the multi-molecule top view are lane line pixels includes:

[0061] Determining a probability that a plurality of pixel points in the front view projection as a top-down scene are non-lane line pixels;

[0062] Pixels whose probability of being non-lane line pixels is greater than a first preset probability are screened out from the multiple pixels and determined as non-lane line pixels.

[0063] In some embodiments, filtering lane line pixels based on the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels and the probability that the plurality of pixels in the molecular top-down views are lane line pixels includes:

[0064] Pixel points having a probability of being lane line pixels greater than a second preset probability are screened out from the plurality of pixel points under the multi-molecule top view and are determined as lane line pixels.

[0065] In some embodiments, filtering lane line pixels based on the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels and the probability that the plurality of pixels in the molecular top-down views are lane line pixels includes:

[0066] Determining a probability that a plurality of pixel points in the front view projection as a top-down scene are non-lane line pixels;

[0067] Filtering pixel points whose probability of being non-lane line pixel points is greater than a first preset probability from the plurality of pixel points, and determining them as a first non-lane line pixel point set;

[0068] Filtering out pixel points whose probability of being lane line pixels is less than a second preset probability from the plurality of pixel points in the multi-molecule top view, and determining them as a second non-lane line pixel point set;

[0069] Pixel points other than the first non-lane line pixel set and the second non-lane line pixel set in the multi-molecule top view are screened and determined as lane line pixel points.

[0070] In some embodiments, the electronic device is further configured to:

[0071] Generate a front view lane line in the front view according to the determined lane line pixel points in the multiple molecular top views.

[0072] In some embodiments, generating a front view lane line in the front view based on the determined lane line pixel points in the plurality of molecular top views includes:

[0073] Generating a plurality of lane lines in each sub-view according to the determined lane line pixel points in the plurality of sub-views;

[0074] The plurality of lane lines are projected back into the front view to generate front view lane lines.

[0075] In a third aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the lane line detection method described in any embodiment of the present application is implemented.

[0076] In a fourth aspect, an embodiment of the present application further provides a computer program product, which includes a computer program stored on a storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any one of the above-mentioned lane line detection methods.

[0077] After acquiring the front view of the mobile platform, the present application divides the front view into multiple sub-front views from near to far, and then further projects the multiple front views into multiple sub-top views of different resolutions, wherein the sub-top view closer to the mobile platform has a higher resolution, thereby realizing lane line recognition with different resolutions for different field of view ranges, reducing the missed detection of lateral lane lines or lane lines with large curvature that are closer to the mobile platform. In addition, since the real-time performance required for autonomous driving limits the input resolution size, the present application does not simply use high-resolution images to improve the accuracy and reliability of lane line detection. Instead, by projecting the multiple front views into multiple sub-top views of different resolutions, different resolutions are used for lane line recognition for different field of view ranges, thereby ensuring the accuracy and reliability of lane line detection while also ensuring sufficient recognition distance to the greatest extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0079] FIG1 is a flow chart of an embodiment of a lane detection method of the present application;

[0080] FIG2 is a flow chart of another embodiment of the lane detection method of the present application;

[0081] FIG3 is a flow chart of another embodiment of the lane detection method of the present application;

[0082] FIG4 is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0083] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0084] It should be noted that, unless there is any conflict, the embodiments and features in the embodiments of this application can be combined with each other.

[0085] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0086] In this application, "module", "device", "system" and the like refer to related entities applied to a computer, such as hardware, a combination of hardware and software, software or software in execution, etc. Specifically, for example, an element can be, but is not limited to, a process running on a processor, a processor, an object, an executable element, an execution thread, a program and / or a computer. In addition, an application or script program running on a server, or a server can all be an element. One or more elements can be in an execution process and / or thread, and an element can be localized on a computer and / or distributed between two or more computers, and can be run by various computer-readable media. An element can also communicate through local and / or remote processes based on a signal having one or more data packets, for example, a signal from a data packet interacting with another element in a local system, a distributed system, and / or a signal from a network on the Internet that interacts with other systems via signals.

[0087] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include" and "comprise" include not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or device. In the absence of further limitations, the elements defined by the phrase "include..." do not exclude the presence of other identical elements in the process, method, article or device that includes the elements.

[0088] In the process of realizing this application, the inventors discovered that in the lane line detection method of the prior art, in the top view, in order to ensure sufficient recognition distance, the vertical pixel density of the top view is often relatively small, resulting in too few pixels for the lateral / large curvature lane lines. This leads to serious omissions of lateral or oblique lane lines (common in large curvature curves) and poor recognition effect. However, in fact, this type of lane line is often located in a close range and does not require a large observation distance. For this reason, the inventors proposed the lane line detection method of the present application, which adopts multi-resolution input and uses different resolutions for recognition of different field of view ranges. The lane line detection method in the embodiment of the present application can be applied to electronic devices, which can be mounted on mobile platforms, for example, cars with different levels of autonomous driving, mobile robots, and logistics robots that automatically deliver packages, etc. This application is not limited to this.

[0089] As shown in FIG1 , a lane line detection method provided in an embodiment of the present application includes:

[0090] S10. Acquire a front view of the mobile platform.

[0091] For example, a front view can be acquired using an image acquisition device that can be used to acquire images. The image acquisition device can be one, two, or more. The image acquisition device can include an infrared camera, a depth camera, a lidar, a millimeter-wave radar, or other image sensors. The image acquisition device can be mounted on a mobile platform to acquire a front view of the mobile platform. The front view is relative to the direction of movement of the mobile platform; the view in the direction of movement of the mobile platform is the front view.

[0092] For example, when the mobile platform is a car, the image acquisition device can be installed on the roof or hood of the car, etc., which is not limited in this application. When the car is moving forward, the image acquisition device acquires a view in the car's forward direction as a front view; when the car is moving backward, the image acquisition device acquires a view in the car's backward direction as a front view. Specifically, the car can be equipped with image acquisition devices specifically for acquiring views in the forward and backward directions to acquire the front view; or an image acquisition device whose direction can be automatically adjusted according to the direction of the car's movement can be installed to acquire the front view. For example, when the car is moving forward, the image acquisition device automatically adjusts to face the front of the car, and when the car is moving backward, the image acquisition device automatically adjusts to face the rear of the car.

[0093] S20, dividing the front view into multiple sub-front views from near to far.

[0094] Exemplarily, after acquiring the front view of the mobile platform, the front view is further segmented into multiple front views from near to far. The front view can be segmented at equal distance intervals or unequal distance intervals to obtain multiple front views. Here, from near to far is relative to the image acquisition device. For example, the front view of the vehicle is captured by a depth camera, and the front view is segmented according to the distance from the depth camera to obtain multiple front views. Here, the distance interval can be set differently according to the resolution of different image sensors. For image sensors with higher resolutions, clear images at longer distances can be captured, and the distance interval when segmenting the corresponding images can be larger (for example, 60m). For image sensors with lower resolutions, clear images can only be captured at shorter distances, and the distance interval when segmenting the corresponding images can be smaller (for example, 30m). The present invention is not limited to this.

[0095] S30: Projecting the multi-molecule front view into multi-molecule top views with different resolutions, wherein a sub-top view closer to the mobile platform has a higher resolution.

[0096] For example, after segmenting to obtain multiple sub-front views, each sub-front view is projected into a sub-top view with different resolutions, with the sub-top view closer to the mobile platform having a higher resolution. The projection method used can be a method in the prior art, and the present invention is not limited thereto.

[0097] S40: Perform lane line recognition based at least on the multi-molecule top view.

[0098] After acquiring a front view of a mobile platform, an embodiment of the present application divides the front view into multiple sub-front views from near to far, and then further projects the multiple front views into multiple sub-top views of different resolutions, wherein the sub-top views that are closer to the mobile platform have higher resolutions, thereby realizing lane line recognition with different resolutions for different field of view ranges, and solving the problem of missed detection of lateral lane lines or lane lines with large curvature that are closer to the mobile platform.

[0099] In addition, since the real-time requirements of autonomous driving limit the input resolution, this application does not simply use high-resolution images to improve the accuracy and reliability of lane line detection. Instead, it projects the multi-molecule front view into multi-molecule top-view views of different resolutions, and uses different resolutions for lane line recognition in different field of view ranges. This ensures the accuracy and reliability of lane line detection while also ensuring sufficient recognition distance to the greatest extent.

[0100] In some embodiments, lane line recognition is performed at least based on the multi-molecule top view, including: identifying lane line pixels in the multi-molecule top view to obtain a lane line recognition result.

[0101] For example, semantic segmentation is performed on the multi-element top-view image to perform binary classification of pixels in the multi-element top-view image to determine whether the pixels in the multi-element top-view image are lane line pixels or background pixels. In some embodiments, binary classification of pixels in the sub-top-view image can be performed to obtain the probability of each pixel being a lane line pixel.

[0102] In some embodiments, the lane line recognition result includes the probability that multiple pixels in the multi-molecule top view are lane line pixel points. The lane line detection method of the present application also includes: when the probability is greater than the preset probability, determining that the corresponding pixel point is a lane line pixel point.

[0103] For example, semantic segmentation is used to perform binary classification on the pixels in the multi-molecule top view, and the probability of each pixel in the multi-molecule top view being a lane line pixel is obtained. When the probability of a certain pixel being a lane line pixel is greater than a preset probability threshold (e.g., 0.8), the pixel is determined to be a lane line pixel. The larger the preset probability threshold, the more accurate and reliable the lane line recognition result obtained. The preset probability threshold can be set differently based on different considerations of lane line recognition accuracy and reliability. When higher accuracy is required, a larger preset probability threshold (e.g., 0.8) can be set. When too high accuracy is not required, a relatively small preset probability threshold (e.g., 0.75) can be set.

[0104] In some embodiments, a deep neural network is used to identify lane line pixels. Exemplarily, identifying lane line pixels in the multi-molecule top view to obtain a lane line recognition result includes: inputting the multi-molecule top view into a deep neural network to identify lane line pixels in the multi-molecule top view to obtain a lane line recognition result.

[0105] Among them, the deep neural network can adopt the LaneNet network structure. LaneNet is a multi-task model that combines semantic segmentation and pixel vectorization. Semantic segmentation is used to separate lane lines from the background, and pixel vectorization is used to cluster pixels belonging to the same lane line. The network used for image segmentation includes two parts, encoding and decoding. LaneNet is no exception. The difference is that LaneNet contains two branches: a segmentation branch and an embedding branch. Among them, the segmentation branch is responsible for semantic segmentation of the input image (such as binary classification of pixels to determine whether the pixels belong to the lane line or the background), and the embedding branch performs embedded representation of the pixels, which can separate the lane lines obtained after segmentation into different lane instances, and the trained vectors are used for clustering. Finally, the results of the two branches are clustered using the MeanShift algorithm to obtain the segmentation results. In some embodiments, the top views of multiple molecules have the same size.

[0106] In this embodiment, the top views of multiple molecules are configured to be of the same size, so that they can be integrated into one input batch when using a deep neural network for semantic segmentation and input into the deep neural network at one time, which helps to reduce processing time and improve lane line recognition efficiency.

[0107] FIG2 is a flow chart of another embodiment of the lane recognition method of the present application. In this embodiment, the multi-molecule top view image is input into a deep neural network to identify lane line pixels in the multi-molecule top view image to obtain a lane line recognition result, including:

[0108] S41. Input the top view of multiple molecules into the deep neural network to extract deep features and shallow features.

[0109] Deep features are used to measure semantic similarity, while shallow features are used to measure fine-grained similarity. For example, deep features are often complex and difficult to describe, such as yellow lane markings, ladybug wings, and colorful flowers. Shallow features are often generalized and easy to express, such as texture, color, edges, and corners.

[0110] S42. Perform lane line pixel recognition based on the deep features and shallow features to obtain a lane line recognition result.

[0111] In this embodiment, the deep neural network gradually extracts deep features by performing convolution and downsampling operations on the input sub-top view. When the network downsampling multiple is too large, although a larger receptive field (Receptive Field refers to the area of ​​the input image that a point on the feature map can see, that is, the point on the feature map is calculated by the receptive field size area in the input image) can be obtained, it will also affect the detection effect of object edges. In order to optimize the detection of small longitudinal objects such as lane lines with large curvature, this application adopts a multi-scale feature fusion neural network structure to fuse deep features with shallow features, extracting global information while better retaining the recognition effect of object edges.

[0112] In some embodiments, the lane line detection method of the present application also includes: obtaining the scene semantic understanding results of the front view; the lane line recognition is performed at least based on the multi-molecule top view, including: lane line recognition is performed based on the scene semantic understanding results and the multi-molecule top view.

[0113] Exemplarily, image semantic segmentation is used to perform semantic analysis on the front view to obtain the corresponding scene semantic understanding results. Image semantic segmentation is a basic technology for image understanding, and plays an important role in autonomous driving systems (specifically street scene recognition and understanding), drone applications (landing point judgment) and wearable device applications. An image is composed of many pixels, and "image semantic segmentation", as the name implies, is to group pixels according to the different semantic meanings expressed in the image. For example, the pixels in the front view that express the semantic meaning of lane lines are divided into a group. The grouping can also include the probability that each pixel belongs to the corresponding group. In this embodiment, the lane line pixel points contained in the front view are obtained by obtaining the scene semantic understanding results of the front view, and then the lane lines are recognized in combination with multiple top views.

[0114] The embodiment of the present application performs lane line detection based on both the front view and the top view. The lane line recognition in the front view can make up for the problem of misidentification of similar lane lines in the top view, while the lane line recognition in the multi-resolution top view improves the problem of missed detection of lateral lanes or lane lines with large curvature.

[0115] FIG3 is a flow chart of another embodiment of the lane line recognition method of the present application. In this embodiment, lane line recognition is performed based on the scene semantic understanding result and the multi-molecule top view, including:

[0116] S41′, projecting the scene semantic understanding result of the front view into a top-down scene semantic understanding result under a top-down perspective.

[0117] For example, scene semantic understanding is generally performed in the front view, which has a better recognition effect for nearby or large objects. Objects such as guardrails and light beams in the front view are easily projected into lane line-like objects in the top view, and existing top view lane line detection schemes are prone to misidentification of them, affecting the stability of vehicle control. Therefore, the embodiment of the present application first projects the scene semantic understanding result of the front view into the top view to obtain the top view scene semantic understanding result, and then based on the top view scene semantic understanding result, it can make a more accurate prediction of objects such as guardrails and light beams, and then filter out the predicted non-lane line pixels to reduce the false detection of lane lines.

[0118] Exemplarily, the scene semantic understanding result of the front view includes the recognition result of each pixel point in the front view (including lane line pixel points and non-lane line pixel points). Projecting the scene semantic understanding result of the front view into the overhead scene semantic understanding result under the overhead perspective is to project the recognition result of each pixel point in the front view into the recognition result of each pixel point under the overhead perspective. The specific projection method can be implemented using existing technology, and this application is not limited to this.

[0119] Pixels in the front view can be projected onto pixels in the top view, and the corresponding scene semantic understanding results of the front view can also be projected onto the scene semantic understanding of the top view (i.e., top-view scene semantic understanding). For example, in a top view, a guardrail on a highway extends along the highway, very similar to a lane line. Therefore, direct recognition in the top view will usually result in the guardrail being mistaken for a lane line. However, in the front view, the guardrail has a height, so the scene semantic understanding of the front view can determine that the corresponding pixel is a non-lane line pixel. Correspondingly, projecting it onto the scene semantic understanding of the top view can determine that the corresponding pixel in the top view is a non-lane line.

[0120] S42′: Determine a first lane line recognition result based on the multi-molecule top view.

[0121] Exemplarily, lane line pixel points in a multi-molecule overhead view are identified to obtain a lane line recognition result (i.e., a first lane line recognition result). Exemplarily, a deep neural network can be used to perform lane line pixel recognition. For example, identifying lane line pixel points in a multi-molecule overhead view to obtain a lane line recognition result (i.e., a first lane line recognition result) includes: inputting the multi-molecule overhead view into a deep neural network to identify lane line pixel points in the multi-molecule overhead view to obtain a lane line recognition result (i.e., a first lane line recognition result). Exemplarily, the overhead scene semantic understanding result includes the probability that the front view is projected as a plurality of pixel points in the overhead scene as non-lane line pixel points; the first lane line recognition result includes the probability that the plurality of pixel points in the multi-molecule overhead view are lane line pixel points.

[0122] S43′: Filter the first lane line recognition result according to the overhead scene semantic understanding result to obtain a second lane line recognition result.

[0123] Exemplarily, the first lane line recognition result includes the probability value of each pixel in the multi-molecule overhead view being a lane line pixel point; the second lane line recognition result includes the lane line pixel point determined after filtering the first lane line recognition result based on the overhead scene semantic understanding result.

[0124] For example, the first lane line recognition result includes pixels a1, a2, and a3 with probabilities of 0.7, 0.8, and 0.85, respectively, indicating they are lane line pixels. However, based on the top-down scene semantic understanding results, it is determined that pixels a1 and a2 are not lane line pixels. Therefore, the second lane line recognition result, determined after screening, includes pixel a3, which is identified as a lane line pixel.

[0125] This application uses multi-resolution input to identify lane lines at different resolutions across different fields of view, addressing issues with missed lane markings or lane markings with large curvature. Furthermore, the front view scene semantic understanding results are projected onto the top view, filtering out non-lane line pixels predicted by the scene semantic understanding to reduce false lane line detections. This application utilizes front view scene semantic understanding to fuse and filter lane lines from the top view, combining the advantages of different perspectives while addressing issues with missed and false lane line recognition in existing technologies.

[0126] In some embodiments, the first lane line recognition result is filtered according to the semantic understanding result of the overhead scene to obtain the second lane line recognition result, including: screening and determining the lane line pixel points under the multi-molecule overhead view based on the probability that the multiple pixel points under the front view projection as the overhead scene are non-lane line pixel points and the probability that the multiple pixel points under the multi-molecule overhead view are lane line pixel points.

[0127] In this embodiment, the probability that the same pixel point is a non-lane line pixel point and the probability that the same pixel point is a lane line pixel point can be considered simultaneously to ultimately determine the option that is most likely to be a lane line pixel point.

[0128] The embodiment of the present application simultaneously considers the probability that a pixel point in the semantic understanding result of the overhead scene is a non-lane line pixel point and the probability that a pixel point in the multi-molecule overhead view is a lane line pixel point to screen and determine the lane line pixel points under the multi-molecule overhead view to obtain a second lane line recognition result. The combination of these two factors improves the accuracy and reliability of the second lane line recognition result, thereby improving the accuracy of lane line detection.

[0129] In some embodiments, based on the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels and the probability that the plurality of pixels in the multi-molecule top-down view are lane line pixels, screening and determining the lane line pixels in the multi-molecule top-down view includes:

[0130] Determining a probability that a plurality of pixel points in the front view projection as a top-down scene are non-lane line pixels;

[0131] Pixels whose probability of being non-lane line pixels is greater than a first preset probability are screened out from the multiple pixels and determined as non-lane line pixels.

[0132] In this embodiment, if the probability that a pixel in the top-down view of the current view is a non-lane line pixel is greater than a first preset probability, it indicates that the pixel cannot be a lane line pixel. Therefore, the pixel in the corresponding top-down sub-view is determined to be a non-lane line pixel and filtered out to avoid false lane line detection. The first preset probability can be configured as 0.8, but is not limited to this. An appropriate first preset probability can be set based on actual needs.

[0133] For example, consider filtering out misidentified lane line category B on scene parsing category A. Assume that the probability that pixel b belongs to scene semantic understanding category A is P A (b) Only when P A (b)>p a (For example, 0.8), filtering is performed, where p a is the preset probability value (0<=p a ).

[0134] In some embodiments, lane line pixel points are screened based on the probability that the multiple pixel points in the front view projection as a top-view scene are non-lane line pixel points and the probability that the multiple pixel points in the multi-molecule top-view are lane line pixel points, including: screening out pixel points whose probability of being lane line pixel points is greater than a second preset probability from the multiple pixel points in the multi-molecule top-view.

[0135] In this embodiment, when the probability that a pixel in the sub-top view is a lane line pixel is greater than a second preset probability, the pixel is determined to be a lane line pixel. The second preset probability can be configured as 0.9, but is not limited to this. An appropriate second preset probability can be set according to actual needs.

[0136] Consider filtering the misidentified lane line category B on the scene semantic understanding category A. Assume that the probability that pixel b belongs to lane line category B is P B (b) Only when P B (b) <p b (For example, 0.9), filtering is performed, where p b is the preset probability value (p b <=1).

[0137] In some embodiments, based on the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels and the probability that the plurality of pixels in the multi-molecule top-down view are lane line pixels, screening and determining the lane line pixels in the multi-molecule top-down view includes:

[0138] Determining a probability that a plurality of pixel points in the front view projection as a top-down scene are non-lane line pixels;

[0139] Filtering pixel points whose probability of being non-lane line pixel points is greater than a first preset probability from the plurality of pixel points, and determining them as a first non-lane line pixel point set;

[0140] Filtering out pixel points whose probability of being lane line pixels is less than a second preset probability from the plurality of pixel points in the multi-molecule top view, and determining them as a second non-lane line pixel point set;

[0141] Pixel points other than the first non-lane line pixel set and the second non-lane line pixel set in the multi-molecule top view are screened and determined as lane line pixel points.

[0142] In this embodiment, the scene semantic understanding result of the front view is first projected onto the top view, and the lane lines on the non-drivable area predicted by the scene semantic understanding are filtered out to reduce the false detection of lane lines. Different from the simple fusion prediction results, this application simultaneously utilizes the observation uncertainty of scene semantic understanding and lane line detection, and combines different scene semantic understanding and lane line categories to obtain a more fine-grained fusion effect. More specifically, consider filtering the lane line category B that is misidentified on the scene semantic understanding category A. Assuming that the probability of pixel b belonging to scene semantic understanding category A is P A (b) The probability of belonging to lane line category B is P B (b) Only when P A (b)>p a (e.g., 0.8) and P B (b) <p b (For example, 0.9), filtering is performed, where p a and p b are the preset probability values ​​(0<=p a ,p b <=1). This post-fusion method of scene semantic information can effectively avoid erroneous filtering of real lane lines.

[0143] The embodiments of the present application optimize the recognition of small objects such as lane lines with large curvature or lateral lane lines in a top view by using multi-resolution top-view input and multi-scale network design. At the same time, with the help of the output and uncertainty of the semantic understanding of the front view scene, false detections of line-like objects in the top view are effectively filtered out, thereby optimizing the lane line recognition effect in the top view from the aspects of missed detection and false detection.

[0144] In some embodiments, the lane line detection method further includes: generating a front view lane line in the front view based on the determined lane line pixel points in the multiple molecular top views.

[0145] Exemplarily, the front view lane lines are generated in the front view according to the determined lane line pixel points in the multi-component top view, including: generating multiple sub-lane lines in each sub-top view according to the determined lane line pixel points in the multi-component top view; and projecting the multiple sub-lane lines back to the front view to generate front view lane lines.

[0146] In this embodiment, a multi-resolution output fusion method is used to perform network inference and scene semantic understanding fusion on inputs of different resolutions to obtain lane line output results of different resolutions. These results are then re-projected and integrated into a complete front view lane line output based on the projection relationship from the front view to the top view.

[0147] In some embodiments, the entire lane detection method from input, network processing, post-fusion to output can be divided into the following steps:

[0148] Step 1: Multi-resolution input

[0149] Using the front view (FV) as input, traditional top-view detection schemes set a uniform projection resolution for the entire image: a longitudinal recognition distance of h meters, a longitudinal pixel density of p pixels / meter, a lateral recognition distance of w meters, and a lateral resolution of q pixels / meter. Thus, the top-view height H = h*p, and the top-view width W = w*q. To ensure sufficient recognition distance, the longitudinal pixel density p is often relatively small, resulting in too few pixels for lateral / high-curvature lane markings, making them difficult to recognize. These lane markings are often located at close range. Therefore, this method projects the front view (FV) into top-views of different resolutions based on distance, such as higher resolution for close proximity and lower resolution for distance. At the same time, these top-views are ensured to have the same input size so that they can be combined into a single input batch and fed into the neural network for inference at once, reducing processing time.

[0150] Step 2: Neural network structure for multi-scale feature fusion

[0151] Deep neural networks gradually extract high-level semantic information by convolution and downsampling their inputs. While downsampling by too much can result in a larger receptive field, it also compromises object edge detection. To optimize the detection of small longitudinal objects, such as lane markings with large curvature, this paper fuses high-level features with shallow-level features, extracting global information while better preserving object edge recognition.

[0152] Step 3: Post-fusion of scene semantic information

[0153] Scene semantic understanding (scene parsing) is generally performed in the front view, and has a better recognition effect for nearby or large objects. Objects such as guardrails and light beams in the front view are easily projected into lane line-like objects in the top view. Existing top view lane line detection schemes are prone to misidentification of them, affecting the stability of vehicle control. Therefore, this method first projects the scene semantic understanding results of the front view into the top view, and filters out the lane lines on the non-drivable area predicted by the scene semantic understanding to reduce the false detection of lane lines. Unlike simple fusion prediction results, this method simultaneously utilizes the observation uncertainty of scene semantic understanding and lane line detection, and combines different scene semantic understanding and lane line categories to obtain a more fine-grained fusion effect. More specifically, consider filtering out the misidentified lane line category B on the scene semantic understanding category A, assuming that the probability of pixel b belonging to scene semantic understanding category A is P A (b), the probability of belonging to lane line category B is P B (b) Only when P A (b)>p a( 0.8 ) And P B (b) <p b (0.9), filtering is performed, where p a and p b are the preset probability values ​​(0<=p a ,p b <=1). This method can effectively avoid incorrectly filtering out the real lane lines.

[0154] Step 4: Multi-resolution output fusion

[0155] By performing network inference and scene semantic understanding fusion on inputs of different resolutions, lane line output results of different resolutions can be obtained. According to the projection relationship in step 1, these results are re-projected and integrated into the complete front view lane line output.

[0156] In some embodiments, the present application further provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the following steps are implemented:

[0157] S10. Acquire a front view of the mobile platform.

[0158] For example, a front view can be acquired using an image acquisition device that can be used to acquire images. The image acquisition device can be one, two, or more. The image acquisition device can include an infrared camera, a depth camera, a lidar, a millimeter-wave radar, or other image sensors. The image acquisition device can be mounted on a mobile platform to acquire a front view of the mobile platform. The front view is relative to the direction of movement of the mobile platform; the view in the direction of movement of the mobile platform is the front view.

[0159] For example, when the mobile platform is a car, the image acquisition device can be installed on the roof or hood of the car, etc., which is not limited in this application. When the car is moving forward, the image acquisition device acquires a view in the car's forward direction as a front view; when the car is moving backward, the image acquisition device acquires a view in the car's backward direction as a front view. Specifically, the car can be equipped with image acquisition devices specifically for acquiring views in the forward and backward directions to acquire the front view; or an image acquisition device whose direction can be automatically adjusted according to the direction of the car's movement can be installed to acquire the front view. For example, when the car is moving forward, the image acquisition device automatically adjusts to face the front of the car, and when the car is moving backward, the image acquisition device automatically adjusts to face the rear of the car.

[0160] S20, dividing the front view into multiple sub-front views from near to far.

[0161] Exemplarily, after acquiring the front view of the mobile platform, the front view is further segmented into multiple front views from near to far. The front view can be segmented at equal distance intervals or unequal distance intervals to obtain multiple front views. Here, from near to far is relative to the image acquisition device. For example, the front view of the vehicle is captured by a depth camera, and the front view is segmented according to the distance from the depth camera to obtain multiple front views. Here, the distance interval can be set differently according to the resolution of different image sensors. For image sensors with higher resolutions, clear images at longer distances can be captured, and the distance interval when segmenting the corresponding images can be larger (for example, 60m). For image sensors with lower resolutions, clear images can only be captured at shorter distances, and the distance interval when segmenting the corresponding images can be smaller (for example, 30m). The present invention is not limited to this.

[0162] S30: Projecting the multi-molecule front view into multi-molecule top views with different resolutions, wherein a sub-top view closer to the mobile platform has a higher resolution.

[0163] For example, after segmenting to obtain multiple sub-front views, each sub-front view is projected into a sub-top view with different resolutions, with the sub-top view closer to the mobile platform having a higher resolution. The projection method used can be a method in the prior art, and the present invention is not limited thereto.

[0164] S40: Perform lane line recognition based at least on the multi-molecule top view.

[0165] After acquiring a front view of a mobile platform, an embodiment of the present application divides the front view into multiple sub-front views from near to far, and then further projects the multiple front views into multiple sub-top views of different resolutions, wherein the sub-top views that are closer to the mobile platform have higher resolutions, thereby realizing lane line recognition with different resolutions for different field of view ranges, and solving the problem of missed detection of lateral lane lines or lane lines with large curvature that are closer to the mobile platform.

[0166] In addition, since the real-time requirements of autonomous driving limit the input resolution, this application does not simply use high-resolution images to improve the accuracy and reliability of lane line detection. Instead, it projects the multi-molecule front view into multi-molecule top-view views of different resolutions, and uses different resolutions for lane line recognition in different field of view ranges. This ensures the accuracy and reliability of lane line detection while also ensuring sufficient recognition distance to the greatest extent.

[0167] In some embodiments, performing lane line recognition based at least on the multi-molecule top view includes: identifying lane line pixels in the multi-molecule top view to obtain a lane line recognition result.

[0168] For example, semantic segmentation is performed on the multi-element top-view image to perform binary classification of pixels in the multi-element top-view image to determine whether the pixels in the multi-element top-view image are lane line pixels or background pixels. In some embodiments, binary classification of pixels in the sub-top-view image can be performed to obtain the probability of each pixel being a lane line pixel.

[0169] In some embodiments, the lane line recognition result includes the probability that multiple pixels in the multi-molecule top view are lane line pixel points, and the processor is further configured to: determine that the corresponding pixel is a lane line pixel point when the probability is greater than a preset probability.

[0170] For example, semantic segmentation is used to perform binary classification on the pixels in the multi-molecule top view to obtain the probability that each pixel in the multi-molecule top view is a lane line pixel, and when the probability that a certain pixel is a lane line pixel is greater than a preset probability threshold (for example, 0.8), the pixel is determined to be a lane line pixel. The larger the preset probability threshold, the more accurate and reliable the lane line recognition result obtained, and vice versa. The preset probability threshold can be set differently based on different considerations of lane line recognition accuracy and reliability. When higher accuracy is required, a larger preset probability threshold (for example, 0.8) can be set. When too high accuracy is not required, a relatively small preset probability threshold (for example, 0.75) can be set.

[0171] In some embodiments, a deep neural network is used to identify lane line pixels. Exemplarily, identifying lane line pixels in the multi-molecule top view to obtain a lane line recognition result includes: inputting the multi-molecule top view into a deep neural network to identify lane line pixels in the multi-molecule top view to obtain a lane line recognition result.

[0172] Among them, the deep neural network can adopt the LaneNet network structure. LaneNet is a multi-task model that combines semantic segmentation and pixel vectorization. Semantic segmentation is used to separate lane lines from the background, and pixel vectorization is used to cluster pixels belonging to the same lane line. The network used for image segmentation includes two parts, encoding and decoding. LaneNet is no exception. The difference is that LaneNet contains two branches, a segmentation branch and an embedding branch. Among them, the segmentation branch is responsible for semantic segmentation of the input image (binary classification of pixels to determine whether the pixels belong to the lane line or the background), and the embedding branch performs embedded representation of the pixels, which can separate the lane lines obtained after segmentation into different lane instances, and the trained vectors are used for clustering. Finally, the results of the two branches are clustered using the MeanShift algorithm to obtain the segmentation results. In some embodiments, the top views of multiple molecules have the same size.

[0173] In this embodiment, the top views of multiple molecules are configured to be of the same size, so that they can be integrated into one input batch when using a deep neural network for semantic segmentation and input into the deep neural network at one time, which helps to reduce processing time and improve lane line recognition efficiency.

[0174] In some embodiments, the multi-molecule top view is input into a deep neural network to identify lane line pixels in the multi-molecule top view to obtain a lane line recognition result, including:

[0175] The multi-molecule top view is input into a deep neural network to extract deep features and shallow features; lane line pixel points are identified based on the deep features and shallow features to obtain lane line recognition results.

[0176] Deep features are used to measure semantic similarity, while shallow features are used to measure fine-grained similarity. For example, deep features are often complex and difficult to describe, such as yellow lane markings, ladybug wings, and colorful flowers. Shallow features are often generalized and easy to express, such as texture, color, edges, and corners.

[0177] In this embodiment, the deep neural network gradually extracts deep features by performing convolution and downsampling operations on the input sub-top view. When the network downsampling factor is too large, although a larger receptive field can be obtained, it will also affect the detection effect of object edges. To optimize the detection of small longitudinal objects such as lane lines with large curvature, this application adopts a neural network structure with multi-scale feature fusion, fusing deep features with shallow features, extracting global information while better preserving the recognition effect of object edges.

[0178] In some embodiments, the processor is further configured to: obtain a scene semantic understanding result of the front view; and perform lane line recognition at least based on the multi-molecule top view, including: perform lane line recognition based on the scene semantic understanding result and the multi-molecule top view.

[0179] Exemplarily, image semantic segmentation is used to perform semantic analysis on the front view to obtain the corresponding scene semantic understanding results. Image semantic segmentation is a basic technology for image understanding, and plays an important role in autonomous driving systems (specifically street scene recognition and understanding), drone applications (landing point judgment) and wearable device applications. An image is composed of many pixels, and "image semantic segmentation", as the name implies, is to group pixels according to the different semantic meanings expressed in the image. For example, the pixels in the front view that express the semantic meaning of lane lines are divided into a group. The grouping can also include the probability that each pixel belongs to the corresponding group. In this embodiment, the lane line pixel points contained in the front view are obtained by obtaining the scene semantic understanding results of the front view, and then the lane lines are recognized in combination with multiple top views.

[0180] The embodiment of the present application performs lane line detection based on both the front view and the top view. The lane line recognition in the front view can make up for the problem of misidentification of similar lane lines in the top view, while the lane line recognition in the multi-resolution top view improves the problem of missed detection of lateral lanes or lane lines with large curvature.

[0181] In some embodiments, performing lane line recognition based on the scene semantic understanding result and the multi-molecule top view includes:

[0182] Project the scene semantic understanding result of the front view into the top-down scene semantic understanding result under the top-down perspective; illustratively, scene semantic understanding is generally performed under the front view, and has a better recognition effect for nearby or large objects. Objects such as guardrails and light beams in the front view are easily formed into lane line-like objects when projected to the top view. Existing top-down view lane line detection schemes are prone to misidentification of them, affecting the stability of vehicle control. Therefore, this method first projects the scene semantic understanding result of the front view to the top view. The top-down scene semantic understanding result under the top-down perspective can make more accurate predictions of objects such as guardrails and light beams, and then filter out the lane lines on the predicted non-drivable areas, reducing false detections of lane lines;

[0183] Determine a first lane line recognition result based on the multi-molecule overhead view; illustratively, identify lane line pixels in the multi-molecule overhead view to obtain a lane line recognition result (i.e., a first lane line recognition result). Exemplarily, a deep neural network can be used to perform lane line pixel recognition. For example, identifying lane line pixels in a multi-molecule overhead view to obtain a lane line recognition result (i.e., a first lane line recognition result) includes: inputting the multi-molecule overhead view into a deep neural network to identify lane line pixels in the multi-molecule overhead view to obtain a lane line recognition result (i.e., a first lane line recognition result). Exemplarily, the overhead scene semantic understanding result includes the probability that the plurality of pixels in the front view projected as the overhead scene are non-lane line pixels; the first lane line recognition result includes the probability that the plurality of pixels in the multi-molecule overhead view are lane line pixels;

[0184] The first lane line recognition result is filtered based on the overhead scene semantic understanding result to obtain a second lane line recognition result. Exemplarily, the first lane line recognition result includes a probability value for each pixel in the multi-molecule overhead view that is a lane line pixel. The second lane line recognition result includes lane line pixels determined after filtering the first lane line recognition result based on the overhead scene semantic understanding result. For example, the first lane line recognition result includes probabilities of 0.7, 0.8, and 0.85 for pixels a1, a2, and a3 being lane line pixels, respectively. However, based on the overhead scene semantic understanding result, it is determined that a1 and a2 are not lane line pixels. Therefore, the second lane line recognition result determined after filtering includes a3, which is determined to be a lane line pixel.

[0185] This application uses multi-resolution input to identify lane markings at different resolutions across different fields of view, addressing issues with missed lane markings. Furthermore, the front-view scene semantic understanding results are projected onto the top-down view, filtering out lane markings in non-drivable areas predicted by the scene semantic understanding to reduce false lane marking detections. This application utilizes front-view scene semantic understanding to fuse and filter lane markings from the top-down view, combining the advantages of different perspectives while addressing issues with missed and false lane markings in existing technologies.

[0186] In some embodiments, filtering the first lane line recognition result according to the overhead scene semantic understanding result to obtain a second lane line recognition result includes:

[0187] The lane line pixels in the multi-molecule top view are screened and determined based on the probability that the multiple pixel points in the front view projection as the top view scene are non-lane line pixels and the probability that the multiple pixel points in the multi-molecule top view are lane line pixels.

[0188] In this embodiment, the probability that the same pixel point is a non-lane line pixel point and the probability that the same pixel point is a lane line pixel point can be considered simultaneously to ultimately determine the option that is most likely to be a lane line pixel point.

[0189] The embodiment of the present application simultaneously considers the probability that a pixel point in the semantic understanding result of the overhead scene is a non-lane line pixel point and the probability that a pixel point in the multi-molecule overhead view is a lane line pixel point to screen lane line pixels to obtain a second lane line recognition result. The combination of these two factors improves the accuracy and reliability of the second lane line recognition result, thereby improving the accuracy of lane line detection.

[0190] In some embodiments, based on the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels and the probability that the plurality of pixels in the multi-molecule top-down view are lane line pixels, screening and determining the lane line pixels in the multi-molecule top-down view includes:

[0191] Determining a probability that a plurality of pixel points in the front view projection as a top-down scene are non-lane line pixels;

[0192] Pixels whose probability of being non-lane line pixels is greater than a first preset probability are screened out from the multiple pixels and determined as non-lane line pixels.

[0193] In this embodiment, if the probability that a pixel in the top-down projection of the current view is a non-lane line pixel is greater than a first preset probability, it indicates that the pixel cannot be a lane line pixel. Therefore, the pixel is determined to be a non-lane line pixel and filtered out to avoid false lane line detection. The first preset probability can be configured as 0.8, but is not limited to this. An appropriate first preset probability can be set based on actual needs.

[0194] For example, consider filtering out the misidentified lane line category B on the scene semantic understanding category A. Assume that the probability that pixel b belongs to the scene semantic understanding category A is P A (b) Only when P A (b)>p a (For example, 0.8), filtering is performed, where p a is the preset probability value (0<=p a ).

[0195] In some embodiments, based on the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels and the probability that the plurality of pixels in the multi-molecule top-down view are lane line pixels, screening and determining the lane line pixels in the multi-molecule top-down view includes:

[0196] Pixel points having a probability of being lane line pixels greater than a second preset probability are screened out from the plurality of pixel points under the multi-molecule top view and are determined as lane line pixels.

[0197] In this embodiment, when the probability that a pixel in the sub-top view is a lane line pixel is greater than a second preset probability, the pixel is determined to be a lane line pixel. The second preset probability can be configured as 0.9, but is not limited to this. An appropriate second preset probability can be set according to actual needs.

[0198] Consider filtering the misidentified lane line category B on the scene semantic understanding category A. Assume that the probability that pixel b belongs to lane line category B is P B (b) Only when P B (b) <p b (For example, 0.9), filtering is performed, where p b is the preset probability value (p b <=1).

[0199] In some embodiments, based on the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels and the probability that the plurality of pixels in the multi-molecule top-down view are lane line pixels, screening and determining the lane line pixels in the multi-molecule top-down view includes:

[0200] Determining a probability that a plurality of pixel points in the front view projection as a top-down scene are non-lane line pixels;

[0201] Filtering pixel points whose probability of being non-lane line pixel points is greater than a first preset probability from the plurality of pixel points, and determining them as a first non-lane line pixel point set;

[0202] Filtering out pixel points whose probability of being lane line pixels is less than a second preset probability from the plurality of pixel points in the multi-molecule top view, and determining them as a second non-lane line pixel point set;

[0203] Pixel points other than the first non-lane line pixel set and the second non-lane line pixel set in the multi-molecule top view are screened and determined as lane line pixel points.

[0204] In this embodiment, the scene semantic understanding result of the front view is first projected onto the top view, and the lane lines on the non-drivable area predicted by the scene semantic understanding are filtered out to reduce the false detection of lane lines. Different from the simple fusion prediction results, this application simultaneously utilizes the observation uncertainty of scene semantic understanding and lane line detection, and combines different scene semantic understanding and lane line categories to obtain a more fine-grained fusion effect. More specifically, consider filtering the lane line category B that is misidentified on the scene semantic understanding category A. Assuming that the probability of pixel b belonging to scene semantic understanding category A is P A (b) The probability of belonging to lane line category B is P B (b) Only when P A (b)>p a (e.g., 0.8) and P B (b) <p b (For example, 0.9), filtering is performed, where p a and p b are the preset probability values ​​(0<=p a ,p b <=1). This post-fusion method of scene semantic information can effectively avoid erroneous filtering of real lane lines.

[0205] This application uses multi-resolution top-view input and multi-scale network design to optimize the recognition of small objects such as lane lines with large curvature or lateral lane lines in top view. At the same time, with the output and uncertainty of the semantic understanding of the front view scene, it effectively filters out false detections of line-like objects in the top view, thereby optimizing the lane line recognition effect in the top view from both missed detection and false detection.

[0206] In some embodiments, the processor is further configured to generate a front view lane line in the front view based on the determined lane line pixel points in the multiple molecular top views.

[0207] Exemplarily, the front view lane lines are generated in the front view according to the determined lane line pixel points in the multi-component top view, including: generating multiple sub-lane lines in each sub-top view according to the determined lane line pixel points in the multi-component top view; and projecting the multiple sub-lane lines back to the front view to generate front view lane lines.

[0208] In this embodiment, a multi-resolution output fusion method is used to perform network inference and scene semantic understanding fusion on inputs of different resolutions to obtain lane line output results of different resolutions. These results are then re-projected and integrated into a complete front view lane line output based on the projection relationship from the front view to the top view.

[0209] In some embodiments, generating a front view lane line in the front view based on the determined lane line pixel points in the plurality of molecular top views includes:

[0210] Generating a plurality of lane lines in each sub-view according to the determined lane line pixel points in the plurality of sub-views;

[0211] The plurality of lane lines are projected back into the front view to generate front view lane lines.

[0212] An embodiment of the present application further provides a mobile platform, characterized in that it is equipped with the electronic device described in any of the aforementioned embodiments.

[0213] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the lane line detection method described in any embodiment of the present application.

[0214] An embodiment of the present application also provides a computer program product, which includes a computer program stored on a storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any one of the above-mentioned lane line detection methods.

[0215] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of combined actions, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0216] FIG4 is a schematic diagram of the hardware structure of an electronic device for performing a lane detection method according to another embodiment of the present application. As shown in FIG4 , the device includes:

[0217] One or more processors 410 and a memory 420 , with one processor 410 being used as an example in FIG4 .

[0218] The device for executing the lane detection method may further include: an input device 430 and an output device 440 .

[0219] The processor 410 , the memory 420 , the input device 430 , and the output device 440 may be connected via a bus or other means. FIG. 4 takes the bus connection as an example.

[0220] Memory 420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the lane detection method in the embodiments of this application. Processor 410 executes the non-volatile software programs, instructions, and modules stored in memory 420 to execute various server functional applications and data processing, thereby implementing the lane detection method in the aforementioned method embodiment.

[0221] Memory 420 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function; the data storage area may store data generated based on the use of the lane detection device. Furthermore, memory 420 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 420 may optionally include memory remotely located relative to processor 410. Such remote memory may be connected to the lane detection device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0222] The input device 430 can receive input digital or character information and generate signals related to user settings and function control of the lane detection device. The output device 440 can include a display device such as a display screen.

[0223] The one or more modules are stored in the memory 420 and, when executed by the one or more processors 410 , perform the lane line detection method in any of the above method embodiments.

[0224] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.

[0225] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0226] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.

[0227] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A lane line detection method, include: Get the front view of the mobile platform; Segmenting the front view into multiple sub-front views from near to far; Projecting the multi-molecule front view into multi-molecule top view with different resolutions, wherein the top view closer to the mobile platform has a higher resolution; Lane line recognition is performed at least based on the multiple molecular top views.

2. The method according to claim 1, It is characterized in that The performing lane line recognition at least according to the multiple molecular top views comprises: Identify lane line pixels in the multiple molecular top views to obtain lane line recognition results.

3. The method according to claim 2, It is characterized in that The identifying lane line pixel points in the multiple molecular top views to obtain lane line recognition results includes: The multi-molecule top view is input into a deep neural network to identify lane line pixels in the multi-molecule top view to obtain a lane line recognition result.

4. The method according to claim 3, It is characterized in that Inputting the multi-molecule top view into a deep neural network to identify lane line pixels in the multi-molecule top view to obtain a lane line recognition result, including: Inputting the multi-molecule top view into a deep neural network to extract deep features and shallow features; Lane line pixel points are identified based on the deep features and shallow features to obtain lane line recognition results.

5. The method according to claim 1, It is characterized in that The method further comprises: obtaining a scene semantic understanding result of the front view; The performing lane line recognition at least according to the multiple molecular overhead views includes: performing lane line recognition according to the scene semantic understanding result and the multiple molecular overhead views.

6. The method according to claim 5, It is characterized in that The performing lane line recognition according to the scene semantic understanding result and the multiple molecular top view includes: Projecting the scene semantic understanding result of the front view into a top-down scene semantic understanding result under a top-down perspective; Determine a first lane line recognition result according to the multiple molecular top views; The first lane line recognition result is filtered according to the overhead scene semantic understanding result to obtain a second lane line recognition result.

7. The method according to claim 6, It is characterized in that The top-view scene semantic understanding result includes the probability that the plurality of pixels in the front view projection as the top-view scene are non-lane line pixels; the first lane line recognition result includes the probability that the plurality of pixels in the multi-molecule top view are lane line pixels; The filtering the first lane line recognition result according to the overhead scene semantic understanding result to obtain the second lane line recognition result includes: The lane line pixels under the multi-molecule top view are screened and determined based on the probability that the multiple pixel points under the front view projection as the top view scene are non-lane line pixels and the probability that the multiple pixel points under the multi-molecule top view are lane line pixels.

8. The method according to claim 7, It is characterized in that The screening and determining of lane line pixels in the multi-molecule top view according to the probability that the plurality of pixels in the top view projection as the top view scene are non-lane line pixels and the probability that the plurality of pixels in the multi-molecule top view are lane line pixels includes: Determine the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels; Pixel points whose probability of being non-lane line pixel points is greater than a first preset probability are screened out from the multiple pixel points and determined as non-lane line pixel points.

9. The method according to claim 7, It is characterized in that The screening and determining of lane line pixels in the multi-molecule top view according to the probability that the plurality of pixels in the top view projection as the top view scene are non-lane line pixels and the probability that the plurality of pixels in the multi-molecule top view are lane line pixels includes: Pixel points whose probability of being lane line pixel points is greater than a second preset probability are screened out from the multiple pixel points under the multiple molecular top-view images, and are determined as lane line pixel points.

10. The method according to claim 7, It is characterized in that The screening and determining of lane line pixels in the multi-molecule top view according to the probability that the plurality of pixels in the top view projection as the top view scene are non-lane line pixels and the probability that the plurality of pixels in the multi-molecule top view are lane line pixels includes: Determine the probability that the plurality of pixels in the front view projection as a top-down scene are non-lane line pixels; Filtering pixel points whose probability of being non-lane line pixel points is greater than a first preset probability among the plurality of pixel points, and determining them as a first non-lane line pixel point set; Filter out pixel points whose probability of being lane line pixel points is less than a second preset probability from the plurality of pixel points in the plurality of molecular top views, and determine them as a second non-lane line pixel point set; Pixel points other than the first non-lane line pixel set and the second non-lane line pixel point set in the multiple molecular top-view images are screened and determined as lane line pixel points.

11. The method according to any one of claims 1 to 10, It is characterized in that The method further comprises: A front view lane line is generated in the front view according to the determined lane line pixel points in the multiple molecular top views.

12. The method according to claim 11, It is characterized in that Generating a front view lane line in the front view according to the determined lane line pixel points in the plurality of molecular top views, comprising: Generate a plurality of lane lines in each sub-view according to the determined lane line pixel points in the plurality of sub-views; The plurality of lane lines are projected back into the front view to generate front view lane lines.

13. An electronic device, include: At least one processor, and a memory in communication with the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the following steps are implemented: Get the front view of the mobile platform; Segmenting the front view into multiple sub-front views from near to far; Projecting the multi-molecule front view into multi-molecule top view with different resolutions, wherein the top view closer to the mobile platform has a higher resolution; Lane line recognition is performed at least based on the multiple molecular top views.

14. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the steps of the method described in any one of claims 1 to 12 are implemented.

15. A mobile platform, It is characterized in that The electronic device according to claim 13 is installed.

Citation Information

Patent Citations

  • Lane line detection method and device and electronic device

    CN110796003A

  • Lane line detection method and lane line detection system

    CN113343742A

  • Lane line identification method and device, electronic equipment and storage medium

    CN114926801A