A road drivable area recognition method and device based on binocular vision

By combining binocular cameras and spatial attention mechanisms with semantic segmentation networks, the problem of identifying drivable areas on unpaved roads using binocular vision was solved, achieving accurate perception and decision support on various road surfaces.

CN115497061BActive Publication Date: 2026-04-17RADAR NEW ENERGY AUTOMOBILE (ZHEJIANG) CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RADAR NEW ENERGY AUTOMOBILE (ZHEJIANG) CO LTD
Filing Date
2022-09-01
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing binocular vision technology has difficulty accurately identifying drivable areas on unpaved roads, making it difficult for autonomous driving decision-making modules to effectively determine drivable areas.

Method used

Road images are acquired using binocular cameras, and features are extracted and matched using a spatial attention mechanism to generate a disparity map. Two coordinate transformations are then performed, and a semantic segmentation network is used to identify the drivable area of ​​the road.

Benefits of technology

It enables accurate identification of drivable areas on both paved and unpaved roads, improving the accuracy and applicability of autonomous driving decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497061B_ABST
    Figure CN115497061B_ABST
Patent Text Reader

Abstract

This application proposes a method for identifying drivable road areas based on binocular vision. The method includes: acquiring left and right images of the road using a binocular camera; extracting and matching features from the left and right images based on a spatial attention mechanism to obtain an attention matrix and image features, and generating a disparity map based on the attention matrix and image features; performing a first coordinate transformation on each pixel in the disparity map to obtain a corresponding depth map, and performing a second coordinate transformation on each pixel in the depth map to obtain a corresponding height map; and identifying the drivable area of ​​the road based on the height map. The technical solution of this application is applicable not only to road surface information perception on paved roads but also to road surface information perception on unpaved roads, and the perception result is the drivable area ahead of the road, which facilitates the autonomous driving decision-making module to make decisions directly based on the drivable area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this application relate to the field of autonomous driving, and more particularly to a method and apparatus for identifying drivable road areas based on binocular vision. Background Technology

[0002] In recent years, with the rapid development of artificial intelligence technology, autonomous driving technology has received increasing attention worldwide. For autonomous vehicles, environmental perception is fundamental to decision-making and control. Currently, autonomous driving perception technology mainly focuses on the detection, segmentation, and tracking of targets such as vehicles, pedestrians, traffic lights, and lane lines in highway scenarios. Meanwhile, due to its lower cost and wider perception range compared to lidar, binocular vision technology is gradually being applied to perception systems in the field of autonomous driving.

[0003] In related technologies, binocular vision technology is used to filter information beyond obstacles such as vehicles and pedestrians, and to obtain the distance between obstacles and vehicles, thereby achieving perception of the road surface. However, this solution is often only applicable to scenarios where vehicles are traveling on paved roads such as cement roads and highways, and is not suitable for unpaved roads with very complex traffic environments, such as unfinished roads or rural roads. Therefore, how to accurately and efficiently perceive road surface information and identify the drivable area ahead is an important problem in the field of autonomous driving. Summary of the Invention

[0004] This application provides a method and apparatus for identifying drivable road areas based on binocular vision, in order to address the shortcomings of related technologies.

[0005] According to a first aspect of one or more embodiments of this application, a method for identifying drivable road areas based on binocular vision is provided, the method comprising:

[0006] Images of the road are captured from the left and right sides using a binocular camera;

[0007] Based on the spatial attention mechanism, feature extraction and matching are performed on the left and right images to obtain the attention matrix and image features, and a disparity map is generated based on the attention matrix and the image features.

[0008] The first coordinate transformation of each pixel in the disparity map is performed to obtain the corresponding depth map, and the second coordinate transformation of each pixel in the depth map is performed to obtain the corresponding height map.

[0009] The drivable area of ​​the road is identified based on the elevation map.

[0010] According to a second aspect of one or more embodiments of this application, a method for training a road drivable area recognition model based on binocular vision is provided, the method comprising:

[0011] Acquire the left and right images of the sample road, the disparity map of the target sample, and the drivable area of ​​the target;

[0012] Based on the spatial attention mechanism, feature extraction and matching are performed on the left and right images of the sample to obtain the attention matrix and sample image features, and a sample disparity map is generated based on the attention matrix and the sample image features.

[0013] The first coordinate transformation is performed on each pixel in the sample disparity map to obtain the corresponding sample depth map, and the second coordinate transformation is performed on each pixel in the sample depth map to obtain the corresponding sample height map.

[0014] The drivable area of ​​the sample road is identified based on the sample height map;

[0015] The road drivable area recognition model is iteratively trained based on the drivable area of ​​the sample road and the target drivable area, as well as the sample disparity map and the target sample disparity map.

[0016] According to a third aspect of one or more embodiments of this application, a road drivable area recognition device based on binocular vision is provided, the device comprising:

[0017] The acquisition unit is used to acquire left and right images of the road using a binocular camera;

[0018] The matching unit is used to extract and match features from the left and right images based on a spatial attention mechanism to obtain an attention matrix and image features, and to generate a disparity map based on the attention matrix and the image features.

[0019] The coordinate transformation unit is used to perform a first coordinate transformation on each pixel in the disparity map to obtain a corresponding depth map, and to perform a second coordinate transformation on each pixel in the depth map to obtain a corresponding height map.

[0020] The identification unit is used to identify the drivable area of ​​the road based on the elevation map.

[0021] According to a fourth aspect of one or more embodiments of this application, an electronic device is provided, comprising:

[0022] processor;

[0023] Memory used to store processor-executable instructions;

[0024] The processor executes the executable instructions to implement the method described in the embodiments of the first / second aspect above.

[0025] According to a fifth aspect of one or more embodiments of this application, a computer-readable storage medium is provided having computer instructions stored thereon that, when executed by a processor, implement the steps of the method as described in the embodiments of the first / second aspects above.

[0026] According to a sixth aspect of one or more embodiments of this application, a vehicle is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method as described in the embodiments of the first / second aspect above.

[0027] As can be seen from the above technical solutions, in one or more embodiments of this application, feature extraction and matching of the left and right images acquired by the binocular camera through a spatial attention mechanism can better capture the correlation between the left and right images, thereby obtaining a more accurate disparity map. The disparity map is then subjected to two coordinate transformations to obtain a height map, which in turn yields the drivable area of ​​the road. The technical solution provided by this application is applicable not only to road surface information perception on paved roads but also to road surface information perception on unpaved roads. Furthermore, the perception results include the road's elevation changes and the drivable area ahead, facilitating direct decision-making by the autonomous driving decision module based on the drivable area.

[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0030] Figure 1 This is a flowchart of an exemplary embodiment of a method for identifying drivable road areas based on binocular vision.

[0031] Figure 2 This is an exemplary embodiment of the left and right images before image correction using pre-calibrated camera information.

[0032] Figure 3 This is an exemplary embodiment of a left and right image after image correction using pre-calibrated camera information.

[0033] Figure 4 This is a schematic diagram of a network structure for feature extraction and matching of left and right images based on a spatial attention mechanism, provided in an exemplary embodiment.

[0034] Figure 5 This is a flowchart of a training method for a road drivable area identification model provided in an exemplary embodiment.

[0035] Figure 6 This is a schematic diagram illustrating the structure of an electronic device in an exemplary embodiment.

[0036] Figure 7 This is a block diagram illustrating a road drivable area recognition device based on binocular vision, as shown in an exemplary embodiment.

[0037] Figure 8 This is a block diagram illustrating a training device for a road drivable area recognition model based on binocular vision, as shown in an exemplary embodiment. Detailed Implementation

[0038] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0039] It should be noted that the steps of the corresponding methods in other embodiments are not necessarily performed in the order shown and described in this application. In some other embodiments, the methods may include more or fewer steps than those described in this application. Furthermore, a single step described in this application may be broken down into multiple steps in other embodiments; and multiple steps described in this application may be combined into a single step in other embodiments.

[0040] In existing technologies, the perception of road information using binocular vision technology only involves detecting and reconstructing the depth of significant obstacles on the road. This makes the relevant technical solutions unsuitable for unpaved roads with very complex traffic environments, such as off-road roads. Furthermore, the perception result is merely a depth map representing the spatial distance between obstacles and the vehicle. The autonomous driving decision-making module cannot directly determine the drivable area of ​​the road based on the depth map, thus failing to make the next driving decision.

[0041] This application provides a training method for a model of road drivable area recognition based on binocular vision, as well as a method for road drivable area recognition based on binocular vision. It is applicable not only to road information perception on paved roads, but also to road information perception on unpaved roads. The perception results include the road's elevation changes and the drivable area ahead of the road, so that the autonomous driving decision module can make decisions directly based on the drivable area.

[0042] Figure 1 This is a flowchart illustrating a method for identifying drivable road areas based on binocular vision, as provided in an exemplary embodiment. Figure 1 As shown, the method may include the following steps:

[0043] S101: Acquire left and right images of the road using a binocular camera.

[0044] Using the binocular cameras that have been installed and calibrated on the vehicle, images of the road environment are collected, and left and right images of the road are obtained.

[0045] In one embodiment, the binocular camera needs to be calibrated first to obtain the intrinsic and extrinsic parameters and relative positional relationship of the left and right cameras, in order to prepare for the subsequent correction of the left and right images and the generation of the depth map. There are many specific camera calibration methods, such as Zhang Zhengyou calibration method for offline calibration and self-calibration algorithm for online calibration. Those skilled in the art can determine the method according to relevant technologies and specific needs, and this application does not limit it. During the calibration process, the role of the camera intrinsic parameters is to determine the projection relationship of the camera from three-dimensional space to two-dimensional image, and the role of the camera extrinsic parameters is to determine the relative positional relationship between the camera coordinate system and the world coordinate system. The conversion principle from the world coordinate system to the image coordinate system is shown in formula (1):

[0046]

[0047] Where α is the proportionality constant, (X W Y W Z W (x, y) is the coordinate representation of a point in the road scene in the world coordinate system, and (x, y) is the coordinate representation of a point in the aforementioned road scene transformed into the image coordinate system. It is the camera's intrinsic parameter matrix, f x f y The focal length is represented in the intrinsic parameters, and (u0, v0) represents the coordinates of the principal point of the image. The principal point is the intersection of the optical axis emitted from the camera's optical center and the imaging plane. [R] 3×3 T 3×1 [] is the extrinsic parameter matrix of the camera 3D coordinate system relative to the world coordinate system. The overall transformation process from the world coordinate system to the image coordinate system is as follows: multiply the world coordinate system with the R and T matrices to transform the world coordinate system to the camera 3D coordinate system; then multiply the camera 3D coordinate system with the camera's intrinsic parameter matrix to transform the camera 3D coordinate system to the image coordinate system.

[0048] Since camera lenses typically exhibit radial and tangential distortion, distortion coefficients need to be obtained during intrinsic parameter calibration. As shown in formula (2):

[0049]

[0050] Where (x′, y′) is the coordinate representation of a point in the aforementioned road scene transformed into the pixel coordinate system, and k1, k2, k3, p1, and p2 are all distortion coefficients. This represents the distance between the point corresponding to a point in the aforementioned road scene and the principal point on the distortion correction map. The transformation from the image coordinate system to the pixel coordinate system is completed through formula (2).

[0051] In one embodiment, after camera calibration is completed, the pre-calibrated camera information can be used to perform image correction on the left and right images to obtain row-aligned corrected left and right images. The pre-calibrated camera information includes the intrinsic and extrinsic parameters and relative positional relationship of the left and right cameras.

[0052] Figure 2 This is a sample embodiment providing left and right images before correction. Wherein, O l It is the optical center of the left camera, O r It is the optical center of the right camera, P l This is the left image of the road before correction, P r This is the right image of the road before correction. Clearly, the left and right images before correction are neither parallel nor coplanar, which is detrimental to the subsequent image matching process. Therefore, the left and right images can be corrected based on the pre-calibrated relative positional relationship between the left and right cameras. Specifically, the positional change matrix of the right camera relative to the left camera is shown in formula (3):

[0053]

[0054] Where R represents the rotation matrix, T represents the translation matrix, and R r T represents the rotation matrix of the right camera. r R represents the translation matrix of the right camera; similarly, R l T represents the rotation matrix of the left camera. l Let represent the translation matrix of the left camera. Rotating the right camera according to formula (3) will make the right image coplanar and parallel to the left image, as shown below. Figure 3 As shown.

[0055] S102: Based on the spatial attention mechanism, feature extraction and matching are performed on the left and right images to obtain the attention matrix and image features, and a disparity map is generated based on the attention matrix and the image features.

[0056] Feature extraction is performed on the corrected left and right images using a Convolutional Neural Network (CNN) to obtain the low-level features of the left and right images. In one embodiment, to reduce the number of parameters in the CNN, a shared-weight CNN can be used to extract features from the left and right images. The specific CNN can be determined by those skilled in the art based on relevant technologies and specific needs; this application does not impose any limitations on it.

[0057] like Figure 4 As shown, after obtaining the low-level features of the left and right images, these features are input into the spatial attention layer and the feature calculation layer, respectively, to obtain the attention matrix and high-level image features. Low-level image features generally refer to features such as contours, edges, colors, and textures, which contain relatively little semantic information. High-level image features, on the other hand, represent what the image expresses that is closest to human understanding. For example, when extracting features from a face, the low-level features extracted are the face contour, nose, glasses, etc., while the high-level features are simply the face itself. Therefore, high-level image features contain richer semantic information.

[0058] In one embodiment, the low-level features of the left and right images are input into a feature calculation layer. This feature calculation layer may include pooling layers and activation function layers. After multiple convolutional calculations, high-level features of the image are extracted. The activation function layer may employ the ReLU activation function to reduce computational load and improve efficiency. After obtaining the high-level features, the high-level features are concatenated with the low-level features along the channel dimension to obtain the image features. Concatenating the features along the channel dimension increases the number of features (channels) in the image itself, enabling subsequent image matching processes to perform feature matching from local details to overall semantics based on the low-level and high-level features, thereby improving the accuracy of image feature matching.

[0059] In one embodiment, the low-level features of the left and right images are input into a spatial attention layer to generate an attention matrix. The spatial attention mechanism aims to improve the feature representation of key regions. Essentially, it transforms the spatial information in the original image into another space, preserving key information, generating weights for each location, and outputting a weighted sum, thereby enhancing the specific target region of interest while weakening irrelevant background regions. The spatial attention layer can be divided into pooling layers, convolutional layers, and computational layers. In neural networks, four commonly used pooling operations are average pooling, max pooling, stochastic pooling, and global average pooling. Pooling operations can reduce the size of the feature map, i.e., reduce the computational load. In this embodiment, the low-level features of the input left and right images are a three-dimensional array of bounding box × height × channels. The low-level features first pass through a pooling layer, using both average pooling and max pooling operations to reduce the bounding box and height of the low-level features, resulting in a dual-channel feature map. The dual-channel feature map is then input into a convolutional layer for convolution operations. The convolutional layer can reduce the number of channels, outputting a single-channel feature map. Then, attention weights are calculated for each pixel in the single-channel feature map at the computational layer, outputting an attention matrix. The sigmoid function or the hyperbolic tangent function (Tanh function) can be used to limit the range of attention weights to 0 to 1, thereby increasing the weight of strong features and decreasing the weight of weak features. This better captures the correlation between the left and right images, improving the accuracy of subsequent feature matching. The attention matrix output from the spatial attention layer is multiplied by the image features obtained by concatenating low-level and high-level features to obtain a feature matching map incorporating spatial attention. This completes the feature extraction and matching of the left and right images of the road.

[0060] The feature matching map is input into a pre-trained disparity map generation network to generate a disparity map. In this embodiment, the disparity map generation network can be a convolutional neural network or a graph neural network (GNN), and the convolutional neural network or the graph neural network can be a single layer or multiple layers. Those skilled in the art can determine this according to specific needs. The training process of the disparity map generation network is detailed in [link to documentation]. Figure 5 In conjunction with the above embodiments, formula (3) yields the change matrix of the right camera relative to the left camera. Therefore, the disparity map generated by the disparity map generation network in this step is an image with the left image as the reference image, the same size as the reference image, and disparity values ​​as its element values. The disparity value represents the difference in the horizontal coordinate of the same point on the corrected left and right images. Of course, in this step, a disparity map with the right image as the reference image, the same size as the reference image, and disparity values ​​as its element values ​​can also be generated.

[0061] S103: Perform a first coordinate transformation on each pixel in the disparity map to obtain a corresponding depth map, and perform a second coordinate transformation on each pixel in the depth map to obtain a corresponding height map.

[0062] The general steps of a binocular vision algorithm are as follows: after acquiring the disparity map, the actual depth is calculated based on the disparity, camera intrinsic parameters, etc., i.e., a depth map is generated based on the disparity map. The value of each pixel in the depth map represents the actual distance from each point in the road scene to the binocular camera. In a calibrated and corrected binocular system, the relationship between the disparity map and the camera's 3D coordinates is shown in formula (4):

[0063]

[0064] Where (u, v) represents the coordinates of a pixel in the disparity map, d represents the disparity value of the same pixel, and matrix Q is the reprojection matrix, which can map two-dimensional points in the image plane back to the three-dimensional coordinate system in the physical world. In matrix Q, T x Let (u0, v0) represent the x-component of the translation vector in formula (3), (u0, v0) represent the principal coordinates of the left image, u′0 represent the principal coordinates of the right image, and f represent the focal length. (X, Y, Z) represent the three-dimensional coordinates of a pixel on the disparity map mapped to the physical world, and W represents the scale factor, which can be canceled out during the calculation.

[0065] Based on the above embodiments, and using formula (4) and the disparity map, the 3D coordinates of each pixel in the disparity map in the left camera coordinate system can be obtained through coordinate transformation, and then the depth map can be obtained. The value of each pixel in the depth map represents the actual distance from each point in the road scene to the binocular camera.

[0066] According to the extrinsic parameter matrix [R] of the left camera 3D coordinate system relative to the world coordinate system in S101 3×3 T 3×1 A second coordinate transformation is performed on each pixel in the depth map to obtain the coordinates of each pixel in the world coordinate system, thus generating a height map from a top-down perspective. The value of each pixel in the height map represents the height of each point in the road scene from the ground.

[0067] S104: Identify the drivable area of ​​the road based on the elevation map.

[0068] The heightmap is input into a pre-trained semantic segmentation network to identify and segment drivable and indestructible regions within the heightmap. For details on the training process of the semantic segmentation network, please refer to [link to training documentation]. Figure 5Semantic segmentation of images assigns a semantic category to each pixel in the input image to obtain a pixelated dense classification. For example, if a pixel is marked as green, it means that the location of this pixel is a tree. However, if there are two green pixels, the semantic segmentation network can only determine that both pixels are located at the location of trees, but cannot determine whether they are the same tree or two separate trees. Based on this characteristic, semantic segmentation is very suitable for identifying drivable and non-drivable areas of roads. A typical semantic segmentation network can be considered as an encoder-decoder network, specifically including fully convolutional networks (FCNs), SegNet networks, U-Net networks, etc. Those skilled in the art can determine the appropriate network according to their actual needs, and this application does not impose any restrictions on this.

[0069] In one embodiment, a pre-trained U-Net network is used for semantic segmentation of the height map. First, the U-Net network's encoder encodes and extracts features from the input height map, progressively reducing the bounding box and height of each feature while increasing the number of channels for each feature. Then, the U-Net network simply concatenates the encoded feature map to the decoder's upsampled feature map. The decoder, similar to feature upsampling, uses deconvolution layers to progressively increase the bounding box and height of features while reducing the number of channels, ultimately generating a semantic segmentation image containing only one channel. This semantic segmentation image accurately segments the drivable and non-drivable areas. By perceiving road information using the above method, not only the drivable area ahead of the road can be perceived, but also the road's undulations, such as stones and depressions on unpaved roads. Therefore, the technical solution of this application is applicable not only to road information perception on paved roads but also to road information perception on unpaved roads. Furthermore, by directly outputting the drivable area ahead of the road, it helps the vehicle's autonomous driving decision-making module to make autonomous driving decisions based on the drivable area, improving the user's autonomous driving experience.

[0070] The above embodiments describe the method for identifying drivable road areas based on binocular vision according to this application. In specific implementation, a drivable road area identification model can be established and trained, thereby accurately identifying the aforementioned drivable road areas based on this model.

[0071] Figure 5 This is a flowchart illustrating a training method for a road drivable area identification model, provided as an exemplary embodiment. Figure 5 As shown, the method may include the following steps:

[0072] S501: Obtain the left and right images of the sample road, the disparity map of the target sample, and the target drivable area.

[0073] In this embodiment, the sample road may include paved surfaces such as cement roads and asphalt roads, or unpaved surfaces such as country roads and rough roads. A sample road is selected, the target drivable area of ​​that sample road is obtained, and sample left and right images of the sample road are acquired using a binocular camera that has been installed and calibrated on the vehicle. The specific calibration method for the camera, and the specific process of image correction of the sample left and right images using pre-calibrated camera information, can be found in embodiment S101 above. There are many ways to generate the target sample disparity map, and those skilled in the art can determine it according to relevant technologies and specific needs; this application does not require detailed limitation. For example, a target sample depth map can be obtained from the sample left and right images using a lidar, and then the corresponding target sample disparity map can be obtained based on the target sample depth map.

[0074] S502: Based on the spatial attention mechanism, feature extraction and matching are performed on the left and right images of the sample to obtain the attention matrix and sample image features, and a sample disparity map is generated based on the attention matrix and the sample image features.

[0075] In one embodiment, a shared-weight convolutional neural network is first used to extract features from the corrected left and right sample images to obtain the low-level features of the left and right sample images. Then, the low-level features of the left and right sample images are input into a spatial attention layer and a feature calculation layer, respectively, outputting an attention matrix and high-level features of the sample images. The high-level features of the sample images are then concatenated with the low-level features of the left and right sample images along the channel dimension to obtain the sample image features. The specific generation method can be referred to in embodiment S102 above, and will not be repeated here.

[0076] In one embodiment, the sample attention matrix is ​​multiplied by the sample image features to generate feature matching maps for the left and right images of the sample. These feature matching maps are then input into a disparity map generation network to output a sample disparity map. The specific generation method can be found in embodiment S102 above, and will not be repeated here.

[0077] S503: Perform a first coordinate transformation on each pixel in the sample disparity map to obtain a corresponding sample depth map, and perform a second coordinate transformation on each pixel in the sample depth map to obtain a corresponding sample height map.

[0078] In one embodiment, a first coordinate transformation is performed on each pixel in the sample disparity map based on a binocular vision algorithm to obtain the 3D coordinates of each pixel in the sample disparity map in the camera coordinate system, thereby obtaining the sample depth map. Then, a second coordinate transformation is performed on each pixel in the sample depth map to obtain the 3D coordinates of each pixel in the sample depth map in the world coordinate system, thereby obtaining the sample height map from a top-down perspective.

[0079] S504: Identify the drivable area of ​​the sample road based on the sample height map.

[0080] In one embodiment, the sample height map is input into a semantic segmentation network to identify and segment drivable and non-drivable areas in the sample height map, and outputs the drivable area of ​​the sample road. The specific identification and segmentation method can be referred to in embodiment S104 above, and will not be repeated here.

[0081] S505: The road drivable area recognition model is iteratively trained based on the drivable area of ​​the sample road and the target drivable area, as well as the sample disparity map and the target sample disparity map.

[0082] The model is trained and supervised based on two aspects: the drivable area of ​​the sample road and the target drivable area output by the model, and the disparity map of the sample road and the target sample road generated by the model. This further ensures the accuracy of the disparity map generated based on the spatial attention mechanism. After multiple iterations of training, the training of the drivable area recognition model is completed when it meets the predefined training objectives or reaches the predefined number of iterations. This model can then be used to achieve, for example... Figure 1 The diagram shows a scheme for identifying drivable road zones.

[0083] Corresponding to the above method embodiments, this application also provides an embodiment of an apparatus.

[0084] Figure 6 This is a schematic diagram illustrating the structure of an electronic device according to an exemplary embodiment of this application. (Reference) Figure 6 At the hardware level, the electronic device includes a processor 602, an internal bus 604, a network interface 606, memory 608, and non-volatile memory 610, and may also include other hardware required for business operations. The processor 602 reads the corresponding computer program from the non-volatile memory 610 into the memory 608 and then runs it. Of course, in addition to software implementation, this application does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.

[0085] Figure 7 This is a block diagram illustrating a road drivable area recognition device based on binocular vision according to an exemplary embodiment of this application. (Refer to...) Figure 7 The device includes a data acquisition unit 702, a matching unit 704, a coordinate transformation unit 706, and a recognition unit 708, wherein:

[0086] Acquisition unit 702 is configured to acquire left and right images of the road using a binocular camera.

[0087] Optionally, the device further includes:

[0088] The image correction unit 710 is configured to perform image correction on the left and right images using pre-calibrated camera information to obtain row-aligned corrected left and right images; the pre-calibrated camera information includes the intrinsic and extrinsic parameters of the binocular camera and the relative positional relationship between the two cameras.

[0089] The matching unit 704 is configured to perform feature extraction and matching on the left and right images based on a spatial attention mechanism to obtain an attention matrix and image features, and generate a disparity map based on the attention matrix and the image features.

[0090] Optionally, the matching unit 704 is specifically used for: extracting features from the corrected left and right images based on a spatial attention mechanism to obtain the low-level features of the corrected left and right images; inputting the low-level features into a spatial attention layer and a feature calculation layer respectively; the feature calculation layer is used to perform convolution calculation on the low-level features to obtain high-level image features; the spatial attention layer includes a pooling layer, a convolution layer, and a calculation layer, wherein:

[0091] The pooling layer is used to pool the underlying features to obtain a dual-channel feature map;

[0092] The convolutional layer is used to convolve the dual-channel feature map to obtain a single-channel feature map;

[0093] The computational layer is used to calculate the attention weights corresponding to each pixel in the single-channel feature map and output the attention matrix.

[0094] The high-level features of the image are concatenated with the low-level features to obtain the image features. The attention matrix is ​​then multiplied by the image features to obtain the feature matching maps corresponding to the left and right images. The feature matching maps are then input into a pre-trained disparity map generation network to generate disparity maps corresponding to the left and right images.

[0095] The coordinate transformation unit 706 is configured to perform a first coordinate transformation on each pixel in the disparity map to obtain a corresponding depth map, and to perform a second coordinate transformation on each pixel in the depth map to obtain a corresponding height map.

[0096] The identification unit 708 is configured to identify the drivable area of ​​the road based on the height map.

[0097] Optionally, the recognition unit 708 is specifically used to: input the height map into a pre-trained semantic segmentation network, the semantic segmentation network including an encoder and a decoder; based on the encoder and decoder, identify and segment the drivable and non-drivable areas in the height map, and output the drivable area.

[0098] Figure 8 This is a block diagram illustrating a training device for a road drivable area recognition model based on binocular vision, according to an exemplary embodiment of this application. (Refer to...) Figure 8 The device includes a sample acquisition unit 802, a sample matching unit 804, a sample coordinate transformation unit 806, a sample recognition unit 808, and an iteration unit 810, wherein:

[0099] The sample acquisition unit 802 is configured to acquire left and right images of the sample road, a target sample disparity map, and a target drivable area.

[0100] Optionally, the device further includes:

[0101] The sample correction unit 812 is configured to perform image correction on the left and right images of the sample using pre-calibrated camera information to obtain row-aligned corrected left and right images of the sample; the pre-calibrated camera information includes the intrinsic and extrinsic parameters of the binocular camera and the relative positional relationship between the two cameras.

[0102] The sample matching unit 804 is configured to extract and match features from the left and right images of the sample based on a spatial attention mechanism to obtain an attention matrix and sample image features, and generate a sample disparity map based on the attention matrix and the sample image features.

[0103] Optionally, the sample matching unit 804 is specifically used to: extract and match features from the left and right images of the row-aligned corrected samples based on a spatial attention mechanism.

[0104] The sample coordinate transformation unit 806 is configured to perform a first coordinate transformation on each pixel in the sample disparity map to obtain a corresponding sample depth map, and to perform a second coordinate transformation on each pixel in the sample depth map to obtain a corresponding sample height map.

[0105] The sample identification unit 808 is configured to identify the drivable area of ​​the sample road based on the sample height map.

[0106] The iteration unit 810 is configured to iteratively train the road drivable area recognition model based on the drivable area of ​​the sample road and the target drivable area, as well as the sample disparity map and the target sample disparity map.

[0107] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0108] The apparatus or unit described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. A typical implementation device is a computer, which can be a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0109] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0110] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0111] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0112] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for identifying drivable road areas based on binocular vision, characterized in that, The method includes: Images of the road are captured from the left and right sides using a binocular camera; Based on the spatial attention mechanism, feature extraction and matching are performed on the left and right images to obtain image features and an attention matrix used to characterize the correlation between the left and right images, and a disparity map is generated based on the attention matrix and the image features. The first coordinate transformation of each pixel in the disparity map is performed to obtain the corresponding depth map, and the second coordinate transformation of each pixel in the depth map is performed to obtain the corresponding height map. The value of each pixel in the height map represents the height of each point in the road from the ground. The drivable area of ​​the road is identified based on the elevation map.

2. The method according to claim 1, characterized in that, The method further includes: using pre-calibrated camera information to perform image correction on the left and right images to obtain row-aligned corrected left and right images; the pre-calibrated camera information includes the intrinsic and extrinsic parameters of the binocular camera and the relative positional relationship of the two cameras; The feature extraction and matching of the left and right images based on the spatial attention mechanism includes: performing feature extraction and matching of the row-aligned corrected left and right images based on the spatial attention mechanism.

3. The method according to claim 1, characterized in that, The spatial attention mechanism is used to extract and match features from the left and right images to obtain an attention matrix and image features, including: Feature extraction is performed on the left and right images respectively to obtain the low-level features of the left and right images; The low-level features of the left and right images are respectively input into the spatial attention layer and the feature calculation layer; the feature calculation layer is used to perform convolution calculations on the low-level features to obtain high-level image features; the spatial attention layer includes a pooling layer, a convolutional layer, and a calculation layer, wherein: The pooling layer is used to pool the bottom-level features to obtain a dual-channel feature map; The convolutional layer is used to convolve the dual-channel feature map to obtain a single-channel feature map; The computational layer is used to calculate the attention weights corresponding to each pixel in the single-channel feature map and output the attention matrix. The high-level features of the image are concatenated with the low-level features to obtain the image features.

4. The method according to claim 3, characterized in that, The step of generating a disparity map based on the attention matrix and the image features includes: The attention matrix is ​​multiplied by the image features to obtain the feature matching maps corresponding to the left and right images; The feature matching map is input into a pre-trained disparity map generation network to generate disparity maps corresponding to the left and right images.

5. The method according to claim 1, characterized in that, The step of identifying the drivable area of ​​the road based on the elevation map includes: The height map is input into a pre-trained semantic segmentation network, which includes an encoder and a decoder; Based on the encoder and decoder, the drivable and non-drivable areas in the height map are identified and segmented, and the drivable area is output.

6. A training method for a road drivable area recognition model based on binocular vision, characterized in that, The method includes: Acquire the left and right images of the sample road, the disparity map of the target sample, and the drivable area of ​​the target; Based on the spatial attention mechanism, feature extraction and matching are performed on the left and right images of the sample to obtain sample image features and an attention matrix used to characterize the correlation between the left and right images of the sample. A sample disparity map is then generated based on the attention matrix and the sample image features. The first coordinate transformation is performed on each pixel in the sample disparity map to obtain the corresponding sample depth map, and the second coordinate transformation is performed on each pixel in the sample depth map to obtain the corresponding sample height map. The value of each pixel in the sample height map represents the sample height of each point in the sample road from the ground. The drivable area of ​​the sample road is identified based on the sample height map; The road drivable area recognition model is iteratively trained based on the drivable area of ​​the sample road and the target drivable area, as well as the sample disparity map and the target sample disparity map.

7. A road drivable area identification device based on binocular vision, characterized in that, The device includes: The acquisition unit is used to acquire left and right images of the road using a binocular camera; The matching unit is used to extract and match features of the left and right images based on a spatial attention mechanism to obtain image features and an attention matrix for characterizing the correlation between the left and right images, and to generate a disparity map based on the attention matrix and the image features. The coordinate transformation unit is used to perform a first coordinate transformation on each pixel in the disparity map to obtain a corresponding depth map, and to perform a second coordinate transformation on each pixel in the depth map to obtain a corresponding height map. The value of each pixel in the height map represents the height of each point in the road from the ground. The identification unit is used to identify the drivable area of ​​the road based on the elevation map.

8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor implements the method as described in any one of claims 1-6 by executing the executable instructions.

9. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, this instruction implements the steps of the method as described in any one of claims 1-6.

10. A vehicle comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Passable road detection method based on binocular vision

    CN106446785A

  • Monocular vision depth estimation system and method, computer equipment and computer readable storage medium

    CN114820745A

  • Vehicle travelable region detection method and detection device

    WO2021159397A1